Independent intelligence for software built with AIResearch edition / 19 Sep 2026
SHIP/PROOFAI CODE FIELD GUIDE

Reading path / Evidence & economics

Measure the work that matters

A leaderboard is not your delivery pipeline. Measure accepted changes, review time and failures under a stated task distribution.

A suggested sequence

  1. 01

    What METR's 2025 developer study can tell us ↗

    A measured slowdown in one setting is a warning about assumptions, not a universal speed estimate.

  2. 02

    Read a SWE-bench score as a test result ↗

    A leaderboard number needs its task set, harness, and failure audit beside it.

  3. 03

    Evaluate an agent on your repository ↗

    A useful trial measures accepted work, failure recognition, and reviewer burden on tasks your team actually owns.

  4. 04

    Count the review work ↗

    Generated output is only productive when the team can accept and maintain it.

Official source-reference image used for orientation; not a result from this reading path. METR ↗ · Original research figure · Owner review pending. Select the image to enlarge.