Define the claim first
AGI can refer to breadth across tasks, human-level performance, autonomy over long horizons, or economic value. Those are different propositions. The Levels of AGI research proposes dimensions of performance and generality to make discussion more operational. A claim about a coding agent therefore needs a threshold: which software tasks, at what quality, with what tools, with how much human help, and over what duration? Without that specification, a striking demo can support an impression but cannot establish the category. The framework is a proposal for measurement, not a certification that a particular system has reached a level.
Inspect the evidence chain
Ask for a representative task distribution, a preregistered or clearly stated success rule, independent grading, and repeated trials. Separate a model's raw response from an agent scaffold that searches, tests, retries, and receives human hints. The SWE-bench literature gives a concrete example of repository issue solving, while later audits show that benchmark tasks and graders themselves need scrutiny. A strong claim should survive changes in task source and evaluation harness. A single leaderboard rank or cherry-picked video does not reveal the failure rate outside the displayed cases.
Probe generality and recovery
Software work includes specification, unfamiliar code, integration, security constraints, and maintenance after release. An evaluation should include ambiguous requirements, missing information, and situations where the correct action is to stop or ask for clarification. It should measure whether the system notices failure and recovers without hiding errors. Human supervision belongs in the accounting: if an expert continually steers the agent, the evidence supports a human-agent workflow, not unattended general capability. This is an editorial evaluation rubric derived from the measurement problem, not a statement about any current model's ultimate status.
Keep forecasts separate
A measured result, a provider's forecast, and a philosophical definition should occupy separate lines in a claim record. Record the source date and whether the evidence was independently replicated. Update a record when new evaluations arrive, but preserve the earlier version so a reader can see what changed. Avoid treating a time horizon extrapolation as a promised arrival date. The practical decision for a team remains narrower than the AGI label: what authority can this system safely receive on this task, and what evidence supports that grant today? This draft makes no broad declaration of current AGI status.
What to carry into the work
- Write an operational definition and threshold.
- Record tools, hints, retries, and grading method.
- Seek representative and repeated tasks with failure analysis.
- Separate measured performance from forecasts and labels.
Sources & dates
- Levels of AGI for Operationalizing Progress on the Path to AGI ↗Google DeepMind researchers · Undated source · Checked 16 Sept 2026
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues? ↗SWE-bench researchers · Undated source · Checked 16 Sept 2026
- Why SWE-bench Verified no longer measures frontier coding capabilities ↗OpenAI · 23 Feb 2026 · Checked 16 Sept 2026
Unknown source dates stay undated. Preparation is not publication; no historical byline or interview is implied.