The measured result
METR published a randomized study on July 10, 2025 involving 16 experienced open-source developers working on issues in repositories they knew well. The study examined tools available during February through June 2025. In the AI-allowed condition, measured completion time was 19% longer than without AI. That is a result for the sampled developers, tasks, and period, not an estimate for every team or today's models. The paper and project write-up also discuss participants' expectations and limitations; use the measured timing as the headline, with its scope attached.
Why the date matters
Tool behavior and developer practice can change quickly. In a February 24, 2026 update, METR described a later experiment that encountered selection effects: some developers did not want to work without AI, and the researchers judged the new data an unreliable signal of the current productivity effect. The update reports suggestive results but explicitly warns against treating them as a clean replacement estimate. It is therefore inaccurate either to freeze the 2025 slowdown into a timeless verdict or to announce a proven reversal from the later update.
What the experiment did not measure
Completion time on assigned repository issues is important, but it does not by itself capture long-term maintainability, security, developer learning, customer value, or the kinds of tasks a team chooses because AI exists. Nor does an experienced maintainer in a familiar codebase represent a novice making a prototype. A local team may face a different mix of tasks and review costs. These are limits on applying the result, not reasons to discard it. Randomization strengthens the causal claim inside the study setting; it does not automatically widen the population.
Use it to design a local check
Before buying a productivity narrative, choose a representative task family and define elapsed time through accepted review, not just first generated patch. Capture the agent setup, number of retries, reviewer minutes, and any defects discovered after the initial answer. Compare like with like and keep the original task difficulty visible. A small local trial will have its own uncertainty, but it can reveal whether the team's bottleneck is code entry, context transfer, review, or integration. This is an editorial recommendation derived from the study's scope, not a result METR reported for this site.
What to carry into the work
- Label the evidence as early-2025 tools and experienced maintainers.
- Keep the 19% result attached to its study setting.
- Note the 2026 update's selection limitations.
- Measure accepted work and review time in any local comparison.
Sources & dates
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ↗METR · 10 Jul 2025 · Checked 16 Sept 2026
- We are Changing our Developer Productivity Experiment Design ↗METR · 24 Feb 2026 · Checked 16 Sept 2026
Unknown source dates stay undated. Preparation is not publication; no historical byline or interview is implied.