METR’s expenditure horizon is not what headlines say
A widely repeated summary of METR’s July 2026 study inverts the definition and narrows the range; the underlying numbers hold, the interpretation does not.
What happened: 21 July 2026 · Written: 31 August 2026
The short answer
METR’s expenditure horizon is the budget at which humans become more cost-effective than agents, not the amount of human work an agent contributes. The published range is $0 to $3,300, not $2,300 to $3,300. Anyone quoting it for a buying case should state the wage assumption behind the figure.
The definition got inverted
A claim now circulating says METR’s expenditure horizon shows agents contribute only about $2,300 to $3,300 of equivalent human optimisation work. METR’s paper, published on 21 July 2026, says something different. It defines the measure as the point where two cost curves cross, described in the paper as the budget level beyond which humans become the more cost-effective option, and restated as the dollar value at which human and AI effort achieve the same progress.
That is a threshold on spending, not a measure of delivered output. Read as output, the number sounds like a small quantity of labour an agent hands over before running dry. Read as METR intends it, the number tells you where to stop spending: under that budget the agent buys more progress per dollar, above it a human does. The two readings point in opposite directions when you are sizing a tool budget.
The published range is wider at the bottom
The range in the paper is $0 to $3,300 across the models tested on the NanoGPT speedrun, not $2,300 to $3,300. Two of the six, GPT-5 and Opus-4.1, made no meaningful progress at all. The remaining four, after re-validation, land between $600 and $3.3K. The tighter range circulating in headlines is simply the two frontier models: GPT-5.5 at $2,300 and Opus-4.8 at $3,300.
Re-validation matters here. METR notes that raw agent trajectories overstate progress because of statistical noise, so the published figures are already a correction downward from what the agents appeared to achieve. Any vendor deck that quotes an uncorrected speedup from a similar run is quoting the inflated version. When a supplier shows you agent benchmark gains, ask whether anyone re-checked the results before they were reported.
Everything rests on one wage assumption
The dollar figures depend on what METR assumes human improvement costs. Its baseline is roughly 16 hours of human labour for each 1% improvement, priced at $150 an hour, rounded to $2,500 per 1% improvement. That is METR’s own estimate, and the paper says plainly that it carries a great deal of uncertainty. Change the wage or the hours and every headline number moves with it.
The paper runs three assumptions and shows how far the numbers travel. Under the easy assumption of $1,000 per 1%, Opus-4.8’s horizon falls to $120 and GPT-5.5’s to $160. Under the hard assumption of $10,000 per 1%, they rise to $14,400 and $9,400. The widely quoted $3,300 and $2,300 sit in the middle column, the medium assumption of $2,500 per 1%, and nowhere else.
METR states that the estimate is fairly sensitive to its assumption about returns to human effort. That is the honest reading: a single figure lifted out of the table without its assumption is close to meaningless, because the same model spans $120 to $14,400 depending on which column you pick. If you are handed one of these numbers, the first question is which wage assumption produced it.
What METR concluded, and what it did not test
METR’s own conclusion is modest. Although some models have expenditure horizons in the thousands of dollars, those sums are small next to the total human labour involved, which the paper reads as autonomous agent optimisation having had minimal effect on AI research progress in NanoGPT so far. That is a finding about one speedrun benchmark, not a general verdict on what coding agents do inside a business.
There is a second qualifier buyers should know. The NanoGPT maintainer estimated the mergeable share of the speedup at roughly 60% for Opus-4.8 and 50% for GPT-5.5, and judged that what the agent mostly explored, hyperparameter tuning, was of low novelty. Roughly half the measured gain would not have survived review. Benchmark progress and accepted work are not the same thing, and the gap is not small.
For anyone budgeting agent spend, the usable move is to copy the method rather than the numbers. Work out your own loaded hourly cost, decide what a unit of progress is worth in your team, then find the budget at which paying people beats paying for tokens. METR’s figures come from a machine learning speedrun with a specific wage assumption, and they will not transfer to your workload unchanged.
What to do about it
Treat the expenditure horizon as a method, not a benchmark score. Before signing an agent contract, calculate your own crossing point: your loaded hourly cost, the hours a unit of progress takes, and the token spend that matches it. If a supplier quotes METR’s $2,300 or $3,300, ask which of the three wage assumptions produced it.
Read our Claude review →Questions readers ask
What is METR’s expenditure horizon?
It is the point where the human and AI cost curves cross: the budget at which humans become more cost-effective than AIs, or equivalently the dollar value at which human and AI effort achieve the same progress. It is a spending threshold, not a count of the human work an agent replaces.
Is the expenditure horizon really $2,300 to $3,300?
No. METR published a range of $0 to $3,300 across the models tested, with $600 to $3.3K for the four that made positive progress. The $2,300 and $3,300 figures are GPT-5.5 and Opus-4.8 alone, and only under the medium assumption of $2,500 per 1% improvement.
Does this mean coding agents are not worth paying for?
It does not settle that. METR measured one benchmark, the NanoGPT speedrun, and concluded that autonomous agent optimisation has so far had minimal effect on progress there. The maintainer also put the mergeable share of the speedup at roughly 60% for Opus-4.8 and 50% for GPT-5.5, mostly low-novelty tuning.
Where every figure came from
Each claim above was checked against a primary source, then checked again by a second reader who had not seen the first check. Open any of them and verify us.
- METR defines the expenditure horizon as the crossing point of the human and AI cost curves — the budget beyond which humans are more cost-effective than AI. It is not the amount of human work an agent contributes. metr.org 2026-07-21
- A second phrasing of the same definition in the paper confirms it marks the budget at which both make equal progress. metr.org 2026-07-21
- Human labour is costed at about $2,500 per 1% improvement, from roughly 16 hours at $150 an hour. metr.org 2026-07-21
- The range METR actually publishes is $0 to $3,300, not $2,300 to $3,300. metr.org 2026-07-21
- GPT-5 and Opus-4.1 made no measurable progress, which is what a $0 horizon means. The other four models have positive horizons between $600 and $3,300. metr.org 2026-07-21
- The sensitivity table gives Opus-4.8 a horizon of $120 / $3,300 / $14,400 and GPT-5.5 $160 / $2,300 / $9,400 under the easy ($1,000 per 1%), medium ($2,500) and hard ($10,000) assumptions — so $3,300 and $2,300 hold only under the medium one. metr.org 2026-07-21
- METR warns the measure is sensitive to the assumed cost of human labour, so a single figure quoted without its assumption is misleading. metr.org 2026-07-21
- METR’s overall conclusion: although some models have horizons in the thousands of dollars, that is small next to total human labour, so autonomous agents have had minimal impact on AI R&D progress in NanoGPT so far. metr.org 2026-07-21
- The NanoGPT maintainer judged the share of agent contributions that were genuinely mergeable at roughly 60% for Opus-4.8 and 50% for GPT-5.5. metr.org 2026-07-21
More on AI agents
- Benchmarks Terminal-Bench 3.0 cuts top agent scores to 34% The same coding agents that passed 83.8% of Terminal-Bench 2.1 now clear 33.8%. Nothing regressed; the exam changed, and so should how you read vendor scores.
- Pricing GitHub Copilot moved to usage-based billing on 1 June Base prices did not move, but the meter did. The promotional credits that cushioned existing Business and Enterprise customers ran out at the end of August.
- Pricing Cursor Teams gains two usage pools and a $96 seat The June restructure separates first-party from third-party spend and adds a $96 seat. The interesting detail is what Cursor has not published about included usage.