OfficeBooks
Benchmarks

Terminal-Bench 3.0 cuts top agent scores to 34%

The same coding agents that passed 83.8% of Terminal-Bench 2.1 now clear 33.8%. Nothing regressed; the exam changed, and so should how you read vendor scores.

What happened: 30 July 2026 · Written: 31 August 2026

The short answer

On 30 July 2026 Terminal-Bench 3.0 replaced a saturated benchmark with 74 harder tasks. Top agent scores fell from 83.8% to 33.8%. Anyone comparing coding CLIs or quoting vendor pass rates is affected. Check which version and date any figure refers to, and run your own tasks before committing.

Terminal-Bench 3.0 pass rate per cent of 74 tasks solved GPT-5.6 Sol, Codex 34.4% Fable 5, Claude Code 33.8% Opus 4.8, Claude Code 21.1% GPT-5.6 Terra, Codex 20.8% Grok 4.5, Cursor CLI 17.8% Sonnet 5, Claude Code 14.6% GPT-5.6 Luna, Codex 14.3% GLM 5.2, Claude Code 5.1%
Scores on Terminal-Bench 3.0 only. Earlier versions are a different test and are not mixed in.

The leaderboard fell through the floor

On 30 July 2026 the team behind Terminal-Bench published version 3.0, and the leaderboard fell through the floor. Fable 5 running in Claude Code scored 83.8% on Terminal-Bench 2.1. On the new set it scored 33.8%. Opus 4.8 in Claude Code went from 78.9% to 21.1%. The maintainers say the first release contains 74 tasks across seven domains, and that the best models reach roughly 34%.

The full table is short enough to read in one go. GPT-5.6 Sol in Codex leads on 34.4%, Fable 5 in Claude Code follows on 33.8%, then Opus 4.8 on 21.1%, GPT-5.6 Terra in Codex on 20.8%, Grok 4.5 in Cursor CLI on 17.8%, Sonnet 5 in Claude Code on 14.6%, GPT-5.6 Luna in Codex on 14.3%, and GLM 5.2 in Claude Code on 5.1%.

Nothing got worse overnight

Nothing about these agents degraded between the two runs. The same models were measured against a harder set of tasks, and the maintainers are explicit about why they built one: many of the older tasks had become saturated, and the leaderboard had compressed into a band too narrow to show real capability gaps. A benchmark where the top six sit between 74.6% and 83.8% has stopped telling buyers much.

The separation argument is the strongest part of the case. On Terminal-Bench 2.1, Fable 5 and Opus 4.8 in Claude Code were 4.9 points apart. On 3.0 the gap is 12.7 points. That is the maintainers’ own framing, and it is the thing worth checking rather than the headline crash: a benchmark earns its keep by spreading models out, not by making them look bad.

What the table does not tell you

Read the labels carefully, because every score is a pairing of a model and a harness. Fable 5, Opus 4.8 and Sonnet 5 were run in Claude Code, the three GPT-5.6 variants in Codex, and Grok 4.5 in Cursor CLI. Nothing here isolates the model from the tool wrapped around it, so 34.4% is a statement about Sol inside Codex. Gemini CLI does not appear in the published table at all.

Scale is the other caveat, and it comes from the maintainers themselves: 74 tasks across seven domains is a first release, not a settled instrument. At that size a handful of awkward tasks carries real weight in the ordering, particularly among the models bunched in the teens. The version churn is real too. Terminal-Bench 4.0 arrived on 28 August 2026, described as calibrating task resources, fixing tasks and removing saturated ones, a month after 3.0 landed.

How to use this when you are buying

The practical reading for anyone choosing a coding agent is narrow. Treat the 3.0 table as a ranking signal, not a success rate you can plan around; roughly a third of tasks passed is not a number to build a business case on. If a vendor quotes 83.8%, or anything in that band, ask which benchmark version it came from and on what date it was measured.

Then do the boring thing. Take five or six jobs your team actually runs in a terminal, script them, and put the shortlist through them in the harness you would deploy. The published tables settle which tools are worth the trouble of testing. They do not settle which one clears your work, and at these pass rates a fair share of it comes back unfinished.

Expect the ground to move. Two versions shipped within a month, and the 4.0 notes say tasks were fixed and removed. Any figure you cite carries a version and a date, so record both wherever your team keeps its shortlist, and revisit that shortlist when the next release lands rather than treating a leaderboard screenshot captured today as a durable fact.

What to do about it

Stop quoting benchmark percentages without a version and a date. Pull the Terminal-Bench 3.0 table for your shortlist, use it only to decide what is worth trialling, then script five or six real terminal jobs your team runs weekly and test each candidate in the harness you would actually deploy. Recheck when the next version ships.

Read our CodeRabbit review →

Questions readers ask

Why did Terminal-Bench scores drop so much in version 3.0?

The models did not get worse. The maintainers replaced a saturated task set with a harder one because the old leaderboard had compressed into a narrow band. Fable 5 fell from 83.8% on 2.1 to 33.8% on 3.0 measured against different tasks.

What is the highest score on Terminal-Bench 3.0?

GPT-5.6 Sol running in Codex leads the published table on 34.4%, with Fable 5 in Claude Code on 33.8%. The maintainers summarise the field as reaching roughly 34%, across 74 tasks in seven domains.

Is Terminal-Bench 3.0 still the current version?

No. Terminal-Bench 4.0 was published on 28 August 2026, described as calibrating task resources, fixing tasks and removing saturated ones. Any pass rate you quote should carry both a version number and a date.

Where every figure came from

Each claim above was checked against a primary source, then checked again by a second reader who had not seen the first check. Open any of them and verify us.

  1. The Terminal-Bench 3.0 announcement was published on 30 July 2026 on the tbench.ai blog. tbench.ai 2026-07-30
  2. The first Terminal-Bench 3.0 release contains 74 tasks across seven domains, and the best model scores about 34%. tbench.ai 2026-07-30
  3. Terminal-Bench 3.0 pass rates: GPT-5.6 Sol on Codex leads at 34.4%, Fable 5 on Claude Code 33.8%, Opus 4.8 on Claude Code 21.1%, GPT-5.6 Terra on Codex 20.8%, Grok 4.5 on Cursor CLI 17.8%, Sonnet 5 on Claude Code 14.6%, GPT-5.6 Luna on Codex 14.3%, GLM 5.2 on Claude Code 5.1%. tbench.ai 2026-07-30
  4. Head to head, Fable 5 scored 83.8% on Terminal-Bench 2.1 and 33.8% on 3.0; Opus 4.8 fell from 78.9% to 21.1%. tbench.ai 2026-07-30
  5. The stated reason for a harder release: many Terminal-Bench tasks had saturated and the leaderboard had compressed into a narrow band. tbench.ai 2026-07-30
  6. Terminal-Bench 3.0 separates frontier from sub-frontier models better than 2.1 — the gap between Fable 5 and Opus 4.8 widens from 4.9 points to 12.7. tbench.ai 2026-07-30
  7. Terminal-Bench 4.0 followed on 28 August 2026, recalibrating task resources, repairing tasks and removing saturated ones. tbench.ai 2026-08-28

More on AI agents