Vals.ai disclosed that Claude Opus 5's Terminal-Bench 2.1 score included nine tasks an older model completed after Opus 5 refused them.
Claude Opus 5 is ranked first on vals.ai's Vals Index, the composite leaderboard the evaluator uses to rank AI models across coding, finance, legal, and other real-world tasks. Part of that ranking is built from a Terminal-Bench 2.1 score that, according to vals.ai's own methodology note, included nine tasks Opus 5 refused to attempt and a different, older model finished instead.
Sources: Vals AI: Claude Opus 5 model page
On August 26, 2026, vals.ai updated its Terminal-Bench 2.1 leaderboard page with a line most readers of the headline percentage will never see. Both Claude Fable 5 and Claude Opus 5 ran the benchmark with Claude Opus 4.8, the Opus-tier model Opus 5 replaced, wired in as what vals.ai calls a refusal fallback: whenever the newer model declined a task outright, Opus 4.8 attempted it instead, and a pass from Opus 4.8 counted toward Opus 5's score.
Sources: Terminal-Bench 2.1 leaderboard, updated 8/26/2026 (vals.ai)
This finding argues for reading a benchmark's methodology page before repeating its headline number, and that habit is close to what we build for a living. TLM Forge exists to gate AI-written code on evidence rather than a green checkmark, and a leaderboard score with an undisclosed asterisk is the same failure in a different discipline: a result that looks verified until someone checks what actually produced it. Read the vals.ai findings as the load-bearing claim here, and read the one product section later knowing where it comes from.
What vals.ai actually disclosed
Vals.ai substituted a different model whenever Claude Opus 5 or Claude Fable 5 refused a task outright, and counted that substitute's success as the flagship model's own pass. Terminal-Bench 2.1 grades an agent on 89 sandboxed command-line tasks, maintained by Stanford, Harbor, and the Laude Institute, and leaderboard submission requires public trajectories so anyone can check how a score was produced. Vals.ai runs each model through the benchmark's harness and scores every task pass or fail. When Opus 5 or Fable 5 refused a task rather than attempting it, the kind of safety-triggered decline that has nothing to do with terminal skill, vals.ai reran that same task on Claude Opus 4.8, the model Opus 5 replaced roughly two months earlier, and recorded a fallback pass as a pass.
Sources: Terminal-Bench 2.1 announcement (tbench.ai)
Nine of Claude Opus 5's passing results on Terminal-Bench 2.1 came from Opus 4.8 this way. Vals.ai's own note states it plainly: counting those nine as failures, the same way a refusal is scored everywhere else on the benchmark, drops Opus 5's Terminal-Bench 2.1 score from 84.64% to 81.27%. Vals.ai did not bury this entirely: the leaderboard page carries a small toggle beside the score that switches between the fallback-included and fallback-excluded numbers. But the number that sorts the leaderboard, gets screenshotted, and shows up in any recap of "how Opus 5 scored on Terminal-Bench" is the one with Opus 4.8's work folded in by default.
Sources: Terminal-Bench 2.1 leaderboard, updated 8/26/2026 (vals.ai)
The number that actually moved
The correction does not just trim Opus 5's Terminal-Bench 2.1 score. It roughly quadruples the gap to first place and nearly erases the cushion over third. GPT-5.6 Sol, OpenAI's top-tier coding model, leads the public leaderboard at 85.77%. At the headline 84.64%, Opus 5 trails Sol by 1.13 points, close enough to read as a near-tie for the top spot. At the corrected 81.27%, the gap to Sol widens to 4.50 points, a clear separation rather than a photo finish.
Sources: Terminal-Bench 2.1 leaderboard mirror, sourced from vals.ai (BenchLM)
| Model, Terminal-Bench 2.1 rank | Headline score | Score with refusals scored as failures |
|---|---|---|
| GPT-5.6 Sol (OpenAI), #1 | 85.77% | 85.77%, no disclosed fallback |
| Claude Opus 5 (Anthropic), #2 | 84.64% | 81.27%, nine fallback-assisted passes removed |
| Kimi K3 (Moonshot AI), #3 | 80.90% | 80.90%, no disclosed fallback |
| Claude Fable 5 (Anthropic), #4 | 80.52% | Also ran with the Opus 4.8 fallback; adjusted figure not published |
| GPT-5.6 Luna (OpenAI), #5 | 79.03% | 79.03%, no disclosed fallback |
Third place belongs to Kimi K3 at 80.90%, unaffected by any disclosed fallback. At the headline number, Opus 5 clears Kimi K3 by 3.74 points, a comfortable margin. At the corrected number, that margin shrinks to 0.37 points, the kind of gap a single difficult task can close. Nine substituted passes are not a rounding error on this leaderboard. They are most of the distance between a clear second place and a lead over third place that barely exists.
A benchmark score is not a fact about a model alone. It is a fact about a model, a harness, and a scoring rule together, and any one of those three can move the headline number without the model itself getting better or worse.
Why a flagship model refuses a benign terminal task
Anthropic built Opus 5 to refuse more, not less. The company's own launch announcement calls Opus 5 "our most aligned model to date," with the lowest rate of deceptive behavior of any Opus release and the least susceptibility to being talked into misuse of any prior model. That kind of tuning does not always distinguish a genuinely dangerous request from a command-line task that merely looks risky: deleting a directory, killing a process, rewriting a configuration file, anything a cautious safety classifier can read as destructive without the full context a human would use. A model tuned to decline more often will decline more often, including on tasks a benchmark grader intended as ordinary. Anthropic's own Opus 5 announcement does not mention Terminal-Bench at all; the only place most people will encounter this specific score is a third-party evaluator, which is exactly why what that evaluator discloses, and how visibly, matters.
Sources: Introducing Claude Opus 5 (Anthropic)
What the number one spot is actually built from
Claude Opus 5's overall standing on vals.ai is not one number. It is an aggregate built from parts. The Vals Index, which vals.ai describes as a single read on AI's economic impact across finance, coding, and legal work, weighted by each sector's share of United States GDP, aggregates seven separate benchmarks and ranks Opus 5 first among every model it tracks. Terminal-Bench 2.1 is one of the coding-sector benchmarks that composite score is built from, alongside Vibe Code Bench and Code Migration. The same fallback recount that pulls Opus 5's Terminal-Bench 2.1 score down to 81.27% also pulls Opus 5's own Vals Index score, the number behind that number one ranking, from 74.82% to 74.47%, according to the same vals.ai methodology note.
Sources: Vals Index methodology (vals.ai), Vals AI: Claude Opus 5 model page
Before quoting a benchmark percentage to justify a model choice, open the methodology page and search it for the words fallback, refusal, and excluded. If a leaderboard has a toggle that changes the score, that toggle is telling you the default number was a choice, and the choice was made by whoever built the leaderboard, not by the model being scored.
This is not one benchmark's problem
The same substitution shows up in more places than Terminal-Bench 2.1, and the tuning decision that caused it is not unique to Anthropic.
Beyond Terminal-Bench 2.1
Vals.ai's note names two more of Opus 5's scores touched by the same Claude Opus 4.8 substitution. MMLU Pro moves from 91.59% to 91.58%, a change small enough to be rounding noise on a test with thousands of questions, but real enough to confirm the fallback reached beyond agentic coding tasks into a knowledge benchmark. The Vals Multimodal Index moves from 73.90% to 73.58%. Neither swing is dramatic on its own. Together with the 3.37-point move on Terminal-Bench 2.1, they show a single substitution decision reaching into at least four different published numbers for the same model release.
Sources: Vals AI: Claude Opus 5 model page
Beyond Anthropic
Over-refusal is a documented, industry-wide tradeoff, not a one-model story. A 2026 audit of 21 language models across four safety benchmarks found that refusal rate alone is a poor stand-in for actual safety: model families tuned to suppress unsafe output tend to also suppress a share of benign requests, while families tuned to stay helpful tend to tolerate more of the harmful requests they were supposed to catch. Vals.ai's fallback exists because that tradeoff is structural, not a bug specific to one release. Separating "can the model do the task" from "did the model's own safety training get in the way" is a reasonable thing for an evaluator to want to measure.
What is not reasonable is folding a substitute model's success into the flagship's headline number without a visible flag on the number itself. Any evaluator facing over-refusal has roughly three options: score every refusal as a failure and risk understating a model that is capable but overcautious, exclude refused tasks from the denominator and risk a smaller, favorable sample, or substitute another model's attempt and disclose that clearly next to the score. Vals.ai chose the third option and disclosed it in a methodology note rather than in the score itself. None of the three choices is wrong on its own. The mistake is letting a headline number imply that the first or second choice was made when the third one actually was.
How to read a leaderboard before you pick a coding model
A single percentage is a starting point for evaluating a coding model, not a decision. Model performance for coding is rarely one clean number even before a fallback mechanism gets involved, since a model that is strong on algorithmic puzzles and weak on multi-file refactors still averages out to a score that hides both facts. Terminal-Bench 2.1 adds a second layer on top of that: the number itself can be a blend of two different models' work, sorted and displayed as though it described one.
- Open the methodology page, not just the leaderboard row, before repeating a benchmark score in a decision document.
- Check whether the benchmark publishes per-task trajectories. Terminal-Bench 2.1 requires them for leaderboard submission, which is what let vals.ai's fallback usage surface at all.
- Look for a toggle, footnote, or asterisk near the score. Its presence means the default number was a choice, and the choice usually favors the higher figure.
- Compare the margin to the model directly below on the leaderboard, not only the model above it. A shrinking cushion over third place can be the real story, not a narrow gap to first.
- Treat a single benchmark, however well built, as one input alongside your own task-specific evaluation before standardizing a team on a model.
None of that replaces running the model on your own repository. Choosing a model for AI coding is a decision that benchmark scores can inform and should never fully make, because Terminal-Bench 2.1's sandboxed command-line tasks are not your codebase, your review process, or your production incident history.
The habit this story argues for, verifying what actually produced a passing result before trusting the top-line number, is the same habit TLM Forge applies to code instead of benchmarks. A pull request showing green tests is not evidence the change is safe, any more than an 84.64% is evidence Opus 5 did the work alone; both numbers can hide a substitution that only surfaces if someone checks. TLM's convergence gate exists to force that check on every diff, with independent reviewers examining what actually happened rather than what a summary claims happened, before a change ships.
Which benchmark claims a team has already checked, and which caveats they found, is exactly the kind of decision that gets lost between one evaluation cycle and the next. That record belongs in a benchmark-tracking doc a team actually maintains, not in whichever engineer happened to read the methodology page first. MemX, from the same team behind TLM Forge, is built for an individual's version of durable memory: a private, persistent memory layer for personal photos, documents, voice notes, and messages, not team benchmark notes.
What actually changed on August 26
Nothing about Claude Opus 5's underlying capability changed when vals.ai published its methodology note. What changed is what the 84.64% is now known to measure: Opus 5's own performance on 89 sandboxed terminal tasks, plus Claude Opus 4.8's performance on nine of them, reported as if it described one model. Opus 5 still leads 14 separate vals.ai leaderboards outright, and it is still, on the Vals Index, the top-ranked model on the site. The specific claim that does not survive the methodology note is that its Terminal-Bench 2.1 score, and the composite score built partly from it, describe Opus 5 alone. They describe Opus 5 with an older model quietly finishing what it would not attempt, and only one of those two models is the one anyone is actually paying to use.
Frequently asked questions
01What is the Claude Opus 4.8 refusal fallback on Terminal-Bench 2.1?
It is a substitution vals.ai built into its evaluation run: when Claude Opus 5 or Claude Fable 5 refused a Terminal-Bench 2.1 task outright, vals.ai reran that task with Claude Opus 4.8, an older Anthropic model, and counted an Opus 4.8 pass toward the newer model's headline score.
02How much did the fallback change Claude Opus 5's Terminal-Bench 2.1 score?
Nine of Opus 5's passing results came from Claude Opus 4.8. Counting those nine as failures, the standard way a refusal is scored elsewhere on the benchmark, drops Opus 5's score from 84.64% to 81.27%, according to vals.ai's August 26, 2026 methodology note.
03Is Claude Opus 5 still ranked first anywhere despite the fallback?
Yes. Vals.ai ranks Opus 5 first on 14 of its leaderboards, including the Vals Index, its composite score across coding, finance, and legal benchmarks. On Terminal-Bench 2.1 specifically, Opus 5 ranks second behind GPT-5.6 Sol at 85.77%, both before and after the fallback correction.
04Why did Claude Opus 5 refuse tasks that an older model then completed?
Anthropic built Opus 5 to be its most aligned model yet, with a lower rate of misuse and deceptive behavior than prior releases. That tuning does not always distinguish a genuinely harmful request from a command-line task that merely looks destructive, which pushes refusal rates up on benign benchmark tasks too.
05How should engineers evaluate a coding model beyond one benchmark score?
Read the methodology page for fallbacks, exclusions, or toggles before repeating a score, compare the margin to the model below it on the leaderboard as well as above, and test the model against your own codebase and review process rather than relying on a single published number.