Claude Opus 5 vs Fable 5 vs GPT-5.6 Sol

Claude Opus 5 vs Claude Fable 5 vs GPT-5.6 Sol benchmark comparison: 43.3 percent on Frontier-Bench, $5 and $25 per million tokens, 57 percent of agentic coding tasks still failed

Claude Opus 5 landed on July 24, 2026 at half the price of Claude Fable 5, and it takes the top score on Anthropic's headline agentic coding benchmark. It is not, however, Anthropic's most capable model, and the company's own documentation says so directly.

Key takeaways

  • Claude Opus 5 shipped July 24, 2026 at $5 per million input tokens and $25 per million output tokens, exactly half of Claude Fable 5's $10 and $50. Both carry a 1 million token context window and 128,000 tokens of maximum output.
  • Anthropic's own documentation is explicit about the hierarchy: Fable 5 is "Anthropic's most capable widely released model," and the guidance is to use it "for workloads that need the highest available capability." Opus 5 is positioned for "complex agentic coding and enterprise work."
  • On the three coding charts Anthropic published, the result is one each. Opus 5 takes Frontier-Bench v0.1, Fable 5 keeps the highest peak on CursorBench 3.2, and GPT-5.6 Sol leads the AA Coding Agent Index.
  • Those charts are score-versus-cost curves across five effort levels, not single scores. Any table that flattens them to one number per model is picking the winner by picking an effort setting.
  • Opus 5's real claim is cost per task. It reaches within 0.5% of Fable 5's CursorBench peak at half the cost, beats Fable 5 on OSWorld 2.0 at just over a third of the cost, and roughly 1.5x's the next-best pass rate on Zapier AutomationBench at the same cost.
  • Its best Frontier-Bench score is 43.3% at max effort. Even at the top of that leaderboard, the strongest agentic coding model available still fails most of the tasks it is given.

Published July 27, 2026. Corrected the same day. An earlier version of this article carried a ten-row benchmark table that mixed Anthropic's published charts with figures from secondary write-ups, including SWE-bench and DeepSWE scores that Anthropic does not publish for Opus 5 anywhere. Those rows have been removed. Everything below is read directly off Anthropic's launch charts, quoted from the launch announcement, or taken from the model comparison table in Anthropic's documentation. Where Anthropic states a claim without publishing a number, we say so rather than sourcing one elsewhere.

Claude Opus 5 vs Claude Fable 5 vs GPT-5.6 Sol benchmark comparison: 43.3 percent on Frontier-Bench, $5 and $25 per million tokens, 57 percent of agentic coding tasks still failed
Claude Opus 5 launched July 24, 2026. It is the cheapest model at the frontier, and by Anthropic's own documentation it is not the most capable one.

What Claude Opus 5 actually is

The model id is claude-opus-5. Per Anthropic's model comparison table, it carries a 1 million token context window, 128,000 tokens of maximum synchronous output (300,000 in batch mode behind a beta header), and a training cutoff of May 2026. It is available through the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry, plus Claude.ai, Claude Code, and Claude Cowork. A Fast mode runs it at roughly 2.5 times the default speed for twice the base price. Two beta API features shipped alongside it: mid-conversation tool changes, and automatic fallbacks between models.

The positioning is unusually candid. Anthropic's launch announcement opens by saying Opus 5 "comes close to the frontier intelligence of Claude Fable 5 at half the price." The documentation is blunter still: Fable 5 is "Anthropic's most capable widely released model," and the recommendation is to reach for it "for workloads that need the highest available capability," while Opus 5 is the pick "for complex agentic coding and enterprise work." Labs do not usually introduce a flagship by telling you another of their models is smarter. That framing is the story, and most of the launch coverage inverted it.

The three coding charts

Anthropic published exactly three coding comparisons on the launch page. Here is what each one shows.

Chart Claude Opus 5 Claude Fable 5 GPT-5.6 Sol Leader
Frontier-Bench v0.1 43.3% at max effort, 44.4% at xhigh about 33.7% at its peak about 37.5% at its peak Opus 5
CursorBench 3.2 within 0.5% of Fable 5's peak, at half the cost per task highest peak score on the chart below both Claude models at peak Fable 5
AA Coding Agent Index peaks just under Sol below both, and the most expensive curve sits above Opus 5 across most of the cost range GPT-5.6 Sol

One each. Opus 5 wins Frontier-Bench outright and by a wide margin, more than doubling Opus 4.8 on the same eval. Fable 5 keeps the highest peak on CursorBench 3.2, which Cursor's own co-founder confirms in a quote Anthropic chose to publish: "On CursorBench it's just under Fable 5." And on the AA Coding Agent Index, GPT-5.6 Sol's curve runs above Opus 5 through most of the cost range.

Why "who wins" is the wrong question here

Every one of those charts plots score against cost per task on a log scale, with five points per model, one for each effort level from low to max. They are not bar charts. There is no single number for any model on any of them.

This matters more than it sounds. Opus 5 scores 43.3% on Frontier-Bench at max effort and 44.4% at xhigh, which means more compute made it slightly worse, and which of those two you quote is a choice. On CursorBench, Opus 5 beats every other model at any given cost on high, xhigh, and max effort, while still finishing below Fable 5's peak. Both of those statements are true at once. A table that reduces each model to one score is choosing a winner by choosing an effort setting, and that is precisely how the launch-day coverage produced a clean sweep that the underlying data does not show.

Where Opus 5 genuinely leads

Beyond coding, Anthropic makes a series of claims about Opus 5 in prose without publishing the underlying scores. We are reproducing them as claims, attributed, rather than borrowing numbers from elsewhere and presenting them as measurements.

Evaluation Anthropic's claim Score published?
ARC-AGI 3 Opus 5 scores three times as high as the next-best model No number published
OSWorld 2.0 Beats Fable 5's best result at just over a third of the cost No number published
Zapier AutomationBench Pass rate about 1.5x the next-best model at the same cost per task No number published
GDPval-AA v2 Opus 5 is the new state of the art on knowledge work No number published
HLE, DeepSearchQA Anthropic's best and most cost-efficient model on both No number published
OSS-Fuzz (cyber) Close to Mythos 5 at finding vulnerabilities, far behind at exploiting them No number published

The ARC-AGI 3 result is the most interesting of these, because ARC-AGI is built specifically to resist memorisation. A three-fold margin over the next-best model on novel problem solving is the strongest available evidence that Opus 5 is doing something genuinely new rather than fitting benchmarks better. Anthropic also reports that Opus 5 is its most aligned model to date, scoring 2.3 on its automated behavioural audit for misaligned behaviour, ahead of Opus 4.8, Sonnet 5, and Fable 5.

Fable 5 is still the capability ceiling

If you take one thing from this release, take the hierarchy Anthropic itself publishes. Fable 5 remains the most capable model the company sells to the general public. Opus 5 gets close to it, beats it on specific agentic and computer-use work, is dramatically better value, and is the new default on Claude Max. None of that makes it the smarter model, and Anthropic never claims it is.

The pattern in the evidence is consistent: Opus 5 wins where planning, autonomy, and cost-efficiency are being measured, and Fable 5 holds on where raw peak capability is. Cognition's Scott Wu, quoted on the launch page, puts the same boundary on FrontierCode 1.1: Opus 5 "approaches Fable-level performance at half the cost." Approaches. If your work sits at the hard end and budget is not the binding constraint, our Claude Fable 5 review still describes the model you want.

The 57% nobody quotes

Two caveats belong on the headline number.

First, the configuration. Anthropic's Frontier-Bench footnote states that results come from an internal run on the mini-SWE-agent harness and a GKE backend, as mean reward over five attempts per task, and that Opus 4.8 served as fallback on safety-classifier refusals for both Opus 5 and Fable 5. That is a reasonable engineering choice for a shipping product, but it means the published figure describes a system rather than a model. Anthropic separately notes that Opus 5's cyber classifiers should intervene around 85% less often than Fable 5's, so the two models are not leaning on that fallback equally.

Second, and more useful: 43.3% is the best score anyone has posted on this benchmark, and it means the leading agentic coding model still fails the majority of the tasks. The practical reading is not that the model is bad. It is that a sub-coin-flip pass rate is the current state of the art, so unsupervised agent runs remain a review problem rather than a solved one. Budget the review time. That is the same conclusion we reached writing about Kimi K3 against Fable 5, and no release since has changed it.

The pricing math

Model Input per 1M Output per 1M Notes
Claude Opus 5 $5.00 $25.00 1M context, 128K max output
Claude Fable 5 $10.00 $50.00 same context and output limits
Claude Opus 5, Fast mode $10.00 $50.00 about 2.5x the default speed
Claude Sonnet 5 $3.00 $15.00 for reference, not in this comparison

Token prices only tell you part of it. Cost per task is what lands on the invoice, and it depends on how many tokens a model burns to finish a job. This is where Opus 5 is genuinely unmatched: within 0.5% of Fable 5's CursorBench peak at half the cost per task, past Fable 5's best OSWorld 2.0 result at just over a third of the cost, and about 1.5 times the next-best AutomationBench pass rate for the same spend. Anthropic's framing that Opus 5 "works more efficiently than other models" is the accurate summary of this release.

What to run for what

Default to Opus 5 for agentic and long-horizon work: multi-step coding agents, computer use, research and browsing, and anything where the model plans rather than transcribes. It is the new default on Claude Max, it is half the price, and on cost-adjusted terms nothing else is close.

Reach for Fable 5 when you need the ceiling and the budget allows it, which is exactly the guidance Anthropic gives. Keep GPT-5.6 Sol in the evaluation set if agent-index-style tasks dominate your backlog, since it leads that chart.

If the motivation is cost rather than capability, note that halving the per-token price does nothing about a workflow that wastes tokens. Fixing token spend on the models you already run usually recovers more budget than a model switch does. And before migrating anything on the strength of a chart, run the contenders on your own tasks, exactly as we argued in our Grok 4.5 versus Claude versus GPT comparison. This article is a second-hand reading of someone else's evaluation harness, and so is every other one you will find this week.

Worth a footnote: we covered the pre-launch leaks in our Claude Opus 5 release date breakdown. The Honeycomb codename was real, the July window was right, and the pricing landed cheaper than the leak threads guessed.

FAQ

Is Claude Opus 5 better than Claude Fable 5?

Not in raw capability. Anthropic's documentation calls Fable 5 "Anthropic's most capable widely released model" and recommends it for workloads needing the highest available capability. Opus 5 comes close at half the price, wins Frontier-Bench outright, and beats Fable 5 on computer use and cost per task, but Fable 5 keeps the higher peak on CursorBench 3.2 and remains the capability ceiling.

How much does Claude Opus 5 cost?

$5 per million input tokens and $25 per million output tokens, the same as Opus 4.8 and exactly half of Claude Fable 5's $10 and $50. Fast mode doubles that to $10 and $50 for about 2.5 times the speed.

Is Claude Opus 5 better than GPT-5.6 Sol for coding?

It splits. Opus 5 wins Frontier-Bench v0.1 by a wide margin, roughly 43.3% against about 37.5% at peak. GPT-5.6 Sol leads the AA Coding Agent Index, where its curve sits above Opus 5 across most of the cost range. Anthropic publishes no head-to-head SWE-bench numbers for Opus 5, so treat any table you see quoting them with caution.

Should I switch to Claude Opus 5?

If you are paying for Fable 5 and your work is agentic, the case is strong: comparable results at half the price, and better results on computer use. If you are at the hard end of capability, Anthropic's own guidance is to stay on Fable 5. Either way, run your own evals before migrating production traffic.

Valletta.Software - Top-Rated Agency on 50Pros

Your way to excellence starts here

Start a smooth experience with Valletta's staff augmentation

Claude Opus 5 vs Fable 5 vs GPT-5.6 Sol: Benchmarks