Best LLM for Coding: Opus 5.5 vs GPT-6 vs Gemini 3.8

The best LLM for coding in September 2026 is Claude Opus 5.5 on most of the public evidence, and the honest version of that sentence is longer. Six frontier models shipped in three weeks, the cheapest of them costs 100 times less per token than the most expensive, and the single most quoted coding benchmark on the internet has not accepted a new submission since February. This page is the current state of the field, with every number's source and configuration attached.
Key takeaways
- Claude Opus 5.5 leads the independent Artificial Analysis Intelligence Index at 58, ahead of GPT-6 Astra and Claude Fable 5.1 at 53. It is also the cheapest of that group to run per task. That combination is unusual and it is the main thing that changed this month.
- More effort can make a model worse. On ARC-AGI-2, Opus 5.5 at high effort scores 93.3% at $0.41 per task. The same model at max effort scores 91.7% at $1.85. That is a lower score for 4.5 times the money, and the same non-monotonic pattern shows up on GPT-6 Astra.
- Stop quoting SWE-bench Verified. The leaderboard has accepted no submission since February 26, 2026. No GPT-6 model, no Opus 5.x, no Fable, no Gemini 3.8 appears on it. The widely repeated "around 79%" figure is an unverified third-party submission from December 2025 that the SWE-bench team never checked. Aider's polyglot board is worse: unmaintained since November 2025.
- GPT-6 Sol is not OpenAI's flagship and most of the coverage says otherwise. OpenAI's own words: "GPT-6 Astra continues to be our best model across the board." Sol is the mid tier and Luna is the volume tier.
- The "50% price cut" on Sol and Luna is half of GPT-5.6's promotional price for the same tier, not half of anything at the frontier. Astra was not repriced. And OpenAI's own arithmetic on Luna is off: $1.20 to $0.50 of output is a 58.3% cut, not 50%.
- Google has no current top-tier entry. Gemini 3.8 Flash is the newest and most capable Gemini, and it sits at 41 on the same index. There is no Gemini 3.8 Pro. Pro is still at 3.1, in preview since February. Flash's price doubles on January 1, 2027.
- Cost to run a fixed evaluation suite spans roughly 30x across vendors, and per-task cost inside a single model spans 4x depending only on the effort setting. The rate card is the smallest part of what you pay.
Last updated September 23, 2026. This page is maintained rather than reposted: it was first published on July 27, 2026 comparing Claude Opus 5, Fable 5 and GPT-5.6 Sol, and has been rewritten for the current lineup. Prices and model IDs come from each vendor's own pricing and documentation pages. Benchmark figures are labelled as either vendor self-reported or independently measured, because in this field those are different kinds of claim. Where a vendor's own chart shows it losing, we keep the row.
The frontier, priced
Every model below is available through a public API today. Prices are per million tokens, in US dollars, taken from each vendor's own rate card: Anthropic, OpenAI and Google. Read the notes column; three of these prices are conditional.
Two observations before the scores. The spread between GPT-6 Luna and the $50-output models is a factor of 100 on output tokens, which is larger than the spread in what they can do. And the context window race is over: almost everything here is at or near a million tokens, and nobody advertised a bigger number this month.
What the independent leaderboards actually say
Vendor charts are marketing artifacts with real numbers in them. The independent boards are the closest thing to a referee, and right now they disagree with each other in instructive ways. The four worth reading are Artificial Analysis, Terminal-Bench, ARC Prize and Arena.
The two dead boards matter more than the live ones, because they are still being quoted. If you read a September 2026 comparison that ranks models by SWE-bench Verified, the author has not opened the leaderboard. The number most often attached to it, somewhere around 79%, comes from an unverified third-party agent submission made in December 2025 that the maintainers never checked. The team-checked ceiling is 74.4%, on models that are now two generations old.
The finding nobody is reporting: more effort, worse results
ARC Prize publishes every run at every effort level with its cost attached, which makes it the most useful board in the field right now, and it shows something that breaks the mental model most teams are using.
Claude Opus 5.5 on ARC-AGI-2
- At high effort: 93.3% correct, $0.41 per task.
- At max effort: 91.7% correct, $1.85 per task.
4.5 times the cost, 1.6 points worse. GPT-6 Astra shows the same shape, scoring the same at medium as at high. Effort is not a quality dial.
This is not a rounding artifact, and it is not unique to one model. It means the standard cost-optimisation advice has the sign backwards: teams routinely raise effort when output quality disappoints, and on at least some task families that spends more money to get a worse answer. The only way to know which family you are in is to run your own evaluation at two or three effort levels and look at the curve.
It also explains why the vendor charts and the independent boards keep disagreeing. Anthropic reports Opus 5.5 at 66.4% on Terminal-Bench 4.0, measured at xhigh effort. The public Terminal-Bench 4.0 leaderboard, last refreshed the day before Opus 5.5 shipped, tops out at 58.18%. Those are not contradictory claims, they are different harnesses and different configurations, and until an independent run lands the 66.4% remains a vendor self-report.
How to read a frontier benchmark in 2026
Every number on this page comes with a configuration attached, and the configuration is doing more work than the number. Four traps account for almost every misleading model comparison published this year, including some of ours.
The four traps
- The effort footnote. A score is published at one thinking-effort setting and the model ships on another. Anthropic's Opus 5.5 chart is run at max and xhigh effort; the model defaults to medium. The headline number is a ceiling you have to pay extra to reach, not the out-of-the-box result.
- The flattened curve. Several vendors publish score-versus-cost curves across effort levels, not single scores. Any table that collapses a curve into one cell has picked the winner by picking an effort setting. If a source cannot tell you which point on the curve it quoted, it does not know what it published.
- The borrowed score. When a lab skips a benchmark, write-ups fill the hole with the previous model's number under the new model's name. Anthropic published no GPQA, AIME, MMMLU, tau-bench or ARC-AGI figures for Opus 5.5. Every table you see carrying those rows has borrowed them.
- Price per token instead of cost per task. The two move independently. A model can be 20% cheaper per token and 40% cheaper per job, or cheaper per token and more expensive per job, depending on how many tokens it burns thinking. Only one of those numbers appears on your invoice.
The fourth trap is the one that changed this month, so it deserves its own section.
Price per token is no longer the price
For two years the frontier was easy to price: multiply tokens by a rate card. That broke once every lab shipped an adjustable thinking budget, because the same model at two effort settings is effectively two products with the same rate card and very different bills.
Claude Opus 5.5 is the cleanest illustration. Its list price is 20% below Opus 5. Anthropic separately states that on typical workloads at default settings it costs 40% less. Both are true, and the gap is made of three things that never appear on a pricing page: the default effort level dropped from high to medium, the cache read multiplier dropped from 0.1x of input to 0.05x, and the model is cheaper to serve. Pushed the other way, Anthropic also says that at a fixed effort level Opus 5.5 thinks more per turn than Opus 5 did, most of all at xhigh and max.
The practical consequence is uncomfortable for anyone building a comparison table, this one included. The cost column depends on your workload, not on the vendor's rate card. A retrieval-heavy agent with a high cache hit rate and a long system prompt gets most of its bill from cache reads, where the multiplier matters more than the headline input price. A code-generation loop at max effort gets its bill from output and thinking tokens, where a lower list price can be swamped by a longer reasoning trace.
So the honest way to use the tables below is: read the rate card as a floor, read the effort setting as the multiplier, and measure your own traffic before you commit a number to a budget.
What OpenAI actually shipped on September 22
GPT-6 Sol and GPT-6 Luna, and neither is the flagship. OpenAI's own sentence is that "GPT-6 Astra continues to be our best model across the board." Sol is the mid tier, built for coding and agentic workflows. Luna is the high-volume tier. Astra, launched on September 3 at $10 and $50, was not repriced at all.
The "50% cheaper" headline needs unpacking too. OpenAI's wording is that it is "reducing API prices for Sol and Luna by 50% compared with their GPT-5.6 promotional pricing." So the comparison is against the previous generation of the same tier, at its promotional rate, not against anything at the frontier. On Sol the arithmetic checks out: $4 to $2 and $20 to $10. On Luna it does not. Output went from $1.20 to $0.50, which is a 58.3% cut, not 50%. The vendor has understated its own discount.
The more interesting number in that launch is one OpenAI published against itself. On AutomationBench 1.0.6, GPT-6 Sol at xhigh scores 33.2% at $0.27 per task, against GPT-6 Astra at low effort on 30.3% for 3.9 times the cost. A mid-tier model beating the flagship on a real task at a quarter of the price is a more useful fact than any of the marketing, and it points the same way as the ARC-AGI result above.
Two things not to repeat from that launch page. GPT-6 Sol loses DeepSWE v1.1 to Claude Fable 5, 68.8% against 69.9%, and its OSWorld 2.0 win over Claude Opus 5 is measured against Opus 5 at a lower effort setting. OpenAI's own footnote concedes that competitor scores were lifted from public reports at whatever version was available, which is why the comparison set changes row to row.
Where Google actually is
Behind, and the shape of the gap is unusual. Gemini 3.8 Flash, released September 2, is Google's newest and by its own description most intelligent model. It is a Flash-tier model. There is no Gemini 3.8 Pro: the Pro line is still at 3.1, which has been in preview since February. There is no Gemini Ultra model at all, despite how often the name appears in comparisons; Ultra is a subscription tier.
On the Artificial Analysis index, Gemini 3.8 Flash at high effort sits at 41 against Opus 5.5's 58, and Google currently has no Pro-tier entry on that index whatsoever. What Google does have is price: $0.75 and $3.75 per million, with cache reads at $0.075. Note the expiry. Those are introductory prices through December 31, 2026, and they double on January 1, 2027. If you are modelling a 2027 budget on today's Gemini rate card, you are modelling half the real number.
Google's own benchmark table also carries a caveat worth reading before you cite it: the non-Gemini results are the vendors' self-reported figures at maximum thinking settings, while Gemini's are computed by Google in its own harness. One row, a DeepSWE score of 74.0% for Claude Opus 5, was a rounding error Google retracted in its methodology document and never corrected on the model card itself.
The rest of the field
Four models belong in a serious comparison and two do not.
xAI Grok 4.7, released September 21, is the strongest non-big-three Western model at index 46 and $2 and $6 per million, with the caveat that the price doubles at 200K tokens and above and the context window stops at 500K rather than a million.
Meta's Muse Spark 1.3 is the surprise. At index 48 and $1.25 and $4.25 it is the best model outside the big three, and it is closed-weight. Meta's current model index lists no Llama model at all. The open-weights story that defined Meta's AI strategy has quietly ended, and any "Llama 5" you read about is unconfirmed.
Moonshot's Kimi K3 remains the best open-weight model on the index at 44, and it is also the most expensive model in that group at $3 and $15, which is an awkward position. We covered it in detail when it launched in our Kimi K3 against Fable 5 comparison.
DeepSeek V4.1 Flash is the price story. Index 39 at $0.30 and $1.20, halving outside peak UTC hours, and on Arena's agent board it ranks third on confirmed task success alone, ahead of Opus 5. It does not belong at the top tier on capability and it belongs in every conversation about cost.
Qwen3.8-Max matches Grok on both price and index and does not double above 200K, which makes it the better deal of the two on long inputs. Alibaba previewed Qwen 4 on September 22 with no published specifications, so anything you read about its performance is invention.
Mistral Medium 3.5 has fallen off the frontier. Index 14 against 58 is not a gap you close with a price argument. It remains interesting as an EU-sovereignty option and it is not a frontier coding model. Our earlier cross-vendor comparison from the Grok 4.5 generation shows how fast this part of the table turns over.
What the cost spread actually looks like
Artificial Analysis publishes the total cost of running its full evaluation suite, which is the closest public proxy for "what would this model cost me to do a fixed amount of real work". The spread is roughly 30x from cheapest to most expensive.
One honest gap: that chart's default selection excludes Claude Opus 5.5, Fable 5.1 and GPT-6 Astra, so there is no suite total for the three models at the top of the index. For those, the per-task figures are the only public comparison, and they tell the same story in miniature. Opus 5.5 spans index 58 down to 51 as cost per task falls from $5.98 to $1.34. Losing 12% of the index score cuts the bill by 78%.
What to run for what
The verdict, with the reasoning attached rather than a star rating.
If you are choosing a tool rather than a model, the model is only part of the answer: our comparison of Codex, Claude Code and Cursor covers the harness, and there is a walkthrough of running an OpenAI model inside Claude Code if you want to mix the two. For what changed in the Anthropic line specifically, see our breakdown of the Opus 5.5 launch and its pricing.
The only benchmark that matters is yours
Six frontier launches in three weeks, two dead leaderboards still being quoted, and an effort setting that can make a model worse for 4.5 times the price. No published table can tell you which model wins on your tasks, your data and your cache hit rate. Only a run against your own workload can.
Valletta Software Development builds that capability into your stack: a provider-agnostic routing layer, an eval suite that runs in CI against your real tasks, cost and latency telemetry per route, and a migration path that does not touch product code. The next launch is about six weeks away. We make it a config change instead of a quarter.
FAQ
What is the best LLM for coding right now?
Claude Opus 5.5 on the balance of public evidence. It leads the independent Artificial Analysis Intelligence Index at 58, ahead of GPT-6 Astra and Claude Fable 5.1 at 53, and it costs less per task than either. The exceptions are real: GPT-6 Astra leads Arena's WebDev board outright and tops Terminal-Bench 4.0, where Opus 5.5 has not yet been independently measured.
Is GPT-6 Sol better than Claude Opus 5.5?
No, and OpenAI does not claim it is. GPT-6 Sol is OpenAI's mid tier, not its flagship; Astra remains that. Sol sits at 48 on the Artificial Analysis index against Opus 5.5's 58. Sol's real argument is cost per task, where it beat OpenAI's own flagship on AutomationBench at roughly a quarter of the price.
Which AI model is cheapest for coding?
GPT-6 Luna at $0.10 per million input tokens and $0.50 per million output, followed by DeepSeek V4.1 Flash at $0.30 and $1.20, which halves outside peak UTC hours. Running a fixed evaluation suite costs about $122 on Luna against roughly $3,658 on Kimi K3, a 30x spread. Neither cheap model is frontier quality, and for classification, extraction and routing that does not matter.
Why should I not use SWE-bench to compare models?
Because the SWE-bench Verified leaderboard has accepted no submission since February 26, 2026. No GPT-6 model, no Claude Opus 5.x, no Fable and no Gemini 3.8 appears on it. The "around 79%" figure widely quoted as state of the art is an unverified third-party submission from December 2025 that the maintainers never checked; their own verified ceiling is 74.4% on much older models. Aider's polyglot benchmark is in the same condition, unmaintained since November 2025.
Does a higher thinking effort setting always give better answers?
No, and this is the most useful thing on this page. On ARC-AGI-2, Claude Opus 5.5 at high effort scores 93.3% at $0.41 per task while the same model at max effort scores 91.7% at $1.85. That is 4.5 times the cost for a worse result, and GPT-6 Astra shows the same non-monotonic pattern. Test your own workload at two or three effort levels before assuming more is better.
Is there a Gemini 3.8 Pro or a Gemini Ultra?
Neither exists. Gemini 3.8 Flash, released September 2, 2026, is Google's newest and most capable model, and it is a Flash-tier one. The Pro line is still at Gemini 3.1, in preview since February. "Gemini Ultra" is a subscription tier, not a model, despite appearing in many comparison articles as one.
How much cheaper are GPT-6 Sol and Luna really?
Both are half the price of the GPT-5.6 model in the same tier, at that model's promotional rate, which is not the same as being cheaper than anything at the frontier. GPT-6 Astra was not repriced and still costs $10 and $50 per million. On Luna, OpenAI's own "50%" understates the cut: output fell from $1.20 to $0.50, which is 58.3%.
Which models have a one million token context window?
Almost all of them. Claude Opus 5.5, Fable 5.1 and Sonnet 5 are at 1M; the GPT-6 family is at 1.05M; Gemini 3.8 Flash, Muse Spark 1.3, Kimi K3, Qwen3.8-Max and DeepSeek V4.1 Flash are at or around 1M. The exceptions are Grok 4.7 at 500K and Mistral Medium 3.5 at 256K. Watch the surcharges instead: GPT-6 charges 2x input above 272K tokens and Grok doubles at 200K.
How often does this page change?
It was first published on July 27, 2026 and has been rewritten as the lineup turned over. Anthropic, OpenAI and Google have each shipped a frontier model roughly every six weeks through 2026, so treat any figure here as accurate to the date at the top and check the vendor's own pricing page before you commit a budget to it.