Best LLM for Coding: Opus 5.5 vs GPT-6 vs Gemini 3.8

Best LLM for coding September 2026: the Anthropic, OpenAI and Google Gemini marks on a repriced frontier

The best LLM for coding in September 2026 is Claude Opus 5.5 on most of the public evidence, and the honest version of that sentence is longer. Six frontier models shipped in three weeks, the cheapest of them costs 100 times less per token than the most expensive, and the single most quoted coding benchmark on the internet has not accepted a new submission since February. This page is the current state of the field, with every number's source and configuration attached.

Key takeaways

  • Claude Opus 5.5 leads the independent Artificial Analysis Intelligence Index at 58, ahead of GPT-6 Astra and Claude Fable 5.1 at 53. It is also the cheapest of that group to run per task. That combination is unusual and it is the main thing that changed this month.
  • More effort can make a model worse. On ARC-AGI-2, Opus 5.5 at high effort scores 93.3% at $0.41 per task. The same model at max effort scores 91.7% at $1.85. That is a lower score for 4.5 times the money, and the same non-monotonic pattern shows up on GPT-6 Astra.
  • Stop quoting SWE-bench Verified. The leaderboard has accepted no submission since February 26, 2026. No GPT-6 model, no Opus 5.x, no Fable, no Gemini 3.8 appears on it. The widely repeated "around 79%" figure is an unverified third-party submission from December 2025 that the SWE-bench team never checked. Aider's polyglot board is worse: unmaintained since November 2025.
  • GPT-6 Sol is not OpenAI's flagship and most of the coverage says otherwise. OpenAI's own words: "GPT-6 Astra continues to be our best model across the board." Sol is the mid tier and Luna is the volume tier.
  • The "50% price cut" on Sol and Luna is half of GPT-5.6's promotional price for the same tier, not half of anything at the frontier. Astra was not repriced. And OpenAI's own arithmetic on Luna is off: $1.20 to $0.50 of output is a 58.3% cut, not 50%.
  • Google has no current top-tier entry. Gemini 3.8 Flash is the newest and most capable Gemini, and it sits at 41 on the same index. There is no Gemini 3.8 Pro. Pro is still at 3.1, in preview since February. Flash's price doubles on January 1, 2027.
  • Cost to run a fixed evaluation suite spans roughly 30x across vendors, and per-task cost inside a single model spans 4x depending only on the effort setting. The rate card is the smallest part of what you pay.

Last updated September 23, 2026. This page is maintained rather than reposted: it was first published on July 27, 2026 comparing Claude Opus 5, Fable 5 and GPT-5.6 Sol, and has been rewritten for the current lineup. Prices and model IDs come from each vendor's own pricing and documentation pages. Benchmark figures are labelled as either vendor self-reported or independently measured, because in this field those are different kinds of claim. Where a vendor's own chart shows it losing, we keep the row.

The frontier, priced

Every model below is available through a public API today. Prices are per million tokens, in US dollars, taken from each vendor's own rate card: Anthropic, OpenAI and Google. Read the notes column; three of these prices are conditional.

Model In Out Cache read Context Note
Claude Opus 5.5 $4 $20 $0.20 1M Sept 22. Defaults to medium effort
Claude Fable 5.1 $10 $50 $0.25 1M Sept 1. Cache read is 0.025x input
GPT-6 Astra $10 $50 $1.00 1.05M Sept 3. OpenAI's flagship, not repriced
GPT-6 Sol $2 $10 $0.20 1.05M Sept 22. Over 272K input costs 2x in, 1.5x out
GPT-6 Luna $0.10 $0.50 $0.01 1.05M Sept 22. Same long-context surcharge
Gemini 3.8 Flash $0.75 $3.75 $0.075 1M Sept 2. Doubles to $1.50 and $7.50 on Jan 1, 2027
Gemini 3.1 Pro (preview) $2 $12 n/a 1M Still preview since Feb. Over 200K: $4 and $18
Grok 4.7 $2 $6 n/a 500K Sept 21. Doubles at 200K and above
Muse Spark 1.3 (Meta) $1.25 $4.25 $0.15 1M Sept 2. Closed weights. Meta lists no current Llama
Kimi K3 (Moonshot) $3 $15 $0.30 1M Jul 16. Best open-weight model on the index
Qwen3.8-Max (Alibaba) $2 $6 n/a 1M Beijing region is cheaper. Qwen 4 has no published specs
DeepSeek V4.1 Flash $0.30 $1.20 n/a 1M Sept 10. Halves off peak UTC hours
Mistral Medium 3.5 $1.50 $7.50 n/a 256K April. No longer competitive on capability

Two observations before the scores. The spread between GPT-6 Luna and the $50-output models is a factor of 100 on output tokens, which is larger than the spread in what they can do. And the context window race is over: almost everything here is at or near a million tokens, and nobody advertised a bigger number this month.

What the independent leaderboards actually say

Vendor charts are marketing artifacts with real numbers in them. The independent boards are the closest thing to a referee, and right now they disagree with each other in instructive ways. The four worth reading are Artificial Analysis, Terminal-Bench, ARC Prize and Arena.

Board Current leader What you must know before quoting it
Artificial Analysis Intelligence Index v4.3 Claude Opus 5.5, index 58 Opus 5.5 occupies four of the top eight slots at different effort levels. It is an index over many evals, not a coding benchmark
Terminal-Bench 4.0 GPT-6 Astra at max, 58.18% Last refreshed Sept 21, so Opus 5.5 is not on it yet. The top five overlap within their error bars
ARC-AGI-2 GPT-6 Astra at max, 95.0% Opus 5.5 at high ties Astra at xhigh on 93.3%, for half the cost per task. ARC-AGI-1 is saturated and ARC-AGI-3 is not cross-vendor comparable
Arena, formerly LMArena Anthropic on text and agents, GPT-6 Astra on WebDev The text top five is a statistical tie. WebDev is the one board where OpenAI leads outright. It moved to arena.ai
SWE-bench Verified Nothing current Frozen since Feb 26, 2026. No 2026 frontier model has been submitted. Do not use it
Aider polyglot Nothing current Unmaintained since Nov 2025. The newest model on it is base GPT-5

The two dead boards matter more than the live ones, because they are still being quoted. If you read a September 2026 comparison that ranks models by SWE-bench Verified, the author has not opened the leaderboard. The number most often attached to it, somewhere around 79%, comes from an unverified third-party agent submission made in December 2025 that the maintainers never checked. The team-checked ceiling is 74.4%, on models that are now two generations old.

The finding nobody is reporting: more effort, worse results

ARC Prize publishes every run at every effort level with its cost attached, which makes it the most useful board in the field right now, and it shows something that breaks the mental model most teams are using.

Claude Opus 5.5 on ARC-AGI-2

  • At high effort: 93.3% correct, $0.41 per task.
  • At max effort: 91.7% correct, $1.85 per task.

4.5 times the cost, 1.6 points worse. GPT-6 Astra shows the same shape, scoring the same at medium as at high. Effort is not a quality dial.

This is not a rounding artifact, and it is not unique to one model. It means the standard cost-optimisation advice has the sign backwards: teams routinely raise effort when output quality disappoints, and on at least some task families that spends more money to get a worse answer. The only way to know which family you are in is to run your own evaluation at two or three effort levels and look at the curve.

It also explains why the vendor charts and the independent boards keep disagreeing. Anthropic reports Opus 5.5 at 66.4% on Terminal-Bench 4.0, measured at xhigh effort. The public Terminal-Bench 4.0 leaderboard, last refreshed the day before Opus 5.5 shipped, tops out at 58.18%. Those are not contradictory claims, they are different harnesses and different configurations, and until an independent run lands the 66.4% remains a vendor self-report.

How to read a frontier benchmark in 2026

Every number on this page comes with a configuration attached, and the configuration is doing more work than the number. Four traps account for almost every misleading model comparison published this year, including some of ours.

The four traps

  1. The effort footnote. A score is published at one thinking-effort setting and the model ships on another. Anthropic's Opus 5.5 chart is run at max and xhigh effort; the model defaults to medium. The headline number is a ceiling you have to pay extra to reach, not the out-of-the-box result.
  2. The flattened curve. Several vendors publish score-versus-cost curves across effort levels, not single scores. Any table that collapses a curve into one cell has picked the winner by picking an effort setting. If a source cannot tell you which point on the curve it quoted, it does not know what it published.
  3. The borrowed score. When a lab skips a benchmark, write-ups fill the hole with the previous model's number under the new model's name. Anthropic published no GPQA, AIME, MMMLU, tau-bench or ARC-AGI figures for Opus 5.5. Every table you see carrying those rows has borrowed them.
  4. Price per token instead of cost per task. The two move independently. A model can be 20% cheaper per token and 40% cheaper per job, or cheaper per token and more expensive per job, depending on how many tokens it burns thinking. Only one of those numbers appears on your invoice.

The fourth trap is the one that changed this month, so it deserves its own section.

Price per token is no longer the price

For two years the frontier was easy to price: multiply tokens by a rate card. That broke once every lab shipped an adjustable thinking budget, because the same model at two effort settings is effectively two products with the same rate card and very different bills.

Claude Opus 5.5 is the cleanest illustration. Its list price is 20% below Opus 5. Anthropic separately states that on typical workloads at default settings it costs 40% less. Both are true, and the gap is made of three things that never appear on a pricing page: the default effort level dropped from high to medium, the cache read multiplier dropped from 0.1x of input to 0.05x, and the model is cheaper to serve. Pushed the other way, Anthropic also says that at a fixed effort level Opus 5.5 thinks more per turn than Opus 5 did, most of all at xhigh and max.

The practical consequence is uncomfortable for anyone building a comparison table, this one included. The cost column depends on your workload, not on the vendor's rate card. A retrieval-heavy agent with a high cache hit rate and a long system prompt gets most of its bill from cache reads, where the multiplier matters more than the headline input price. A code-generation loop at max effort gets its bill from output and thinking tokens, where a lower list price can be swamped by a longer reasoning trace.

So the honest way to use the tables below is: read the rate card as a floor, read the effort setting as the multiplier, and measure your own traffic before you commit a number to a budget.

What OpenAI actually shipped on September 22

GPT-6 Sol and GPT-6 Luna, and neither is the flagship. OpenAI's own sentence is that "GPT-6 Astra continues to be our best model across the board." Sol is the mid tier, built for coding and agentic workflows. Luna is the high-volume tier. Astra, launched on September 3 at $10 and $50, was not repriced at all.

The "50% cheaper" headline needs unpacking too. OpenAI's wording is that it is "reducing API prices for Sol and Luna by 50% compared with their GPT-5.6 promotional pricing." So the comparison is against the previous generation of the same tier, at its promotional rate, not against anything at the frontier. On Sol the arithmetic checks out: $4 to $2 and $20 to $10. On Luna it does not. Output went from $1.20 to $0.50, which is a 58.3% cut, not 50%. The vendor has understated its own discount.

The more interesting number in that launch is one OpenAI published against itself. On AutomationBench 1.0.6, GPT-6 Sol at xhigh scores 33.2% at $0.27 per task, against GPT-6 Astra at low effort on 30.3% for 3.9 times the cost. A mid-tier model beating the flagship on a real task at a quarter of the price is a more useful fact than any of the marketing, and it points the same way as the ARC-AGI result above.

Two things not to repeat from that launch page. GPT-6 Sol loses DeepSWE v1.1 to Claude Fable 5, 68.8% against 69.9%, and its OSWorld 2.0 win over Claude Opus 5 is measured against Opus 5 at a lower effort setting. OpenAI's own footnote concedes that competitor scores were lifted from public reports at whatever version was available, which is why the comparison set changes row to row.

Where Google actually is

Behind, and the shape of the gap is unusual. Gemini 3.8 Flash, released September 2, is Google's newest and by its own description most intelligent model. It is a Flash-tier model. There is no Gemini 3.8 Pro: the Pro line is still at 3.1, which has been in preview since February. There is no Gemini Ultra model at all, despite how often the name appears in comparisons; Ultra is a subscription tier.

On the Artificial Analysis index, Gemini 3.8 Flash at high effort sits at 41 against Opus 5.5's 58, and Google currently has no Pro-tier entry on that index whatsoever. What Google does have is price: $0.75 and $3.75 per million, with cache reads at $0.075. Note the expiry. Those are introductory prices through December 31, 2026, and they double on January 1, 2027. If you are modelling a 2027 budget on today's Gemini rate card, you are modelling half the real number.

Google's own benchmark table also carries a caveat worth reading before you cite it: the non-Gemini results are the vendors' self-reported figures at maximum thinking settings, while Gemini's are computed by Google in its own harness. One row, a DeepSWE score of 74.0% for Claude Opus 5, was a rounding error Google retracted in its methodology document and never corrected on the model card itself.

The rest of the field

Four models belong in a serious comparison and two do not.

xAI Grok 4.7, released September 21, is the strongest non-big-three Western model at index 46 and $2 and $6 per million, with the caveat that the price doubles at 200K tokens and above and the context window stops at 500K rather than a million.

Meta's Muse Spark 1.3 is the surprise. At index 48 and $1.25 and $4.25 it is the best model outside the big three, and it is closed-weight. Meta's current model index lists no Llama model at all. The open-weights story that defined Meta's AI strategy has quietly ended, and any "Llama 5" you read about is unconfirmed.

Moonshot's Kimi K3 remains the best open-weight model on the index at 44, and it is also the most expensive model in that group at $3 and $15, which is an awkward position. We covered it in detail when it launched in our Kimi K3 against Fable 5 comparison.

DeepSeek V4.1 Flash is the price story. Index 39 at $0.30 and $1.20, halving outside peak UTC hours, and on Arena's agent board it ranks third on confirmed task success alone, ahead of Opus 5. It does not belong at the top tier on capability and it belongs in every conversation about cost.

Qwen3.8-Max matches Grok on both price and index and does not double above 200K, which makes it the better deal of the two on long inputs. Alibaba previewed Qwen 4 on September 22 with no published specifications, so anything you read about its performance is invention.

Mistral Medium 3.5 has fallen off the frontier. Index 14 against 58 is not a gap you close with a price argument. It remains interesting as an EU-sovereignty option and it is not a frontier coding model. Our earlier cross-vendor comparison from the Grok 4.5 generation shows how fast this part of the table turns over.

What the cost spread actually looks like

Artificial Analysis publishes the total cost of running its full evaluation suite, which is the closest public proxy for "what would this model cost me to do a fixed amount of real work". The spread is roughly 30x from cheapest to most expensive.

Model (effort) Cost to run the full suite
Kimi K3 (max) $3,658
GPT-5.6 Sol (max) $3,465
Gemini 3.8 Flash (high) $1,623
GPT-6 Sol (max) $1,550
DeepSeek V4.1 Flash (max) $477
GPT-6 Luna (max) $122

One honest gap: that chart's default selection excludes Claude Opus 5.5, Fable 5.1 and GPT-6 Astra, so there is no suite total for the three models at the top of the index. For those, the per-task figures are the only public comparison, and they tell the same story in miniature. Opus 5.5 spans index 58 down to 51 as cost per task falls from $5.98 to $1.34. Losing 12% of the index score cuts the bill by 78%.

What to run for what

The verdict, with the reasoning attached rather than a star rating.

The job Run this Why
Long agentic coding sessions Claude Opus 5.5 at high, not max Top of the index, cheapest cache reads in its class, and high beat max on ARC-AGI-2
Front-end and web UI generation GPT-6 Astra The one board where OpenAI leads outright and not within the error bars: Arena WebDev
High-volume classification, extraction, routing GPT-6 Luna or DeepSeek V4.1 Flash $0.50 output against $20 to $50 at the top. Frontier quality is not what this work needs
Mid-tier agents where cost per task rules GPT-6 Sol Beat OpenAI's own flagship on AutomationBench at a quarter of the cost
Very long single-pass inputs Gemini 3.8 Flash No long-context surcharge, unlike GPT-6 above 272K and Grok above 200K. Budget for the 2027 doubling
Self-hosting or data residency Kimi K3 Best open-weight model available. Meta no longer ships one
Highest capability, cost no object Opus 5.5 at max, or Fable 5.1 Verify on your own tasks first. Max is not reliably better than high

If you are choosing a tool rather than a model, the model is only part of the answer: our comparison of Codex, Claude Code and Cursor covers the harness, and there is a walkthrough of running an OpenAI model inside Claude Code if you want to mix the two. For what changed in the Anthropic line specifically, see our breakdown of the Opus 5.5 launch and its pricing.

The only benchmark that matters is yours

Six frontier launches in three weeks, two dead leaderboards still being quoted, and an effort setting that can make a model worse for 4.5 times the price. No published table can tell you which model wins on your tasks, your data and your cache hit rate. Only a run against your own workload can.

Valletta Software Development builds that capability into your stack: a provider-agnostic routing layer, an eval suite that runs in CI against your real tasks, cost and latency telemetry per route, and a migration path that does not touch product code. The next launch is about six weeks away. We make it a config change instead of a quarter.

Book a scoping call

FAQ

What is the best LLM for coding right now?

Claude Opus 5.5 on the balance of public evidence. It leads the independent Artificial Analysis Intelligence Index at 58, ahead of GPT-6 Astra and Claude Fable 5.1 at 53, and it costs less per task than either. The exceptions are real: GPT-6 Astra leads Arena's WebDev board outright and tops Terminal-Bench 4.0, where Opus 5.5 has not yet been independently measured.

Is GPT-6 Sol better than Claude Opus 5.5?

No, and OpenAI does not claim it is. GPT-6 Sol is OpenAI's mid tier, not its flagship; Astra remains that. Sol sits at 48 on the Artificial Analysis index against Opus 5.5's 58. Sol's real argument is cost per task, where it beat OpenAI's own flagship on AutomationBench at roughly a quarter of the price.

Which AI model is cheapest for coding?

GPT-6 Luna at $0.10 per million input tokens and $0.50 per million output, followed by DeepSeek V4.1 Flash at $0.30 and $1.20, which halves outside peak UTC hours. Running a fixed evaluation suite costs about $122 on Luna against roughly $3,658 on Kimi K3, a 30x spread. Neither cheap model is frontier quality, and for classification, extraction and routing that does not matter.

Why should I not use SWE-bench to compare models?

Because the SWE-bench Verified leaderboard has accepted no submission since February 26, 2026. No GPT-6 model, no Claude Opus 5.x, no Fable and no Gemini 3.8 appears on it. The "around 79%" figure widely quoted as state of the art is an unverified third-party submission from December 2025 that the maintainers never checked; their own verified ceiling is 74.4% on much older models. Aider's polyglot benchmark is in the same condition, unmaintained since November 2025.

Does a higher thinking effort setting always give better answers?

No, and this is the most useful thing on this page. On ARC-AGI-2, Claude Opus 5.5 at high effort scores 93.3% at $0.41 per task while the same model at max effort scores 91.7% at $1.85. That is 4.5 times the cost for a worse result, and GPT-6 Astra shows the same non-monotonic pattern. Test your own workload at two or three effort levels before assuming more is better.

Is there a Gemini 3.8 Pro or a Gemini Ultra?

Neither exists. Gemini 3.8 Flash, released September 2, 2026, is Google's newest and most capable model, and it is a Flash-tier one. The Pro line is still at Gemini 3.1, in preview since February. "Gemini Ultra" is a subscription tier, not a model, despite appearing in many comparison articles as one.

How much cheaper are GPT-6 Sol and Luna really?

Both are half the price of the GPT-5.6 model in the same tier, at that model's promotional rate, which is not the same as being cheaper than anything at the frontier. GPT-6 Astra was not repriced and still costs $10 and $50 per million. On Luna, OpenAI's own "50%" understates the cut: output fell from $1.20 to $0.50, which is 58.3%.

Which models have a one million token context window?

Almost all of them. Claude Opus 5.5, Fable 5.1 and Sonnet 5 are at 1M; the GPT-6 family is at 1.05M; Gemini 3.8 Flash, Muse Spark 1.3, Kimi K3, Qwen3.8-Max and DeepSeek V4.1 Flash are at or around 1M. The exceptions are Grok 4.7 at 500K and Mistral Medium 3.5 at 256K. Watch the surcharges instead: GPT-6 charges 2x input above 272K tokens and Grok doubles at 200K.

How often does this page change?

It was first published on July 27, 2026 and has been rewritten as the lineup turned over. Anthropic, OpenAI and Google have each shipped a frontier model roughly every six weeks through 2026, so treat any figure here as accurate to the date at the top and check the vendor's own pricing page before you commit a budget to it.

Vibe coded an app that needs to become a product?

We audit AI generated codebases and bring them to production quality. Book a free 30 minute call and get a senior review of your project.

Valletta.Software - Top-Rated Agency on 50Pros

Talk to the engineers behind this blog

Valletta Software builds and staffs dedicated development teams for companies across the EU and US. Tell us about your project and get a reply within one business day.