Gemini 4 Argon Is Here: 1M Output Tokens, $2/$10, and a Catch

Gemini 4 Argon is here: Google's new frontier model with a 1M token output limit, against Claude and OpenAI

Gemini 4 is here, and it is called Gemini 4 Argon: Google's new frontier model, announced on September 30, 2026, can write up to 1 million tokens in a single response, scores 77.9% on the DeepSWE v1.1 long-horizon coding benchmark, and launches at $2 per million input tokens and $10 per million output. The catch is in the last line of the announcement: almost nobody can use it yet.

Key takeaways

  • The headline spec is output, not context. Google raised the output limit to 1M tokens, up from 64K on its previous models. Google has not published the input context window, so ignore the "2M context" figure circulating on X.
  • By our count of Google's own chart, Argon posts the top score on 13 of 19 benchmark rows, ties on one and loses on five. Claude Opus 5.5 still wins Terminal-Bench 4.0 by nine points, 66.4% to 57.4%.
  • $2 and $10 is an introductory price. Google's footnote says it rises to $4 and $20 afterwards, which is exactly what Claude Opus 5.5 costs today. No end date for the introductory period has been published.
  • Access starts with vetted cyber defenders in Google's Fairwind Program, then paid API customers and Google AI Ultra subscribers. There is no public model ID, no date for general availability, and no Argon entry on the Gemini API pricing page.
  • Two of the viral numbers need context. The 2.7x speedup applies to one video decoder compared with its Rust port, and the 300 TiB of freed memory is a figure Google quotes for "once rolled out".
  • Google's chart and Anthropic's chart report different scores for the same Claude models on the same benchmarks. Vendor tables are marketing with footnotes. Read the footnotes.

Published October 1, 2026. Every number below comes from Google's announcement, Introducing Gemini 4 Argon, and the evaluation methodology Google DeepMind published alongside it. Where the X trend summary or secondary coverage says something different, we follow the primary source and say so.

What Google actually announced

Google describes Gemini 4 Argon as a model that "delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense." Before launch, Argon was the codename that leaked on X. Google kept it as the product name. It is the first model in the Gemini 4 generation, and it is not Gemma 4, which is Google's separate family of open-weight models.

The launch was coordinated across Google, Google DeepMind and Google AI, with Sundar Pichai posting that there was "lots of discussion out there about our next model(!)" and that he wanted to give an early look as soon as possible. The trend reached about 62,000 posts on X within ten hours. The best reply was a pun: the name was chosen "because after 4 prompts all your tokens Argon."

The joke has a point. Artificial Analysis, which benchmarks models independently, already lists Argon as one of the more verbose models it has tested. Combine a verbose model with a 1M output ceiling and the bill becomes the first engineering question, ahead of benchmarks.

The 1M token output limit, explained

Most model launches compete on context window, meaning how much a model can read at once. Argon's headline number is how much it can write. In Google's words, "we are significantly expanding the model's output token limit to an industry-leading 1M tokens, up from the previous 64K tokens."

A million output tokens is roughly 750,000 words of English, or tens of thousands of lines of code, in one uninterrupted response. That matters for a specific class of work:

  • Large code migrations. Rewriting a library from C++ to Rust used to mean chunking the job into dozens of calls and stitching the results back together. Google says Argon agents are working on migrations from "tens of thousands of lines in core libraries" up to "800K+ lines for the Fuchsia Zircon kernel."
  • Long-horizon agent runs. Agents that plan, write, test and fix for hours generate far more tokens than chat does. A higher output ceiling means fewer forced restarts mid-task.
  • Full documents in one pass. Contract sets, financial models and research reports that previously broke into fragments can come back whole.

Now the cost. At the introductory price, a response that uses the full 1M output budget costs $10. At the post-introductory price it costs $20. That is per response. An agent loop that runs a few of those a day is a real line item, and the verbosity Artificial Analysis measured makes it more likely, not less.

Gemini 4 Argon benchmarks: the full table

Google compared Argon against GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5. It did not include any Gemini 3.x model. The best score in each row is highlighted. The rows Argon loses are kept in, because Google kept them in too.

Benchmark Gemini 4 Argon GPT-6 Astra Claude Fable 5.1 Claude Opus 5.5
Knowledge work
Vals Index 68.9% 63.1% 65.8% 67.0%
AutomationBench 51.3% 41.4% 31.4% 42.5%
Vals Finance Agent v2 65.4% 53.5% 58.9% 58.6%
Harvey Legal Agent Benchmark 19.6% 5.4% 6.7% 3.8%
Agentic coding
DeepSWE v1.1 77.9% 74.1% 67.4% 74.2%
FrontierSWE v2 55.0% 65.5% 56.3% 62.3%
Vibe Code Bench 91.9% 89.6% 90.3% 90.3%
Terminal-Bench 4.0 57.4% 58.2% 57.9% 66.4%
ML engineering, science and math
PostTrainBench 45.3% 44.3% 40.2% 49.3%
Terminal-Bench Science 0.1 57.6% 68.1% 52.6% 63.3%
LABBench 2 88.8% 85.4% 68.6% 73.1%
RiemannBench 76.0% 72.0% 65.6% 69.6%
Long context
GraphWalks BFS, up to 128K 99.7% 98.7% 91.4% 90.6%
GraphWalks BFS, 256K to 1M 84.2% 71.8% 65.0% 66.8%
Computer use
Agent's Last Exam 39.5% 34.2% n/a 38.2%
OSWorld 2.0 (offline subset) 69.2% 72.6% n/a n/a
Multimodal
Chartography 71.6% 71.0% 46.2% 66.3%
LVBench (video) 91.7% 87.5% 79.7% 83.7%
Cybersecurity
CWE-bench v1 68.0% 68.0% 58.0% 67.0%

Source: Google's launch chart and evaluation methodology, September 30, 2026. Argon run via the Gemini API at the highest thinking setting, pass@1. Rival scores are mostly self-reported by their vendors or taken from public leaderboards. CWE-bench is a tie that Google breaks on pass@4. Fable 5.1 has no Agent's Last Exam score, and Anthropic models are excluded from the OSWorld 2.0 offline subset.

Three things stand out once you read the table row by row instead of skimming the blue cells.

Where Argon wins clearly

Knowledge work is the real story. On Harvey's Legal Agent Benchmark, Argon scores 19.6% against 6.7% for the best Claude model and 5.4% for GPT-6 Astra. On AutomationBench it leads by almost nine points. Long context is the other clear lead: 84.2% on GraphWalks from 256K to 1M tokens, where every rival sits between 65% and 72%. If your workload is legal review, financial analysis or reasoning over very long documents, this is the model to test first once you can get access.

Where Claude and GPT still lead

On agentic coding the picture is mixed, and coding is where most of our readers spend their tokens. Argon takes DeepSWE v1.1 and Vibe Code Bench. Claude Opus 5.5 takes Terminal-Bench 4.0 by nine points and PostTrainBench by four. GPT-6 Astra takes FrontierSWE v2 by 10.5 points. If your team lives in a terminal agent such as Claude Code or Codex, the Terminal-Bench gap is the number to watch, not DeepSWE. We break down how those tools differ in Codex vs Claude Code vs Cursor.

The 77.9% DeepSWE score, in context

DeepSWE v1.1 measures long, multi-step software engineering tasks. Argon's 77.9% is 3.7 points ahead of Opus 5.5 and 10.5 ahead of Fable 5.1. Google's footnote says it computed Argon's score itself with a mini-swe agent harness at the highest thinking setting, while the competitor numbers come from a public leaderboard and from Anthropic's system cards. That is normal practice, but it means the comparison is not a single controlled experiment.

The numbers do not match Anthropic's numbers

This is the part almost nobody has noticed yet. Google and Anthropic both published scores for Claude Opus 5.5 on some of the same benchmarks, eight days apart. They disagree.

Benchmark and model Google's chart (Sept 30) Anthropic's chart (Sept 22) Gap
AutomationBench, Opus 5.5 42.5% 40.0% 2.5 pts
Terminal-Bench Science, Opus 5.5 63.3% 58.7% 4.6 pts
Terminal-Bench Science, GPT-6 Astra 68.1% 64.6% 3.5 pts
Terminal-Bench 4.0, Fable 5.1 57.9% 55.8% 2.1 pts
Terminal-Bench 4.0, GPT-6 Astra 58.2% 57.9% 0.3 pts

Google's numbers are from the Gemini 4 Argon launch chart. Anthropic's are from the Claude Opus 5.5 launch page. Same models, same benchmark names, different results.

Neither company is necessarily wrong. Google's methodology says that for rival models it reports "maximum thinking/reasoning settings available, but when reported results are not available we use best available reasoning results." Anthropic ran Terminal-Bench 4.0 for Opus 5.5 at xhigh effort, while the model ships defaulting to medium, as we covered in our Claude Opus 5.5 launch breakdown. Different harnesses, effort levels, timeouts and sample sizes all move scores by a few points. The video benchmark is the starkest example: Argon was tested at one frame per second, while the rivals got 300 to 800 frames "due to API limitations."

The practical rule is the one we apply to our own best LLM for coding leaderboard: treat any gap under about three points as a tie, prefer independent boards over vendor charts, and run your own tasks before you switch.

Gemini 4 Argon pricing vs Claude and GPT

Google's exact wording: "Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price." A footnote adds: "After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply."

Model Input / 1M Output / 1M Cache read / 1M Status
Gemini 4 Argon (intro) $2 $10 $0.10 Fairwind Program only
Gemini 4 Argon (after intro) $4 $20 $0.20 No date published
Claude Opus 5.5 $4 $20 $0.20 Generally available
Claude Fable 5.1 $10 $50 $0.25 Generally available
GPT-6 Astra $10 $50 $1.00 Generally available
GPT-6 Sol $2 $10 $0.20 Generally available
Gemini 3.8 Flash (intro) $0.75 $3.75 $0.075 Doubles Jan 1, 2027

Argon cache prices are our arithmetic from Google's "95% off input" wording. Other prices from each vendor's rate card, as tracked in our model comparison.

Read that table carefully and the pricing strategy becomes obvious. The introductory price matches OpenAI's mid-tier GPT-6 Sol. The long-term price matches Claude Opus 5.5 to the cent, including the $0.20 cache read. Google is pricing its flagship at a fifth of Fable 5.1 and GPT-6 Astra, and betting that the benchmark table will do the rest.

Two cautions before you model a budget on it. First, the introductory period has no published end date, and Google's last discounted model, Gemini 3.8 Flash, doubles in price on January 1, 2027. Second, list price is not cost per task. A model that writes more tokens to reach the same answer can be cheaper per token and more expensive per job. That was the whole story of the ChatGPT Pro Max repricing, and it applies here too.

What Argon did inside Google

Google's most persuasive evidence is not a benchmark. It is a list of jobs Argon agents have already done on Google's own infrastructure. Two of those claims went viral in shortened form, so here is the precise wording.

  • 300 TiB of memory. A team of Argon agents analysed fleet-wide profiling telemetry and applied memory optimisations across Google's data centers, "freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings." That is a projection tied to a rollout, not a completed result.
  • 2.7x faster. This is one project. Argon agents took an existing Rust port of the libgav1 video decoder and replaced 32K lines of SIMD code, producing "a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output." It is 2.7x faster than the Rust port, still behind the optimised C++ original, and it is not a general speedup figure.
  • Quantum circuits. On a qubit and gate-count optimisation task, Argon beat the published baseline by 40% "in a matter of minutes."

Google says the migrated code is "undergoing rigorous automated and manual auditing, emulation testing, and review before rolling out to production." That sentence is worth copying into your own process for any AI-written port.

Why cyber defenders get it first

Argon is rolling out first "to a set of trusted cyber defenders through our Fairwind Program." Those defenders, and Google's internal teams, get a version "without cyber guardrails." Everyone else will get a version with them. Google says it is also taking part in the U.S. government's voluntary process for pre-release model access.

The reasoning is the same one Anthropic and OpenAI have been dealing with all year. A model that is very good at finding vulnerabilities is very good at exploiting them, and the last few months showed how quickly that becomes real: three researchers used Claude Opus 5 to chain an image decoder bug into OpenAI's private monorepo in 72 hours, which we reconstructed step by step in OpenAI hacked in 72 hours. Giving defenders a head start is an attempt to tilt that race.

Argon ties GPT-6 Astra for first on CWE-bench v1 at 68%. Google also says Wiz, through its Scan for Good initiative, used Argon to find "a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide" that previous frontier models had missed. Help Net Security notes that the vulnerability results come from internal testing without independent verification, which is fair, and true of every lab's security claims so far.

Who can use Gemini 4 Argon, and when

As of October 1, 2026:

  • Now: cyber defenders and trusted testers in the Fairwind Program, plus Google's internal teams.
  • Next: "paid API customers and Google AI Ultra subscribers," with no date given.
  • Later: "developers, enterprises, and consumers as soon as possible."

There is no published model ID, no Argon entry on the Gemini API pricing page or changelog, and no Vertex AI documentation yet. VentureBeat put it accurately: Google is "retaking benchmark lead over OpenAI and Anthropic, but in limited release." TNW, citing Bloomberg, reports that some Google employees question how well the coding results hold up in real use. Until paid API access opens, nobody outside Google can check.

What to do this week

  1. Do not migrate anything yet. You cannot, and the price after the introductory period is the same as Opus 5.5, which you can use today.
  2. Write your evaluation set now. Pick twenty real tasks from your backlog, with known good answers, so you can run Argon against your current model on the day access opens. Our model comparison lists what to measure besides accuracy.
  3. Keep your stack model-agnostic. If switching models means rewriting prompts, tools and parsers, every launch like this is a project instead of a config change. A router layer, like the one in our Claude Code Router guide, keeps your options open.
  4. If you do legal, finance or long-document work, get on the waitlist through your Google Cloud account team. The Harvey and GraphWalks gaps are large enough to be worth testing early.

Pick the right model on your own workload, not on a vendor chart

Three frontier launches in a little over a week, and each vendor's table says it won. The only benchmark that matters is your own backlog: your codebase, your documents, your latency budget and your bill at the end of the month.

Valletta Software Development builds production AI systems on Claude, GPT and Gemini, and designs them so the model can be swapped without a rewrite. We will build an evaluation harness on your real tasks, run the current frontier models through it, and tell you which one earns its price. When Argon opens up, you will have the answer in a day, not a quarter. You can also hire AI developers who already ship on all three APIs.

Talk to an engineer

FAQ

What is Gemini 4 Argon?

Gemini 4 Argon is Google's newest frontier AI model and the first of the Gemini 4 generation, announced on September 30, 2026. It targets software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense, and it has a 1 million token output limit.

When is the Gemini 4 release date?

Google announced Gemini 4 Argon on September 30, 2026, and began rolling it out the same day to cyber defenders and trusted testers in its Fairwind Program. Paid API customers and Google AI Ultra subscribers are next, but Google has not published a date for general availability.

How much does Gemini 4 Argon cost?

The introductory price is $2 per million input tokens and $10 per million output tokens, with cached input at 95% off. After the introductory period ends, the price rises to $4 per million input and $20 per million output. Google has not said when the introductory period ends.

Is Gemini 4 Argon better than Claude Opus 5.5?

On Google's own chart, Argon beats Claude Opus 5.5 on 14 of the 18 benchmarks where both have a score, with the largest leads in legal, finance and long-context tasks. Opus 5.5 wins Terminal-Bench 4.0, FrontierSWE v2, PostTrainBench and Terminal-Bench Science. Opus 5.5 is available today at the same price Argon will cost after its introductory period.

What is the Gemini 4 Argon context window?

Google has not published the input context window. The 1 million token figure in the announcement is the output limit, up from 64K tokens on earlier Gemini models. Figures of 2 million tokens that circulate online have no primary source.

What is the Fairwind Program?

Fairwind is Google's early-access program for trusted cyber defenders. Members get Gemini 4 Argon first, in a version without the cyber guardrails that general users will get, so defenders can find and patch vulnerabilities before the model is widely available.

Is Gemini 4 the same as Gemma 4?

No. Gemini 4 Argon is Google's closed frontier model, available through Google's API and apps. Gemma 4 is Google's separate family of smaller open-weight models that developers can download and run themselves.

Can I use Gemini 4 Argon in the Gemini API today?

Not unless you are in the Fairwind Program. As of October 1, 2026, Argon has no public model ID and does not appear on the Gemini API pricing page or changelog. Google says paid API customers will get access first when it expands.

Need a senior engineering team behind your product?

Valletta Software delivers custom development and staff augmentation for companies across the EU and US. Book a free 30 minute call to discuss your project.

Valletta.Software - Top-Rated Agency on 50Pros

Talk to the engineers behind this blog

Valletta Software builds and staffs dedicated development teams for companies across the EU and US. Tell us about your project and get a reply within one business day.