
For the first time since February, Google has a new top-tier AI model. It is called Gemini 4 Argon. Google calls it a frontier model: the most capable system an AI company has, the one it puts up against its rivals’ best.
Google’s last model above its cheaper Flash line was Gemini 3.1 Pro Preview, released on 19th February. In May, Google said it was hard at work on Gemini 3.5 Pro and would roll it out the following month. That model never came. Google kept shipping Flash models instead, while its rivals moved.
Earlier this month we covered Anthropic’s Claude Fable 5.1 and, two days later, OpenAI’s GPT-6 Astra. Anthropic then released Claude Opus 5.5 on 22nd September.
So does Argon put Google back on top? And when can anyone outside Google use it?
How Argon compares
Google’s benchmark table compares Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 on 19 tests. A benchmark is a fixed set of tasks that every model attempts, so their scores can be compared. Argon has the top score on 13, ties on one and loses on five. Its biggest lead is on a legal test built by Harvey, a legal AI company, where Argon passes 19.6% of tasks against 6.7% for the next model in the table. A task only counts there if every requirement is met, which is why the scores look low. Among its losses, GPT-6 Astra beats it by 10.5 points on FrontierSWE v2, a set of very long software engineering tasks, and Opus 5.5 beats it by 9 points on Terminal-bench 4.0, a test of multi-step work in a command line.
| Test | What it measures | Google’s table | Public leaderboard |
|---|---|---|---|
| Vals Index | Finance, coding, legal and tax work, weighted by each sector’s share of the US economy | Argon first, 68.9% | Argon first, 68.9% |
| Harvey’s Legal Agent Benchmark | Legal work produced from a set of files | Argon first, 19.6% | Argon fifth; four Meta Muse Spark models score higher |
| Vibe Code Bench | Building a working web app from a written brief | Argon first, 91.9% | Argon second; Claude Sonnet 5.5 leads with 92.4% |
| CWE-bench | Finding and patching security flaws in real code | Argon tied first with GPT-6 Astra, 68% | Three-way tie with xAI’s Grok 4.7, which wins the tie-break Google’s methodology names |
| Terminal-Bench Science | Research tasks written by scientists | Argon 57.6%, with six times the usual time limit | Argon 44.3% when Vals AI ran it |
Independent testers who had early access give a more even picture than Google’s table. Argon is first on the Vals Index, about two points clear of Claude Sonnet 5.5. The Artificial Analysis Intelligence Index, which combines ten tests the firm runs itself, gives Argon 53. That ties it with GPT-6 Astra and Claude Fable 5.1 and leaves it behind both of Anthropic’s newest models. Artificial Analysis says the result puts Google back among the top three AI labs.
| Model | Vals Index | Artificial Analysis Intelligence Index |
|---|---|---|
| Gemini 4 Argon | 68.90% | 53 |
| Claude Sonnet 5.5 | 67.04% | 56 |
| Claude Opus 5.5 | 66.97% | 58 |
| Claude Fable 5.1 | 65.83% | 53 |
| GPT-6 Astra | 63.13% | 53 |
| GPT-6.1 Sol | 61.15% | 52 |
Artificial Analysis scores are for each model at its highest setting.
What’s new in Argon
The change Google leads with is output length. Argon can write up to 1 million tokens in one go, up from 64,000. One token is roughly three-quarters of an English word, so 1 million tokens is about 750,000 words, and that count includes the reasoning a model does before it answers. Every rival in the price table below stops at 128,000. So an Argon agent working through a long coding or research job has about eight times the room before it gets cut off mid-task.
Google says thousands of its own staff already use Argon. In one of its examples, Argon agents are moving more than 800,000 lines of code in the kernel of Google’s Fuchsia operating system from C and C++ to the Rust programming language.
Developers, who build the model into their own apps through Google’s API, get an introductory price that matches Claude Sonnet 5.5 and GPT-6.1 Sol. It then doubles to Opus 5.5’s price, and Google hasn’t said when that happens. On Artificial Analysis’s index, Argon cost USD 1.99, or about KES 260, per task at the introductory price, about 60% of GPT-6 Astra’s cost at its maximum setting.
| Model | Input, per million tokens | Output, per million tokens | Longest single response |
|---|---|---|---|
| Gemini 4 Argon, introductory | USD 2 (KES 260) | USD 10 (KES 1,300) | 1,000,000 tokens |
| Gemini 4 Argon, afterwards | USD 4 (KES 520) | USD 20 (KES 2,600) | 1,000,000 tokens |
| Claude Sonnet 5.5 | USD 2 (KES 260) | USD 10 (KES 1,300) | 128,000 tokens |
| GPT-6.1 Sol | USD 2 (KES 260) | USD 10 (KES 1,300) | 128,000 tokens |
| Claude Opus 5.5 | USD 4 (KES 520) | USD 20 (KES 2,600) | 128,000 tokens |
| GPT-6 Astra | USD 10 (KES 1,300) | USD 50 (KES 6,500) | 128,000 tokens |
| Claude Fable 5.1 | USD 10 (KES 1,300) | USD 50 (KES 6,500) | 128,000 tokens |
Who can use it?
Almost nobody yet. Argon is going first to cyber defenders in Google’s Fairwind Program, which Google set up on 2nd September for security teams at governments, critical infrastructure and technology companies. The programme has more than 650 partners, among them CrowdStrike, Palo Alto Networks and Wiz, and only some of them get Argon. They get it without cyber guardrails, so it can find, check and patch software vulnerabilities for them. Google says it is also taking part in the US government’s voluntary process for testing models before release.
OpenAI and Anthropic have done the same. As we reported at the time, OpenAI gave security professionals first access to GPT-6 Astra through a programme called Daybreak, and Anthropic has limited Mythos 5.1, the version of Fable 5.1 with its cyber filters loosened, to a set of US organisations.
The next group in line is paying API customers and subscribers to Google AI Ultra, Google’s most expensive Gemini plan, “as soon as possible”. Google has given no date for any of that. It also hasn’t said how much text Argon can read at once, and it has not published a model card, the document where Google normally sets out a model’s safety test results.
Ordinary Gemini users got something else on the same day. Google started rolling out skills in the Gemini app worldwide, for users aged 18 and over. Skills are saved instructions you run by typing “/” and the skill’s name, which Gemini can also apply on its own when a prompt matches. Skills will replace Gems, the custom assistants Gemini has had until now. Google will move existing Gems over automatically and stop supporting Gems on personal accounts from November.





Join the discussion