Skip to content
News

Google is so back! Gemini 4 Argon is its first top-tier AI model since February

For the first time since February, Google has a new top-tier AI model. It is called Gemini 4 Argon. Google calls it a frontier model: the most capable system an AI company has, the one it puts up against its rivals’ best.

Google’s last model above its cheaper Flash line was Gemini 3.1 Pro Preview, released on 19th February. In May, Google said it was hard at work on Gemini 3.5 Pro and would roll it out the following month. That model never came. Google kept shipping Flash models instead, while its rivals moved.

Earlier this month we covered Anthropic’s Claude Fable 5.1 and, two days later, OpenAI’s GPT-6 Astra. Anthropic then released Claude Opus 5.5 on 22nd September.

So does Argon put Google back on top? And when can anyone outside Google use it?

How Argon compares

Google’s benchmark table compares Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 on 19 tests. A benchmark is a fixed set of tasks that every model attempts, so their scores can be compared. Argon has the top score on 13, ties on one and loses on five. Its biggest lead is on a legal test built by Harvey, a legal AI company, where Argon passes 19.6% of tasks against 6.7% for the next model in the table. A task only counts there if every requirement is met, which is why the scores look low. Among its losses, GPT-6 Astra beats it by 10.5 points on FrontierSWE v2, a set of very long software engineering tasks, and Opus 5.5 beats it by 9 points on Terminal-bench 4.0, a test of multi-step work in a command line.

TestWhat it measuresGoogle’s tablePublic leaderboard
Vals IndexFinance, coding, legal and tax work, weighted by each sector’s share of the US economyArgon first, 68.9%Argon first, 68.9%
Harvey’s Legal Agent BenchmarkLegal work produced from a set of filesArgon first, 19.6%Argon fifth; four Meta Muse Spark models score higher
Vibe Code BenchBuilding a working web app from a written briefArgon first, 91.9%Argon second; Claude Sonnet 5.5 leads with 92.4%
CWE-benchFinding and patching security flaws in real codeArgon tied first with GPT-6 Astra, 68%Three-way tie with xAI’s Grok 4.7, which wins the tie-break Google’s methodology names
Terminal-Bench ScienceResearch tasks written by scientistsArgon 57.6%, with six times the usual time limitArgon 44.3% when Vals AI ran it

Independent testers who had early access give a more even picture than Google’s table. Argon is first on the Vals Index, about two points clear of Claude Sonnet 5.5. The Artificial Analysis Intelligence Index, which combines ten tests the firm runs itself, gives Argon 53. That ties it with GPT-6 Astra and Claude Fable 5.1 and leaves it behind both of Anthropic’s newest models. Artificial Analysis says the result puts Google back among the top three AI labs.

ModelVals IndexArtificial Analysis Intelligence Index
Gemini 4 Argon68.90%53
Claude Sonnet 5.567.04%56
Claude Opus 5.566.97%58
Claude Fable 5.165.83%53
GPT-6 Astra63.13%53
GPT-6.1 Sol61.15%52

Artificial Analysis scores are for each model at its highest setting.

What’s new in Argon

The change Google leads with is output length. Argon can write up to 1 million tokens in one go, up from 64,000. One token is roughly three-quarters of an English word, so 1 million tokens is about 750,000 words, and that count includes the reasoning a model does before it answers. Every rival in the price table below stops at 128,000. So an Argon agent working through a long coding or research job has about eight times the room before it gets cut off mid-task.

Google says thousands of its own staff already use Argon. In one of its examples, Argon agents are moving more than 800,000 lines of code in the kernel of Google’s Fuchsia operating system from C and C++ to the Rust programming language.

Developers, who build the model into their own apps through Google’s API, get an introductory price that matches Claude Sonnet 5.5 and GPT-6.1 Sol. It then doubles to Opus 5.5’s price, and Google hasn’t said when that happens. On Artificial Analysis’s index, Argon cost USD 1.99, or about KES 260, per task at the introductory price, about 60% of GPT-6 Astra’s cost at its maximum setting.

ModelInput, per million tokensOutput, per million tokensLongest single response
Gemini 4 Argon, introductoryUSD 2 (KES 260)USD 10 (KES 1,300)1,000,000 tokens
Gemini 4 Argon, afterwardsUSD 4 (KES 520)USD 20 (KES 2,600)1,000,000 tokens
Claude Sonnet 5.5USD 2 (KES 260)USD 10 (KES 1,300)128,000 tokens
GPT-6.1 SolUSD 2 (KES 260)USD 10 (KES 1,300)128,000 tokens
Claude Opus 5.5USD 4 (KES 520)USD 20 (KES 2,600)128,000 tokens
GPT-6 AstraUSD 10 (KES 1,300)USD 50 (KES 6,500)128,000 tokens
Claude Fable 5.1USD 10 (KES 1,300)USD 50 (KES 6,500)128,000 tokens

Who can use it?

Almost nobody yet. Argon is going first to cyber defenders in Google’s Fairwind Program, which Google set up on 2nd September for security teams at governments, critical infrastructure and technology companies. The programme has more than 650 partners, among them CrowdStrike, Palo Alto Networks and Wiz, and only some of them get Argon. They get it without cyber guardrails, so it can find, check and patch software vulnerabilities for them. Google says it is also taking part in the US government’s voluntary process for testing models before release.

OpenAI and Anthropic have done the same. As we reported at the time, OpenAI gave security professionals first access to GPT-6 Astra through a programme called Daybreak, and Anthropic has limited Mythos 5.1, the version of Fable 5.1 with its cyber filters loosened, to a set of US organisations.

The next group in line is paying API customers and subscribers to Google AI Ultra, Google’s most expensive Gemini plan, “as soon as possible”. Google has given no date for any of that. It also hasn’t said how much text Argon can read at once, and it has not published a model card, the document where Google normally sets out a model’s safety test results.

Ordinary Gemini users got something else on the same day. Google started rolling out skills in the Gemini app worldwide, for users aged 18 and over. Skills are saved instructions you run by typing “/” and the skill’s name, which Gemini can also apply on its own when a prompt matches. Skills will replace Gems, the custom assistants Gemini has had until now. Google will move existing Gems over automatically and stop supporting Gems on personal accounts from November.

The Analyst

The Analyst delivers in-depth, data-driven insights on technology, industry trends, and digital innovation, breaking down complex topics for a clearer understanding. Reach out: Mail@Tech-ish.com

Join the discussion

0 comments
posting as Simba Jasiri

Anonymous by default — no sign-up or email needed. Prefer to be recognised? Add a name or email above, your call. We don't email you about replies, so do check back.

protected, no CAPTCHAs
Back to top button