Google has a new frontier model, and the most interesting thing about Gemini 4 Argon isn’t the model itself — it’s who gets to use it first.
On September 30, 2026, Google announced Gemini 4 Argon, which the company is positioning as its most capable model to date. The official Google blog carries the full announcement post.

Argon, Google writes in its announcement, “delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.” But instead of a general-availability launch, Google is staggering access: the first group outside Google gets it through a program called Fairwind, reserved for “a set of trusted cyber defenders.”
That’s a genuinely unusual launch policy for a flagship model, and it tells you a lot about where frontier AI is heading. Let’s break down what shipped, who gets it, and when you’ll be able to try it.
Google’s Flagship Returns: Gemini 4 Argon Ships September 30
First, the backstory, because it matters for how you read the launch.
Google first announced plans for Gemini 3.5 Pro at its I/O developer conference in May 2026, with a planned June launch. That launch never happened — instead, Google shipped a series of Flash models, and the flagship Pro-tier model went quiet. Bloomberg and Reuters reporting at the time attributed the cancellation to weak pre-launch results. What we can verify today, via The New Stack’s coverage: Argon is “essentially its replacement” — the flagship that 3.5 Pro was supposed to be.
So the comeback framing isn’t just spin; there’s a real gap between Google’s flagship ambitions in May and what actually shipped until now.
The announcement itself came from Koray Kavukcuoglu, Google’s chief AI architect and the SVP who took over leadership of Google DeepMind in August. He’s quoted in the announcement describing the model this way:
“Built to sustain deep reasoning across complex, long-horizon workflows, Argon is fundamentally changing the way we work and build at Google.”
That “long-horizon” framing shows up everywhere in the launch details. Argon’s output token limit jumps to an industry-leading 1 million tokens, up from 64K for previous Gemini models. Google’s argument: when a model has headroom to think and generate across hundreds of thousands of tokens in a single trajectory, it can solve hard problems in one pass instead of stitching together context across sessions.
Google also points to internal deployment as evidence the model is production-grade, not just a benchmark machine. A few of the announced examples, straight from Google’s own post on the Google blog:
- Quantum research: Argon helped Google’s quantum team optimize the spacetime resources (qubits × gates) of subroutines that bottleneck key applications — in one example beating the published baseline by 40%, “in a matter of minutes.”
- Data-center memory: A team of Argon agents analyzed fleet-wide profiling telemetry and autonomously applied memory optimizations across Google’s data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
- Codebase migration: Argon agents are migrating C/C++ codebases to Rust across Google, scaling from tens of thousands of lines in libraries like
re2andlibgav1up to 800K+ lines for the Fuchsia Zircon kernel. Forlibgav1, Argon’s agents replaced 32K lines of SIMD code with safe Rust — producing a memory-safe video decoder that runs 2.7x faster than the prior Rust port, with identical output.
Those are Google’s own claims about its own internal use, so treat them accordingly — but they’re concrete, specific, and checkable in a way most launch-day anecdotes aren’t.

Fairwind: Why Cyber Defenders Get Argon First
Here’s the genuinely novel part of this launch. Most frontier models go GA first and get safety-critiqued later. Google is doing the opposite with Argon.
What the Fairwind Program actually is
Fairwind isn’t a new thing invented for this launch — Google runs a Fairwind Program page describing it as a way to give “high-priority defenders” a head start. Per that page, the program targets defenders whose work protects critical systems:
“The Fairwind Program gives high-priority defenders (like governments, healthcare providers, and telecommunications services) early access to advanced models that help them build better defenses, before new threats arrive.”
The logic: frontier models that can find and patch vulnerabilities are powerful tools — for both sides. Give them to defenders first, before attackers can get their hands on equivalent capability. Fairwind also extends to Google’s own CodeMender agent — the automated vulnerability-patching system — per the announcement, making the defender-first principle cover tooling, not just model access.
Who qualifies, and under what rules
The program has real teeth, at least on paper. According to Google’s Fairwind page:
- Due diligence: Google runs background checks on applicant organizations to verify security history and their record of ethical operations.
- Managed access: Partners can’t share, redistribute, or sell access to the models.
- Hardened orgs: Participating organizations must use user-level authentication, phishing-resistant MFA, and applicable access controls — and may only grant Argon access to internal cybersecurity, incident response, or penetration testing teams, with employee access tracked.
- Scope limits: Partners are permitted only restricted dual-use tasks — authorized threat simulation, reverse engineering, and malware analysis for defensive and academic research purposes.
The guardrails-optional tier
Perhaps the most striking detail in the whole announcement, buried in Google’s cybersecurity section:
“For trusted defenders and our own internal teams at Google, we’ll be releasing Argon without cyber guardrails so they can leverage its full frontier-level cybersecurity defense capabilities.”
That’s a deliberate policy choice, not an oversight. The guardrails exist — Google says the model “is designed to refuse harmful requests” for cyber and CBRN misuse under its Frontier Safety Framework — but trusted defenders get a version without them, because defensive work (finding real vulnerabilities, validating exploits before patching) looks a lot like offensive work from the model’s perspective. Weakening those refusals would blunt the model’s usefulness for exactly the audience Fairwind serves.
Google says it’s “actively engaged in the U.S. government’s voluntary process for pre-release model access while we gradually expand access” — meaning the phased rollout isn’t just a marketing decision but folds into an existing government-led pre-release process.

What this looks like in practice already
Wiz — the cloud security company Google acquired for $32 billion in March — is already using Argon through its Scan for Good initiative, which finds and remediates high-risk exposures in critical public infrastructure for free. Google says that in an early demonstration, Argon “uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide,” identifying a risk that previous frontier models had missed.
What Gemini 4 Argon Claims to Be Good At (and Where It’s Mixed)
As with any launch-day benchmark slide, Google’s numbers are Google’s numbers. Here’s what the announcement claims, with context from independent coverage.
Knowledge work: the headline strength
This is where Google’s claims are strongest, and where third-party coverage agrees the model stands out:
| Benchmark | Argon’s score | Context |
|---|---|---|
| Vals Index | Leading | Measures economic impact across finance, coding, legal, and tax work, weighted by U.S. GDP contribution |
| Zapier AutomationBench | 51.3% (#1) | TechCrunch-adjacent coverage in The New Stack notes this is nearly 9 points ahead of the next model |
| Harvey’s Legal Agent Benchmark | 19.6% | Nearly triple the next-closest competitor per The New Stack — though that still means full completion on only about one in five tasks |
| LVBench (long video understanding) | 91.7% (state of the art) | Long-video comprehension; chart analysis is a separate benchmark |
| GraphWalks (256K–1M input tokens) | 84.2% | The New Stack notes this is more than 12 points ahead of the next model |
The Harvey number deserves the caveat The New Stack applies: tripling a competitor’s score sounds dramatic, but 19.6% is still roughly one-in-five full task completions. Legal agent workflows are hard, and nobody — including Argon — is closing them end to end reliably yet.
Coding: strong in places, honestly mixed elsewhere
Google leads with a genuine state-of-the-art claim: 77.9% on DeepSWE v1.1, which measures real-world long-horizon software engineering tasks.
But The New Stack’s read of Google’s own benchmark chart tells a more nuanced story. By its analysis, Argon “outright or tied” leads in 13 of 18 tests against Anthropic’s and OpenAI’s top models — but on two coding benchmarks, FrontierSWE v2 and Terminal-Bench 4.0, The New Stack reports Argon comes in last among the compared models, trailing the leaders by 10.5 and 9 points respectively. Its other coding highlight, Vibe Code Bench at 91.9%, is a win where all compared models score above 89%.
We’re not doing a full head-to-head breakdown here — that’s a separate post for another day — but the honest takeaway from the launch material is: Argon looks exceptional at knowledge work and long-horizon business tasks, strong-but-not-dominant at coding, with the coding picture depending heavily on which benchmark you weight most.
Cybersecurity: the reason for the rollout order
Google’s cyber claims center on autonomous defensive capability. Per the announcement, Argon “can autonomously find, validate, and patch critical software vulnerabilities.” Specific numbers Google cites:
- CWE-bench v1 (vulnerability remediation): ties for first place at 68%, building on Gemini 3.8 Flash Cyber’s v0 result. The New Stack adds a fairness caveat worth noting: the OpenAI and Anthropic models run in their own agent harnesses (Codex and Claude Code), so that leaderboard measures each model plus its tooling.
- Google’s internal vulnerability discovery benchmark: 85.8%, vs. 71.0% for 3.8 Flash Cyber — Google’s own internal benchmark, compared only against Google’s prior model.
- Wiz’s internal black-box penetration testing benchmark: 70.9% vs. 58.2% for 3.8 Flash Cyber, covering attack-surface discovery, vulnerability identification, and proof-of-concept generation.
Gemini 4 Argon Pricing and Availability Timeline
The pricing is, for a frontier flagship, aggressive. From Google’s announcement:
| Tier | Price |
|---|---|
| Input tokens (introductory) | $2 per million |
| Output tokens (introductory) | $10 per million |
| Cached input tokens | 95% off input token price |
The “introductory” qualifier matters. The New Stack notes pricing rises to $4 per million input and $20 per million output after the introductory period — the latter matching what Anthropic charges for its Opus-tier model.
No model card or developer docs page was reachable at the time of writing — worth watching if you’re evaluating the model for production use, since context windows, rate limits, and fine print on the “introductory” pricing period will matter as much as the headline rates.
When can you actually get it?
Google’s stated rollout sequence is:
- Now: A set of trusted cyber defenders via the Fairwind Program, plus Google’s own internal teams.
- Next: Paid API customers and Google AI Ultra subscribers — “as soon as possible,” per the announcement, once guardrail feedback from early testers is folded in.
- After that: Broader developer, enterprise, and consumer availability.
Google’s exact language: “We’ll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible.”
If you want in early, here’s your checklist
- Defense org? Apply via the Fairwind Program form. Expect background checks and restrictions on who inside your org can touch the model.
- Developers: Watch for paid API availability. If your use case leans on long-horizon agentic workflows — migrations, audits, multi-step research — the 1M output limit alone may justify being first in line.
- Budget holders: Model your costs at the $4/$20 post-introductory rates, not the $2/$10 teaser. Cached input at 95% off could make high-volume repeated-context workflows dramatically cheaper.
What It Signals for the Industry
Strip away the benchmark theater and two things stand out about the Gemini 4 Argon launch.
First, security-gated releases are becoming a real pattern, not a press-release flourish. Gemini 4 Argon’s phased rollout is folded into the U.S. government’s voluntary pre-release process, and Fairwind formalizes “defenders first” as distribution policy — down to specific terms like phishing-resistant MFA and background checks on applicant organizations. This launch also arrives in a notable context window: it came a day after OpenAI’s DevDay, and just a day after Google CEO Sundar Pichai co-signed a “self-police” commitment alongside other AI industry leaders following a meeting with President Trump, per The New Stack. The Verge also notes OpenAI has said it won’t release its planned GPT-6.1 Astra model due to safety worries — a caution shaped by the recurring misalignment incidents documented across the industry. Frontier labs are visibly experimenting with how to ship capability, not just what to ship — and as the synchronized outage that hit ChatGPT, Claude, and Grok showed, the whole ecosystem now hangs on a handful of providers getting things right.
OpenAI’s own Astra call-off is a case study in why OpenAI held back Astra over safety concerns.
Second, the “without guardrails for trusted defenders” tier is a fascinating experiment in itself. Google is betting that verified identity, contractual scope limits, and organizational hardening can substitute for model-level refusals for a narrow, audited audience. If that works — if no Fairwind partner leaks the unguarded model or the access itself — it becomes a template. If it doesn’t, expect the next flagship’s rollout to tighten further — what happened last time frontier AI cyber capability broke out is precisely the kind of incident that would force Google’s hand.
For now, the practical takeaway is simple: the most capable version of Gemini 4 Argon is, deliberately, not the one you can buy yet. Defense teams doing critical-infrastructure work have a real head start. Everyone else is waiting on APIs and AI Ultra — and watching how Google’s phased experiment plays out will tell us a lot about how the next model launches everywhere.
References and further reading
- Gemini 4 Argon announcement — Google DeepMind section of the official Google blog
- Google blog — official Google blog
- Google DeepMind — official Google DeepMind site
- The Fairwind Program — program details and eligibility rules
- Google and Wiz acquisition announcement — official Google blog
- The New Stack — independent launch coverage and benchmark analysis
- TechCrunch — tech industry news
- The Verge — coverage of OpenAI’s Astra decision
- Bloomberg — reporting on the cancelled Gemini 3.5 Pro launch
- Reuters — reporting on the cancelled Gemini 3.5 Pro launch
Please let us know if you enjoyed this blog post. Share it with others to spread the knowledge! If you believe any images in this post infringe your copyright, please contact us promptly so we can remove them.