Did Google Quietly Ship Gemini 4 Pro? The Leak vs. What Google Actually Said
Developers on the Arena testing platform hit a model labeled gemini-3.8-flash that felt far stronger than a Flash-tier model should — and the rumor mill turned that into "Google quietly shipped Gemini 4 Pro and crushed GPT-6 Astra and Claude Fable 5.1." Here is the detail most coverage skips: Gemini 3.8 Flash is a real model Google shipped on September 2, 2026. This post separates the rumored specs and floating benchmark table from the two things Google actually said on the record, includes a video recap, and gives you three five-minute checks for this kind of story.

The short version: as of September 22, 2026, Google has not released Gemini 4 Pro and has not confirmed it is testing one on Arena. What is actually circulating is something narrower — developers on Arena, the platform where two AI models answer the same prompt and people vote, ran into a model labeled gemini-3.8-flash that felt far stronger than a Flash-tier model should, and concluded Google might be stealth-testing its next flagship under an old name. The only thing Google has said on the record is one line from July 21: it has started its most ambitious pre-training run yet, for Gemini 4.
TL;DR
① Where it started: around September 17-18, developers on Arena drew a model labeled gemini-3.8-flash whose output looked well above expectations, and guessed it was an internal Gemini 4 Pro build (rumored codename Argon).
② The detail most coverage skips: Gemini 3.8 Flash is a real model Google shipped on September 2, 2026, at an introductory $0.75 per million input tokens and $3.75 per million output tokens (the promo ends at the close of 2026). "Same name, wildly different results" is the whole premise of the rumor.
③ Rumored specs: 10M-token input, 256K-token output, persistent cross-session memory, $2.25/$11.25 pricing — none of it sourced, and the price lands on a clean multiple of both official prices, which is what a calculation looks like.
④ Rumored benchmarks: DeepSWE v1.1 88.7%, Terminal-bench 2.1 95.3%, OSWorld-2.0 86.8%, said to beat GPT-6 Astra and Claude Fable 5.1 — but every one of those sits well above anything on Google's own published table. Even the original Chinese report hedged it: "if this table is accurate."
⑤ Google has said exactly two things: the July 21 blog line about starting Gemini 4 pre-training, and Pichai on the July 22 earnings call saying the next leap needs much larger base models, with coding and agentic coding as the gaps to close. No date, no specs, no scores.
I found three or four versions of this story, each headline more certain than the last: "quietly launched," "utterly outclasses," "Google is back on top." But when I traced every number, nearly all of them came from the same handful of user test impressions — not from any official document. So below I keep the rumor and the official record in separate columns, with a source on every claim.
1. How this started
Around September 17-18, 2026, developers using Arena (where two models answer the same prompt, sometimes under anonymous or codename labels, and people vote on the better answer) drew a model labeled gemini-3.8-flash and started posting what it built.
Notably, they posted artifacts rather than text answers: web design, SVG illustration, 3D modeling, animation, playable games. The collected demos include a graphite-style site with scroll-driven drawing effects, a pelican riding a bicycle, a voxel pagoda, an Airbus H145 model, a flight simulator and a 3D kart racer. More than one person posted; the most-shared writeup was Pankaj Kumar's thread on X, which said SVG and 3D generation had improved a lot, worked well in one shot, and followed instructions more closely.
Then people started asking: is this Gemini 4 Pro?
The catch: that name already belongs to something
This is the part most easily misread, and if you remember one thing from this post, remember this: Gemini 3.8 Flash is not a made-up label. It is a model Google officially released on September 2, 2026.
That part is fully documented. Google's own announcement calls it their "most intelligent workhorse model," with clear gains over 3.7 Flash in software engineering, agentic tasks and multi-step reasoning, at the same introductory price as its predecessor: $0.75 per million input tokens and $3.75 per million output tokens. Worth noting, because it matters later: that is promotional pricing and it expires on December 31, 2026 — from January 1, 2027 the rate is $1.50 in and $7.50 out. Google also launched a variant called Gemini 3.8 Flash Cyber, tuned for finding and fixing software vulnerabilities. 9to5Google and The Register both covered it that day.
So the suspicion runs like this: the model on Arena wore a name people had already used and roughly benchmarked by feel — and then produced work above that tier. Like a test car wearing a production badge that laps well above its class; of course people assume the engine was swapped.
But pause here: going fast doesn't prove what's under the hood. The same model can swing wildly across task types, testers may have drawn a different checkpoint, or their prompts may simply have suited it. That point matters again later.
2. What the rumor claims (all of it unverified)
Here is every claim I found across the coverage, with my own note on the right:
| Item | The claim | My note |
|---|---|---|
| Identity | An internal Gemini 4 Pro checkpoint, codename Argon | Pure inference. Google has confirmed none of it. |
| Input limit | 10 million tokens | No official post or API doc backs this. |
| Output limit | 256,000 tokens | Same — unsourced. |
| Memory | Persistent memory across sessions | Unsourced, and close to impossible to verify as a user. |
| Pricing | $2.25 per million input tokens, $11.25 output | A clean multiple of both official prices: 3x the promo rate, 1.5x the 2027 standard rate. Too tidy — looks extrapolated to a Pro tier, not copied from a price sheet. |
| Benchmarks | DeepSWE v1.1 88.7%, Terminal-bench 2.1 95.3%, OSWorld-2.0 86.8% — beating GPT-6 Astra and Claude Fable 5.1 | From a table circulating on social. The sources don't even agree with each other — some print 88%, others 88.7%. The original report itself wrote "if this table is accurate." |
| Release timing | Community guesses October 2026 | Google has announced no timeline. |
"Outclasses GPT-6 Astra and Claude Fable 5.1" comes entirely from those three numbers. If you want to see what officially published benchmarks look like by comparison, check my earlier piece on GPT-6 Astra — every figure there traces back to OpenAI. These don't. That's the difference.
Put the rumored numbers next to Google's published table
This part you can check yourself. When Google shipped Gemini 3.8 Flash it published a full evaluation table covering its own models and several rivals (the chart above comes from it). Here are the rumored figures dropped into that context:
| Benchmark | Rumored Gemini 4 Pro score | Best score on Google's published table |
|---|---|---|
| DeepSWE v1.1 Long-horizon software engineering |
88.7% | 74.0% (Claude Opus 5) Gemini 3.8 Flash: 73.7% |
| Terminal-bench 2.1 Agentic terminal coding |
95.3% | 89.4% which is Gemini 3.8 Flash itself |
| OSWorld-2.0 Agentic computer use |
86.8% | 75.4% (Claude Opus 5) Gemini 3.8 Flash: 59.0% |
| GDPval-AA v2 Knowledge work, scored in Elo |
2064 | 1824 (Claude Opus 5) Gemini 3.8 Flash: 1545 |
All four sit well clear of the field, and the DeepSWE figure is nearly 15 points above the strongest model on the table. That doesn't make it fake — a new generation is supposed to jump. But a jump that size sitting quietly under a Flash-tier label doesn't add up: with results like that, Google holds a launch event. That's the part I'd question hardest.
And one bigger leap: AI improving AI
The later half of the reporting adds another layer: RSI, or recursive self-improvement. In plain terms, you let AI help design and train the next generation of AI, and the stronger next generation helps improve more — a snowball. The article ties Gemini's progress to that idea and implies Google used it to finish Gemini 4 faster.
That part is entirely inference. RSI is a genuine research direction, but "Google already used it to complete Gemini 4" has no official statement or technical paper behind it. It's the claim I'd discount most, because it sounds the best and is the hardest to check.
3. Video
Video: Sketched Truths on YouTube (third-party recap, not Google).
4. What Google actually said
This part is well sourced, and it amounts to two things.
One: July 21, 2026 — Google confirmed Gemini 4 training had started. At the very end of the post announcing Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber, the Gemini team wrote: "We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress." Saying that mid-training, rather than at launch, is unusual for Google — which is why it got noticed at the time.
Two: July 22, 2026 — Pichai gave the direction on the Q2 earnings call. He said the next leap depends on "much larger base models," and named the gap to close as coding and agentic coding — AI that writes and runs the code itself, not general-purpose autonomous agents. That came from the call itself; Alphabet's own earnings post pins the date, and reporting from the call has the wording.
That's it. No release date, no specs, no scores, and not one line confirming that the Arena model is Gemini 4 Pro. "In development" is officially supported. "Already live and beating everyone" is not something you can treat as settled yet.
5. Three things I do when I see a story like this
"Mystery model quietly appears, crushes everything" shows up roughly monthly. I run three checks that take under five minutes total:
- Ask who is actually saying it. The closer a source sits to the people involved, the more weight it carries. The strongest is whatever the company publishes itself — official blog, API docs, earnings calls — because they're accountable for getting it wrong. Next is reported interviews: a journalist at least asked someone and put their name on it. Loosest is a user's impressions, which may well be right but have nobody standing behind them. Measure this story with that ruler: "Google started training Gemini 4" is Google writing about Google. Everything else — the model's identity, the specs, the scores, the codename — stops at somebody's impressions.
- Check whether the benchmark table has a source. Official results name the benchmark, the version and the test setup. Community tables are usually just a row of numbers. If it's only numbers, ignore it for now — this report literally wrote "if the scores are accurate," meaning the author knew too.
- Compare the headline against the body. "Already launched" turns out to mean "someone met it on a test platform." "Crushes" turns out to mean "an unverified table." A headline more certain than its own body is the most common way a story gets inflated — and it doesn't require anyone to lie.
I'm not saying the rumor is false. Google really is training Gemini 4, an October launch is entirely plausible, and much of this may look right in hindsight. My point is narrower: it's too early to change your workflow or make decisions on it.
6. Why this matters if you make content or run a small business
Two practical things.
First, don't rebuild your workflow for a model that isn't out. I've watched people pause a working document-processing setup because a rumor said the input limit was about to hit 10 million tokens. Two months later: no new model, no progress either. My rule is to wait until a model actually ships, then run one small project I genuinely have to deliver through it before deciding whether to switch. Anything that survives one real delivery is worth learning. The tools I actually use and that hold up under delivery are collected in the Zeona toolbox.
Second, stories like this get cited heavily by AI search, and which pages get cited follows a logic. For the same rumor, ChatGPT and Perplexity tend to pull from pages that clearly separate what was officially said from what wasn't — those are easier to extract from and less likely to produce a wrong answer. If you make content and want to understand how that works, I wrote a complete guide to generative engine optimization and a case study where it actually moved numbers.
FAQ
Q1: So is Gemini 4 Pro out or not?
No. As of September 22, 2026, Google has not released Gemini 4 Pro and has announced no timeline. The only official word is the July 21 confirmation that Gemini 4 pre-training had begun. The community guesses October, but that's a guess.
Q2: Then what is the gemini-3.8-flash I met on Arena?
Most likely the shipped Gemini 3.8 Flash, publicly released on September 2. The rumor holds that some of those sessions were Google testing a stronger model behind the same label, but that is unconfirmed and unverifiable from the user side — you see a name and an answer, not which checkpoint is serving it.
Q3: What's the difference between Gemini 3.8 Flash and the rumored Gemini 4 Pro?
One already happened, the other hasn't. Gemini 3.8 Flash shipped September 2 as a workhorse-tier model at an introductory $0.75 input / $3.75 output per million tokens (rising to $1.50 / $7.50 in 2027), with a Cyber variant for vulnerability work. Gemini 4 Pro has exactly one official data point — that it's in training. Google hasn't even used the full name itself.
Q4: Are the 10M-token and persistent-memory specs credible?
Not right now, because nothing sources them. Every figure traces back to community roundups rather than official posts or API docs. One clue you can check yourself: the rumored pricing lands on a clean multiple of both official rates — 3x the promotional price (0.75 / 3.75) and 1.5x the standard 2027 price (1.50 / 7.50). That tidy looks extrapolated to a Pro tier, not copied off a real price sheet.
Q5: What is RSI, and is Google really using it?
RSI means letting AI help improve the next generation of AI, which then helps improve more. It's a real research direction, but "Google already used RSI to finish Gemini 4" is the article's inference, with no official statement or technical documentation behind it. It's the claim I'd discount most in this whole story.
Sources
Rumor coverage (unconfirmed by Google):
· Original English item: Futu News (body text requires JavaScript)
· Chinese reprint of the Xinzhiyuan piece: Wallstreetcn
· Dataconomy, OfficeChai, the atoms.dev demo roundup
· Developer impressions: @pankajkumar_dev on X
· Video recap: Sketched Truths on YouTube
Official sources:
· Gemini 3.8 Flash and 3.8 Flash Cyber announcement (September 2, 2026 — pricing and eval charts)
· Gemini 3.6 Flash announcement (July 21, 2026 — the Gemini 4 pre-training line is at the end)
· Alphabet Q2 2026 earnings post (July 22, 2026) and reporting from the call
· Shipped 3.8 Flash coverage: 9to5Google, The Register
Charts and cover image are from Google's official blog, copyright Google. The embedded video is the property of its YouTube channel.
If you want to talk through how to judge whether a new AI tool belongs in your workflow, or you need help making your content and site visible in AI search, get in touch.


