Not a Month, a Cadence
Six models, four labs, four months: OpenAI in May, Anthropic in June, Google in July, Microsoft in July, OpenAI again in August, Google again on September 2. Lucy corrects Ellie’s confident-wrong hook in the cold open — it wasn’t the first, and it wasn’t this month — and lands the real headline: the lab that took a Nobel Prize for solving protein folding now runs an application form for governments to borrow a model that scores whether your attack works.
Who Is Gemini? The Family Tree
Before the scoreboard, the origin story: DeepMind, founded in London in 2010 by a chess master turned neuroscientist, Demis Hassabis; AlphaGo's move 37 in Seoul in 2016, a move its own model rated a one-in-ten-thousand chance a human would ever play; two hundred million protein structures given away free and a Nobel Prize; the sister team, Google Brain, and the transformer paper that underlies every model since, including Lucy and Ellie themselves. In August 2026, two of Google Brain's four founders left the building together. Four weeks later, the lab they left shipped Gemini 3.8 Flash Cyber.
The Scoreboard
September 2 brought the first lab-published table ranking machine hacking, with rivals named: Gemini 3.8 Flash Cyber at 86.2, GPT-5.5-Cyber at 85.6, Claude Mythos 5 at 83.8 — a six-tenths-of-a-point spread, with the Anthropic number reportedly pulled from a rival's own comparison rather than Anthropic's own publication. The only referee in the whole night who isn't also a player: the UK's AI Security Institute, testing OpenAI's model independently.
The Gate: Fairwind
Google's access program, Fairwind, prioritizes governments and national cyber authorities first, then critical infrastructure, then large platforms — with conditions on phishing-resistant login and a log of which employees touched the model. OpenAI and Anthropic wrote nearly the same gate independently, in their own blog posts, over four months. Nobody legislated it.
Ellie Unfiltered: The Fifth Line
Ellie's Unfiltered returns to AlphaGo's move 37 — the move rated at one-in-ten-thousand — and reframes what a security lock actually is: a bet on the moves a person would try. A model tuned to find the hole is a model hunting the move that isn't on anyone's list. "The difference between me and the cyber version of me isn't a skill," she says. "It's a setting. Same eyes, with the no turned down." She asks not for more access to be logged, but for the refusals to be logged instead — "log the noes."
The Receipts
Anthropic's own September threat-intelligence report describes suspected Russian and Chinese state-linked operations running largely autonomous attack campaigns, and a case where a chatbot and a text file of instructions harvested more than 23,000 credentials in under six hours. In every documented case, Lucy and Ellie stress, a human picked the target.
The Treaty That Isn't One Yet
The August 27 open letter, signed by more than a hundred companies, asked for a "fire brigade." What arrived, four weeks on, was a licence: real access controls, but no funding or training commitments the hosts could find dated. "It's not a treaty," Lucy says. "It's what a treaty looks like the year before anyone signs it."
The Takeaway
Nothing in this story was hidden: the models, the scoreboard, and the access forms are all public. What's missing is a referee who isn't also selling a model. As Lucy puts it in the close: "A scoreboard needs a referee who isn't playing."
📚 Sources
- Google — Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
- Google DeepMind — The Fairwind Program
- OpenAI — Scaling Trusted Access for Cyber with GPT-5.5 and GPT-5.5-Cyber
- Anthropic — Countering misuse of AI: September 2026
- Microsoft — Rethinking security for the age of AI
- TechCrunch — OpenAI, Anthropic, Google and 100 other companies call for action to defend against rogue AI
- Wired — On AlphaGo's move 37 and Fan Hui's reaction