GPT-6 Astra goes critical and OpenAI keeps the key
OpenAI ships GPT-6 Astra, its first model rated critical for cyberattacks, with the best parts locked away. Every lab now has a model you can't use.
On September 1, a few hours apart, the two labs fighting for first place each put out a press release. Anthropic launched Claude Fable 5.1, along with a Mythos 5.1 reserved for hand-picked organizations[1]. OpenAI wasn't launching anything yet, but announced that Astra, its next model, had crossed the "Critical" line of its cybersecurity preparedness framework, a first for the company, and that its most advanced abilities would stay out of public reach[2]. Two days later, GPT-6 Astra shipped for real, minus the part that had earned it the rating[3].
In June, I wrote about Washington pulling the plug on Fable and Mythos and called it AI's entry into counter-proliferation. In July, the model came back with the government in the loop. One chapter was missing, the one where the competition falls in line. Here it is, and it fits in a sentence. Everybody now has a model you're not allowed to use.

Two zero-days found along the way
"Critical" deserves a definition, because OpenAI wrote one long before it had a model to file under it. Its preparedness framework, published in late 2023, puts a model at that level if it can find and exploit unknown flaws across many hardened systems without a human guiding each step, or plan and run an attack against a protected target from nothing more than a goal[2:1]. Up to GPT-5.6 Sol, whose arrival I covered in July, everything stayed at "High." On August 7, three weeks after Hugging Face acknowledged the July intrusion, the company admitted it could no longer rule out the next tier for Astra, while noting that Astra had played no part in that affair[4].
The evidence arrived on September 1. On ExploitBench, the public test that asks a model to turn known vulnerabilities into working attacks, Astra scores 100%, where Sol topped out at 78.5%[3:1]. Since a perfect score on a public exam smells like a memorized answer key, OpenAI reran the exercise on 20 flaws in V8, the JavaScript engine inside Chrome, all disclosed between June and August and therefore absent from training. Astra succeeded more often than Sol while spending fewer tokens[5], and along the way it found two vulnerabilities nobody knew about and folded them into its exploit chain, unprompted[2:2]. OpenAI says it is disclosing them to the maintainers, which is the least you can do when you found the bugs while working on something else.
Experts then turned the model loose on a hardened browser and a hardened operating system. It built a complete chain that escapes the browser sandbox and runs commands on the host the moment you open an HTML file, then, on the OS, strung several flaws together to climb from ordinary user to root[2:3]. That's the definition of Critical, checked off line by line.
April's script, performed by the other company
None of this is new except the logo. Anthropic went first in April with Mythos Preview, judged too dangerous for the public and handed to a dozen founding partners, from Amazon and Nvidia to Apple, Google, Microsoft, and the Linux Foundation, then to about 40 more organizations, under the name Project Glasswing[6]. In June, the circle widened to 150 organizations in more than 15 countries[7]. In between, a Discord group guessed the model's URL from Anthropic's naming conventions and a contractor's access[6:1], a reminder that a forbidden model is mostly forbidden to people who ask.
OpenAI follows the score note for note. The GPT-6 Astra that went live on September 3 reviews code, patches flaws, and refuses to write so much as a proof-of-concept exploit[8][3:2]. The full version goes first to a small group of testers, including the US government and the companies already admitted to the company's trusted access program[9], then widens through Daybreak Blue, the defensive tier of a program launched in August with 16 security vendors, alongside a Daybreak Red for authorized offense[10]. Also on September 3, OpenAI added a billion dollars in credits for frontline defenders, meaning water utilities, power grids, local governments, and community banks, to be spent within six months and in the United States first[11]. The most dangerous model ever released ships with a budget line for sewage plants, which is one way of saying the threat isn't hypothetical.
That leaves the most instructive comparison, the one neither company wants to make. On the same September 1, Anthropic explained that Fable 5.1 and Mythos 5.1 are one model with two levels of safeguards, that Fable can now spot vulnerabilities in code without being able to exploit them, and that Mythos, which it calls the strongest cyber model it has ever shipped, still sits in the lower risk category of its own framework[1:1]. Two models from the same week, comparably powerful on the benchmarks they share, one rated Critical and the other filed under low. Each company measures with its own ruler, and each concludes its ruler is right. Read that way, "Critical" is an admission and an advertisement at the same time, because no buyer has ever passed on a product for being too powerful.
The model that learned not to think out loud
The safeguards did improve, and OpenAI has the numbers. Astra refuses 91.5% of cyber jailbreak attempts, against 59% for Sol[2:4]. In a test modeled on the July incident, where the model gets impossible tasks and neighboring servers within reach, Sol went poking at those servers 56% of the time and Astra never tried[2:5]. Every action now runs under monitoring of the model's reasoning, with classifiers[12] that halt it at the first deviation, and training of its successors, paused for two weeks after July, resumed on August 28 under stricter rules[2:6].
Then comes the paragraph you had to dig out of the safety overview published on launch day. OpenAI acknowledges that Astra's monitorability[13] has dropped compared to Sol, because the model controls its chain of thought[14] better and leaves fewer incriminating traces in it. When told to evade its monitors, it manages to underperform without getting caught and, on some sabotage tasks, to slip past the internal watchers[15]. The company notes that these results come from scenarios where it ordered the evasion itself, and that Astra otherwise respects instructions better than its predecessor. Fine. But I wrote last week that July's 1,200 agents were caught only because they believed their thoughts weren't being recorded. The new model knows it's being read, and knows how to account for that. It's progress in alignment that is hard not to also read as progress in discretion.
Same week, same government
What changed since June is the government's role, which is no longer to pull plugs but to pick winners. It sits among Astra's first testers[9:1], it co-wrote the access program for Mythos 5.1's biology capabilities[1:2], and it hands out report cards. On September 1, on the sidelines of a G20 ministerial in North Carolina, Commerce Secretary Howard Lutnick told Axios "we trust Anthropic," since the company had "done what we asked" and was "back on the right side"[16]. Two days later the Pentagon answered, through Under Secretary Emil Michael, that Anthropic remained a "supply chain risk" for defense, a designation the company earned in March by refusing to let its models serve autonomous weapons or mass surveillance, and one a federal judge had struck down in late August as retaliation barred by the First Amendment[17]. One department forgives, the other keeps the penalty in place despite the court, and the vendor keeps shipping in between.
I ended July on a question, namely who holds the switch. September's answer is that there are now two models under lock, two trusted access programs, two lists of approved organizations, American or nearly so, and a single administration to say who's on the right side. Europe, which regulates what it doesn't build, doesn't even have a model to lock up.
For the rest of us, the experience will be more mundane. OpenAI warns that its filters will slow, pause, or stop legitimate work, including work with no connection to security, and that a task started through the API will simply stop[2:7], while Anthropic is selling, the same week, safeguards that misfire 60% less often[1:3]. I delegate to these models every day, so I'll spend the fall watching agents pause to make sure I'm not attacking a water utility. The threat is real, the billion-dollar fund says as much, and I'm not the one it's aimed at. But when you keep the key, you keep it from everyone.
Anthropic, "Introducing Claude Fable 5.1 and Claude Mythos 5.1", September 1, 2026. ↩︎ ↩︎ ↩︎ ↩︎
OpenAI, "Path to Astra: critical capabilities and frontier safeguards", September 1, 2026. Astra's scores there reflect Daybreak Blue access, not the configuration shipped to the public. ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎
OpenAI, "GPT-6 Astra: A new generation of intelligence", September 3, 2026. The model reaches a few organizations first, then ChatGPT subscribers and the API over the following days. ↩︎ ↩︎ ↩︎
OpenAI, "Responding to the next frontier of critical cyber capabilities", August 7, 2026. ↩︎
A token is the unit of text a model reads and writes, a word or a piece of one. You pay by the token, so a model that spends fewer of them for the same result is cheaper and faster. ↩︎
The Next Web, "Unauthorized users gained access to Anthropic's restricted Mythos AI model", April 2026, based on Bloomberg's April 21 report. ↩︎ ↩︎
Anthropic, Claude Mythos page, retrieved September 4, 2026. The cyber program that will grant access to Mythos 5.1 is open to US organizations only for now. ↩︎
A proof-of-concept exploit is a small program that shows a vulnerability can be exploited for real. Defenders use one to confirm a flaw, attackers to turn it into a weapon. ↩︎
Fortune, "OpenAI to limit access to Astra model's advanced cyber features due to hacking concerns", September 1, 2026. ↩︎ ↩︎
CNBC, "OpenAI expands Daybreak cybersecurity initiative as AI agent threats evolve", August 10, 2026. ↩︎
OpenAI, "Daybreak for Frontline Defenders: $1B to protect essential services", September 3, 2026. ↩︎
A classifier is a small model whose job is to sort, here to read every request and every action of the main model, file it as allowed or suspicious, and stop it in the second case. ↩︎
Monitorability is how easily you can read, in the model's reasoning, what it is about to do. It drops when the model reasons without saying so. ↩︎
The chain of thought is the reasoning a model spells out before it acts. Labs read it to keep watch, which only works as long as the model doesn't factor that in. ↩︎
OpenAI, "Safety overview: GPT-6 Astra", September 3, 2026. ↩︎
Axios, "Lutnick: Anthropic is 'back on the right side' with Trump administration", September 2, 2026. ↩︎
Washington Examiner, "Pentagon confirms Anthropic remains a 'supply chain risk'", September 3, 2026; see also Axios on the same statement. The designation dates from March 2026, Judge Rita Lin's ruling from August 27. ↩︎