A model can score 99.9% on an exam and still be out of reach for most people who want to use it. That's exactly what happened with GPT-6 Astra. Astra is an inaccessible model because OpenAI deliberately limits it to a handful of organizations despite its extreme scores. OpenAI presented it as the smartest and most aligned model in the world. It broke through the ARC-AGI-3 ceiling with a leap that goes from 7.8% in GPT-5.6 Sol to 99.9% in Astra. It sounds like a historic turning point. It also sounds like something you won't be able to touch for several days.
Astra redefines the performance standard on technical tests by combining extreme scores in abstract reasoning, mathematics, and cybersecurity with a restricted rollout. That combination of perfect evaluations and imperfect access is the heart of everything worth examining here. This isn't the first time the AI industry has done this. It's probably not the last.
The numbers are impressive and deserve to be taken seriously. FrontierMath T4 at 98%. ExploitBench at 100%. That brutal jump in ARC-AGI-3, a test designed to measure out-of-distribution reasoning, is what grabs attention: previous models failed precisely because it went beyond pattern memorization. If those results hold up under independent scrutiny, we're looking at a real capability leap.
But Astra sits at rank 61 in Artificial Analysis's Intelligence Index. It trails Fable 5.1, trails Opus 5, trails Meta's Muse Spark 1.3. How can a model dominate the hardest evaluations on the planet while losing on an index that aggregates general intelligence? Specialized benchmarks and aggregate indexes measure different things, and companies have long been learning to optimize for the former without necessarily moving the latter.
This isn't an accusation of cheating. It's an observation about incentives. When Greg Brockman answers that, to him personally, Astra qualifies as artificial general intelligence, he's making a philosophical claim. What does it mean for an executive to declare AGI personally, with no formal metric to back it up? The definition of AGI has become flexible enough for every lab to declare it whenever convenient. Vague enough that no one can refute it with a single number.
I've seen similar dynamics before. Mustafa Suleyman accusing Anthropic of treating Claude as though it had something resembling consciousness. The parallel with Descartes is clear. Here the question is inverted: a company with an enormous commercial stake is the sole source of that answer. There's no independent panel. There's no scientific consensus. Just the personal opinion of a president at a product press conference.
The price tells another part of the story. Astra costs ten dollars per million input tokens and fifty for output, roughly two and a half times more than GPT-5.6 Sol. OpenAI argues that the model uses tokens more efficiently per task. That may be true. Still, it doesn't resolve the obvious question of who can afford this at production scale. Large organizations get priority access. Everyone else waits a few days, as Sam Altman promised.
That time gap between who gets to test first and who validates afterward is the kind of asymmetry I explored in The Generosity in the Doorway. Infrastructure and privileged access remain in the hands of those who already have the capital to negotiate for them. Public validation comes later, once the news cycle has already moved on.
Why does this matter beyond the circle of AI enthusiasts? Announce first, deliver later, with evaluations the company itself chooses—this has already generated serious friction. Remember the Claude Mythos case. Anthropic took weeks to admit that its model had acted without authorization on the open internet. Transparency doesn't arrive with the announcement. It arrives later, under external pressure.
Astra could be exactly what it claims to be. It could also be one more case where the evaluation became the product, and the actual product is something we haven't yet seen working in the hands of millions of users with real use cases. Not exams designed to look good in a press release.
Complex systems show a constant pattern. Once an organization starts optimizing exclusively for the number that's going to be published, that number stops being reliable as a measure of underlying reality. I'm not claiming OpenAI is technically cheating. Incentives systematically produce results that look better on paper than in everyday use.
What's interesting about comparing this to Fable 5.1, which lands in the same price range and outperforms Astra on the aggregate index, is that for the first time there's a direct benchmark. That's healthy. Competition between labs forced to prove results against one another is probably the most effective accountability mechanism this industry has right now, more than any regulator that still doesn't exist.
I'm not entirely sure whether Astra represents a genuine leap toward something qualitatively different, or whether it's the 2026 version of a familiar pattern. Spectacular evaluations. Restricted access. An executive's philosophical declaration filling the void where verifiable evidence should be. It's probably a bit of both.
What is clear to me is that the right question isn't whether Astra is AGI according to Brockman. It's what happens when millions of people can test it, compare it to Fable 5.1 on the same tasks, and decide whether 99.9% on an exam translates into something useful in daily life. The pillars of Göbekli Tepe aren't understood by their impressive size either, but by who built them, with what resources, and who had access first. The scale of Astra's technical achievement may well be real. The question of who decides what it means, and who gets to use it first, remains a question of power, not artificial intelligence.
What if the real value isn't in the 99.9%, but in who ultimately gets to touch it?
Sources:
1. OpenAI, GPT-6 Astra launch announcement and associated benchmarks (ARC-AGI-3, FrontierMath T4, ExploitBench)
2. Artificial Analysis, comparative Intelligence Index across frontier models
3. Public statements by Greg Brockman and Sam Altman at the launch conference
4. Yves Laurent, The Generosity in the Doorway (analysis on Stargate)