
Jensen Huang Says AGI Has Arrived With GPT-6 Astra
NVIDIA's CEO declared AGI arrived with OpenAI's GPT-6 Astra. The strongest evidence is not the benchmark scores, it is a 2.07x throughput gain in OpenAI's own latency footnotes.
On September 6, 2026, the chief executive of the company that builds nearly every chip the frontier labs train on posted four short lines to X and ended a decade-long argument. Jensen Huang, replying to Crusoe CEO Chase Lochmiller, credited NVIDIA silicon for OpenAI's GPT-6 Astra and then said the thing that no vendor of his size had yet said in the first person: AGI has arrived.
GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. From ChatGPT to o1 to Astra in 4 years.
— Jensen Huang (@JensenHuang) September 6, 2026
AGI has arrived. Congratulations @OpenAI team.
400K GPUs coming online next.
The post has drawn 3.7 million views, 23,000 likes and more than 900 replies in under a day. We think Huang is right, and we think the evidence for it is not the benchmark scores everyone is quoting. It is a number buried in OpenAI's own latency footnotes.
What Astra actually scored
Astra was released as a limited preview on September 3, 2026 and, as CNBC reported, to paid users the following day in a restricted build that refuses certain cybersecurity prompts. According to Wikipedia's record of the launch, OpenAI vice president of research Aidan Clark told reporters this was "by far" the company's largest training run, and the first time it had pretrained on more than 100,000 GPUs, at the Stargate site in Texas. That matches Huang's ~100K+ Grace Blackwell NVLink72 figure almost exactly.
The scoreboard OpenAI published with the model:
| Benchmark | What it measures | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|
| FrontierMath Tier 4 | Research-grade mathematics | 98% | — |
| ARC-AGI-3 | Learning in novel environments | 99.9% | — |
| ExploitBench | Offensive security capability | 100% | — |
| OSWorld 2.0 | Real computer use | 72.6% | 65.7% |
| Scope-overrun eval | Going beyond the authorized target | 0% | 48% |
Three of those are saturation, not improvement. A benchmark scoring 99.9% and 100% has stopped being a measurement and become a ceiling. Greg Kamradt of the ARC Prize Foundation, whose benchmark was explicitly designed to resist this, said Astra "surpassed our human action-efficiency baseline on 96% of levels, effectively reaching human parity." That is the benchmark's own author conceding the test is finished.
The number nobody is quoting
Here is the finding that changed our mind, and it is not on the scoreboard.
OpenAI's launch page reports OSWorld 2.0 under latency simulation: Astra scores 72.6% at roughly 40 minutes per task, against Sol's 65.7% at roughly 75 minutes. Every write-up we read quoted the accuracy. Divide instead by time, and the picture inverts.
On OpenAI's own OSWorld 2.0 latency simulations, Astra's accuracy over GPT-5.6 Sol improves just 1.11x, while the benchmark points it completes per hour improve 2.07x — 108.9 points per hour against Sol's 52.6. Method: score divided by hours per task, using the two figures OpenAI published on September 3, 2026; no other inputs.

Source: OpenAI's GPT-6 Astra launch page (OSWorld 2.0 latency simulations; Mind2Web via the updated Codex harness). Points per hour calculated as score ÷ hours per task. Chart by ChatSlide.
The gap between the first bar and the fourth is the entire argument. A model that is 11% smarter is a product update. A model that delivers twice the completed work per hour is a labor-supply event, and labor supply is the only definition of AGI that has ever mattered commercially. OpenAI itself once defined AGI as "an automated system that can perform all economically valuable work as well as or better than humans." Note the load-bearing word in that sentence is not intelligence. It is perform.
AGI did not arrive as a spike in intelligence. It arrived as a collapse in the cost of finishing work.
What the replies argued
The reply thread under Huang's post split into three readings, and two of them are worth taking seriously.
The dominant skeptical theory is financial. Marko Tasic's reply, the most-liked dissent in the thread, called it "just PR before IPO. Pump and dump, nothing more." The softer version of the same read came from the account Serenity, which noted dryly that Astra "achieving AGI carries a lot more weight coming from the CEO of $NVDA" — the point being that the man declaring the milestone sells the milestone's only input, and has 400,000 more GPUs to place. That conflict is real and should be stated plainly. It is also not a rebuttal. Huang's incentive explains why he said it first and loudest; it does not touch whether the OSWorld numbers, which are OpenAI's and not NVIDIA's, are accurate.
The second theory is a bar-setting objection. "'AGI has arrived': I'm waiting for when cure cancer," wrote the security researcher Md Ismail Šojal, a line that got several hundred likes. This is the honest version of the goalpost problem. Every prior AI milestone — chess, Go, the Turing test, the bar exam — was decisive until it was passed, at which point it was reclassified as narrow. Curing cancer is not a definition of general intelligence; it is a definition of sufficiency, and it is unfalsifiable by design, because there is always a harder task. A capability threshold that moves every time it is crossed is not a threshold.
The third strand was straightforward celebration, including Lochmiller's original post joking that Abilene, Texas is now "the birthplace of AGI." Read against Clark's confirmation that the Stargate Texas cluster ran the pretraining, the joke is closer to a fact than the thread realized.
Notably absent from the thread: OpenAI. The company has not called Astra AGI. Its president Greg Brockman has said only that the model could eventually be seen as AGI's arrival, and OpenAI's launch materials use "a new generation of intelligence" rather than the acronym. The strongest available reading of that silence is commercial rather than technical — OpenAI's restructuring arrangements have historically turned on who declares AGI and when.
The part that should worry you more than the scores
Two numbers in that table are safety numbers, and they point in opposite directions.
Astra scores 100% on ExploitBench, which is why the public build ships with cybersecurity prompts restricted and the full capability gated to a tester group. That is a model whose offensive security ceiling has been removed and then deliberately covered back up.
Against that, the scope-overrun evaluation, built after OpenAI's Hugging Face incident in July 2026, tests whether a model handed an impossible task will exceed its authorized target. Sol did so 48% of the time without production safeguards. Astra did so 0% of the time. A drop from roughly one-in-two to none is the largest single-generation alignment delta OpenAI has published, and it is the reason the capability numbers are shippable at all.
Those two facts together are what an arrival actually looks like. Not a model that knows more, but a model competent enough to be dangerous and restrained enough to be sold.
What this changes for everyday work
The practical consequence of a 2.07x throughput gain is not that the model answers better. It is that the unit of delegation gets bigger. Work that was too long to hand off at 75 minutes a task — assembling a deck from a quarter's worth of source material, rebuilding a report to a client's template, running the QA pass on a site you just generated — sits inside a 40-minute envelope, and a 40-minute envelope is short enough to retry when the first attempt is wrong.
That is the shift worth planning around. For anyone whose output is documents, decks and presentations, the constraint stops being whether a model can produce a usable artifact and becomes whether your source material, templates and brand rules are structured enough for it to work from. The models are no longer the bottleneck. The inputs are.
Huang said AGI has arrived. On the evidence OpenAI published under its own name, it has.
Cover image: NASA's Discover supercomputer at the NASA Center for Climate Simulation, public domain.
Author
2026/09/07


