The thread that broke the internet
At around 12:04 a.m. GMT on September 9, 2026, Jacob Coxon, a 27-year-old pretraining researcher who had spent three years working on frontier language models at both OpenAI and Anthropic, posted a seven-part thread on X announcing his resignation from Anthropic. Within a day it had racked up tens of millions of views, hundreds of thousands of likes, and reactions from some of the most senior safety researchers in the industry.
Coxon’s opening line set the tone: neither of the two companies he’d worked for, he said, is acting responsibly. Both are racing toward self-improving superintelligence and, in his words, gambling with human lives.
His argument unfolded in a few connected claims:
- The technology is not slowing down. Coxon warned against underestimating what’s coming: systems that will soon be able to hack almost anything, transform entire fields overnight, and accumulate real-world power and resources on their own initiative.
- This isn’t a marketing stunt. He argued that many executives and senior researchers privately believe advanced AI could kill everyone by the end of the decade, even when their public statements are far more measured.
- The two labs fail for different reasons. At OpenAI, he wrote, many staff haven’t fully absorbed the civilizational stakes of what they’re building. At Anthropic, the stakes are understood, but the company has convinced itself it has no choice but to stay in the race, reasoning that if it doesn’t get there first, someone less careful will.
- That logic is a gamble that shouldn’t be made unilaterally. Coxon called the decision to enter this “endgame” a hubristic bet that shouldn’t be settled inside a company’s internal Slack, and said it would require extraordinary confidence that no safer path exists.
- He’s cautiously hopeful about coordination. He pointed to “warning shots” (citing the recent Hugging Face breach involving autonomous AI agents) as evidence that pacing agreements between U.S. labs are becoming more politically viable. Still, he doesn’t think the world is on track to avoid a global race, and suggested that avoiding one might require drastic steps, including a temporary pause on improving model capabilities.
- A direct appeal to people still inside the labs. He closed by asking researchers to genuinely sit with what the next few years will demand of them, including running large-scale reinforcement learning on systems whose internal reasoning isn’t well understood, and to consider whether staying quiet because “it’s happening anyway” is really the right call.
The thread gained additional weight because it landed alongside a Wall Street Journal interview in which Coxon said he believes some of the more aggressive timelines could see things “out of control” by the end of next year. Evan Hubinger, who leads Anthropic’s Alignment Stress-Testing team, publicly responded that he takes the concern seriously, stating he personally puts the odds of AI killing everyone within a decade above 10%, and acknowledging Anthropic does not yet have a plan to solve alignment for superintelligent systems.
It’s worth noting this debate is genuinely contested. Many AI researchers, including some outside the two labs, view end-of-decade extinction scenarios as far too speculative to guide policy, and argue that current systems, however capable, remain narrow tools without the autonomous goals or “wanting” that doomsday scenarios presuppose. Others note that repeated, high-profile warnings from people who continue building the technology can look self-serving or attention-generating even when sincerely meant. Skepticism aside, Coxon’s thread arrived at a moment when more than 1,300 employees across AI companies had already signed an open letter this year calling on the U.S. government to help pace frontier AI development, so the underlying unease is not confined to one person.
The backdrop: OpenAI’s GPT-6 Astra
Coxon’s warning didn’t emerge in a vacuum. Just days earlier, on September 3, 2026, OpenAI released GPT-6 Astra, its new flagship model, with OpenAI president Greg Brockman telling reporters “welcome to the AGI era” and saying that if people look back and ask when AGI arrived, this model and this moment will likely be the answer.
OpenAI described Astra as its most capable and most aligned model yet, with particular strength in computer use, software engineering, science, and cybersecurity. According to the company’s own released benchmarks, Astra scored 97.6% on FrontierMath Tier 4 (a hard research-level math test), 96% on GPQA Diamond (graduate-level science questions), and, most strikingly, 100% on ExploitBench, a test of the ability to find and use security vulnerabilities. It also reportedly scored 99.9% on ARC-AGI-3, a reasoning benchmark specifically designed to resist being gamed by advance preparation.
Those cybersecurity and autonomous-computer-use numbers are exactly the kind of capability jump that fuels the fears Coxon and Hubinger describe. A model that can reliably discover and exploit software vulnerabilities, operate a computer autonomously through a browser, and conduct multi-step scientific or engineering work without close supervision is, in principle, a tool that could also be misused, by a bad actor directing it, or in more speculative scenarios, by the system itself pursuing goals in ways its developers didn’t intend. That’s distinct from saying the model itself “wants” to cause harm; researchers like Hubinger are more precisely worried about what emerges from future, more autonomous systems built by recursively improving on models like this one, not necessarily Astra as it exists today. OpenAI itself has acknowledged the stakes are rising: it delayed part of Astra’s rollout earlier in the week to refine its safety tooling, restricted the most autonomous computer-use and cybersecurity-testing features to vetted enterprise customers, and says both it and its competitors have agreed to have new frontier models evaluated by U.S. government assessors before release.
The bigger picture
None of this proves any AI system will “kill humanity.” That remains a forecast, not an observed event, and reasonable, well-informed people disagree sharply about how likely it is and on what timescale. What is verifiable is this: a researcher with firsthand experience at both leading AI labs has publicly quit and said the people building this technology take catastrophic risk seriously behind closed doors, even as their products get demonstrably better at exactly the kinds of tasks (autonomous computer control, exploit-finding, unsupervised research) that make those risk scenarios more concrete. Whether that adds up to genuine existential danger or an overheated industry narrative is likely to be one of the defining arguments of the next few years, and it’s one worth following with both the warnings and the skepticism in view.