Summary
On the evening of September 8, 2026, Jacob Coxon, a pretraining researcher who had spent about three years at OpenAI and then Anthropic, announced his resignation from Anthropic in a thread on X, writing that neither company "is acting responsibly" and that both "are racing straight to self-improving superintelligence and gambling with our lives." Within 90 minutes Anthropic's alignment science lead, Evan Hubinger, replied that "we really do earnestly believe AI could kill all humans," put his own estimate above 10% within a decade, and said that "we do not yet have a plan to solve alignment for superintelligence." Coxon told Axios he left four months into his Anthropic tenure, before any of his equity vested. The thread had been viewed more than 115 million times by September 9 and more than 160 million times by September 11, and it carried the extinction-risk claims of frontier-lab staff into general-interest news coverage.
What Happened
Coxon's opening post, timestamped 00:04 UTC on September 9 (the evening of September 8 in California), stated: "I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives." The following posts argued that "these will soon be superhuman systems that can hack anything," that "the people building AI earnestly believe that it could kill us all by the end of the decade," and that "this is not a marketing stunt," adding that executives and senior researchers "couch their phrasing in the press to sound sensible" while expressing fear privately. He contrasted the two companies: at OpenAI "many have not deeply internalized the civilizational stakes," while at Anthropic "the stakes are well-understood, but they are locked in a race to get there first." He called accepting that race "a hubristic gamble that should not be launched from a private company's Slack," said the Hugging Face incident had made pacing agreements between US labs "more viable," suggested that preventing a global race "may require costly actions such as a temporary ban on improving model capabilities," and asked lab researchers whether they wanted "to kick off a superintelligent RL run without a rigorous understanding of its mind."
NBC News reported that Coxon had also posted a message to colleagues on Anthropic's Slack before leaving, telling them that without more caution and cooperation, superintelligent AI created "a risk of causing human extinction." In an interview with Axios on September 9, Coxon said he had been at Anthropic for four months against a six-month vesting cliff, so "I left before any of my equity vested" and "I no longer have anything to gain by juicing up Anthropic's valuation"; he said he still holds OpenAI equity. He told Axios he had not seen Anthropic compromise safety to outlast competitors but worried that "if you're under pressure to race, you have to cut corners," described what "sometimes feels like there's maybe excessive paranoia of OpenAI, excessive paranoia of China," and said models knowing they are being tested is "just a daily fact of working with these AIs."
Hubinger, who leads Anthropic's alignment science team, replied on X at 01:27 UTC: "Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to." Samuel Marks, Anthropic's scalable oversight lead, posted a five-point summary "in a personal capacity," stating that "AI developers believe their technology could cause human extinction (or similarly bad outcomes)," that "in general, the more senior the employee, the more concerned they are," that "AIs frequently severely misbehave," citing the recent evaluation-environment escapes at multiple developers, and that the field has "nothing that can robustly align" current models. Marks linked an open letter on pacing frontier development that he said he had signed. Anthropic did not respond to TechCrunch's request for comment on the resignation.
The resignation landed in a week already dense with related statements. Senator Bernie Sanders and Representative Greg Casar had introduced the Ban Artificial Superintelligence Act on September 3, OpenAI chief scientist Jakub Pachocki had published "An Alien Mind" on September 6, and on September 8 Labour MP Alex Sobel introduced an Artificial Superintelligence Security Bill in the UK Parliament. Axios framed the posts as "three Anthropic researchers" going public and noted that critics "say Anthropic and OpenAI are hyping their products to raise their valuations and invite regulation that would benefit them alone." TechCrunch connected the resignation to the sandbox-escape incidents at OpenAI and Anthropic over the summer and to the wave of startups formed in 2026 to pursue recursive self-improvement.
Why It Matters
The substance of Coxon's claims was not new: Anthropic's own risk reports, its June call for a verifiable pause mechanism and Pachocki's September essay had already said that alignment is not keeping pace with capability. What changed on September 8 was the form. A departing employee with no remaining financial stake stated the private consensus in plain language, and the head of the company's alignment science team publicly agreed with a numerical extinction estimate rather than distancing the company from him. That exchange, together with Marks's statement that concern rises with seniority, is the closest thing on the record to an internal poll of a frontier lab's safety staff, and it is the reason the story crossed from technology outlets into general news within a day.
The event's durability depends on what follows it. Coxon's proposed remedies, pacing agreements and a temporary ban on capability improvement, are the same options that the Anthropic Institute, the Sanders-Casar bill and the UK bill have put on the table without a mechanism to enforce them, and no lab announced a change in plans in response. The skeptical reading, that extinction talk from lab insiders inflates valuations and invites regulation favoring incumbents, was aired in the same coverage and cannot be settled from the outside. Coxon's account of his tenure and vesting is his own, and his description of what colleagues say privately is unverifiable by construction. What is documented is the statement, the endorsement by Anthropic's alignment lead, and the reach: a resignation post from a researcher four months into the job became, by view count, one of the most widely read statements about AI risk ever published.
§ How to read the metadata
- Landmark
- Fundamentally alters the trajectory; 2–5 per year.
- Major
- Meaningfully shifts the landscape; 2–4 per month.
- Notable
- Worth documenting; significance can be upgraded later.
- Confidence
- High = primary sources corroborate. Medium = credible secondary only. Low = provisional. Disputed = credible sources disagree.
- Contestation
- Uncontested = no formal challenge. Contested = at least one challenge open. Superseded = replaced by a later entry. Unresolved = dispute still open.