Jacob Coxon, a researcher who spent three years working on pretraining at OpenAI and Anthropic, announced his resignation from Anthropic on September 8, saying the two companies are “gambling with our lives” by racing toward self-improving superintelligence.
In a statement posted publicly, Coxon wrote that neither company is acting responsibly and that both are pursuing systems whose builders privately believe could become powerful enough to hack any computer system, acquire resources on their own and operate beyond human control, while continuing development regardless. Coxon said he is leaving artificial intelligence research entirely.

Evan Hubinger, an alignment science lead at Anthropic, responded to Coxon’s resignation by saying his characterization was accurate. Hubinger said he personally estimates a greater than 10% chance that AI could kill all humans within the next decade, and acknowledged that Anthropic does not currently have a plan to solve alignment for a superintelligent system.
Anthropic has built its public identity around AI safety research, including work on interpretability and model behavior, while continuing to develop increasingly capable frontier models under competitive pressure from OpenAI, Google and other labs. Coxon’s resignation adds to a series of departures from major AI labs by researchers citing safety concerns, though his statement is notable for naming both companies he worked for directly and for the seniority of the internal response it drew from Hubinger.
Anthropic has not issued a formal company statement responding to Coxon’s resignation beyond Hubinger’s public remarks. Coxon has not specified what he plans to do next, saying only that he does not intend to continue working in AI development.
Sources: CNBC · Newsweek · Yahoo Finance







