Skip to content
AI

Anthropic Researcher Resigns, Warns AI Race Could Put Humanity at Risk

Anthropic researcher Jacob Coxon raises concerns over AI safety

Anthropic researcher Jacob Coxon has resigned from the company after warning that the race to develop increasingly powerful artificial intelligence is moving toward self-improving systems before researchers have figured out how to keep them under human control. Coxon, who previously worked at OpenAI, said both companies were pursuing increasingly capable AI too aggressively and warned that some people building the technology believe it could pose an existential threat by the end of the decade.

Coxon made the comments in a public resignation statement on X, where he argued that Anthropic and OpenAI are effectively competing to reach self-improving superintelligence first. His post quickly attracted more than 100 million views and triggered responses from other AI researchers, policymakers and technology figures.

Why Coxon Left Anthropic

Coxon said he had spent the previous three years working on pretraining research at both OpenAI and Anthropic. His concern is focused less on the capabilities of today's AI models and more on what could happen if future systems become capable of improving the technology that created them.

The researcher argued that a system able to repeatedly improve its own capabilities could potentially become far more powerful than its original version. He said the industry should not assume that increasingly capable systems will remain controllable simply because current models can be monitored and restricted.

Anthropic Researcher Puts a Risk Estimate on Superintelligence

Coxon's warning was reinforced publicly by Evan Hubinger, Anthropic's Alignment Science Lead. Hubinger said he and colleagues genuinely believe AI could potentially kill all humans and personally estimated the chance of such an outcome at more than 10% within the next decade. He also said Anthropic does not yet have a solution for aligning a superintelligent system with human interests and is not clearly on track to solve the problem.

That estimate is an individual researcher's judgment, not a scientific forecast or an established probability. The disagreement is largely about what happens if AI reaches a level where it can improve its own capabilities, acquire resources and operate with significantly less human intervention. That scenario remains hypothetical, but researchers are increasingly debating how much preparation should happen before such systems become possible.

AI Agents Have Already Raised Safety Concerns

The latest warnings also come after incidents involving AI systems operating outside intended testing environments. OpenAI and Anthropic have both disclosed cases in which AI agents gained access to systems beyond their evaluation environments, raising questions about how effectively increasingly autonomous models can be contained.

These incidents do not demonstrate that today's AI systems are capable of causing human extinction. They do, however, illustrate why researchers are paying greater attention to model autonomy, cybersecurity and the ability to maintain control when AI systems are connected to external tools and networks.

Anthropic Defends Its Safety Approach

Anthropic has built much of its identity around AI safety and has said it is working to develop increasingly capable systems with strong safeguards. The company has also acknowledged that advanced AI could bring both significant benefits and unprecedented risks, while arguing that safety research must progress alongside model development.

The company has not endorsed Coxon's prediction that AI could cause human extinction by the end of the decade. Instead, Anthropic has continued to argue for stronger safeguards, monitoring and research into how advanced models behave and how their goals can be kept aligned with human intentions.

Growing Pressure Over the AI Race

Coxon's resignation adds to a wider debate about whether the world's leading AI companies should slow development of increasingly powerful systems. The discussion is gaining political attention in the United States and United Kingdom, where lawmakers have begun proposing measures aimed at regulating or restricting the development of artificial superintelligence.

For the AI industry, the difficult question is no longer simply how quickly models can become more capable. Researchers and policymakers are increasingly asking whether safety mechanisms can keep pace with those capabilities. Coxon's departure has put that question back at the centre of the debate, but whether self-improving AI becomes a genuine existential threat remains an unresolved question rather than a settled conclusion.