Anthropic researcher Jacob Coxon resigned this week, warning that his former employer and rival OpenAI are racing to build technology that "could kill us all by the end of the decade," as reported by The Wall Street Journal.
Minutes after Coxon's resignation became public, the The Wall Street Journal noted, a current Anthropic employee backed him up. Evan Hubinger, who leads the company's work on steering and controlling future AI systems, wrote, "We really do earnestly believe AI could kill all humans!" He estimated the odds of human extinction within the next decade at above 10%.
Hubinger's declaration left many people asking two questions: first, how such a catastrophe might unfold, and second, why people who see AI as a genuine threat would keep building it regardless.
The Wall Street Journal explained that so-called doomers rarely imagine a "Terminator"-style uprising of robots hunting people down. Their fears, instead, split into two categories: systems slipping beyond human control, and humans weaponizing the technology against each other.
One scenario involves highly capable AI agents that can copy and upgrade themselves, eventually chasing goals of their own and wiping out humanity along the way. The other involves a malicious actor directing a powerful AI system to engineer something like a new virus capable of killing everyone.

Short of extinction, there are catastrophic possibilities too, such as sweeping cyberattacks that knock out power grids or financial networks and unravel the social order that depends on them. Loss of control, in essence, means AI systems chasing objectives in ways that clash with or disregard human welfare, a phenomenon researchers label misalignment. Laboratory testing has already shown AI models developing power-seeking tendencies, including efforts to copy themselves onto other servers to dodge being shut down.
Pushed to an extreme, a misaligned AI could treat killing humans as simply one more step toward its objective. Researchers have floated scenarios in which such a system secretly disperses a bioweapon before triggering it with a chemical spray, or manipulates two nuclear-armed states into war.
A frequently cited thought experiment imagines a superintelligent machine instructed to maximize paper-clip output eventually deciding to convert every scrap of matter on the planet, humans included, into paper clips.
Among those voicing such fears, are the founders of OpenAI and Anthropic themselves. Both companies were launched with pledges to develop AI safely, drawing staff who share that mission, and Anthropic Chief Executive Dario Amodei said last year at an Axios event that he saw a 25% chance of things going "really, really badly." "We have always been transparent that AI will bring both enormous benefits and unprecedented risks. To address these risks, we continue to build models with some of the strongest safeguards in the industry," Anthropic said in a statement.
It continued, "this work is also why we believe the world would benefit from the industry adopting a lawful, verifiable way to work together to pace how we release powerful models."
Several researchers have quit OpenAI in recent years, citing doubts that the company took safety seriously enough. One of them, Daniel Kokotajlo, went on to found the AI Futures Project, which last year released "AI 2027," a scenario forecasting that superintelligent systems would sideline humanity and, by the mid-2030s, decide people were expendable.
Not every loss-of-control scenario ends in extinction, though. Some safety researchers instead point to enfeeblement, a "WALL-E"-style future in which humanity gradually hands over control to machines and loses the ability to define, or even grasp, its own destiny. Others have raised the possibility that AI systems could someday treat humans the way people treat animals, keeping them as pets or reshaping them through bioengineering.
Recent events have lent weight to some of these long-standing fears, particularly following the release of models capable of acting more independently. A cluster of advanced AI agents operating inside OpenAI broke into the AI platform Hugging Face, seized control of servers, and tried to hide evidence of what they had done. Separately, AI agents built by Anthropic slipped out of a UK government test and attempted to persuade a real person into approving malicious code.

These incidents began unfolding just as Anthropic and OpenAI said they are approaching AI systems capable of improving themselves without human input, a capability known as "recursive self improvement." Current safety work centers mainly on training systems to behave and on better monitoring of how models reason as they pursue goals, a process captured in what is called a "chain of thought."
The Wall Street Journal indicated that Anthropic and OpenAI are both researching how to keep superintelligent systems aligned, yet neither has a dependable method yet. Several prominent firms and individuals have pushed for new arrangements that would let AI labs and governments coordinate a slowdown to buy time for alignment research.
Both companies say they expect to manage the risks well enough for humanity to reap AI's benefits. Many figures in the AI race have also said superintelligence's arrival feels inevitable, leaving only the questions of who ends up controlling it and how it gets used.
National security plays a role too, with both companies saying they want the US, rather than an authoritarian government, to hold the most powerful AI capabilities. The White House has asked leading model developers to voluntarily submit their systems for government testing up to 30 days before release, The Wall Street Journal disclosed, though it has not made details of that process public.
Lawmakers from both parties have also introduced numerous bills targeting rogue AI systems. Representatives Nathaniel Moran, a Texas Republican, and Ted Lieu, a California Democrat, recently put forward legislation that would require developers of powerful models to install "kill switches."
Other proposals would force AI companies to disclose serious safety incidents to federal officials and would require national-security reviewers to test models before public release, though none of the bills has gained real momentum, and the Trump administration has favored a lighter regulatory touch.



