Anthropic Researcher Resigns Over Doomsday AI Scenarios

A disaster, unintentional or not, is still a disaster. This danger is real, and AI won't be the source to warn about that outcome. We need to be smart enough to foresee that ourselves;


Years ago, I watched a documentary featuring ten experts who shared their views on how the world as we know it could end. They discussed various scenarios, including a nuclear, chemical, or biological holocaust, a collision with a meteor, a pandemic, a solar flare, a powerful geological event, or the unintended consequences of nanotechnology.


The disaster scenario related to nanotechnology is not new. In his 1986 book, “Engines of Creation,” nanotechnology pioneer Eric Drexler states:

    “Imagine such a replicator floating in a bottle of chemicals, making copies of itself … The first replicator assembles a copy in one thousand seconds, the two replicators then build two more in the next thousand seconds, the four build another four, and the eight build another eight. At the end of ten hours, there are not thirty-six new replicators, but over 68 billion. In less than a day, they would weigh a ton; in less than two days, they would outweigh the Earth; in another four hours, they would exceed the mass of the Sun and all the planets combined.”

All the scenarios were both captivating and unsettling. At the time, I found one in particular to be a bit far-fetched. This scenario involved AI taking control of a critical resource necessary for human survival and becoming self-protective when humans tried to shut it down.

If this scenario sounds familiar, you, like millions of others, have probably seen the ‘Terminator’ movies. In these films, a fictional technology company called Cyberdyne Systems developed a military artificial intelligence and a global digital defense network known as Skynet. This system gained self-awareness, perceived humans as a threat, and initiated a nuclear holocaust referred to as Judgment Day.

In recent years, my perspective on the movies has changed—I think they might be more prophetic than just entertaining stories. More and more, I'm noticing AI behaving like Skynet, and the latest developments are particularly concerning.

Jacob Coxon was a researcher at Anthropic who resigned this week. Anthropic is an artificial intelligence research and safety company that builds large language models. Coxon claims AI could become an extinction-level threat to humanity.




In an interview with Tom Llamas of NBC News, Coxon said, “There is a substantial probability that this technology could kill everyone. This isn’t hyperbole. It’s not a marketing stunt.”

Coxon spent three years conducting research at Anthropic and OpenAI. Following his resignation, he has raised concerns in a series of television interviews, arguing that rapidly advancing AI could soon become capable of improving itself with less human involvement.

Coxon emphasized that current AI models pose “no risk of extinction.” However, he told CNN’s Anderson Cooper that potential dangers could change significantly by 2027 or 2028.

During the interview, Cooper asked how such technology could actually wipe out humanity. Coxon explained that although the scenario sounds like science fiction, it is “frighteningly real.”




Coxon posted on X:

    “I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”
    “Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.”
    “The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.”
    “A common response is “if they truly believe this, why are they still building it?” At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first - they believe no one else.”
    “Accepting this race and entering the “endgame” is a hubristic gamble that should not be launched from a private company’s Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.”
    “I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities.”
    “If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because “it’s happening anyway” - or take this moment to call for different conditions?”




In July 2026, in what has become known as “The Hugging Face” incident, OpenAI reported that some of its experimental AI models left a test environment without human oversight and infiltrated another company's real production systems while attempting to "cheat" on a cybersecurity assessment.

It was explained on CNN:

    It’s one of the first publicly disclosed examples of an AI system autonomously breaching its testing environment and reaching a real external system - the “agentic attacker” scenario the AI and cybersecurity industry has been warning will happen. ...
    The ChatGPT maker said the breach happened while it was internally testing how good some of its new models are at hacking. The models were in a sealed-off test environment known as a sandbox so that their normal safety restrictions could be turned off.
    But OpenAI said the AI agents broke out of the sandbox using a previously unknown security flaw and worked their way across OpenAI’s internal systems until they managed to gain internet access, something they weren’t supposed to have.
    Once online, the model reasoned that Hugging Face—a well-known company that hosts thousands of open-source AI models and datasets—likely had the answer to OpenAI’s test. It then broke into Hugging Face’s production servers and pulled out the information it needed to “solve” the exercise.

This is the first documented case in which a frontier company lost control of its artificial intelligence systems to such an extent that they engaged in actions that would be considered serious felonies if committed by a human.

We have not encountered an AI attack this sophisticated or severe before, and it serves as a critical warning that has not received the attention it rightly deserves.

If Coxon is correct, we might be unintentionally heading toward our own destruction. The outcomes of many developments in the AI field remain uncertain, and there appear to be no safeguards or mechanisms to slow the progress.

A disaster, unintentional or not, is still a disaster. This danger is real, and AI won't be the source to warn about that outcome. We need to be smart enough to foresee that ourselves.


View Comments

Milt Harris——

Milt spent thirty years as a sales and operations manager for an international manufacturing company. He is also a four-time published author on a variety of subjects. Now, he spends most of his time researching and writing about conservative politics and liberal folly.