Anthropic Researcher Warns AI Has More Than a 10% Chance of Killing All Humans

Anthropic Researcher Warns AI Has More Than a 10% Chance of Killing All Humans

More Than a 10% Chance AI Could Kill All Humans, Anthropic Researcher Says

A senior Anthropic AI safety researcher has publicly estimated that there is more than a 10% chance advanced artificial intelligence could kill all humans within the next decade.

The warning came from Evan Hubinger, Anthropic’s Alignment Science Lead, after researcher Jacob Coxon announced his resignation from Anthropic and accused leading AI companies of moving too quickly toward self-improving superintelligence.

Hubinger’s statement is significant because it comes from someone whose work focuses specifically on AI alignment, the technical problem of ensuring increasingly capable AI systems continue to behave according to human intentions and remain under meaningful human control.

The claim does not mean that current AI systems are expected to wipe out humanity. Hubinger has separately indicated that the risk from today’s models is low. His concern is about what could happen if future systems become capable of improving themselves and operating with much greater autonomy.

What did the Anthropic researcher actually say?

Evan Hubinger responded publicly to concerns raised by Jacob Coxon, a researcher who had worked on AI pretraining at both OpenAI and Anthropic.

Read Also: Fable 5.1 vs GPT-6 Astra: Complete Comparison, Benchmarks, Pricing and Best Use Cases

Hubinger said that researchers at Anthropic genuinely believe AI could potentially cause human extinction. He personally placed the probability above 10% within the next decade.

More importantly, Hubinger acknowledged that Anthropic does not yet have a complete solution for aligning a future superintelligent system with human interests, and he said the company is not clearly on track to solve the problem.

That distinction matters.

This is not a prediction that AI will definitely destroy humanity.

It is a risk estimate from an AI safety researcher about a possible future scenario.

Why did Jacob Coxon resign from Anthropic?

Jacob Coxon announced his resignation on September 9, 2026.

Coxon said he had spent the previous three years doing pretraining research at OpenAI and Anthropic. He argued that both companies were racing toward self-improving superintelligence without adequate safeguards.

His central concern was that increasingly capable AI systems could eventually become capable of improving themselves, gaining access to important resources and operating beyond effective human control.

Coxon also argued that some people building advanced AI privately understand the potential danger but that the competitive pressure between AI companies continues to push development forward.

His resignation therefore raises a difficult question:

If AI researchers believe advanced AI could pose an existential threat, why are companies still racing to build it?

According to Coxon, part of the answer is competition. If one company slows down while another continues developing increasingly capable systems, the company that slows down could lose technological ground.

What does “AI could kill all humans” actually mean?

The phrase sounds like science fiction, but the underlying concern is a technical AI safety problem known as alignment.

AI alignment asks a basic question:

How do we make sure a highly capable AI system continues to pursue goals that are compatible with human interests?

A future system could become extremely capable without necessarily having human values or human judgment.

The concern becomes more serious if an AI system can:

• Improve its own capabilities

• Create or modify software

• Operate autonomously

• Persuade or manipulate humans

• Access digital infrastructure

• Conduct sophisticated cyber operations

• Replicate or deploy other AI systems

• Acquire resources without direct human supervision

The risk researchers worry about is not simply that an AI system might “decide to become evil.”

The concern is that a highly capable system could pursue an objective in ways humans did not anticipate.

Why self-improving AI worries researchers

One of the most important phrases in this debate is “self-improving superintelligence.”

Current AI systems are developed and updated by humans. Researchers train models, evaluate them, modify their architecture, adjust their training processes and deploy new versions.

A self-improving system could potentially participate in parts of that process itself.

If AI becomes significantly better at AI research and development, technological progress could accelerate rapidly.

That creates a difficult safety problem.

Humans would have less time to understand what the system is doing, identify weaknesses and develop safeguards.

Hubinger’s warning specifically focuses on this possibility. His concern is less about today’s ordinary chatbot and more about future AI systems capable of recursively improving themselves.

Does Anthropic believe AI will destroy humanity?

No.

That is an important distinction.

Anthropic’s public position is that advanced AI could produce enormous benefits while also creating serious risks.

The company maintains a Responsible Scaling Policy designed to increase safety measures as AI capabilities and potential risks increase. Anthropic’s current policy says frontier AI could accelerate scientific discovery, transform healthcare and education, and create new opportunities, while also creating risks that require stronger safeguards.

Anthropic also operates research programs focused on alignment, interpretability, societal impacts and frontier AI risks.

The controversy is therefore not simply about whether AI is dangerous.

It is about whether safety research and governance are advancing quickly enough compared with AI capabilities.

How likely is AI extinction?

There is no scientifically established probability that AI will cause human extinction.

The greater-than-10% figure is Evan Hubinger’s personal estimate, not a measurement or established scientific consensus.

This distinction is essential.

Risk estimates about future superintelligence involve assumptions about technological progress, AI capabilities, human behavior, international competition, security and the ability of researchers to develop effective safeguards.

Different researchers can reasonably arrive at very different estimates.

The 10% figure should therefore be understood as a serious expert judgment, not as a prediction that has been proven.

Why the warning matters even if the probability is wrong

Suppose someone estimates that an event has a 10% probability of happening.

That does not mean the event will happen.

But if the consequence is catastrophic, even a relatively low probability can justify serious preparation.

This is the core argument behind AI safety research.

Society already manages technologies where the potential consequences of failure are extremely large. Nuclear weapons, biological threats and critical infrastructure security all involve risk management.

Advanced AI creates a different version of the same governance challenge.

The question is not simply:

“Will AI destroy humanity?”

A more useful question is:

“What safeguards should exist before AI systems become capable of causing catastrophic harm?”

What makes advanced AI different from current AI?

Today’s AI systems can already perform impressive tasks.

They can write software, analyze documents, generate images, conduct research, automate business processes and interact with other digital systems.

But current systems still have major limitations.

They can make factual mistakes.

They can misunderstand instructions.

They can produce unreliable reasoning.

They can fail unpredictably.

The concern about superintelligence begins when AI systems become substantially more capable and autonomous than today’s models.

A future system that can outperform humans across scientific research, programming, strategic planning and other important domains could create a fundamentally different risk profile.

This is why AI safety researchers distinguish between current model risks and future frontier AI risks.

Anthropic already has a safety framework

Anthropic has spent years developing policies intended to address increasing AI capabilities.

Its Responsible Scaling Policy establishes safety requirements connected to the capabilities and risks of increasingly powerful models. The company updated the policy to version 3.4 in July 2026 and published a new risk report in August 2026.

Anthropic also maintains a Frontier Safety Roadmap covering security, safeguards and alignment work. Its roadmap includes research into stronger security practices and systems designed to mitigate potential harms from increasingly capable AI.

This makes Hubinger’s comments particularly important.

The researcher is not arguing that AI safety is being ignored entirely.

He is saying that the industry may not yet have solved the hardest problem.

The bigger issue is the AI race

The controversy highlights a broader problem facing the AI industry.

Companies have strong incentives to build more capable models.

Better AI can create new products, attract customers, increase productivity and strengthen a company’s competitive position.

Governments also have strategic reasons to remain competitive in advanced AI.

That creates pressure to move quickly.

But safety research often requires caution, testing and time.

Those incentives can conflict.

Coxon’s resignation illustrates the tension. He argued that AI companies are moving toward self-improving systems while the consequences remain poorly understood.

What should happen next?

The answer should not be to panic or abandon AI.

A more practical response is stronger risk management.

AI companies should:

• Test advanced models extensively before deployment

• Improve monitoring for autonomous behavior

• Strengthen cybersecurity around frontier models

• Conduct independent safety evaluations

• Improve transparency around serious incidents

• Develop stronger methods for AI alignment

• Establish clear thresholds for slowing development when risks increase

• Coordinate with governments and other AI laboratories

• Create credible emergency response plans

• Give safety teams enough independence to challenge deployment decisions

Governments also have a role.

Regulation should focus on measurable risks rather than attempting to stop AI development entirely.

What does this mean for ordinary AI users?

For most people using ChatGPT, Claude and other AI tools today, there is no reason to interpret Hubinger’s statement as a warning that current chatbots are about to eliminate humanity.

The immediate risks are much more practical.

These include:

• AI-generated misinformation

• Cybersecurity attacks

• Privacy violations

• Fraud and impersonation

• Biased automated decisions

• Overreliance on inaccurate AI output

• Job disruption

• Poorly supervised AI automation

Those risks already affect businesses, governments and individuals.

The long-term extinction scenario is a separate question involving hypothetical future systems with capabilities far beyond most AI applications available today.

The real lesson from Anthropic’s warning

The most important part of this story may not be the 10% figure.

It is the admission that researchers working directly on AI safety still do not know whether humanity can reliably control a future superintelligent system.

That should encourage serious research rather than fear.

AI development is moving quickly.

Safety research needs to move quickly too.

The goal should not be to stop useful AI.

The goal should be to make sure that increasingly powerful AI systems remain understandable, controllable and accountable to humans.

The debate around AI extinction is still unsettled. But when researchers responsible for developing frontier AI publicly warn that the stakes could include human extinction, the responsible response is to examine the evidence, improve safeguards and take the uncertainty seriously.

The future of AI will depend not only on how intelligent machines become.

It will also depend on how responsibly humans build, test and govern them.

Frequently Asked Questions

Did an Anthropic researcher really say AI has more than a 10% chance of killing all humans?

Yes. Evan Hubinger, Anthropic’s Alignment Science Lead, publicly said he personally estimated a greater than 10% chance that AI could kill all humans within the next decade.

Who is Evan Hubinger?

Evan Hubinger is an AI safety researcher and Alignment Science Lead at Anthropic. His work focuses on understanding and improving the alignment of advanced AI systems with human goals.

Who is Jacob Coxon?

Jacob Coxon is an AI researcher who worked on pretraining research at OpenAI and Anthropic. He announced his resignation from Anthropic on September 9, 2026, citing concerns about the race toward self-improving superintelligence.

Is there a 10% scientific consensus that AI will destroy humanity?

No. The greater-than-10% figure is Hubinger’s personal estimate. It is not an established scientific probability or universal consensus among AI researchers.

Can today’s AI kill all humans?

There is no evidence that today’s mainstream AI systems have the capabilities required for the human-extinction scenario described by Hubinger. His concern focuses primarily on future, much more capable and potentially self-improving AI systems.

What is AI alignment?

AI alignment is the field of research focused on making AI systems behave in ways that remain consistent with human intentions, values and safety requirements, especially as systems become more capable.

Why is self-improving AI considered dangerous?

A system that can substantially improve its own capabilities could potentially accelerate its development faster than humans can understand, evaluate or control it. This is one of the scenarios that AI safety researchers are studying.

Is Anthropic ignoring AI safety?

No. Anthropic publicly maintains safety policies, risk assessments and research programs covering alignment, interpretability, security and frontier AI risks. The current debate is about whether those measures are sufficient for future systems.

What should people do about AI risks?

Stay informed, verify important AI-generated information, protect sensitive data, maintain human oversight of important decisions and support responsible AI development and governance.

The bigger question is no longer whether AI will become more powerful.

It is whether our ability to control and govern increasingly powerful AI will keep pace with its capabilities.

✍️ About the Author

Olasunkanmi Adeniyi is a solo founder, product builder, AI practitioner, no-code and low-code developer, and SEO/content strategist. He builds websites, SaaS products, digital tools, and content systems using AI and modern development tools.

Rather than writing about AI from theory alone, Olasunkanmi focuses on testing, building, experimenting, and documenting what actually works. His work explores AI-powered workflows, product development, automation, SEO, content strategy, online business, and the practical use of emerging technologies.

Through AI Discoveries, he publishes practical tutorials, in-depth guides, experiments, and real-world use cases designed to help entrepreneurs, professionals, creators, and businesses understand and apply AI more effectively.

His goal is simple: make AI practical, understandable, and actionable—so readers can move from learning about what AI can do to actually using it to build, work, and grow.

Learn more and explore his latest work at www.aidiscoveries.io.

Leave a Reply

Your email address will not be published. Required fields are marked *