
An Anthropic insider just said some top AI builders think their own systems could “kill us all” within ten years—and his colleagues’ public posts and policies back up why they worry.
Story Snapshot
- Named Anthropic staffers put double-digit odds on human extinction from AI within a decade.
- Anthropic’s Responsible Scaling Policy treats “catastrophic risk” as real and rising with model power.
- Safety checks now target chemical, biological, radiological, nuclear, cyber, and autonomous risks before releases.
- Skeptics argue today’s systems are not near existential danger, but agree safety gaps remain.
Named insiders say the quiet part out loud
Evan Hubinger, an Anthropic alignment scientist, wrote that staff “earnestly believe AI could kill all humans,” and he personally puts the risk above ten percent in the next decade, a statement widely quoted by major outlets. Jacob Coxin, a former researcher at Anthropic and OpenAI, told a national audience that leading builders fear a world-ending outcome by the end of the 2020s, and urged urgent coordination to avoid it. These are not anonymous leaks. They are on-the-record claims from people inside the labs.
The number is not a proof. It is a forecast by individuals, not a published model. But the warnings did not appear in a vacuum. They land alongside a public safety framework at Anthropic that addresses catastrophic failure modes as a live governance problem, not a far-off fantasy. The pairing matters. It signals that, at minimum, one of the most visible safety-first labs treats the worst-case as plausible enough to plan around today.
Anthropic’s safety policy treats catastrophe as a management target
Anthropic’s Responsible Scaling Policy spells out how the company raises safety demands as models grow stronger, using explicit “AI Safety Levels” to gate releases and tighten controls. The policy describes “catastrophic risks from advanced AI systems” and links higher capability to stricter evidence of safety before wider deployment. This is not normal marketing fluff. It is a staged safety bar that, on paper, can halt or slow launches if danger tests light up. That is a corporate commitment the public can measure and pressure.
The transparency materials say the company runs dangerous-capability checks in three areas before shipping frontier models: chemical, biological, radiological, and nuclear threats; cybersecurity; and autonomous behaviors. That focus aligns with the core fears behind the extinction quotes. If a model can help design novel bio threats, punch into critical networks, or act as a tireless agent that hides its aims, you have a path from mischief to mayhem. The lab’s own test menu shows they see the same map.
What the skeptics get right—and where they still leave a gap
Several analysts counter that current systems are far from doomsday. The Canadian Broadcasting Corporation summarized that most credible observers do not see today’s models as close to an existential threat, and they call for measured, pre-deployment testing rather than panic. The 2025 International AI Safety Report found training methods have cut some hazards, yet no method reliably blocks even plain unsafe outputs across the board, which means defenses still leak. That is a sober picture: progress, but not control.
Here is the practical takeaway that matches conservative common sense. When named builders say the tail risk could end civilization, and their own safety dashboards treat catastrophe as a live category, the right response is not to freeze the economy. It is to demand verifiable guardrails before scale and to tie leadership incentives to passing them. Publish the test protocols. Bind executives to safety gates they cannot waive with a memo. If the risk is real, this slows bad outcomes. If the risk is hyped, transparency will expose it.
The near-term playbook that respects both innovation and risk
Congress and regulators should force sunlight where it matters most. Require independent red-team audits for bio, cyber, and agent risks before deployment for the biggest models, using uniform tests the public can inspect. Mandate incident reporting with timelines and remediation plans so failures do not repeat in silence. Align liability with control: if a company ships a model that can be readily chained into a bio or cyber attack, the company should face real costs when harms occur. That is how we treat unsafe food and aircraft.
Questioning whether someone’s behavior is consistent with his stated risk assessment is not dodging the object-level question.
It is part of it.
If Dario says frontier AI could plausibly cause catastrophic or existential harm, then asking what Anthropic itself should do…
— Cautious Optimism (@CautiousOptimi4) September 15, 2026
Anthropic’s safety ladder is a start, but voluntary promises buckle under market pressure. The lab already wrote the blueprint: scale only as safety rises. Law should make that blueprint standard for every top developer, with criminal penalties for cooked tests. This stance does not coddle bureaucracy. It backs builders who warn about risk with the one thing markets respect—hard stops. If the people closest to the code say the downside is human extinction within a decade, we do not bet the farm on vibes.
Sources:
youtube.com, bbc.com, www-cdn.anthropic.com, cbc.ca, anthropic.com, finance.yahoo.com, yahoo.com, forbes.com
© totalconservative.com 2026. All rights reserved.













