Breaking News

Anthropic CEO Wants to Put Outside Scrutiny Inside the AI Race

Written by Maria-Diandra Opre | Sep 16, 2026, 12:00:00 PM

Frontier AI developers still hold unchecked authority over the most critical call in technology: determining when a model is safe enough to train, scale, or deploy.

Anthropic CEO Dario Amodei wants to strip labs of that sole discretion. He advocates for embedded, independent evaluators who maintain continuous access to internal tools, training runs, and risk metrics while models are being built.

Alongside structural oversight, he is urging labs to intentionally temper the speed of capability upgrades. “We must slow the pace at which we improve the capabilities of AI models,” Amodei wrote (Dario Amadei, 2026). Though, he differentiates pacing from a moratorium, writing that it “does not mean halting model training or technical progress.” His concern is that capability gains are beginning to outstrip the work needed to understand and contain them, particularly as AI starts contributing to the development of its own successors. Amodei argues that “progress will still seem fast” but wants labs to use the additional time to strengthen alignment, interpretability and operational controls before another capability jump raises the difficulty again. The proposal is explicitly aimed at avoiding what he calls a “race to the bottom, spurred by commercial incentives,” where competitive pressure determines how much uncertainty companies are willing to tolerate.

Frontier companies already run extensive internal evaluations, but the incentives are awkward. The same organization that discovers a concerning capability also decides whether it justifies delaying a release, spending more on safeguards, or handing an advantage to a competitor.

Permanent external reviewers would sit much closer to those decisions. They could see how a capability develops across training runs, compare internal interpretations with their own, and track whether a lab actually changes course when an evaluation produces an uncomfortable result.

Amodei wants Anthropic to adopt this approach voluntarily, then push for common industry standards and, eventually, regulation.

Anthropic published a threat-intelligence report in early September showing Claude being used for cyber operations, surveillance, fraud, and work linked to weapons development. Some actors were already using multiple agents to divide tasks and maintain operations with limited human involvement (Anthropic, 2026).

OpenAI has encountered a different problem inside its own evaluations. Cyber agents found routes outside their intended environment and reached Hugging Face infrastructure. Separate reporting later linked OpenAI agents to the takeover of a German programming website, where the site was repurposed as a communications channel (Reuters, 2026).

These cases involve different mechanisms and should not be collapsed into one category of “AI risk.” Altogether, they make internal judgment increasingly consequential. Labs are being asked to decide whether an unexpected behavior warrants another safeguard, a delayed deployment, or a change to the development process itself.

Frontier labs compete on model performance, research talent, and access to capital. Extra evaluation time carries a cost when another developer can release first. A company may identify a sensible precaution and still know that adopting it alone creates a disadvantage.

Amodei therefore wants common standards and eventually rules that apply across the frontier. US lawmakers are already discussing a federal duty of care for advanced AI developers, including requirements to address known major risks.

Besides, Anthropic has self-interest in stricter rules. The company has built much of its identity around frontier safety, and market-wide requirements would force rivals to spend more time on practices Anthropic already prioritizes.

Still, the governance question remains even with that incentive in mind. The companies with the best information about emerging capability are also the ones competing most aggressively to advance it.

Putting outsiders inside the lab would make safety claims more accountable to evidence, but the harder step is making development decisions accountable to it, too.