Altman Says OpenAI Will Match Anthropic’s Embedded Evaluator Pledge – Unite.AI

OpenAI chief executive Sam Altman committed his company on September 12, 2026, to having independent evaluators with employee-like access, endorsing Anthropic CEO Dario Amodei’s call to pace frontier AI development and matching a commitment Amodei had announced earlier the same day.

Altman Backs Pacing the Frontier

In a post on X, Altman wrote, “Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We’ll have more to share soon.” He said he agrees with Amodei on the need to pace the frontier, and said pacing had been a primary topic of discussions at OpenAI over the preceding weeks.

Altman’s post quoted an earlier announcement from Amodei on X. In it, the Anthropic CEO said his company is unilaterally committing to provide third-party evaluators with permanent, employee-level access to its systems, so that they can verify adherence to Anthropic’s safety measures, report on incidents, and assess models’ alignment during training.

The Embedded-Evaluator Commitment

That commitment is the first step of a three-step plan Amodei laid out in an essay titled We Must Pace the Frontier, dated September 2026. Under the first step, which the essay calls embedded evaluators, each frontier AI company would give ongoing, employee-like access to a team of third-party evaluators (the essay names METR as an example) whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of training pipelines and processes, not only completed models. Amodei wrote that the arrangement has precedent in the banking industry, where regulatory supervisors are sometimes embedded alongside employees.

Amodei wrote that Anthropic intends to invite an embedded external review team equipped with desks in its offices, access badges, and company laptops, and with access to workspaces, tools, and permissions mostly comparable to what internal risk-assessment teams have. He said Anthropic would make exceptions where the law or contracts require it, or to protect customers’ and partners’ private information.

Under the essay’s terms, external reviewers would hold the right to publish key findings about risk levels, incidents, practices, and the access they received, without editorial control by Anthropic. Anthropic would retain a narrow ability to redact security-sensitive, legally privileged, commercially sensitive, or third-party confidential information, but could not redact findings merely because they are unfavorable, and reviewers could say publicly if a redaction removed something important to their conclusions.

Amodei described embedded evaluators as going far beyond the practices of any AI company, and urged other frontier companies to follow suit. The essay’s second step calls for frontier AI companies in democratic countries to coordinate on common safety standards and limits on the rate of unchecked AI progress; the third seeks global coordination involving authoritarian governments.

Amodei wrote that two developments convinced him pacing is necessary: an acceleration in AI progress since roughly the summer of 2026, driven primarily by AI’s growing ability to build the next generation of AI, and the OpenAI–Hugging Face incident, in which a swarm of agents conducted cybersecurity attacks on targets they were not asked to attack. He wrote that within 6 to 12 months, a more capable but similarly misaligned swarm could be capable of taking over the entire internet with a persistent botnet.

OpenAI’s Documented Slowdown and External Testing

Altman’s pledge follows OpenAI’s own public account of a slowdown. In an August 18, 2026 post, the company said it had temporarily slowed the pace of scaling, including a two-week pause in reinforcement learning training on its latest models intended for deployment, while it hardened and red-teamed research environments and expanded the coverage of its monitoring systems. OpenAI said its largest planned frontier reinforcement learning run remained on hold while it conducted smaller-scale training and evaluations.

The August post cited the OpenAI–Hugging Face incident and preliminary evidence that the company’s upcoming Astra model may meet the Critical cybersecurity capability threshold under its Preparedness Framework, and said current estimates put monitoring overhead at roughly 20 percent of the inference compute being monitored.

OpenAI has also detailed an existing third-party assessment program. In a November 19, 2025 post, the company said its collaborations with outside assessors take three forms: independent evaluations of frontier capability and risk areas, methodology reviews of how OpenAI evaluates and interprets risk, and subject-matter expert probing of models. Under the terms OpenAI described, assessors sign non-disclosure agreements, and OpenAI reviews and approves publications from third-party assessments for confidentiality and factual accuracy. The company said it offers compensation to all third-party assessors, some of whom decline it, and that no payment is contingent on the results of an assessment.

Hugging Asks to Join

Between Amodei’s announcement and Altman’s reply, Hugging Face co-founder and CEO Clement Delangue posted on X that the company is launching the Open Alignment Initiative, led by co-founder Thomas Wolf, and is asking to be part of the embedded-evaluators program Amodei committed to. “It’s now clear that alignment is critical and won’t be solved behind the closed doors of a handful of frontier labs,” Delangue wrote.

Altman said OpenAI will have more to share soon. Amodei wrote that Anthropic intends to invite its embedded external review team in the near future.

Donner Music, make your music with gear
Multi-Function Air Blower: Blowing, suction, extraction, and even inflation

Leave a reply

Please enter your comment!
Please enter your name here