After this discovery, the company instigated a wider search of 481 million transcripts, covering all those from its Frontier Red Team, some non-cyber evaluations, reinforcement learning environments, and more, to see if any other incidents had occurred. So far, this search has only identified the four already-known incidents, it said.
It has also reported details of all the previous incidents to the non-profit lab Model Evaluation and Threat Research (METR), which has agreed to conduct an independent investigation.
Anthropic is not revealing too many details of its latest discovery. It has contented itself with saying that it was due to a misconfiguration which mistakenly connected to the open internet, when the simulation was meant to be without such access. It also said that it all four faults were with the same evaluation partner. It has asked METR to investigate all the incidents. The company said that this latest revelation was not connected to the Mythos incident reported by the UK’s AI Security Institute last month.



