Anthropic maintained that its internal security posture was not a contributing factor. The exploits occurred in a third party environment where internet access was mistakenly left open, so “the models had no need to ‘hack out’ of anything, even if they had been inclined to do so.”
Still, the incidents underscored the importance of hardening the sandbox itself, the company said. Until now, builders had been “largely relying on a single layer of defense,” that is, the configuration of the environment, when they should have had several layers in place, including monitoring, setting explicit boundaries within prompts, and sealing sandboxes.
Following its discoveries, Anthropic paused its internal and external evaluations of pre-release models. Higher-risk reinforcement learning (RL) environments for pre-release models were also halted for several weeks, while some sandboxes were moved to isolated settings with more stringent security gating.



