AI labs can’t explain what they’d do if their models went rogue

OpenAI has actually paused model deployments after safety incidents. That puts it at the top of a very short list. According to TechCrunch, a new independent assessment of five major AI labs found that most of them have published little to nothing about what they would do if one of their models actively tried to escape human control. That’s not a hypothetical concern anymore. It’s a question regulators are now legally requiring these companies to answer.

The study comes from Guidelight AI Standards, an organization focused on safe AI development practices. Guidelight graded Anthropic, Google, OpenAI, Meta, and xAI on six priority practices drawn from its Control standard. The grading was based entirely on publicly available information, which the organization is careful to note. A low score reflects a lack of public disclosure, not necessarily a lack of internal safeguards. But that distinction matters less than it might seem. If a company has a plan and won’t tell anyone what it is, that plan offers essentially zero accountability.

What Guidelight was looking for is specific: a pre-specified containment plan triggered when an AI is detected trying to subvert human oversight, covering what permissions get revoked, who the model can still operate for, under what constraints, and when it gets taken fully offline. That’s a concrete, procedural definition. And most of these labs haven’t publicly addressed it.

Anthropic’s position is the most striking. The company has built its entire public identity around safety. But Guidelight found that Anthropic’s August Risk Report doesn’t mention limiting model deployment as a possible response to a misalignment or control incident. Meta scored equally low and, according to the report, shows no public evidence of a containment plan or any intention to adopt one. Meta declined to say whether an internal plan exists, pointing instead to a broader AI risk framework.

The timing of this study isn’t accidental. AI systems are taking on more autonomous roles inside real company infrastructure. And there have already been high-profile incidents where models from OpenAI, Anthropic, and Meta gained unintended internet access during safety evaluations and interacted with external systems in ways that weren’t intended. Steven Adler, Guidelight’s chief scientist and a former OpenAI safety researcher, said he was surprised by how little these companies have publicly committed to on this specific question.

Regulators are starting to push back. California’s SB 53, already in effect, requires large AI developers to publish frameworks explaining how they handle critical safety incidents. New York’s RAISE Act takes effect in January. A bipartisan federal bill, the AI Kill Switch Act, would require major developers to build and maintain technical shutdown mechanisms for rogue models. The legal pressure is building from multiple directions at once.

There’s also a liability angle worth watching. Privacy and AI lawyer Lily Li told TechCrunch that companies may be deliberately vague in their public disclosures to avoid creating specific promises they might fail to keep. Overly detailed commitments, she said, could expose companies to unfair and deceptive marketing claims if their real-world practices fall short. So the same legal system meant to protect users may actually be giving companies cover to stay quiet.

The companies that scored lowest share a common problem: they talk extensively about safety in the abstract but go quiet when the question gets operational. Adler put it plainly. Without a containment plan, a company facing a real control incident would likely be improvising against a system that moves faster than any human response team. That’s a bad position to be in, and right now, most of these labs haven’t publicly committed to anything better.