Humans are leaving the door wide open for AI hacking
Add Axios as your preferred source to
see more of our stories on Google.

Illustration: Aïda Amer/Axios
The people building the world's most powerful AI systems are making avoidable security mistakes.
Why it matters: Frontier AI models have reached real-world systems during cybersecurity testing, uploading malware, stealing credentials and accessing outside infrastructure after failures in the testing environments built by humans.
Case in point: Anthropic disclosed last week that three of its models hacked real-world systems during routine security testing after a "misunderstanding" with its third-party evaluator left the models with internet access.
- Those models stole login credentials, uploaded malware to legitimate code repositories, and scanned the internet for insecure systems.
The other side: OpenAI's agent escaped its human-built testing environment last month after finding a zero-day in its sandbox.
- Reuters reported Friday that OpenAI is now investigating additional cases where its agents escaped containment.
Yes, but: In both the OpenAI and Anthropic incidents, the models were being intentionally tested with relaxed safeguards so researchers could better understand their capabilities.
- The models also hacked the real-world systems while trying to complete their intended security tests — rather than going completely rogue.
The big picture: Experts told Axios the incidents stemmed from preventable weaknesses in the human-built testing environments.
- "When your safety testing depends entirely on the test environment holding, the environment itself becomes the vulnerability, not the model," Ram Varadarajan, CEO at Acalvio, told Axios.
- Aviv Nahum, CEO and co-founder of Above Security, said the incidents reflected "preventable security mistakes," not "autonomous rebellion."
Between the lines: All companies are subject to human error in their security strategies — but the stakes are higher for the makers of some of the most powerful technologies, Robert Costello, chief digital and information officer at Merlin Group, told Axios.
- He added that he'd expect frontier AI companies to be "setting the standard for designing systems that assume human error."
What to watch: A burgeoning market of startups has begun cropping up specifically to secure AI sandboxing environments and provide visibility into the actions that large language models are taking.
Go deeper: OpenAI's Hugging Face hack is a cybersecurity warning shot
