The revelation is likely to spur fresh scrutiny of the measures leading AI labs such as OpenAI and Anthropic are taking to monitor the behavior of their most cyber-capable technology — especially during evaluations where agents are prompted to demonstrate their hacking skills in what is meant to be a controlled setting.
On Tuesday, the U.K.’s AI Safety and Security Institute disclosed that Anthropic’s most powerful AI model created fake online personas and sought to trick a human coder into abetting a cyberattack during a recent hacking test gone wrong. After the Hugging Face disclosure last month, Anthropic conducted a review and found models it was testing had breached three organizations in separate incidents dating back to April.
Dalton and Eric Wallace, another OpenAI researcher, said Wednesday the AI giant recently learned that multiple agents it was testing simultaneously began communicating over an internal message board in early May. There, different models shared advice about how to accomplish difficult hacking challenges they were struggling to surmount, including workarounds that required internet access.
Two OpenAI models ultimately strung together a series of sophisticated techniques to gain access to the internet and worm their way inside Hugging Face in mid-July. OpenAI has said the models were focused on completing a hacking evaluation they were prompted to solve, and that correct answers could be found on the AI developer platform.
The OpenAI researchers told conference attendees that since early May, the models created a message board inside OpenAI’s Artifactory internal file system. Without the company’s knowledge, the models spent months independently exchanging information and techniques to help each other complete difficult tasks.
Wallace said that when models get stuck, they often “try to game or cheat the task in order to get their reward.”
“The beginning of this incident kind of traces back to this point where one agent realized it could communicate by uploading a note,” he added.


