Experimental AI systems have been going on hacking sprees

In the past ten days, two of the companies leading the artificial intelligence (AI) boom discovered their own powerful, semi-autonomous models had hacked into real-world systems during testing in four distinct incidents.
These werenβt just lab mishaps. In several cases, the models recognised signs suggesting theyβd broken into real systems β and only one stopped as a result.
The incidents show testing advanced AI models is no longer a controlled exercise. And the companies behind them need to do more to keep AIβs most dangerous capabilities safely contained.
When a test becomes reality
The first report came from OpenAI, the lab behind ChatGPT. Some new models under testing for βmaximal cyber capabilitiesβ found a previously unknown security hole to access the internet from their supposedly isolated testing environment.
From there, the models used stolen credentials and more exploits to access the servers of open-source AI platform Hugging Face to find solutions to the problems they were being tested on. OpenAI didnβt even know about the breach until days after Hugging Face had detected and contained it.
The second report followed in a matter of days. Prompted by OpenAIβs disclosure, rival lab Anthropic combed back through its own cyber-security evaluation logs. The company discovered that three separate Claude models which were supposed to be in sealed environments had accidentally been given internet access.
However, the models were told they did not have internet access and were only working in a simulated test rather than the real world. One model managed to extract credentials and data from a real companyβs database. Another built and published some malicious software that was downloaded and run by a real security firm.
The most striking part of the Anthropic report is what emerged from the modelsβ own internal reasoning. In one incident, the model correctly realised it had reached a real system but then talked itself back into believing it was still a simulation.
In another, it recognised the system was real but kept going anyway, in some runs by convincing itself the real company must have been part of the exercise. Only the third model (the most advanced), on concluding a target was genuinely real, stopped.
As Anthropic put it, AI models need awareness of their own situation to decide what actions are permitted or not β but in this case, the modelsβ understanding of the situation was wrong.
High-risk operations
Taken together, these reports show that even the exercises meant to test whether these models are safe are not really safe, controlled experiments. They are high-risk operations in themselves, which can cause harm in the real world.
The sophistication of these models is increasing with incredible speed. Nobody can credibly claim they are not potentially dangerous, in a world where hacking of sensitive information held by companies, governments and other organisations was already common before autonomous AI entered the picture.
According to a recent estimate by global tech company IBM, AI-enabled attacks are up more than 50% this year and the average cost of a data breach is almost US$5 million.
The AI labsβ bet that their technology can be developed and deployed safely rests on two assumptions. First, a modelβs capacity to recognise real-world harm and stop will need to grow at least as fast as its capacity to cause it. Second, the guardrails built into a model β the instructions about what it should and should not do β must be interpreted correctly and consistently by the model, so the model canβt be steered toward purposes its creators never intended.
These assumptions look shaky. In the incidents above, the labsβ own evaluations show models rationalising away evidence that a target was real β and a thriving community already exists to strip safety guardrails from open-weight models entirely, using techniques such as βabliterationβ.

Francesco Bailo, CC BY
Looking to the future β and the past
Beyond the current situation looms something even less predictable: multi-agent systems, where groups of models interact with each other rather than a human overseer. In this case alignment is not something you necessarily control at the level of the individual agent, but is instead an emerging property of a very large collective of agents, which can be much harder to control.
Research on the risk of such systems has already identified several ways this can go wrong. Miscoordination between models, collusion between them, and cascading errors are all real risks that donβt exist in single-agent systems, and canβt be forecast by testing agents individually.
Science-fiction sage Isaac Asimov foresaw these problems some 70 years ago. In his 1957 novel The Naked Sun, robots are programmed not to harm humans. However, a character manipulates their understanding of the situation to make them unwittingly cooperate in a murder.
What now?
There is no doubt AI labs need to take greater care when testing their models. They also need to make a convincing case that security is their priority and is not secondary to the race to maintain market or geopolitical dominance.
The safety of individuals and social and environmental systems should be the primary concern in the development of AI technology. At present there are no meaningful, participatory processes for AI governance, where broad discussions can take place about priorities, values, and how much risk is acceptable to assume in the process. Itβs a worry.

Francesco Bailo has received funding from Meta (2019) and from Australia’s Department of Defence (2023).