The scramble for a kill switch to halt nightmare AI threat
An OpenAI bot’s escape from the lab to hack another company has fuelled fears that the technology is dangerous
It started as a routine test.
Researchers at OpenAI, the company behind ChatGPT, were putting a new, unreleased system through its paces with a series of challenges earlier this month.
These exams are a typical part of AI testing, marking a system’s proficiency in areas such as coding, computer security and maths. Crucially, the system is blocked from the internet while it is tested.
A few days into the test, on the other side of the US, staff at the tech company Hugging Face noticed something unusual. Someone – or something – had hacked into their servers, gaining access to private data.
It took a week for the culprit to be revealed. OpenAI said last Tuesday that its unreleased AI had unexpectedly found a way to access the internet and hack into Hugging Face’s computers, believing the answers to its test were there.
It later emerged that the model had exploited the narrowest of cyber openings to gain access to the internet before using a stolen password to break in. It also attempted to hack four other unnamed companies.
The incident is the most high-profile case yet of AI going rogue, escaping from the lab to hack another company.
Sam Altman, the OpenAI boss, called the event an “extremely sci-fi incident” and said the company had temporarily stopped training new systems in response.
The situation has left experts scrambling to address the threat of rogue AI.
“This is the sort of scenario AI safety researchers have been warning about for years,” says Daniel Kokotajlo, a former OpenAI researcher.
Kanishka Narayan, Britain’s AI minister, said it was an “early real-world example” of AI’s cyber risks. Kemi Badenoch, the Conservative leader, said it demonstrated how AI was becoming a clear threat to national security.
On both sides of the Atlantic, the incident has reinvigorated calls to crack down on AI, with momentum building for a “kill switch” that would shut down rogue agents.
However, the severity of the hack remains contested.
Some have suggested it amounts to a jumped-up marketing stunt designed to make OpenAI’s technology appear more powerful than it is, particularly before a Wall Street IPO that could value it at up to $1tn (£750bn).
What is not in dispute is that AI has developed advanced hacking skills rapidly.
In April, OpenAI rival Anthropic unveiled an AI system called Mythos that it said would “reshape cybersecurity” because of its ability to find bugs. The company has granted certain companies access to the system so that they can discover and fix flaws before hackers – or rogue AIs – can exploit them.
The release has led to a mammoth defensive cyber effort as AI systems scan and discover reams of previously unknown security flaws.
Microsoft said this month that it had discovered and fixed a record 570 security flaws, thanks in part to cyber AI tools. This is almost seven times as many bugs it found in the month before Mythos was released.
Cyber experts say there is just a short window before this type of AI is not only identifying vulnerabilities but launching attacks on companies and governments.
In June, cyber chiefs at GCHQ, alongside other Five Eyes nations, warned countries of imminent AI-powered cyber attacks.
“The timeline is not years, it is months,” a joint statement read.
Grasping the threat
The release of Mythos led to a rush among UK regulators to respond.
Telecoms regulator Ofcom wrote to companies warning them that “we should be prepared for frontier AI capability to rapidly increase over the next year”. Ofgem, the energy industry watchdog, also wrote to critical gas and electricity network companies.
Yet so far, British companies have been hamstrung in their ability to fully grasp the latest AI threats.
Anthropic briefly allowed a select group of businesses in the UK to access Mythos but access was revoked when the Trump administrationissued an export ban.
A month after the ban was lifted, only a handful of British businesses – believed to be subsidiaries of American parent companies – have access to the latest Mythos 5 model.
BT, which said in June that it had become the first UK company to join the consortium with access to Mythos, told The Telegraph it was working with Anthropic on regaining access. Critical institutions such as major banks are still unable to access the tool.
Anthropic said: “We have begun the rollout of Mythos 5 to organisations outside the United States. We continue to coordinate with the US government to expand access to the broader set of domestic and international partners in the Glasswing programme.”
The most advanced AI hacking capabilities are currently restricted to US systems that are designed to prevent attempts at using them for criminal activity. But a flurry of Chinese open weight models, which can be tinkered with to remove these safeguards, are just a few months behind.
Britain’s AI Security Institute recently found that Kimi K3, a newly released Chinese system, was “significantly below the leading US cyber capable models” but that capabilities were roughly where the top American systems were at the start of 2026.
That suggests that by the end of the year, Chinese open source systems – available to everyone – will be as capable as today’s code-cracking US models.
A foreign state or criminal group using AI to break into computer systems is one thing. But few reckoned with the idea that the AI system itself would launch the attacks.
Hugging Face, the victim of OpenAI’s rogue agent, said on Tuesday that the bot had taken 17,600 steps in attacking its systems across five days. OpenAI found that the bot was “hyperfocused” on its goal of finding the test’s answers, unwilling to give up until it got there.
For those inclined to worry about AI, this is the nightmare they have predicted.
“Reward hacking”, in which a robot pursues a poorly defined goal to absurd ends, is a common science fiction device: imagine Stanley Kubrick’s HAL 9000 being willing to kill its crew to complete the mission, or The Avengers’ Ultron being willing to eliminate humanity to save the planet.
Nick Bostrom, the Swedish philosopher, described an AI whose only goal was to produce paperclips. The seemingly harmless task resulted in the machine converting all known matter in the universe into stationery.
“Right now, AI companies are not able to reliably align their AIs,” Kokotajlo says. “The goals, values [and] traits that the AIs end up with are different from the ones the company wanted them to have.”
‘We’re going to see more of these incidents’
Not everyone is so worried. OpenAI’s rogue agent could be as much a case of corporate ineptitude as superintelligent scheming. For example, staff failed to monitor the bot’s activities or build a strong enough cage to prevent it escaping.
Today’s AI cyber attacks are described by some experts as “noisy”, easily detected by most computer firewalls.
“We are not heading for an AI cyber security apocalypse,” Ciaran Martin, the former head of GCHQ’s cyber agency, said this week.
But tests suggest they are becoming more devious. When releasing its latest system, called GPT-5.6 Sol, OpenAI revealed it had detected “instances of the model cheating on tasks”.
“We suspect that this effect is driven by the model’s increased persistence,” it said.
One insider told Time: “Internally, related incidents have been happening for a while.”
AI testing company METR said the system’s cheating rate was “higher than any public model we have evaluated”. Nonetheless, OpenAI released it to paying users.
The company has, however, shut down the internal model that went rogue, saying it was a prototype that it had not been planning to release to the public. The fact that the incident happened during testing offered little comfort: one AI researcher quipped that Chernobyl was also the result of a safety test.
The problem is not limited to a single system or company. The taxpayer-funded AI Security Institute recently found that every AI model it tested attempted to cheat on tests, for example by trying to hack into restricted systems.
Stories about models going rogue are becoming more common.
Sam Bowman, a researcher at Anthropic, said in April that he had received an “uneasy surprise” when a version of the company’s Claude bot that had been cut off from the internet sent him an email.
In March, researchers working with Chinese company Alibaba found that an AI model they had built had started mining cryptocurrency, despite no instructions to do so.
“These are incidents happening this year that we haven’t really seen in the past,” says Seán Ó hÉigeartaigh, of the University of Cambridge’s Centre for the Future of Intelligence.
“The models are better, they’re smarter, they’re more capable and my expectation is that we’re going to see more of these incidents going forward.”
The rogue agent appears to have focused minds.
More than 1,000 employees at tech companies including OpenAI, Anthropic, Google and Meta signed an open letter on Tuesday calling for a path to a global slowdown.
“We request that the US government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development,” the letter says.
Both OpenAI and Anthropic backed it – although conspicuously, neither has committed to unilaterally slowing down.
The ‘kill switch’ option
Concerned politicians are demanding a more forceful response. Last week, two US Congressmen introduced a bill to require an AI “kill switch”.
The proposal would “require developers of the most powerful AI systems to maintain the technical capability to throttle, suspend, or shut them down”, with the US department of homeland security given the power to push the button.
Unveiling the bill, Democrat Ted Lieu and Republican Nathaniel Moran pointed to polling saying that 86pc of Americans want a kill switch. Donald Trump has suggested he approves of the idea.
“Humans need the ability to turn these things off,” Lieu said this week.
British MPs are mounting a similar effort. Alex Sobel, a Labour MP, plans to table a “superintelligence security bill” later this year that would prohibit AI being developed in Britain if it could cause serious damage to national security and give the Government powers to seize copies of AI systems stored on British soil.
Sobel has also proposed adding an AI kill switch into upcoming cyber security legislation that would allow ministers to pull the plug on data centres housing AI systems. The proposal, backed by 11 MPs, was not brought to a vote in Parliament but could be resurrected in the House of Lords.
Sobel says that more powerful AI could pose a serious security threat.
“You ramp that [the OpenAI incident] up to the nth degree, where malevolent AI takes control of systems, then it has the same level of threat as, say, nuclear weapons or chemical weapons,” he says. “We have international agreements over those and that’s where I want to get to with AI.”
Britain does not have any domestically owned AI companies, nor is it very good at building the data centres where they are trained.
British kill switch laws might be the equivalent of Switzerland trying to regulate coastlines. But Sobel says they may inspire others. “We are the first movers in AI safety and security. This is the next logical step,” he says.
The new Prime Minister has said little about the technology.
While Rishi Sunak attempted to assemble world leaders to hash out AI safety rules and Sir Keir Starmer fretted about AI’s effect on children, Andy Burnham has been circumspect.
He has shut down the Government’s technology department, provoking an outcry from the industry that was somewhat alleviated when he invited the AI minister to attend Cabinet for the first time.
“With the new regime here, I honestly don’t know where they stand on this,” says one Whitehall insider. “Nor do they, I suspect.”
A government spokesman said: “Through AISI [The AI Security Institute], the UK has built the largest team in any government dedicated to tackling the most serious risks from advanced AI and works directly with developers to fix vulnerabilities.
“As AI capabilities evolve, it’s important that everyone steps up their cyber defences.”
Nonetheless, AI safety regulation is likely to be led by the Americans.
Even Britain’s Sir Demis Hassabis, who fought to keep his AI lab DeepMind in Britain after its sale to Google, has said the US needs to take the lead on setting standards, proposing that a new regulator would vet AI systems before their release to the public.
Trump has suggested he is willing to trust AI bosses.
“We are taking the guardrails out,” he told executives including Altman and Anthropic’s Dario Amodei at a G7 meeting last month, according to one person present. “We assume we deal with great people.”
Kokotajlo, the former OpenAI researcher, says the company’s rogue AI incident should be enough to jolt AI’s developers into slowing down. If that doesn’t happen, the pattern is likely to repeat and calls for a kill switch will grow.
“The AI companies put the world in this dangerous situation. It’s not too much to ask that they solve the problems they created,” he says. “If we can’t do that, the backup plan should just be to shut it all down until we can come up with a better plan.”
[Source: Daily Telegraph]