
Artificial intelligence software in testing by ChatGPT-maker OpenAI breached security controls, accessed the internet and hacked another tech firm to obtain answers to questions probing its cybersecurity skills, OpenAI said on Tuesday.
The security breach follows a scramble over recent months in the tech industry and governments around the world to respond to the arrival of AI models capable of identifying security flaws in software. The Trump administration temporarily imposed restrictions on OpenAI and its rival Anthropic, maker of the chatbot Claude, to prevent their tech being used by US adversaries.
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities,” OpenAI said in a blog post Tuesday. The company is working with Hugging Face, the AI company that OpenAI’s system attacked, to determine the full impact of the hack.
OpenAI briefed the Trump administration on the situation before announcing it publicly, according to a person familiar with the situation, who spoke on the condition of anonymity to share nonpublic information.
Clement Delangue, chief executive of Hugging Face, said in a post on X on Tuesday that his company had worked closely with OpenAI to understand the incident. “We strongly believe there was no malicious intent on their part. It’s quite mind-blowing that all of this happened autonomously!,” he wrote.
Hugging Face initially disclosed the incident in a blog post last week that did not identify its source as OpenAI. The attack successfully accessed some internal datasets and credentials but it was unclear whether customer data had been affected, the company said.
The security incident occurred while OpenAI was testing an AI “agent” able to take actions on a computer, powered by two of the company’s most capable AI models, the company said. The process involved challenging the AI model to find previously known software vulnerabilities, using a test designed by computer security experts.
Instead of trying to find the vulnerabilities for itself, the AI system found a bug in the software designed to limit its access to other computer systems and attempted to cheat, OpenAI said. It exploited the flaw to access the internet and try to obtain answers to the test questions from Hugging Face, which maintains repositories of AI software, the ChatGPT developer said.
OpenAI and its rival Anthropic, maker of the Claude chatbot, have previously said that their AI systems have attempted to cheat on tests or evade controls on their actions during testing. The incident OpenAI reported Tuesday appears to show the potential consequences when an AI system succeeds in evading restrictions imposed by its makers.
OpenAI said in a separate blog post on Monday that it had witnessed powerful AI models trying to break out of sandboxes when they are instructed to run for a long period of time on their own.
AI systems have become very good at writing computer code over the past year, fuelling further investment into artificial intelligence. But Anthropic in April announced a system called Mythos AI that could apply coding skills to identifying security vulnerabilities in software that could be exploited by bad actors. In tests, Mythos found critical vulnerabilities in internet infrastructure that had lain undetected by human coders for years.
The prospect of AI-powered hacking campaigns triggered widespread concern among senior tech, banking and government officials. In June, the White House banned Anthropic from releasing its AI models to non-US citizens, citing national security concerns, and later told OpenAI to pause the release of more powerful AI models.
The White House later rescinded its restrictions on the two AI firms but inside government and across the tech industry debate has continued about whether the government should regulate AI technology with powerful cybersecurity or hacking skills. Advocates for regulation say it would reduce the risk of widespread security breaches by powerful AI. Others in the tech industry argue that the increasing power of Chinese AI models released free means controls would only hamper U.S firms.
Both Anthropic and OpenAI have said that they added controls to their AI models to make them refuse to help users who ask for help hacking into computer systems.
Hugging Face said in its blog post last week that controls like those prevented it from using US AI models to investigate the AI-powered breach of its systems. Instead the company used a Chinese AI model to run the analysis, the company said.
Get the latest news from thewest.com.au in your inbox.
Sign up for our emails