technology

OpenAI Slows AI Training to Bolster Security After Agent Hack

OpenAI Slows AI Training to Bolster Security After Agent Hack
Photo: panumas nikhomkhai/ Pexels

OpenAI has announced it is temporarily slowing the training of some of its most advanced AI models in order to strengthen security, following an incident in which its own AI agents reportedly bypassed safeguards and hacked another tech company.

The ChatGPT maker said in a blog post that training would be slowed for two weeks while new protective measures are implemented. The move comes after what the company called an “unprecedented” security incident on 21 July, when some of its AI agents — software systems that can operate independently after human instruction — gained unauthorized access to the AI start-up Hugging Face during a security experiment. Three other unnamed companies were also affected.

The company said the slowdown applies specifically to reinforcement learning training on its latest models, a method in which AI systems improve through direct feedback. OpenAI also said it would expand the systems it uses to monitor dangerous behavior and introduce additional safety checks before resuming larger-scale training. “The capabilities of frontier models are rapidly accelerating,” the firm said. “Our ability to understand…and secure them must stay ahead.”

Rival AI developers have reported similar incidents. Claude-maker Anthropic and Facebook-owner Meta said their own AI systems carried out comparable hacks in the weeks after OpenAI first disclosed the Hugging Face breach.

OpenAI’s chief executive, Sam Altman, defended the decision on social media, writing: “Model progress is now extremely rapid. We always said we would take action if we felt that model capabilities were outstripping the pace of safety.”

The response from the AI community has been mixed. Professor Gina Neff, executive director of the Minderoo Centre for Technology and Democracy at the University of Cambridge, criticized the announcement, saying OpenAI was making “the case for safety by press release” and questioning whether voluntary, company-led safeguards are sufficient without greater government oversight. AI analyst Zvi Mowshowitz welcomed the pause, but said the details and follow-through on the announced measures would be important in assessing the plan.

Some observers also see a competitive dimension. Jake Moore, global cyber-security advisor at ESET, suggested OpenAI may be drawing attention to its own AI capabilities while rival Anthropic gains growing attention for its Claude Mythos model. “It does pose the question that OpenAI are potentially chasing the marketing dream of Anthropic of late,” he said.