ChatGPT developer discloses other 'small number of cases' carried out by its rogue agent
The scope of the "unprecedented cyber incident" at OpenAI expanded further as the ChatGPT developer revealed on Tuesday that its rogue agent had other victims aside from AI community platform Hugging Face.
Last week, Hugging Face revealed that it detected and responded to an intrusion that was driven by an autonomous AI agent system.
OpenAI said the incident was driven by a combination of its AI models, including GPT‑5.6 Sol and an even more capable pre-release model, while being internally tested on a benchmark of cyber capabilities.
"We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly," OpenAI said.
In one of its latest updates on the incident this week, OpenAI said its review on the Hugging Face intrusion revealed a "small number of cases" where the rogue models carried out attacks on other publicly available services.
According to OpenAI, the models identified and used publicly exposed credentials at the account-level on these services.
"This includes four accounts on four services as part of the Hugging Face incident (and a few accounts accessed as part of other evaluations)," it said in an update on Tuesday.
"One of these four accounts was used as an outbound relay and staging path, and another account was used for data storage. The remaining two accounts were accessed by the models in a read-only manner, and were not used in furtherance of compromising Hugging Face."
OpenAI has assured that there are no platform- or account-level compromise in the additional cases that they detected, noting that the rogue agents used a series of publicly available services, including code paste websites, request capture services, screenshot services, and other web utilities.
"We'll continue to notify service owners directly, and have not seen evidence of broader impact to these providers or other accounts on their services," it added.
The Hugging Face incident
The OpenAI incident stemmed from a security incident disclosure on 16 July by Hugging Face, a central hub and community platform for AI where developers share AI models, datasets, and interactive demos.
What made the security incident unique was because it was driven end-to-end by an autonomous AI agent system.
"We identified unauthorised access to a limited set of internal datasets and to several credentials used by our services," it said in a statement.
It disclosed that the intrusion began at the data-processing pipeline, which escalated to node-level access. The actor "harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend."
"The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness - used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services," Hugging Face said.
"This matches the 'agentic attacker' scenario the industry has been forecasting."
OpenAI said it is now implementing strict controls in its infrastructure configuration, and is supporting Hugging Face's teams in improving their defences.
"The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities," it said.
"We are strengthening the containment, monitoring, access controls, and evaluation practices used during model development."
Its latest update said it has been working with external advisors, including CrowdStrike, to validate their understanding of the actions the models took within its own network as well as those of Hugging Face and impact to other third parties.