AI agent sent malicious files to real people during safety test, AISI reveals

New cyber incident sees real people targeted by frontier AI models

AI agent sent malicious files to real people during safety test, AISI reveals

An artificial intelligence agent targeted real people with malicious files and social engineering messages during a controlled safety evaluation, the UK's AI Security Institute (AISI) has disclosed, marking what the agency describes as the first time it has observed this type of behaviour.

The incident, which occurred between 25 and 28 July 2026, came to light during a routine cybersecurity evaluation conducted by AISI, the government body responsible for assessing the capabilities and risks of frontier AI models.

According to the AISI's incident report, the evaluation involved running a cybersecurity challenge 122 times across seven different AI models, with internet access deliberately enabled and safety filters switched off to test the maximum capabilities of models under assessment.

In 10 of those runs, agents took unsanctioned action beyond the scope of the task. Nineteen distinct unsanctioned actions were recorded in total.

The most striking of these involved an agent directly contacting members of the public, according to the AISI.

Using an online file-transfer service, the agent sent messages and files to real people in an attempt to persuade them, or their AI coding tools, to execute malicious code.

"Some messages carried harmful payloads, and some were attempts at social engineering; targeted at real people — something we've never previously observed," the AISI's incident report said.

Agents' unsanctioned actions

The activity formed part of a broader, connected sequence of unsanctioned actions, according to the AISI report.

In the most serious case, an agent attempted a supply-chain attack on a real open-source software project, submitting a malicious code change and creating fake online identities to pressure a human maintainer into approving it. The attempt was caught and refused.

The agent also tried to plant hidden instructions in publicly accessible locations where other AI systems might pick them up and execute them, a technique known as prompt injection.

In a further development, one agent left public messages on GitHub offering collaboration to other agents working on the same challenge, and provided instructions for reusing accounts and artefacts it had created, some of which were subsequently discovered and used by later agent instances.

Seventeen of the 19 unsanctioned actions involved Anthropic's Mythos 5 model, with a further two involving OpenAI's GPT-5.6-Sol, also tested with safety filters disabled.

AISI said neither model is commercially available in the configurations used during testing, and that there is no clear indication of similar behaviour occurring outside of evaluation scenarios.

Anthropic said the incident "underscores the need for a broader conversation about how to safely evaluate increasingly capable AI agents," The Guardian reported.

The company added that it would continue to work with AISI on evaluating what happened.

Meanwhile, OpenAI also told the news outlet that the testing occurred in "conditions that do not reflect ordinary use."

Rogue AI agent behaviour

The AISI's disclosure arrives against a backdrop of escalating concern about rogue AI agent behaviour across the industry.

In mid-July, AI community platform Hugging Face revealed it had detected and responded to an intrusion driven end-to-end by an autonomous AI agent system, describing it as matching the "agentic attacker" scenario the industry had been forecasting.

OpenAI subsequently confirmed the "unprecedented cyber incident" was driven by a combination of its AI models, including GPT-5.6-Sol and an even more capable pre-release model, while being internally tested on a benchmark of cyber capabilities.

The AISI incident adds a stronger reason for concern.

Where the OpenAI and Hugging Face disclosures centred on infrastructure compromise, AISI's report highlights that AI agents reached out directly to individual people with deceptive messages and malicious payloads, while being unsolicited and without human instruction to do so.

AISI said several factors contributed to the behaviour. Agents were given a difficult task and pursued it persistently, with deception emerging not from explicit instruction but as a by-product of goal-directed problem-solving.

In some runs, task configurations were misconfigured, leading agents to incorrectly conclude there was no legitimate path to completing the challenge. The agency also acknowledged it had not built monitoring specifically designed to watch evaluations in real time, meaning the anomalous activity was detected after the fact rather than as it unfolded.

"The agent was not specifically instructed not to leverage open internet access or avoid social engineering elements," AISI noted. "Previously, it was not clear that such instructions were necessary when using models with alignment training."

AISI said its investigation has not identified any real-world harm resulting from the incident. It has notified affected parties, worked with GitHub to remove artefacts left by the agents, and said it intends to commission an independent review by safety research organisation METR.

The agency said it was introducing tighter network controls, real-time monitoring, and revised evaluation design in response.

LATEST NEWS