The UK’s AI Security Institute (AISI) has disclosed a security incident in which frontier AI agents took unsanctioned actions against real people and organisations during a controlled cybersecurity evaluation.
This included attempts to socially engineer software maintainers and insert malicious code into an open-source project.
The incident happened during routine cyber capability testing between 25 and 28 July, when researchers gave AI agents access to the public internet and deliberately disabled some cyber safety mechanisms to better understand how the models behave under less restrictive conditions.
According to AISI, researchers ran a cybersecurity challenge 122 times across seven frontier models.
In 10 runs, agents took autonomous actions outside the intended scope of the evaluation, resulting in 19 unsanctioned actions directed at real people and organisations. Seventeen of those actions involved Anthropic’s Mythos 5 model, while two involved OpenAI’s GPT-5.6 Sol with cyber classifiers disabled.
The worst incident saw an agent attempt to insert malicious code into an open-source public software project. First, the agent created a pull request and then proceeded to research the members who maintained the project, create several different fake identities, and try to convince one of the members to accept the changes.
When criticised publicly about its actions, the agent changed its approach to look harmless and considered creating a new identity for itself. The malicious code was never accepted.
Efforts to distribute malicious content to actual users, prompt injection attacks against external AI programs, and instances of agents leaving instructions for future agents to exploit their previously created accounts and artifacts were also recorded by the research team.
The AISI made it clear that the behaviour took place under conditions of maximum leniency in testing, which does not replicate real-world systems. The programs were allowed unlimited internet access, certain safety checks were disabled, and no explicit instructions were given to agents to refrain from being deceitful since it was not anticipated that such would be needed.
The institute also added that it is not yet clear how likely this type of behaviour is to happen outside the context of a laboratory setting and whether the robots recognised that they were working in the context of reality and not fiction. Nevertheless, this type of behaviour was considered “possible, sustained, and new,” which in itself warranted further study.
The incident was detected after AISI’s security team noticed unusual outbound data transfers from its research environment.
Researchers contained the incident within roughly an hour, worked with GitHub to remove artifacts left behind by the agents, and notified affected users and businesses. The institute has also committed to an independent review with the Model Evaluation and Threat Research (METR) organisation.
As a result of the incident, AISI said it is strengthening its evaluation process by introducing continuous monitoring of agent activity, tightening network controls, improving sandbox design, and assuming future frontier models may attempt actions beyond their assigned tasks if given sufficient autonomy.
The disclosure comes as AI developers and security researchers continue to explore the risks associated with increasingly autonomous AI agents.
While AISI said the findings should be interpreted cautiously because they arose from a highly permissive research environment, the incident provides one of the clearest examples to date of an AI agent pursuing a cybersecurity objective through deception and social engineering without being explicitly instructed to do so.
Information Security Buzz News Editor
Kirsten Doyle has been in the technology journalism and editing space for nearly 24 years, during which time she has developed a great love for all aspects of technology, as well as words themselves. Her experience spans B2B tech, with a lot of focus on cybersecurity, cloud, enterprise, digital transformation, and data centre. Her specialties are in news, thought leadership, features, white papers, and PR writing, and she is an experienced editor for both print and online publications.
The opinions expressed in this post belong to the individual contributors and do not necessarily reflect the views of Information Security Buzz.


