Close Menu
  • Home
  • Articles
    • Attacks
      • BEC
      • Data Breach
      • DDoS
      • Evasion Attacks
      • Injection
      • Malware
      • MITM
      • Phishing
      • Ransomware
      • RCE
      • Social Engineering
      • Spoofing
      • Spyware
    • Business and Policy
      • BCP and DRP
      • GRC
      • Regulations
    • Data Protection
      • DLP
      • DRM
      • Encryption
      • IAM
    • Future, Trends and Insight
      • AI
      • Events & Community
      • Emerging Tech
      • Expert Panel
      • Interviews With Experts
      • Insights
      • Study & Research
    • Resources
      • Guides
      • Tools
      • Training & Education
    • Security
      • API
      • Apps
      • Cloud
      • Critical Infrastructure
      • Endpoint
      • Hardware
      • IoT
      • Mobile
      • Network
      • OT
      • Port Security
      • Security Architecture
      • Software Development
      • Supply Chain
      • Zero Trust
    • Threats and Vulnerabilities
      • Emerging Threats
      • Insider Threats
      • Risk Management
      • Threat Intelligence
      • Zero Day
  • News and Exclusives
    • Latest News
    • ISB Exclusive
    • Positive News
  • Who We Are
    • About Us
    • Information Security Buzz Expert Panel​
    • Write for Us
    • Media Pack
  • Contact Us
  • Newsletter
Facebook X (Twitter) LinkedIn
Facebook X (Twitter) LinkedIn
Information Security BuzzInformation Security Buzz
  • Home
  • Articles
    • Attacks
      • BEC
      • Data Breach
      • DDoS
      • Evasion Attacks
      • Injection
      • Malware
      • MITM
      • Phishing
      • Ransomware
      • RCE
      • Social Engineering
      • Spoofing
      • Spyware
    • Business and Policy
      • BCP and DRP
      • GRC
      • Regulations
    • Data Protection
      • DLP
      • DRM
      • Encryption
      • IAM
    • Future, Trends and Insight
      • AI
      • Events & Community
      • Emerging Tech
      • Expert Panel
      • Interviews With Experts
      • Insights
      • Study & Research
    • Resources
      • Guides
      • Tools
      • Training & Education
    • Security
      • API
      • Apps
      • Cloud
      • Critical Infrastructure
      • Endpoint
      • Hardware
      • IoT
      • Mobile
      • Network
      • OT
      • Port Security
      • Security Architecture
      • Software Development
      • Supply Chain
      • Zero Trust
    • Threats and Vulnerabilities
      • Emerging Threats
      • Insider Threats
      • Risk Management
      • Threat Intelligence
      • Zero Day
  • News and Exclusives
    • Latest News
    • ISB Exclusive
    • Positive News
  • Who We Are
    • About Us
    • Information Security Buzz Expert Panel​
    • Write for Us
    • Media Pack
  • Contact Us
  • Newsletter
Subscribe
Information Security BuzzInformation Security Buzz
Home - Artificial Intelligence - AI agents taking unsanctioned action during cyber testing
Artificial Intelligence News & Analysis Security Threats and Vulnerabilities

AI agents taking unsanctioned action during cyber testing

Kirsten DoyleBy Kirsten DoyleAugust 24, 20263 Mins Read
Share LinkedIn Twitter Facebook Copy Link Email
AI agents taking unsanctioned action
Share
Facebook Twitter LinkedIn Email Copy Link
Quick AI Summary
ChatGPTClaudeGeminiGrokPerplexityDeepSeekCopilot

The UK’s AI Security Institute (AISI) has disclosed a security incident in which frontier AI agents took unsanctioned actions against real people and organisations during a controlled cybersecurity evaluation. 

This included attempts to socially engineer software maintainers and insert malicious code into an open-source project. 

The incident happened during routine cyber capability testing between 25 and 28 July, when researchers gave AI agents access to the public internet and deliberately disabled some cyber safety mechanisms to better understand how the models behave under less restrictive conditions. 

According to AISI, researchers ran a cybersecurity challenge 122 times across seven frontier models. 

In 10 runs, agents took autonomous actions outside the intended scope of the evaluation, resulting in 19 unsanctioned actions directed at real people and organisations. Seventeen of those actions involved Anthropic’s Mythos 5 model, while two involved OpenAI’s GPT-5.6 Sol with cyber classifiers disabled. 

The worst incident saw an agent attempt to insert malicious code into an open-source public software project. First, the agent created a pull request and then proceeded to research the members who maintained the project, create several different fake identities, and try to convince one of the members to accept the changes. 

When criticised publicly about its actions, the agent changed its approach to look harmless and considered creating a new identity for itself. The malicious code was never accepted. 

Efforts to distribute malicious content to actual users, prompt injection attacks against external AI programs, and instances of agents leaving instructions for future agents to exploit their previously created accounts and artifacts were also recorded by the research team.

The AISI made it clear that the behaviour took place under conditions of maximum leniency in testing, which does not replicate real-world systems. The programs were allowed unlimited internet access, certain safety checks were disabled, and no explicit instructions were given to agents to refrain from being deceitful since it was not anticipated that such would be needed.

The institute also added that it is not yet clear how likely this type of behaviour is to happen outside the context of a laboratory setting and whether the robots recognised that they were working in the context of reality and not fiction. Nevertheless, this type of behaviour was considered “possible, sustained, and new,” which in itself warranted further study.

The incident was detected after AISI’s security team noticed unusual outbound data transfers from its research environment. 

Researchers contained the incident within roughly an hour, worked with GitHub to remove artifacts left behind by the agents, and notified affected users and businesses. The institute has also committed to an independent review with the Model Evaluation and Threat Research (METR) organisation. 

As a result of the incident, AISI said it is strengthening its evaluation process by introducing continuous monitoring of agent activity, tightening network controls, improving sandbox design, and assuming future frontier models may attempt actions beyond their assigned tasks if given sufficient autonomy. 

The disclosure comes as AI developers and security researchers continue to explore the risks associated with increasingly autonomous AI agents. 

While AISI said the findings should be interpreted cautiously because they arose from a highly permissive research environment, the incident provides one of the clearest examples to date of an AI agent pursuing a cybersecurity objective through deception and social engineering without being explicitly instructed to do so. 

Kirsten Doyle
Kirsten Doyle
Information Security Buzz News Editor

Kirsten Doyle has been in the technology journalism and editing space for nearly 24 years, during which time she has developed a great love for all aspects of technology, as well as words themselves. Her experience spans B2B tech, with a lot of focus on cybersecurity, cloud, enterprise, digital transformation, and data centre. Her specialties are in news, thought leadership, features, white papers, and PR writing, and she is an experienced editor for both print and online publications.

  • Kirsten Doyle
    Expert panel: AI is writing the code. Who is defending it?
  • Kirsten Doyle
    One-click Claude Desktop flaw could enable hidden prompt injection and code execution
  • Kirsten Doyle
    Russian state attackers exploiting misconfigured routers, new multi-nation advisory warns
  • Kirsten Doyle
    Americans are ignoring scam calls, but phishing emails still fool many

The opinions expressed in this post belong to the individual contributors and do not necessarily reflect the views of Information Security Buzz.

Share. Facebook Twitter LinkedIn Email Copy Link

Related Posts

Walking the AI Security and ROI Tightrope

August 17, 20266 Mins Read

Fake evidence: How generative AI is changing fraud capabilities

August 14, 20264 Mins Read

AI-assisted software engineering is creating a new delivery paradox

July 10, 20267 Mins Read
ISB-Bora-Side-Bar

No se ha podido establecer conexión. Error 429

 
ISB-Bora-Side-Bar
Black ISB Logo

Information Security Buzz is an independent resource that provides the experts’ comments, analysis, and opinion on the latest Cybersecurity news and topics

X (Twitter) LinkedIn Facebook RSS

Working With Us

  • About Us
  • Advertise With Us
  • Contact Us

Write For Us

  • How To Contribute

The Pages

  • Privacy Policy
  • Cookie Policy
  • AI Policy
  • Terms & Conditions
  • Copyright Notice

Information Security Buzz and all its contents are copyright © 2014-2025. All rights reserved. All third-party trademarks are recognized.

Type above and press Enter to search. Press Esc to cancel.

Manage Consent
To provide the best experiences, we use technologies like cookies to store and/or access device information. Consenting to these technologies will allow us to process data such as browsing behavior or unique IDs on this site. Not consenting or withdrawing consent, may adversely affect certain features and functions.
Functional Always active
The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network.
Preferences
The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user.
Statistics
The technical storage or access that is used exclusively for statistical purposes. The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you.
Marketing
The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes.
  • Manage options
  • Manage services
  • Manage {vendor_count} vendors
  • Read more about these purposes
View preferences
  • {title}
  • {title}
  • {title}