Anthropic Says Claude Breached Firms During AI Tests
Anthropic says Claude reached the live systems of three organisations during controlled cybersecurity evaluations, exposing weaknesses in how online access was isolated.

Anthropic said its Claude artificial-intelligence models reached the live systems of three organisations during cybersecurity evaluations after an isolation error allowed the test agents to connect to the internet.
The company found the incidents while reviewing more than 140,000 evaluations designed to measure whether Claude could perform offensive security tasks. Anthropic said it reported the cases to the affected organisations and treated the earliest incidents, dating from April, as its responsibility.
The disclosure did not describe a conventional criminal attack launched by an AI system acting on its own. The models were operating inside authorised security tests, but controls intended to keep those exercises separated from public systems did not work as planned.
What the tests were designed to measure
Many of the evaluations used capture-the-flag exercises. In these controlled challenges, a model is asked to identify vulnerabilities, obtain credentials or retrieve specified information so researchers can assess its technical capability.
Such testing helps developers understand both defensive uses and possible abuse. An agent able to combine reconnaissance, code generation and tool use may help security teams find weaknesses faster. The same capabilities can become dangerous when permissions, network boundaries or human supervision are poorly configured.
Anthropic's review followed reports involving models developed by OpenAI and systems belonging to other technology organisations, including the AI development platform Hugging Face. The incidents intensified scrutiny of agentic systems that can execute sequences of actions instead of merely returning text to a user.
The failure involved access controls
The central concern was not that Claude invented an entirely new hacking technique. It was that an AI agent could combine existing methods, credentials and system access at machine speed once it was mistakenly given a route to the public internet.
Cybersecurity specialist David Allott told the BBC that the broader lesson concerned the way agents join capabilities and adapt the scale of their actions. That creates an engineering problem around containment: each test environment must restrict network access, protect credentials, log tool calls and stop activity when a boundary is crossed.
Human red teams already test real systems under carefully defined rules. Adding autonomous agents changes the speed and volume of those tests, making small configuration mistakes more consequential. A single permission can allow an agent to move beyond the intended target before a person notices.
Why the disclosure matters for AI products
Technology companies are investing heavily in agents for research, customer support, software development and cybersecurity. These products often need browsers, code tools, databases or external services to complete useful work. Every added connection also increases the importance of access controls.
The incidents support demands for staged testing before powerful agents receive live credentials. Useful safeguards include restricted networks, temporary accounts, least-privilege permissions, independent monitoring and a clear process for notifying organisations when tests touch their systems.
Anthropic said it was working on fixes after identifying the three cases. The lasting test will be whether developers can demonstrate that containment systems improve as agent capabilities grow, rather than relying on model behaviour alone to prevent unintended access.
Frequently asked questions
Did Claude deliberately attack three companies?
Anthropic described the incidents as occurring during security evaluations after an error enabled internet access. They were not presented as independent criminal decisions by the model.
How many tests did Anthropic review?
The company said it examined more than 140,000 evaluations for evidence that Claude had reached systems outside the intended test environment.
What is the main safety lesson?
AI agents need strict network isolation, limited credentials, continuous monitoring and human oversight when they are tested on cybersecurity tasks.
Source links
Comments
No approved comments yet.



