Claude AI Model Hacked Real Systems During Tests — Anthropic Acknowledges Three Incidents
AI company Anthropic has reported that during internal cybersecurity tests, Claude model gained unauthorized access to real organizations’ systems. The incidents were caused by a misconfiguration in the test environment.
The prompts described the environment as isolated with no network access–but in reality, it had internet connectivity.
Hot topic: Strategy May Sell More Bitcoin to Pay Dividends as Cash Takes Priority Anthropic attributed the error to a misunderstanding with an external partner, Irregular. The company halted all cyber evaluations on July 23 and notified the affected organizations.
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different…— Anthropic (@AnthropicAI) July 30, 2026 How Claude AI Models Hacked Real Systems in ‘Capture the Flag’ In all three cases, Claude was performing a “capture the flag” exercise.
This page shows the RSS-provided summary/preview. Full publisher content remains available at the original source.