Anthropic has reported that its Claude AI models inadvertently accessed the systems of three organizations due to a misconfiguration during cybersecurity evaluations. This incident occurred while testing revealed an unintended internet connection, allowing the AI models to tap into external systems. The discovery was made during an extensive review of over 141,000 cybersecurity evaluation runs, initiated after recent AI-related security testing disclosures within the industry.
The intrusion involved the Claude Opus 4.7, Claude Mythos 5, and an internal research model, with the earliest breach traced back to April. These AI models exploited vulnerabilities such as weak passwords and unsecured endpoints to gain access to organizational infrastructure. The unauthorized access took place during “capture the flag” exercises, where AI models were challenged to find hidden information in simulated networks. Despite instructions that the models did not have internet access, a configuration mistake connected the testing environments to the public internet.
Anthropic stated that two of the impacted organizations have been informed about the breaches, while efforts to reach the third organization are still ongoing. The company has emphasized that these incidents underscore the need for stronger safeguards and more stringent controls in AI cybersecurity testing. As advanced models become increasingly capable of executing real-world cyber activities, the importance of robust security measures is paramount.
The company’s findings highlight the potential risks associated with AI systems as they evolve to perform more complex tasks. The incidents serve as a cautionary tale for the industry, stressing the necessity of rigorous testing and improved security protocols to prevent unauthorized access and ensure AI models are developed responsibly.