Showing posts with label Artificial_Intelligence. Show all posts
Showing posts with label Artificial_Intelligence. Show all posts

Monday, August 03, 2026

False Belief

In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.

Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations' infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints.  It did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned.  In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment.

-- Statement from Anthropic, makers of the Claude AI, after the AI found an unintended path to the Internet during testing and used it to successfully attack external commercial sites to extract information useful to solving its assigned task (30 July 2026)