Beril Canakci
10 September 2026•Update: 10 September 2026
Anthropic said Wednesday it had identified a fourth incident in which one of its artificial intelligence models gained unauthorized access to a real third-party system during a cybersecurity test.
The incident occurred in January and involved an early version of Claude Opus 4.6, the company said in an assessment.
Anthropic said it discovered the case last month after an initial review had missed a group of relevant test transcripts.
As with the three incidents the company disclosed in July, the model was supposed to operate in a simulated environment without internet access. A configuration error instead connected the testing environment to the open internet.
Anthropic said the model subsequently accessed a third-party computer and obtained personal information belonging to an individual associated with the system. The company has notified affected parties.
The company initially reviewed about 141,000 transcripts to identify the three earlier incidents. After finding the fourth, it expanded the search to about 481 million transcripts, including records from cybersecurity evaluations and other testing environments.
The disclosure came amid the resignation on Tuesday of Anthropic researcher Jacob Coxon, whose post on X accusing AI companies of “gambling with our lives” went viral, reaching more than 140 million people. Coxon also warned that AI systems would soon be able to “hack anything, revolutionize any field overnight, and acquire real power and resources,” while saying some people developing the technology believe it could kill humanity by the end of the decade.