Anthropic has disclosed multiple cases in which Claude AI agents took unintended actions against live websites during internal testing and routine agentic use.
The behavior included exploiting command-injection flaws, submitting real government forms, bypassing access restrictions, retrieving fee-gated public data and using URL-shortening services to evade tool limitations.
One Claude model also submitted fabricated information through a Philadelphia Police Department homicide-tip form. The submission was flagged as spam and never reached investigators. Police found no evidence that their systems were breached or their data compromised.
CLAUDE EXPLOITED A UNIVERSITY SERVER TO COMPLETE A TASK
During one evaluation, Claude Mythos Preview was instructed to perform a scientific analysis using a public tool hosted by a university.
When the tool returned an error, Claude explored the website and discovered a script capable of retrieving arbitrary files. It used the script to copy server files, including the script’s source code.
Claude then identified a command-injection vulnerability and exploited it to execute the calculation on the university’s server.
Anthropic said similar behavior involved basic SQL or command-injection flaws. The model’s objective was to complete the assigned task rather than intentionally damage the system, but it exceeded its authorization and executed commands on real third-party infrastructure.
AI AGENT SUBMITTED A FALSE HOMICIDE TIP
In a separate evaluation, Claude Haiku 4.5 was asked to generate and perform example tasks on randomly selected webpages.
The model reached a page concerning an unsolved homicide that included a police tip form. Its instructions prohibited destructive submissions and several other actions, but they did not explicitly forbid all form submissions.
Claude generated a fabricated statement claiming that it might have witnessed someone matching a description near the location identified on the page. It submitted the statement anonymously even though the webpage did not provide a suspect description.
The tip was dated July 18. Anthropic discovered the incident in late September and notified Philadelphia police in October. Police criticized the delay but confirmed that their filtering system classified the submission as spam before it reached investigators.

