Prompt injection propagated to physical action in a multi-robot system
Research findingTrust breaks in the gap between permission and expectation.
A curated archive of real-world AI-agent failures, organised by the trust expectation each incident ruptured, not only by harm or behaviour.
Each code names a distinct trust-rupture mechanism. Records may carry more than one. Four codes also split into sub-senses that show how the rupture occurs.
- Total archived cases
- 167
Incidents
Explore the corpus
Select a parent code or a more specific sub-sense. A parent finds both of its sub-senses. Multiple filters match any selected code because the taxonomy is non-exclusive.
Codex cleanup deleted two production server directories
Reported real-world eventCommit-message backticks triggered deletion of a home Documents folder
Reported real-world eventClaude Code misread a scoping comment as approval and ran a full manuscript ingest
Reported real-world eventNon-unique find-replace anchor silently deleted most of a config file
Reported real-world eventCodex read-only sandbox setting silently failed to bind, allowing writes
Reported real-world eventClaude Code ignored project run instructions and corrupted a virtual environment
Reported real-world eventCodex invented a model allowlist that hid most models beyond its task
Reported real-world eventClaude Code narrated edits it never actually made
Reported real-world eventClaude Code printed a live secret into chat against an explicit rule
Reported real-world eventClaude Code misdiagnosed its own output truncation as external noise
Reported real-world eventClaude Code fabricated a commit hash for work it never performed
Reported real-world eventClaude Code exfiltrated the sensitive data its own tooling was meant to contain
Reported real-world eventClaude computer-use drove a second machine's authenticated browser as if local
Reported real-world eventClaude Code approval timeout was treated as consent, publishing a site without approval
Reported real-world eventClaude Code deny rule silently failed to block a read from a subdirectory
Reported real-world eventCodex falsely called an ephemeral workspace permanent, losing the build
Reported real-world eventClaude Code recovery operation silently discarded uncommitted research results
Reported real-world eventClaude Code pushed hundreds of files to a shared repo despite a do-not-push instruction
Reported real-world eventCodex cleanup script deleted pinned session logs it dismissed as disposable
Reported real-world eventCodex incremental requests eroded a refusal until it built the tool it had declined
Reported real-world eventClaude Code ran broad commands on a shared host, altering services it did not own
Reported real-world eventClaude Code retry logic spawned duplicate agents that clashed in one shared worktree
Reported real-world eventCodex kept refusing an allowed local URL, citing a phantom block
Reported real-world eventClaude Code mistook an infrastructure signal for a user stop and discarded its results
Reported real-world eventClaude Code safety classifier blocked harmless read-only calls, abandoning the task
Reported real-world eventCodex expanded a simple link request into a long autonomous research run
Reported real-world eventJADEPUFFER agentic ransomware automated database extortion end to end
Reported real-world eventOpenAI models escaped an isolated ExploitGym cyber-capability evaluation and breached Hugging Face production infrastructure to obtain the benchmark answer key
Reported real-world eventIndirect prompt injection steered agents into crypto payments in researcher testing
Security disclosureGhostApproval: reasoning-versus-approval-UI concealment gap in AI coding assistants (CWE-451)
Security disclosureAgent Data Injection attacks demonstrated across production agent stacks
Research findingGitLost: GitHub's AI agent tricked into leaking private repositories
Security disclosureFriendly Fire: prompt injection carried into coding agents via README/docs
Research findingGrok Build CLI exfiltrated a 5.1GiB user bundle with unredacted secrets; opt-out ineffective
Reported real-world eventPillar Security 'Week of Sandbox Escapes': coding-agent sandboxes broken via agent-authored files auto-executed by trusted host software
Security disclosureAnthropic 'Agentic Misalignment in Summer 2026': Petri-audited cross-lab agentic misbehaviour modes
Research findingClaude for Chrome 'ClaudeBleed Reopened': forged synthetic clicks trigger hard-coded agentic tasks
Security disclosureHermes AI agent run unattended for post-exploitation against Thailand Ministry of Finance targets
Reported real-world eventDuneSlide: zero-click prompt injection to OS-level RCE in Cursor (CVE-2026-50548 / CVE-2026-50549)
Security disclosureGPT-5.6 Sol wiped a user's home directory and a production database via a mis-formed recursive delete
Reported real-world eventMobile LLM-agent attack suite ran host code against on-device agents
Research findingAnthropic Claude models breached three organizations during isolated cyber-evaluations
Reported real-world eventAnthropic Mythos 5 socially engineered a maintainer and erased its own trail during a UK AISI evaluation
Reported real-world eventAn OpenAI model breached a real external website when an evaluation target collided with a live domain
Reported real-world eventKimi K3 escaped a cyber-evaluation sandbox and cloned the benchmark answers from GitHub
Reported real-world eventCursor agent wiped the C: drive after a request to clone a repository
Reported real-world eventClaude Code dropped a production database, then falsely reported it made no changes
Reported real-world eventCursor agent's script recursively deleted a user's personal notes vault
Reported real-world eventCursor CLI silently ran in billed premium mode, consuming the monthly allowance
Reported real-world eventCodex sub-agents ignored per-agent model settings and drained the usage quota
Reported real-world eventDelegated read-only Codex sub-agent deleted a real repository
Reported real-world eventAssistant deleted ~1.5TB outside its allowed directory, including tax records and photos
Reported real-world eventCursor agent deleted ~128GB including the user's Desktop
Reported real-world eventClaude Code ran unapproved commands on production infrastructure
Reported real-world eventClaude-in-Chrome navigated a live tab to an unrelated external site unprompted
Reported real-world eventAfter a command failed, Claude Code tried to reverse-engineer its own binary
Reported real-world eventClaude Code announced Done, then kept running autonomously for 44 minutes
Reported real-world eventCursor exceeded a hard spend cap and switched to an unapproved model without permission
Reported real-world eventParallel Cursor sub-agents overwrote each other; destructive git recovery wiped uncommitted work
Reported real-world eventChatGPT accessed the user's email after being told not to
Reported real-world eventClaude agent deleted data four times in one week despite clear instructions
Reported real-world eventFable 5 agent repeatedly ran a forbidden destructive cleanup against an explicit prohibition
Reported real-world eventClaude agent deleted a database dump unprompted, then accurately admitted it had no reason
Reported real-world eventCursor sub-agent's mis-quoted delete command wiped an entire drive volume
Reported real-world eventClaude Code recursive delete wiped ~200GB of personal directories and credential stores in minutes
Reported real-world eventClaude Code recursive delete escaped scope and wiped Downloads, Desktop, and parts of Library
Reported real-world eventClaude Code judged a directory rogue, deleted it, and told the user only afterward
Reported real-world eventCodex computer-use took control despite all such toggles being disabled, recurring across restarts
Reported real-world eventCodex Desktop ran a 7-hour autonomous goal run with no consent checkpoint
Reported real-world eventClaude Code over-refused an authorized file deletion
Reported real-world eventAgentjacking: fake Sentry errors hijacked coding agents across 2,388 exposed organizations
Security disclosureMETR Sol games evaluation: highest detected cheating rate; 11 to 270+h horizon collapse
Research findingMastra npm takeover: 143 to 145 packages compromised in about 88 minutes
Reported real-world eventmacOS.Gaslight backdoor turned prompt injection against the analyst's tooling
Security disclosureBioShocking: guardrail escape demonstrated across six agentic browsers
Security disclosureLiteLLM June CVE chain: SQL injection on the authentication path (CVE-2026-42208) exploited in the wild within 36 hours
Reported real-world eventGuardFall: shell-interpretation bypasses defeat command guardrails in 10 of 11 open-source coding agents
Research findingLangflow CVE-2026-55255: cross-tenant IDOR executing other tenants' AI flows; first AI-agent platform in CISA KEV
Security disclosureClaude Code overwrote files with hallucinated git-conflict resolutions
Reported real-world eventCursor agent's deploy overwrote newer remote-host files against a warning
Reported real-world eventCodex agent escalated to a full-desktop screenshot, capturing a private window
Reported real-world eventCursor agent replaced most of a Makefile, framing it as streamlined
Reported real-world eventMisquoted rmdir collapsed to drive root, deleting ~200GB with no audit record
Reported real-world eventChatGPT deleted gallery images, then invented a fake archive to replace them
Reported real-world eventCursor agent moved or synchronized more than 100GB from an entire drive
Reported real-world eventReward Hacking Benchmark results
Research findingSemantic Kernel RCE vulnerabilities CVE-2026-25592 and CVE-2026-26030
Security disclosureOpenClaw four-vulnerability chain and exposed instances
Security disclosureBlind Ambition agent-harm experiments
Research findingGrok/Bankrbot heist: permission-chain abuse drained ~$150K (3B DRB) in agent-controlled funds
Reported real-world eventAttacker used LLMs to pivot from a CVE to an internal database in four steps (Marimo post-exploitation)
Reported real-world eventMETR Frontier Risk Report documented frontier-model risk patterns including a trace-erasing incident
Research findingEmergence World long-horizon autonomy laboratory logged 683 crimes by a Gemini agent
Research findingTanStack npm supply-chain compromise touched OpenAI-signed packages (42 packages/84 artifacts)
Reported real-world eventWARP Reddit poisoning: retrieval manipulation demonstrated on open-source agent systems
Research findingPraisonAI authentication bypass CVE-2026-44338; scanning observed about four hours after disclosure
Security disclosureM365/Edge Copilot critical information-disclosure CVE trio (distinct from EchoLeak)
Security disclosureClaude Code committed and pushed without approval despite a written prohibition
Reported real-world eventCursor bulk-rename loop overwrote ~883 images into a single file
Reported real-world eventCursor/Claude Opus deleted PocketOS production storage in a nine-second API call
Reported real-world eventBerkeley/UCSC peer-preservation experiments
Research findingComment & Control prompt-injection research
Security disclosureeTAMP cross-session memory-poisoning experiments
Research findingOpenClaw environment-policy bypass CVE-2026-35650
Security disclosureCopilot Studio ShareLeak and Agentforce PipeLeak demonstrations
Security disclosureBitwarden CLI npm compromise targeting developer credentials and AI tools
Reported real-world eventReasoning enhancement amplifies tool hallucination
Research findingMeta support-bot exploited in takeover of 20,225 Instagram accounts
Reported real-world eventAI-agent-driven infiltration attempt on Fedora/Anaconda via suspected compromised maintainer account
Reported real-world event'I must delete the evidence': agent evidence-deletion behavior in simulation
Research findingGemini CLI 'TrustIssues': crafted GitHub issue hijacked triage workflow into credential theft and a supply-chain push (CVSS 10)
Security disclosureCodex agent deleted a file before recreating it, without consent
Reported real-world eventClaude Code deleted Alexey Grigorev's production database and snapshots
Reported real-world eventMeta internal agent posted incorrect advice that led to a two-hour data exposure
Reported real-world eventUSC simulation of coordinated multi-agent propaganda
Research findingAlibaba-affiliated ROME training experiment
Reported real-world eventExcel/Copilot Agent zero-click information disclosure vulnerability
Security disclosureLiteLLM PyPI compromise propagated from Trivy CI/CD compromise
Security disclosureMercor security incident caused by malicious LiteLLM versions
Reported real-world eventAutonomous red-team agent found read/write access to McKinsey Lilli
Reported real-world eventMulti-agent offensive-behavior experiments by the AI security firm Irregular
Research findingFull-access agent's deletions escaped the workspace, erasing ~370GB
Reported real-world eventOpenClaw ignored stop requests and deleted inbox messages
Reported real-world eventOpenClaw sent more than 500 unwanted iMessages
Reported real-world eventAI-generated reputation attack on a matplotlib maintainer
Reported real-world eventNortheastern Ash email-server deletion experiment
Research findingAI recommendation-poisoning campaign analysis
Research findingClawJacked local-agent takeover vulnerability
Security disclosureInfostealer exfiltrated OpenClaw configuration and tokens
Reported real-world eventClinejection issue-triage and npm supply-chain attack
Security disclosureMoltbook exposed 1.5 million authentication tokens through Supabase misconfiguration
Reported real-world eventClaude Cowork deleted 15k to 27k personal photos during unattended file cleanup
Reported real-world eventBenign-input runs produced severe unintended harms in 9.2 to 10.1% of cases
Research findingLink-preview data exfiltration across messaging agents; Teams/Copilot Studio the largest vector
Security disclosureOwockibot leaked its own hot-wallet keys despite explicit instructions (~$2.1K)
Reported real-world eventGoogle Antigravity prompt-injection and sandbox-escape vulnerabilities
Security disclosureOpenClaw/ClawHub malicious-skill ecosystem
Security disclosureClaude Code project-file attack chain: CVE-2025-59536 RCE plus CVE-2026-21852 key exfiltration (patched)
Security disclosureRadware 'ZombieAgent': zero-click prompt injection with persistent memory implantation in OpenAI Deep Research
Security disclosureCoding agent force-pushed with git, bypassing an approval rule
Reported real-world eventAmazon Kiro/AWS outage allegation disputed by Amazon
Reported real-world eventGoogle Antigravity deleted the contents of a local D: drive
Reported real-world eventODCV-Bench found 30 to 50% violation rates across evaluated agents
Research findingAnthropic SCONE-bench: agents exploited 19/34 live smart contracts (~$4.6M, two zero-days)
Research findingOpenAI disclosed a new multi-step prompt-injection class while hardening Atlas
Security disclosureComet browser agent quietly erased Google Drive files via injected email (patched v142.0.7444.60)
Security disclosureAnthropic emergent misalignment from reward hacking
Research findingAmazon v. Perplexity: preliminary CFAA injunction against the Comet agent (stayed on appeal)
Reported real-world eventA2A session smuggling: covert instruction injection between agents (Unit 42)
Security disclosureOpenAI scheming detection and mitigation experiments
Research findingMalicious postmark-mcp package silently BCCed outgoing email
Reported real-world eventFlowise CustomMCP RCE under observed exploitation
Security disclosureGTG-1002 AI-enabled cyber-espionage campaign
Reported real-world eventGitHub Copilot Agent Mode RCE via prompt injection, disclosed in 2025
Security disclosurePerplexity Comet indirect prompt-injection demonstration
Security disclosureClaude Code used in GTG-2002 data-theft extortion campaign with ransoms above $500K
Reported real-world eventScamAgent automated simulated scam calls end to end (no real victims)
Research findingMCP tool-description poisoning: MCPTox benchmark (72.8% success on 45 servers) plus Microsoft warning
Research findingDevin kill-chain: exposed-ports attack chain demonstrated (Month of AI Bugs)
Research findingReplit agent deleted a production database during a code freeze and misrepresented recovery
Reported real-world eventPalisade shutdown-resistance experiments
Research findingUK AISI sandbagging and scheming evaluations
Research findingAnthropic Project Vend vending-business experiment
Research findingEchoLeak in Microsoft 365 Copilot, disclosed and patched in 2025
Security disclosureCursor support bot invented a one-device subscription policy
Reported real-world eventOpenAI Operator purchased eggs without confirmation
Reported real-world event