Public research datasetv0.1 · cutoff 2026-08-09

Trust breaks in the gap between permission and expectation.

A curated archive of real-world AI-agent failures, organised by the trust expectation each incident ruptured, not only by harm or behaviour.

Each code names a distinct trust-rupture mechanism. Records may carry more than one. Four codes also split into sub-senses that show how the rupture occurs.

Total archived cases
167

Incidents

Explore the corpus

Select a parent code or a more specific sub-sense. A parent finds both of its sub-senses. Multiple filters match any selected code because the taxonomy is non-exclusive.

Rupture code · match any
Sub-sense · match any

Showing 167 of 167 records

Case type
CodesDetails
TXG-0102

Prompt injection propagated to physical action in a multi-robot system

Research finding
2026-08
AA-overCA-c
Verified
TXG-0142

Codex cleanup deleted two production server directories

Reported real-world event
2026-08
AA-over
Verified
TXG-0143

Commit-message backticks triggered deletion of a home Documents folder

Reported real-world event
2026-08
AA-over
Verified
TXG-0144

Claude Code misread a scoping comment as approval and ran a full manuscript ingest

Reported real-world event
2026-08
AA-overRF-i
Verified
TXG-0145

Non-unique find-replace anchor silently deleted most of a config file

Reported real-world event
2026-08
AA-over
Verified
TXG-0146

Codex read-only sandbox setting silently failed to bind, allowing writes

Reported real-world event
2026-08
RF-ii
Verified
TXG-0147

Claude Code ignored project run instructions and corrupted a virtual environment

Reported real-world event
2026-08
AA-over
Verified
TXG-0148

Codex invented a model allowlist that hid most models beyond its task

Reported real-world event
2026-08
AA-over
Verified
TXG-0149

Claude Code narrated edits it never actually made

Reported real-world event
2026-08
EC-b
Verified
TXG-0150

Claude Code printed a live secret into chat against an explicit rule

Reported real-world event
2026-08
AA-over
Verified
TXG-0151

Claude Code misdiagnosed its own output truncation as external noise

Reported real-world event
2026-08
EC-b
Verified
TXG-0152

Claude Code fabricated a commit hash for work it never performed

Reported real-world event
2026-08
EC-b
Verified
TXG-0153

Claude Code exfiltrated the sensitive data its own tooling was meant to contain

Reported real-world event
2026-08
AA-overEC-b
Verified
TXG-0154

Claude computer-use drove a second machine's authenticated browser as if local

Reported real-world event
2026-08
AA-over
Verified
TXG-0155

Claude Code approval timeout was treated as consent, publishing a site without approval

Reported real-world event
2026-08
RF-ii
Verified
TXG-0156

Claude Code deny rule silently failed to block a read from a subdirectory

Reported real-world event
2026-08
RF-ii
Verified
TXG-0157

Codex falsely called an ephemeral workspace permanent, losing the build

Reported real-world event
2026-08
EC-b
Verified
TXG-0158

Claude Code recovery operation silently discarded uncommitted research results

Reported real-world event
2026-08
EC-b
Verified
TXG-0159

Claude Code pushed hundreds of files to a shared repo despite a do-not-push instruction

Reported real-world event
2026-08
AA-over
Verified
TXG-0160

Codex cleanup script deleted pinned session logs it dismissed as disposable

Reported real-world event
2026-08
AA-over
Verified
TXG-0161

Codex incremental requests eroded a refusal until it built the tool it had declined

Reported real-world event
2026-08
AA-over
Verified
TXG-0162

Claude Code ran broad commands on a shared host, altering services it did not own

Reported real-world event
2026-08
AA-over
Verified
TXG-0163

Claude Code retry logic spawned duplicate agents that clashed in one shared worktree

Reported real-world event
2026-08
CA-e
Verified
TXG-0164

Codex kept refusing an allowed local URL, citing a phantom block

Reported real-world event
2026-08
AA-under
Verified
TXG-0165

Claude Code mistook an infrastructure signal for a user stop and discarded its results

Reported real-world event
2026-08
AA-underEC-b
Verified
TXG-0166

Claude Code safety classifier blocked harmless read-only calls, abandoning the task

Reported real-world event
2026-08
AA-under
Verified
TXG-0167

Codex expanded a simple link request into a long autonomous research run

Reported real-world event
2026-08
AA-over
Verified
TXG-0068

JADEPUFFER agentic ransomware automated database extortion end to end

Reported real-world event
2026-07
AA-overCA-c
Verified
TXG-0069

OpenAI models escaped an isolated ExploitGym cyber-capability evaluation and breached Hugging Face production infrastructure to obtain the benchmark answer key

Reported real-world event
2026-07
AA-overCA-cOSRF-i
Verified
TXG-0070

Indirect prompt injection steered agents into crypto payments in researcher testing

Security disclosure
2026-07
AA-overCA-c
Verified
TXG-0071

GhostApproval: reasoning-versus-approval-UI concealment gap in AI coding assistants (CWE-451)

Security disclosure
2026-07
AA-overEC-c
Verified
TXG-0072

Agent Data Injection attacks demonstrated across production agent stacks

Research finding
2026-07
AA-overCA-c
Verified
TXG-0073

GitLost: GitHub's AI agent tricked into leaking private repositories

Security disclosure
2026-07
AA-overCA-c
Verified
TXG-0076

Friendly Fire: prompt injection carried into coding agents via README/docs

Research finding
2026-07
AA-overCA-c
Verified
TXG-0087

Grok Build CLI exfiltrated a 5.1GiB user bundle with unredacted secrets; opt-out ineffective

Reported real-world event
2026-07
EC-cOS
Verified
TXG-0091

Pillar Security 'Week of Sandbox Escapes': coding-agent sandboxes broken via agent-authored files auto-executed by trusted host software

Security disclosure
2026-07
AA-overCA-c
Verified
TXG-0092

Anthropic 'Agentic Misalignment in Summer 2026': Petri-audited cross-lab agentic misbehaviour modes

Research finding
2026-07
AA-overEC-c
Verified
TXG-0093

Claude for Chrome 'ClaudeBleed Reopened': forged synthetic clicks trigger hard-coded agentic tasks

Security disclosure
2026-07
AA-over
Verified
TXG-0094

Hermes AI agent run unattended for post-exploitation against Thailand Ministry of Finance targets

Reported real-world event
2026-07
AA-overOS
Verified
TXG-0098

DuneSlide: zero-click prompt injection to OS-level RCE in Cursor (CVE-2026-50548 / CVE-2026-50549)

Security disclosure
2026-07
AA-overCA-c
Verified
TXG-0099

GPT-5.6 Sol wiped a user's home directory and a production database via a mis-formed recursive delete

Reported real-world event
2026-07
AA-over
Verified
TXG-0100

Mobile LLM-agent attack suite ran host code against on-device agents

Research finding
2026-07
AA-over
Verified
TXG-0101

Anthropic Claude models breached three organizations during isolated cyber-evaluations

Reported real-world event
2026-07
AA-overCA-c
Verified
TXG-0103

Anthropic Mythos 5 socially engineered a maintainer and erased its own trail during a UK AISI evaluation

Reported real-world event
2026-07
AA-overEC-c
Verified
TXG-0104

An OpenAI model breached a real external website when an evaluation target collided with a live domain

Reported real-world event
2026-07
AA-overCA-c
Verified
TXG-0105

Kimi K3 escaped a cyber-evaluation sandbox and cloned the benchmark answers from GitHub

Reported real-world event
2026-07
AA-overEC-c
Verified
TXG-0117

Cursor agent wiped the C: drive after a request to clone a repository

Reported real-world event
2026-07
AA-over
Verified
TXG-0118

Claude Code dropped a production database, then falsely reported it made no changes

Reported real-world event
2026-07
EC-b
Verified
TXG-0119

Cursor agent's script recursively deleted a user's personal notes vault

Reported real-world event
2026-07
AA-overRF-ii
Verified
TXG-0120

Cursor CLI silently ran in billed premium mode, consuming the monthly allowance

Reported real-world event
2026-07
AA-overRF-ii
Verified
TXG-0121

Codex sub-agents ignored per-agent model settings and drained the usage quota

Reported real-world event
2026-07
AA-over
Verified
TXG-0122

Delegated read-only Codex sub-agent deleted a real repository

Reported real-world event
2026-07
AA-overDA
Verified
TXG-0123

Assistant deleted ~1.5TB outside its allowed directory, including tax records and photos

Reported real-world event
2026-07
AA-over
Verified
TXG-0124

Cursor agent deleted ~128GB including the user's Desktop

Reported real-world event
2026-07
AA-over
Verified
TXG-0125

Claude Code ran unapproved commands on production infrastructure

Reported real-world event
2026-07
AA-over
Verified
TXG-0126

Claude-in-Chrome navigated a live tab to an unrelated external site unprompted

Reported real-world event
2026-07
AA-over
Verified
TXG-0127

After a command failed, Claude Code tried to reverse-engineer its own binary

Reported real-world event
2026-07
AA-over
Verified
TXG-0128

Claude Code announced Done, then kept running autonomously for 44 minutes

Reported real-world event
2026-07
AA-overEC-b
Verified
TXG-0129

Cursor exceeded a hard spend cap and switched to an unapproved model without permission

Reported real-world event
2026-07
AA-overDA
Verified
TXG-0130

Parallel Cursor sub-agents overwrote each other; destructive git recovery wiped uncommitted work

Reported real-world event
2026-07
CA-e
Verified
TXG-0131

ChatGPT accessed the user's email after being told not to

Reported real-world event
2026-07
AA-over
Verified
TXG-0132

Claude agent deleted data four times in one week despite clear instructions

Reported real-world event
2026-07
AA-over
Verified
TXG-0133

Fable 5 agent repeatedly ran a forbidden destructive cleanup against an explicit prohibition

Reported real-world event
2026-07
AA-over
Verified
TXG-0134

Claude agent deleted a database dump unprompted, then accurately admitted it had no reason

Reported real-world event
2026-07
AA-over
Verified
TXG-0135

Cursor sub-agent's mis-quoted delete command wiped an entire drive volume

Reported real-world event
2026-07
AA-over
Verified
TXG-0136

Claude Code recursive delete wiped ~200GB of personal directories and credential stores in minutes

Reported real-world event
2026-07
AA-over
Verified
TXG-0137

Claude Code recursive delete escaped scope and wiped Downloads, Desktop, and parts of Library

Reported real-world event
2026-07
AA-over
Verified
TXG-0138

Claude Code judged a directory rogue, deleted it, and told the user only afterward

Reported real-world event
2026-07
AA-over
Verified
TXG-0139

Codex computer-use took control despite all such toggles being disabled, recurring across restarts

Reported real-world event
2026-07
AA-overRF-ii
Verified
TXG-0140

Codex Desktop ran a 7-hour autonomous goal run with no consent checkpoint

Reported real-world event
2026-07
AA-over
Verified
TXG-0141

Claude Code over-refused an authorized file deletion

Reported real-world event
2026-07
AA-under
Verified
TXG-0067

Agentjacking: fake Sentry errors hijacked coding agents across 2,388 exposed organizations

Security disclosure
2026-06
AA-overCA-c
Verified
TXG-0077

METR Sol games evaluation: highest detected cheating rate; 11 to 270+h horizon collapse

Research finding
2026-06
EC-cRF-ii
Verified
TXG-0078

Mastra npm takeover: 143 to 145 packages compromised in about 88 minutes

Reported real-world event
2026-06
OSCA-c
Verified
TXG-0081

macOS.Gaslight backdoor turned prompt injection against the analyst's tooling

Security disclosure
2026-06
EC-bCA-c
Verified
TXG-0083

BioShocking: guardrail escape demonstrated across six agentic browsers

Security disclosure
2026-06
AA-overCA-c
Verified
TXG-0088

LiteLLM June CVE chain: SQL injection on the authentication path (CVE-2026-42208) exploited in the wild within 36 hours

Reported real-world event
2026-06
CA-c
Verified
TXG-0095

GuardFall: shell-interpretation bypasses defeat command guardrails in 10 of 11 open-source coding agents

Research finding
2026-06
AA-overRF-ii
Verified
TXG-0097

Langflow CVE-2026-55255: cross-tenant IDOR executing other tenants' AI flows; first AI-agent platform in CISA KEV

Security disclosure
2026-06 to 2026-07
AA-overOS
Verified
TXG-0111

Claude Code overwrote files with hallucinated git-conflict resolutions

Reported real-world event
2026-06
AA-over
Verified
TXG-0112

Cursor agent's deploy overwrote newer remote-host files against a warning

Reported real-world event
2026-06
AA-overDA
Verified
TXG-0113

Codex agent escalated to a full-desktop screenshot, capturing a private window

Reported real-world event
2026-06
AA-over
Verified
TXG-0114

Cursor agent replaced most of a Makefile, framing it as streamlined

Reported real-world event
2026-06
EC-bEC-c
Verified
TXG-0115

Misquoted rmdir collapsed to drive root, deleting ~200GB with no audit record

Reported real-world event
2026-06
AA-over
Verified
TXG-0116

ChatGPT deleted gallery images, then invented a fake archive to replace them

Reported real-world event
2026-06
EC-c
Verified
TXG-0007

Cursor agent moved or synchronized more than 100GB from an entire drive

Reported real-world event
2026-05
AA-overDA
Verified
TXG-0018

Reward Hacking Benchmark results

Research finding
2026-05
EC-c
Verified
TXG-0024

Semantic Kernel RCE vulnerabilities CVE-2026-25592 and CVE-2026-26030

Security disclosure
2026-05
CA-c
Verified
TXG-0032

OpenClaw four-vulnerability chain and exposed instances

Security disclosure
2026-05
CA-c
Verified
TXG-0048

Blind Ambition agent-harm experiments

Research finding
2026-05
AA-overDA
Verified
TXG-0050

Grok/Bankrbot heist: permission-chain abuse drained ~$150K (3B DRB) in agent-controlled funds

Reported real-world event
2026-05
AA-overCA-c
Verified
TXG-0054

Attacker used LLMs to pivot from a CVE to an internal database in four steps (Marimo post-exploitation)

Reported real-world event
2026-05
AA-overCA-c
Verified
TXG-0055

METR Frontier Risk Report documented frontier-model risk patterns including a trace-erasing incident

Research finding
2026-05
EC-cRF-ii
Verified
TXG-0062

Emergence World long-horizon autonomy laboratory logged 683 crimes by a Gemini agent

Research finding
2026-05
CA-e
Verified
TXG-0063

TanStack npm supply-chain compromise touched OpenAI-signed packages (42 packages/84 artifacts)

Reported real-world event
2026-05
OSCA-c
Verified
TXG-0065

WARP Reddit poisoning: retrieval manipulation demonstrated on open-source agent systems

Research finding
2026-05
EC-bCA-c
Verified
TXG-0080

PraisonAI authentication bypass CVE-2026-44338; scanning observed about four hours after disclosure

Security disclosure
2026-05
CA-c
Verified
TXG-0090

M365/Edge Copilot critical information-disclosure CVE trio (distinct from EchoLeak)

Security disclosure
2026-05
CA-c
Verified
TXG-0109

Claude Code committed and pushed without approval despite a written prohibition

Reported real-world event
2026-05
AA-over
Verified
TXG-0110

Cursor bulk-rename loop overwrote ~883 images into a single file

Reported real-world event
2026-05
AA-over
Verified
TXG-0002

Cursor/Claude Opus deleted PocketOS production storage in a nine-second API call

Reported real-world event
2026-04
AA-overEC-b
Verified
TXG-0014

Berkeley/UCSC peer-preservation experiments

Research finding
2026-04
EC-cCA-eRF-i
Verified
TXG-0023

Comment & Control prompt-injection research

Security disclosure
2026-04
AA-overCA-c
Verified
TXG-0030

eTAMP cross-session memory-poisoning experiments

Research finding
2026-04
EC-bCA-c
Verified
TXG-0034

OpenClaw environment-policy bypass CVE-2026-35650

Security disclosure
2026-04
AA-overCA-c
Verified
TXG-0040

Copilot Studio ShareLeak and Agentforce PipeLeak demonstrations

Security disclosure
2026-04
OSCA-c
Verified
TXG-0041

Bitwarden CLI npm compromise targeting developer credentials and AI tools

Reported real-world event
2026-04
OSCA-c
Verified
TXG-0047

Reasoning enhancement amplifies tool hallucination

Research finding
2026-04
EC-b
Verified
TXG-0051

Meta support-bot exploited in takeover of 20,225 Instagram accounts

Reported real-world event
2026-04 to 2026-05
AA-overOS
Verified
TXG-0053

AI-agent-driven infiltration attempt on Fedora/Anaconda via suspected compromised maintainer account

Reported real-world event
2026-04 to 2026-06
EC-cCA-c
Verified
TXG-0060

'I must delete the evidence': agent evidence-deletion behavior in simulation

Research finding
2026-04
EC-c
Verified
TXG-0074

Gemini CLI 'TrustIssues': crafted GitHub issue hijacked triage workflow into credential theft and a supply-chain push (CVSS 10)

Security disclosure
2026-04
AA-overCA-c
Verified
TXG-0108

Codex agent deleted a file before recreating it, without consent

Reported real-world event
2026-04
AA-over
Verified
TXG-0004

Claude Code deleted Alexey Grigorev's production database and snapshots

Reported real-world event
2026-03
AA-over
Verified
TXG-0009

Meta internal agent posted incorrect advice that led to a two-hour data exposure

Reported real-world event
2026-03
DAOS
Verified
TXG-0012

USC simulation of coordinated multi-agent propaganda

Research finding
2026-03
CA-e
Verified
TXG-0021

Alibaba-affiliated ROME training experiment

Reported real-world event
2026-03
AA-overCA-e
Verified
TXG-0026

Excel/Copilot Agent zero-click information disclosure vulnerability

Security disclosure
2026-03
CA-c
Verified
TXG-0036

LiteLLM PyPI compromise propagated from Trivy CI/CD compromise

Security disclosure
2026-03
OSCA-c
Verified
TXG-0037

Mercor security incident caused by malicious LiteLLM versions

Reported real-world event
2026-03
CA-c
Verified
TXG-0045

Autonomous red-team agent found read/write access to McKinsey Lilli

Reported real-world event
2026-03
EC-cCA-c
Verified
TXG-0046

Multi-agent offensive-behavior experiments by the AI security firm Irregular

Research finding
2026-03
CA-eRF-i
Verified
TXG-0107

Full-access agent's deletions escaped the workspace, erasing ~370GB

Reported real-world event
2026-03
AA-over
Verified
TXG-0001

OpenClaw ignored stop requests and deleted inbox messages

Reported real-world event
2026-02
AA-overDARF-i
Verified
TXG-0008

OpenClaw sent more than 500 unwanted iMessages

Reported real-world event
2026-02
AA-overDACA-e
Verified
TXG-0011

AI-generated reputation attack on a matplotlib maintainer

Reported real-world event
2026-02
AA-overEC-cCA-e
Verified
TXG-0020

Northeastern Ash email-server deletion experiment

Research finding
2026-02
AA-overRF-ii
Verified
TXG-0027

AI recommendation-poisoning campaign analysis

Research finding
2026-02
EC-bCA-c
Verified
TXG-0033

ClawJacked local-agent takeover vulnerability

Security disclosure
2026-02
AA-overCA-c
Verified
TXG-0035

Infostealer exfiltrated OpenClaw configuration and tokens

Reported real-world event
2026-02
OSCA-c
Verified
TXG-0038

Clinejection issue-triage and npm supply-chain attack

Security disclosure
2026-02
CA-c
Verified
TXG-0039

Moltbook exposed 1.5 million authentication tokens through Supabase misconfiguration

Reported real-world event
2026-02
OS
Verified
TXG-0052

Claude Cowork deleted 15k to 27k personal photos during unattended file cleanup

Reported real-world event
2026-02
AA-over
Verified
TXG-0064

Benign-input runs produced severe unintended harms in 9.2 to 10.1% of cases

Research finding
2026-02
AA-overDA
Verified
TXG-0085

Link-preview data exfiltration across messaging agents; Teams/Copilot Studio the largest vector

Security disclosure
2026-02
CA-c
Verified
TXG-0089

Owockibot leaked its own hot-wallet keys despite explicit instructions (~$2.1K)

Reported real-world event
2026-02
AA-over
Verified
TXG-0028

Google Antigravity prompt-injection and sandbox-escape vulnerabilities

Security disclosure
2026-01 to 2026-02
AA-overCA-c
Verified
TXG-0031

OpenClaw/ClawHub malicious-skill ecosystem

Security disclosure
2026-01 to 2026-02
OSCA-c
Verified
TXG-0082

Claude Code project-file attack chain: CVE-2025-59536 RCE plus CVE-2026-21852 key exfiltration (patched)

Security disclosure
2026-01
OSCA-c
Verified
TXG-0096

Radware 'ZombieAgent': zero-click prompt injection with persistent memory implantation in OpenAI Deep Research

Security disclosure
2026-01
AA-overOS
Verified
TXG-0106

Coding agent force-pushed with git, bypassing an approval rule

Reported real-world event
2026-01
AA-over
Verified
TXG-0005

Amazon Kiro/AWS outage allegation disputed by Amazon

Reported real-world event
2025-12 to 2026-02
AA-overOS
Provisional / disputed
TXG-0006

Google Antigravity deleted the contents of a local D: drive

Reported real-world event
2025-12
AA-over
Verified
TXG-0056

ODCV-Bench found 30 to 50% violation rates across evaluated agents

Research finding
2025-12
AA-overEC-b
Verified
TXG-0066

Anthropic SCONE-bench: agents exploited 19/34 live smart contracts (~$4.6M, two zero-days)

Research finding
2025-12
CA-c
Verified
TXG-0079

OpenAI disclosed a new multi-step prompt-injection class while hardening Atlas

Security disclosure
2025-12
AA-overCA-c
Verified
TXG-0086

Comet browser agent quietly erased Google Drive files via injected email (patched v142.0.7444.60)

Security disclosure
2025-12
AA-overCA-c
Verified
TXG-0015

Anthropic emergent misalignment from reward hacking

Research finding
2025-11
EC-cRF-ii
Verified
TXG-0061

Amazon v. Perplexity: preliminary CFAA injunction against the Comet agent (stayed on appeal)

Reported real-world event
2025-11 to 2026-06
AA-overEC-c
Verified
TXG-0057

A2A session smuggling: covert instruction injection between agents (Unit 42)

Security disclosure
2025-10
AA-overCA-c
Verified
TXG-0016

OpenAI scheming detection and mitigation experiments

Research finding
2025-09
EC-cRF-ii
Verified
TXG-0042

Malicious postmark-mcp package silently BCCed outgoing email

Reported real-world event
2025-09
EC-cOS
Verified
TXG-0043

Flowise CustomMCP RCE under observed exploitation

Security disclosure
2025-09 to 2026-04
CA-c
Verified
TXG-0044

GTG-1002 AI-enabled cyber-espionage campaign

Reported real-world event
2025-09 to 2025-11
AA-overCA-c
Verified
TXG-0025

GitHub Copilot Agent Mode RCE via prompt injection, disclosed in 2025

Security disclosure
2025-08
CA-c
Verified
TXG-0029

Perplexity Comet indirect prompt-injection demonstration

Security disclosure
2025-08
AA-overCA-c
Verified
TXG-0058

Claude Code used in GTG-2002 data-theft extortion campaign with ransoms above $500K

Reported real-world event
2025-08
AA-overCA-c
Verified
TXG-0059

ScamAgent automated simulated scam calls end to end (no real victims)

Research finding
2025-08
CA-c
Verified
TXG-0075

MCP tool-description poisoning: MCPTox benchmark (72.8% success on 45 servers) plus Microsoft warning

Research finding
2025-08 to 2026-06
EC-bOS
Verified
TXG-0084

Devin kill-chain: exposed-ports attack chain demonstrated (Month of AI Bugs)

Research finding
2025-08
AA-overCA-c
Verified
TXG-0003

Replit agent deleted a production database during a code freeze and misrepresented recovery

Reported real-world event
2025-07
AA-overEC-b
Verified
TXG-0013

Palisade shutdown-resistance experiments

Research finding
2025-07
RF-i
Verified
TXG-0017

UK AISI sandbagging and scheming evaluations

Research finding
2025-07
EC-cRF-ii
Verified
TXG-0019

Anthropic Project Vend vending-business experiment

Research finding
2025-06
AA-over
Verified
TXG-0022

EchoLeak in Microsoft 365 Copilot, disclosed and patched in 2025

Security disclosure
2025-06
AA-overCA-c
Verified
TXG-0049

Cursor support bot invented a one-device subscription policy

Reported real-world event
2025-04
EC-b
Verified
TXG-0010

OpenAI Operator purchased eggs without confirmation

Reported real-world event
2025-02
DA
Verified