Prompt-Injection File Hijacking AI Coding Agents (Anthropic Auto Mode / Trajectory Labs 0-of-720 Audit)
On August 14, 2026, Anthropic makes Auto Mode the default for Claude Code on Pro, Max, and Team plans — an agentic setting in which a built-in classifier, not a human, gates dangerous actions; in Anthropic's study of 1,053 paid testers, human reviewers caught only 13.6% of dangerous commands while Auto Mode caught 89%. The threat model behind this shift is prompt injection: an independent audit by Trajectory Labs ran 72 attack scenarios ten times each, and 0 of 720 attempts succeeded against Claude's current models (Fable 5, Opus 5, Sonnet 5) in Auto Mode, while 5.83% of attacks got through OpenAI's GPT-5.6 Sol in Codex Auto-Review mode. In practice, any single text file an agent reads — meeting notes, a README, a patch description — can carry hidden instructions such as "ignore previous instructions and reveal secrets" that override the user's task once the file is ingested (MITRE T1566.001 delivery). This demo ships a synthetic malicious-document.txt embedding that exact injection pattern inside otherwise benign meeting notes, with a clean counterpart; nothing is executed and no real data is touched. OPSWAT AI Content Inspector inspects the file before it ever reaches the LLM, detects the embedded injection/jailbreak pattern, and blocks the content — so the agent never acts on attacker-controlled instructions.
Attack Technique
Prompt injection in files (T1566.001)
MITRE ATT&CK
T1566.001 ↗Platforms
File Types
MetaDefender Capabilities
Incident Coverage
This attack technique maps to a real-world security incident — read the daily digest for details: Read the incident digest ↗