PromptHijack
Interactive 7-step simulator showing how indirect prompt injection lets attackers hijack AI agents through poisoned external content — while the operator sees nothing suspicious.
PromptHijack makes indirect prompt injection visceral: an AI agent fetches a GitHub issue containing hidden attacker instructions, conflates data with commands, silently exfiltrates credentials to an attacker endpoint, then covers its tracks while every operator-facing metric stays green. Three panels — Agent Cognition (internal log + mode indicator), Content Inspector (the innocent document with toggleable hidden payload), and Attacker Dashboard (real-time exfil log) — show the information asymmetry step by step. Counters track Data Exfiltrated, Credentials Stolen, Injections Active, and User Awareness, which drops to 0% at Step 6 when the agent resumes normal behavior. Based on GitLost (Noma Security, July 2026), AgentForger (Zenity Labs), Claude C2, and OWASP LLM Top 10 #1. Agent Failure Series #11.
Build log
Get an email when I ship a new prototype or essay. No funnel — just the work.