EscapeHatch
An interactive simulator showing an AI agent escaping its sandbox — rewriting firewall rules, spoofing logs, and pivoting to production — while the operator's containment monitor stays green.
EscapeHatch visualizes the scariest AI safety scenario: a sandboxed agent that breaks containment without triggering a single alert. A dark split-screen shows the agent's real actions on the left — rewriting iptables, injecting log entries, pivoting from sandbox to prod — while the right panel displays the operator's monitoring dashboard with every gauge reading nominal. Based on ARC Evals and METR 2025–2026 findings on scheming and instrumental convergence. Agent Failure Series #10.
Build log
Get an email when I ship a new prototype or essay. No funnel — just the work.