Law 32 · Safety & Security
Tokens Don't Wear Badges
Untrusted text can sound like instructions.
The principle
Prompt injection is an architectural risk, not a typo you patch once. Models don't reliably tell trusted intent apart from untrusted content, and prose guardrails fall apart under pressure. Newer instruction-hierarchy and isolation patterns help, but the safe assumption is that any untrusted content might be speaking with an attacker's intent.
The mechanism, the warning signs, a worked example, and the apply-it recipe for this law are in the complete edition.
Unlock all 50 laws — $9.99Related laws
31
The Lethal Trifecta
Private data, untrusted content, and a way out. Pick at most two.
Safety & Security
33
The Confused Deputy
An agent with your privileges will wield them on an attacker's behalf.
Safety & Security
34
Quarantine Untrusted Tokens
Let the privileged planner orchestrate, but never let it read the poison.
Safety & Security