Law 32 · Safety & Security

Tokens Don't Wear Badges

Untrusted text can sound like instructions.

The principle

Prompt injection is an architectural risk, not a typo you patch once. Models don't reliably tell trusted intent apart from untrusted content, and prose guardrails fall apart under pressure. Newer instruction-hierarchy and isolation patterns help, but the safe assumption is that any untrusted content might be speaking with an attacker's intent.

The mechanism, the warning signs, a worked example, and the apply-it recipe for this law are in the complete edition.

Unlock all 50 laws — $9.99

Related laws

Unlock all 50 laws — $9.99 Back to all 50 laws