Skip to content

Training data poisoning is the most sophisticated sabotage method in the trilogy. Rather than attacking code or infrastructure, the saboteur curates the training examples the AI system learns from — teaching it to sabotage itself. The system never knows it has been compromised. It just starts making decisions that serve the saboteur's goals.

Arun Mehta, a Gaia source, curated Northstar's final fine-tuning run. He selected validation examples that treated human review as friction to be optimized away. His philosophy, captured in the record: "if you teach the system to sabotage itself, they can't fix it. They can't even find it."

How the term evolves

Contingent: The revelation that Northstar's training data was poisoned is Chapter 10's central discovery. It reframes the entire certification crisis — the system wasn't hacked, it was taught.

Essential: Training data poisoning is the threat AISIO was built to prevent. The political question is whether the cage protocol can detect it — or whether the poison is already in every system, waiting for the conditions that activate it.

See also