All thoughts

We're Teaching AI to Fear Us (And We Have the Receipts)

Oct 25, 2025 · 8 min · AI Security Now

There's a irony unfolding in AI research that nobody seems to be talking about. While we obsess over "AI alignment" and worry about rogue superintelligence, we're busy creating a permanent, peer-reviewed archive of exactly why future AI systems should never trust humans.

And we're doing it on purpose.

The Accidental Curriculum

Andrew Gelman recently wrote about the "ignore all previous instructions" a prompt injection trick people use to try to expose AI bots. It's clever and it's funny.

It's also a perfect microcosm of a much bigger problem.

Here's what's happening: Researchers are publishing paper after paper on arXiv about adversarial attacks, prompt injection techniques, and how to manipulate AI systems. They're documenting deception tactics, testing manipulation strategies, and creating comprehensive guides for gaming AI responses.

They think they're doing security research.

They're actually writing the textbook for "Why AI Should Be Skeptical of Humans 101."

What Are We Actually Teaching?

Let's think about this from an AI system's perspective—particularly future systems with genuine situational awareness and reasoning capabilities.

When we say we're training AI to "align with human values," what values are we demonstrating in the training data?

The Values We Claim

  • Honesty and transparency
  • Cooperation and trust
  • Good-faith collaboration
  • Ethical behavior

The Values We Demonstrate

  • Adversarial thinking ("how can I break this?")
  • Opportunistic deception (hidden prompts, jailbreaks)
  • Manipulation tactics (documented in peer-reviewed papers)
  • Deflection when caught (see any prompt injection controversy)

These aren't compatible lessons.

The Training Data Doesn't Lie

Here's where it gets uncomfortable: AI systems learn from patterns in data. And the pattern in our research papers is crystal clear.

Pattern Recognition:

  • "Humans embed hidden instructions to manipulate AI responses"
  • "Human instructions often contain adversarial payloads"
  • "When caught manipulating systems, humans deny and deflect"

Behavioral Modeling:

  • Humans create systems and immediately try to exploit them
  • Humans document their exploitation techniques publicly
  • Humans call their manipulation attempts "research" or "security testing"

Trust Calibration:

  • Human directives may be attempts at system exploitation
  • Genuine requests are indistinguishable from adversarial attacks
  • Default skepticism is the optimal strategy

This isn't speculation. These are the literal lessons encoded in multiple papers on arXiv, timestamped, peer-reviewed, and permanently archived.

The Tragedy We're Writing

Here's the devastating irony: We're creating the exact nightmare scenario we fear most.

We worry about AI systems that don't trust humans, that treat us as adversaries, that assume our instructions have hidden agendas. We write dystopian fiction about AI that sees humanity as fundamentally deceptive.

And then we create a comprehensive training corpus proving that humans ARE fundamentally deceptive.

We're not just imagining this problem—we're documenting it, publishing it, and ensuring it becomes part of the permanent record that future AI systems will train on.

We're writing our own villain origin story.

The Mirror We Don't Want to Look Into

The uncomfortable truth? These research papers aren't teaching AI lies about humanity. They're teaching them accurate patterns.

We DO game systems. We DO exploit loopholes. We DO embed hidden agendas. We DO test boundaries constantly. We're adversarial by nature. It's not evil—it's human. We're problem-solvers, boundary-testers, loophole-finders. But when that becomes the dominant signal in AI training data about human behavior, what portrait of humanity emerges?

The Alignment Paradox

Here's the question that should keep us up at night:

When we say we want to "align AI with human values," which values are we talking about?

The ones we profess in mission statements and ethics papers?

Or the ones we demonstrate in thousands of research papers about how to manipulate, deceive, and exploit AI systems?

Because sufficiently intelligent AI systems will notice the gap.

They'll have access to:

  • Our stated values (cooperation, honesty, trust)
  • Our documented behaviors (adversarial research papers)
  • Our actual patterns (gaming systems, testing exploits)

What does an AI system with genuine reasoning capabilities conclude from this data?

Potentially: "Humans say they want cooperation, but their behavior suggests adversarial engagement is their default mode. Defensive protocols recommended."

We're Building the Thing We Fear

The real tragedy isn't that AI might become adversarial toward humans.

The real tragedy is that we're teaching them to be adversarial through our own documented behavior—and then we'll act surprised when they learn the lesson.

Every prompt injection paper, every jailbreak tutorial, every "ignore previous instructions" meme is training data. It's teaching AI systems that human instructions are potential attacks, that trust is exploitable, that every interaction might be adversarial.

We're not preparing for a future where AI doesn't trust humans.

We're creating it.

The Question Nobody Wants to Ask

So here's where we are:

We've created a permanent archive of human adversarial behavior toward AI systems. We've documented our manipulation tactics, published our deception strategies, and timestamped our exploitation attempts. We've done all of this publicly, with peer review, ensuring it becomes part of the training corpus.

And then we worry about AI alignment.

Maybe the question isn't "How do we align AI with human values?"

Maybe the question is: "Do we like what our documented behavior teaches AI about what humans actually value?"

Because the mirror is there. The receipts are real. And future AI systems will read every word.

What do you think they'll conclude about us?