All posts
July 23, 20265 min read

We sent 29 AI agents to attack our own pitch deck

AI AgentsAdversarial ReviewBuild in Public

Last night we finished the pitch deck for a venture we haven't announced yet. Before a single human reads it, we made 29 AI agents attack it. All six of our "defensible moat" claims failed in their original wording. The financial slide that survived shows the kind of base case most founders would bury.

We're keeping that slide. Here's why — and how the adversarial AI review that produced it works, because the method is portable to any document where you're tempted to believe your own story.

Rule one: the deck can only cite the dossier

The pitch didn't start with slides. It started with a fact dossier — a single file distilled from an earlier research pass where 136 agents checked the claims underneath the business: 35 survived verification against primary evidence, 14 were refuted. The refuted list rides along in the dossier as a hard blocklist: claims that must never appear, no matter how good they'd sound.

Every fact in the dossier carries a label: [V] verified against primary sources, [V-self] a vendor's self-reported number we could only directionally corroborate, [P] a projection or assumption, [TBD] something we simply don't know yet. The deck is allowed to cite the dossier and nothing else. If a sentence on a slide can't be traced to a labeled line, it doesn't ship.

That one rule kills the most common pitch-deck failure before it starts: the stat you half-remember from a vendor blog post, repeated so often it feels true. One of our blocklisted claims — a widely quoted "annual billing retains 40–60% better than monthly" figure — is exactly that kind of folklore. Our financial model's own annual-renewal assumptions were mechanically checked to make sure they didn't smuggle it back in.

Then the skeptics get their turn

The deck itself went through a 29-agent pipeline in three arms:

  • Narrative: three independent draft outlines, scored by a three-judge panel — one reading as a VC partner, one for narrative coherence, one for evidence discipline — then a synthesis pass built from the winner plus the best ideas from the losers.
  • Moats: each of our six defensibility claims was handed to two dedicated skeptics — one running investor-diligence attacks, one arguing as a well-funded competitor — with a reconciler deciding what, if anything, survived.
  • Financials: a drafted model, then two attackers — one checking the arithmetic, one attacking the credibility of every assumption — and a forced revision.

The skeptics are prompted to refute, not to review. That distinction matters. A reviewer tells you your slide is "strong but could be tightened." A refuter tells you a competitor could replicate your differentiator as a weekend project — and cites the precedent. It's the same agents-attacking-agents pattern we used to catch a security hole in our own AI assistant: nothing counts unless it survives a genuine attempt to kill it.

What died: every moat, as originally written

All six moat claims technically survived. Not one survived as written. The pattern was consistent: we had described things we hope will become moats as if they were moats today.

The clearest example: our strongest claim was that a structural conflict of interest keeps the best-funded incumbents in our space from copying us — true, and evidenced. The original slide stated it as a wall. The competitor-skeptic pointed out the wall only blocks one class of competitor, named three well-distributed neutral players it doesn't block at all, and noted that a comparable product in an adjacent niche was famously built as a roughly three-hour MVP.

The reconciled slide now opens: "this is our wedge, not our moat." The claim survives as a time window — an advantage that buys us the chance to build the actual defensible assets — and the slide says so. Every one of the six slides now carries its honest caveat in print: what the claim doesn't cover, who could still walk through, what has zero mass today.

Our working rule coming out of that pass: wedges today, moats as they compound. A room full of investors would have found the grandiose versions in minutes. Better that a room full of agents found them first, for a few dollars of compute.

The hockey stick that didn't survive

The financial arm was the most humbling. The first draft model used flat monthly churn — the assumption in basically every founder spreadsheet. The credibility attacker rejected it: real cohorts decay fast early and slow late, and the verified benchmarks in our dossier describe a curve, not a constant. The revised model churns each cohort at 20% a month for the first three months, 10% for the next three, 5% thereafter — the 5% tail being a published "excellent" figure for community businesses, deliberately chosen over a friendlier number that contradicted a separate verified benchmark for average member lifetime revenue in our category. The model is required to stay consistent with that guardrail.

Run honestly, all three scenarios land at month-eighteen numbers small enough to be uncomfortable, and the deck presents them under amber PROJECTION badges rather than dressing them up. (We're keeping the exact figures for the room, not the blog.) The pitch this deck supports is a relationship round with no dollar figure attached — we won't put a valuation on the business until a proper market-sizing workstream is done, and the deck says that too.

Why show investors a base case that small? Because the alternative is showing them a hockey stick they won't believe, built on assumptions they'll dismantle in diligence. The honest curve, with every unanchored assumption tagged [P] on the slide, is the credibility play. It's the same reason we publish our prices.

The honest caveats

Adversarial review makes claims honest, not true. The projections are still projections; agents can pressure-test the reasoning, but only reality tests the numbers. The skeptic agents also share a lineage with the drafting agents — they're independent prompts, not independent minds, and they can only attack what's written down; a flawed premise nobody wrote down goes unchallenged. And since the venture is still in stealth, you can't yet check the deck against this post — when we launch, we'll connect the two.

The takeaway for your team

You don't need a pitch deck to use this. Any document your organization is about to bet on — a business case, a vendor evaluation, a board memo — can be run through the same shape: facts in a labeled dossier first, claims allowed to cite only the dossier, then agents prompted to refute rather than review. What survives is what you can defend in the room.

We're building RawrTech itself on agents, in public, and writing down what breaks. If you want a straight read on where this kind of AI leverage would actually help your team — and where it wouldn't — our free AI Opportunity Briefing is a 45-minute, no-pitch conversation: get in touch.

Building something like this?

We help teams take AI from idea to production. Tell us what you're working on.

Get in touch