MEASUREMENT

How to measure whether your AI is actually working

The short answer

You need a holdout: a randomly chosen slice of your own deals or accounts that the AI deliberately does not touch, running at the same time as the slice it does. Compare the two and the difference is the AI's effect. Without a concurrent control group, every other number available to you, a vendor ROI calculator, a before-and-after comparison, a benchmark from someone else's data, is an estimate of what the tool might do rather than a measurement of what it did.

The four mistakes that void the result

A holdout is easy to describe and easy to get wrong in ways that leave you with a number that looks rigorous and means nothing. These are the four that matter, roughly in order of how often they happen.

1. Randomizing reps instead of deals or accounts

This is the one that quietly ruins most rollouts. Put half your reps on the AI and half without it, and the intervention leaks through the human: a rep in the treatment arm reads a card that says chase the champion when they go quiet, learns the lesson, and then applies it to the deals in their own held-out group too. The control arm is contaminated by the treatment, so the measured lift is biased toward zero by an amount you cannot recover. It looks like the AI did nothing. Randomize the deal or the account instead, and the rep is no longer the carrier.

2. Analyzing at a different grain than you randomized

If you assigned accounts, one account is one row. If you assigned deals, one deal is one row. Randomize accounts and then analyze individual emails and your sample looks several times larger than it is, your confidence interval shrinks to something you have not earned, and a null result reads as a win. Keeping the unit of randomization and the unit of analysis identical is what lets you skip clustering corrections entirely, because there is nothing nested to correct for.

3. Peeking, then stopping when the number looks good

Checking the result repeatedly and calling it as soon as it crosses significance is optional stopping, and it manufactures significance out of noise. With enough looks, a coin will eventually appear biased. The fix is to write down, before the experiment starts, which endpoint is primary and at what point you will read it, then only let the verdict render at that point. Watching the point estimate and its interval move in the meantime is fine and useful. Acting on a crossing that arrives early is not.

4. Comparing against a vendor calculator instead of a control

An ROI calculator takes your inputs and multiplies them by the vendor's assumed improvement. Whatever it prints is the assumption, restated. It cannot tell you what your pipeline would have done without the tool, which is the only comparison that answers the question. Neither can a before-and-after: your pipeline in Q1 and your pipeline in Q3 differ by seasonality, headcount, pricing, and market, and the AI is one of many things that changed. You need a control group running at the same time as the treatment.

How SealDeal measures it

SealDeal runs two independent holdout arms, because outbound and the next-best-action engine are separate interventions and averaging them would hide which one worked.

The account arm measures grounded outbound. Accounts are randomly assigned, and reply and meeting outcomes are compared between arms. The deal arm measures the next-best-action engine: deals are randomly assigned, and a held-out deal simply never receives a card, so the comparison is a card against no card rather than one flavor of AI copy against another.

Both arms randomize at the grain they analyze, so one deal or one account is one row and no clustering correction is needed. Both use Fisher's exact test, which stays correct at the small samples real pipelines produce, reported with a Newcombe 95% confidence interval on the difference between arms. The primary endpoint and the analysis horizon are fixed before the experiment starts, and the significance verdict renders only once that horizon is reached, so stopping early on a favorable peek is not possible. The point estimate and its interval are visible throughout.

Rep-level randomization was considered and rejected for the contamination reason above. The deal arm accepts a smaller, bounded cost instead: a rep may carry a generic habit across their own deals, but not the specific grounded content of a card they never saw.

None of this is on by default. With no active experiment, everything is in the treatment arm and nothing is withheld.

Common questions

What is a holdout group in sales?+

A holdout is a randomly selected slice of your own deals or accounts that the tool deliberately does not touch, kept running at the same time as the group it does. Because assignment is random and the two groups run concurrently, the difference in outcome between them is attributable to the tool rather than to seasonality, headcount changes, or market conditions. It is the sales version of a control group.

How big does the holdout need to be?+

Big enough that a real effect would be visible, which depends on your baseline conversion rate and the size of the effect you care about, not on a fixed percentage. The more useful discipline is to pick a leading endpoint that accrues in weeks rather than quarters, so you learn something before the full sales cycle completes. SealDeal uses stage advancement as that leading endpoint and closed-won as the slower gold endpoint.

Does holding deals back cost me pipeline?+

That is the honest tradeoff, and it is why the holdout is off by default rather than always on. What you buy with it is the ability to know. Without a control group, you can never distinguish a tool that works from a tool you are paying for during a good quarter, which is a more expensive mistake over any real time horizon. The cost is also bounded: a held-out deal still gets your rep, your CRM, and your process, just not the AI card.

How does SealDeal measure lift?+

Two independent arms. The account arm measures grounded outbound: accounts are randomly assigned, and reply and meeting outcomes are compared. The deal arm measures the next-best-action engine: deals are randomly assigned, and a held-out deal simply never receives a card. Both use Fisher's exact test, which stays correct at the small samples real pipelines produce, with a Newcombe 95% confidence interval on the difference. The primary endpoint and the analysis horizon are both fixed before the experiment starts.

What do I actually see while the experiment is running?+

One of three states, deliberately including the unfinished ones. Not started, meaning nothing has accrued yet. Running, which shows how many deals or accounts have accrued per arm against the horizon you pre-registered, so an experiment that will take months to conclude is not invisible in the meantime. Concluded, where the verdict and the confidence interval render. The verdict is withheld until the horizon; the numbers are not.

Is the holdout on by default?+

No. With no active experiment, every deal and account is in the treatment arm, nothing is withheld, and the lift analysis is empty. You turn it on when you want a measured answer. That default is deliberate: a holdout that ran without your knowing would be withholding your AI from your deals without your consent.

Does SealDeal publish its own lift numbers?+

No, because it does not have any yet. SealDeal is pre-launch, so there is no measured lift figure, no case study, and no customer count to quote. The claim on this site is about the instrument, that your lift is measurable on your own pipeline, not about a result. Any specific percentage you see attributed to SealDeal is not from us.

SealDeal is pre-launch. This page describes the measurement instrument, not a measured result: there is no published lift figure, case study, or customer count to quote, and any specific percentage attributed to SealDeal did not come from us.

Related: vs HubSpot Breeze · vs Salesforce Agentforce · Security and compliance · Outbound plays