All Measure What Matters

Testing Without a Data Team

The holdout method: a real experiment any practice can run with a calendar and honesty. How to pause, what to watch, how long to wait, and reading the result without fooling yourself.

The one question — would it have happened anyway? — has an answer that doesn't require believing anyone: run the world both ways and compare. Big companies do this with test markets and statisticians. A practice does it with a calendar: turn the thing off, watch what changes, turn it back on if you miss it. That's the holdout method, and a small practice can run it better than most corporations — fewer moving parts, cleaner signal, and nobody's quarterly bonus riding on the answer coming out right.

The holdout, complete: pick one spend; note the honest baseline (the last 3–6 months of the real numbers it claims to move); pause it fully for a meaningful window (usually 4–8 weeks — long enough for your enquiry rhythm to speak); change nothing else; then compare against baseline, not against hope. If the needle didn't move, you've learned the spend's true price. If it dropped, you've learned the spend works — restart it with confidence you couldn't have bought any other way.

Running it clean

The method is simple; the failure modes are human:

One variable. Pause the ads or the directory or the retainer — never two at once. Overlapping tests answer nothing. A practice can cycle through its whole marketing budget in a year, one clean test at a time.

Whole weeks, fair seasons. Enquiries have weekly rhythms; test in full-week blocks. And don't pause the tax accountant's ads in April and conclude they were worthless in the quiet of June — compare like seasons, or use last year's same-months as the baseline where your records allow.

Decide the verdict line before you start. Write it down: "if enquiries stay within the normal range, the spend stops permanently." Deciding after seeing the data is how motivated reasoning wins — pre-commitment is the entire discipline, and it's free.

Expect the wobble. Enquiries are lumpy; a quiet fortnight happens with or without the ads. This is why the window is weeks and the comparison is baseline range, not last month's exact number. Small practices read direction and magnitude, not decimals — a spend whose absence is invisible inside normal noise is, for decision purposes, not working.

Beyond pausing: the other cheap experiments

The holdout's logic — compare against the counterfactual — powers smaller tests too. The A/B you can actually run: two months of quote follow-up with the sequence, two without, same numbers watched. The geographic split, if you serve multiple areas: run the spend in one, not the other, and let the areas referee. The before/after with teeth: any change (the booking link, the price on the page) becomes an experiment the moment you note the baseline first and the date of the change — which is the difference between "we redesigned and things feel better" and knowing.

The identity worth adopting: a practice that changes things on dates, against baselines is running experiments continuously without ever calling them that. The spreadsheet row that says "March 4: added deposits at booking" next to the no-show column is a controlled trial wearing work clothes.

Questions practices actually ask

Isn't pausing working spend risky? You're testing precisely because you don't know it's working — the risk being managed is permanent waste, priced against a bounded few weeks. Genuinely can't afford the test? Note what that says about how certain you actually are.

My agency says pausing will 'reset the algorithm' and hurt long-term. Sometimes partially true (ad platforms do relearn), mostly deployed as test-repellent. The reply: "then let's test somewhere contained — one region, one campaign." A partner who resists every falsifiable version of their value has answered the question.

What can't the holdout test? Slow-compounding assets: pausing content or reviews for six weeks shows nothing — their effects arrive and decay on year timescales. Holdouts are for spends with claimed ongoing effect: ads, directories, retainers, subscriptions.

Do I need statistical significance? You need decision-grade honesty, not journal-grade proof: pre-committed verdict lines, fair windows, baseline ranges. A practice choosing where next month's dollars go can act on "clearly no visible effect" — waiting for p-values is how the question never gets asked at all.


Part of Measure What Matters — the instrument panel.

Explore further

Based on the themes in this article, you might find these topics interesting.