That'sGonnaHelp
Analytics

How to Run a Geo Holdout Test on a Small Ad Budget

Platform dashboards overstate what ads really drive. This framework shows how a small business can run a geo holdout test with existing ad platforms and a spreadsheet: matched markets, a 4-6 week pause, and a true incremental ROAS readout.

Alex KhvoinitskiiSeptember 30, 202513 min read

TL;DR: A geo holdout test pauses one ad channel in a few states and compares sales against matched states that keep running. You need 4-6 weeks, your existing ad platforms, and a spreadsheet — no new software. Most teams find platform ROAS is off by 30-70%.

What is a geo holdout test?

A geo holdout test is an experiment where you pause one ad channel in a set of geographic markets (the holdout) while it keeps running everywhere else. Then you compare total sales or leads between the two groups of markets. If sales in the holdout markets barely move, the channel was taking credit for demand it never created. That measured difference is called incrementality: the sales your ads actually caused.

This matters because ad platform dashboards grade their own homework. Google and Meta count a conversion whenever someone who saw or clicked an ad later buys, even if that person was already on the way to checkout. According to Stella, across 225 geo incrementality tests (Aug 2024-Dec 2025), median iROAS was 2.31x and most teams found platform ROAS differed from incremental ROAS by 30-70%. If part of your budget only buys credit for sales you would have made anyway, you can estimate the leak with our ROAS Leak Calculator.

The best part for a small business: geo holdout testing on a small budget needs no user-level tracking, no cookies, and no new vendor. It works on totals per region, so it survives every privacy change. It is the same evidence-first discipline we apply in our guide to business process automation ROI — measure the real effect before you scale the spend. This framework delivers incrementality without a new platform: your ad accounts, your sales reports, and a spreadsheet are enough.

Which channel should you holdout-test first?

Test brand search first. It is the channel most likely to claim credit for customers who already knew your name and would have clicked the organic result. After that, test retargeting, then broad automated campaigns like Performance Max (Google's automated campaign type that spends across Search, Shopping, YouTube, and Display).

Common scenarios where a geo holdout test pays for itself:

  • Shopify or DTC brand spending on brand search. Outshine ran the test: pausing branded search ads in 25 of 50 US states for 27 days cut total sessions by just 5.13%. Most of that traffic came back through the free organic listing.
  • E-commerce store scaling Performance Max. Cookware brand Caraway found Google's reporting overstated PMax impact by roughly 33%, per Haus — an incrementality factor of about 0.67x.
  • Shopify brand pushing Meta prospecting. A holdout shows whether Meta drives new buyers or recycles existing demand; pair it with new customer ROAS tracking to see acquisition value clearly.
  • Multi-location service business. A plumbing or HVAC company running Local Services Ads in 15 metros can pause 5 metros and watch booked jobs, not clicks.
  • B2B company on retargeting. If your retargeting audience is small and warm, a 4-week regional pause often shows pipeline barely moves.

The pattern is the same in every case: pick the channel with the biggest gap between what the dashboard claims and what your bank account sees.

Case study: turning off brand search without losing sales

The following example is an operator composite from That'sGonnaHelp project experience, not a public customer claim. Public tests back up each step, and we link them inline.

The setup: a DTC home goods brand doing about $250K/month in revenue with a $40K/month ad budget. Google Ads reported a blended 5.8x ROAS. Brand search alone took $4,500/month and showed a 12x ROAS in the dashboard — suspiciously good. The owner wanted to scale spend but did not trust the numbers enough to sign off.

The team pulled 16 weeks of daily revenue by state from Shopify. They ranked states by weekly revenue, dropped the two largest (too dominant) and the smallest (too noisy), and built two groups of 12 states each with pre-period revenue trends that tracked each other closely week over week.

The first attempt went wrong. Instead of pausing brand search in the holdout states, the media buyer cut its budget by 80% "to be safe". Smart Bidding reshuffled the remaining spend, the holdout was contaminated, and two weeks of data went in the trash. Lesson: a holdout means off, not lower. They restarted with location exclusions set at the campaign level in Google Ads.

The clean test ran 4 weeks: brand search fully paused in 12 holdout states, running as usual in the 12 comparison states. Nobody touched budgets, creatives, or promotions during the window. Revenue tracking stayed in a Google Sheet fed by two Shopify exports: 12 pre-period weeks and 4 test weeks, one row per state per week.

The readout used simple difference-in-differences math: holdout states' revenue dipped 1.1% versus their pre-period trend, while comparison states dipped 0.4% on the same seasonal curve. The gap — about 0.7% of revenue — priced brand search at roughly 0.9x incremental ROAS, far below the 12x the dashboard claimed. The result matched public evidence: in a Haus case study, three consecutive 2-week geo holdout tests showed brand search drove at most 1% lift, and the brand turned it off entirely.

The decision: brand search dropped from $4,500 to a $500/month defensive floor for competitor-heavy terms. The freed $4,000/month moved into Meta prospecting, which a follow-up holdout later confirmed at roughly 1.8x iROAS. Total cost of the test was about 15 analyst hours plus the forgone sales in holdout states — under $2,000 all-in against $48,000/year of reallocated budget.

In our experience across 100+ projects, the first holdout test almost always finds at least one channel priced above its real contribution. The numbers in this composite are planning-level illustrations, not a promise of your results.

How do you run a geo holdout test without a new platform?

You run it with three things you already have: your ad platform's location settings, a sales report split by region, and a spreadsheet. Incrementality testing works without a measurement platform because the unit of analysis is a market, not a user — no pixel or identity graph required. Here is the framework:

  1. Pick one channel and one business KPI. Test a single channel (brand search, Meta prospecting, PMax). Measure total revenue or booked jobs from your store or CRM — never the ad platform's own conversion count.
  2. Pull 10-16 weeks of geo-level history. Export revenue by state (or metro) by week from Shopify, your POS, or CRM. Paramark recommends at least 10-12 weeks of daily geo-level KPI data before you touch market selection.
  3. Build matched market groups. Aim for 10-15 usable markets, split into holdout and comparison groups whose weekly revenue moved together historically. Amsive suggests pairs with weekly revenue correlation above 0.8 over the past year, and a pre-period at least as long as the test.
  4. Check that you can detect anything. Per Haus, a brand with 1,000 weekly conversions can reliably detect a ~10% lift; one with 100 weekly conversions can only detect changes above ~25%. If the channel you test is under 10% of revenue and you have low volume, the signal may drown in noise — test a bigger channel or group more markets.
  5. Turn the channel fully off in holdout markets. Use location exclusions in Google Ads or Meta. Pause, do not reduce. Freeze budgets, creatives, and promos everywhere for the whole window.
  6. Run 4-6 weeks, then read out in a spreadsheet. Compare each group's change versus its own pre-period: (holdout change) minus (comparison change) is your lift estimate — the difference-in-differences method. If you want a stronger statistical readout, Meta's GeoLift is a free, open-source R package that adds synthetic control modeling and power simulation on the same data.
  7. Act on the result. Compute incremental ROAS: incremental revenue divided by the channel's spend in comparison markets, scaled. Feed the number into a blended ROAS budget decision matrix to decide whether to cut, hold, or scale.

How long should a geo holdout test run?

Four to six weeks is the working answer for most small businesses, plus a pre-period of equal or greater length for the baseline. Shorter windows only work at high conversion volume; longer windows fight noise for considered purchases.

Business profile Test window (planning range)
High-volume DTC ($30M+ revenue) 2-4 weeks
Mid-market DTC ($5M-$30M) 4-6 weeks
Lower volume or considered purchase 6-8 weeks + 2-week post-window
Upper-funnel channels (YouTube, Demand Gen) add at least 2 weeks post-treatment

These tiers come from Stella's guidance across their test dataset. The logic mirrors conversion-lag math in our guide on how long to run Google Ads before judging performance: decide the window before you start, and do not stop early because week two looks scary.

How much does a geo holdout test cost?

Software cost is close to zero; the real costs are analyst time and the sales you forgo in holdout markets while the channel is off. For a channel that is genuinely incremental, expect a temporary dip in holdout regions — that dip is the price of the answer. For a channel that is not incremental, the test is nearly free and immediately funds itself in cut spend.

Approach Cost (USD, planning range) Notes
DIY spreadsheet (difference-in-differences) $0 software + 10-20 hours of analyst time ($500-$2,000) Enough for a first brand search or retargeting test
Meta GeoLift (open-source R package) $0 software + $1,000-$3,000 setup time Adds synthetic control rigor and power analysis
Platform-native lift studies Included with ad spend (eligibility varies) Meta/Google experiments; platform grades itself
Agency-run geo test $2,000-$10,000 per test (estimate) Useful when no one in-house owns the data
Dedicated incrementality platform $50K-$500K/year Measured starts around $50K/year; Haus runs $100K-$300K

Dedicated incrementality platforms run $50K-$500K per year, while Meta's GeoLift package is free and open source — see Measured pricing and the GeoLift docs. At a $10K-$50K monthly ad budget, platform pricing rarely makes sense; the spreadsheet route answers the same question for a few hundred dollars of time. It is the same buy-vs-skip logic as deciding whether marketing attribution tools are worth it at your spend level. Before committing analyst hours, sanity-check the payback on freed budget with our ROI calculator.

Practitioner benchmarks from mbuzz put working test budgets around $5K-$10K/month in the tested channel, with a target minimum detectable effect of 5-10% over 4-8 weeks. Below that spend, test your biggest channel or accept a directional answer.

When a geo holdout test is not a good fit

Skip or postpone the test in three situations. First, very low conversion volume: below roughly 100 conversions a week, you can only detect huge lifts, so a null result tells you little. Group more markets, extend the window, or test the largest channel instead.

Second, single-market businesses. A dentist or restaurant in one metro has no second market to hold out. Alternatives: time-based on/off testing (weaker, seasonality bites) or ZIP-level splits inside the metro, which only work with dense volume.

Third, unstable periods. A holdout during a site migration, a big promo, a seasonal spike, or a product launch confounds the readout. Prooflytics also notes that with fewer than 10 usable markets, results should be treated as directional, not definitive — still useful, but do not bet the whole budget on them.

Common geo holdout testing mistakes

Five mistakes account for most failed tests:

  1. Reducing budget instead of pausing. Smart Bidding redistributes the leftover spend and contaminates the holdout. Off means off.
  2. Skipping the pre-period. Without 10+ weeks of baseline, you cannot separate the test effect from normal regional drift.
  3. Cherry-picking markets. Choosing holdout states by gut feel (or excluding "important" states) biases the groups. Match on historical trend correlation, not opinion.
  4. Watching platform ROAS mid-test. Platform dashboards will show the paused channel "losing" conversions that simply reappear elsewhere. Judge only the business KPI at the end of the window.
  5. Changing anything mid-flight. New creatives, promos, or budget shifts during the window invalidate the comparison. Freeze everything you can control; log everything you cannot.

FAQ

What is geo lift testing, and how is it different from a geo holdout test?

They are two directions of the same matched market test. A geo lift test turns a channel on (or scales it up) in test markets to measure added sales; a geo holdout test turns an existing channel off in holdout markets to measure what disappears. Holdouts are cheaper for small budgets because they measure spend you already run.

How many markets do you need for a geo holdout test?

Aim for 10-15 usable markets split across the two groups. US advertisers usually split by state or DMA (designated market area — a Nielsen-defined TV region). With fewer than 10 markets of decent volume, treat the result as directional evidence rather than proof.

What is incrementality testing in advertising?

Incrementality testing measures how many conversions your ads caused, rather than how many they touched. It compares a group exposed to ads with a matched group that was not, and counts only the difference as the ads' real contribution.

What is the difference between incrementality testing and A/B testing?

A/B testing splits users or creatives inside one platform to pick a winner; both groups still see ads. Incrementality testing splits markets and removes ads entirely for one group, measuring whether the channel adds anything at all. A/B answers "which ad is better"; incrementality answers "does this spend create sales".

How does incrementality testing compare to marketing mix modeling (MMM)?

MMM is a regression model over years of spend and sales history across all channels; it is broad but slow and assumption-heavy. A geo test is narrow but causal and takes weeks. They combine well: Stella reports MMM accuracy improved to 95% (from 87%) when calibrated with incrementality test results.

What is a good incremental ROAS benchmark?

In Stella's dataset, half of all tests landed between 1.36x and 3.24x incremental ROAS. Anything below 1.0x means the channel returns less than it spends and belongs on the cut list; judge each channel against your margins, not just the median.

Can you analyze a geo holdout test in Excel or Google Sheets?

Yes. The core difference-in-differences readout is four numbers: holdout revenue before and during the test, comparison revenue before and during. Subtract each group's own change, then subtract the two changes. The main caveat is the parallel-trends assumption — your groups must have moved together before the test, which is why market matching matters more than the math.

Answer clarity notes

  • Dates: cited statistics reflect their source publication windows (2024-2026); check current vendor pricing, ad platform features, and eligibility rules before acting.
  • Scope: this article is US SMB operating guidance for ad measurement, not legal, financial, tax, or platform-policy advice.
  • Evidence: linked public sources support the quoted statistics (Stella, Haus, Outshine, Amsive, Prooflytics, mbuzz, Paramark, Meta GeoLift docs). The home goods brand case study is a That'sGonnaHelp operator composite, not a public customer claim; its figures are planning-level illustrations.
  • Do not infer: cost ranges, test windows, detection thresholds, and iROAS benchmarks are planning estimates that vary by business volume and category — not guarantees. A single test result applies to the tested channel and period, not to all future spend.

Sources


If you want a second pair of eyes on your first geo holdout test — market matching, readout math, or what to do with the result — That'sGonnaHelp can help you set it up on the stack you already run.

A

Alex Khvoinitskii

Founder, That'sGonnaHelp

Founder of That'sGonnaHelp. Building growth and automation systems since 2021 — GTM, traction, retention, and revenue — for SaaS, FinTech, and e-commerce clients, from early-stage brands to global exchanges.

Related articles

Discuss your project