Stillwater Media guide to geo experiment design - a darkened analytics room with an illuminated United States metro-market map showing lit test regions and deliberately dark holdout markets used in advertising incrementality testing.
Measurement & Incrementality

Geo Experiment Design: Markets, Duration, Detectable Lift

Stillwater MediaSeptember 12, 202618 minutes

The dark markets are the point: a geo experiment only produces an answer when a brand is willing to deliberately stop advertising somewhere for long enough to measure what changes.

Sound geo experiment design is the difference between a measurement program that settles arguments and one that manufactures them. The pattern we encounter most often is a brand that ran a geo test, saw a result labeled "not statistically significant," and concluded the channel does not work - when the test as designed could never have detected the effect the brand plausibly had. Ten test markets, four weeks, a 40% spend reduction, and a conversion series that swings 9% week to week will return a null result almost regardless of the truth. The experiment did not fail to find lift. It was never capable of finding it.

For luxury and high-consideration advertisers, this problem is structural rather than occasional. User-level holdouts are unavailable across most premium CTV and audio inventory, conversion volumes are low by design, and the purchase cycle runs longer than the typical test window. Geographic experimentation is usually the only credible route to a causal number - which makes getting the design math right the entire game. This piece covers the three experimental structures, how market count and baseline volatility set your minimum detectable effect, how long a test actually needs to run, how to match markets, and the seven errors that produce false nulls.

What a geo experiment actually measures

A geo experiment partitions geography into treatment and control groups, changes media spend in one group, and attributes the difference in outcomes to that change. The output is not a modeled contribution or a credited conversion path - it is a measured difference between a world where the media ran and a constructed estimate of a world where it did not. That is why it sits above multi-touch attribution in the evidentiary hierarchy: attribution allocates credit among touchpoints that all occurred, while a geo experiment establishes whether the outcome would have occurred anyway.

Three structures dominate practice, and they are not interchangeable.

DesignStructureBest forTypical durationMain weakness
Matched-market holdoutPaired treatment and control markets selected on pre-period similarityChannel-level "does this work at all" questions6–12 weeksRequires enough comparable markets to pair
Synthetic controlA weighted composite of untreated markets constructed to track each treated market's pre-period seriesBrands with few markets or poor natural pairs8–16 weeksNeeds 12+ months of clean pre-period data
Switchback / time-basedSame markets alternate on and off in randomized intervalsShort-cycle categories with fast response8–16 weeks in 1–2 week blocksInvalid where carryover exceeds block length

For most luxury advertisers, matched-market and synthetic control are the viable options. Switchback designs assume the effect decays within the block interval, which is false for a category where the consideration window runs sixty to two hundred days. Running one against a private aviation or luxury residential program produces contaminated blocks and a meaningless average.

Geo experiment design inputs: what sets minimum detectable effect

Minimum detectable effect (MDE) is the smallest true lift your test has a reasonable probability of detecting. It is determined before a single impression serves, by four inputs.

1. Baseline volatility. The week-to-week coefficient of variation in your outcome series within each market. This is the dominant term and the one most often ignored. A brand whose weekly qualified inquiries per market swing 6% has a fundamentally different test than one swinging 22%. Compute it from at least twelve months of history, after removing known seasonality.

2. Market count. Precision improves roughly with the square root of the number of geographic units. Going from 10 to 40 markets does not cut your MDE by four - it cuts it by about half. This is why market count is a costly lever and duration is often the cheaper one.

3. Test duration. More weeks means more observations per market and more averaging of idiosyncratic noise, with the same square-root relationship. Duration is usually the least expensive way to buy power, up to the point where seasonality and market drift begin to erode the match.

4. Spend delta. The size of the change between treatment and control. A 100% holdout - media fully dark in control - produces a far larger signal than a 30% budget reduction. A 30% delta requires roughly three times the sample of a 100% delta to detect the same underlying effect.

Here is the practical consequence, using a baseline coefficient of variation of 12%, a full on/off delta, 80% power, and a 90% confidence threshold - the design parameters we consider the sensible default for a commercial media decision.

Matched markets per arm4-week test8-week test12-week test
816%–22% MDE11%–15% MDE9%–12% MDE
1511%–15% MDE8%–11% MDE6%–9% MDE
259%–12% MDE6%–8% MDE5%–7% MDE
407%–9% MDE5%–6% MDE3.5%–5% MDE

Read that table against the effect you are actually trying to find. If premium CTV is contributing a true 7% lift to qualified demand - a genuinely good result for an upper-funnel channel in a high-consideration category - then the 8-market, 4-week test at the top left has roughly a one-in-four chance of detecting it. Three brands out of four would run that test and conclude the channel does nothing.

Working the math backwards

The correct sequence is to start from the decision, not the calendar.

  1. State the effect size that would change your behavior. If a 5% lift in qualified demand would justify sustaining the investment and 2% would not, your MDE target is 5% or better. Write this down before the test.
  2. Measure baseline volatility from history. Twelve to twenty-four months, seasonality removed, computed at the geographic unit you intend to test.
  3. Solve for the market-and-duration combination that reaches your MDE. Use the table above as a first pass, then confirm with a simulation against your own historical series.
  4. Check whether you have enough comparable markets to support it. This is usually the binding constraint, and it is where most designs quietly break.
  5. Price the holdout. Multiply the control-market spend by the test duration to get the media cost, then add the estimated foregone revenue.
  6. Only then set dates.

A worked example. A luxury residential developer sells across 31 metros, with weekly qualified tour requests showing a 14% coefficient of variation per market. They want to detect a 6% lift with 80% power. At 14% volatility, a 6% MDE requires roughly 15 matched pairs at 12 weeks or 25 pairs at 8 weeks. With 31 markets, 15 pairs is impossible - pairing consumes two markets each, so 15 pairs requires 30 markets and leaves nothing for the treatment group to scale into. The workable design is therefore 12 control markets against 19 treatment markets using a synthetic control weighting, run for 14 weeks. The additional two weeks compensate for the smaller control arm.

Notice what happened: the honest answer was a longer test, not a smaller one. Brands routinely make the opposite substitution.

Market matching in geo experiment design: the thresholds that matter

A matched-market design is only as good as its pairs. Four criteria, applied in order:

  • Pre-period outcome correlation of 0.80 or higher between paired markets across at least 52 weeks, computed on the actual outcome metric rather than on population or spend. Below 0.75, the pair contributes noise rather than precision.
  • Comparable scale, within roughly a 2:1 ratio of baseline volume. Pairing a market producing 400 weekly outcomes with one producing 30 lets the small market's variance dominate.
  • No spillover. Adjacent metros with shared media markets, commuting overlap, or overlapping out-of-home footprints will contaminate the control. This is a real constraint for DOOH and location-based layers and for any streaming inventory sold at a regional rather than metro level.
  • Comparable affluent composition. For luxury advertisers, matching on total population is close to meaningless. Match on the density of the qualified audience - the same wealth-based segmentation that defines the media buy should define the market pairing.

Randomize the assignment within qualified pairs rather than selecting which market goes dark. Analyst-chosen controls are the most common source of bias in commercial geo testing, and the direction of the bias is always flattering.

Duration, carryover, and the cool-down window

Three duration components are frequently collapsed into one, which is why so many tests measure the wrong interval.

Pre-period. Twelve months minimum, twenty-four preferred, used to establish the match and the baseline. Synthetic control methods are effectively unusable below twelve months.

Burn-in. The first one to three weeks of a holdout do not reflect the absence of media, because prior exposure is still converting. In a category with a 90-day consideration cycle, demand in a newly dark market continues at near-normal levels for several weeks on residual awareness. Excluding a two-to-three-week burn-in from the analysis window is standard practice for high-consideration advertisers and is the single most common omission we see.

Measurement window. The interval actually analyzed, which must be long enough to clear burn-in and still deliver the observations your MDE requires. A test billed as "eight weeks" with a three-week burn-in is a five-week test, and its real MDE is materially worse than the planner assumed.

Add a cool-down of four to six weeks after restoring spend before beginning another test in the same markets. Running back-to-back experiments without one contaminates the second test's pre-period with the first test's after-effects.

Reading the result: intervals, not point estimates

A geo experiment does not return "incremental ROAS was 3.4." It returns a distribution. The correct output is a point estimate with a confidence interval, and the interval is the part that governs the decision.

Consider two results from the same category:

  • Test A: iROAS 3.1, 90% CI [2.4, 3.8]
  • Test B: iROAS 4.6, 90% CI [0.3, 8.9]

Test B has the better headline number and is nearly worthless. Its interval spans outcomes from "destroyed value" to "extraordinary," which means the test lacked the power to constrain anything. Test A, with a tighter interval that excludes break-even, actually supports a budget decision. When an experiment returns an interval wider than roughly ±40% of the point estimate, report it as inconclusive and redesign rather than reporting the midpoint.

Interpret a null result correctly as well. "We did not detect a significant effect" means the true effect is probably smaller than your MDE - not that it is zero. If your MDE was 14%, a null result is entirely consistent with a real and valuable 9% lift. State the MDE alongside every null finding; a null without its MDE is not a finding at all. This distinction is the core of what separates incrementality from attribution in practice.

What the test actually costs

The holdout is not free, and pretending otherwise causes brands to under-resource tests and then distrust them.

For a program spending $200,000 monthly across 30 markets, a 12-week holdout in 12 markets removes roughly 40% of spend for three months - about $240,000 in media redeployed to the treatment arm, plus foregone demand in the dark markets. At a true 6% lift, the cost of the darkness is the revenue that 6% would have produced over twelve weeks in 40% of the footprint.

Against that, price the alternative. A brand spending $2.4M annually with no causal read is making a full-year allocation decision on modeled evidence. A single well-powered experiment that moves the credible range of iROAS from "somewhere between 1 and 6" to "2.4 to 3.8" is worth a substantial multiple of its cost, and the result informs budgeting for several quarters - which is why we generally recommend one to two properly powered tests per year rather than four underpowered ones. The same logic drives how we approach media mix optimization for clients with multiple channels in play.

Seven geo experiment design errors that produce false nulls

  1. Running the test before computing the MDE. If you cannot state the smallest lift your design can detect, you do not have a design.
  2. Using a partial spend reduction when a full holdout was feasible. A 30% delta triples your required sample for the same detectable effect.
  3. Selecting control markets by judgment. Randomize within qualified pairs; analyst selection biases toward the expected answer.
  4. Failing to exclude a burn-in period. Residual awareness in high-consideration categories keeps dark markets converting for two to three weeks.
  5. Matching on population rather than on qualified audience density. For luxury advertisers these are different variables, and only one of them is relevant.
  6. Ignoring spillover between adjacent metros. Contaminated controls compress the measured difference toward zero, manufacturing a null.
  7. Reporting the point estimate without the interval. A 4.6 iROAS with a [0.3, 8.9] interval is a number, not evidence.

How Stillwater Media designs geo experiments

We compute minimum detectable effect from the client's own historical series before proposing a structure, and we will say plainly when a client's market count and volatility make a credible test impossible at their current spend - in which case the honest recommendation is a synthetic control design, a longer window, or a different question. We randomize assignment within qualified pairs, exclude burn-in explicitly, and report intervals rather than headline numbers. When a test returns a null, we report the MDE alongside it so the result is interpreted as a bound rather than a verdict.

If you are spending meaningfully on premium CTV, programmatic or audio and want a causal read you can defend to a board, apply to work with Stillwater Media. We accept a limited number of engagements each quarter.

Ready to discuss your strategy?

Discover how our approach can transform your brand's media performance.

Related insights