Two identical city grids at night, one lit and one dark, illustrating Stillwater Media's guide to evaluating an independent agency for CTV incrementality testing.
Measurement & Attribution

Independent Agency for CTV Incrementality Testing

Stillwater Media•2026-09-25•13 min read

A platform reporting its own lift is not the same thing as an independent measurement of incrementality. The difference is the holdout.

Independent Agency for CTV Incrementality Testing: What "Platform-Agnostic" Actually Means

Finding an independent agency for CTV incrementality testing has become a harder search than it should be, because nearly every media agency now claims to "prove incrementality" while very few actually run the test that word requires. For a CMO or growth lead spending $50,000 to $200,000 a month across CTV and programmatic, the distinction matters enormously: a platform telling you it drove incremental lift, using its own attribution and its own reported baseline, is not an independent measurement of anything. It is the platform grading its own homework. Independent, in the sense that actually protects a budget decision, means a party with no financial stake in which channel gets credit is running the test, controlling the holdout, and reporting the result — whether that result flatters the media plan or not.

What "Independent" Actually Means in Incrementality Testing

The word gets used loosely enough that it is worth defining precisely before evaluating anyone against it. An independent incrementality test has three characteristics that a platform-reported lift study does not.

  • No financial interest in the outcome. The party designing and reading the test does not have its own media spend or platform relationship riding on the result showing a positive lift.
  • Holdout control outside the platform's own systems. The exposed and holdout groups are defined and enforced by the agency or a third-party measurement partner, not by a platform's internal experimentation tool that also decides what counts as "exposed."
  • A pre-registered methodology. The test design — market selection, holdout size, duration, success metric — is set before the campaign runs and reported regardless of outcome, rather than a post-hoc analysis run only when the numbers look favorable.

A platform's own conversion lift or brand lift study can be a useful input, but it fails all three tests: the platform has a direct interest in the result, it controls both the exposed and the measurement layer, and the methodology is rarely disclosed in enough detail to audit. That is the specific gap an independent, platform-agnostic measurement partner is meant to fill.

Why Platform-Reported Lift Isn't Incrementality

Platform-reported lift studies compare users who happened to see an ad to users who did not, using the platform's own audience assignment, and then attribute the difference in conversion rate to the ad. The problem is selection bias: the users a platform's delivery algorithm chooses to show an ad to are frequently already more likely to convert, because the algorithm is optimizing delivery toward exactly that population. A true holdout removes that bias by defining the control group independently of the platform's delivery decisions — typically at the geographic market level, where an entire region is withheld from exposure rather than individual users within a shared market, which also avoids the cross-contamination that happens when "exposed" and "control" users live in the same household or see the same organic and earned-media environment.

What a Geo-Holdout Test Actually Requires

A defensible geo-holdout test has specific mechanical requirements, and this is the part of the evaluation where most agencies claiming incrementality expertise fall short in practice.

  • Matched market pairs. Exposed and holdout markets need to be statistically matched on population, historical sales trend, seasonality and existing media weight before the test starts, not selected for convenience.
  • A sufficient holdout size. Typically 15 to 25 percent of comparable markets, large enough to detect the minimum lift that would justify the spend, calculated in advance rather than assumed.
  • A clean media boundary. Other channels — paid search, paid social, email — need to run identically in exposed and holdout markets during the test, or their variation has to be explicitly controlled for, or the test measures channel interaction rather than CTV's own effect.
  • First-party outcome data, not platform-reported conversions. The read should come from CRM, point-of-sale or e-commerce data matched back to exposed and holdout geographies, not from the ad platform's own pixel.
  • A pre-set test duration tied to the actual sales cycle. A high-consideration purchase with a 30-to-90-day cycle needs a test window that runs at least that long past the media flight, or the read captures only the fastest-converting fraction of influenced buyers.
  • Statistical significance reporting, including a null result if that's what the data shows. A credible measurement partner reports a "no detectable lift" result exactly as readily as a positive one.

Marketing Mix Modeling vs. Incrementality Testing: When to Use Each

These two methodologies get talked about as competitors when they are better understood as complementary tools that answer different questions on different timelines.

DimensionGeo-holdout incrementality testingMarketing mix modeling (MMM)
Question answeredDid this specific channel or campaign cause a measurable lift versus a true controlHow has spend across all channels historically correlated with revenue over time
Data requiredFirst-party outcome data by geography, a controlled holdoutHistorical spend and outcome data across all channels, typically 18-24+ months
Time to first read6 to 12 weeks per testRequires a substantial historical data set before the first model is usable
GranularityPrecise for the specific channel or tactic under testDirectional at the channel level; weaker at the campaign or creative level
Best forValidating a new channel like CTV, or settling a specific budget-reallocation decisionOngoing, top-down budget-allocation guidance once a program is mature and has enough history
Common failure modeUnderpowered test (too few markets, too short a window) produces a false negativeModel collinearity between correlated channels produces unstable or misleading channel-level attribution

The practical answer for most brands evaluating a new channel or a specific agency claim is to start with geo-holdout testing, because it produces a fast, causal answer to a specific question, and to layer in marketing mix modeling once there is enough historical data across a mature, multi-channel program to make a model-based approach reliable.

What to Ask Before Hiring an Agency That Claims to "Prove Incrementality"

  • "Who controls the holdout — you, or the platform?" If the answer is the platform's own experimentation tool, this is platform-reported lift with different branding, not an independent test.
  • "What happens if the test shows no lift?" An agency that cannot describe how it would report and act on a null result is not running a pre-registered test; it is running an analysis designed to find a positive number.
  • "What outcome data does the read use — your pixel, or my CRM/POS data?" First-party outcome data matched back to geography is the standard; platform-side conversion data is not.
  • "How many markets, and how were exposed and holdout markets matched?" A specific, defensible matching methodology should be describable in one or two sentences; a vague answer is a signal the test wasn't designed rigorously.
  • "Is there a third-party measurement partner involved, or is the agency both running the media and grading its own result?" Not disqualifying on its own, but it changes how much independent verification the result carries.
  • "Can I see a redacted example of a past holdout test design and result, including one that didn't work?" An agency with real experience here should have at least one test that produced a negative or inconclusive result and can explain why.
  • "What's the minimum spend and market count you need to run a test that would actually be readable for my business?" A confident, specific answer — not "it depends" with no follow-up — indicates the agency has actually sized tests before.

What Incrementality Testing Should Cost and How Long It Should Take

Monthly media spendRecommended measurement investmentTypical test durationMarkets needed for a clean holdoutWhat's realistically answerable
$50,000 to $100,000/month3% to 6% of media spend8 to 10 weeks6 to 10Channel-level incrementality (CTV vs. no CTV)
$100,000 to $200,000/month3% to 5% of media spend8 to 12 weeks10 to 15Channel-level incrementality plus a first creative or daypart read
$200,000 to $500,000/month2% to 4% of media spend10 to 12 weeks, then rolling12 to 20Ongoing channel and tactic-level incrementality, feeding a maturing MMM
$500,000+/month1.5% to 3% of media spendRolling, continuous20+Full measurement stack: rolling holdouts plus MMM cross-validation

Measurement spend below roughly 2 percent of media spend at any tier is a warning sign that the test is being run as a checkbox rather than as a decision-grade instrument; below that level, agencies typically cut corners on market count or test duration in ways that quietly undermine statistical power.

Certifications and Third-Party Validation: What They Actually Verify

Badges like a Measured certified partner designation or similar third-party measurement certifications verify that an agency has demonstrated proficiency using a specific measurement platform's methodology and tooling — they are a meaningful signal of technical competence, but they are not, on their own, proof that a given agency runs independent tests for its own clients rather than platform-reported studies. The certification answers "can this agency operate this measurement tool correctly." It does not answer "does this agency structure holdouts that protect the client's interest over its own media commission," which is the question that actually matters and has to be answered through the due-diligence questions above, not the badge alone.

Common Mistakes When Evaluating Incrementality Claims

  • Accepting a platform-reported lift study as "incrementality" without asking who controlled the holdout.
  • Hiring based on a case study with no visible methodology, market count or test duration.
  • Assuming a bigger agency brand name implies a more rigorous test design. Test rigor is a function of methodology and market count, not agency size.
  • Skipping the "what if it shows no lift" question. This is the single fastest way to identify an agency that only reports favorable results.
  • Treating MMM output as a substitute for testing a brand-new channel. MMM needs history a new channel doesn't have yet; it's the wrong tool for a first CTV test.
  • Under-funding the measurement budget below the level needed for statistical power, then blaming the channel when the test is inconclusive.

A Due-Diligence Sequence for Evaluating Independent Incrementality Testing Partners

  • Define the specific decision the test needs to inform — a new-channel go/no-go, a budget reallocation, or ongoing optimization — before evaluating any agency.
  • Ask each candidate agency the seven due-diligence questions above and document the specific, non-generic answers.
  • Request a redacted example of a past test, including its market count, duration, matching methodology and at least one result that wasn't uniformly positive.
  • Confirm whether the read will use first-party CRM or POS data matched to geography, not platform-reported conversions.
  • Size the test against your actual monthly spend using the cost and duration table above, and flag any proposal that comes in meaningfully under the recommended measurement investment.
  • Confirm the test is pre-registered — methodology and success metric locked before launch — rather than designed after the campaign to explain results already in hand.
  • Set expectations in writing for how a null or negative result will be reported and what decision it will trigger, before the test begins.
  • Run the first test at the minimum viable scale from the table above, and treat it as a decision instrument, not a marketing claim to be repeated regardless of outcome.

Where Stillwater Media Fits

Stillwater Media runs geo-holdout incrementality testing as a standard part of every engagement, not as an add-on service — holdouts are built and controlled independently of the platforms we buy media on, read against first-party CRM and sales data, and reported whether the result is favorable to our own media recommendation or not. We work with brands spending $50,000 to $500,000-plus a month across premium CTV and programmatic who need a measurement partner that has no financial interest in which channel wins the comparison. We take a limited number of new engagements each quarter. If you've been told your media is incremental and want to see the holdout that proves it, [apply to work with us](https://stillwatermedia.io/apply).

---

*Stillwater Media is a selective performance media agency for luxury and high-consideration brands, based in Charlotte, North Carolina and working nationally. We plan and buy premium CTV, programmatic, digital out-of-home, streaming audio and YouTube Select for clients including JetLinx, W Hotels, PXG, FLY Exclusive and Financial Independence Group, and we measure everything against holdouts rather than platform-reported lift. Signal. Strategy. Scale.*

Ready to discuss your strategy?

Discover how our approach can transform your brand's media performance.

Related insights