A CMO at a luxury hospitality brand once asked us a fair question: "If someone watches a 30-second ad on Hulu and then books a room two weeks later by typing our brand name into Google, how do I know the CTV ad had anything to do with it?" The honest answer is that last-click attribution will never tell you. Brand lift measurement CTV campaigns require an entirely different measurement framework — one built around comparing exposed and unexposed groups, not tracking clicks that connected TV environments mostly don't produce.
This is the central tension in CTV measurement: it's one of the highest-impact channels for luxury and high-consideration brands, reaching audiences in a premium, lean-back environment that builds the kind of brand consideration search and social can't replicate — yet it's also the channel most likely to get under-credited in attribution models built for direct-response. This post walks through how brand lift is actually measured for CTV, what benchmark ranges to expect, and how to combine methodologies into something a CFO will trust.
Why CTV Breaks Traditional Attribution
Connected TV ads run on a television, in a living room, often on a second screen from where the eventual conversion happens. There's no click. View-through attribution — crediting a conversion to anyone who saw an ad without clicking — exists as a workaround, but it's notoriously unreliable: it tends to over-credit CTV for conversions that would have happened anyway, simply because CTV reaches such a broad swath of the population that some viewers convert by coincidence.
The result is a measurement gap that causes real damage: brands that evaluate CTV using the same last-click or view-through logic as search end up systematically underfunding it, even when it's driving meaningful brand consideration and downstream conversions. Brand lift measurement exists to close this gap by measuring what CTV is actually good at — shifting awareness, consideration, and purchase intent — using methodologies designed for that purpose.
The Three Core Methodologies
1. Survey-Based Brand Lift Studies
This is the most direct method: serve a brand awareness or favorability survey to two groups — people who were exposed to your CTV campaign, and a matched control group who weren't — and compare the results. Platforms like Disney+, Netflix's ad-supported tier, and Prime Video, along with measurement partners like Innovid and DoubleVerify, offer survey-based lift products that integrate directly with CTV buys.
Typical metrics measured include:
- Ad recall — "Do you recall seeing an ad from [Brand]?"
- Brand awareness — "Have you heard of [Brand]?"
- Brand favorability — "How favorable is your impression of [Brand]?"
- Purchase intent — "How likely are you to consider [Brand] for your next [category] purchase?"
Benchmark ranges vary significantly by category and exposure level, but for luxury and high-consideration brands running CTV, ad recall lift in the range of 4-12 percentage points and purchase intent lift in the range of 2-6 percentage points over a matched control group are common for campaigns with meaningful frequency (3+ exposures). Campaigns below this range often indicate insufficient frequency, weak creative, or audience targeting too broad to produce a measurable shift.
2. Geo-Holdout Incrementality Testing
Survey-based lift tells you about perception; geo-holdout testing tells you about behavior. The method: select a set of geographic markets (DMAs, zip code clusters, or states) and run your CTV campaign in some of them while deliberately holding it out of matched, comparable markets. After the campaign period, compare actual business outcomes — bookings, leads, sales, web traffic — between exposed and holdout markets.
The key to a valid geo-holdout test is matching: holdout markets need to be statistically comparable to exposed markets on the metrics that matter (historical sales trends, seasonality, demographic composition) before the test begins, not just "similar-looking" markets chosen casually. Without rigorous matching, you're measuring market differences, not media impact.
For luxury brands, geo-holdout testing is particularly valuable because it captures the long sales cycle that survey-based methods often miss — a holdout test run over 8-12 weeks can capture conversions that happen well after initial exposure, which is exactly the window in which premium CTV campaigns tend to influence high-consideration purchases.
3. Marketing Mix Modeling (MMM)
MMM takes a step back from individual campaigns and analyzes the relationship between aggregate media spend across channels (including CTV) and business outcomes over time, using statistical regression to isolate each channel's contribution. Unlike survey-based lift or geo-holdout testing, MMM doesn't require a live experiment — it works retrospectively on historical spend and outcome data.
MMM is most useful as a quarterly or annual validation layer: it won't tell you which specific CTV creative or platform drove results, but it will tell you whether CTV as a channel is pulling its weight relative to its share of budget — and it's increasingly important as cookie deprecation makes individual-level tracking less reliable across the board.
Comparing the Three Methods
| Method | What It Measures | Timeframe | Best For | Limitation |
|---|---|---|---|---|
| Survey-based brand lift | Awareness, favorability, intent shift | Real-time, during campaign | Creative and messaging validation | Doesn't measure actual conversions |
| Geo-holdout incrementality | True incremental sales/leads | 8-12+ weeks | Proving CTV drives business outcomes | Requires sufficient scale and market-level data |
| Marketing mix modeling | Channel contribution to outcomes over time | Quarterly/annual | Budget allocation across channels | Doesn't isolate individual campaigns or creative |
The mistake we see most often is brands picking exactly one of these and treating it as the complete answer. Survey-based lift alone can show favorable perception shifts that never translate to revenue. Geo-holdout testing alone, run once, can't tell you whether results are consistent across seasons or creative refreshes. MMM alone is too slow and too aggregate to inform in-flight optimization. The methodologies are complementary, not redundant.
Designing a Geo-Holdout Test for CTV: Step by Step
- Define the outcome metric. For luxury brands, this is often leads, qualified inquiries, or bookings rather than immediate sales — choose a metric with enough volume in each market to detect a meaningful effect.
- Select and match markets. Use 12-24 months of historical data to identify pairs or groups of markets with similar trends, seasonality, and baseline volume. Statistical matching (not just "these cities seem similar") is essential.
- Allocate markets to test and holdout groups. A common split is 70-80% of markets exposed, 20-30% held out — enough holdout volume to detect lift without sacrificing too much overall reach.
- Run the campaign for a sufficient duration. For high-consideration categories, 8-12 weeks minimum is typical to allow the sales cycle to play out.
- Measure the delta. Compare the percentage change in your outcome metric between exposed and holdout markets over the test period versus the pre-period baseline.
- Calculate incremental ROI. Translate the lift percentage into incremental units (bookings, leads, revenue) and compare against media spend to calculate true incremental ROAS — not blended ROAS, which includes conversions that would have happened anyway.
What Counts as a "Good" Result?
Benchmark ranges depend heavily on category, baseline brand awareness, and campaign scale, but a few general patterns hold across luxury and high-consideration verticals:
- Established brands with high existing awareness tend to see smaller percentage lifts in awareness metrics (since there's less room to move) but can still see meaningful purchase intent and consideration lift.
- Challenger or newer brands in a category typically see larger relative awareness lifts but need sustained frequency — often 4-6+ exposures over a campaign — before intent metrics move meaningfully.
- Incremental lift from geo-holdout tests in the 5-15% range relative to holdout markets is a reasonable target for well-targeted CTV campaigns in high-consideration categories; lift below 3-5% often suggests either insufficient reach/frequency or an audience that overlaps too heavily with people who would have converted regardless.
Common Mistakes in CTV Brand Lift Measurement
- Mistake 1: Running a lift study without enough scale. Survey-based lift studies need a minimum sample size in both exposed and control groups to detect statistically significant differences — campaigns with limited reach often produce inconclusive results that get misread as "no lift" rather than "insufficient sample."
- Mistake 2: Treating view-through conversions as proof of lift. View-through metrics are directionally informative but easily inflated, especially for broad-reach CTV campaigns. They should never be the sole evidence presented to justify CTV budget.
- Mistake 3: Testing during an atypical period. Running a geo-holdout test during a major sale, a competitor's campaign, or a seasonal anomaly contaminates the result. Test periods should reflect "normal" conditions as closely as possible, or account explicitly for known anomalies in the analysis.
- Mistake 4: Ignoring frequency in the analysis. A campaign that delivered an average of 1.5 exposures per household and a campaign that delivered 5 exposures will produce very different lift results — aggregating across frequency levels can mask whether the campaign actually had enough weight to move perception.
- Mistake 5: Comparing CTV lift to search/social conversion rates directly. These channels operate on fundamentally different mechanisms. CTV's value shows up in consideration and downstream conversion lift, not in a directly comparable "conversion rate."
Benchmark Ranges by Vertical
Brand lift benchmarks vary meaningfully by category, and luxury verticals tend to follow different patterns than mass-market consumer goods. While every brand's results depend on baseline awareness, creative quality, and frequency, these directional ranges reflect what we typically see across high-consideration categories:
- Private aviation and luxury travel: Lower baseline awareness for challenger brands often produces larger ad recall lifts (8-15 points), but purchase intent lift takes longer to materialize given multi-week booking decision cycles — geo-holdout windows of 10-12 weeks are typical.
- Luxury real estate: Awareness lift tends to be highly localized and geo-holdout testing is especially effective here, since real estate decisions are inherently tied to specific markets — incremental lift in inquiry volume of 10-20% in exposed markets is a reasonable target for a well-targeted campaign.
- Wealth management and financial services: Brand favorability and trust metrics often move more than raw awareness, reflecting the relationship-driven nature of the category; survey-based lift studies should weight favorability and "would consider for financial advice" intent questions over simple recall.
- Premium automotive: Purchase intent lift is highly frequency-dependent, with meaningful movement typically requiring 4-6+ exposures over a campaign — lower-frequency campaigns often show recall lift without corresponding intent lift.
- Luxury hospitality: Seasonal timing strongly affects both survey-based and geo-holdout results; tests should either run within a single season or explicitly control for seasonality in the analysis.
These ranges are starting points for setting expectations and designing tests, not guarantees — the only way to know what's realistic for a specific brand is to run a baseline study and use it to calibrate future campaigns.
Building a Measurement Framework That Holds Up
For our clients, we typically recommend a layered approach: survey-based brand lift studies running continuously (or for each major creative refresh) to validate messaging and frequency, a geo-holdout incrementality test run at least twice per year to validate actual business impact, and an annual MMM refresh to confirm CTV's overall contribution relative to other channels in the mix. No single study needs to carry the entire burden of proof — together, they build a body of evidence that survives scrutiny from finance.
The Bottom Line
Brand lift measurement for CTV isn't about finding a replacement metric for last-click conversions — it's about measuring what CTV actually does: shift awareness, build favorability, and create the consideration that shows up as conversions weeks later through other channels. Brands that insist on judging CTV by direct-response standards will consistently conclude it "doesn't work," not because it isn't working, but because they're measuring the wrong thing. The brands that build a layered measurement framework — survey-based lift, geo-holdout incrementality, and periodic MMM — are the ones that can defend CTV budget with evidence instead of intuition.
If your CTV measurement currently consists of a view-through pixel and a hopeful narrative, it's worth a conversation. Stillwater Media accepts a limited number of new client engagements per quarter.
