TL;DR
Incrementality asks whether the sale would have happened without the ad, which is a different question from whose attribution number is right. Meta runs formal Conversion Lift studies through representatives rather than as a self-serve product [1], so the practical design for a Shopify store is a geo split you design and read yourself. Both are governed by sample size before anything else, and the arithmetic is unforgiving: below a few hundred purchases in the measurement window, no split will return an answer you can act on.
Key Takeaways
- Meta's lift studies run through a representative rather than as a self-serve product [1]. Plan around that; there is no button to find.
- A lift study creates "a randomized test group of Accounts Center accounts that see your ads and control group who don't see your ads", then reports the difference between them [1]. Your own geo holdout copies that logic with worse randomization and better transparency.
- Sample size decides feasibility. By the arithmetic below, detecting a 20% relative lift takes about 475 purchases per arm and a 10% lift takes about 1,730.
- Freeze your conversion definition before the window opens. A lift number is a difference between two counts, so a value convention or a deduplication rule that changes mid-flight produces lift that is really instrumentation.
- Read total store revenue and order count for the test region, not platform-reported conversions. The platform number moves for reasons that have nothing to do with incrementality.
- Incrementality survives consent loss and click-ID loss better than attribution does, because both groups are measured the same imperfect way.
How many purchases does a lift test need?
Start here, because the answer disqualifies most tests before they are designed.
Treat purchases in each arm as counts, assume an even split between test and holdout, two-sided 95% confidence, 80% power, and one measurement window. Under those assumptions, detecting a relative lift L needs roughly n >= 2 * (2.8 / ln(1+L))^2 purchases per arm. This is our arithmetic from standard sample-size conventions, not a figure Meta publishes.
| Lift you want to detect | Purchases per arm | Purchases in the window, both arms |
|---|---|---|
| +10% | ~1,730 | ~3,460 |
| +20% | ~475 | ~950 |
| +30% | ~230 | ~460 |
| +50% | ~95 | ~190 |
Read the top row and then look at your own monthly Meta-influenced order count. A store doing 400 purchases a month across all channels cannot resolve a 10% lift in a four-week window, and cannot resolve a 30% one either. What it can resolve is a difference so large that the P&L already told you about it.
That is the uncomfortable part of incrementality work, and it is why the vendor sales pitch skips it. The stores with enough volume to measure lift precisely are the ones already spending enough to have an agency doing it. Everyone else is choosing between a test that returns noise and a cheaper decision rule.
Design the geo holdout so it can be read
A geo split is the design available to any store, and its quality comes from four decisions made before launch.
Pick regions that already behave alike. Pull 8 to 12 weeks of history and compare weekly order counts and revenue for candidate regions. Two US states whose weekly order counts track within a few percent of each other are a usable pair. Two whose curves diverge seasonally are not, and no statistical adjustment afterwards will rescue them.
Change exactly one thing. Meta ads off in the control region, everything else identical: email sends, promotions, other channels, site changes. A promotion that lands mid-window in one region and not the other ends the test.
Fix the window before you start, and do not look early. Divide the per-arm number from the table by your regional weekly purchase rate to get the number of weeks required. If that arithmetic says nine weeks, the test is nine weeks. Stopping when the gap looks good is how these tests produce a number that does not replicate.
Read Shopify, not Ads Manager. Your outcome variable is total orders and total revenue for each region, taken from your own order records. Platform-reported conversions are the wrong instrument here, because they change with signal quality, modelling and attribution windows, none of which are the thing you are testing. Whose number is right is a separate argument, covered in how to calculate Shopify ROAS and why platforms disagree and ROAS accuracy across ad platforms.
For on-site variant tests, which are a different tool for a different question, Shopify rollouts A/B testing covers the tracking side.
What has to be true about your conversion counter first
A lift result is a subtraction between two counts. Anything that makes one count drift relative to the other becomes lift, and you will never see it in the output.
Three properties matter, and each is a one-time verification.
One purchase per order, counted once. If a purchase can arrive twice for some orders and once for others, the imbalance lands in your result. In WeltPixel Conversion Tracking the Meta purchase is delivered from the Shopify order webhook with the browser and server event sharing an identifier, which lets Meta collapse the pair into one conversion, and the webhook path is idempotent across retries. The mechanics of that shared identifier are in how event id prevents double counting and the two-delivery-path reasoning is in CAPI versus browser pixel.
A value definition that does not move. Every channel in the app sends the order's total price with tax and shipping included. Whatever convention you use, freeze it for the window. Switching from total to subtotal halfway through a nine-week test produces a revenue "lift" of exactly the size of your average tax and shipping.
Matching quality that is stable, not necessarily high. A test tolerates mediocre match rates as long as they are the same in both arms. It does not tolerate improving match quality mid-window, which is why the week you follow the event match quality guide is the wrong week to start measuring lift.
The consent angle cuts in your favour here, which is unusual. Browser events respect Shopify's Customer Privacy API signals, so some share of sessions goes unmeasured in the browser layer no matter what you install. For attribution that is a hole. For a holdout test it mostly cancels, because the same measurement gap applies to both groups and the comparison is between them. Incrementality is the one measurement approach that gets more reliable as attribution gets harder.
Can you ask Meta to run a Conversion Lift study?
You can ask. Whether you get one depends on your account's access, and Meta's documentation is direct about the limit.
The developer guide for lift studies carries a warning: "Conversion Lift Measurement is currently limited. Please contact your Meta Representative for information about obtaining access" [1]. So a Conversion Lift study is not a self-serve feature you enable, and if nobody at Meta is assigned to your account, the practical answer is the geo design above.
The mechanics are still worth understanding, because they define what a good holdout looks like. Meta creates "a randomized test group of Accounts Center accounts that see your ads and control group who don't see your ads", then "calculates the difference between the test groups and control groups so that you evaluate the impact of your Facebook ads towards business goals" [1]. The holdout size is set as a control_percentage, described as "a holdout percentage of the Accounts Center accounts who will not see ads", with treatment and control percentages required to sum to 100 in a single-cell study [1]. Results include test, control and incremental conversions, confidence, cost per incremental conversion and buyer metrics on older studies, with breakdowns by cell, age, gender and country [1].
One number in that documentation gets quoted out of context. Meta states that breakdowns need "at least 100 conversions from test and control groups combined for results to display" [1]. That governs whether a slice renders, and it is not a claim that 100 conversions make a test conclusive. Compare it against the table above and the gap is obvious.
Meta's Business Help Center is where any self-serve eligibility thresholds would live, and it did not render for us. So this article states no minimum spend, minimum audience size, holdout percentage or study duration for a Meta-run study. The numbers that circulate on vendor blogs for those thresholds are uncited, and repeating them here would be guessing with a citation style.
For campaign types where Meta's own automation controls delivery, Advantage+ shopping campaigns and CAPI covers what the platform needs from your event feed.
How to read the result without fooling yourself
Run the arithmetic on an illustrative store: four-week window, one matched region holding Meta ads off, $4,000 of Meta spend in the test region during the window.
The test region records 520 purchases and the control region 470. That is 50 extra purchases, a 10.6% relative difference, and a cost per incremental purchase of $80 if you take the numbers at face value.
Do not take them at face value. From the table, resolving a 10% lift needs about 1,730 purchases per arm, and this test has 470 and 520. The observed gap sits comfortably inside the range that ordinary week-to-week variation produces, so the correct conclusion is that this test failed to measure anything, and the $80 figure is a number the spreadsheet computed rather than a fact about the business.
What that store learns is still useful. It now knows its window was too short by roughly a factor of three to four, and it can either commit to a longer hold or change instruments. Reporting the $80 to a board would be the actual error.
When your volume is too low for any of this
Below roughly 150 to 200 purchases in the entire measurement window, both arms combined, no split test will return a usable answer. Two cheaper decision rules exist.
The first is a full-channel pause read against total store revenue. Turn Meta off entirely for two to four weeks, watch total orders and revenue in your Shopify reports, then turn it back on and watch again. It is crude, it confounds seasonality, and it answers a blunt question that many stores genuinely have: does the channel move the top line at all. Run it twice in opposite directions before believing it.
The second is the same geo split held for a full quarter, accepting that a quarter of drift in promotions and seasonality is now inside your measurement. Neither is a lift study. Both beat a four-week test that reports noise with a confidence interval attached.
FAQ
Can I run a Meta Conversion Lift study on a self-serve ad account?
Meta's documentation says access to Conversion Lift Measurement is currently limited and directs you to a Meta representative [1]. Without that access, a geo split you design yourself is the route this article covers.
How long should a geo holdout run on a Shopify store?
Long enough to accumulate the per-arm purchase count for the lift you want to detect, which is about 475 per arm for a 20% lift by the arithmetic in this article. Divide that by your regional weekly purchase rate and fix the window before launch.
Should I measure lift using Meta's reported conversions or my Shopify orders?
Shopify orders and revenue, by region. Platform-reported conversions move with signal quality and attribution modelling, which would put your instrument inside the thing you are measuring.
Does poor event match quality invalidate a holdout test?
Not by itself, as long as it stays the same in both arms and across the window. What invalidates the test is changing your tracking setup, your value definition or your match rate mid-flight.
A holdout is only as good as the counter underneath it. WeltPixel Conversion Tracking [2] sends the Meta purchase from your Shopify order record with a shared deduplication identifier and one fixed value convention across all nine integrations, so the conversion count stays stable while your test window runs.
One thing to write down before you launch: the per-arm purchase count you need, the date the window closes, and your commitment not to read the result before then. The third line is the one that takes discipline.
Sources
- Meta for Developers, "Conversion Lift Studies" (warning that Conversion Lift Measurement is currently limited and to contact a Meta Representative for access; randomized test and control groups of Accounts Center accounts;
control_percentageholdout definition with cell percentages summing to 100; reported metrics including incremental conversions, confidence and cost per incremental conversion; breakdowns require at least 100 conversions from test and control groups combined to display), developers.facebook.com/docs/marketing-api/guides/lift-studies/, accessed September 10, 2026 - WeltPixel Conversion Tracking, Shopify App Store listing, apps.shopify.com/weltpixel-conversion-tracking, accessed September 10, 2026