Journal
Holdouts after the release
A two-week test is a poor match for a navigation change that will live in the binary for a year.
Feature flags make it easy to call something “shipped” when 100% of new sessions see the new tab bar. Novelty effects fade. Power users adapt. A metric that looked healthy in week two can sag in week eight. If you have no holdout, you have no counterfactual, only a time series and a story.
Holdouts are politically expensive. Support will hear from users who still have the old navigation. Marketing will want the screenshot of the new UI. A/B Test Analytics for Apps has to include the memo that says: we are keeping 5% on the old information architecture until week six because purchase rate is a slow metric. That memo belongs in the ship gate, next to crash and latency.
Operationally, holdouts fail when assignment is not sticky across app updates, or when a marketing campaign deep-links people into a flow that bypasses the flag. Treat those as leakage, the same way you would in a classic A/B. The Assignment Integrity Workshop exists because we kept seeing holdouts quietly dissolve.
You do not need a holdout on every colour tweak. You need one when the change is expensive to reverse or when the primary metric needs a longer window than your patience. Write the criterion down so the next PM does not have to rediscover it in a post-mortem.
If leadership insists on 100% immediately, record that as a decision with a named owner, not as “the data said ship”. Data cannot say that if you refused the comparison group.