Journal
Peeking and the false lift
Written for people who already have a dashboard and a Friday habit.
The most common way an app team invents a win is not fraud. It is a well-meaning look. Someone opens the experimentation tool, sees a 6% lift on trial start, screenshots it into Slack, and the room treats the screenshot as a decision. Two days later the interval has walked back toward zero. Nobody updates the thread.
Optional looks spend statistical budget whether you admit it or not. If you planned a two-week window and you also checked on day three, day five, and day nine, you no longer have the error rate you wrote in the pre-analysis plan. A/B Test Analytics for Apps has to include this as an operational problem, not a footnote for specialists.
A stopping rule that survives a sceptical engineer is boring: fixed horizon, or a sequential method you named before the first look, with the software actually implementing it. “We’ll stop if it looks good” is not a method. Group sequential boundaries exist; so do always-valid intervals. Pick one and write the name in the plan. If your vendor cannot do it, you still do not get to peek for free — you just have a worse tool.
There is a cultural piece. Product reviews in GB companies often reward speed. The person who waited until day fourteen looks slower than the person who shipped on day four with a pretty chart. Readout Craft spends a full module on this because the math is short and the social pressure is not.
If you take nothing else: the screenshot is evidence of a look, not evidence of an effect. Put the planned end date in the channel topic. When someone pastes an early chart, ask which boundary they used. The silence that follows is diagnostic.