Tests statistical literacy (power, peeking, SRM) plus decision-making under uncertainty - the stats are a means, the decision is the job.
Interview prompt
Your A/B test shows a +2% conversion lift, but it's not statistically significant. The team wants to ship. What do you do?
What interviewers evaluate
Do you ask whether the test was powered to detect the effect size at all?
Do you check validity first (sample-ratio mismatch, peeking, novelty effects)?
Do you treat post-hoc subgroups as hypotheses, not results?
Do you frame shipping as a decision with asymmetric costs, not a stats ritual?
Do you protect the integrity of the experimentation culture (no re-running until significant)?
A framework to structure your answer
Separate the questions - 'what does the data say?' and 'what should we do?' are different questions.
Power check - was the test capable of detecting a 2% lift? Compute the required sample; 'not significant' from an underpowered test means 'didn't look hard enough.'
Validity checks - sample-ratio mismatch, repeated peeking, novelty/decay effects over the window, guardrail metrics.
Subgroups - post-hoc segments where it 'worked' are hypotheses to retest, never results.
Decide on asymmetric costs - cost of shipping a truly-neutral change vs cost of skipping a truly-positive one, weighted by prior plausibility and maintenance cost.
Path forward - extend to power, or ship with a monitored holdback; never re-run until significant.
Strong sample answer
Try structuring your own answer first, then reveal a strong worked example.
Common variants
The test is significant but the effect is tiny - ship it?
Two metrics moved in opposite directions in the same test. Decide.
Your experiment platform shows a sample-ratio mismatch. What does it mean and what do you do?
Pitfalls to avoid
Treating 'not significant' as 'no effect' without asking whether the test was powered.
Skipping validity checks (SRM, peeking) and debating a possibly-broken result.
Shipping on a post-hoc subgroup where the change 'worked.'
Making it a pure stats argument with no decision framework (costs, prior, reversibility).
Re-running the test until it reaches significance.
Likely follow-ups
Traffic is too low to ever reach power for a 2% effect. How do you make decisions at this company?
The +2% is significant at 90% confidence but not 95%. Does that change anything?
How would you design the holdback so it actually settles the question?