A/B Testing Fundamentals
A/B Testing Fundamentals empowers teams to design, execute, and analyze experiments that drive informed product decisions. By applying the rigorous principles of the scientific method, this skill addresses common pitfalls in testing—such as early result peeking and hypothesis changes—ensuring reliable, actionable insights. It streamlines workflows for conversion optimization, enabling teams to efficiently validate hypotheses on user engagement and conversion rates through robust statistical analysis. With detailed guidance on sample size requirements and test design, users can confidently make data-driven choices that enhance user experiences and ultimately boost product performance. This structured approach guarantees high-quality output that brings clarity and precision to the decision-making process.
Spec
A/B Testing Mastery
Core Principles
The Scientific Method in Product
- Hypothesis: "Removing the optional fields will increase signups by 15%"
- Design: 50% see old form, 50% see new form
- Measure: Track conversion rate, statistical confidence
- Conclude: Is the difference real, or random variation?
Test Design
Sample Size Calculator
You need enough traffic to detect your desired effect at 95% confidence:
- 10% lift, 10K conversions baseline → 140K users needed
- 5% lift, 10K conversions baseline → 560K users needed
- 20% lift, 10K conversions baseline → 35K users needed
Duration
- Run at least 1 full week (captures day-of-week variation)
- Never run less than 100 conversions per variant
- Stop early only if you hit statistical significance AND practical significance
Common Mistakes
❌ Peeking at Results Early: Introduces selection bias, inflates false positive rate ❌ Changing Hypothesis Mid-Test: P-hacking—always define hypothesis before running ❌ Running Too Many Tests: Multiple comparisons inflate false positive rate (Bonferroni correction) ❌ Ignoring Segments: A winning variant might hurt a key segment ❌ Not Considering Externalities: Traffic source changes, seasonal effects, marketing campaigns
Statistical Significance
- p-value < 0.05: Less than 5% chance results are due to random variation
- Confidence Interval: "Lift is between 8-18% with 95% confidence"
- Sample Ratio Mismatch: Detect if variant assignment is skewed
Real Example: Checkout Flow
- Control: 3-step checkout with mandatory fields
- Variant A: 2-step checkout, optional fields
- Variant B: 1-step checkout, inline validation
- Traffic: 50K users over 2 weeks
- Result: Variant B increases conversion by 22%, p<0.01
- Decision: Roll out Variant B, measure long-term retention impact
Tools
- VWO, Optimizely, LaunchDarkly for infrastructure
- Built-in analytics (Google Analytics experiments)
- Statistical calculators (Evan Miller's A/B test calculator)

