Playbook / Measurement

How to Design a Startup A/B Test Around a Real Decision

An A/B test should compare a defined change under conditions that make the result interpretable. Simply showing two versions and comparing their totals does not establish a reliable experiment.

Begin with the decision the test will inform and the mechanism you expect to change behavior.

Define the hypothesis and unit

State what changes, for whom, and why it might improve the outcome. Decide the unit of assignment: a user, account, session, or another appropriate unit. Consider whether exposure in one group can affect the other.

For a fictional onboarding test, assigning individual users may be inappropriate if coworkers in the same account share the experience. The design needs to reflect how the product is used.

Choose the outcome before launch

Select one primary outcome tied to the decision, plus guardrails for important harms or tradeoffs. Define the observation window and the minimum effect that would matter commercially.

A higher button-click rate may not justify the change if activation or customer quality deteriorates. Follow the outcome far enough to answer the actual question.

Plan interpretation

Choose a statistical approach and stopping rule suitable for the design. Repeatedly checking results and stopping at the first favorable moment can change the error behavior of a test. Use a method that accounts for how the experiment will actually be monitored.

Keep the hypothesis and analysis plan available to the team before results arrive.

Check implementation

Verify assignment, exposure, event collection, and major differences between groups that should not exist. A tracking defect or an uneven rollout can invalidate a clean-looking result.

Record concurrent changes and unusual events. They may affect interpretation even when the experiment itself is functioning.

Make the commercial decision

Consider effect size, uncertainty, implementation cost, and guardrails together. A statistically detectable difference may be too small to matter. A promising but uncertain result may justify another bounded test rather than full adoption.

Preserve the learning

Document the design, result, limitations, and decision. Make the next action explicit. The experiment brief is the operational companion to this design process.

A startup does not need to call every change an experiment. When a formal comparison is useful, the extra design work should make the resulting decision more credible.

Co-founder and CEO of Stackmatix, startup advisor, and former Head of Sales at MightyHive. · More about Matt →