Design your experiment
Experi helps startup PMs design statistically rigorous experiments before they launch. Enter your metrics on the left and get sample size, runtime, CUPED variance reduction, a risk score out of 100, and a full pre-launch checklist.
⬡
Loading example experiment
Pre-loaded: SaaS onboarding at 5% CVR, 8K users/week. Edit the sidebar to use your own numbers.
—
/ 100
—
experiment quality
Score breakdown
Per variant
—
users needed
Runtime
—
weeks raw
—
With CUPED
—
weeks
—
MDE
—
abs. pp lift
Industry benchmarks
| Your metric | Your value | Industry median | Assessment |
|---|
Statistical power curve 80% threshold marked
Without CUPED
—
users per variant
—
With CUPED
—
users per variant
—
How CUPED works
CUPED strips pre-experiment noise by regressing your outcome on pre-period behavior:
Because X is independent of treatment assignment, subtracting θX doesn't change the expected treatment effect - it only removes variance. The result: same signal, less noise, tighter confidence intervals.
Variance reduction = rho^2 - the square of the correlation between your pre and in-experiment metric. At rho=0.5, that's 25% reduction. At rho=0.7, it's 49%.
Y_adj = Y − θ(X − X̄) where θ = Cov(Y,X) / Var(X)
Because X is independent of treatment assignment, subtracting θX doesn't change the expected treatment effect - it only removes variance. The result: same signal, less noise, tighter confidence intervals.
Variance reduction = rho^2 - the square of the correlation between your pre and in-experiment metric. At rho=0.5, that's 25% reduction. At rho=0.7, it's 49%.
Python implementation
—
Bayesian A/B analysis Beta-Binomial model
—
probability variant B beats control
Variant B winsControl wins
—
expected uplift
—
expected loss if wrong
95%
ship threshold
Frequentist vs Bayesian - when to use which
| Criterion | Frequentist | Bayesian |
|---|---|---|
| Question answered | P(data | null true) | P(B wins | this data) |
| Early stopping | Inflates false positive | Natural with loss functions |
| Communicating to PMs | Harder - p-value | Easier - "73% likely to win" |
| Prior knowledge | Not used | Can incorporate |
| Industry standard | More common | Netflix, Booking.com |
Sequential testing Prevents early-stopping bias
Peeking at experiment results before the planned end date and stopping early inflates your false positive rate - often to 20-30% when you think it's 5%. Sequential testing sets an always-valid confidence boundary that you can check at any time without penalty.
—
Nominal alpha
—
Sequential boundary z*
—
Safe peeks allowed
mSPRT implementation
# mixture Sequential Probability Ratio Test import numpy as np def msprt_boundary(alpha, n_max, n_current): # O'Brien-Fleming spending function info_frac = n_current / n_max z_boundary = ppf(1 - alpha/2) / np.sqrt(info_frac) return z_boundary # Check at each peek z_obs = (p_b - p_a) / se z_bound = msprt_boundary(0.05, n_max=n_target, n_current=n_so_far) if abs(z_obs) > z_bound: print("Safe to stop - significance reached")
Multi-metric experiment plan Primary + guardrails
A complete experiment tracks one primary metric and 2-3 guardrail metrics. The primary drives the ship/kill decision. Guardrails are automatic kill switches — if any regresses, the experiment stops regardless of the primary result.
Metric hierarchy framework
| Level | Type | Examples | Action on regression |
|---|---|---|---|
| L1 | North star | Revenue, DAU, retention | — |
| L2 | Primary | CVR, activation rate | Ship/kill decision |
| L3 | Guardrail | p99 latency, error rate | Auto-kill |
| L4 | Diagnostic | Click depth, scroll | Inform only |
Week-by-week roadmap
Pre-launch checklist
Traffic growth forecast 0% monthly growth
Configure traffic growth in the sidebar to see how your experiment runway changes over time.
Experiment spec - Notion / Slack ready
—
Python analysis script
—
Live experiment analyzer Paste real data
Your experiment is running. Paste in the current numbers — Experi will tell you where you stand, whether there's an SRM, your current Bayesian probability, and whether it's safe to stop.
Control (A)
Variant (B)
CVR Control
-
CVR Variant
-
Observed lift
-
p-value
-
Bayesian prob B wins
-
Data collected
-
SRM check
-
Enter visitors and conversions above to analyze your live experiment.
Segment explorer Heterogeneous treatment effects
Different user segments often respond differently to the same treatment. A feature that lifts desktop CVR by 20% might hurt mobile. Define 3 segments below — Experi calculates independent sample sizes, runtimes, and flags which segments are likely underpowered.
Segment analysis
| Segment | Traffic/wk | Baseline CVR | Required n | Runtime | Power status |
|---|
Why segment analysis matters
Running one experiment on your full population and calling it done misses heterogeneous treatment effects (HTE). A variant that's neutral overall might be a strong winner for power users and a loser for new users — averaged together, you'd ship something that hurts new user retention.
The right workflow: run your primary experiment on the full population, then use segments as a diagnostic layer. If segments diverge significantly, consider separate feature flags per cohort rather than a single global ship.
The right workflow: run your primary experiment on the full population, then use segments as a diagnostic layer. If segments diverge significantly, consider separate feature flags per cohort rather than a single global ship.
Stakeholder digest generator Week-N update
Auto-generates a weekly experiment update you can paste directly into Slack or email. Fill in where you are — Experi writes the update for you.
Your digest Slack