Evidence before
the experiment.
Predict A/B test outcomes from your own user data — direction, effect size, and calibrated confidence. In minutes, not weeks.
Predictions locked before you reveal the answers.
Learning is rationed by traffic.
Three numbers set the ceiling on how much any team can learn in a quarter.
experiments that succeed at Google, Bing, and Booking.com1
of live traffic for one answer
experiments per quarter, ceiling, at 10% allocation
The failed ones aren’t free. For weeks, real users live inside the failed variant — slower checkouts, worse rankings, prices that feel wrong. That’s not just spent traffic; it’s spent trust.
An average test needs four weeks and a 10% slice of traffic. That’s a hard ceiling of ~30 experiments a quarter — before holdouts, exclusion groups, and re-runs bring it closer to 20. Your backlog is ten times that. Ideas were never the bottleneck. Traffic is.
1 Kohavi & Thomke, Harvard Business Review, 2017; Thomke, HBR, 2020.
How it works
The events you already collect become a behavior model. The model produces simulated traffic. The traffic produces a report. Scroll to watch it run.
Connect your data.
Read-only access to the events you already collect. No SDK, no instrumentation project.
We learn your users.
Behavior models fine-tuned on your event stream — revealed preferences, not personas. Segments emerge from behavior, not from a dropdown. Calibrated to your product in weeks, not quarters.
Simulate the test.
Your experiment runs against simulated traffic. Out comes a full A/B-style report: direction, effect size, confidence interval, segment cuts, guardrail movements — with a plain-language summary grounded in the simulated outcomes.
Many trajectories, one readout.
The simulated runs collapse into a distribution: direction, effect size and interval, segment cuts, guardrails.
Plain language: the single-page variant is predicted to win, driven almost entirely by new users; returning users are near flat.
Our models never read your hypothesis. They can’t be talked into a winner.
Graded on your data,
not our benchmark.
Anyone can quote an accuracy percentage from their own benchmark. We’d rather be graded blind. You pick experiments you’ve already run — and keep the answer key. We lock timestamped predictions before you reveal a single outcome. Then we report directional accuracy and calibration on your metrics. If we’re wrong, you’ll know exactly how wrong. So will we.
- 01You pick past tests
- 02We predict and lock
- 03You reveal outcomes
- 04Accuracy and calibration, scored
by the customer
A calibration read is the whole scorecard: predicted effect on one axis, revealed effect on the other. Points on the diagonal mean the prediction landed. Distance from it is the error, in your metric, on your tests.
No measured accuracy is claimed on this page. The diagram shows the approach, not a result.
From copy tweaks to algorithm swaps.
Simpirical isn’t limited to what fits in a screenshot. Our models learn from your event stream, not from rendered pages — so they can simulate whatever your logs can express. If your users’ behavior records it, you can test it here first.
Built by engineers who’ve lived the queue.
Years running experimentation-heavy product teams at Meta, Uber, Intuit, and PayPal — where the backlog was always ten times the traffic.
We’re early — and looking for partners to build this with.
Simpirical is early. The product is built on synthetic data created from published research — and we’re looking for partners to grow with, teams willing to point it at their own event history and hold us to the answer.
Partners help set the roadmap and get early access as the product hardens, a direct line to the team building it, and pricing locked from day one.