Free tool

Time-to-outcome checker

Ranking form variants on what closed instead of what completed is a better idea than it is a practical one, and whether it works for you is an arithmetic question rather than a philosophical one. Volume, close rate and cycle length decide it. For a lot of funnels the honest answer is no, and this will say so.

Runs in your browser. Nothing you type is sent anywhere, stored, or logged — there is no request to send it in. Built by the team behind the dishonest dashboard.

Your volume

Leads that go into the CRM and never get a disposition are not gradeable, however real they were.

Your outcomes

20% means going from a 3% close rate to 3.6%. Smaller differences cost dramatically more traffic to see.

Time for one outcome-weighted test to conclude

5.6 years

5.5 years accumulating enough submissions, plus 4 weeks waiting for the last of them to resolve into a won or lost.
Gradeable submissions a month
420

Submissions that eventually get a disposition. The rest are invisible to a test.

Per variant, per month
210

Split evenly across the variants in the test.

Submissions needed per variant
13,914

To detect a lift this size on the close rate at 95% confidence and 80% power.

The difference being tested for
3.00% → 3.60%

Baseline close rate against the improved one you asked to be able to detect.

Finding

No — outcome-weighted split testing will not work at this volume

This is the honest answer and it is not the answer that sells software. At your volume and cycle length the test cannot conclude before it is meaningless. The outcome data is still worth collecting — knowing which leads turned into money is useful on its own — but it should grade your traffic sources, not adjudicate a form variant.

How this is calculated

Every figure above comes from the arithmetic below. No weighting, no model, no numbers of ours mixed into yours.

How fast outcomes arrive

gradeable per month = submissions × gradeable share per variant per month = gradeable per month ÷ variants

Submissions that never get a disposition cannot be part of a test, however real they were. A CRM full of leads nobody ever marked won or lost has a gradeable share near zero.

How many you need

pA = close rate pB = close rate × (1 + lift) n = ( z₀.₉₇₅ × √(2p̄(1−p̄)) + z₀.₈ × √(pA(1−pA) + pB(1−pB)) )² ÷ (pB − pA)²

The standard two-proportion sample size at 95% confidence and 80% power. Note how it scales: halving the lift you want to detect roughly quadruples the sample.

How long that takes

months accumulating = n ÷ per variant per month months of lag = median days to disposition ÷ 30.44 total = months accumulating + months of lag

The lag is added once rather than per submission. Submissions resolve while later ones are still arriving, so only the tail of the sample is still open when collection ends — adding one median cycle at the end is the reasonable approximation, and it is an approximation.

Where the defaults come from

The defaults describe a small B2B funnel: 600 submissions a month, a 3% close rate, a month to disposition, and 70% of leads eventually getting marked one way or the other. Nothing about it is a benchmark — it exists so the page does something before you touch it.

The 20% default lift is the smallest improvement most people would bother shipping a form change for. Try lowering it to 10% and watch what happens to the sample size; that behaviour is the most useful thing on this page and it is a property of the statistics, not of our opinion.

The verdict bands — under six weeks, under four months, under a year — are our judgment about what a marketing team can actually wait for, not a finding. Someone with a different testing cadence would draw them elsewhere, and the number above them is the part that matters.

What this cannot tell you

It assumes your traffic, offer and sales team hold still for the duration. Over a test measured in quarters they will not, and every change is a confound. This is the main reason a long answer here should be read as “no” rather than as “yes, but be patient.”

It also treats disposition lag as a single median. Real time-to-outcome distributions have long tails, and a median of 30 days often hides deals that take six months. If your distribution is skewed, the answer here is optimistic.

And it says nothing about whether outcome data is worth collecting — only whether it can adjudicate a split test. Those are different questions with different answers. Knowing which of your traffic sources produce leads that close is useful at any volume, including volumes far too low to test on. If the answer here is a flat no, that is the version to build.

Once you do have two variants with outcomes attached, the outcome-weighted split test calculator is where you find out whether the difference holds up.

Why we built this

Every number on this page is one your form builder could have told you and didn’t.

Endpoint Forms is an open-source form builder for marketers: forms built to convert, data that goes wherever you need it, and every submission carrying what it turned out to be worth. It is not shipped yet. The waitlist is where we tell you when it is.

Waitlist