The complete date-selection package had the highest observed opt-in rate, but the experiment did not establish that radio buttons alone were the winning change. Replacing three date CTAs with a selector moved the reported rate from 49.12% to 49.47%; adding a waitlist and clearer instructions together reached 50.14%. The distinction matters when deciding what to build next.

Four experiences, three different questions

The page invited visitors to choose a date for a live video event. Control offered three separate CTA buttons. We wanted to examine the selection format, capture people unavailable for the listed dates, and make the next action clearer. Instead of three independent tests, the treatments formed a cumulative ladder.

ExperienceDate selectionWaitlistTop-bar copy
ControlThree CTA buttonsAbsentList of event dates
Group 1Radio-button selectorAbsentList of event dates
Group 2Radio-button selectorPresentList of event dates
Group 3Radio-button selectorPresentPick The Date You Want To Attend Below!

Control versus Group 1 addresses selection format. Group 1 versus Group 2 addresses the added waitlist within the selector experience. Group 2 versus Group 3 addresses the instruction within the selector-plus-waitlist experience. Group 3 versus control evaluates the whole package, not the top-bar wording alone.

The observed results

The historical article reports 34,946 users over six days. The rates below reproduce its reported result table. Per-arm counts and the full statistical configuration have not been recovered for this revision, so the table is an observed comparison rather than a new inferential analysis.

ExperienceReported opt-in rateDifference from control, percentage points
Control49.12%Baseline
Group 149.47%+0.35
Group 249.22%+0.10
Group 350.14%+1.02

The article’s historical model output reports +0.73% relative lift for Group 1 and +2.08% for Group 3 against control. Calculations from rounded rates can differ slightly from the model display. The original “win probability” column mixes values whose reference comparison is not fully documented; it should not be read as four mutually exclusive probabilities of being the best arm. This revision does not use it to declare a component winner.

Why adjacent comparisons change the interpretation

From the displayed rates, adding the waitlist changes Group 1 to Group 2 by −0.25 percentage points. Changing the top-bar instruction changes Group 2 to Group 3 by +0.92 points. These are descriptive differences. The probability or interval for Group 3 versus control cannot be reused as the uncertainty for Group 3 versus Group 2.

A cumulative design is not automatically invalid. With appropriate random assignment and planned contrasts it can estimate incremental effects in those component combinations. What it cannot show here is whether the instruction also helps the original three-button layout, or whether the waitlist behaves differently without the selector. Those interactions require additional comparisons.

A large total sample does not settle those questions by itself. To reconstruct the analysis, recover participants and outcomes per arm, assignment and exposure rules, conversion windows, the probability definition and the prespecified comparisons. Apply the same outcome definition to all arms; see choosing A/B testing metrics and PostHog’s statistical interpretation.

A waitlist selection is useful information, not proof of better leads

The original account reports about 1,200 waitlist selections, but does not establish whether that count covers one treatment or both treatments offering a waitlist. It also does not report attendance, later registration or purchases for those people. We therefore cannot conclude that the waitlist created warmer leads or improved satisfaction.

The choice may still serve users who cannot attend. Evaluate it separately: distinguish a confirmed event registration from a waitlist request, honor the stated follow-up, and measure how many waitlisted people later select a date and attend. Do not count both as equivalent successful registrations merely to increase the primary metric.

The next test and the implementation checks

The clearest next question is whether the instructional top bar helps with the original layout. Compare that instruction against the existing date list while retaining the same buttons, dates and event availability. This is a proposed follow-up, not an experiment already completed. It tests whether the promising copy difference transfers beyond the cumulative package.

Before launch, define valid registrations and unique exposure, verify date and timezone labels, and test the selection using keyboard and mobile interactions. A selector should retain explicit labels and a clear submit action; adding a waitlist must not silently book a date. Preserve the selected option across validation errors and confirm the downstream registration record matches it.

The CTA animation experiment is a separate motion question. Likewise, segment breakdowns investigate populations within results; they do not replace the between-arm comparisons needed here.

What is worth keeping from this result

The complete package is a promising direction for this page, and the experiment provides a useful map of unresolved questions. The practical lesson is to name the contrast behind each conclusion. A bundle’s improvement is not proof that every component helped, and a small observed difference is not proof of a reliable component effect.

Author

  • Amin Heshmati

    Amin Heshmati is an author at 99Ways focused on PostHog, conversion optimization, experimentation, attribution, and analytics implementation. He writes about building reliable measurement systems and using data to improve digital performance.

Related blog posts

Leave a Reply

Your email address will not be published. Required fields are marked *