This four-arm subscription test produced an early signup leader, but the later revenue winner is not yet proven by equal-age cohort evidence. The $17 plan with a seven-day trial led the first ten-day experiment display; the surviving account says $27 without a trial later generated the most total income. Treat that later statement as an unverified follow-up observation until original assignments and billing records are reconciled.

The four offers

ArmMonthly priceTrial
Control$27Seven days
Group 1$27None
Group 2$17Seven days
Group 3$17None

The initial experiment began October 25, 2024, ran for ten days and included 10,222 users. Group 2—the lower price with a trial—had a reported 79.18% probability of being the top early conversion option. That result concerns the initial conversion metric, not mature net revenue.

Why the later comparison is difficult

A coworker reportedly ended the experiment early and rolled out Group 2. That decision may have changed later allocation and cohort sizes. The article does not say whether the two- and three-month revenue totals include only originally randomized users, whether each arm had equal follow-up age, or whether refunds and failed payments were deducted.

“Total income per group” is not a fair comparison when groups differ in size or observation time. The decision metric should be cumulative net revenue per originally assigned visitor at the same cohort age, supported by first payments, renewals, cancellations, refunds and retained subscribers.

Separate price, trial and interaction questions

The four-arm design can estimate price, trial availability and whether their effects interact when allocation and outcomes are retained. Picking the single highest arm does not by itself explain whether trial removal helps equally at both prices or whether the price effect changes with a trial.

Mechanism claims are not results

The account says trial users were less committed or satisfied and compares the outcome with a painting-choice study. This experiment did not measure commitment or satisfaction, and a separate intervention cannot establish the cause here. Onboarding, acquisition mix, billing friction, offer expectations and chance are other possibilities.

The statement that PostHog “cannot track” revenue is also too broad. In this implementation, billing data was reviewed in ThriveCart and had not been connected to the experiment analysis. That is an integration limitation of the historical setup, not a universal product limit.

What a defensible revenue readout needs

Create 30-, 60- and 90-day cohorts from original assignment. For each arm report assigned visitors, paid conversions, renewals, refunds, cancellations and cumulative net revenue per assigned visitor. Explain censoring and exclude post-rollout users from the randomized comparison unless they remained assigned under the original protocol.

Decision from Phase 1

The early signup result is documented; the mature-revenue conclusion remains provisional. Do not present $27 without a trial as the “true winner” until the assignment and billing reconciliation is complete. The durable lesson today is narrower: subscription pricing decisions need a prespecified maturity horizon and revenue denominator, and an early signup leader can be the wrong rollout choice.

Use the currency/revenue tutorial for normalized implementation concepts, the event tracking plan for the billing handoff, and the experimentation-system audit for stopping and assignment checks.

Author

  • Amin Heshmati

    Amin Heshmati is an author at 99Ways focused on PostHog, conversion optimization, experimentation, attribution, and analytics implementation. He writes about building reliable measurement systems and using data to improve digital performance.

Related blog posts

Leave a Reply

Your email address will not be published. Required fields are marked *