Metric breakdowns can show whether an experiment’s observed effect differs across event-property values, but a positive subgroup is not automatically the source of the lift. Use breakdowns to investigate heterogeneity and data quality, then validate any rollout decision with a planned comparison or a follow-up test.

What a metric breakdown changes

An aggregate result combines everyone eligible for a metric. A breakdown displays that same metric across values such as device type, country, plan or acquisition source. This can reveal measurement gaps and plausible differences that the average hides. It also creates more comparisons, so some attractive-looking segments will appear by chance.

In the experiment results view described by the original PostHog release, open the metric and choose Breakdowns, then select an event property. The exact interface and limits can change; confirm the current product behavior before documenting a fixed number of properties as a permanent rule.

A synthetic example

Suppose an experiment has the following invented results. These numbers illustrate interpretation only; they are not customer data or a historical 99Ways experiment.

DeviceControl conversions / usersTreatment conversions / usersObserved rates
Desktop120 / 1,000135 / 1,00012.0% vs 13.5%
Mobile180 / 2,000174 / 2,0009.0% vs 8.7%
All users300 / 3,000309 / 3,00010.0% vs 10.3%

The table suggests a desktop increase and a small mobile decrease, while the aggregate is nearly flat. It does not by itself prove the treatment works on desktop. Check uncertainty for the interaction, whether device was recorded before treatment, whether missing values differ by arm, and whether this segment was chosen before seeing the result.

A safe interpretation workflow

First confirm that eligibility, exposure and conversion use the same definitions across arms. Reconcile each segment’s assigned users and outcomes to the aggregate, including an explicit missing or unknown bucket. A breakdown that drops unclassified events can tell a misleading story.

Next distinguish a prespecified comparison from an exploratory finding. A segment named in the analysis plan can inform the original decision if its statistical method was also planned. A segment discovered after trying many properties is a hypothesis for validation, not a new winner.

Finally ask whether the property could be changed by the treatment. Device type observed before exposure is usually easier to interpret than a property recorded after the tested experience. Post-treatment properties can select people based on behavior caused by the variation.

What to check before a segmented rollout

Do not ship to one subgroup merely because its displayed rate is positive. Review sample size, uncertainty, multiple comparisons, practical effect size, guardrail metrics and implementation feasibility. Then run a follow-up test targeted to the proposed population when the decision is material.

Use the A/B testing metrics guide to define the outcome before analysis. PostHog’s experiment metrics documentation is the appropriate source for current metric behavior; the original changelog entry records the feature’s release context.

The practical takeaway

Breakdowns are most useful for finding questions: Is instrumentation missing on one platform? Is a prespecified audience responding differently? Does an apparent effect survive a direct comparison? They add context to an experiment result; they do not replace the aggregate analysis or authorize selective winner-picking.

Author

  • Moein Heshmati

    Moein Heshmati is Co-Founder of 99Ways, specializing in analytics engineering, tracking, and experimentation systems. He writes about data quality, measurement, technical implementation, and reliable experimentation.

Related blog posts

Leave a Reply

Your email address will not be published. Required fields are marked *