The situation
A moderation bug changed how a subset of listings was displayed. The business needed to know whether it had cost engagement and revenue. The headline metrics showed nothing, which is the point at which most investigations stop and conclude there was no harm.
What I did
- Reconstructed the true onset of the bug, which was a full day earlier than it had been reported.
- Isolated the affected cohort and built a control group of comparable unaffected listings.
- Measured harm per item rather than in aggregate, which is where the effect was actually visible.
- Expressed the result against its Poisson standard error so the finding carried an explicit confidence statement rather than an impression.
- Explained why the aggregate had not moved — demand had shifted to substitute listings — and named the metric the business should monitor for this class of failure in future.
What it means for you
Aggregate metrics hide substitution effects. When a change harms part of a marketplace and demand simply moves elsewhere, the total looks fine and the harm is still real — measuring it needs a cohort, a control and a stated confidence level.
Context
These are from eight years owning the data warehouse and reporting function of a consumer marketplace platform, in-house rather than as an outside consultant. The employer and the internal system names are withheld; the numbers are the real ones.