In-house · consumer marketplace platform

The metric had not collapsed. Its attribution had.

A key metric appeared to fall off a cliff and stakeholders feared a real drop in user activity. The events were still arriving — they had lost their attribution upstream.

~75% → ~100%Completeness restored
Backfilled unsampledHistory
Structural fix shippedRecurrenceMoved off the sampled source rather than patching it

The situation

A metric the business watched closely appeared to collapse. The reasonable interpretation was that user activity had genuinely dropped, and the organisation was preparing to respond to that. Confirming or refuting it needed to happen quickly and be provable either way.

What I did

  1. Proved the underlying events were still present in the source at the expected volume, so the drop was not a change in user behaviour.
  2. Traced the cause upstream: a removed retry window combined with processing lag meant events were arriving too late to be attributed to an individual item.
  3. Pinpointed the day the divergence began, rather than accepting the date the problem was first noticed.
  4. Backfilled the broken partitions unsampled, so history was corrected rather than approximated.
  5. Opened structural follow-up work to move off the sampled source entirely, so the same failure could not recur.

What it means for you

A metric falling is not the same as the thing it measures falling. Establishing which one you are looking at, before the business reacts, is often worth more than the fix itself.

Context

These are from eight years owning the data warehouse and reporting function of a consumer marketplace platform, in-house rather than as an outside consultant. The employer and the internal system names are withheld; the numbers are the real ones.

Tell me what is broken.

A short call is usually enough to tell whether this is a two-week fix or a two-month one — and whether I am the right person for it.