Metric Design
North-star metrics, metric trees, guardrails, counter-metrics, proxies, and Goodhart's law.
11 questions
JuniorDesignVery commonYour team owns part of an e-commerce checkout and is told to 'grow revenue', but revenue is too broad to act on. Build a metric tree that decomposes revenue into inputs, drilling down until you reach levers your team can directly move. Show the decomposition and name the leaf metric your team would actually own.
Your team owns part of an e-commerce checkout and is told to 'grow revenue', but revenue is too broad to act on. Build a metric tree that decomposes revenue into inputs, drilling down until you reach levers your team can directly move. Show the decomposition and name the leaf metric your team would actually own.
Decompose multiplicatively: Revenue = visitors x conversion x average order value, and conversion into add-to-cart x checkout-start x payment-success. Traffic is not its lever, so the checkout team owns payment-success rate — a leaf it moves — not revenue.
Common mistakes
- ✗Decomposing revenue additively when the natural structure is multiplicative
- ✗Assigning the team a leaf it cannot move, like traffic, instead of payment-success
- ✗Stopping at revenue itself without drilling down to an actionable leaf lever
Follow-up questions
- →Where would a discount-rate lever attach in this multiplicative tree?
- →How do you keep sibling leaves from overlapping or double-counting?
JuniorTheoryVery commonWhat is a north-star metric, and why should a company have exactly one?
What is a north-star metric, and why should a company have exactly one?
The single metric that best captures the core value users get and that most closely leads long-term revenue. One shared number aligns every team on the same outcome and stops each team optimizing a local metric that pulls against the others.
Common mistakes
- ✗Equating the north-star with a revenue or vanity number instead of a value proxy
- ✗Thinking several north-stars, one per team, beat a single shared metric
- ✗Picking it for dashboard visibility rather than its predictive link to value
Follow-up questions
- →How do you tell a genuine value proxy from a vanity metric?
- →What breaks when two teams each pick their own north-star?
SeniorDesignVery commonYou are launching a brand-new 'saved carts' feature and must define its metric set before rollout: one primary metric, two secondary metrics, and two guardrails, plus the decision rule that says whether the launch is a success. Design the set for this feature and state the exact rule under which you would ship it to everyone, hold, or roll it back.
You are launching a brand-new 'saved carts' feature and must define its metric set before rollout: one primary metric, two secondary metrics, and two guardrails, plus the decision rule that says whether the launch is a success. Design the set for this feature and state the exact rule under which you would ship it to everyone, hold, or roll it back.
Primary: repeat-purchase rate of savers vs control. Secondary: save-to-checkout conversion and saves per active user. Guardrails: checkout latency and revenue per user, against cannibalization. Rule: ship only if the primary rises significantly with all guardrails in threshold; else roll back.
Common mistakes
- ✗Choosing adoption or raw saves as the primary instead of a downstream value metric
- ✗Shipping on any uptick without a significance check on the primary
- ✗Omitting guardrails, so cannibalization or a latency regression goes unnoticed
Follow-up questions
- →Why is a value metric, not adoption, the right primary for a new feature?
- →How long must the guardrails run before the decision rule can fire?
JuniorTheoryCommonWhat are guardrail metrics, and why should you never optimize them directly?
What are guardrail metrics, and why should you never optimize them directly?
Guardrails are metrics you watch so improving your primary metric does not harm something else — latency, errors, refunds, cost. You hold them within a threshold, not maximize them; they are constraints, not targets — optimizing one is a category error.
Common mistakes
- ✗Treating a guardrail as another target to maximize rather than a threshold to hold
- ✗Confusing guardrails with post-incident alerts you only react to after a break
- ✗Believing you can improve a guardrail without hurting the primary it protects
Follow-up questions
- →For an engagement primary, which two guardrails would you set and why?
- →How do you pick the threshold that counts as a guardrail breach?
JuniorDesignCommonYou are the first analyst at a food-delivery app (customers order from restaurants; couriers deliver). Leadership wants a single north-star metric to align product, growth, and operations. Two candidates are on the table — total app downloads, and gross merchandise value (GMV) this quarter. Propose the metric you would make the north-star, and explain why it beats both of those alternatives as a guiding number for every team.
You are the first analyst at a food-delivery app (customers order from restaurants; couriers deliver). Leadership wants a single north-star metric to align product, growth, and operations. Two candidates are on the table — total app downloads, and gross merchandise value (GMV) this quarter. Propose the metric you would make the north-star, and explain why it beats both of those alternatives as a guiding number for every team.
Pick a value-and-frequency metric like monthly delivered orders per active user. It captures realized value and repeat usage, and every team can move it. Downloads is vanity top-of-funnel that ignores retention; GMV lags and is inflated by discounts and one-off orders.
Common mistakes
- ✗Choosing downloads, a top-of-funnel vanity number that ignores whether users return
- ✗Choosing quarterly GMV, a lagging figure inflated by discounts and one-off orders
- ✗Refusing to pick one and blending three metrics into a vague combined target
Follow-up questions
- →How would you turn orders per active user into a metric tree of team levers?
- →What guardrail would you pair with it to protect unit economics?
MiddleDesignCommonA delivery team is measured on '% of orders delivered on time', and gaming appears — couriers quote padded ETAs and mark orders delivered before handoff so the clock stops early. You want to keep the on-time target but stop it from being hit dishonestly. What counter-metric do you add, and where in the scorecard does it sit relative to the primary metric?
A delivery team is measured on '% of orders delivered on time', and gaming appears — couriers quote padded ETAs and mark orders delivered before handoff so the clock stops early. You want to keep the on-time target but stop it from being hit dishonestly. What counter-metric do you add, and where in the scorecard does it sit relative to the primary metric?
Add a counter-metric that worsens exactly when on-time is gamed — the quoted-vs-actual ETA gap or the customer-reported late rate — plus a real delivery-confirmation check. Pair it as a guardrail beside the primary, so on-time counts only while the counter stays flat.
Common mistakes
- ✗Just raising the on-time bar, which changes the number, not the incentive to game
- ✗Picking a counter-metric that does not move when the primary is actually gamed
- ✗Hiding the counter on another dashboard instead of pairing it beside the primary
Follow-up questions
- →How do you set the counter-metric's threshold without punishing honest delays?
- →Would a delivery-confirmation signal from the customer close the handoff loophole?
MiddleDesignCommonA support org sets one target — 'tickets closed per hour' per agent — and ties bonuses to it. Applying the idea that a measure under pressure stops measuring what you wanted (Goodhart's law), predict the specific ways agents will hit the number without actually helping customers, and name the real outcome the target has stopped tracking.
A support org sets one target — 'tickets closed per hour' per agent — and ties bonuses to it. Applying the idea that a measure under pressure stops measuring what you wanted (Goodhart's law), predict the specific ways agents will hit the number without actually helping customers, and name the real outcome the target has stopped tracking.
Agents optimize the proxy, not the goal: cherry-pick easy tickets, split one issue into several, close prematurely or reopen, and rush customers off. Closed-per-hour climbs while the real outcome — problems solved and customers satisfied — stops improving.
Common mistakes
- ✗Assuming a clear numeric target cannot be gamed and only drives real effort
- ✗Predicting agents slow down, missing that gaming inflates the count instead
- ✗Naming symptoms but not the lost outcome — resolved problems and satisfaction
Follow-up questions
- →Which paired guardrail would catch the split-and-reopen gaming fastest?
- →Would measuring first-contact resolution instead remove the incentive?
MiddleDesignCommonLast quarter total revenue rose 15% but ARPU (average revenue per user) fell 8%, because a cheap acquisition channel brought in many low-spend users. Leadership asks which number belongs in next quarter's OKR — the absolute revenue total or the per-user ratio — and why. Give your recommendation and the reasoning, including what each choice would incentivize the team to do.
Last quarter total revenue rose 15% but ARPU (average revenue per user) fell 8%, because a cheap acquisition channel brought in many low-spend users. Leadership asks which number belongs in next quarter's OKR — the absolute revenue total or the per-user ratio — and why. Give your recommendation and the reasoning, including what each choice would incentivize the team to do.
Use both in different roles — absolute revenue as the growth target and ARPU as a guardrail. Absolute alone rewards buying unprofitable volume; the ratio alone can be 'improved' by shedding low-value users. Pairing them keeps growth healthy without eroding per-user value.
Common mistakes
- ✗Picking the ratio alone, which a team can 'improve' by shedding low-value users
- ✗Picking the absolute alone, which rewards buying unprofitable low-spend volume
- ✗Treating it as one-or-the-other instead of target-plus-guardrail together
Follow-up questions
- →What ARPU threshold would trip the guardrail and pause the cheap channel?
- →How does segmenting ARPU by channel change the diagnosis here?
SeniorTheoryOccasionalA good metric is sensitive, attributable, hard to game, and value-aligned — judge 'sessions per user' on all four.
A good metric is sensitive, attributable, hard to game, and value-aligned — judge 'sessions per user' on all four.
Mixed. Sensitive: yes, reacts fast to product changes. Attributable: fair, sliceable by cohort and surface. Hard to game: weak — inflatable by fragmenting sessions or nagging pushes. Value-aligned: weak, since more sessions can mean confusion, not value. A useful signal, poor target.
Common mistakes
- ✗Calling it game-proof, missing that fragmenting sessions or pushes inflate it
- ✗Assuming more sessions always means more value rather than possible confusion
- ✗Passing or failing all four together instead of scoring each criterion apart
Follow-up questions
- →Which single change would most improve its resistance to gaming?
- →What value-aligned metric would you pair it with to offset its weakness?
SeniorDesignOccasionalYour team's true goal metric is 7-day retention, but at your team's small weekly grain it is far too noisy to detect any change you ship. Propose a more sensitive proxy metric you could steer by instead, and describe how you would prove that the proxy is actually valid — that moving it really does move 7-day retention rather than drifting on its own.
Your team's true goal metric is 7-day retention, but at your team's small weekly grain it is far too noisy to detect any change you ship. Propose a more sensitive proxy metric you could steer by instead, and describe how you would prove that the proxy is actually valid — that moving it really does move 7-day retention rather than drifting on its own.
Pick an earlier, higher-frequency leading indicator upstream of retention — day-1 activation or first-week core-action count — with more events and less noise. Prove validity: show it historically predicts retention, and that past experiments moving the proxy also moved retention.
Common mistakes
- ✗Confusing a high-volume metric with a valid proxy without checking the link
- ✗Treating a present-day correlation as full proof the proxy is causally valid
- ✗Giving up and averaging the noisy target longer instead of finding a leading proxy
Follow-up questions
- →How would past A/B tests serve as the validation set for the proxy?
- →What would tell you the proxy has decoupled from retention over time?
SeniorDesignRareDesign a seller rating for a marketplace. The naive raw average is broken — a seller with three 5-star reviews shows 5.0 and outranks a seller with 500 reviews averaging 4.8. Design a rating that accounts for the volume and reliability of the evidence, so that a tiny sample cannot outrank a large well-reviewed seller, and explain the mechanism that fixes it.
Design a seller rating for a marketplace. The naive raw average is broken — a seller with three 5-star reviews shows 5.0 and outranks a seller with 500 reviews averaging 4.8. Design a rating that accounts for the volume and reliability of the evidence, so that a tiny sample cannot outrank a large well-reviewed seller, and explain the mechanism that fixes it.
Don't rank on the raw average; shrink it toward the global mean by review count — a Bayesian smoothed average or a lower confidence bound. Few reviews stay near the prior; many approach it. So three 5-star reviews sit below a 4.8 from 500, as thin evidence is discounted.
Common mistakes
- ✗Adding a minimum-review floor but still ranking on the raw mean above it
- ✗Ranking on review volume alone, ignoring the actual star scores
- ✗Bolting an additive volume bonus on instead of shrinking toward a prior
Follow-up questions
- →How does the prior's strength set how fast the score trusts new reviews?
- →Why does a lower confidence bound naturally penalize a tiny sample?