BootcampInterview prep

Interview Drills — Metric Definition & Analytical Judgment

6 drills with frameworks and rubrics.

Interview Drills — Metric Definition & Analytical Judgment

Open-ended interview questions. Each has a Framework (the structure a strong answer follows), a Model answer (a concise example), and a Rubric (what an interviewer listens for). These drill the judgment that separates a real analyst from someone who just runs queries: defining a metric so two people get the same number, choosing the right summary statistic, and refusing to confuse correlation with causation. Practice thinking aloud and always pin down a definition, a number, and the caveat. The app can role-play these as mock interviews (see mock-interview.md).

The universal definition structure: State the decision the metric serves → name the exact event/population that counts → fix the time window → list what's excluded (refunds, bots, internal accounts) → say how you'd sanity-check it. Use it on any "how would you define X" question.

D1

  • difficulty: easy
  • concept: metric-definition How would you define "active user" for a mobile app, and why does the definition matter?
  • Framework: Ask what decision the metric drives → pick a qualifying action that reflects real value (not just an app-open) → fix the window (DAU/WAU/MAU) → state exclusions (bots, internal/test accounts) → note you'd write the definition down so everyone reports the same number.
  • Model answer: "It depends on the decision. For an app whose value is messaging, I'd define a daily active user as a unique account that sends or reads a message on a given calendar day, excluding internal and test accounts. I'd pick daily because engagement is meant to be daily; for a tax app I'd use monthly instead. The definition matters because 'opened the app' and 'actually used it' give very different numbers — and if it isn't written down, two dashboards will disagree."
  • Rubric: Strong answers tie the definition to the product's value and the decision, name a concrete qualifying action (not a bare open), fix a window with a reason, and list exclusions. Weak answers say "anyone who logs in" with no window, no exclusions, and no link to what the business cares about.

D2

  • difficulty: medium
  • concept: metric-definition Define "churn" for a subscription business. What edge cases would you resolve before reporting a churn rate?
  • Framework: State churn = customers (or revenue) lost over a period ÷ base at the start → pick the unit (logo churn vs. revenue churn) → fix the period → resolve the messy cases: voluntary vs. involuntary (failed payment), downgrades, pauses, win-backs, free trials → say which you include and why → sanity-check against net revenue.
  • Model answer: "Logo churn = subscribers who cancel in a month ÷ subscribers at the month's start. Before reporting it I'd resolve: do failed-payment cancellations count as churn or get a grace period? Do pauses count, or only true cancels? Do downgrades count as partial churn — if so I'd report revenue churn separately so a small account leaving and a whale downgrading aren't treated the same. I'd exclude free trials that never converted, since they were never customers. Then I'd cross-check: revenue churn should reconcile with the MRR movement."
  • Rubric: Strong answers distinguish logo vs. revenue churn, name the base and period explicitly, and surface real edge cases (involuntary churn, pauses, downgrades, trials) with a defensible choice. Weak answers give "people who leave" with no base, no period, and no awareness that downgrades and failed payments need a ruling.

D3

  • difficulty: medium
  • concept: mean-vs-median A stakeholder asks for the "average order value." When would reporting the mean mislead them, and what would you report instead?
  • Framework: Recall mean is distorted by extreme values, median resists them → diagnose the distribution (is it skewed? a few huge orders?) → pick the summary that represents the typical customer → add a measure of spread → if both are useful, report both and explain the gap.
  • Model answer: "If order sizes are skewed — most orders are $30 but a handful of bulk orders are $5,000 — the mean gets dragged up to, say, $80, which describes nobody. I'd report the median (the typical order, maybe $35) alongside the mean, and flag the gap as evidence of a long tail. I'd add the range or a couple of percentiles so they see the spread. I'd keep the mean only where it's the right tool — e.g., total-revenue planning, where the big orders genuinely count."
  • Rubric: Strong answers connect skew/outliers to the mean-vs-median choice, report a measure of spread too, and know the mean still has legitimate uses (totals). Weak answers reflexively say "always use the median" or report a lone average with no sense of distribution.

D4

  • difficulty: medium
  • concept: confounder Our data shows users who use Feature X have 2x the retention of users who don't. Should we push everyone to use Feature X? Walk me through your reasoning.
  • Framework: Resist the causal leap → name the likely confounder (self-selection: already-engaged users adopt features) → state correlation ≠ causation → propose how to actually test it (A/B test, or at least matched cohorts / a holdout) → say what you'd recommend doing before acting on the correlation.
  • Model answer: "I'd push back on the 2x before we act on it. The most likely explanation is self-selection: highly engaged users adopt Feature X and retain well — engagement is the confounder driving both, so the feature may not cause retention. The correlation is a reason to investigate, not a reason to ship a push. To establish cause I'd run an experiment — randomly prompt one group to try Feature X and compare retention — or, if we can't, compare matched cohorts with similar prior engagement. Only if the experiment shows lift would I recommend rolling it out."
  • Rubric: Strong answers immediately flag the correlation-causation trap, name the specific confounder (self-selection / prior engagement), and propose an experiment or matched comparison to test cause. Weak answers take the 2x at face value and recommend the rollout, or vaguely "look at more data" without naming the confounder or a test.

D5

  • difficulty: hard
  • concept: spotting-confounders A report claims that customers who contact support churn more, concluding "support is driving people away." How would you investigate before accepting that?
  • Framework: Separate the correlation from the causal claim → generate the confounder: customers contact support because they already hit a problem, so the problem (not support) may drive churn → consider reverse causation and timing (did churn intent precede contact?) → propose a check: segment by issue type/resolution, compare resolved vs. unresolved, look at timing → state what evidence would actually support each story.
  • Model answer: "The claim confuses a symptom with a cause. People usually contact support after something breaks, so the underlying problem is the likely confounder — support contact is a marker of trouble, not its source. I'd check timing (did dissatisfaction precede the contact?), then segment: do customers whose issue was resolved quickly churn less than those left unresolved? If good support resolution lowers churn, the 'support drives people away' story collapses and the real driver is the unresolved problem. I'd only blame the support experience if poor resolution, not contact itself, tracks with churn."
  • Rubric: Strong answers identify reverse causation / a confounding problem, distinguish 'contacted support' from 'had a bad support experience,' and design a segmentation (resolution, timing) that can tell the stories apart. Weak answers accept the headline causal claim, or reject it without proposing how to actually adjudicate it.

D6

  • difficulty: hard
  • concept: metric-definition You're asked to build the single "north-star" metric for a two-sided marketplace. How do you define it, and what failure modes would you guard against?
  • Framework: Anchor on the value exchanged (a completed transaction), not a vanity proxy → write the exact definition (event, both sides, window, exclusions) → check it can't be gamed and isn't a vanity metric → name what it hides (so you'd pair it with guardrail/segment metrics) → state how you'd validate it reflects real value.
  • Model answer: "I'd anchor on delivered value: 'completed transactions per week,' where completed means the buyer received the good/service and it wasn't refunded or cancelled — counting unique matched buyer-seller pairs, excluding internal test accounts. I'd avoid vanity proxies like 'listings created' or 'sign-ups,' which rise without value changing. Failure modes: it can hide a collapsing supply side, so I'd pair it with seller-side and buyer-side health as guardrails; and it can be gamed by tiny self-deals, so I'd add minimum-value and fraud filters. I'd validate it by checking it moves with revenue and repeat usage."
  • Rubric: Strong answers pick a value-bearing, hard-to-game event, write a precise definition (both sides, window, exclusions), explicitly reject vanity metrics, and pair the north star with guardrails for what it hides. Weak answers choose a vanity count (sign-ups, GMV with no completion check), give no exclusions, and ignore gaming or one-sided blind spots.