← All playbooks Playbook 02 · From My Mentoring Sessions

Product Execution

The round that tests whether you can reason from a metric to a decision - precisely, exhaustively, and with commitment.

~12 min read · Updated June 2026

MetricsDiagnosisTrade-offsExperimentation

What this round is

Product Execution is the round where the interviewer hands you a metric, a goal, or a sudden change in the data and asks you to reason your way to a decision. The format is open-ended, but the prompts cluster into a few recognisable shapes: "Define the goals and metrics for Instagram Reels," "Daily active users dropped 10% overnight - diagnose it," "Should we show more ads in WhatsApp or protect engagement?", "How would you test this?"

There is rarely a single correct answer. The interviewer wants to see how you turn a vague metric question into a precise one, how you reason from user behaviour to numbers and back, how you isolate a cause without flailing, and whether you can converge on a decision you'd actually defend in a planning meeting.

The defining difference from Product Sense: this round is numbers-first. You won't be scored on the elegance of a feature. You'll be scored on whether your metric is the right metric, whether your diagnosis is exhaustive before it is narrow, and whether your trade-off call holds up when the interviewer pushes back - and they will push back. Google interviewers are known to respond to your first hypothesis with a flat "No. Go deeper."

What they're testing

Not your math. Your judgment with numbers.

Anyone can recite "DAU, retention, engagement." The candidates who clear this round are the ones who define a metric precisely before reasoning about it (does "engagement" mean a like or a comment? those lead to completely different analyses), who exhaust the problem space before narrowing, who name the counter-metric that the headline metric could quietly destroy, and who commit to a decision instead of analysing forever.

The five dimensions being scored:

  1. Goal & metric definition - Can you anchor to what the product is for and pick the metric that actually reflects it?
  2. Metric selection & trade-offs - Can you choose a North Star, support it, and name the guardrail that stops you from gaming it?
  3. Systematic root-cause analysis - When a number moves, can you methodically cover bugs, releases, segments, seasonality, and external shocks before guessing?
  4. Data-driven decision-making - Do you ground every claim in a metric, an assumption, or a logical deduction - or do you hand-wave?
  5. Execution reasoning & commitment - Can you converge on a call, say what you'd not do, and define how you'd validate it?

The four things that win the round

  1. Precision - You define the metric before you diagnose or optimise it. Sloppy definitions ("more engagement") are the fastest way to look junior.
  2. Exhaustiveness, then focus - You map the whole space (internal/external, every funnel step, every segment) and then prioritise where to dig. Jumping straight to one cause reads as luck, not method.
  3. Trade-off honesty - Every metric you move has a counter-metric you risk. Naming it unprompted is the senior signal.
  4. Commitment - You land a recommendation. "It depends" without a decision is a fail; "it depends on X - and given the likely value of X, here's my call" is a pass.

The mental model - 6 questions

Most prep teaches Product Execution as a framework you recite. Senior PMs don't recite - they think. And what they're thinking is a small set of questions, asked in whatever order the prompt demands.

  1. What exactly am I being asked - and what type of question is this? Restate the prompt. Decide whether it's a metrics/goal question, a diagnosis (metric-change) question, a trade-off/decision question, or an experiment question. The type determines the playbook.
  2. What's the goal, and how does this product create value? Anchor to the product's job and the company's mission. State one primary objective clearly; mention secondary ones briefly. Don't spend 25% of the interview here - show you can prioritise.
  3. What's the user journey, and where in the funnel does this live? Sketch the funnel: acquisition → activation → engagement → retention → monetisation → referral. Everything downstream - your metric, your diagnosis, your bet - hangs off this map.
  4. What's the right metric - and what's the counter-metric? Pick a North Star. Define its inputs. Name supporting metrics and, critically, the guardrail/counter-metric the North Star could cannibalise (e.g. notifications up, but session time down).
  5. Is the change real, and where is it concentrated? - or - What are the options, and what's the right bet? For a metric drop: confirm it's not a logging artefact, then segment hard (geo, platform, version, cohort, time) to localise it. For a trade-off: lay out the genuinely different options, size each, and pick.
  6. How would I validate it, and what's my recommendation - and what could break? Name the experiment or the data you'd pull to confirm. Commit to a recommendation. Name the risk and the second-order effect.

The order doesn't matter. The rigour does.

The question shapes

Product Execution prompts fall into four (sometimes five) recognisable shapes. Classifying the shape in the first 30 seconds tells you where to focus.

Shape What it sounds like
Metrics / Goal-setting "Define the goals and metrics for X." "How would you measure success of Y?"
Metric change / Diagnosis "DAU dropped 10% - what do you do?" "Comments up, watch-time down - why?"
Trade-off / Prioritisation "Should we add more ads or protect engagement?" "What would you prioritise for Live?"
Experiment / A-B design "How would you test this feature?" "Is this result trustworthy?"
Estimation / Sizing "How many X happen per day?" "Size the market for Z."

Sample questions

Metrics & Goal-Setting. The trap: listing every metric. The win: one North Star, a tight driver tree, and a guardrail.

  • If you're the PM for Instagram Reels, define the goals and metrics.
  • Define success metrics for Facebook Marketplace.
  • What's the North Star metric for WhatsApp? Defend it.
  • How would you measure the success of Facebook's "Save Item" feature?
  • How would you measure the success of Amazon Prime Video's recommendation algorithm?
  • You're the PM for Horizon Worlds - what metrics do you focus on first?
  • How would you measure retention for the Facebook Feed?

Metric-Change / Diagnosis. The trap: guessing a cause immediately. The win: confirm it's real, segment exhaustively, then narrow with hypotheses.

  • A Meta product shows a 10% drop in newly registered users - what data would you need to understand and fix it?
  • Daily active usage of Facebook Dating dropped 25% overnight. How would you diagnose it?
  • YouTube comments are up, but watch time is down. What do you do?
  • Notification engagement is rising but time-on-site is falling. What's going on?
  • Sign-up conversion fell 8% week-over-week in one region only. Walk me through it.
  • App store rating dropped from 4.6 to 4.1 in a month. What happened?

Trade-off & Prioritisation. The trap: refusing to decide. The win: size the options, pick, and name what would change your mind.

  • A messaging app wants to add more ads - but ads may reduce engagement. What do you do?
  • Should Uber Eats be a separate app from regular Uber?
  • You can ship feature A (lifts revenue, risks retention) or feature B (lifts retention, no revenue impact). Which, and why?
  • Engagement is up but creator satisfaction is down. Which do you optimise for?
  • Personalisation increases time-spent but reduces content diversity. Where do you land?

Experimentation & A/B Testing. The trap: vague "I'd A/B test it." The win: hypothesis, randomisation, guardrails, duration, pre-committed decision rule.

  • How would you A/B test a new onboarding flow?
  • You ran a test and the treatment won by 2% - would you ship it? What would you check first?
  • How would you test whether a higher ad load hurts retention?
  • How do you account for novelty effects and network effects in a social-feature test?

Estimation / Sizing. Often inside another question. The trap: hand-waving. The win: explicit assumptions, math out loud, sense-check against reality.

  • How many photos are uploaded to Instagram per day?
  • Estimate the daily revenue of Facebook Marketplace in one country.
  • How much storage does WhatsApp need for a year of messages?

How to structure your answer

The structure of your answer is half the signal. A candidate who guesses a root cause in the first 30 seconds loses even if the guess is right. A candidate who runs the playbook transparently wins even if the final number is debatable.

  • Talk while you think. The interviewer needs to follow your reasoning. Silent thinking reads as no structure.
  • Use the whiteboard. Sketch the funnel, the metric tree, the segmentation cuts. Visual structure beats verbal every time in this round.
  • Number your sections. "First I'll anchor to the goal, then map the funnel, then build a metric tree, then land on a North Star and guardrails." Then stick to it.
  • Define before you reason. Always nail down what a metric means before optimising or diagnosing it.
  • Be exhaustive, then focused. In diagnosis, map the whole space out loud, then prioritise where to dig. In metrics, list candidates, then pick one.
  • Time it loosely (for a ~35-min core): ~5 min clarify + anchor, ~10 min funnel + metric tree (or exhaustive diagnosis space), ~10 min the decision/experiment, ~5 min summary and follow-ups.
  • Commit, then hedge. End with a recommendation and then say what would change it. Not the other way round.
  • Recover gracefully. If a clarifying answer changes your direction, say so and pivot visibly. Course-correcting in real time is a senior signal.

Common traps

  • Guessing the root cause immediately. The single fastest way to lose a diagnosis question. Confirm it's real, then segment, then hypothesise.
  • Listing twenty metrics with no North Star. A metric dump signals you can't prioritise. One North Star, a tight driver tree, one guardrail.
  • Sloppy metric definitions. "More engagement" / "better experience" - undefined metrics make every downstream step mush.
  • Forgetting the counter-metric. Every metric you push has one it can quietly destroy. Naming it unprompted is the senior tell.
  • Refusing to decide. "It depends" with no landing is a fail. State the dependency, estimate it, and commit.
  • "I'd just A/B test it." Without randomisation unit, guardrails, duration, and a decision rule, this is hand-waving.
  • Skipping the data-integrity check. Many "drops" are logging bugs or metric-definition changes. Rule that out first.
  • Ignoring scale and segments. At FAANG scale, aggregate numbers hide everything. If you never segment, you never localise.
  • Reciting the framework. Run the steps silently; narrate the thinking, not the acronym.
  • Trailing off at the end. The summary is where they decide. Land it like a recommendation to a VP.

Strong vs weak answers

Prompt: "Daily active users dropped 10% week-over-week. What do you do?"

A weak answer sounds like:

"A 10% DAU drop is concerning. It's probably a bug in the latest release, or maybe users are losing interest. I'd tell engineering to check the code and maybe send a re-engagement push notification to bring people back. We could also run a survey to ask users why they left."

What's wrong: guesses a cause instantly, no check that the drop is even real, no segmentation, no structure, jumps to solutions before diagnosis.

A strong answer sounds like:

"Before diagnosing, I want to confirm the drop is real and not a logging or metric-definition change - I've seen 'drops' that were pipeline bugs. So step one: cross-check DAU against an independent source like server-side session logs.

Assuming it's real, I'd scope it: is it a sudden cliff on a specific day, or a gradual slope? A cliff points to a release or outage; a slope points to behaviour or competition.

Then I'd segment to localise it - by platform and app version, geography, acquisition cohort, new vs. existing users, and time-of-day. The shape of where the drop concentrates usually names the cause. If it's isolated to one Android version, my lead hypothesis is the v8.2 release; I'd line up crash rates and the release timestamp against the drop curve. If it's concentrated in one country, I'd check for an outage, a local competitor launch, or a holiday. If it's spread evenly across everyone and started gradually, I'd suspect a recommendation or ranking change reducing content quality.

I'd rank those hypotheses, pull the specific data to confirm the top one, and act: roll back if it's a release, file with infra if it's an outage, escalate to the ranking team if it's quality. And I'd set a guardrail to catch this faster next time - an automated DAU anomaly alert segmented by version and geo."

What's right: confirms the metric is real, scopes the shape, segments exhaustively, forms ranked hypotheses tied to where the drop lives, names the confirming data, commits to an action per case, and closes with a systemic fix.

How to prepare

Preparation for Product Execution is mostly pattern recognition. The method is constant; what changes is the product and the metric. The more products and metric trees you've reasoned through, the faster you move under pressure.

  • Build 10-15 metric trees. For products you use daily, write the North Star, the driver metrics, supporting metrics, and the guardrail/counter-metric. This is the single highest-ROI rep for this round.
  • Run "the number moved" drills. Pick a metric, invent a drop, and force yourself through real-check, scope, segment, hypothesise - without skipping to a cause. Time-box to 8 minutes.
  • Pre-build a segmentation checklist. Geography, platform/OS, app version, device, cohort (new/existing/power), acquisition channel, time-of-day. Memorise it so you never freeze in a diagnosis.
  • Pre-build a counter-metric reflex. For every common North Star, know its enemy: engagement ↔ wellbeing/quality; revenue/ad-load ↔ retention; notifications ↔ session time; growth ↔ activation quality.
  • Practice experiment design out loud. Hypothesis → randomisation unit → primary metric + guardrails → MDE/power → duration → pre-committed decision rule. Five reps and it's automatic.
  • Drill estimation. Practice making assumptions explicit and sense-checking against reality. Numbers fluency removes the panic.
  • Practice with a partner who pushes back. Google-style "No, go deeper" pressure is the real test. Have someone reject your first hypothesis and make you produce a second and third.
  • Time yourself. Land the core answer in 25-30 minutes with a clean summary.
  • Study real teardowns (Lenny Rachitsky, Exponent, RocketBlocks, IGotAnOffer, Reforge) - steal the structure of how senior PMs reason with metrics, not the conclusions.

Define precisely, diagnose exhaustively, trade off honestly, commit clearly - and the round is yours.

Design
Color