Define the decision and the north-star outcome
I start by asking which recurring decision the metric should support. Examples include:- whether to expand a pilot to another team;
- whether a release is safe to deploy;
- whether to improve retrieval, workflow design, or user onboarding;
- whether a model-routing policy saves money without reducing quality;
- whether the product should continue receiving investment.
Protect the outcome with guardrail metrics
Optimizing one outcome creates pressure elsewhere. I surround the north star with four metric families:| Family | Questions |
|---|---|
| Quality | Is the answer, prediction, or action correct enough for the workflow? |
| Reliability | Is the service available, stable, fast, and recoverable? |
| Risk | Are authorization, privacy, safety, and escalation controls working? |
| Economics | What does each successful outcome cost, including human review and rework? |
Measure adoption as behavior, not exposure
Login counts and prompt volume show exposure, not sustained value. Adoption becomes informative when it describes the intended workflow:- eligible users who complete the target task;
- repeat usage after the novelty period;
- successful task rate by cohort;
- abandonment and reformulation;
- fallback or human escalation;
- correction and rework after the AI output;
- time from request to accepted outcome.
Connect technical metrics to product action
Every recurring metric should have an owner, review cadence, and expected response. A scorecard that cannot change a priority is reporting, not management. A compact monthly review can ask:- Did the north-star outcome improve for the intended population?
- Did any quality, reliability, risk, or cost guardrail deteriorate?
- Which workflow segment explains the movement?
- What evidence supports the diagnosis?
- Which product or engineering action follows?