Define successful work before measuring it

A running process does not prove a customer placed an order. HTTP 200 can contain an incomplete result or stale data. Begin with important journeys: confirming an order, searching, uploading or completing a report.

Choose an SLI for quality and an SLO target. An example service might require 99.9% correct order confirmations over 30 days. That is the example's promise, not a standard for every service. Define which attempts enter the denominator and which outcomes count as bad.

Alongside the user-facing measure, we track request rate, errors, latency and resource saturation. They show what is happening to the workload and where headroom is running out. High CPU is useful context, but paging depends on product impact and how quickly the situation can worsen.

For checkout we distinguish request acceptance from completed outcome. Returning 202 while a task stays in a queue is not a successful order. An SLI can measure operations completed correctly within the required time rather than only count responses without server errors. The definition follows the promise the customer receives.

Measurement location changes coverage. An internal service never sees traffic rejected at the edge. Client instrumentation sees network failures but may fail to send telemetry. We choose a source that fits the objective and record what part of the path it misses, rather than assume one instrument describes the whole experience.

Exclusions need meaning. Invalid customer input may not count as a service failure; our faulty validation does. Excluding every 4xx response can hide broken authorization or lost access. The classification should match responsibility and product behavior, and remain stable enough to compare periods honestly.

Critical workflows need their own view. An availability figure dominated by catalog browsing can hide broken payments. We select boundaries by product importance rather than create an SLO for every endpoint. Key promises must remain visible without turning the monitoring system into an unmanageable catalogue of targets.

We agree on the objective with its cost and business need. Another nine requires resources, time and operational change. A working-hours product may have different conditions from continuous payment processing. The SLO is an agreement about quality and response, not an impressive percentage detached from how the service is used.

Check what the measurement can hide

Separate successful operations from failures. Fast rejections can improve average latency during an outage. Percentiles describe a distribution, but p99 on a small sample is unstable. Interpret latency alongside request volume and important routes.

Averaging the p99 values from several instances does not give us the service p99. We need to combine observations appropriately, for example through compatible histograms. Bucket boundaries should match the objective: coarse buckets around a one-second limit make that limit difficult to assess.

Queues need arrival rate, completion rate and oldest-job age. A hundred jobs can mean seconds or hours. Scheduled work also needs the last successful completion time; no new error does not prove a job ran.

A counter accumulates events, so rate calculations need correct handling of time windows and resets. A process restart should not appear as negative workload. When a graph changes after deployment, we inspect the measurement semantics before concluding that user behavior or system performance changed.

Histogram buckets approximate a distribution. Boundaries matching the decision make the result more useful, but unbounded precision increases series and storage. A one-second threshold needs a different layout from a broad diagnostic range. Accuracy follows the question the team is trying to answer.

Latency is measured across the appropriate boundary. SQL takes thirty milliseconds, while an order waits five minutes in a queue. The handler looks fast and the user result remains slow. For background paths we measure waiting, processing and acceptance-to-completion separately rather than mistake one for the other.

Missing data is not zero. A collector failure can remove a line without recording a new error value. We need absence detection and an expected interval. A daily job requires different conditions from continuous HTTP traffic. Otherwise the quietest dashboard can belong to the least observable incident.

We check window and sample size before drawing conclusions. Short windows react quickly and fluctuate; long windows smooth away beginnings. Release comparisons need comparable traffic. One p99 without observation count, route and period is insufficient evidence that the application became faster for customers.

Collect explanations at a controlled cost

Add pool waits, locks, dependency duration, memory and disk use. These help trace slow checkout to a particular part of the path. Each metric needs an owner, unit and understood observation point.

Label cardinality counts combinations. Twenty routes, five statuses and ten instances allow 1000 combinations for one metric. A user identifier makes growth depend on audience size; histogram buckets multiply the resulting series further.

For route metrics, we keep a template such as /orders/:id. Individual IDs and request context belong in logs or traces when needed. We collect context that helps investigation, rather than copying bodies and secrets because they are easy to capture. Monitoring has a storage and processing budget too.

Pools benefit from occupancy, waiters and wait duration. Upstreams need latency, outcome and timeout share. Memory needs distinctions between useful state, cache and leaks. These measurements guide the next check, but none establishes a diagnosis alone. A good panel helps investigate rather than merely assign a label to an incident.

Labels are a contract with bounded, understood values. Full URLs, arbitrary customer names and error text grow without control. Small current traffic does not guarantee low future cost. We review the set of possible values before the metric becomes widely used and expensive to change.

Logs and traces need budgets too. Recording every request body is costly and risks data exposure. Sampling saves resources but can miss rare events. Important errors and operation correlation need explicit rules that preserve useful context without assuming unlimited retention or copying everything by default.

A correlation ID links steps of an individual request; metrics show prevalence. An alert about rising failures should lead toward examples that test a cause. Dashboards and investigation tools need a usable connection. Neither a lone trace nor an aggregate graph supplies the full picture.

We check the instrumentation's effect on the service. Frequent updates, synchronous log delivery and expensive handler calculations can slow the path being measured. Observability should degrade predictably under overload, rather than become another dependency the team must rescue during the same outage.

Use SLOs and error budgets to decide when to alert

An SLO of 99.9% allows 0.1% bad operations over its window. In an example with a million requests, that means 1000 bad ones. An observed 1% failure ratio consumes the allowance ten times faster than its permitted pace: burn rate 10.

With uniform traffic over 30 days, that pace would spend a full budget in roughly three days. Actual traffic varies, some budget may already be spent and a short window can be noisy. Use the calculation with observations and service conditions.

Combining a longer window with a shorter one helps: the first shows meaningful budget consumption, and the second checks whether the problem is still happening. Low traffic needs different conditions. One failure out of ten requests is already 10%, so sample size, operation importance and additional checks matter.

Burn rate relates the bad-event share to the allowed share. The same one percent means an allowed pace under a 99% target and ten times the pace under 99.9%. Duration and scale still influence urgency. The ratio creates a useful basis for response, not a complete incident classification.

The budget window and alert windows have different jobs. A month evaluates the promise; shorter periods detect ongoing trouble. The condition should show meaningful consumption and evidence the failure continues. Otherwise a short burst can page someone after the service has already recovered on its own.

Low traffic can benefit from a synthetic workflow check alongside real events. It does not replace the user SLI, but can expose a fully broken path when few requests arrive. It uses safe data and avoids real charges or unintended business actions. Its coverage is stated rather than assumed.

Error budgets support release decisions too. If quality stays below target, adding risky features may worsen it. The team can address failure causes and change rollout plans. Business participants need that agreement beforehand; otherwise the metric becomes a report that has no influence over work.

A budget is not permission for every kind of bad result. Financial errors, data exposure or corruption can require immediate action despite a low aggregate percentage. An SLO describes a chosen aspect of quality. Other critical guarantees remain independent conditions that the monitoring and response policy must preserve.

Every notification needs a useful next action

Page when delay materially increases harm and a person can intervene. Gradual disk growth may initially need a ticket; a few hours until exhaustion can require immediate action. Severity depends on remaining time and available mitigation.

An alert needs the affected workflow, scale, start time, owner and a runbook. The runbook explains how to check the cause, limit damage and confirm recovery. "Restart the service" is risky without context: a restart may interrupt work, clear local state or increase load on dependencies.

Group symptoms of one failure and define repeat and escalation rules. A separate page from every instance about the same dependency impedes response. Verify suppression relationships so that grouping does not hide an independent problem.

We separate urgent pages, working-hours notifications and planned tasks. One channel for everything loses attention. Each level has an expected response time and an owner able to act. Posting into a shared chat does not assign responsibility or establish that a useful intervention will follow.

The runbook begins by checking the incident: relevant graphs, affected scope and recent changes. Then it explains safe damage reduction and the limits of each action. If the instruction needs a permission unavailable at night, that is an operational risk to resolve before it is needed.

Service and environment names prevent confusion. A host name, opaque code and generic dashboard link make the responder construct context from scratch. The page should explain customer impact and offer an entry into diagnosis, instead of simply forwarding the internal condition that happened to become true.

Escalation handles both no response and no progress. Someone may acknowledge but lack access or get stuck in a hypothesis. Rules for bringing in another specialist and handing over context avoid treating an acknowledged notification as proof that the incident is already being resolved.

We test grouping and repetition using actual scenarios. Related symptoms can share one incident, but independent failures must not vanish behind common silence. Recovery also needs an appropriate signal. Closing a notification does not establish that customers once again receive the correct result.

Test the whole path from failure to recovery

Create a known failure in a safe environment. Confirm the measurement changes, the rule triggers, the on-call receives it, instructions help and recovery is detected. Test missing data and monitoring failure separately.

After an incident, we review missed signals, noise and time to useful action. A noisy alert might have a bad threshold, a bad measurement or a condition that no longer matters. Raising the threshold without finding the cause risks hiding the next failure. Monitoring earns its place when it helps us understand and recover the product.

We test with an expected timeline. We introduce a known fault, mark its start and observe measurement change, rule firing and delivery. This separates collection delay from notification delay. "The alert arrived" alone cannot show where time was lost or whether the response target was met.

Delivery is checked outside the comfortable desk setup. The responder may not have a laptop open, notification permissions may be wrong or a phone may be silent. A realistic agreed channel and fallback matter. The exercise belongs to an agreed on-call process, rather than uncontrolled test noise.

Incident reviews compare existing signals with those actually used. Dozens of graphs can still fail to answer the central question. We remove irrelevant views and add missing context. Dashboards develop around diagnostic tasks instead of displaying every available metric because collection was easy.

Threshold changes are checked against history and the failure the rule should catch. Reducing noise must not remove its original purpose. If workload changed, we revise the condition and explanation. Unowned rules slowly become background noise, especially when nobody can explain what useful action they expect.

Monitoring has its own maintenance cycle: ownership, tests, usefulness review and updates after system changes. New workflows and limits arrive as the product grows. A good launch configuration does not guarantee that next year's team will see the same classes of failure in time.

Building checkout monitoring

Checkout success means a correct order within the agreed duration, not only a fast edge response. We distinguish invalid input from service error, choose measurement location and identify failures it cannot see. Payment stays visible even when fast catalog reads dominate total traffic.

Diagnosis includes queues, pools, upstreams and completion time. Metrics use bounded route templates; individual histories use IDs in logs and traces. We check buckets near the target and absence detection. The dashboard connects the problem's scale to examples useful for an investigation.

Paging requires meaningful budget consumption and evidence of an ongoing fault. Low traffic gains a safe synthetic check, while critical invariants keep independent signals. The notification contains user impact, ownership and a runbook with available actions. Delivery and escalation are exercised rather than assumed.

We introduce a test upstream failure and mark measurement, firing, delivery and useful intervention times. Recovery checks correct orders, not only a closed notification. The exercise shows missing context and delay. Monitoring becomes an operational tool through those checks, rather than through the count of graphs that have been added.

Back to articles