Describe one change
A useful runbook names a specific change and a specific system. “Update the server” is too broad to verify. “Replace the application release while keeping the previous release available” gives the operator something concrete to observe.
Begin with the current version, target version, affected service and person responsible. Include the maintenance window if there is one. List dependencies such as configuration, database changes or a package repository. Treat these as separate moving parts rather than assuming they roll back together.
Define success from the reader’s side
A running process is only one signal. Check the request path that people use, along with error rate and latency where measurements exist. The Google SRE monitoring chapter offers a useful foundation for distinguishing internal signals from externally visible symptoms.
Write three explicit outcomes: normal behaviour, a warning worth investigating, and a failure that stops the change. Choose thresholds from the service’s baseline and requirements. Do not copy arbitrary numbers from an unrelated system.
Rehearse the recovery path
Confirm access to the previous release and the configuration it expects. A deployment rollback may not reverse data transformations. If a migration is not reversible, plan a compatible application path or a separate data recovery procedure before proceeding.
After the change, retain the observations and note any unexpected step. Update the runbook while the details are fresh. A short, accurate document is more valuable during an incident than a long checklist nobody has tried. This is a planning pattern, not a substitute for the specific product’s upgrade instructions.
Before you finish
- Scope and dependencies named
- Stop condition written
- Recovery path rehearsed
- User-facing checks included
Technical reference
Google SRE: monitoring distributed systems
