Skip to content

Google's SRE Principles for Development Teams: Concerns and Answers

By BPMBI Engineering

This is a practical walkthrough of the core ideas in Google’s Site Reliability Engineering (Beyer, Jones, Petoff, Murphy; O’Reilly, 2016), free at sre.google/books, paraphrased for development teams who are being asked to adopt SRE practices and have reasonable doubts about it. Examples below are illustrative and drawn from the kind of large-enterprise systems we work on — retail checkout, banking, insurance — not any specific client’s architecture.

The short version: this deck is for development teams, not SRE teams. It explains the principles first, then spends real time on the concerns developers most often raise — speed, pager duty, process, blame, and legacy systems. The goal is honest answers, not a sales pitch for SRE.

Why SRE exists

Two goals pull against each other inside most engineering organizations. Product development wants to ship features quickly and is rewarded for velocity. Operations wants stable services and is rewarded for uptime — to that team, change is the enemy. SRE’s answer is that engineers on both sides use the same engineering approach and agree on one measurable reliability target that says how much risk is acceptable: a shared error budget.

That’s the idea the rest of this piece builds on. The traditional split between development and operations is a conflict of incentives; SRE aligns them by making reliability a measured target both sides own.

The vocabulary: SLI, SLO, and SLA form a ladder

  • SLI (Service Level Indicator) — a measurement. Example: the share of checkout API requests that succeed within 500ms.
  • SLO (Service Level Objective) — a target. Example: 99.9% of checkout requests succeed within 500ms, over a rolling 30-day window.
  • SLA (Service Level Agreement) — an agreement with consequences. A contract with customers or partners, with penalties, usually looser than the SLO.

Keep the SLO tighter than any SLA so you have room to react before penalties apply. The checkout numbers above are illustrative — pick SLIs by what the system actually does. For each of your services, ask what the user would complain about, then find the measurement that tracks it. A few meaningful indicators beat many: start from what a customer or support agent notices, then find the metric that moves when they would complain.

The math: what a target allows in a 30-day window

30 days is 43,200 minutes. The allowed downtime at a given reliability target drops fast:

Target Informally Allowed downtime / 30 days
99% Two nines 7.2 hours
99.9% Three nines 43.2 minutes
99.99% Four nines 4.32 minutes
99.999% Five nines ~26 seconds

A deploy that fails for five minutes exceeds the entire four-nines monthly budget. That’s why release safety matters more as the target climbs — at four nines, one slow rollback can consume the whole month’s budget.

Error budgets in practice

Picture the remaining budget as a line that burns down over a 30-day rolling window. It burns steadily day to day, an incident consumes a chunk of it, and if it hits zero before the window resets, feature releases pause under a policy the team agreed to before the incident — not during it:

  • Budget remaining? Ship at a normal pace, take measured risks, experiment.
  • Budget spent? Pause feature releases (urgent and security fixes only), prioritize reliability work, and run a postmortem. Disagreements go to a named escalation owner, not to whoever argues loudest.

The policy is deliberately boring. That’s the point — a pre-agreed rule removes the emotion from feature-versus-stability arguments. Your organization should write its own thresholds and name its own escalation owner; the pattern above is just the shape of it.

SRE operational flow: telemetry collection feeds SLO and error-budget evaluation, which drives alerting on budget burn, which triggers incident response and runbooks, which feeds a post-mortem and review step that loops back into SLO evaluation.
The same operational loop error budgets, alerting, and postmortems all feed into — see our SRE Practice Modernization page for the full picture.

Eight concerns development teams raise, and straight answers

These are legitimate concerns, and the book has a reasonable answer to each one.

Will SRE slow our releases? Error budgets exist to let you ship — while budget remains, you decide when to release, and releases are checked against data, not queued for a person to approve. Only when the budget is spent does change slow, and by a rule you agreed to in advance. You keep release control while budget remains; a spent budget triggers a rule, not a debate.

Who carries the pager? The book caps SRE on-call at about a quarter of an engineer’s time, with a design goal of no more than two incidents per shift so each one gets time to be handled and reviewed properly. Who holds the pager for a given service is negotiated, not assumed, and can go back to the developers. By default you keep first response for your own service; pager handover needs readiness first, not just willingness.

Isn’t this just ops with a new name? SRE is engineering applied to operations, with a hard ceiling on manual work. Toil — repetitive, automatable work with no lasting value that grows with the service — is capped at roughly half of an SRE’s time; teams that can’t stay under the cap push work back or automate it. Expect software engineers, not ticket handlers, and expect toil counts to be visible to everyone, not hidden in a backlog.

When it breaks, who gets blamed? Postmortems focus on causes in systems and processes, not on individuals, because people have to feel safe reporting what happened or learning stops. Every significant incident gets a written review with tracked actions. In a blameless culture the team reconstructs the timeline together, fixes systemic causes, and tracks action items with owners and dates — the opposite of hunting for a culprit while people quietly hide details.

Shouldn’t we aim for 100% uptime? The book argues 100% is the wrong target for nearly every service. Past a certain point, users can’t tell the difference — their own network and devices are less reliable than that anyway — while extra reliability keeps costing more and slowing feature work. Pick a target that matches actual user need, spend the surplus reliability budget on speed, and revisit the target as needs change.

Our legacy system can’t meet an SLO. Start by measuring what the service delivers today, then set an SLO the system can actually achieve — SLOs are refined with experience, not set once and frozen, and even a rough indicator beats none. Legacy and vendor-run systems are common across large enterprises — core banking platforms, retail inventory systems, and insurance policy-administration systems chief among them. Baseline first, give old systems realistic targets, and schedule improvement work from the data rather than starting with an aspirational number the system can’t meet.

More alerts means more noise. Every page should be urgent, actionable, and need a human. Alert on user-visible symptoms and SLO burn rate, not on every internal cause, and remove or downgrade alerts that don’t need action. A page means a human must act now; a ticket means act within days; a log or dashboard entry means diagnose with no action needed. Noisy alerts get treated as bugs, and every alert should link to a runbook.

Reviews are bureaucracy before launch. Launch and readiness reviews are meant to be a partnership that starts at design time, not a gate at the end — late reviews are painful because changes cost more the later they arrive. A short checklist finds gaps early, when they’re cheap to fix. Invite reviewers at design time, use the checklist instead of a meeting, and close the review when the gaps are closed.

Monitoring: the four golden signals

If you can only monitor four things about a user-facing service, monitor these:

  • Latency — time to serve a request; track failed and successful requests separately. Example: checkout response time.
  • Traffic — demand on the system. Example: requests per second at peak.
  • Errors — rate of failed requests, explicit or implicit. Example: failed payment authorization.
  • Saturation — how full the service is. Example: queue depth, memory, database load.

The goal is to detect symptoms users actually feel, then use richer data to diagnose once you’ve been alerted.

Safe change: small, progressive, reversible releases

About 70% of outages come from changes to a live system, per the book’s introduction — which is why release engineering is where reliability is won. A progressive rollout (1% → 5% → 25% → 100%, with automated checks at each stage and a fast rollback if signals degrade) gives you a chance to catch problems while the audience is still small:

  • Automate the pipeline. Repeatable builds and deploys.
  • Limit blast radius. Canary first, expand on healthy signals.
  • Make rollback easy. Fast, rehearsed, low drama.

Incident management: declared roles keep incidents calm

The book adapts an emergency-response model: one person coordinates, others do specific jobs. Roles are declared early and handed over explicitly, and developers are often the operations lead for their own service:

  • Incident commander — holds the big picture, delegates, decides.
  • Operations lead — does the technical work of mitigation.
  • Communications lead — keeps stakeholders and customers informed.
  • Planning lead — tracks actions, notes, and handoffs.

This is where developers and SRE work most closely together. Practice with drills so the roles are familiar before a real incident — not during one.

Designing for resilience

A regional outage, a seasonal sales event, or a breaking news cycle can send a sudden traffic surge at your service — the same dynamic that hits a retail platform on a peak shopping day, a bank’s app after a rate-change announcement, or an insurer’s portal after a natural disaster. Without designed defenses, a system that doesn’t shed load simply collapses under offered load past its capacity limit; one that sheds excess load stays steady. Three concrete defenses:

  • Shed load early. Reject some requests cleanly instead of failing all of them.
  • Retry with care. Backoff and limits — naive retries amplify outages, they don’t fix them.
  • Degrade gracefully. Serve a reduced experience rather than none at all.

The partnership

A working agreement runs both directions. What developers get from SRE: shared tooling for monitoring, rollout, and incidents; coaching on SLOs and design for failure; data to justify reliability work to product; and relief from repetitive operations. What SRE asks of developers: a defined SLO and error-budget policy; instrumented services with actionable alerts; ownership of postmortem actions; and a readiness review for new services. Shared ground: the SLOs themselves, the error budget, on-call load, and postmortem learning.

Get started this quarter

A concrete first quarter that needs no new tooling beyond what you already have:

  1. Map user journeys. List the top journeys your service supports.
  2. Choose 2 or 3 SLIs. Pick measurements a user would actually notice.
  3. Set a first SLO. Base it on measured data, and start loose.
  4. Add burn-rate alerts. Alert on SLO burn, not on every underlying cause.
  5. Run a blameless review. Review the next incident together, as a team.

Choose one service — ideally one with a clear user journey — and do all five steps before scaling to more services.

Where to read more

Reliability is a feature you can measure and budget for. The book’s chapters map directly to the concerns above: Embracing Risk, Service Level Objectives, Eliminating Toil, Monitoring Distributed Systems, Being On-Call, Managing Incidents, Postmortem Culture, and Handling Overload and Cascading Failures. It’s free online at sre.google/books, alongside the companion workbook.

If you’re weighing where to start on your own services, that’s exactly what our SRE Practice Modernization engagements are built around — baselining your current incident load and telemetry, then setting the first SLOs with your team rather than for them.

↑ Back to top

Ready to talk architecture?

Request a technical discovery call with our engineering team.

Request Discovery Call