← Back to catalog
Site Reliability Engineering Practices cover image
DevOps Intermediate

Site Reliability Engineering Practices

Define Service Level Objectives (SLOs), manage error budgets, and reduce operational toil.

Instructor Naomi Takahashi
Duration 180 minutes (4 lessons)
Estimated Effort 3 hours total (1.5 hrs/week over 2 weeks)
Price USD 65.00
USD 65.00 Full Lifetime Access

Sign in to track your learning progress.

Course Overview

Translate business expectations into measurable SLIs, set up error budget burn rate alerts, and implement automated toil reduction practices.

What You Will Learn

Formulate meaningful Service Level Indicators (SLIs) and Service Level Objectives (SLOs)
Manage error budgets to balance rapid feature delivery with system stability
Design actionable error budget burn rate alert policies
Identify, measure, and automate operational toil out of daily workflows
Conduct blameless post-mortem incident retrospectives that drive architectural fixes

Tools & Technologies Used

Prometheus Grafana Python 3.11 PagerDuty Yaml

Structured Curriculum

2 Modules  ·  4 Lessons  ·  180 Minutes Total

Module 1

Module 1: SLIs, SLOs & Error Budget Governance

2 lessons

Quantify user experience expectations into measurable service level objectives.

  • 📄

    Defining Meaningful SLIs & Target SLOs

    Select availability and latency metrics that directly correlate with user satisfaction.

    Strategy Session 45 min
  • 📄

    Error Budget Calculation & Policy Enforcement

    Calculate monthly error budgets and establish freezes when budgets are exhausted.

    Hands-on Exercise 45 min
Module 2

Module 2: Burn-Rate Alerting & Toil Automation

2 lessons

Build non-fatiguing alert rules and eliminate repetitive manual toil.

  • 📄

    Multi-Window Multi-Burn-Rate Alerting

    Configure Prometheus alert rules that fire only when error budgets burn at dangerous rates.

    Alerting Lab 45 min
  • 📄

    Toil Audit & Automation Blueprinting

    Identify repetitive manual tasks and replace them with self-healing Python scripts.

    Code Workshop 45 min

Practical Project & Capstone Outcome

🚀 Capstone Project

Production Service SLO Framework & Blameless Post-Mortem

Define complete SLI/SLO specs for a multi-service web backend, configure Grafana error budget dashboards, write multi-burn rate alerts, and author a blameless post-mortem report.

Prerequisites

  • Basic understanding of web backend operations and monitoring

Intended Audience

  • SREs and Operations Engineers formalizing reliability practices
  • Engineering Managers adopting Google SRE framework principles

Instructor Information

N

Naomi Takahashi

Course Author & Industry Expert

Naomi Takahashi is a Lead SRE Consultant who has spent over 12 years building reliable infrastructure systems for high-availability tech companies.

Frequently Asked Questions

Are these practices based on Google SRE principles?

Yes! The course adapts Google SRE principles into practical patterns for teams of any size.