New: Roadmaps ordered paths through our cheat sheets and flashcards, so you always know what to study next.
Explore themSee what's new on GitHubFrom your first error budget to the metrics that prove your reliability work is real.
A 12-step learning path. Follow it in order, or jump to what you need.
This path is for developers or DevOps engineers ready to specialize in keeping production systems reliable instead of just shipping to them. Plan on about 6 to 8 weeks at 3 to 5 hours a week, moving from your first error budget to the five DORA numbers that prove reliability work actually pays off. It goes deep on the incident, toil, and chaos-engineering practices that DevOps only sketches, and it does not cover container orchestration or infrastructure provisioning at all, see the DevOps Engineer and Platform Engineer paths for that ground. By the end you can define an SLO and defend an error budget in a planning meeting, run an incident from first alert to blameless postmortem, and run a chaos experiment that proves your system survives what you think it survives.
Expected: comfort with Linux, Git, CI/CD pipelines, Docker, and Kubernetes, the ground the DevOps Engineer path covers. Helpful but not required: some time already carrying a pager.
Open this and you'll finally have the vocabulary, error budgets, SLOs, toil, that turns 'keep it up' into a discipline you can actually practice, the map every later step in this path fills in.
Once the map from step 1 makes sense, ground it in real signals: logs, metrics, and traces are what an SLO in the next step actually gets measured against.
This is where reliability stops being a feeling and becomes a number you can defend in a meeting; expect to come back to error budgets more than once as the idea settles in.
You can name the four golden signals and defend an SLO with a real error budget behind it instead of a vague uptime promise. Next up: what happens when that budget actually runs out.
Finish this section to unlock.
+100 XP
With a real error budget from step 3 to protect, the next skill is running the actual response when something burns through it: detection, triage, and recovery in a structured order.
Every incident you just ran through in step 4 leaves a choice: write it down and learn, or repeat it next quarter, this is how SRE teams choose the former.
The postmortem habit from step 5 keeps surfacing the same manual fixes, so learn to name that pattern as toil before it quietly eats your whole week.
Take this if you want the golden signals from step 2 wired into a real time-series database and PromQL queries instead of staying theoretical.
Toil identified in step 6 is only half the job: turn the repetitive fixes into executable runbooks so the next page fixes itself instead of waking you up at 3am.
You can run an incident from first alert through a blameless writeup and turn the fix into a runbook that handles it on its own next time; a few minutes of due flashcards keeps the error budget math from step 3 sharp while you build on it here. Next up: breaking things on purpose to prove all of this actually holds.
Finish this section to unlock.
+100 XP
Automated runbooks from step 8 only help after an incident starts; this is where you flip the model and break things on purpose, expect it to feel reckless at first until the first caught bug proves otherwise.
Head here if you want the architectural fixes for what step 9's experiments expose: circuit breakers and backoff strategies that stop one failure from cascading into ten.
Chaos experiments in step 9 test how you fail; this is about making sure you have enough room to not fail in the first place when traffic doubles overnight.
Every practice in this path, monitoring, SLOs, incident response, toil reduction, finally shows up in five numbers leadership actually reads, proof that reliability work moves the business and not just the pager.
You can break your own system on purpose, plan for the traffic spike before it happens, and put a number on how reliable and fast your team actually is, the complete loop from error budget to evidence. That's the Site Reliability Engineer badge, earned.
Finish this section to unlock.
+100 XP
Finish every required step, at least 70% of them genuinely done (not skipped), to earn this badge and 500 XP.