New: Roadmaps ordered paths through our cheat sheets and flashcards, so you always know what to study next.
Explore themSee what's new on GitHubFrom your first prompt to running a full AI red-team engagement.
A 12-step learning path. Follow it in order, or jump to what you need.
This path is for security practitioners and AI engineers who want to attack LLM systems on purpose: probing prompt injection, jailbreaks, and agentic exploits before real attackers get the chance. It runs roughly six to ten weeks at a few hours a week. This path stays on the offensive, technical side of AI security; for the policy, bias, and regulatory side of responsible AI, see the AI Governance & Ethics path instead. By the end you can craft and defend against prompt injection and jailbreak attacks, stress-test an LLM's guardrails and agentic tool access, and scope and run a real AI red-team engagement from rules of engagement through to a written report.
Expected: comfortable using LLM-based tools and writing prompts, plus a general software or IT background. Helpful but not required: hands-on penetration testing or application security experience.
You write real, working prompts today, and step 4 shows exactly how attackers twist that same craft into a jailbreak.
Opening the hood on what step 1's prompts actually trigger inside the model sets up why attention and instruction-following break down under the attacks in step 4.
You can write working prompts and explain what's actually happening inside the model that receives them. Next up: mapping every way that machinery can be turned against itself.
Finish this section to unlock.
+100 XP
STRIDE gives you the vocabulary security teams use to map what can go wrong before touching a live system, the same structured habit step 12's engagement depends on.
This is where it clicks that the model itself, not the code around it, is the attack surface; prompt injection and jailbreaks are where most security veterans get stuck relearning instinct, and you'll keep this sheet open as your working reference from here on.
You can name exactly how a model gets attacked, from prompt injection to data poisoning, and it clicks that the model itself, not just the code around it, is the attack surface. Next up: building and testing the defenses that are supposed to stop it.
Finish this section to unlock.
+100 XP
Every attack vector step 4 named has a matching defense layer here, so you can test whether a system's guardrails actually stop what you just learned to throw at them.
Guardrails from step 5 assume a single prompt-response exchange, but agents that call tools and hold memory open entirely new doors: tool misuse, memory poisoning, and MCP rug-pulls that a guardrail alone won't catch.
You now turn what steps 4 through 6 uncovered into repeatable evidence: benchmarking and adversarial evaluation are how a single successful jailbreak becomes a defensible, measured finding.
You can point to specific guardrail layers, agent-specific weak points, and hard evaluation numbers instead of gut feel. Next up: turning all of it into a real, authorized engagement, with a few minutes of due flashcards each day keeping steps 1 through 7 sharp while you build toward it.
Finish this section to unlock.
+100 XP
Everything so far has been technique; before you point any of it at a real system, you need the authorization, scope, and rules of engagement that keep testing legal and safe.
With scope and authorization from step 8 in hand, this is where you pick black-box, white-box, or grey-box testing, load PyRIT or Garak, and practice on labs like Gandalf before touching a real target.
Take this if you're heading toward securing the APIs and tool endpoints your red-teamed agents call, not just the model sitting behind them.
Take this route if you're heading toward a formal threat-intel practice: mapping AI-specific findings onto ATT&CK's tactics gives your reports a vocabulary defenders already trust.
This is the whole path in one engagement: threat models, prompt attacks, guardrail gaps, and agent risks all come together into the recon-exploit-report cycle real red teams run.
You can scope an authorized AI red-team engagement, throw prompt attacks, agent exploits, and guardrail bypasses at a real target, and turn what you find into a report defenders can act on. That's what the AI Red Teamer badge certifies: you don't just use AI systems, you know exactly how they break.
Finish this section to unlock.
+100 XP
Finish every required step, at least 70% of them genuinely done (not skipped), to earn this badge and 500 XP.