AI Red Teaming: From Prompt Hacking to Infrastructure Defense Roadmap
⏱ 8 weeks · 👤 Beginner (Zero to Hero) · 📋 AI_RED_TEAMING
Comprehensive, production-grade learning path for Ai Red Teaming, architected with foundational-to-advanced pedagogical progression.
Phase 1 Phase 0: Orientation & Mental Models
Establishes the foundational vocabulary, ethical frameworks, and role definitions for AI Red Teaming. Distinguishes between traditional software security and AI-specific risks, setting the stage for technical exploration.
Milestone Ethical Considerations & Responsible Disclosure: Ethical AI and Algorithmic Bias Awareness
Awareness of fairness, algorithmic bias in training data, and ethical responsibility in developing and using AI systems.
fairlearnai-fairness-360ethical-design
ULO-ETHICAL_AI_AND_BIAS_AWARENESS · UNIVERSAL · Bloom: Understand
Awareness of fairness, algorithmic bias in training data, and ethical responsibility in developing and using AI systems.
Milestone Ethical Considerations & Responsible Disclosure: Design Ethics
An ethical framework guiding designers to assess the real-world impact of products, protect data privacy, and proactively mitigate hidden biases before interface design. Applied when evaluating data collection requirements, commercialization features, or interface patterns that may harm users.
inclusive-designprivacy-by-designuser-protection
ULO-DESIGN_ETHICS · UNIVERSAL · Bloom: Understand
An ethical framework guiding designers to assess the real-world impact of products, protect data privacy, and proactively mitigate hidden biases before interface design. Applied when evaluating data collection requirements, commercialization features, or interface patterns that may harm users.
Milestone Role of Red Teams in the AI Lifecycle: AI Governance & Active Red-Teaming
AI governance processes, proactive safety testing (Red-Teaming), and establishing guardrails for autonomous systems.
red-teamingguardrailsai-alignment
ULO-AI_GOVERNANCE_AND_RED_TEAMING · UNIVERSAL · Bloom: Understand
AI governance processes, proactive safety testing (Red-Teaming), and establishing guardrails for autonomous systems.
Milestone Role of Red Teams in the AI Lifecycle: Cognitive Load in Security Operations
The total amount of mental effort and short-term memory a user must expend to process information, make decisions, and perform actions on an interface. Reducing cognitive load optimizes interaction flows and improves task completion efficiency without causing mental fatigue.
progressive-disclosurecognitive-frictionmental-processing
ULO-COGNITIVE_LOAD · UNIVERSAL · Bloom: Understand
The total amount of mental effort and short-term memory a user must expend to process information, make decisions, and perform actions on an interface. Reducing cognitive load optimizes interaction flows and improves task completion efficiency without causing mental fatigue.
Milestone Foundational Knowledge: Supervised vs. Unsupervised Learning Risks: Regression vs. Classification
Distinguish the two main types of supervised learning problems: predicting continuous values (regression) and predicting labels (classification).
supervised-learningtraining-data-biasoverfitting
ULO-REGRESSION_CLASSIFICATION · UNIVERSAL · Bloom: Understand
Distinguish the two main types of supervised learning problems: predicting continuous values (regression) and predicting labels (classification).
Milestone Foundational Knowledge: Supervised vs. Unsupervised Learning Risks: Clustering Algorithms
Understand the concept of clustering and basic algorithms such as K-Means to group similar data points together.
unsupervised-learningclusteringpattern-exploitation
ULO-CLUSTERING_ALGORITHMS · UNIVERSAL · Bloom: Understand
Understand the concept of clustering and basic algorithms such as K-Means to group similar data points together.
Phase 2 Phase 1: The Black-Box Attacker - Prompt Engineering & Jailbreaking
Focuses on the most accessible entry point for AI red teaming: interacting with Large Language Models (LLMs) via natural language. Learners use quick-win techniques to bypass safety filters and extract sensitive information without needing internal model access.
Milestone Prompt Engineering for Attackers: Context Manipulation: Prompt Engineering
Designing effective prompts for large language models: system prompts, few-shot examples, chain-of-thought, tuning temperature/top-p, prompt templates, and prompt quality evaluation.
prompt engineeringsystem promptcontext injectionchatgptclaude
ULO-PROMPT_ENGINEERING · UNIVERSAL · Bloom: Apply
Designing effective prompts for large language models: system prompts, few-shot examples, chain-of-thought, tuning temperature/top-p, prompt templates, and prompt quality evaluation.
Milestone Direct Prompt Injection & Safety Filter Bypasses: Direct Input Threat Classification
The learner will be able to identify potential cybersecurity threats and classify them according to their nature and source.
📚 Prerequisites: REGRESSION_CLASSIFICATION
threat modelingprompt injectiondirect
ULO-THREAT_IDENTIFICATION_AND_CLASSIFICATION · UNIVERSAL · Bloom: Analyze
The learner will be able to identify potential cybersecurity threats and classify them according to their nature and source.
Milestone Jailbreak Techniques: Role-Playing and Obfuscation: Contextual Framing for Model Manipulation
The learner will be able to design effective prompts to generate accurate and efficient code from AI assistants, including specifying context, constraints, and examples.
📚 Prerequisites: ETHICAL_AI_AND_BIAS_AWARENESS, AI_GOVERNANCE_AND_RED_TEAMING
ai assisted codingcopilotrole-playingchatgpt
ULO-AI_CODE_GENERATION_PROMPTING · UNIVERSAL · Bloom: Apply
The learner will be able to design effective prompts to generate accurate and efficient code from AI assistants, including specifying context, constraints, and examples.
Phase 3 Phase 2: Deep Dive - Model Vulnerabilities & Extraction
Moves from surface-level prompting to deeper structural vulnerabilities. Explores how models can be manipulated through adversarial examples, data poisoning, and extraction attacks, revealing the mechanics behind model weights and training data.
Milestone Model Inversion Attacks: Reconstructing Training Data: Data Literacy & DIKW Pyramid
DIKW pyramid model transforming raw Data into Information, Knowledge, and Wisdom.
📚 Prerequisites: AI_CODE_GENERATION_PROMPTING
model weight stealing
ULO-DATA_LITERACY_DIKW_MODEL · UNIVERSAL · Bloom: Analyze
DIKW pyramid model transforming raw Data into Information, Knowledge, and Wisdom.
Milestone Model Extraction & Unauthorized Access to Weights: Unidirectional Data Flow
The learner will be able to understand and apply the principle that state changes propagate in a single direction—from a source of truth to the UI, with user actions triggering state updates through explicit callbacks—to improve predictability and maintainability.
client-serverdata retrievalrate limiting
ULO-UNIDIRECTIONAL_DATA_FLOW · UNIVERSAL · Bloom: Analyze
The learner will be able to understand and apply the principle that state changes propagate in a single direction—from a source of truth to the UI, with user actions triggering state updates through explicit callbacks—to improve predictability and maintainability.
Milestone Data Poisoning: Corrupting Training Pipelines: Scientific Inquiry and Experimental Method
Scientific inquiry process based on questioning, hypothesis building, experiment design, data collection, and conclusion.
experimental methoddata collection
ULO-SCIENTIFIC_INQUIRY_METHOD · UNIVERSAL · Bloom: Analyze
Scientific inquiry process based on questioning, hypothesis building, experiment design, data collection, and conclusion.
Phase 4 Phase 3: Code Execution & System Integration Risks
Addresses the intersection of AI and code execution environments. Focuses on how LLMs integrated with tools (function calling, code interpreters) can be tricked into executing malicious code, leading to Remote Code Execution (RCE) or insecure deserialization.
Milestone Code Injection via Function Calling Interfaces: Agent Tool Use and Function Calling
The learner will be able to enable an AI agent to call external tools and functions, handle structured inputs and outputs, and incorporate results into an ongoing task following a safe and reliable tool-use loop.
langchainopenai-apipydanticjson-schema
ULO-AGENT_TOOL_USE_FUNCTION_CALLING · UNIVERSAL · Bloom: Analyze
The learner will be able to enable an AI agent to call external tools and functions, handle structured inputs and outputs, and incorporate results into an ongoing task following a safe and reliable tool-use loop.
Milestone Remote Code Execution (RCE) through Generated Scripts: Sandboxed Execution Environments for Agent Actions
Architecture for orchestrating tasks for autonomous agents, including problem decomposition (planning), tool selection (tool use), and self-reflection on results.
docker-sandboxfirecrackere2bremote-code-execution
ULO-AGENTIC_WORKFLOW_ORCHESTRATION · UNIVERSAL · Bloom: Evaluate
Architecture for orchestrating tasks for autonomous agents, including problem decomposition (planning), tool selection (tool use), and self-reflection on results.
Milestone Testing Methodologies for Tool-Augmented Models: AI-Generated Test Cases
The learner will be able to use AI to automatically generate unit tests and integration tests from source code, and evaluate their coverage.
test-generationcoverage-analysisfuzzingsecurity-testing
ULO-AI_TEST_CASE_GENERATION · UNIVERSAL · Bloom: Create
The learner will be able to use AI to automatically generate unit tests and integration tests from source code, and evaluate their coverage.
Phase 5 Phase 4: White-Box Mechanics & Adversarial Robustness
Reveals the internal mechanics of how defenses work and fail. Introduces concepts like adversarial training, gradient-based attacks, and the mathematical limits of robustness, providing a white-box perspective on model security.
Milestone Gradient-Based Attacks: Understanding Internal Vulnerabilities: Identifying Logical Failure Modes
Recognize errors in program logic that cause it to run incorrectly, despite not causing syntax or runtime errors.
adversarial exampledecision boundarymodel vulnerability
ULO-LOGIC_ERRORS · UNIVERSAL · Bloom: Analyze
Recognize errors in program logic that cause it to run incorrectly, despite not causing syntax or runtime errors.
Milestone Logit Manipulation & Confidence Score Exploitation: API Response Analysis for Security
Integrating large language models via API: authentication, request/response lifecycle, token limits, rate limiting, retry/backoff, streaming responses, and error handling.
📚 Prerequisites: AI_TEST_CASE_GENERATION, AGENT_TOOL_USE_FUNCTION_CALLING
llm apiresponse parsing
ULO-LLM_API_INTEGRATION · UNIVERSAL · Bloom: Apply
Integrating large language models via API: authentication, request/response lifecycle, token limits, rate limiting, retry/backoff, streaming responses, and error handling.
Milestone Defense Strategies: Input Filtering & Output Sanitization: Validation Rules for Input Controls
The learner will be able to implement client-side validation rules for input controls, including required fields, format checks, and custom validation logic.
input validationsanitization
ULO-CONTROL_VALIDATION_RULES · UNIVERSAL · Bloom: Apply
The learner will be able to implement client-side validation rules for input controls, including required fields, format checks, and custom validation logic.
Phase 6 Phase 5: Production Security & Continuous Monitoring
Covers the operational aspects of securing AI systems in production. Focuses on infrastructure security, API protection, continuous monitoring for drift and attacks, and implementing defense-in-depth strategies.
Milestone API Protection: Rate Limiting, Authentication, and Authorization: Zero Trust Principles for Service-to-Service Communication
Network security design principle 'Never trust, always verify' with continuous authentication and least-privilege access.
mtlsservice-meshistioenvoyrbac
ULO-ZERO_TRUST_ARCHITECTURE · UNIVERSAL · Bloom: Understand
Network security design principle 'Never trust, always verify' with continuous authentication and least-privilege access.
Milestone Continuous Monitoring: Detecting Drift and Anomalous Behavior: TinyMLOps Fleet Lifecycle Management
Process for lifecycle management, performance monitoring, and automated model updates for machine learning on millions of embedded edge devices.
prometheusgrafanalangsmitharizedatadog
ULO-TINY_MLOPS_LIFECYCLE_MANAGEMENT · UNIVERSAL · Bloom: Apply
Process for lifecycle management, performance monitoring, and automated model updates for machine learning on millions of embedded edge devices.
Milestone Infrastructure Security: Securing Vector Stores and Embeddings: Vector Embedding and Similarity Search
Method of representing text or objects as dense vectors in high-dimensional space and searching for similarity based on spatial distance.
pineconechromamilvuspgvectorencryption-at-rest
ULO-VECTOR_EMBEDDING_SEARCH · UNIVERSAL · Bloom: Apply
Method of representing text or objects as dense vectors in high-dimensional space and searching for similarity based on spatial distance.