AI Safety
Understand the risks of AI systems and the practices that keep them aligned, robust and fair — from evals to governance.
Why AI Safety Matters
The stakes, the risks, and the field that studies how to build AI that behaves.
Bias & Fairness
Measure and mitigate bias: metrics, audits and the human choices behind them.
Interpretability & Explainability
Understand why models decide: feature importance, saliency and post-hoc explanations.
Alignment
Make AI goals match human values — and the tricky parts of defining those values.
Robustness & Adversarial Examples
Models that break under tiny perturbations — and how to harden them.
Privacy & Data Protection
PII, memorization and differential privacy in AI systems.
Hallucination & Factualness
When models make things up — measure it, reduce it, and tell users about it.
AI Governance & Policy
Laws, standards and internal governance for responsible AI.
Transparency & Disclosure
Tell people when AI is involved: disclosure, provenance and model cards.
Safety Evaluations
Test systems for harmful behavior before release — systematically.
Red Teaming
Adversarial testing by skilled humans — find failures before users do.
Guardrails & Content Moderation
Layer protections on outputs: classifiers, filters and policy enforcement.
Auditing AI Systems
Independent review of data, models and processes — internal and external.
Societal Impact of AI
Jobs, information ecosystems, inequality and the second-order effects of AI.
AI Safety Case Studies
Learn from real incidents: what broke, why, and what changed.
Data Governance for AI
Source, quality, consent and retention — the data practices under safe AI.
Designing for Human Values
Turn values into requirements: privacy, fairness, dignity as design inputs.
Emerging & Frontier Risks
Synthetic media, autonomous systems and the longer horizon of powerful AI.
Building Responsible AI Products
Integrate safety into the product lifecycle from day one.
The AI Safety Community & Research
Where to learn, contribute and stay current in safety research.
AI Safety Roadmap
Synthesize the course into a plan: apply safety to every system you build.

