Skip to main content

Compiler Engineering Roadmap

From Source Code to IR, Optimization, Runtime, and Hardware

Part 1: Compiler Worldview and Engineering Mental Model

Chapter 1: Compiler Transformation Model

Chapter 2: Source Code as a Structured System

Chapter 3: Compilation as a Sequence of Representations

Chapter 4: Interpreter, Compiler, Transpiler, and JIT

Chapter 5: Compiler Correctness, Performance, and Developer Experience

Part 2: Source Text, Tokens, and Lexical Analysis

Chapter 6: Source Text as Raw Characters

Chapter 7: Tokens as the First Structured Representation

Chapter 8: Lexical Rules, Keywords, Identifiers, and Literals

Chapter 9: Whitespace, Comments, Newlines, and Layout-Sensitive Syntax

Chapter 10: Lexer Errors, Diagnostics, and Source Locations

Part 3: Grammar, Parsing, and Abstract Syntax Trees

Chapter 11: Grammar as the Shape of Valid Programs

Chapter 12: Parsing Token Streams into Structure

Chapter 13: Parse Trees vs Abstract Syntax Trees

Chapter 14: Operator Precedence, Associativity, and Ambiguity

Chapter 15: Parser Error Recovery and Developer Feedback

Part 4: Semantic Analysis, Symbol Tables, and Scopes

Chapter 16: Names, Bindings, and Program Meaning

Chapter 17: Symbol Tables and Scope Chains

Chapter 18: Declarations, Definitions, and Resolution

Chapter 19: Modules, Imports, and Visibility Rules

Chapter 20: Semantic Errors Beyond Syntax

Part 5: Type Systems, Type Checking, and Type Inference

Chapter 21: Types as Compile-Time Program Facts

Chapter 22: Static Typing, Dynamic Typing, and Gradual Typing

Chapter 23: Type Checking Expressions, Statements, and Functions

Chapter 24: Generics, Templates, and Parametric Polymorphism

Chapter 25: Type Inference and Constraint Solving

Part 6: Intermediate Representation and Program Lowering

Chapter 26: Intermediate Representation as Compiler Working Language

Chapter 27: High-Level IR vs Low-Level IR

Chapter 28: Lowering as Controlled Loss of Abstraction

Chapter 29: IR Verification and Structural Invariants

Chapter 30: Designing IR for Analysis and Transformation

Part 7: Control Flow Graphs, Basic Blocks, and Program Structure

Chapter 31: Basic Blocks as Straight-Line Code Regions

Chapter 32: Control Flow Graphs and Branch Structure

Chapter 33: Dominators, Loops, and Reachability

Chapter 34: Structured Control Flow vs Unstructured Jumps

Chapter 35: Control Flow Evidence in Real Compiler Pipelines

Part 8: Data Flow Analysis, Def-Use Chains, and Program Facts

Chapter 36: Program Facts and Fixed-Point Thinking

Chapter 37: Reaching Definitions and Live Variables

Chapter 38: Use-Def Chains and Value Tracking

Chapter 39: Forward vs Backward Data Flow Analysis

Chapter 40: Data Flow as the Basis of Optimization

Part 9: SSA Form and Modern IR Design

Chapter 41: Static Single Assignment and Value Identity

Chapter 42: Phi Nodes, Block Arguments, and Value Merging

Chapter 43: SSA Construction and Destruction

Chapter 44: SSA-Based Optimizations

Chapter 45: SSA as the Language of Modern Optimizers

Part 10: Optimization Passes and Transformation Pipelines

Chapter 46: Semantics Preserving Transformation Criteria

Chapter 47: Analysis Passes vs Transform Passes

Chapter 48: Constant Folding, DCE, CSE, and Inlining

Chapter 49: Pass Ordering and Optimization Pipelines

Chapter 50: Miscompilation as the Dark Side of Optimization

Part 11: Memory, Alias Analysis, and Side Effects

Chapter 51: Memory as the Hard Part of Program Analysis

Chapter 52: Pointers, References, and Aliasing

Chapter 53: Escape Analysis and Object Lifetime

Chapter 54: Side Effects, Volatile, and Observable Behavior

Chapter 55: Optimization Under Memory Uncertainty

Part 12: Loop Optimization, Vectorization, and Parallelism

Chapter 56: Loop Dominance in Performance Engineering

Chapter 57: Loop Invariant Code Motion and Strength Reduction

Chapter 58: Loop Unrolling, Fusion, Fission, and Tiling

Chapter 59: Auto-Vectorization and SIMD Code Generation

Chapter 60: Parallelism, Dependence Analysis, and Safety

Part 13: Backend Fundamentals and Instruction Selection

Chapter 61: From IR Operations to Target Instructions

Chapter 62: Instruction Selection and Pattern Matching

Chapter 63: Legalization and Target Constraints

Chapter 64: Machine IR and Target-Specific Lowering

Chapter 65: Backend Correctness and Target Semantics

Part 14: Register Allocation, Stack Frames, and Calling Conventions

Chapter 66: Registers as Scarce Execution Resources

Chapter 67: Liveness, Interference, and Register Allocation

Chapter 68: Spilling, Reloading, and Stack Slot Management

Chapter 69: Stack Frames, Prologues, and Epilogues

Chapter 70: Calling Conventions as Binary-Level Contracts

Part 15: Object Files, Relocation, Linking, and Debug Information

Chapter 71: Object Files as Partially Built Programs

Chapter 72: Symbols, Sections, and Relocation Records

Chapter 73: Static Linking, Dynamic Linking, and Loaders

Chapter 74: ELF, Mach-O, COFF, and Platform Formats

Chapter 75: Debug Information and Source-Level Observability

Part 16: Runtime Systems, ABI, Exceptions, and Garbage Collection

Chapter 76: What the Compiler Leaves to the Runtime

Chapter 77: ABI, Runtime Helpers, and Language Support Libraries

Chapter 78: Exception Handling and Stack Unwinding

Chapter 79: Garbage Collection and Memory Safety Support

Chapter 80: Runtime Systems as Execution Partners

Part 17: Interpreters, Bytecode VMs, and JIT Compilation

Chapter 81: AST Interpreters and Direct Execution

Chapter 82: Bytecode as a Compact Execution Format

Chapter 83: Virtual Machines and Evaluation Loops

Chapter 84: JIT Compilation and Runtime Specialization

Chapter 85: Deoptimization, Inline Caches, and Dynamic Feedback

Part 18: LLVM Infrastructure and Pass Engineering

Chapter 86: LLVM as a Modular Compiler Infrastructure

Chapter 87: LLVM IR, Modules, Functions, and Basic Blocks

Chapter 88: The LLVM Pass Manager and Analysis Preservation

Chapter 89: Building and Testing LLVM Passes

Chapter 90: Reading Optimized LLVM IR

Part 19: MLIR, Dialects, and Multi-Level Compiler Infrastructure

Chapter 91: Multi-Level IR and Progressive Lowering

Chapter 92: Dialects as Domain-Specific Compiler Languages

Chapter 93: Operations, Regions, Attributes, and Types

Chapter 94: Dialect Conversion and Progressive Lowering

Chapter 95: MLIR in Heterogeneous and Domain-Specific Compilation

Part 20: CPython Compilation Pipeline and Bytecode VM

Chapter 96: From Python Source to Tokens

Chapter 97: Parsing Python into AST

Chapter 98: Symbol Table, Scopes, and Code Objects

Chapter 99: AST to CFG to Bytecode

Chapter 100: Evaluation Loop and Runtime Objects

Part 21: WebAssembly and Portable Runtime Targets

Chapter 101: WebAssembly as a Portable Compilation Target

Chapter 102: Stack Machine, Linear Memory, Tables, and Modules

Chapter 103: Validation, Security, and Sandboxing

Chapter 104: WASI and Host Environment Interfaces

Chapter 105: JIT, AOT, and Runtime Embedding

Part 22: Shader Compilers, SPIR-V, and GPU Code Generation

Chapter 106: Shader Source as a Specialized Program

Chapter 107: GLSL, HLSL, WGSL, and Frontend Differences

Chapter 108: SPIR-V and GPU-Oriented Intermediate Representation

Chapter 109: Driver Compilers and GPU Machine Code

Chapter 110: Graphics Pipeline State and Shader Optimization

Part 23: AI Graph Compilers, Tensor IR, and Operator Fusion

Chapter 111: Neural Networks as Computation Graphs

Chapter 112: Tensor Shapes, Layouts, and Type-Like Information

Chapter 113: Graph Optimization and Operator Fusion

Chapter 114: Lowering Tensor IR to Kernels

Chapter 115: Runtime Scheduling, Memory Planning, and Hardware Backends

Part 24: Database Query Optimizers as Compiler Systems

Chapter 116: SQL as a Declarative Source Language

Chapter 117: Parsing, Binding, and Logical Query Plans

Chapter 119: Physical Operators and Execution Plans

Chapter 120: Query Compilation, JIT, and Vectorized Execution

Part 25: Compiler Diagnostics, Error Recovery, and Developer Experience

Chapter 121: Diagnostics as a Compiler User Interface

Chapter 122: Source Ranges, Notes, Hints, and Fix-Its

Chapter 123: Syntax Error Recovery

Chapter 124: Semantic Error Reporting

Chapter 125: Designing Diagnostics for Human Understanding

Part 26: Compiler Testing, Fuzzing, Verification, and Miscompilation Debugging

Chapter 126: Golden Tests, Unit Tests, and Integration Tests

Chapter 127: IR Tests and Pass-Level Verification

Chapter 128: Fuzzing Parsers, Optimizers, and Code Generators

Chapter 129: Differential Testing and Miscompilation Detection

Chapter 130: Reducing and Debugging Compiler Bugs

Part 27: Building a Small Compiler from Scratch

Chapter 131: Designing a Tiny Language

Chapter 132: Implementing Lexer, Parser, and AST

Chapter 133: Building Semantic Analysis and IR

Chapter 134: Adding Optimizations and Bytecode Execution

Chapter 135: Generating LLVM IR or WebAssembly