Module 121 min read
Inference engines
What you’ll learn: Name the four major inference engines, identify what's built-in versus opt-in for each, and predict how the engine choice changes per-replica throughput.
The software layer that turns hardware into a serving cluster. vLLM, TensorRT-LLM, SGLang, Triton: what each is, what they share, and how the engine choice changes the optimization math.