aisle
Module 121 min read

Inference engines

What you’ll learn: Name the four major inference engines, identify what's built-in versus opt-in for each, and predict how the engine choice changes per-replica throughput.

The software layer that turns hardware into a serving cluster. vLLM, TensorRT-LLM, SGLang, Triton: what each is, what they share, and how the engine choice changes the optimization math.