File 002 · Local Intelligence
Record 043 · Runtimes · 2023
23.43.UvLLM
vLLM team

“The serving engine that made a single box look like a proper LLM API.”
A local server that feels like an API farm.
PagedAttention, continuous batching, OpenAI-compatible endpoints. Homelabs and offices run vLLM when they outgrow a single chat window.
Opened high-throughput local serving.
Filed notes
- PagedAttention
- OpenAI-compatible server
- Multi-GPU local serve
SIC
7372
Prepackaged Software
NAICS
511210
Software Publishers




