File 002 · Local Intelligence

Record 043 · Runtimes · 2023

23.43.U

vLLM

vLLM team

vLLM (2023) by vLLM team

The serving engine that made a single box look like a proper LLM API.

A local server that feels like an API farm.

PagedAttention, continuous batching, OpenAI-compatible endpoints. Homelabs and offices run vLLM when they outgrow a single chat window.

Opened high-throughput local serving.

Filed notes

  • PagedAttention
  • OpenAI-compatible server
  • Multi-GPU local serve

SIC

7372

Prepackaged Software

NAICS

511210

Software Publishers

Official

Also in this file