File 002 · Local Intelligence

Record 001 · Runtimes · 2023

23.01.U

llama.cpp

Georgi Gerganov

llama.cpp (2023) by Georgi Gerganov

The C++ engine that made a 7B model a laptop object, not a cluster job.

The runtime that fit in RAM.

Gerganov quantized and ported Llama to plain C/C++ with GGML, then GGUF. CPU, Metal, CUDA, Vulkan — one repo. Every later local UI is a frontend to this idea.

Opened local LLM inference as something you run, not rent.

Filed notes

  • GGML/GGUF quantization
  • CPU-first then GPU backends
  • Single-binary local inference

SIC

7372

Prepackaged Software

NAICS

511210

Software Publishers

Official

Also in this file