File 002 · Local Intelligence
Record 056 · Vision · 2023
23.56.ILLaVA
UW / Wisconsin + collaborators

“The open vision-language model that made 'look at this screenshot' a local call.”
A pair of eyes on the local model.
LLaVA connected a vision encoder to a Llama-class LLM. Moondream, Llama 3.2 Vision, and Qwen-VL sit on this idea.
Opened local multimodal chat.
Filed notes
- Open VLM architecture
- Visual instruction tuning
- Screenshot/document VQA
SIC
7372
Prepackaged Software
NAICS
511210
Software Publishers




