LLM on a Raspberry Pi
How useful a language model can you fit on a Raspberry Pi?
Builder
- Pi 5 + Hailo
- Runs on
- ~15 tok/s
- Throughput
- 0
- Cloud calls
I had a Pi 5, a Hailo accelerator, and a question: could this little board run an assistant I'd actually want to use?
A small, local assistant
The Pi runs a small model on the Hailo NPU, with Open WebUI providing the chat interface and conversation history. I saw around fifteen tokens per second. Inference stays on the device.
Fitting the model into the available memory and compute was the interesting constraint. The setup, Docker configuration, and control scripts are open source.
The Ollama-shaped trap
Hailo's server exposes an Ollama-compatible API, but it isn't Ollama. There is no ollama command, and its compiled HEF models aren't GGUF files. Those distinctions cost me some time.
The guide records the steps that actually worked, including package-installation gotchas. A health-polling script manages the services, with systemd handling startup. If you have the same hardware, it should save you some of the detours.
A closer look
On the parts list