← Back to the workbench

Vision Pipeline

A camera feed that more than one model can make sense of at once.

Builder

Redis perception bus
Architecture
YOLO + Moondream
Live demo
Mac + Jetson
Runs on
Where it started

I wanted a reusable starting point for robotics: capture a scene once, then let different models work from the same feed.

One feed, several subscribers

A camera captures frames and publishes their file locations over Redis. Models subscribe independently, so I can add another detector or vision-language model without rebuilding the capture path.

The current demo runs YOLO for bounding boxes and Moondream for questions about the scene. A continuous-narration mode repeats a chosen question and reads new descriptions aloud. Actuator control is a future step, not a capability of this build.

The Mac GPU wrinkle

Docker on the Mac couldn't use the GPU for this workload. The models run natively with Metal/MPS acceleration, while Redis, the API, and the React frontend stay in containers.

There's also a Jetson Orin Nano version using a Raspberry Pi camera. Both repositories are linked above.

On the parts list

PythonPyTorchYOLO11Moondream2RedisFastAPIReactJetson
More from the collection