Computer vision · Edge hardware

SentinelVision

Multi-camera detection running on a Raspberry Pi. It watches for weapons, people and vehicles, streams every camera live, keeps an event log, and tells you about the things that matter, in tiers, so an urgent thing looks urgent.

Python YOLO / NCNN OpenCV FastAPI SQLite Tailscale systemd
person 0.90
car 0.83
CAM-01 · NCNN · LIVE

How it decides

Three tiers, three different urgencies

Not everything deserves the same alarm. A weapon should interrupt you. A car passing should not.

Tier 1

Weapons

Firearms and knives. Urgent-flagged notification, an annotated snapshot, and a short video clip covering the seconds either side, because a still frame only shows that a box was drawn.

Tier 2

People

A bounding box around a human shape, nothing more. Alerts fire on arrival, then stay quiet: someone standing in view for an hour produces one notification, not sixty.

Tier 3

Vehicles

Cars, trucks, motorcycles. Gated on arrival rather than presence, so a car parked in the driveway is reported once instead of every two minutes for as long as it sits there.


Architecture

The small computer watches. The big one thinks.

Neural networks were the entire thermal budget of a Raspberry Pi. It ran hot enough to throttle itself off the network. So inference moved to a machine with headroom, and the Pi kept the jobs it is good at.

Cameras USB · CSI Raspberry Pi capture · stream events · alerts always on Inference host YOLO · clips storage sometimes on frames detections host unreachable → the Pi falls back to running the models itself

The fallback is the important part. The host is a desktop that sleeps and reboots; a monitoring system that quietly stops watching whenever that happens would be worse than one that never offloaded at all. So the Pi keeps its own copy of the models, notices when the host stops answering, and carries on alone until it comes back.


Measured, not estimated

Numbers from the hardware it runs on

2.58
inferences per second on the Pi alone, NCNN, one camera
65%
higher throughput from NCNN over ONNX at three cameras
26°C
cooler once inference moved off the Pi, 82°C down to 56°C
204
automated tests, including guards on the privacy rules

A deliberate limit

It does not know who you are

The system answers is someone there, never who is there. That was a design decision taken at the start, and it is enforced by tests that fail the build rather than by good intentions.

One rule was later relaxed by the owner, on their own hardware: the system may group repeat sightings so it stops re-alerting about the same visitor. It does that by comparing clothing colour, it forgets within hours, and it still has no idea whose clothes they are.

blockNo face detection or face matching
blockNo biometric identification libraries
blockNo names, and no link to any outside record
blockNo footage leaving the private network

checkSnapshots and clips expire on a schedule
checkDashboard bound to localhost and a private tailnet only
checkCameras can be switched off from a chat command


Not yet available

Hear about the public beta

SentinelVision runs on one household's hardware today. There is no download, no signup and no product to buy yet. Leave an email and you'll get a single message when there is something worth trying, nothing else.

This is a working system with known limitations, not a finished product. A beta would ship with the same false positives described above.

You're on the list. You'll get one email when the beta opens.
That didn't send. Email support@sentinelgrid.us instead.

One email, no newsletter, no reselling your details. See our privacy policy.


What it is not

The limitations, stated plainly

Small models on cheap hardware get things wrong, and pretending otherwise would make the alerts useless. Early on it classified a gaming chair as a person and a head of curly hair as a grenade. Both were fixed by measuring the failure and changing the design, not by hoping.

False positives happen

Treat a weapon alert as “open the clip and look”, never as confirmation. Detections must persist across frames before anything is sent.

It is not fast

A couple of detection passes per second. Something crossing the frame quickly can pass between them unseen.

Lighting matters

Coloured lighting and near-darkness both degrade it badly. The models were trained on ordinary daylight.