Why we built our own uptime monitor instead of paying for one

Off-the-shelf uptime tools check that a page returns 200. Ours checks that the page is actually right. Here's why that difference matters.

uptime-monitoring cloudflare-workers status-pages msp

A 200 response doesn't mean your site works. We learned this the annoying way: a page can return a perfectly healthy status code while serving an error template, an empty shell, or last month's "under construction" placeholder. Every commodity uptime checker we tried would have called that "up."

So SentinelGrid Monitor, the thing behind status.sentinelgrid.us, asserts content, not just status codes.

What a real check looks like

For every service we watch, the monitor can require three things:

  • The HTTP status we expect, which is not always 200. One of our client APIs is checked by asserting it returns a 401 with the exact "Not signed in" body. That single check proves DNS, routing, TLS, the worker, and the auth layer are all alive. A 200 there would actually be the bug.
  • A string that must appear in the body. A car shop's website is only "up" if the response actually contains the shop's name.
  • Timing. Slow is a symptom; we record response times on every check, not just failures.

Heartbeats: catching the things that go quietly

HTTP checks catch things that break loudly. The quieter failure mode is a scheduled job that just… stops running. Backups are the classic case: nothing errors, nothing pages, and you find out the night you need the backup.

For those we use dead-man's switches: the job pings the monitor when it finishes, and the monitor alerts when the pings stop. Silence becomes the alarm.

Why Cloudflare Workers

The monitor runs entirely on Cloudflare Workers with a D1 database. No VM to patch, no server that can itself go down and take the monitoring with it, and the free-tier economics mean we can run checks every minute without thinking about cost. A monitoring system you hesitate to add checks to is a monitoring system with blind spots.

There's one design constraint worth mentioning: Workers cap subrequests per invocation, so alert fan-out has to be deliberate. When an incident opens, watchers get one bcc'd email per incident, not one email per watcher. The budget shaped the design, and honestly shaped it better.

Alerts that respect a schedule

Alerts go out over email and Pushover. One detail we're proud of: quiet windows are enforced by the sender, not the phone. Pushover's high-priority messages deliberately ignore the phone's own quiet hours, so "don't make noise during school hours, do wake us up at 2am for a critical" can only be honored server-side. The monitor knows the schedule and drops priority inside the window. Weekends stay loud on purpose.

The takeaway

If you run anything people depend on, monitor the meaning of the response, not the status code, and monitor the silence, not just the noise. If you'd rather someone else carry the pager, that's literally what our Managed IT beta is.

Want this handled for you?

SentinelGrid Managed IT is in open beta: monitoring, remote support, patching and a real helpdesk for one flat monthly fee, plus a free infrastructure audit whether you sign or not.