Uptime and monitoring

Uptime is a number that is very easy to quote and surprisingly hard to define. This hub is about what it really means.

"99.9% uptime" sounds like a guarantee. It is roughly forty-three minutes of permitted downtime a month, measured however the vendor chooses to measure it, with a remedy of a partial service credit you have to ask for.

These guides cover the arithmetic, the contractual reality, how status pages are built and why they lag, and how to set up monitoring that tells you something before your customers do.

If you are not sure where to start, what uptime percentages mean is the one most people need first, and service level agreements explained follows on from it directly. If you would rather just find out whether a particular site is working, check it from our servers or look at what is failing right now.

All 7 guides in uptime and monitoring

What uptime percentages mean

The uptime nines table, what each level costs to achieve, and why how uptime is measured matters more than the number itself.

4 min readUpdated 2026-09-19

Service level agreements explained

What service level agreements really commit a provider to, how credits are calculated and claimed, and the exclusions that matter most.

4 min readUpdated 2026-09-19

How status pages work

Why official status pages lag user reports, how they are built, and how to read one properly - Including what a green banner does not mean.

4 min readUpdated 2026-09-19

Monitoring your own site

How to set up uptime monitoring that actually catches problems: What to check, from where, how often, and what to alert on.

4 min readUpdated 2026-09-19

Synthetic vs real user monitoring

The difference between synthetic checks and real user monitoring, what each catches, and why serious operations run both.

4 min readUpdated 2026-09-19

Incident metrics that matter

What MTTR, MTBF, MTTD and MTTA actually measure, why means are misleading for incident data, and which metrics change behaviour.

4 min readUpdated 2026-09-19

Alert fatigue

Why too many alerts is more dangerous than too few, and how to design alerting that people still respond to after six months.

4 min readUpdated 2026-09-19