Statuspy

About

How Statuspy decides whether something is down

Statuspy checks 127 online services from our own infrastructure and publishes what we measured, with the time we measured it. It is run by one person, independently of every service listed here.

What a verdict actually claims

Every couple of minutes we send one unauthenticated HTTP request to a chosen endpoint for each service and record the status code and how long it took. That is the whole measurement, and being precise about its limits is the point: “up” means one request from one machine succeeded in under five seconds. It does not mean every feature works, or that the service works from where you are.

Up
The endpoint answered normally in under five seconds.
Degraded
It answered, but took five to ten seconds. Reachable and slow.
Down
Repeated checks failed. Never published on the strength of a single failure.
Unverifiable
We could not get a usable reading, so we say so rather than guess. A service that blocks automated requests lands here, and a block is never reported as an outage.

Why we are slow to say “down”

A wrong outage report is the one mistake that would make this site not worth reading, so publishing one requires corroboration. A single failed check is treated as a candidate, not a verdict.

  • Three consecutive failed checks are needed before the public status becomes down, which takes about six minutes at our normal cadence.
  • One failed check plus the service’s own status page reporting a major or critical incident is also enough, because that is independent confirmation. Those incidents carry a shield on the badge.
  • The recheck button on each page cannot create an outage. Only our scheduled checks count toward that rule, so activity on a page can never talk the verdict into changing.

Where the numbers and the history come from

Two sources, kept visibly separate, because they are different kinds of claim.

  • Our own checks produce every response time, uptime percentage, and daily trace on the site. Those are measurements, and they are set in a monospace typeface throughout to mark them as such.
  • Official status pages for the 82 services that publish a machine-readable one. We poll them to corroborate our own findings, to fall back on when a service blocks us, and to import the incident history from before we were watching. Anything sourced that way is labelled as reported by the vendor, never presented as something we measured.

Days we could not measure are hatched rather than left blank or coloured in, because a gap must not look like a healthy day:

  • up
  • slow
  • down
  • no reading

What this cannot tell you

  • We check from one location. A regional outage that does not affect our vantage point will read as healthy here.
  • We check one endpoint per service. A broken feature behind a login can fail while the endpoint we watch answers perfectly.
  • History for most services begins when we started monitoring. Anything older comes from the vendor’s own status page and is labelled as such.
  • The service’s own status page remains the authority on its status. Where one exists we link to it on that service’s page.

Corrections

If a verdict here is wrong, that is a bug and worth reporting. Issues can be raised on the project’s issue tracker. Being told we published something false is more useful than being told nothing.