Skip to main content

Uptime Kuma

External availability watchdog. Answers "is it up, and did the job run?"

CategoryObservability
Service IDuptime-kuma
Namespacemonitoring
Imagelouislam/uptime-kuma:2.5.0
Port3001
URLhttp://uptime-kuma.localhost
Websitehttps://uptime.kuma.pet

Deploying it also deploys AutoKuma, which turns monitor definition files into monitors. It is part of this service, not a separate thing to install.

⚠️ Run this outside the platform it monitors

A monitor inside a cluster cannot report that cluster being down — which is the one situation you need it for.

./uis deploy uptime-kuma will happily install it onto the cluster you intend to monitor. It will look healthy and be useless when it matters. UIS cannot express "deploy this elsewhere", so this is a convention you honour, not a rule the tooling enforces.

For local development that does not matter and you should ignore it — see the developer path below.

You do not write monitors

This is the part worth understanding, because it is different from how most people use Uptime Kuma.

Both halves of a monitor are already known to UIS:

NeededWhere it comes from
How to probe — path, keyword, which secretships with the service, as a probe artifact
Where the service isdiscovered from the Service and Ingress that uis deploy created

So you deploy services, deploy the watchdog, and run:

uis monitors apply

Nothing asks you for a hostname. You never had to know one to deploy the service, and you should not have to know one to monitor it.

probe artifacts (ship with each service)
+
cluster discovery (Service / Ingress)
│ uis monitors apply

Secret uptime-kuma-monitors
│ AutoKuma reconciles

Uptime Kuma ──alerts──▶ your phone

The three commands

uis monitors render    # what would be monitored. Changes nothing
uis monitors apply # write the definitions, restart AutoKuma, attach alerting
uis monitors check # compare intent against what Kuma is actually running

check is not decoration. AutoKuma reconciles continuously, so drift heals itself — which means a stopped AutoKuma looks exactly like a healthy one. check compares against Kuma itself, not just against the definitions.


Path 1 — a developer on Rancher Desktop

You have one cluster. Uptime Kuma runs in it. That is fine: you are not trying to detect your laptop being off.

uis deploy uptime-kuma
uis monitors apply

That is the whole thing. Open http://uptime-kuma.localhost, log in as admin with the shared DEFAULT_ADMIN_PASSWORD, and the services you have deployed are already listed.

Because the watchdog is inside the cluster, endpoints resolve to in-cluster DNS:

http://litellm.ai.svc.cluster.local:4000/v1/models
postgresql.default.svc.cluster.local:5432

You get more coverage than production does, not less: in-cluster-only services like PostgreSQL are reachable here and are not reachable from an external watchdog.

Deploy another service and re-run uis monitors apply.

Path 2 — an ops engineer running it in production

The watchdog runs on a separate machine — a Pi, a NAS, a small VM — with its own tiny cluster. The services run somewhere else. So discovery has to read one cluster and write to another:

uis monitors apply --from production --to watchdog
--fromwhere the services are. Read-only
--towhere Uptime Kuma runs. Written to

Omit both and you get Path 1 behaviour.

Give discovery a read-only identity

Do not copy an admin kubeconfig onto the watchdog host. Discovery needs three read verbs on three resource types; an admin credential would put full control of your platform on the machine whose entire purpose is to sit outside it.

kubectl --context production apply -f manifests/230-uptime-kuma-discovery-rbac.yaml

That creates a monitor-discovery ServiceAccount which can list services, endpoints and ingresses, and nothing else. Add it to the watchdog host's kubeconfig as the --from context.

Verify it is as powerless as intended:

kubectl --context production get secrets -A    # must be Forbidden

What resolves, and what cannot

From outside the cluster, in order of preference:

Gives
Ingress with ingressClassName: tailscalea real FQDN with a real certificate — the best case
A shim Service whose Endpoints point outside the pod networkthe external address, straight from the Endpoints object
LoadBalancerits external address
ClusterIP onlynothing. Not reachable from another machine

That last row is reported, never silently dropped:

NOT MONITORED (1):
postgresql: only reachable inside the cluster (ClusterIP). An external
watchdog cannot probe it - expose it, give it a shim, or rely on the
services that depend on it failing their own probes

A deployed service that quietly gets no monitor is indistinguishable from one that is passing. Read that list.


Targets UIS did not deploy

A hypervisor, a NAS, a laptop, a printer, a nightly job. UIS cannot discover these because it did not create them, so they go in .uis.extend/monitors.yaml:

defaults:
interval: 60
maxretries: 2
notify: true

monitors:
- name: hypervisor-ui
type: port
hostname: 192.168.1.10
port: 8006
why: If this is down, everything running on it is too.

- name: vault
type: http
url: https://192.168.1.20:8200/v1/sys/health
ignore_tls: true # internal certificate
accepted_statuscodes: ["200-299"] # a SEALED vault answers 503
why: Sealed counts as down - it is useless to consumers even when running.

- name: heartbeat-nightly-backup
type: push
interval: 93600 # 26h - headroom for a late run
why: Catches a backup that silently stopped.

This file is empty or absent on a stock install. If you find yourself adding something UIS did deploy, that is a bug in discovery or a missing shim — and uis monitors will refuse it as a duplicate rather than let two definitions silently overwrite each other.

why: is not decoration. A monitor nobody can justify is a monitor nobody maintains.

Heartbeats — the part nothing else provides

A push monitor gives you a URL. A job calls it after its success check. If the call stops arriving inside the expiry window, you are alerted.

That detects the absence of work, which metric alerting handles badly: nothing crashes, no error rate moves, work simply stops.

# in a shell script under `set -e`, the last line is the success check
curl -fsS "https://kuma.example/api/push/<token>?status=up&msg=ok"
# in a systemd unit - ExecStartPost only runs if ExecStart exited 0
ExecStartPost=/usr/local/sbin/kuma-push.sh <token> "backup ok"
Never push unconditionally

A job that calls its heartbeat regardless of outcome reports success when it fails. That is worse than no heartbeat, because it manufactures confidence.

Tokens are derived, not randomHMAC-SHA256(salt, monitor name) from uptime-kuma-push-salt. A rebuilt watchdog therefore reissues identical URLs and no job needs rewiring. Back that salt up: lose it and every heartbeat caller is silently orphaned, still exiting 0 while nothing records that it ran.

Alerting — getting it onto your phone

This is the one step UIS cannot do for you: an app has to be installed on a device, and only you can do that.

1. Generate a topic

echo "uis-$(head -c 18 /dev/urandom | base64 | tr -dc 'a-z0-9')"
warning
On public ntfy.sh the topic name is the credential

There are no accounts and no passwords — anyone who knows the topic can read your alerts and publish fake ones. Use a generated value, never something guessable like uis-alerts or your company name. Treat it like a password: do not paste it into a ticket, a chat, or a screenshot.

2. Tell UIS about it

UPTIME_KUMA_NTFY_TOPIC=uis-xxxxxxxxxxxxxxxxx

in the secrets template, then uis secrets generate && uis secrets apply and deploy or re-deploy uptime-kuma. Leave it empty and the watchdog still monitors — it just tells nobody, and says so during deploy rather than being quietly silent about it.

3. Install the app and subscribe

ntfy — free, open source, no account:

iOSApp Store — ntfy
AndroidPlay Store, or F-Droid
Desktop / browserhttps://ntfy.sh/<your-topic>

In the app: + → paste the topic → leave the server as ntfy.sh → Subscribe. There is no sign-up step; subscribing to the topic is the whole configuration.

4. Prove it arrives — before you rely on it

curl -H "Title: test" -d "if you can read this, alerting works" \
https://ntfy.sh/<your-topic>

If that does not reach the phone, alerting does not work, and you will not find out later — a watchdog with a broken channel is indistinguishable from a quiet week. Check it now, and re-check after changing the topic.

Then test the real path, which is not the same thing: point a throwaway monitor at a closed port and confirm Uptime Kuma itself pushes.

Things that will bite you

Android battery optimisation can delay or drop notifications. Exclude ntfy from it, or the alerts arrive an hour late — which for an outage is the same as not arriving.

Alerts are sent at priority 5 so they break through Do Not Disturb. That is deliberate: an availability alert you sleep through has not done its job. If that is too aggressive, lower it in the notification's settings rather than muting the app.

iOS and self-hosting do not mix freely. The topic-is-the-credential problem above makes self-hosting tempting, but iOS push has to travel via Apple's APNs, which the ntfy iOS app reaches through ntfy.sh's infrastructure. A self-hosted server needs upstream-base-url configured before iOS delivery works at all. Android has no such constraint. Check this before migrating, not after.

One phone is one point of failure. If it is off, lost, or with someone on a plane, nobody is alerted. Subscribe a second device, or a second person.

Every monitor pages by default. Set notify: false on anything whose failure is not both real and actionable:

  • a laptop or workstation that sleeps — it will page every nap
  • a heartbeat whose job does not exist yet — it pages once and stays down

An alert you learn to swipe away is worse than no alert: it trains you to ignore the real one.

Adding a probe to a service

For service authors. Create provision-host/uis/services/<category>/probes/<id>.yaml:

# Only what is true of EVERY install of this service.
# No hostnames - those are discovered. No secret values - name the key.
service: my-service-web # only if the k8s Service name differs from the id
probes:
- id: gateway
type: http # http | tcp
path: /healthz
keyword: ready # match content, not just status
auth: bearer:MY_API_KEY # a KEY in urbalurba-secrets, never a value
interval: 60
maxretries: 2

⚠️ service: matters. The probe is matched to a Kubernetes Service by name. If your service id is temporal but the Service is temporal-web, omitting this means the probe matches nothing and the service goes unmonitored with no error.

⚠️ Prefer a keyword to a status code. A gateway can return 200 while its database is dead, because the response never touched the database.

Deploy, verify, remove

uis deploy uptime-kuma            # Kuma + AutoKuma + admin + retention + alerts
uis verify uptime-kuma
uis undeploy uptime-kuma # keeps monitors, history, notification config
uis undeploy uptime-kuma --purge # deletes all of it

No first-run wizard. UPTIME_KUMA_DB_TYPE=sqlite skips the database screen, and the playbook seeds the admin account — needSetup is just "is the user table empty?". Seeding is idempotent, so redeploying never clobbers a password you changed in the UI.

Use --purge when you want to prove an install works. Keeping the volume means every reinstall after the first lands on existing state, so a broken first-install path stays hidden.

Storage and secrets

Embedded SQLite on a ReadWriteOnce volume, StatefulSet, one writer. History retention defaults to 30 days rather than Kuma's 180, because a watchdog usually runs on flash storage.

On a Pi or similar, put the volume somewhere other than the boot card — not for speed, but so wearing it out costs a replaceable device rather than a rebuild.

Key in urbalurba-secrets
uptime-kuma-admin-useradmin
uptime-kuma-admin-passwordinherits ${DEFAULT_ADMIN_PASSWORD}
uptime-kuma-push-saltderives heartbeat URLs — back this up
uptime-kuma-ntfy-server / -topicthe alert channel

It shares the platform admin password rather than defining its own, so there is one credential to rotate.

warning

Three things write to Kuma's own database, which is a private interface rather than a published API: the admin seed, retention, and the alert channel with its attachments. Kuma has no API for any of them. Each is guarded — the schema is checked first, writes are insert-only, and each is verified afterwards rather than assumed. Monitors are not in that list; AutoKuma owns those.

Who watches the watchdog

If Uptime Kuma dies you get silence, and silence looks exactly like health.

The arrangement that works: a second machine polls the watchdog and, if it cannot reach it, alerts your phone directly — going through Kuma would be pointless when Kuma is what is down. Latch the state so you get one alert, not a storm, plus one on recovery.

⚠️ That still does not cover everything. Two machines in the same building share a power cut and an internet outage. Full coverage needs something off-site: a free dead-man's-switch service, or a cron job somewhere you do not own.

Decide where you put that boundary deliberately, rather than discovering it during an incident.

When something is wrong

The failure modes here are quiet ones, so check in this order:

SymptomLikely cause
apply succeeded, no new monitorsAutoKuma has not re-read the definitions. apply restarts it; a manual Secret edit does not
Monitor count doubled after a restart/data is not persisted — it holds AutoKuma's id map
A monitor is DOWN with 401the probe's auth: key is missing from urbalurba-secrets
A deployed service has no monitorread the NOT MONITORED list from uis monitors render
check reports drift that is not realwrong --from/--to contexts
Everything green, nothing ever alertsno channel configured, or notify: false