All incidents
resolvedSep 9, 20267 min read

Cluster-wide shutdown, no power failure

At 03:05 every node in the cluster powered itself off and stayed off for three and a half hours. Mains power never faltered. The UPS monitor talked itself into a shutdown using a flag it had been holding for a full day.

Date
Sep 9, 2026
Duration
3h 39m
Status
Resolved
Impact
Full outage — every self-hosted service
Proxmox clusterKubernetesUPS / NUTNetwork

At 03:05 on 9 September 2026, all three Proxmox nodes shut themselves down, cleanly and within three seconds of each other. Nothing came back until the rack was powered on by hand at 06:45 — three hours and thirty-nine minutes later.

The obvious explanation was a power cut. It wasn’t. Utility power never dropped, the UPS never went to battery, and the battery was sitting at 100 % the whole time. What actually happened is that a routine firmware update on a network switch made the UPS unreachable for about two minutes, and the UPS monitor — still holding a stale flag from the previous night’s self-test — concluded the UPS was dead and ran the emergency power-fail shutdown.

Impact

  • Everything. Three Proxmox nodes, eight VMs, two containers, the whole Kubernetes cluster and every service on it were down for 3h 39m.
  • No data loss. The shutdown was orderly — guests stopped through the normal path, the cluster cordoned itself, filesystems synced. Nothing came back dirty.
  • No self-recovery. A power-fail shutdown is a power-off. The machines sat waiting to be switched on; nothing was going to bring them back automatically.
  • No alert of any kind. A cluster-wide outage produced not one notification.

Timeline

TimeEvent
03:03:43NIC link drops on every node; bonds drop their slaves
03:04:53–55Link returns after ~70 s; corosync re-forms, quorum restored at 03:05:00
03:05:39UPS poll fails → assuming dead → Executing automatic power-fail shutdown
03:05:42The same decision, reached independently, on the other two nodes
03:05:47Guests stopped via the normal shutdown path, 180 s timeout each
03:05:51Kubernetes cordons its own nodes as they go down
03:07:33Filesystems synced, journals closed, machines powered off
03:09The second switch completes its own firmware reboot
06:45:18–30All three nodes powered back on by hand; same kernel as before
06:58Cluster fully reconverged; transient startup errors cleared on their own

Root cause

Five links, each one verified independently. The third is the one that made an ordinary network blip fatal.

1. Firmware auto-update is scheduled for 03:00

The network controller had automatic firmware updates enabled with the update hour set to 3 AM. Both switches took a new firmware build and rebooted, a few minutes apart.

2. The nodes lost their link for about seventy seconds

Every node logged NIC Link is Down at 03:03:43 and Link is Up around 03:04:54. Cluster communication recovered on its own. But the UPS lives on a different VLAN and is reached through a routed hop, and that path stayed broken past link-up.

3. The monitor was still flagged “calibrating” — from the night before

This is the whole incident in one detail. The UPS runs a weekly self-test. It ran at 03:08:56 the previous night, logging on battery and calibration in progress, then on line power ten seconds later.

It never sent calibration finished. So upsmon kept the calibrating flag set for a full twenty-four hours, with the UPS sitting perfectly healthy on line power the entire time.

4. NUT’s rule for that state is to assume the worst

Unreachable, plus last known to be calibrating, equals dead UPS — and a dead UPS during what might be a power failure means shut down immediately. With DEADTIME at its stock 15 seconds, that verdict took three failed polls.

Had the last known state simply been OL (on line), losing contact with the UPS would not have triggered anything at all.

03:08:56 (Sep 8)  UPS ups.homelabnj.tech on battery
03:08:56 (Sep 8)  UPS ups.homelabnj.tech: calibration in progress
03:09:06 (Sep 8)  UPS ups.homelabnj.tech on line power
                  ... 24 hours pass, the CAL flag is never cleared ...
03:05:39 (Sep 9)  Communications with UPS ups.homelabnj.tech lost
03:05:39 (Sep 9)  UPS [...] was last known to be calibrating and currently is
                  not communicating, assuming dead
03:05:39 (Sep 9)  Executing automatic power-fail shutdown

5. The shutdown itself worked perfectly

Which is the frustrating part. Every piece of the emergency path did its job: guests stopped gracefully, the cluster drained itself, filesystems synced, the journals close normally. The mechanism is sound. Only the trigger was wrong.

How I know the power never went out

Three independent witnesses, any one of which settles it.

Two devices on that same UPS never restarted

The gateway’s uptime is unbroken since 27 August, and the UPS itself has been up since 5 September. If mains had failed for three and a half hours, a twenty-minute battery would not have carried either of them through it. Only the two switches rebooted — and they rebooted into a new firmware version, which is its own confession.

The battery is full

Read straight off the UPS thirteen minutes after the nodes came back: status OL, charge 100 %, runtime 1235 s, input 118.2 V, load 34 %. A UPS that had carried a rack and then sat dead for hours does not report a full charge and twenty minutes of runtime.

There is no “on battery” line for that morning

The only transfer-to-battery event anywhere in the journal is the self-test the night before. A genuine power failure would have been logged the same way — and would have produced an ordinary low-battery shutdown, not an “assuming dead” one.

Recovery

Everything came back on its own once the machines were powered on. The cluster went quorate, all eight VMs and both containers started, secrets unsealed automatically, and 33 of 34 GitOps applications returned to healthy without intervention.

The one exception was Home Assistant, which came up in a crash loop — and that turned out to be an unrelated bug the restart merely exposed. Its deployment tracked the mutable :stable image tag, which means the effective pull policy is IfNotPresent and the version that starts is whichever image the scheduled node happened to have cached. It landed on a node whose cache predated the build that wrote its config, and the older version refused to read newer config files. That failure was latent for every reschedule; this restart just rolled the dice.

What changed

ChangeWhy
DEADTIME 15 s → 180 s The single change that would have prevented this. A UPS reached over a routed network hop needs a grace period longer than a switch’s firmware reboot. In a real outage this value is never consulted — the monitor acts on the low-battery signal directly — so widening it costs nothing against a ~20 minute battery and a ~2 minute shutdown.
Pin the Home Assistant image An explicit version instead of a floating tag, so version changes are deliberate commits and rescheduling is deterministic.
Move the firmware update hour Unattended 3 AM reboots of core network gear are a poor default when something downstream treats a network blip as a power emergency.

Still open

  • The UPS has still never been tested on battery. This proved the shutdown mechanism works and proved the trigger was too tight. It did not test what happens when mains power actually fails.
  • Shutdown ordering with the NAS is unverified. It backs every network volume; if it stops before the VMs do, guests hang instead of shutting down.
  • Alerting. A cluster-wide shutdown generated no notification whatsoever.

What I’d take from it

  • A monitor’s memory is state, and state goes stale. The UPS was healthy for twenty-four hours while the thing watching it believed otherwise. Nothing surfaced that disagreement, because nothing was looking for it.
  • Protective systems fail toward action. Anything that can shut down your infrastructure to save it will, given an ambiguous signal, do exactly that. The question to ask of any such system is not “will it fire when it should” but “what else looks like the thing it fires on.”
  • Uptime is a witness. The fastest way to bound the blast radius of any infrastructure event is to ask every device when it last booted. Whatever didn’t reboot tells you what didn’t happen.
  • Mutable image tags are a downgrade waiting for a reschedule. With a cached image and a floating tag, the version you get is a property of the node, not of your repository.
All incidents Home Lab NJ · homelabnj.tech