Node Lifecycle¶
Everything you do to one machine: reboot it, add one, replace a dead one, or rebuild one from scratch. The four procedures share a shape — drain, act, wait for Ceph, move on — and the waiting is the part people skip.
Two of these destroy data and two do not
Rebooting and adding are safe and routine. Replacing destroys one node's OSD. Rebuilding destroys the disk it runs against, and doing it to every node in turn destroys every replica of everything. Read the danger block in that section before you start it.
Rebooting a node¶
Kured reboots nodes on its own once an update has been staged, one at a time, and refuses while Ceph or etcd is unhealthy. The procedure below is for the cases it does not cover: rebooting a node ahead of its next 30-minute check, or rebooting one for a reason nothing set a sentinel for.
A reboot is just a reboot. Flatcar is installed to disk, the node boots from it
without the boot server, and /etc/kubernetes, /var/lib/etcd and
/var/lib/rook are all still there when it comes back. A change made with
make config is not picked up here — that takes a
rebuild.
One node at a time, and both gates at the end are mandatory rather than advisory:
-
Drain it.
-
Reboot it, and wait for it to come back
Ready. -
Uncordon it.
-
Wait for Ceph before touching the next node. Draining a second node while the first is still backfilling can take a placement group below its minimum replica count. Ceph is patient; impatient operators are how "one node down" becomes "read-only cluster".
-
Unseal OpenBao if the node hosted a replica. It came back sealed, and a sealed OpenBao looks exactly like a healthy one — see below.
After any node reboot: unseal OpenBao¶
OpenBao seals itself whenever its pods restart, and while it is sealed no
ExternalSecret resolves — which means cert-manager cannot renew certificates.
Nothing unseals it for you, so this is a chore rather than a check:
make bao-unseal # unseals whatever is sealed
for pod in openbao-0 openbao-1 openbao-2; do
kubectl -n openbao exec "$pod" -- bao status | grep Sealed # false
done
make bao-unseal reads the key shares from output/credentials/openbao-init.json,
checks each replica, and feeds three shares to any that is sealed. It is
idempotent — on an unsealed cluster it reports that and stops.
If you have moved the keys into a password manager and deleted that file, which is where they belong, it cannot help: each pod then wants 3 of the 5 shares by hand. OpenBao → Unsealing after a restart has the loop.
It is worth actually running, every time. A cluster that comes back with OpenBao
still sealed looks entirely healthy — every pod Ready, OpenBao's included — and
the consequence surfaces sixty days later when a certificate expires on a
Sunday, with nothing connecting it to the reboot that caused it. See
the limitation this creates.
Adding a node¶
The bootstrap token generated at provisioning time has a 24 hour TTL, so it has long expired on an established cluster. This trips up everyone exactly once. Generate a fresh join command on a control-plane node:
For a new control-plane node you also need a current certificate key, which expires after two hours:
Add the host to ansible/inventory.yaml, re-run make config and make serve
so it can PXE boot, then run the printed join command on it. The
bootstrap-k8s.service unit only fires when /etc/kubernetes/kubelet.conf is
absent, so it will not interfere with a node that has already joined.
Replacing a failed node¶
-
Remove it from the cluster:
-
Let Ceph re-replicate. With
useAllNodes: truethe OSD on that disk is gone for good; checkceph statusreturns toHEALTH_OKbefore continuing. This is not a step to rush — Ceph will tell you when it is done, and it is never as fast as you would like. -
If it was a control-plane node, remove its etcd member. A dead member left in the list still counts toward quorum, which is a delightful way to lose a cluster that is otherwise fine:
-
Reprovision the replacement following Adding a node.
Danger
On a cluster provisioned before the Control Plane VIP, odin is not an interchangeable control-plane node: its address is baked in as the API endpoint and as Cilium's k8sServiceHost, so losing it breaks node joins and Cilium's API connection on every other node. Check which endpoint your kubeconfig uses before assuming otherwise.
Rebuilding or repartitioning a node¶
The disk layout lives in ansible/templates/butane_node_config.yaml.j2 and is applied
by Ignition, which runs once — on the first boot after a node is installed.
Changing a size or order in it is not an edit you roll out. Neither is
changing anything else under ansible/: the node's Ignition config is embedded
in its OEM partition at install time, so a running node will never see the new
one.
rook-osd is the last partition and is deliberately raw: Ceph owns the bytes,
and there is no filesystem or label inside it that would let anything relocate
them. Move its start offset by so much as a sector — which is what inserting or
resizing any partition above it does — and every OSD on that node is gone.
The backups are inside the thing being wiped
Velero and the etcd snapshot CronJob both write to the Ceph object store this destroys — see Backups do not leave the cluster. Copy anything you intend to restore from off-cluster before you start. This is the failure mode that limitation was written about, arriving in person.
The procedure is a reinstall:
- Copy what matters off the cluster — Velero backups, the latest etcd snapshot, and anything in a PVC that is not reproducible from Git.
- Edit whatever needs editing under
ansible/. make configto regenerate both Ignition configs, thenmake serve.-
Arm the nodes you are rebuilding:
That flips
DEFAULT localboottoDEFAULT installinoutput/tftp/pxelinux.cfg/01-<mac>. The template is untouched, so the nextmake configputs the safe default back — including over anything armed and not used.make reinstall-canceldoes the same deliberately, and the boot server does it for you once the node has the OS image, which is what keeps the reboot at the end of the install from starting a second one: see Boot Server → Switching back to local boot.Arming is the only way in: the menu shows no prompt, so there is nothing to pick at the console and nothing to mistype at three in the morning. 5. Network-boot the node. The installer wipes the disk — every partition signature, the GPT, and a device-level discard where the hardware supports it — runs
flatcar-install, and reboots into the freshly installed system, which then runskubeadm. 6. Leave the boot server up until the node isReady: the sysext images are still fetched from it on that first boot. Then stop it.
One node at a time is safe if you are rebuilding rather than repartitioning — etcd keeps quorum and Ceph backfills, exactly as in Replacing a failed node. Repartitioning is different only in that it destroys every OSD as it goes, so a rolling rebuild across all six nodes eventually destroys all replicas of everything. Step 1 is not optional for that case.
The menu will not do this by accident
The PXE entry that installs has to be chosen. The menu's default is
LOCALBOOT, so a node that reboots while the boot server happens to be
running boots what it already has. The install does not touch the
firmware's boot order, which means the menu decides on every boot, on
every node — see Boot order.
Only the last partition can grow without a reinstall. Shrinking rook-osd to
make room for something else cannot be done in place either, because Ceph has
already written across the space you would be taking back.