Security Posture¶
This is a homelab on a private network, and several deliberate shortcuts follow from that. They are listed here so the assumptions are explicit rather than implied — someone reading the manifests should be able to tell a decision from an oversight.
That distinction is the whole reason this page exists. Every system carries weaknesses; the dangerous ones are the weaknesses nobody chose.
Assumed trust boundary¶
The cluster assumes a trusted L2 network segment. Anything with a port on that segment is treated as friendly. There is no VPN requirement, no mutual TLS between components, and no network segmentation inside the cluster.
Worth being clear-eyed about what "trusted" means in a house: it includes the guest laptop, the smart TV, and the doorbell running firmware from 2019 that nobody has thought about since. The boundary is real, it is just not as tidy as the phrase suggests.
The published hostnames are a partial exception. Certificates are issued by
Let's Encrypt through a DNS-01 challenge against a public zone, so
argo.infra.k8s.wlkr.ch and its siblings are publicly resolvable names and
appear in Certificate Transparency logs, even though they point at RFC1918
addresses that are unreachable from outside the LAN. Your internal hostnames are
public knowledge the moment you request a certificate for them; plan names
accordingly.
Provisioning¶
Provisioning is the least protected phase, by design — it has to work before any of the cluster's own security exists. This is the classic bootstrap problem, and everyone solves it the same way: briefly, and with the door open.
| Property | Detail |
|---|---|
| Ignition configs are served unauthenticated over HTTP | Anything on the segment can fetch http://<boot-server>:8000/ignition-<host>.json while the boot server is running |
| Those configs embed join credentials | The inlined kubeadm config carries the bootstrap token and the certificateKey, which together are enough to join a new control-plane node |
Nodes join with --discovery-token-unsafe-skip-ca-verification |
A joining node does not verify the API server's CA |
Sysext transfers set Verify=false |
System extension images are fetched over HTTPS but their signatures are not checked |
The OS image is the exception, and deliberately so. flatcar-install is given
-b/-V rather than a local file, so the node downloads the image and its
detached signature from the boot server and checks both against Flatcar's
signing key before writing anything. Serving it over plain HTTP on a segment
this page calls only conditionally trusted is fine precisely because a
substituted image fails the signature check. It is the one artifact on that
server that becomes the operating system, which is why it gets the treatment the
sysexts still do not.
The practical mitigation is time: make serve is a foreground command, the
bootstrap token has a 24 hour TTL, and the uploaded certificate key expires
after two hours. Stop the boot server when provisioning is finished — it is
the only thing keeping those credentials off the network. A make serve left
running in a forgotten tmux session for three months is a genuinely bad outcome,
and it is an easy one to reach.
Stopping it is now free
Flatcar is installed to disk, so a running node reboots, updates and rejoins with the boot server switched off — the exposure above exists only during a build. That was not true when the nodes ran from RAM and PXE-booted on every restart, which made "stop the boot server" and "let Kured reboot a node at 02:00" mutually exclusive instructions. See Boot & Bootstrap Process.
output/credentials/ holds the generated bootstrap token and certificate key in
plaintext. The directory is 0700 and output/ is gitignored, but the values
are reused across make config runs — the Ansible password lookup reads back
an existing file rather than regenerating. Delete them to force new ones.
Secrets¶
Secrets live in OpenBao and reach workloads as native
Kubernetes Secret objects through the
External Secrets Operator. Nothing sensitive
is committed to Git.
Two consequences worth knowing:
- OpenBao is sealed with Shamir and unsealed by hand, so the 5 key shares are the root of trust for every other secret and the only thing that brings the store back after a restart. They exist only wherever the operator put them; losing all of them loses everything, and nothing outside the cluster holds a copy. The cost is a manual step after every reboot — see the resulting limitation.
- A Kubernetes
Secretis base64, not encryption. Anyone withget secretsin a namespace can read what ESO materialised there. Encryption at rest, below, does nothing about this — it protects the bytes in etcd, not the API. If you remember one thing from this page, make it this one; the number of people who believe otherwise is remarkable.
Encryption at rest¶
The API server is configured with an EncryptionConfiguration that encrypts
secrets with secretbox before they reach etcd
(ansible/templates/kubeadm.yaml.j2, and the key file in
ansible/templates/butane_node_config.yaml.j2). Without it a Secret sits in the etcd
data directory as plaintext, so an etcd backup, a stolen disk, or read access to
/var/lib/etcd yields every credential the cluster holds. strings on an
unencrypted etcd file is a memorable demonstration, and one worth doing exactly
once, on a cluster you do not care about.
| Property | Detail |
|---|---|
| Provider | secretbox, with identity listed after it |
| Key | 32 random bytes, generated once by make config into output/credentials/encryption_key |
| Scope | secrets only; ConfigMaps and other resources are unencrypted |
| Distribution | The same key on every control-plane node, written by Ignition to /etc/kubernetes/enc/encryption-config.yaml (mode 0600) |
Two things follow from identity being listed last. New writes are encrypted,
and Secrets written before this was enabled stay readable — they are not
rewritten automatically. To encrypt what already exists, rewrite every Secret
in place once the API servers have restarted:
The key is a single static key with no rotation, and it lives beside the
kubeadm token and certificate key in output/credentials/. That directory is
now the thing to protect: it holds the material that decrypts etcd. A KMS
provider would remove the static key, at the cost of a dependency the cluster
must reach before it can serve Secrets.
Audit logging¶
The API server records who did what, to which object, and whether it was
allowed — and the audit half of the Pod Security Admission labels has nowhere
to go without it. The policy, the retention, and the LogQL to query it are in
Audit Logging.
Authorization¶
ArgoCD AppProjects do not constrain much. payload/argocd/argocd-projects.yaml
defines three projects, but apps and infra both allow sourceRepos: "*" and
a clusterResourceWhitelist of every group and kind, in every namespace. Only
system restricts its destination namespace.
They are useful as grouping and as a place to add restrictions later. They are
not an isolation boundary today: an Application in the apps project can create
cluster-scoped RBAC. Which is to say, a workload's manifest directory can quietly
grant itself the keys to the cluster, and nothing would object.
Network policy covers eight namespaces. openbao, cert-manager,
external-secrets, monitoring, external-dns, kubelet-csr-approver,
kured and logging have default-deny ingress
CiliumNetworkPolicy rules; every other namespace, and all egress everywhere,
is still unrestricted. See Security Policies.
Pod Security Admission is on, but mostly auditing. Every platform namespace
carries enforce at the level it demonstrably needs and warn/audit at a
stricter one, so violations are visible without breaking what runs today. This is
the sane order of operations: measure first, enforce second. Enforcing first is
how you end up disabling the control entirely at 2am. The audit half of that
now has a destination — see Audit logging.
Both Gateways admit routes from every namespace (allowedRoutes.namespaces.from: All).
Any namespace can attach an HTTPRoute to infra-gateway and claim a hostname
under *.infra.k8s.wlkr.ch.
Exposed interfaces¶
Every platform UI on the infra gateway is behind Authentik, by one of two routes:
| Service | Authentication |
|---|---|
| ArgoCD | Authentik OIDC; local admin disabled. The server runs with --insecure because TLS terminates at the Gateway |
| Grafana | Authentik OIDC; login form disabled |
| Hubble UI | Authentik proxy outpost |
| Rook dashboard | Authentik proxy outpost |
| Prometheus | Authentik proxy outpost |
| Alertmanager | Authentik proxy outpost |
| OpenBao UI | Token or configured auth method — not behind Authentik |
Two things follow. Authentik is now a dependency of reaching any of them, so the break-glass paths in When Authentik is down matter. And OpenBao is deliberately left out: putting the thing that holds Authentik's own database password behind Authentik would be a loop, and circular dependencies in an auth stack are only funny from a distance.
What would tighten this up¶
Roughly in order of value against effort:
- Extend default-deny ingress to the namespaces not yet covered, then start on egress — the larger and more breakable half.
- Narrow
sourceReposon the AppProjects to this repository and the Helm repositories actually in use. - Restrict
allowedRoutesoninfra-gatewayto the platform namespaces. - Add a second Alertmanager receiver on a different transport. One receiver, one mailbox and one SMTP provider means a failure of the mail path is itself unmonitored — see Alerting reaches one mailbox.
- Move etcd encryption to a KMS provider, removing the static
encryption_keythat currently sits inoutput/credentials/with no rotation. - Get the backups out of the cluster. Velero and the etcd snapshots write to the Ceph object store they are backing up — see Backups do not leave the cluster. With etcd now persisting across reboots there is more worth losing than there used to be.
See Known Limitations for the operational counterparts.