Skip to content

The home lab

A platform engineer's claims are cheap without somewhere to make them break. This is the lab behind the talks and the writing: a small cluster that has been rebuilt twice, failed in instructive ways, and is run the way the production systems I talk about are run, with every change through Git.

Hardware

MachinesEight Beelink mini PCs
ProcessorsIntel N100 and Celeron N5105, four cores each
Memory16 GB per machine
NetworkTwo wired interfaces per machine: one for applications, one for storage
Storage hostsTwo Ubuntu servers, 9.1 TB and 7.4 TB, serving NFS and Samba
EdgeA pfSense firewall, with Pi-hole DNS and ISC DHCP behind it

Two generations

To spring 2025

Generation one: virtualization

Eight Proxmox VE 8.4 hosts in one cluster, with Kubernetes running as virtual machines on top of them.

Platform
Proxmox VE 8.4, eight hosts, one cluster
Storage
Ceph Quincy on its own network, over the second interface
Kubernetes
Three control plane and several worker VMs spread across the hosts
Change control
Ansible playbooks for the hosts, Ceph and the VMs; nothing edited by hand
Templates
VM templates exported and restored between hosts through an NFS share

It taught the hard parts of running a hypervisor cluster on small machines, and it put a layer between Kubernetes and the metal that the second generation removed.

Mid 2025 to April 2026

Generation two: Kubernetes on the metal

Kubernetes 1.33 built with kubeadm directly on the machines: three control plane nodes and six workers, one of them added for USB-attached storage.

Platform
Kubernetes 1.33 by kubeadm, three control plane nodes and six workers
Networking
Cilium in place of kube-proxy, kube-vip for a highly available API, MetalLB for load balancer addresses
Storage
Longhorn, with Velero for backups
Change control
ArgoCD GitOps: every production change through Git, kubectl for emergencies only and followed by a pull request, Sealed Secrets for anything secret
Observability
Prometheus, Grafana, Loki, OpenTelemetry and Jaeger
Security
Falco, the Trivy operator, cert-manager, and a scored hardening program in phases
Also running
Harbor, KubeVirt, Kata Containers, Authentik, MinIO, PostgreSQL, ingress-nginx and self-hosted GitHub Actions runners

Removing the hypervisor made a node failure a machine failure, which is the failure mode the talks are about.

The network

Each machine has two wired interfaces. The first carries applications and the cluster's own traffic, behind the firewall, with DNS and DHCP on the same segment. The second carries storage traffic between the nodes on a separate network, so a storage rebuild cannot starve the applications.

The lab's two networksNetwork diagram: a firewall at the edge; below it an application network joining all nodes, the DNS and DHCP host, and the two storage hosts; a separate storage network joining the nodes to each other over their second interface.FirewallApplication network, first interfaceControl plane, 3Workers, 6DNS and DHCPStorage hosts, 2Storage network, second interfaceCeph in generation one, Longhorn replication in generation two

What it taught

  1. Test a kernel on one node first

    A kernel update took two control plane nodes down together. The first investigation blamed the kernel, correctly, and missed that one of the two machines also had failing hardware underneath, which kept it down after the kernel was fixed. The lab now takes a new kernel on one node, the one with the different processor, before the rest.

  2. A DHCP pool that overlaps static addresses is an outage waiting

    A virtual machine was assigned an address at the start of the DHCP pool, and unidentified devices turned up inside it on a scan. The fix was boring and worth writing down: a reserved range for infrastructure, a reserved range for services, and the pool kept out of both.

  3. Score security and raise it in phases

    The hardening work was scored before and after each phase rather than described. Phase two, which brought in Sealed Secrets, Falco and the Trivy operator, moved the score from 35 to 48 out of 100 in an afternoon, and the number made the next phase's priorities an argument about evidence instead of taste.