Problem Statement
Enterprise Kubernetes fleets drift. A node provisioned in January and one provisioned in June rarely carry identical kernel patches, container runtime versions, or CIS benchmark hardening — even when both were “built from the same base image.” That drift is where outages and failed audits come from: an incident response team re-imaging a node under pressure, using whatever AMI/image happened to be default that day, reintroducing a CVE that was patched out three months earlier.
The fix is treating the node image itself as a build artifact — versioned, scanned, and rebuilt on a schedule — rather than a long-lived, hand-patched machine.
Architecture Blueprint
- Define the image declaratively. A HashiCorp Packer template describes the base OS, kernel parameters, container runtime, and CIS hardening scripts in one versioned file.
- Build in CI, not by hand. Every merge to the image repository triggers a Packer build, producing a new, immutable AMI/image tagged with the commit SHA.
- Scan before promotion. The freshly built image is scanned (e.g., Trivy/Grype) for known CVEs; a failing scan blocks promotion to the “approved” image pool.
- Reference by immutable image ID, not by
latest. Terraform modules that provision node groups pin the exact AMI ID produced by step 2 — never a mutable tag. (This is an AMI ID, not a container digest — the two are distinct identifiers even though both serve the same “pin the exact artifact” purpose.) - Roll, don’t patch. When a new golden image is approved, roll it out via node-group replacement (surge + drain), not in-place patching of running nodes.
Code Example
source "amazon-ebs" "golden" {
ami_name = "bpmbi-eks-node-${var.image_version}"
instance_type = "m6i.large"
source_ami_filter {
filters = {
name = "amazon-eks-node-al2023-x86_64-standard-*"
}
owners = ["amazon"]
most_recent = true
}
}
build {
sources = ["source.amazon-ebs.golden"]
provisioner "shell" {
script = "scripts/cis-harden.sh"
}
provisioner "shell" {
inline = [
"curl -sfL https://raw.githubusercontent.com/aquasecurity/trivy/main/contrib/install.sh | sh -s -- -b /usr/local/bin v0.56.2",
"/usr/local/bin/trivy rootfs --exit-code 1 --severity CRITICAL,HIGH /"
]
}
}This is an illustrative excerpt, not a complete reproducible project: it omits the
variable "image_version" {} declaration, the contents of cis-harden.sh, and the
ssh_username the Amazon EBS builder requires to connect to the build instance. The scan
step installs a specific, pinned Trivy release to /usr/local/bin (the upstream installer
defaults to ./bin, which usually isn’t on PATH) and runs trivy rootfs against the
instance’s root filesystem (/) — scanning the provisioned node’s actual filesystem before
it’s imaged, rather than treating the current directory as a container image reference. Pin
whatever Trivy version your team has actually tested, not necessarily the one above.
Get the full checklist
Want the full checklist — including the CIS hardening script list and the Terraform node-group rollout pattern we use in production? Get the Kubernetes Golden Image Checklist or reach out directly at [email protected]. You can also find us at BPMBI @ The Clubhouse, Augusta, GA.