Skip to content

The Kubernetes Golden Image Checklist

By BPMBI Engineering

Problem Statement

Enterprise Kubernetes fleets drift. A node provisioned in January and one provisioned in June rarely carry identical kernel patches, container runtime versions, or CIS benchmark hardening — even when both were “built from the same base image.” That drift is where outages and failed audits come from: an incident response team re-imaging a node under pressure, using whatever AMI/image happened to be default that day, reintroducing a CVE that was patched out three months earlier.

The fix is treating the node image itself as a build artifact — versioned, scanned, and rebuilt on a schedule — rather than a long-lived, hand-patched machine.

Architecture Blueprint

  1. Define the image declaratively. A HashiCorp Packer template describes the base OS, kernel parameters, container runtime, and CIS hardening scripts in one versioned file.
  2. Build in CI, not by hand. Every merge to the image repository triggers a Packer build, producing a new, immutable AMI/image tagged with the commit SHA.
  3. Scan before promotion. The freshly built image is scanned (e.g., Trivy/Grype) for known CVEs; a failing scan blocks promotion to the “approved” image pool.
  4. Reference by immutable image ID, not by latest. Terraform modules that provision node groups pin the exact AMI ID produced by step 2 — never a mutable tag. (This is an AMI ID, not a container digest — the two are distinct identifiers even though both serve the same “pin the exact artifact” purpose.)
  5. Roll, don’t patch. When a new golden image is approved, roll it out via node-group replacement (surge + drain), not in-place patching of running nodes.

Code Example

golden-image.pkr.hcl
source "amazon-ebs" "golden" {
ami_name      = "bpmbi-eks-node-${var.image_version}"
instance_type = "m6i.large"
source_ami_filter {
  filters = {
    name = "amazon-eks-node-al2023-x86_64-standard-*"
  }
  owners      = ["amazon"]
  most_recent = true
}
}

build {
sources = ["source.amazon-ebs.golden"]

provisioner "shell" {
  script = "scripts/cis-harden.sh"
}

provisioner "shell" {
  inline = [
    "curl -sfL https://raw.githubusercontent.com/aquasecurity/trivy/main/contrib/install.sh | sh -s -- -b /usr/local/bin v0.56.2",
    "/usr/local/bin/trivy rootfs --exit-code 1 --severity CRITICAL,HIGH /"
  ]
}
}

This is an illustrative excerpt, not a complete reproducible project: it omits the variable "image_version" {} declaration, the contents of cis-harden.sh, and the ssh_username the Amazon EBS builder requires to connect to the build instance. The scan step installs a specific, pinned Trivy release to /usr/local/bin (the upstream installer defaults to ./bin, which usually isn’t on PATH) and runs trivy rootfs against the instance’s root filesystem (/) — scanning the provisioned node’s actual filesystem before it’s imaged, rather than treating the current directory as a container image reference. Pin whatever Trivy version your team has actually tested, not necessarily the one above.

Get the full checklist

Want the full checklist — including the CIS hardening script list and the Terraform node-group rollout pattern we use in production? Get the Kubernetes Golden Image Checklist or reach out directly at [email protected]. You can also find us at BPMBI @ The Clubhouse, Augusta, GA.

↑ Back to top

Ready to talk architecture?

Request a technical discovery call with our engineering team.

Request Discovery Call