Never Reuse a Release Tag: One Overwritten Image, One Burned Version
Reusing a Docker image tag is dangerous because Kubernetes treats an unchanged tag as unchanged code: nodes with a cached copy keep the old build, and helm sees no diff to roll out. Make tags immutable at the registry, mint a new tag per build, deploy by digest, and let each pipeline push only to its own repository.
What is a release tag supposed to promise?
Everyone who ships containers believes a version tag identifies exactly one build. That is the point of writing v1.0.1 on a release: same label, same bits, on every node and every rollback.
We ran a Payment Hub EE production build for an agricultural-finance organisation in Kenya from January to August 2022, with the first live M-Pesa payment going through on 1 February 2022. By September 2022 our release checklist listed 17 repositories. Seventeen images, each carrying the same quiet promise.
Then the tag said v1.0.1, the image behind it was another bank's connector, and the first fix on the table was reinstalling the platform.
Why does Kubernetes not notice a changed image?
Two defaults combine. Most registries let you push a new image under an existing tag by default; the tag then points to the new digest. And when a pod spec names an image with a specific tag, Kubernetes defaults imagePullPolicy to IfNotPresent (the Kubernetes image pull policy rules). A node that already holds v1.0.1 will not ask the registry again.
So an overwritten tag goes unnoticed in three places. The registry accepted it. The node trusts its cache. And helm upgrade sees the same image string and rolls nothing out. Only a pod rescheduled onto a fresh node pulls the new build, so two replicas can report the same version and run different code.
There is a cruel detail: Kubernetes defaults to Always only for :latest or untagged images. So the trap opens exactly when a team does the responsible thing and pins a real version for production.
What happened when the tag moved?
On 13 October 2022 I spotted another bank's connector being released under our payment channel connector's name. Our release pipeline for that bank's connector was pushing its image into this client's container registry, and its configuration looked right at first glance. A colleague's first reply was one word: "how?"
No person had overwritten anything, and no client could have: clients versioned their own chart configuration on top of our releases, they did not publish images. Our own pipeline had. It held push rights to a registry that was not its own, so the push succeeded. The prevention is cheap: give each pipeline a registry credential that can push only to its own repository, and a misrouted push is refused instead of accepted.
I fixed the pipeline and wrote down the cost: on clusters without imagePullPolicy: Always, the whole helm chart would have to be uninstalled and reinstalled in production, causing confusion because it would have to be reverted again later.
The workaround avoided the reinstall: an engineer re-released v1.0.1 as v1.0.2, a new tag every node would pull. Except that burned a number. Ninety minutes later I wrote: "This means we can never release channel connector v1.0.2 for [the client]." The upstream project would ship its own v1.0.2 one day, with different code, and this client could never take it under that name.
Better fixes existed, and we rolled them in over the following releases rather than that afternoon. A reinstall alone would not have cleared the images cached on each node. A digest pin, ph-ee-connector-channel:v1.0.1@sha256:…, needs no downtime and no new number: a changed image string is a changed spec, and a changed spec is a rolling update that pulls.
Why did the problem keep coming back?
Because the pressure to re-release never went away. On 26 February 2023 a large pull request, the migration of the channel connector's API to Spring, arrived carrying documentation for features from several earlier releases. The release manager's question was whether to fold it into the current release, since customers often expect all documentation and tests for a released feature to be part of that release, even though in agile teams they rarely are. My reply: "There is no way to mutate images in non Fynarfin container registries. Please suggest tactical change of how to address this recurring issue which [the client] requested to fix as top priority in May 2022."
The pull is stronger at a small vendor. When one team builds the product and also customises it for each client, clients see no reason to tell a product release apart from their own customisation. The release manager is left balancing what the client expects a release to contain against how the product can technically be released.
Nine months after the client called it top priority, the same choice was back: re-release an old version, or cut a new one. I chose to fix the cause rather than rule on the tag: "documentation and tests are not features or bugs and we have to strive to include them in the same sprint." Late docs and late tests are what create the urge to re-push a version. Ship them with the feature. If they arrive late, ship them in a new patch release, even one with no code change, and leave the old tag alone.
That holds for internal release candidates too: a shared test cluster caches images exactly like production. Docs that land after rc.1 mean rc.2. A team that re-pushes candidates in test will re-push releases under a deadline.
Why couldn't a rollback or a delete undo it?
And the damage could not be rolled back. A rollback restores a label, not the bits. If v1.0.1 has pointed at two digests, helm rollback to a revision that says v1.0.1 gives you whichever copy each node holds. The old build survives at the registry only as an untagged digest until cleanup, and nothing in the chart recorded which one.
Nor could the number be freed by deleting it. A version is released to every system that has resolved its name to bits and kept a copy, announced or not, and deleting it removes only the registry's copy. The edge cases:
- A git tag with no artifact, never fetched: not released; Git's documentation says just re-tag it. Once anyone fetches it, it is.
- An artifact published but never announced: released. In October 2022 a colleague spotted
ph-ee-engine 1.2.0in our chart repository while "the tag is not yet released". - Only ephemeral pipelines consumed it: the runners are gone, their outputs are not. Pull-through caches, long-lived clusters and SBOMs keep the name.
So reusing a deleted release's tag is valid in two cases: nothing outside the build ever resolved it, or the new bits are identical, such as a reproducible rebuild or a release page re-created on the same commit. Many ecosystems refuse even those: npm never reuses an unpublished version, PyPI never reuses a filename, Go's checksum database rejects a moved tag, and a deleted GitHub immutable release retires its tag for good. Our v1.0.2 failed the first test the moment a node pulled it.
The deeper mistake was sharing upstream's numbering at all. A fork build should never carry a bare upstream number: tag it v1.0.1-acme.2, upstream version plus your build counter, and v1.0.2 stays free for upstream. One catch: SemVer reads anything after a hyphen as a pre-release, so 1.0.1-acme.2 sorts below 1.0.1 and version ranges skip it. Pin fork builds exactly.
The same habit reached the charts. In June 2022 a deploy failed because ph-ee-engine 1.1.3-SNAPSHOT was not found in our chart repository. SNAPSHOT versions are mutable by design.
To see how common this is, Claude ran a release-safety audit on Avik's behalf against the public openMF/ph-ee-env-template Helm charts (commit e8bc9ef, July 2024). It wrote 12 findings before seeing the answer key. Among them: three connector templates set a pinned imageTag in values and then render the image without it, so the pod runs an implicit :latest while the values claim a pin. Against the four incident types the team actually hit: 3 hits, 1 partial. The limit matters more than the score: the run was primed with those four categories, it read templates without rendering them through helm, and the public template is not the production fork where the incidents happened. Read it as "the same risks exist upstream", not as a benchmark.
That template finding deserves a test of its own, because helm lint and schema validation both pass a template that ignores its value. Render the chart with every image tag overridden to a sentinel such as override-check, and fail the build if any rendered image line lacks it. helm-unittest can assert the same per template.
Which of your tags can be overwritten today?
Not every label should be frozen. latest, python:3.12, a GitHub Action's @v4 and a DNS record are pointers: moving them is their job. A release version is an identifier: one name, one set of bits, for good. The failure is deploying a pointer as if it were an identifier, or moving an identifier as if it were a pointer. Attackers use the second: in March 2025 someone moved the tj-actions/changed-files version tags to a malicious commit, and GitHub's advisory counts over 23,000 repositories impacted. Workflows pinned to a commit SHA were not.
Closed-source teams file a misrouted pipeline under "can't happen here". It happens wherever one CI credential can push to every repository and the image name is a value someone typed: a service scaffolded from a template, a pipeline file copied from a sibling, a promotion job with a hand-kept mapping table. The new thing's first release lands on the old thing's tag, and nothing alerts.
The list the pipeline should enforce, not the team:
- Turn on tag immutability at the registry (ECR tag mutability settings, Harbor immutability rules, Docker Hub immutable tags), so a re-push fails loudly.
- Scope each pipeline's push credential to its own repository, and derive the image name from the source repository, not a typed variable.
- Tag every build with a git SHA or a semver that is never reused; deploy by digest (
image@sha256:…) where you can. - Suffix fork builds. Bare numbers belong to upstream.
- Fail CI on
:latest, untagged images, or a chart whose render ignores an override. - Bump the chart version per release and commit the lock file. No SNAPSHOT charts in production, and release candidates count as releases.
Run the audit on your own estate this week: list every tag in production, then ask which ones the registry would let you overwrite, and which pipelines could. Next week: the incident where the first count of affected customers was the smallest one.
Straight answers, marked up for Google.
- Is imagePullPolicy: Always enough on its own?
- No. It pulls only when a container starts, so running pods keep their digest until they restart. A moved tag still leaves rollback without the old build. Unique tags or digests fix the cause.
- Why not just use the latest tag and always pull?
- Because then the running version depends on when each pod started. Replicas can run different code, and nothing in the chart or values records which build was live; only each pod's status does, until the pod is gone.
- Can helm rollback recover from an overwritten tag?
- Not reliably. Helm restores the old spec, but the spec names the tag, and the tag now points at the new build. The old build survives as an untagged digest until registry cleanup; you can reach it only if you recorded that digest.
- When is it safe to reuse a deleted release's tag?
- Only when nothing outside the build ever pulled, fetched or cached it, or when the new bits are identical. npm, PyPI, Go's module proxy and GitHub immutable releases refuse it regardless.
Two ways to stop this happening to your releases.
Immutable tags at the registry, push credentials scoped per pipeline, CI gates for :latest and ignored overrides, and a rollback you can trust.
A read of your charts, pipelines and registries that ranks the ten release risks most likely to cost you a production night.
Where the same lessons came from.
Last week's post covered the other half of release safety: why payment releases pass testing, then break production. On the GovStack payments building blockcase study shows the wider release setup these components shipped in. And a release that replays payments needs idempotency keys enforced at the ledger write.
- Founder, Fynarfin — scaled to $1M revenue (2019–2024)
- Apache Fineract — SDE & Solution Architect (ledger, line of credit, loan restructuring)