Passed All Tests in Integration and Acceptance, Then Broke Production
Payment releases that pass integration and acceptance testing still break in production when configuration, not code, differs between environments that different teams own. A RACI matrix for deployment steps, a canary on the riskiest workflows, a written UAT parity spec and a core-job health probe independent of third parties close the gap.
What does a passed test promise?
A payment release that passes testing is safe to ship. Your release plan says so. So do your release manager, your risk register and the rubber stamp on the change form.
Our release passed all tests in our integration environment and all tests in the client's acceptance environment. Then, every fortnight, it took production payments down for half a day.
The code was fine. Whatever broke production was hiding where none of those tests were allowed to look.
The platform was Payment Hub EE, the open-source Mifos payment stack, collecting loan repayments over M-Pesa for a farm-finance organisation in Kenya, with my company, Fynarfin, as the vendor. Each release climbed from our system integration test environment (SIT) to the client's user acceptance environment (UAT), then production.
When production fell, the dashboards lit up with cascading payment failures, and we rolled back with scripts nobody else had reviewed. Every outage got a fresh theory, a frantic rollback and a four-leaf clover taped to the deploy button.
Where was the bug hiding?
Every vendor's pitch, ours included: test the code twice and the third run, in production, is safe. Except the ladder never runs it with the same configuration twice.
SIT was ours, and every third party in it was a mock, because the real ones refuse connections from our network. UAT was the client's, and their SREs chose which mocks became real services. Production was a cluster both teams managed. Same code, three configurations, three owners.
On 20 January 2022, before go-live, the Message Gateway could not resolve a config placeholder, callbackconfig.address, and the notification service was calling the SMS endpoint on the wrong port. I wrote: "livenessProbe is failing but it's green."
Three weeks later one of our engineers wrote: "Operations app not running has resurfaced in production… We are not able to reproduce the issue consistently so we will have to take educated guesses…" Resurfaced. Not able to reproduce. Educated guesses. That is drift, heard from the wrong environment. The release carried its code from SIT to production. It left behind the callback addresses, the credentials and the coffee-stained sticky note saying which database was which.
What did the canary catch?
The textbook fix was blue-green: two production stacks, flip the traffic, flip back if it burns. That needed new infrastructure and a client sign-off that had not landed yet. So we fell back to release engineering's cheapest experiment, the canary deployment.
At each deployment only the highest-risk payment workflows switched to the new release, under 5% of payment volume. A canary is supposed to catch bad code. Ours caught good code in a bad environment: production configuration had drifted from dev and UAT.
The charts drifted too. By March 2022 one platform had three Helm charts, and I wrote that we needed "one base helm chart that we can do helm lint against to prevent from breaking." A payment hub with three birth certificates eventually turns up to production holding the wrong one.
Why not test like production?
My stop-gap was a spec: test the payment channel migration in UAT with production-like mock data, proving each release stable and isolating production-only configuration issues on paper.
The spec moved the fight rather than ending it. We had to prove the behaviour in our own SIT first, on mocks. Then the client's SREs decided how many mock responses became real UAT services, theirs or another vendor's, none of which took calls from our SIT. Every swap was a sensible local decision, and a fresh chance to drift.
On 28 March 2022 one of our engineers, chasing a production fault, wrote: "I am unable to proceed with recreation of issue or testing the fix because I still do not have access to [the client's] Qa…"
SIT said: works on my mocks. UAT said: works on the mocks we kept. Production said nothing for half a day, which is how production says no.
Why couldn't anyone fix the config?
Because no single company could. Production Kubernetes and its Helm charts were co-managed by our SREs and the client's SREs. On 28 May 2022 I ran a helm upgrade in the client's cluster myself and got back "secrets is forbidden": the identity deploying the platform could not list the platform's own secrets. Two SRE teams on one cluster is a three-legged race in which each runner assumes the other knows the way.
That is the gap a RACI matrix closes: for each deployment step, who is Responsible, Accountable, Consulted and Informed, with exactly one Accountable name for canary routing, the blue-green switch and the rollback call. Startup environments often run without one, or with a foggy one: nobody is responsible, or several people are, each assuming someone else will decide and put the implementation tickets into the Program Increment or sprint plan. A canary with three owners is a canary nobody feeds.
So the fix was never a patch. It was process, persuasion and paperwork: cycles of joint triage, influencing a team I had no authority over, and in the end a small change to the accountability mapping for production processes, because several people were responsible for each process group. The next two major releases shipped with no operational downtime.
In November 2023 a joint call agenda still asked how to tell "whether the problem is on Payment Hub, Safaricom, Other system". The releases had stopped breaking. Nobody could yet say whose system had.
Where else does this repeat?
Any vendor shipping into a client's infrastructure climbs the same ladder: three configurations, three owners.
Even AI-written probes reach for the wrong dependency. I had Claude (Opus 5.5) build a synthetic payment connector that queues a notification after each payment, then write its probes. Its readiness probe failed whenever the notification service was down, which would have stopped payment intake; its write-up admits "all pods go unready together and the connector is effectively down." One run, scaled thresholds, no real cluster, and the model graded its own work: not a benchmark.
Payment intake should never stop because an SMS did not go out. Notifications are critical but downstream: give the messaging layer its own 99.99% uptime target and redundancy, a second aggregator and a retrying queue, so no upstream payment service takes forced downtime for it. Spring Boot's probe guidance agrees: external systems the application can work without "definitely should not be included" in readiness.
- One Accountable name per deployment step, with its tickets in the plan.
- A UAT parity spec signed before each release: every third party, mock or real, its owner, who holds its credential, and whether SIT can reach it.
- A core-job health probe: one synthetic payment job through your own engine and connectors, every third party stubbed at the boundary. Red means your release or your configuration broke.
- Provider error rates as alerts, never as probes, so a dead SMS aggregator pages someone while payments keep flowing.
Your tests prove the code. Production also runs the config, the credentials and the chain of command. Next week: the release tag that looked identical and wasn't.
Straight answers, marked up for Google.
- Why do releases that pass UAT still fail in production?
- Production runs configuration the test environments never had: real third-party endpoints, credentials, secrets and database settings. Test the differences, not only the code.
- What is a core-job health probe?
- A synthetic job through your own payment engine and connectors, every third party stubbed. Red means your release or configuration broke, not your provider.
- Should a readiness probe check SMS or mobile money providers?
- Not if the platform can take payments without them. Keep them out of readiness, give messaging its own redundancy and uptime target, and alert on each provider's error rate.
- Who should own canary and rollback decisions?
- One Accountable person per deployment step, named in a RACI matrix, with the implementation tickets in the sprint or Program Increment plan.
Two ways to stop this happening to your releases.
A core-job health probe that tells a broken release from a broken provider, plus provider alerts that page someone while payments keep flowing.
Canary or blue-green deploys, rollback on failed health checks and a release runbook for one service.
Where the same lessons came from.
On the GovStack payments building block, the test strategy ran BDD suites against mocked banks in CI, the same mocks-versus-real-services line this post is about. And a rollback that replays payments needs idempotency keys enforced at the ledger write so the replay does not charge twice.
- Founder, Fynarfin — scaled to $1M revenue (2019–2024)
- Apache Fineract — SDE & Solution Architect (ledger, line of credit, loan restructuring)