From Rolling Deployments to Zero-Downtime Releases

Oct 6, 2026

From Rolling Deployments to Zero-Downtime Releases

This blog explains how Blue-Green deployment helps reduce downtime in online banking releases through traffic switching, pod readiness, static-resource versioning, and rapid rollback.

Executive Summary

In digital banking, deployment availability is a business requirement as much as a technical one. A deployment that interrupts access to online banking can affect customers attempting to log in, view balances, initiate transfers, or perform other critical operations.

Our earlier rolling deployment model introduced problems around version coexistence. Every new frontend build generated build-specific resources such as runtime.1234.js and main.1234.js. During a rolling restart, old and new pods could coexist. A request for a resource from the new build could reach an older pod where that resource did not exist, producing a 404 response.

Static resources were cached at the CDN for about four minutes. We needed to consider application availability and version consistency across the application shell, static resources, CDN, and backend pods.

We moved to a Blue-Green deployment strategy. The new version is deployed into a separate environment, its pods start and pass readiness checks, and the environment is validated before customer traffic is switched. The existing environment continues serving production traffic until the cutover.

In our environment, the observed ALB traffic transition from Blue to Green was about 40โ€“180 ms. The key benefit is not faster application startup. Pod creation, image pulling, startup, and readiness occur outside the customer-facing cutover window. The previous environment also remains available so that traffic can be routed back to it if a rollback is required.

A race condition remained during the ALB transition under high request volumes. A request for a build-specific static resource could reach the previous environment and receive a 404 response. We addressed this with a targeted NGINX retry for the affected static resources, with a delay of about one second before the retry.

1. Why We Needed a Different Deployment Strategy

1.1 The Rolling Deployment Model

A rolling deployment replaces application instances in stages. This avoids replacing every instance at once, but it also means different application versions can coexist during the deployment.

Rolling deployment with Build 100 and Build 101 pods behind the same load balancer

For backend APIs, teams can manage this coexistence through backward-compatible API contracts. For a frontend application whose resources are build-specific, the situation is more sensitive.

1.2 Build-Specific Static Resources

Every application build generates resources associated with that build. For example:

Build 100:
runtime.100.js
main.100.js
styles.100.css

Build 101:
runtime.101.js
main.101.js
styles.101.css

The browser receives an application shell that references a particular build's resources. If a request for runtime.101.js reaches a Build 100 pod, the resource does not exist there, and the server can return a 404 response.

Build 101 runtime request routed to a Build 100 pod and returning a 404 error

The pod itself could be healthy. The failure was caused by a version mismatch between the resource requested by the client and the application instance handling the request.

1.3 CDN Caching Added Another Timing Dimension

Static resources were cached at the CDN for about four minutes. The application origin could move to a new version while cached resources from the previous version remained available at the edge.

CDN-cached Build 100 resource routed through an ALB to Blue and Green environments

Deployment correctness had to account for three independent timing domains: application and pod readiness, ALB traffic convergence, and CDN cache lifetime.

2. Blue-Green Deployment Architecture

Blue-Green deployment introduced two complete application environments. Blue is the active version, and Green is the new version.

Customer traffic routed to active Blue Build 100 while Green Build 101 remains idle

The principle is simple: deploy and validate Green before allowing it to receive production traffic. This separates deployment preparation from production traffic cutover.

3. What Happens During a Blue-Green Deployment

3.1 Build and Image Creation

The CI/CD pipeline creates an immutable image for the release. The frontend resources generated by that build are part of the same release artifact.

CI/CD pipeline creating Build 101 static resources and an immutable container image

This preserves a clear relationship between the application version and the resources it serves.

3.2 Creating the Green Pods

The new release is deployed to Green without terminating the active Blue environment.

Blue and Green Kubernetes pods running in parallel before production traffic cutover

The Green pods go through image pull, container startup, application initialization, and readiness checks before they become eligible to receive traffic.

3.3 Pod Readiness

Readiness is a key mechanism in this architecture. Kubernetes distinguishes a running container from a container that is ready to accept traffic. A failed readiness probe causes the pod to be treated as not ready for normal service traffic. A startup probe can be used when an application needs more initialization time before liveness or readiness checks begin.

Kubernetes readiness probe keeping failed pods unavailable and admitting ready pods

For a banking application, readiness should represent the minimum state in which the application can process production requests without creating service risk. Depending on the implementation, this can include application initialization, configuration loading, dependency availability, and other required startup conditions.

3.4 Green Validation

Once the Green pods are Ready, the new environment can be validated before customer traffic is moved.

  • Confirm that the pods and application are healthy.
  • Run API smoke tests against the new environment.
  • Confirm that the required static resources are available.
  • Check connectivity to critical dependencies.
  • Test business-critical flows.
  • Review application and infrastructure metrics.

This allows the team to validate Green while Blue continues serving production traffic.

4. How the ALB Traffic Switch Works

The Application Load Balancer (ALB) provides the boundary for switching production traffic between Blue and Green. An ALB listener evaluates routing rules and forwards traffic to target groups. Target groups contain registered targets and can have their own health checks.

Application Load Balancer routing through separate Blue and Green target groups

Before the cutover, Blue is the active target group and Green is prepared without receiving normal production traffic. Once Green has passed the deployment gates, the routing configuration changes so that production requests are directed to Green.

5. The Critical Difference: Deployment Time vs Customer-Impact Time

A key distinction in Blue-Green deployment is the difference between deployment time and customer-impact time. The new pods still need to pull the image, start the container, initialize the application, and pass readiness checks.

What changes is where that work occurs.

Rolling deployment path compared with Blue-Green validation before customer traffic cutover

Blue continues serving customers while Green is being prepared. This removes the time-consuming preparation work from the customer-facing cutover path.

The customer-facing transition is limited to the routing window rather than the full duration required to create and warm the new pods.

The improvement should not be described as โ€œthe deployment became X% faster.โ€ The deployment preparation time still exists, but it does not need to translate into customer downtime.

6. Rollback: Turning Recovery Into a Traffic Switch

Blue-Green changes the rollback mechanism.

Normal state:

BLUE  = Build 100
GREEN = Build 101
Traffic = 100% Green

If monitoring identifies a production problem after the cutover, Blue can remain available as the previous known-good version.

ALB rollback routing production traffic from Green Build 101 to Blue Build 100

Instead of rebuilding and redeploying the previous image, the deployment can route traffic back to the running Blue environment.

This changes rollback from a deployment operation into a traffic-routing operation, which can reduce mean time to recovery (MTTR).

7. The Production Challenge After Blue-Green

7.1 The ALB Transition Window

Blue-Green solved the major version-coexistence problem, but production testing exposed a smaller race condition during the traffic switch.

In our environment, the observed ALB transition from Blue to Green was about 40โ€“180 ms.

Before:
Blue  = 100%
Green =   0%

Transition:
Blue  = 100% -> 0%
Green =   0% -> 100%

After:
Blue  =   0%
Green = 100%

Under high request volume, requests can arrive during this short interval. A request that arrives during the transition can encounter the previous routing state.

7.2 Static Resource 404 During the Transition

Build 101 static-resource request reaching Blue Build 100 during traffic transition and returning 404

This was not a persistent application failure. It was a transient race condition caused by the interaction between build-specific resource names and traffic convergence.

8. NGINX Retry as the Last-Mile Resilience Layer

To address this transient condition, a targeted retry mechanism was introduced at NGINX for the affected static resources.

NGINX retrying a transient static-resource 404 after the ALB traffic transition

The one-second delay was longer than the observed ALB transition window. This gave the traffic transition time to complete before the retry.

This mechanism was scoped to static-resource requests. A generic retry policy should not be applied to state-changing banking operations such as fund transfers without idempotency and transaction-safety guarantees.

9. CDN and Resource Versioning

Build-specific resource names remain part of the design. Instead of overwriting a stable path such as /main.js, each build produces a unique resource path.

/main.100.js
/main.101.js
/main.102.js

This allows multiple versions of a resource to coexist in the CDN. The application shell should reference the resource set belonging to its build.

The CDN cache lifetime of about four minutes is treated as a separate concern from the ALB transition. Build-specific resources provide coexistence, while the NGINX retry addresses the short traffic-convergence race.

10. Zero-Downtime as an End-to-End Property

The implementation showed that zero-downtime (ZDT) is not a property of Kubernetes or the load balancer alone.

Client
  |
CDN
  |
ALB
  |
NGINX
  |
Kubernetes
  |
Pod readiness
  |
Application
  |
Dependencies / data layer

A system can have healthy pods and still return customer-visible errors if resource versioning, caching, or routing transitions are not handled correctly.

Our final approach combines several mechanisms:

Blue-Green deployment
        +
Pod readiness
        +
Build-specific static resources
        +
CDN caching strategy
        +
ALB traffic cutover
        +
NGINX static-resource retry
        +
Fast rollback
        =
Zero-Downtime Release Architecture

11. Business and Operational Impact

Before the move toward ZDT, deployments were scheduled during low-traffic periods such as night-time maintenance windows. This was a risk-management decision. If a deployment could cause service interruption, a period with fewer customers using the platform reduced the number of users exposed to that interruption.

Before:
Release Ready
    |
    v
Wait for Night
    |
    v
Maintenance Window
    |
    v
Deployment

Blue-Green changed the release model.

After:
Release Ready
    |
    v
Deploy Green
    |
    v
Validate
    |
    v
Switch Traffic

Because Blue remains available while Green is prepared, a new deployment does not require the application to be taken offline. Subject to change-management, monitoring, and business controls, teams can perform deployments during normal operating hours, including the morning, instead of restricting them to night-time maintenance windows to avoid deployment downtime.

This changes the operational model: application availability no longer dictates the deployment window.

12. Deployment Sequence

1. Developer merges release
          |
2. CI/CD builds application
          |
3. Build-specific static resources generated
          |
4. Immutable container image created
          |
5. Image pushed to registry
          |
6. Green deployment created
          |
7. Green pods start
          |
8. Startup / initialization
          |
9. Readiness probes pass
          |
10. Green validation / smoke tests
          |
11. Green target group confirmed healthy
          |
12. ALB routing switched
          |
13. Traffic moves Blue -> Green
          |
14. Monitor application + business metrics
          |
15. Keep Blue available for rollback
          |
16. Decommission Blue after confidence window

13. Key Engineering Lessons

13.1 Blue-Green Requires Coordinated Deployment Controls

Blue-Green depends on the interaction between deployment orchestration, readiness, traffic management, resource versioning, CDN behavior, observability, and rollback. Two application environments alone do not address the dependencies across the release path.

13.2 Move Preparation Work Outside the Cutover Path

Image pulling, pod creation, startup, and readiness can take minutes. Blue-Green allows the previous version to continue serving traffic while the new environment completes this work.

13.3 Rollback Should Be Designed Up Front

Keeping the previous environment available allows the system to route traffic back to a known-good version without rebuilding and redeploying that version.

13.4 Short Transition Windows Can Expose Race Conditions

A short routing transition can still expose a race condition when request volume is high. Deployment design must account for requests that arrive while traffic moves between environments.

13.5 Static Resources Are Part of the Release Architecture

Frontend resources form part of the production deployment architecture. Their versioning, CDN lifetime, and relationship with the application shell must be considered as part of the release design.

14. Final Architecture

Zero-downtime Blue-Green architecture with pod readiness, ALB traffic switching, monitoring, and fast rollback

15. Conclusion

The move from rolling deployments to Blue-Green deployment addressed the version mismatch that could occur when build-specific frontend resources reached pods running an older build. Blue-Green separates release preparation from traffic cutover by allowing the new environment to complete startup, readiness checks, and validation while the existing environment continues serving customers.

The implementation showed that zero-downtime deployment depends on controls across the release path. Resource versioning, CDN caching, traffic switching, readiness checks, retry behavior, and rollback design must work together to protect customer requests during a release.

Keeping the previous environment available also allows recovery through a traffic switch instead of another deployment. This release model reduces the need to schedule deployments around maintenance windows while preserving a path back to the previous version.

Building Release Architectures for Production Reliability

Zero-downtime deployment depends on more than switching traffic between two environments. Pod readiness, static-resource versioning, caching, traffic routing, monitoring, and rollback must work together to protect customer-facing applications during a release. For teams building or modernizing applications with these production requirements, GeekyAnts provides custom software development services covering architecture, CI/CD, deployment, and post-launch support.

References

Subscribe to Our Newsletter

More from the engineering frontline.

Dive deep into our research and insights on design, development, and the impact of various trends to businesses.
Insight
The Deprecation Notice Nobody Reads
Oct 5, 2026

The Deprecation Notice Nobody Reads

This blog explains why API deprecation notices often fail and how staged retirement, telemetry, brownouts, and scheduled enforcement can make API migrations more controlled and predictable.

Insight
Why Legacy Systems Make Business Growth More Expensive: Navigate A Smarter Path to Legacy Modernization
Sep 23, 2026

Why Legacy Systems Make Business Growth More Expensive: Navigate A Smarter Path to Legacy Modernization

Learn how legacy systems make business growth more expensive and how edge-first modernization can remove constraints without replacing the existing system.

Insight
SSO, Audit Logs and RBAC: The Enterprise Features AI Prototyping Tools Do Not Cover | Sarika Gautam
Sep 23, 2026

SSO, Audit Logs and RBAC: The Enterprise Features AI Prototyping Tools Do Not Cover | Sarika Gautam

Why AI-generated prototypes fail enterprise review: the context behind SSO, the cost of skipping audit logs, and how role explosion makes RBAC a product of its own.

Insight
Feature Flags as Technical Debt: The Cleanup Nobody Schedules
Sep 21, 2026

Feature Flags as Technical Debt: The Cleanup Nobody Schedules

This blog explains how unmanaged feature flags create technical debt and how teams can detect, manage, and remove them safely.

Insight
Building AI-First Enterprises: Why System Design Matters More Than AI Adoption
Sep 21, 2026

Building AI-First Enterprises: Why System Design Matters More Than AI Adoption

This blog explores how system design, architecture, and validation shape AI-first enterprises, while examining AIโ€™s impact on software engineering and human decision-making.

Insight
The Product Studio in the AI Era: What Actually Changes | Sarika Gautam
Sep 21, 2026

The Product Studio in the AI Era: What Actually Changes | Sarika Gautam

What changes in product development when AI writes the code: the shift to architecture, the token cost of unplanned builds, and why juniors still matter.

Insight
AI and the Future of Digital Customer Experience: Where Technology Meets Human Creativity
Sep 18, 2026

AI and the Future of Digital Customer Experience: Where Technology Meets Human Creativity

A discussion on how AI, human creativity, research, and cross-functional collaboration are shaping the future of digital customer experience.

Footer

The Right Conversation Can

Save You Six Months.

Book a Call