Back to the notebook
Oct 09, 20265 min read

Deployment Strategies: Choosing How a Release Reaches Production

DevOpsCloudEngineering

A deployment strategy answers two practical questions: how does the new version replace the old one, and how do we limit the damage if something is wrong?

There is no universal winner. A small internal service with an agreed maintenance window has different needs from a public application processing requests continuously. The best choice is the one a team can operate, observe, and recover from.

I find it useful to separate application replacement, traffic exposure, and feature activation. They can be combined, but they are not the same decision.

Recreate: stop the old version, start the new one

A recreate rollout removes the old application instances before bringing up the replacement. It is straightforward, but introduces an availability gap for the affected service unless another layer covers that gap.

I would consider it for a noncritical internal tool with an acceptable maintenance window. It can also be appropriate when concurrent application versions are undesirable, though existing connections, jobs, and external clients still need a plan.

Recovery usually requires restoring the previous version and waiting for it to become ready. Simplicity does not eliminate the need for a tested restart procedure.

Rolling: replace instances progressively

A rolling rollout replaces portions of the running application while the remaining instances continue serving. Old and new versions overlap, so both must understand the shared database, APIs, and messages during that overlap.

Kubernetes Deployments support RollingUpdate and Recreate. For rolling updates, maxSurge controls how far the replica count may exceed the desired count, and maxUnavailable limits how many desired replicas may be unavailable. The Deployment documentation describes these controls.

For a stateless application with several replicas, this would often be my starting point. I would verify readiness checks, graceful shutdown, available capacity, and compatibility before treating it as an uninterrupted release.

A standard Kubernetes Deployment reports a stalled rollout; it does not automatically roll back merely because the progress deadline is exceeded. Recovery needs an operator or additional automation.

Blue/green: prepare a replacement environment

Blue/green keeps the existing environment and a replacement environment available. Validate the replacement, then move traffic to it. Keeping the old environment ready can make switching traffic back fast.

The tradeoff is additional capacity and coordination. Scheduled jobs and queue consumers do not necessarily follow an HTTP traffic switch. They can continue running in both environments unless explicitly managed.

I would choose it when a clean traffic cutover and a warm recovery target are valuable enough to justify that complexity. It is particularly important to examine shared state: switching requests back cannot undo data already written by the new version.

Canary: expose a limited audience, then evaluate

A canary introduces the new version to a controlled portion of production traffic or users before wider exposure. Define what makes the release acceptable, observe that audience, and either proceed or stop.

The audience could be a percentage of requests, a tenant group, or a region. If a user journey needs a consistent version, cohort assignment should reflect that rather than randomly switching every request.

Google's SRE workbook on canarying releases explains why evaluation is part of the release process. The comparison should account for the workload: an apparent improvement is unconvincing if the canary happens to receive easier requests.

I would choose a canary for a consequential change when there is enough representative traffic and useful telemetry. Five quiet minutes are not meaningful evidence for a task that runs once a day.

Linear: increase exposure in equal steps

A linear traffic rollout shifts exposure in equal increments at regular intervals. For example, move another portion of traffic every observation window until the replacement handles everything.

AWS distinguishes this timed progression from its canary configurations in its deployment strategies overview, which also covers blue/green and all-at-once switching.

I would still pair a timed ramp with health gates. A schedule should not continue promoting a version whose relevant signals are deteriorating. The percentage and interval are operational choices, not proof of safety.

These labels overlap: a blue/green environment can receive traffic gradually, and a canary can run alongside a rolling process. All-at-once exposure also differs from recreate: a prepared replacement environment can receive all traffic in one switch without first tearing down the live environment.

Feature flags and shadow traffic solve adjacent problems

Feature flags let code be deployed while activation is controlled separately. They can target a cohort or disable a feature quickly, but cannot erase incompatible writes or make a destructive database change reversible. Plan their ownership and eventual removal.

Shadow traffic sends a copy of requests to a candidate without using its responses for the live user journey. It can be useful for comparison, provided candidate-side effects are prevented or isolated. Duplicating real message delivery or other mutations would create a second production action rather than a harmless test.

A/B testing asks which experience performs better. A canary asks whether a release is acceptable to expand. They may use similar audience routing, but need different evaluation criteria.

Choose the strategy around the failure you need to contain

For a modest stateless service, I would begin with a rolling release and reliable health checks. For a release needing a ready replacement and fast traffic reversal, I would consider blue/green. For a high-impact behavioural change with observable traffic, I would add a canary. For an internal service with an explicit downtime agreement, recreate may be sufficient.

Across all of them, database compatibility is a separate obligation. Expand, Migrate, Contract keeps the schema compatible while application versions change. None of these rollout strategies repairs a column that old code can no longer write.

A good rollout has an identifiable version, a defined audience, meaningful acceptance signals, and a recovery path that includes shared state. The strategy gives those decisions a structure; the team's preparation makes it dependable.