A Deployment Is Not Done When CI Turns Green
A green pipeline is a good signal. The build succeeded, the configured checks passed, and an artifact is ready. It is tempting to treat that as the end of the release.
I would treat it as the handoff to a different kind of evidence. Does the application behave correctly in the environment where people actually use it?
State the expected behaviour before shipping
A release should have a short, concrete description of what changes and what must continue to work. If it adds a new sign-in flow, the checks should include signing in and reaching the protected screen. If it changes message processing, they should include a message moving through the relevant stages.
A process being alive is useful information. It is not proof that the complete journey works. An application can serve its health endpoint while a background worker cannot reach its queue.
I would choose a few representative journeys and keep the checks small enough that somebody can understand a failure quickly.
Separate availability from correctness
A successful HTTP response does not necessarily mean the expected result was produced. For a report export, the useful outcome is a readable file with the intended contents. For an asynchronous request, acceptance is only the beginning of processing.
That changes the signals worth watching. Alongside errors and latency, I would consider queue age, completion status, and the specific business outcome the release touches. The right checks depend on what changed.
Synthetic checks should use controlled test data and have a cleanup plan. A deployment verification that creates confusing production records is making its own operational problem.
Release to a smaller audience when it helps
A canary rollout exposes a change to a limited portion of a service and evaluates it before wider rollout. Google's SRE workbook chapter on canarying explains the importance of evaluation and integrating that decision into the release process.
For a small team, I would keep the first version understandable. Decide who receives the change, which signals are compared, how long the observation lasts, and what stops the rollout. A percentage slider without those decisions is only a different deployment button.
Low traffic also changes the interpretation. No observed errors from a handful of requests is weaker evidence than the same result across a representative workload.
Recovery should be a plan, not a hope
I would ask what happens if the new application version is removed but its database changes remain. Can the previous version still read the data? Have background jobs started using a new format? Does a configuration change need to be reversed separately?
For a schema change, an additive migration followed by a later cleanup may be easier to recover from than a single destructive step. The particular approach depends on the application and its deployment constraints.
The useful definition of done is straightforward: the change is running, the relevant journeys have been checked, the signals are understood, and the recovery path is clear. CI gets the release to the starting line. Production behaviour decides whether it stays there.