Most production incidents aren’t solved by the cleverest engineer in the room. They’re solved by whoever calmly works out what is actually failing, and when it started. This is the checklist we use on Python backends, in the order we use it.
1. Confirm the impact
Before touching anything, write down three things: who is affected, what they can’t do, and since when. “Checkout fails for card payments since 14:05” leads you somewhere. “The site is slow” doesn’t. Keep notes with timestamps in one channel. They become your timeline later.
2. Ask what changed
Systems that worked yesterday usually break because something changed. Check, roughly in this order:
- deployments and configuration changes
- database migrations and scheduled jobs
- third-party services: payment providers, email, identity, and their status pages
- certificates, domains and API keys that may have expired
- traffic: a campaign, a partner integration or a bot
If a release lines up with the start of the problem, rolling back is usually faster and safer than a hotfix. Stabilise first, then find the cause.
3. Know which release is running
“Did the deploy go out?” shouldn’t take ten minutes to answer. Pass the commit hash into the build as a RELEASE variable, and let the app report it on every response, in every log line and on its health endpoint:
import logging
import os
from fastapi import FastAPI, Request
RELEASE = os.environ.get("RELEASE", "unknown")
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s %(levelname)s release=" + RELEASE + " %(name)s: %(message)s",
)
logger = logging.getLogger("app")
app = FastAPI()
@app.middleware("http")
async def add_release_header(request: Request, call_next):
response = await call_next(request)
response.headers["X-Release"] = RELEASE
return response
@app.get("/api/health")
async def health() -> dict:
return {"status": "ok", "release": RELEASE}In a Docker build that is an ARG RELEASE plus ENV RELEASE=$RELEASE, built with --build-arg RELEASE=$(git rev-parse --short HEAD). Now curl -sI against any endpoint tells you exactly which code answered.
4. Follow one failing request
Pick a single failed request and follow it end to end: the error, the log lines around it and every service it called. This is where correlation IDs pay for themselves. One ID takes you through the API, the background workers and the calls to third parties.
5. Check the database
When requests hang rather than fail, the database is a common culprit: a long transaction, a migration holding a lock, or a pool with no free connections. In PostgreSQL, start with what is running now:
SELECT pid,
now() - query_start AS running_for,
state,
wait_event_type,
left(query, 60) AS query
FROM pg_stat_activity
WHERE state <> 'idle'
AND pid <> pg_backend_pid()
ORDER BY query_start;A wait_event_type of Lock means a query is waiting for another one. This shows who is blocking whom:
SELECT pid,
pg_blocking_pids(pid) AS blocked_by,
now() - query_start AS waiting_for,
left(query, 60) AS query
FROM pg_stat_activity
WHERE cardinality(pg_blocking_pids(pid)) > 0;Both queries were tested on PostgreSQL 17 with one transaction holding a row lock and a second update waiting behind it. The waiting query shows up with the blocker’s pid. Ending the blocker with pg_cancel_backend or pg_terminate_backend is sometimes the right call, but check what it is doing first. If it’s a migration or a payment job, stopping it halfway can cause more trouble than waiting.
6. Look at queues and external services
Growing queues, retries piling up and timeouts to one provider tell you where the slowdown sits. If a payment provider is having a bad day, make sure your retries have a limit and use idempotency keys, so a recovery doesn’t turn into duplicate charges or double-processed webhooks.
7. Afterwards: fix the cause, not just the symptom
Once things are stable, write a short, blameless review: the timeline, the cause, how you found out and what would have made it quicker. Then do one or two concrete things, such as the missing alert, the test that would have caught it or the runbook step you wished you’d had. A review that leads to no changes is just a story.
The checklist
- Who is affected, what can’t they do, and since when?
- What changed: release, config, migration, job, provider, traffic?
- Which release is running? Can we roll back?
- What does one failing request look like, end to end?
- Is the database waiting on locks or out of connections?
- Are queues growing, or is one external service slow?
- After: timeline, cause, and one or two changes that stick.