django-deferred-migrations: safe destructive schema changes in one PR and one deploy (PostgreSQL)

We’ve open-sourced the migration tooling we run at Breww. Introducing django-deferred-migrations :rocket:

The problem. In a rolling deploy, migrations run first while the old release still serves every request, and then old and new processes overlap until the rollout finishes. Anything that breaks the old code’s view of the schema errors for the length of that overlap - RemoveField, DeleteModel, a new NOT NULL column with no database default, tightening a column to NOT NULL, a rename, or an index build that holds a write lock.

The standard answer is to split the change across two deploys and remember to ship the second one. In practice the second one gets forgotten, and it can’t go out until every environment has taken the first.

What this does instead. It splits each migration run into two phases around the rollout, and defers only the destructive SQL:

Step Database Old code New code
migrate_pre_deploy legacy_ref relaxed so inserts may omit it; DROP COLUMN queued Works Not running yet
Rollout Works: the column is still there Works: its model has no such field
migrate_post_deploy Waits for terminating pods, then DROP COLUMN legacy_ref Gone Works

DeferredRemoveField updates Django’s migration state immediately, so nothing in the migration graph ever depends on unrun work. That’s the main difference from the whole-migration approach taken by django-syzygy and django-safemigrate, and it’s what stops a deploy blocking when two PRs touching the same app land between releases. It also means it works perfectly with django-linear-migrations (which we love, if you don’t use django-linear-migrations, take a look).

What’s in the box:

  • DeferredRemoveField, DeferredDeleteModel, DeferredRenameModel, SetNotNull, trigger-synced column reshaping with batched backfills, and concurrent field/index/constraint builds.

  • A makemigrations that writes the deploy-safe operation for you, and fix_deploy_safety to rewrite existing ones.

  • check_deploy_safety for CI, with per-app baselines so an existing project can adopt it without touching history.

  • A short lock_timeout with jittered retry on every DDL statement, so a migration never queues behind a long transaction.

  • Differential tests: the suite generates the same change with Django’s makemigrations and with ours, applies both to separate databases, and compares columns, types, nullability, defaults, collations, comments, indexes, constraints, sequences, triggers, functions, views and row contents - at every step, including after a rollback and after full reversal.

Requirements: Python 3.12+, Django 5.2 / 6.0 / 6.1, PostgreSQL 14+, psycopg 3. No other backend - migrate_pre_deploy refuses to run on one.

It needs a post-rollout hook. If your platform only gives you a pre-deploy release phase, you can’t get the full benefit. Helm pre-upgrade/post-upgrade Jobs are what we use; there’s a generic CI recipe in the readme.

The readme has a comparison table covering django-syzygy, django-safemigrate, pgroll and django-pg-zero-downtime-migrations, including what each of them does better and which ideas this borrows from them. Thank you to the authors of those for their ideas and insights - without those, we probably would never have built django-deferred-migrations (or at least it wouldn’t have been as good as we think it is).

Feedback is very welcome, especially from anyone doing this at larger scale than us. We hope this can use useful to others.

2 Likes

EDIT: A moderator has merged my post on Let's talk zero-downtime migrations into this thread. I’m not really sure why, but I’ve added this line here as my post below is now out of context. It relates to that other thread.

Reviving this because the thread shaped what I ended up building, and because @boxed’s question upthread never really got an answer.

Late to this one, but the thread is a good record of the problem, so I’ll add a data point rather than reopen the argument.

@boxed asked for the simple version: “a simple check command I can run in my deploy script that fails the deploy if I’m in one of those situations”, covering the two cases that cause most real transitory downtime, which are adding a non-nullable column with no db_default and dropping a column. That turned out to be most of what we needed day to day, and it’s cheap to do, because the check reads the operations in unapplied migrations rather than asking anyone to annotate them.

We’ve been running the rest of it in production at Breww for a while, and have just released it as django-deferred-migrations. PostgreSQL only.

The design decision that differs from what’s already out there: it defers the destructive SQL, not the migration. DeferredRemoveField takes the field out of Django’s migration state straight away, relaxes the old column so the old release keeps working (drop NOT NULL, or set a database default where there’s a safe one), and queues the DROP COLUMN for a post-deploy step that runs once the rollout has finished and the old pods have gone.

The reason that’s worth doing: if you defer whole migrations, then as soon as two PRs touching the same app land between deploys, the second one’s migration depends on an unapplied post-deploy migration, and the pre-deploy pass has to refuse or guess. django-syzygy raises AmbiguousPlan here, and django-safemigrate’s strict mode blocks. Nothing here is ever deferred at the migration level, so the graph never contains unrun work and the situation can’t come up. That’s also what lets it sit alongside the very helpful django-linear-migrations with one enforced leaf per app.

In practice it means a destructive change is one PR and one deploy. No remembering to ship the second half next week, and no waiting for every environment to take the first half.

Roughly what’s in it:

  • Operations: DeferredRemoveField, DeferredDeleteModel, DeferredRenameModel (the old table name survives as a security_invoker view for the length of the rollout), SetNotNull with a batched backfill, trigger-synced column reshaping, and AddFieldConcurrently / AddIndexConcurrently / AddConstraintConcurrently.
  • check_deploy_safety for CI, which is @boxed’s command plus about twenty other rules, with per-app baselines so you can adopt it on an existing project without rewriting history.
  • fix_deploy_safety rewrites the mechanical cases, and a makemigrations wrapper writes the safe operation in the first place, so most of the time you don’t think about any of this.
  • Every DDL statement runs under a short lock_timeout with jittered retry, so a migration can’t queue up behind a long transaction and take the site down with it. That failure mode has nothing to do with old code vs new code, but it bites just as hard.

On @adamchainz’s point before about platforms only giving you a pre-deploy hook: that’s still true, and it’s a real limit. This needs somewhere to run a post-rollout command. We use Helm pre-upgrade and post-upgrade hook Jobs (we run on Kubernetes), and there’s a generic CI recipe in the readme. If all you have is a release phase, it falls back to a migrate_full that runs both phases back to back, which gets you the lock safety and the checks but not the overlap safety.

Prior art it leans on, with thanks: django-syzygy for pre-drop preparation and for the idea that checks should inspect operations rather than trust annotations, django-safemigrate for grouped deploy logs and for failing loudly on genuine conflicts, and pgroll for the two-phase lifecycle and the lock-timeout-with-retry pattern. The readme has a comparison table that’s explicit about what each of them does better.

Very happy to hear the tradeoff is wrong, particularly from anyone who’s run the whole-migration approach at scale and found the multi-PR case less painful than we did.