Standby and failover
deploy/scripts/failover.sh keeps a second server as a warm standby and
promotes it when the primary dies. It’s for servers behind
Cloudflare Tunnel and runs on the standby, from
its checkout of the repository.
How it works
Section titled “How it works”- Database: the standby’s Postgres is a streaming replica of the
primary’s (hot standby, read-only). It connects through an SSH link, so
the primary’s Postgres stays unpublished.
setupcreates a replication login and a slot namedreplica_<hostname>on the primary, and caps the WAL a slot may hold back at 8 GB (STANDBY_MAX_WAL) so a standby that stays offline can’t fill the primary’s disk. Past the cap the standby needsfailover.sh setup --rebuild. - Files: with
BLOB_DRIVER=s3both servers use the same bucket and nothing is copied. With the volume, missing files are copied every few minutes (they never change once stored). - Config:
env/*.envandsecrets/are copied from the primary; a standby needs the same database passwords,pepper_keyandjwt_signing_key. The tunnel token comes fromenv/tunnel.envor the primary’scloudflaredservice. - Release: the API image the primary runs is copied (or pulled, or built for another CPU architecture from the same commit), and so is the live web release.
- Cutover: the standby runs a tunnel connector with the same token. Cloudflare sends requests to any connected connector of a tunnel, so the standby answers as soon as it connects. No DNS change.
- Write pause: the same switch as the admin dashboard’s write pause.
Writes wait up to 10 seconds, then get a retryable
maintenance_write_pause; reads keep working.
Everything privileged goes through Docker, so no sudo is needed, except to
stop a cloudflared installed as a system service.
Requirements
Section titled “Requirements”- The standby is prepared like the primary (Docker, the repository cloned,
the user in the
dockergroup, the data volume at the sameDATA_ROOT), but without the tunnel setup’s last steps: nocloudflared service install, no Actions runner and no update timer. Those would serve or deploy from it. - Keep its checkout on a recent
main; the scripts don’t update it. - SSH from the standby to the primary. The first
setupcreatesdeploy/standby/id_ed25519and prints thessh-copy-idcommand that authorizes it. That key can use Docker on the primary, which is root-equivalent: keepdeploy/standby/private (git ignores it). The primary’s sshd must allow TCP forwarding (on by default on Debian and Ubuntu). - Both servers run the same Postgres major version.
Set up a warm standby
Section titled “Set up a warm standby”On the standby:
deploy/scripts/failover.sh setup --primary deploy@primary.example --dry-rundeploy/scripts/failover.sh setup --primary deploy@primary.exampledeploy/scripts/failover.sh statussetup copies config and secrets, sets up replication, copies the database,
the release and the files, and installs a cron job that runs
failover.sh sync every 5 minutes (--every N to change it). It asks before
replacing a database on the standby and refuses one that holds accounts
(--replace-local-database overrides). Running it again is safe.
Other options: --port (SSH port), --peer-dir (the checkout on the
primary, default the same path), --yes.
status shows replication, how far the standby is behind, the slot on the
primary, the last file sync and whether the primary and the public URL
answer. The sync log is deploy/standby/sync.log.
What a failover can lose: replication is asynchronous, so the last
moments before a crash may not have reached the standby (status shows the
lag). With the volume, files uploaded since the last sync can be missing
too.
Fail over
Section titled “Fail over”When the primary is gone, on the standby:
deploy/scripts/failover.sh promote --dry-rundeploy/scripts/failover.sh promoteIt refuses while the primary still answers (use
Moving to a new server for a planned move;
--force takes over anyway and stops the primary first if it can reach it).
It promotes Postgres, turns the write pause off, starts the API, worker and
Caddy with the copied release and connects the tunnel. If no release was
copied, deploy afterwards with deploy/tunnel/update.sh --force.
Afterwards
Section titled “Afterwards”-
Fence the old server before it comes back. It would reconnect to the same tunnel with old data. As soon as SSH works:
deploy/scripts/failover.sh fence --host deploy@old.exampleThis stops its tunnel connector, API and worker and turns on the write pause in its database. The surest fence is a new tunnel: create one in Cloudflare, move the public hostname to it, put the new token in
env/tunnel.envon the new primary and recreate itscloudflaredcontainer. -
Move deploys: install the Actions runner on the new primary (step 5 of the tunnel guide) and remove the old one.
-
Set up a new standby, for example the repaired old server:
failover.sh setup --primary <new primary>there.