Running the server

Running the server as a systemd service or a container, its files and permissions, environment variables, and several nodes on one store.

The systemd service

The packages install versealx-server.service. It runs the server as the versealx user and confines it tightly.

SettingEffect
Type=notifysystemctl start returns once every listener is bound, so services ordered after it start after mail can actually be accepted.
ExecStartPre=… check --config /etc/versealx-server/versealx-server.tomlThe configuration is checked before every start; a refusal stops the start with the reason in the journal.
ExecStart=… run --config /etc/versealx-server/versealx-server.tomlThe server.
Restart=on-failure, RestartSec=2sRestarted if it exits with an error.
User=versealx, Group=versealxNever root.
AmbientCapabilities=CAP_NET_BIND_SERVICEAllows ports below 1024 and nothing else.
StateDirectory=versealx-server (mode 0750)/var/lib/versealx-server, the only place it can write.
ConfigurationDirectory=versealx-server (mode 0700)/etc/versealx-server.
ProtectSystem=strict, ProtectHome=true, PrivateTmp=true and further sandboxingThe rest of the system is read-only or hidden.
LimitNOFILE=65536Room for many connections.
Environment=VERSEALX_CONFIG=/etc/versealx-server/versealx-server.tomlThe configuration path.

Everyday commands

TaskCommand
Start, and start at bootsudo systemctl enable --now versealx-server
Stopsudo systemctl stop versealx-server
Restart after changing the configurationsudo systemctl restart versealx-server
Statussystemctl status versealx-server
Follow the logjournalctl -u versealx-server -f

The configuration file is read at start, so restart the service after changing it. Settings kept in the store change without a restart; see Runtime settings.

On SIGTERM or SIGINT the server stops accepting new work, lets its tasks finish, and exits. Mail in the queue stays there and is picked up at the next start.

Environment variables

Put secrets such as a relay password or a database URL in a drop-in rather than in the configuration file:

sudo systemctl edit versealx-server
[Service]
Environment=VERSEALX_RELAY_PASSWORD=YOUR_PASSWORD
Environment=RUST_LOG=info
VariableUsed for
VERSEALX_CONFIGThe configuration file, when --config is not given.
Any name you reference as ${NAME}node, [store] url, and relay passwords.
AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEYCredentials for an S3-compatible bucket.
RUST_LOGThe log level: error, warn, info (the default), debug or trace.
HOSTNAMECommonly referenced as node = "${HOSTNAME}" in containers.

Files and permissions

With the layout init creates:

PathOwner and modeContents
/etc/versealx-server/versealx-server.tomlversealx, 0600The configuration.
/var/lib/versealx-server/versealx, 0700The data directory.
/var/lib/versealx-server/store.dbversealxThe SQLite store.
/var/lib/versealx-server/blobs/versealx, 0700Message bodies, encrypted.
/var/lib/versealx-server/kekversealx, 0600The key-encryption key. The server refuses a file others can read.
/var/lib/versealx-server/tls/versealx, 0700The certificate and key.
/var/lib/versealx-server/admin.sockversealx, 0600The admin socket, while the server runs.

Run every versealx-server command that touches these as the service user: sudo -u versealx versealx-server <command> --config /etc/versealx-server/versealx-server.toml.

In a container

The container image holds the same binary, the CA certificates and time zone data, and nothing else. It runs as user and group 10001.

ItemValue
Entry pointversealx-server, with run as the default command.
VERSEALX_CONFIG/etc/versealx-server/versealx-server.toml
Volumes/etc/versealx-server and /var/lib/versealx-server
Exposed ports25, 465, 587, 143, 993, 110, 995, 4190, 443, 80

Set up once, then run:

docker run --rm -v versealx-etc:/etc/versealx-server -v versealx-var:/var/lib/versealx-server \
  <versealx-server image> init --hostname mail.example.com --domain example.com --self-signed

docker run -d --name versealx \
  -p 25:25 -p 465:465 -p 587:587 -p 143:143 -p 993:993 -p 110:110 -p 995:995 -p 4190:4190 -p 443:443 \
  -v versealx-etc:/etc/versealx-server -v versealx-var:/var/lib/versealx-server \
  <versealx-server image>

The first command prints the administrator’s password once. Run other commands inside the running container, for example:

docker exec versealx /versealx-server admin get tenants
docker exec versealx /versealx-server doctor

Stop the container before backup and restore, and run them in a new container with the same volumes.

Several nodes on one store

Nodes that point at the same PostgreSQL database and bucket behave as one installation. You can run every role on each, or give nodes different roles — inbound edge nodes running mx, relay and filter, and mailbox nodes running store, deliver, dav, serve and admin, for example.

What every node must share

  • The same [store] and [blobs]. With a filesystem blob store, that means one shared filesystem, such as NFS, EFS or CephFS, mounted at the same path on every node. Each node leaves a small mark at the top of the blob directory when it starts. A node that can’t see the marks of the other running nodes refuses to start and names the node it can’t see, because the two would each write mail the other can’t open.
  • The same key file. Copy it to every node. A node with a different key cannot read any message.
  • A unique node name. In an orchestrator, node = "${HOSTNAME}" gives each replica its own.
  • The same hostname on nodes that answer JMAP or OAuth behind one name. A sign-in token names the host that issued it, and a node with another hostname refuses it.

What the nodes coordinate through the store

ConcernHow it works
The queueA node leases an entry for 15 minutes while working on it; if the node stops, another takes the entry over.
PacingEach destination’s allowance is shared, so the nodes together stay within one ceiling.
Hosted domainsEvery node rereads the domain list every second, so a domain added on one node is local on all of them.
PushWith PostgreSQL, a delivery on one node wakes IMAP IDLE sessions and JMAP push connections on the others.
CertificatesAn ACME certificate one node obtains is kept in the store and served by every node; challenge answers are shared too. See TLS certificates.
Daily reportsOne node takes a lease for the day and sends the DMARC and TLS reports.
SettingsRuntime settings live in the store, so every node reads the same values.

Domain ownership checks may run on more than one node at once; that costs a few extra DNS queries and nothing else.

Nodes of a cluster

The first node to start on an empty store gives the cluster its identity: a random id, and the name from [cluster] name (the host name unless set). A node whose configuration names a different cluster from its store refuses to start, naming both. Every node then keeps a record of itself in the store: its roles, release, the listeners it bound, its state, and a fingerprint of its configuration (never the configuration itself). It also renews a heartbeat every 10 seconds. A second node started under a name whose heartbeat is live refuses to start.

versealx-server cluster status
versealx-server cluster nodes
versealx-server cluster node node-2

A node not heard from for 30 seconds is missing, and the server’s node-missing alert goes off. After 10 minutes it is gone. cluster nodes marks a node whose release, store format or configuration differs from the others, and lists which node holds each shared job now. A gone node’s record stays until you remove it with cluster forget node-3; a node that is live or missing can’t be forgotten.

cluster node node-2 also shows how many sessions the node has open on each of its mail listeners, such as SMTP, IMAP and POP3, as of its last heartbeat. The console’s Cluster page shows the same in its Sessions open column.

To take a node out of service before maintenance, drain it:

versealx-server cluster drain node-2
versealx-server cluster undrain node-2

Within 10 seconds the drained node’s /readyz answers 503, so a load balancer sends it no new connections. A connection that reaches it anyway is turned away, and mail apps and senders try another node:

ProtocolNew connections are told
SMTP421 4.3.2, try again later
IMAP* BYE
POP3-ERR [SYS/TEMP]
ManageSieveBYE (TRYLATER)

Connections already open get 30 seconds to finish, then each is closed where nothing is lost: an IMAP session while idle or between commands, an SMTP session between messages, with a message being sent answered first, and POP3 and ManageSieve between commands. Mail apps reconnect, and reach another node. Set the time with drain_grace_seconds under [cluster]; 0 leaves open connections alone, and the node only turns new ones away.

cluster status, cluster node node-2 and the console’s Topology page show how many connections a drained node still has open, so you know when it is empty. undrain puts the node back in service, and so does restarting it. The operator sees the same at GET /api/v1/cluster/nodes.

Sites

The console’s Topology page draws the sites with their nodes, store and blob store, the links between them with each link’s certificate expiry, and the routes beyond the cluster: the edge passing mail in, a premises route, and the relay out. A node that describes the sites differently from the others is marked Sites differ, and a standby outside an organisation’s residency is marked too. Each link and route also says when mail last passed over it, as the newest time any node saw, kept when a node restarts; one no node has used says nothing. Every part of the drawing is reached with the keyboard and said in words.

Say where each node, its store and its blob store are, so the server can answer where an organisation’s mail is. Describe each site once, in the configuration every node shares, and say where this node’s parts are:

[sites.frankfurt-a]
kind = "cloud"
provider = "aws"
region = "eu-central-1"
zone = "eu-central-1a"
country = "DE"

[sites.head-office]
kind = "on-premises"
place = "Head office, Hamburg"
country = "DE"

[placement]
node = "frankfurt-a"
store = "frankfurt-a"
blobs = "head-office"

check refuses:

  • a cloud site without a provider and region;
  • a kind other than cloud or on-premises;
  • a country that isn’t an ISO 3166 code;
  • a placement naming a site that no [sites] table describes.

Each node publishes its sites. versealx-server cluster sites (and GET /api/v1/cluster/sites) lists each site with its nodes and what rests there.

doctor compares the declaration with the machine’s own firmware (the DMI vendor strings on Linux). A node declared on premises on a machine that reports Amazon EC2, or declared on one cloud on a machine that reports another, is a warning naming both. A hypervisor you run yourself, such as VMware, Proxmox, Hyper-V or QEMU, is consistent with on premises. The check asks no network. It only warns, because firmware on your own hardware can say anything.

Upgrading a cluster, one node at a time

versealx-server cluster upgrade --to 1.6.0 --channel https://downloads.example.org/versealx

The upgrade first checks the cluster can go to that release, and refuses with every reason at once:

  • a node on a release it doesn’t upgrade from;
  • a store format it doesn’t read;
  • no good backup from the last day, unless you add --without-backup;
  • a node that is missing, since it could come back on the old release.

Then it upgrades one node at a time:

  1. It drains the node and waits for its connections to close.
  2. It asks the node to upgrade itself.
  3. It waits for the node to come back serving on the new release.
  4. It undrains the node and goes on to the next one.

Before the first node runs the new release, the store is prepared for it: whatever the release changes that somebody could notice is kept as it was for each organisation until it decides. See Upgrades and the changes they bring. On Kubernetes this is the chart’s hook, so run helm upgrade before cluster upgrade.

A node that fails stops the upgrade there, and the nodes after it are left alone. cluster upgrade --resume tries that node again, and --abort undrains it and ends the upgrade. cluster upgrade status says where it stands, from any node.

Installed from a package, a node upgrades through versealx-server upgrade, which its service manager runs. On Kubernetes, with the chart, each role group is a StatefulSet whose pods are replaced only when deleted. A node whose turn it is sets its StatefulSet’s image to the new release and deletes its own pod, and Kubernetes makes it again under the same name on the new release. A StatefulSet set to replace pods on its own schedule is refused, since it would replace every pod at once without draining them. The chart gives the pods a service account that may read and change their StatefulSets’ image and delete their pods, and nothing else.

The store’s format

Every node records itself in the store every 30 seconds, with its release and the store formats it reads. A node that starts on a store written in a format it cannot read refuses to start, and names both.

versealx-server cluster upgrade status

This lists every node, its release, the formats it reads, and whether it is live, missing (unseen for about a minute and a half) or gone (unseen for a day). It ends with the cluster’s position, such as “every node reads format 2; the cluster still writes format 1”.

When a release’s notes say it brings a new store format, upgrade every node first, then switch the cluster to it:

versealx-server cluster finalize --to 2

This is refused while a live node cannot read the new format, or while any node is missing, since a missing node might come back on the old release. Once it is accepted, the older records are converted a page at a time, and cluster upgrade status shows the progress. Going back to a release that cannot read the new format is then impossible. The operator sees the same at GET /api/v1/cluster/upgrade and POST /api/v1/cluster/finalize.

Behind a load balancer

  • For SMTP, turn on proxy_protocol on [listeners.mx], [listeners.submission] and [listeners.submissions], and configure the load balancer to send the PROXY header, so connection limits and recorded addresses use the real client.
  • IMAP, POP3 and ManageSieve do not read the PROXY header; balance them at the TCP level without it.
  • Send HTTPS to any node running the roles its paths need: /api/ to admin, /jmap/, /oauth/ and /.well-known/jmap to store, /dav/ and /.well-known/caldav to dav, and autoconfig, Autodiscover and MTA-STS to serve.
  • Point /readyz health checks at each node’s metrics address.

Each node has its own admin socket; run versealx-server admin on any node.

Moving the store and the mail

A node starts on SQLite and a disk. To grow past that, move the store to PostgreSQL and the mail to a bucket while the node keeps serving:

versealx-server store move --to postgres://versealx@db.internal/versealx --live --apply
versealx-server blobs move --to s3://versealx-mail --region eu-central-1 --apply

The store. store move copies every row as it is stored, so sealed rows stay sealed. Its target must be empty. With --live, it copies while mail flows. Then it holds the whole installation still: new mail is answered “try again later”, and writes wait. It copies what changed meanwhile and compares row counts and a hash for each kind of record on both sides. It estimates the hold from the first copy’s speed and refuses to hold for longer than --max-hold (60 seconds unless you say). When it finishes, restart the node on the new store; --apply writes the new [store] section into the configuration file and keeps the old file beside it. A move that stops part way carries on with --resume, or is undone with --abandon. store move status says how far it got. Moving from PostgreSQL back to SQLite is the same command.

The mail. blobs move copies each message as stored, still encrypted, skipping those already there. It then reads every one back from the target and checks it before giving the new [blobs] section. After you restart on the new place, --delete-source removes the old copies, and refuses to until a check has found every one. Mail that arrived during the copy is picked up by running it again with --from the old place.

To move without stopping, start the node on the new place with the old one named as previous, in the same form --from takes. Every message opens, whichever place holds it: the node reads the new place first and the old one only when a message isn’t there yet. New mail is written to the new place only, and a deleted message goes from both.

[blobs]
kind = "s3"
bucket = "versealx-mail"
region = "eu-central-1"
previous = "/var/lib/versealx-server/blobs"

Then run blobs move --from the old place until it says nothing is left only there, remove previous, restart, and run --delete-source. doctor says a move is still being read through, and warns if the old place stops answering.

The operator sees how the last moves stand at GET /api/v1/storage/moves.

Old mail on a cheaper disk

Most mail is read in its first weeks and rarely after. On a node that keeps its mail on a disk, name a second, cheaper disk for old mail:

[blobs]
kind = "filesystem"
path = "/var/lib/versealx-server/blobs"
cold = "/srv/cold/versealx"
cold_after_days = 90
price_per_gb_month = { main = 0.10, cold = 0.02 }

The server moves messages older than cold_after_days there, a little at a time. Each one moves as stored, still encrypted, and is read back and checked before the copy on the main disk is removed, so a move that stops part way loses nothing. People and apps notice nothing: a message is read from wherever it is, and new mail is always written to the main disk.

versealx-server storage status shows how much is on each disk, when the last move ran, and, with prices given, what each disk costs a month. The operator’s Cluster page in the console shows the same. doctor warns when the cold directory is missing or not writable, or is on the same disk as the main one, where it would save nothing.

On a node whose mail is in a bucket, use the bucket’s own lifecycle rules to move old objects to a cheaper storage class instead, and choose one that can be read at once, such as an infrequent-access class. A class that needs a restore request first would leave old mail unopenable for hours. versealx-server doctor reads the bucket’s lifecycle rules: a rule that moves mail to Glacier Flexible Retrieval or Deep Archive, or that expires objects, is reported as broken, naming the rule.

Choosing where the data lives

Every data layer option is offered:

  • The store: SQLite on one node, PostgreSQL you run, or managed PostgreSQL.
  • The mail: a local disk, or any S3-compatible bucket.
  • The key-encryption key: a file, or a key service.

What differs is the risk. An option that cannot work is refused, with the reason. An option that works with less safety has a named risk, and you accept it in writing before the node runs on it.

Qualifying your own database

Before trusting mail to a database, ask it every requirement the store has:

versealx-server store test postgres://versealx@db.internal/candidate

The candidate must be empty. The report names each requirement and whether it passed:

  • reading its own writes;
  • rolling back;
  • ordered, bounded scans;
  • values of 4 MB;
  • two conflicting transactions not both kept, which serializable transactions guarantee;
  • leases held by one node at a time;
  • for PostgreSQL, notifications between connections.

A failed requirement the store can’t work without means the database can’t be used. A failed notification check is the notify-missing risk. The checks remove their own rows afterwards. A SQLite file works the same way: versealx-server store test /srv/candidate.sqlite.

Qualifying a blob store

Before moving mail to a bucket or a disk, ask it every requirement the node has of a blob store:

versealx-server blobs test s3://versealx-mail --region eu-central-1

The report names each requirement and whether it passed:

  • reading its own writes, taking the same blob twice, and deleting;
  • answering a blob that was never written as absent, not as an error;
  • blobs of 32 MB;
  • sixteen blobs written at once.

The checks write a few blobs of their own and remove them afterwards. A directory works the same way: versealx-server blobs test /srv/candidate-blobs.

Risks

RiskWhat it risks
blobs-across-sitesThe blob store is at another site from the node, so every message opened crosses the link between them.
key-single-copyThe key-encryption key is a file on the node. If that is its only copy, losing the disk loses every message: keep a copy apart, or use a key service.
notify-missingThe database doesn’t deliver notifications between nodes, so a change reaches clients on another node only at its next poll.
dr-writes-in-turnThe site takes part in disaster recovery on PostgreSQL. While it is the primary and a standby copies it, transactions that write run one at a time. PostgreSQL’s own streaming replication to the standby site doesn’t slow writes.
dns-changed-by-server[dr.dns] changes lets the server rewrite the record mail is pointed at when it is promoted, so whoever breaks into it could rewrite the domain’s DNS too.
dr-automatic-failover[dr.failover] automatic lets the standby promote itself and the primary stop taking changes on a witness’s word, so the witness’s availability becomes the primary’s.

check refuses a configuration carrying blobs-across-sites, notify-missing, dr-writes-in-turn, dns-changed-by-server or dr-automatic-failover until you accept it, and prints the line that does:

[risks]
accept = ["blobs-across-sites"]

key-single-copy is reported by doctor on every run rather than refused, so a single node set up by init still starts. doctor lists every risk the node carries, accepted or not. Each node publishes the risks it accepted, and a node whose accepted risks differ from the others’ is marked in GET /api/v1/cluster/data-layer.

The console’s System › Risks page lists every risk, says whether check refuses it until accepted, and names the nodes that carry it and the nodes that accepted it. Over the API it is GET /api/v1/cluster/risks.

Planning a deployment

The Risks page also has a planner. Pick the deployment’s shape: in a cloud, on two premises, a cloud edge in front of premises, or one site. The planner lists each choice to make:

  • the store
  • where messages’ content is kept
  • the encryption key
  • how mail finds the site that leads
  • what happens when a site is lost
  • how upgrades roll

For each choice it gives every option, recommended one first, with the risk or cost of each, and marks the options your nodes use now. It changes nothing: what each node runs on is in its configuration. Over the API it is GET /api/v1/cluster/planner?type=cloud, with two-premises, hybrid or single for the other shapes.

Refusals can’t be accepted. SQLite on a network filesystem, such as NFS, SMB or CephFS, can’t work, because its write-ahead log needs shared memory that such a filesystem doesn’t provide. check refuses it. Put the file on a local disk, or use PostgreSQL.

A standby site

A standby at a second site keeps a copy of the primary’s store, and takes over when the primary is lost. Name each site’s part in its configuration:

# At the primary.
[dr]
role = "primary"
# At the standby.
[dr]
role = "standby"
primary = "https://mail.example.org"

primary is the address of the primary’s admin API. The standby reaches it as you: sign in to it once from the standby with versealx-server admin login --server https://mail.example.org.

Two things live outside the store, and the standby needs both:

  • The key-encryption key. Give the standby the same [keys]: the same key service, or a copy of the key file.
  • The mail itself. Copy the blob store with its own replication, such as S3 replication or MinIO site replication, and point the standby’s [blobs] at the copy. Or follow by the link, below, which copies it too.

Instead of signing in to the primary as yourself, the standby can follow it by the link between the two sites’ clusters. Link them, then at each site write the link’s files:

versealx-server link files <the other site> --dir /etc/versealx/links/<the other site>

and name that directory in [dr]. The primary serves the standby on a listener of its own:

# At the primary.
[dr]
role = "primary"
link_dir = "/etc/versealx/links/standby"

[listeners.dr]
bind = "0.0.0.0:7443"
# At the standby.
[dr]
role = "standby"
primary = "https://mail.example.org:7443"
link_dir = "/etc/versealx/links/primary"

The dr listener demands a certificate from the standby’s link authority and serves nobody else, and it answers disaster recovery and nothing else. The standby trusts only the primary’s link authority. dr follow, dr switchover and dr promote then go this way with no sign-in, and the audit log names the link. The dr listener needs an address no other listener uses.

Following by the link also copies the mail itself: each message the copied records refer to that the standby doesn’t hold is fetched, checked against its hash, and kept before the records that name it. The standby needs a blob store of its own.

It can take the primary’s key-encryption key the same way, when both keep it in a file:

versealx-server dr take-key

The standby makes a one-time key, kept only in memory, and the primary seals its key-encryption key to it, so only that standby can open what crosses. The standby writes it as its [keys] file, readable by its owner only, and refuses if the file already exists. The primary’s audit log records that the key was sent. A key held by a key service isn’t sent: give the standby the same [keys].

Keeping the copy

On the standby, run:

versealx-server dr follow

The first time, it copies every row the primary holds. The primary keeps taking mail while it does. After that it applies each change the primary commits, in order, and keeps running until stopped. Run it as a service at the standby. --once stops once it has caught up.

A standby’s store must start empty: dr follow refuses a store that already holds mail. A copy that was interrupted is taken away and started again. The standby’s own nodes don’t start while it is a standby.

The standby tells the primary which of your sites its store and blob store rest in, from its [placement]. The primary checks that against every organisation’s residency before copying anything. A standby outside an organisation’s countries, or one that names no site while any organisation has a residency, is refused, naming the organisation, the site and its country. To copy anyway, run dr follow --outside-residency; the organisations it breaks are recorded in the audit log. Each organisation’s Where your mail is then lists the standby’s countries too, and doctor warns about it at both sites.

versealx-server dr status, at either site, shows the site’s role and how far behind the standby is. The operator also sees it at GET /api/v1/cluster/dr.

Switching over on purpose

To move the primary’s part to the standby with nothing lost, for maintenance or to move sites, run this on the standby:

versealx-server dr switchover

The standby catches up while mail keeps flowing. The primary then pauses all changes for a few seconds, at most --max-hold (60 by default, up to 300). Mail senders are asked to try again shortly, and apps wait. The standby takes the last changes, the old primary is marked as replaced, and the standby becomes the primary. If the standby can’t catch up in time, the pause ends and nothing changes.

Afterwards, start the new primary’s nodes and point your mail’s names at them, then bring the old site back as the standby, as below.

Taking over

When the primary is lost, promote the standby:

versealx-server dr promote

The standby becomes the primary, and a promotion mark records it. If the old primary can be reached, it’s told at once. From then on it takes no change and none of its nodes starts. Start the new primary’s nodes with versealx-server run, then point your mail’s names (MX and client host names) at them.

If the old primary couldn’t be told, run this there before any of its nodes starts again, with the epoch dr promote printed:

versealx-server dr fence --epoch 2

The standby as a second MX

The standby site can take mail while it isn’t the leader, so senders who can’t reach the primary still hand their mail over. List it as the second MX record, then run on the standby:

versealx-server dr hold

with, at the standby:

[dr]
hold_spool = "/var/lib/versealx/held"
hold_bind = "0.0.0.0:25"

It takes mail only for addresses its copy knows, with STARTTLS from the certificate in [tls] (which must come from files), and keeps each message in the spool, not in its copy. When following by the link, it hands each message to the primary every half minute. The primary checks it as if it had come straight from its sender: its address, its authentication and the filter. When the standby is promoted, dr hold stops, and the site’s nodes, started with the same hold_spool, take whatever it still holds.

Taking over automatically

By default a promotion is yours to start. With a witness, the standby can take over by itself when the primary is gone, and the two sites never take mail at once. It is off unless you turn it on, and check refuses it until you accept its risk in writing:

  • A primary that can’t reach the witness for a lease’s length stops taking changes, even if nothing else is wrong.
  • The witness must be at a third place that each site reaches by its own path, or one fault can stop both.

Run the witness at that third place, with a link to each site and an address per site:

versealx-server witness serve --data /var/lib/versealx-witness \
  --site primary=/etc/versealx/links/primary@0.0.0.0:7444 \
  --site standby=/etc/versealx/links/standby@0.0.0.0:7445

Each address answers only the certificate of its own site’s link. Then at both sites:

[dr.failover]
automatic = true
witness = "https://witness.example.org:7444"   # this site's address at the witness
witness_link_dir = "/etc/versealx/links/witness"
lease_seconds = 30

[risks]
accept = ["dr-automatic-failover"]

The primary renews its lease at the witness every third of lease_seconds (from 10 to 300, 30 by default). When the lease runs out unrenewed, the primary stops taking changes: mail senders are asked to try again later and apps wait, until it is renewed. On the standby, run versealx-server dr follow --automatic. When the primary hasn’t answered for twice the lease’s length, the standby asks the witness for the lease at the next epoch. Only if the witness grants it does the standby promote itself, and it points the mail as [dr.dns] says. The old primary, when it can reach the witness again, is told it was promoted over and stays fenced.

To have a person confirm every takeover, start the witness with --approvals. The standby then waits, and logs the command to run at either site:

versealx-server dr approve --epoch 2

Rehearsing a takeover

On the standby, a drill promotes a copy of it and reads it back, timed, without touching mail:

versealx-server dr drill

The standby’s copy is copied into a scratch store beside its own (or where --into <file> says), the copy is promoted and read back as a starting node would read it, and the report says how long promoting took and which epoch and organisations it would hold. The standby keeps following its primary throughout, and the primary is not asked anything. The scratch store is deleted afterwards unless you add --keep.

Bringing the old site back

The old primary comes back as the new primary’s standby:

versealx-server dr step-down
versealx-server dr follow --server https://mail-b.example.org

dr step-down works only on a site that a newer promotion has fenced. dr follow then copies the new primary in full.

How mail finds the site that leads

Every node answers https://<node>/health/leader with 200 leader while its site takes mail, and 503 standby or 503 fenced while it doesn’t. A site with no standby always answers 200. The answer is about now and is never cached. A DNS provider’s health check asks it, beside port 25, to decide which site’s address to hand out.

Plan how mail reaches the leading site for your deployment:

versealx-server dr dns plan --type cloud --provider route53 --domain example.com --primary 192.0.2.10 --standby 198.51.100.20

--type is cloud, two-premises, hybrid or single. The plan lists every way, the one recommended for that shape first, with what each needs and risks, and the records the recommended one needs:

ShapeRecommendedHow it works
cloudProvider failover recordsThe name the MX points at is a failover record at your DNS provider, which checks each site’s port 25 and /health/leader and hands out the leading site’s address within a minute or two. --provider route53 prints the health checks and records as CDK; cloudflare, azure and google print what to set up there.
two-premisesA second MXBoth sites are MX records, the standby after the primary. A sender that can’t reach the primary tries the standby, which takes the mail from the moment it is promoted, with no DNS change.
hybridThe cloud edge in frontThe edge is the MX and holds mail while the premises behind it is down. A promotion changes only where the edge relays: its [relay_domains.mailboxes] route.
singleGuided changeThe only way for one site: after a promotion, change the record by hand and watch for it.

Every shape also offers guided change, and having the server change DNS itself.

Name the record and this site’s address in [dr.dns], and dr promote and dr switchover end by saying exactly what to point where. Then watch until resolvers hand out the new address:

versealx-server dr dns watch --name mx.example.com --expect 198.51.100.20
versealx-server dr dns check --domain example.com --leader 198.51.100.20

dr dns watch asks every 15 seconds for up to 15 minutes (--for <seconds> changes that) and says what it last saw if it gives up. dr dns check lists the domain’s MX records as resolvers answer them now, and says whether the first a sender tries leads to the leading site.

Having the server change DNS itself

With changes = "route53" in [dr.dns], a promotion changes the Route 53 record to this site’s address itself. It uses the AWS credentials of the role the node runs as, from its environment, its container or its instance.

Whoever breaks into the mail server could then rewrite your domain’s DNS: its mail, its web site, its proof of ownership. So it is a risk check refuses until you accept it in writing, and never the default. Give the node credentials that can change that one record and nothing else. If Route 53 refuses, the promotion says why and what to change by hand.

Something unclear or out of date on this page? Tell us.