Monitoring

Watching Nixt Server — Prometheus metrics, health and readiness endpoints, logs, the doctor command and its every check, and what to alert on.

Nixt Server gives you four ways to see how it is doing: Prometheus metrics, health and readiness endpoints for your load balancer or orchestrator, logs, and doctor, which checks the things that stop mail arriving or being believed.

The metrics and health endpoint

The endpoint is off until you give it an address:

[telemetry]
metrics_bind = "127.0.0.1:9187"

It serves plain HTTP with no authentication, so bind it to a loopback or private address only. Without metrics_bind, metrics are still counted but not served, and the log says no metrics endpoint configured; metrics are kept but not served.

PathAnswer
/metricsEvery metric, in the Prometheus text format.
/healthz200 ok whenever the process is running. Use it for liveness.
/readyz200 when every readiness check passes, 503 otherwise, with one line per check. Use it for readiness.
Anything else404 This endpoint serves /metrics, /healthz and /readyz.

Readiness

/readyz has one check, store, sampled every 15 seconds:

BodyMeaning
store: okThe store answered the last read.
store: not ready (not checked yet)The node started moments ago.
store: not ready (<error>)The store did not answer; the log says the store did not answer a health read.

Metrics

No metric label ever carries an address, a mailbox or anything else belonging to one person.

MetricTypeLabelsMeaning
vsx_build_infogaugeversion, nodeAlways 1. Tells a dashboard which version each node runs.
vsx_messages_accepted_totalcounterrole: mx, submission, jmapMessages accepted onto the queue.
vsx_messages_accepted_bytes_totalcounterroleBytes of those messages.
vsx_messages_refused_totalcounterrole, reasonMessages refused, counted as role="mx", reason="filter" for the filter’s refusals.
vsx_authentications_totalcounterprotocol: submission, imap, pop3, jmap, managesieve, dav; outcome: ok, failed, unavailable, refusedSign-in attempts. unavailable is a sign-in that could not be decided because the directory it checks against could not be reached; refused is a right password or token the access rules kept out.
vsx_sending_refusals_totalcounterpolicy: send-as, sending-limit, sending-ceiling, unproved-domainMessages the organisation’s sending policies refused, each also in the audit log.
vsx_sessions_closed_totalcounterlistener: mx, submission, submissions, imap, imaps, pop3, pop3s, managesieve; reason: harvesting, network-refused, per-address, per-network, address-rate, network-rateConnections closed for guessing at addresses, connections refused before their greeting because their network kept doing so, and connections over a listener’s ceilings.
vsx_filter_stages_cut_totalcounterstageScoring stages the filter’s time budget cut before a message was accepted, and left for delivery to finish.
vsx_filter_verdicts_totalcounterverdict: deliver, tag, quarantine, hold, reject, tempfail; stage: the filter stage that decided, or thresholds when the score didMessages arriving on port 25, by what the filter decided and which stage decided it.
vsx_outbound_recipients_totalcounteroutcome: delivered, bounced, deferred; reply: 2xx, 4xx, 5xx, none; provider: google.com, outlook.com, yahoo.com, icloud.com or otherEach recipient of each delivery attempt, by what it came to and the kind of reply that decided it. Mail that could not be sent at all, such as to a domain with no mail server, is counted too.
vsx_outbound_connections_totalcountertls: verified, unverified, none; providerConnections opened to other mail servers, by how they were encrypted.
vsx_connections_opengaugelistener: mx, submission, submissions, imap, imaps, pop3, pop3s, managesieve, httpsConnections each listener has open now. https counts JMAP, CalDAV, CardDAV and the console together, since they share a port.
vsx_store_transaction_secondshistogramkind: transact, snapshot; outcome: ok, conflict, errorHow long the database’s transactions take, as the server waited for them, retries included.
vsx_queue_depthgaugequeue: transport, deliveryEntries waiting, sampled every 15 seconds.
vsx_queue_oldest_secondsgaugequeueAge in seconds of the entry that has been due longest, sampled every 15 seconds.

A metric appears once something has been counted in it; a new node may show only some of them.

# HELP vsx_queue_depth Entries waiting in a queue.
# TYPE vsx_queue_depth gauge
vsx_queue_depth{queue="transport"} 3
vsx_queue_depth{queue="delivery"} 0

What to alert on

AlertSuggested ruleWhy
Mail is stuckvsx_queue_oldest_seconds{queue="transport"} > 3600 for 15 minutesA deep queue that moves is healthy; one whose oldest entry keeps ageing is not.
Delivery into mailboxes is stuckvsx_queue_oldest_seconds{queue="delivery"} > 300No node is running the deliver role, or delivery is failing.
Node not ready/readyz not 200The store is unreachable.
Password guessingA sharp rise in rate(vsx_authentications_total{outcome="failed"}[5m])Someone is guessing; lockout is working, but you want to know.
Someone is harvesting addressesAny vsx_sessions_closed_total{reason="network-refused"}A network kept guessing at your addresses and is refused for an hour. The audit log’s receiving lines name it.
A filter stage is slowA steady rise in rate(vsx_filter_stages_cut_total[15m])A blocklist or a resolver answers slowly, and scoring is being finished at delivery.
Mail to a provider is failingA rise in rate(vsx_outbound_recipients_total{outcome="deferred",provider="google.com"}[15m]) against its delivered rateA provider is putting your mail off. Sending and delivery explains its pace.
Mail leaving unencryptedAny rise in vsx_outbound_connections_total{tls="none"}Some receiving servers offer no encryption. A rise for a named provider is worth a look.
Directory unreachableAny vsx_authentications_total{outcome="unavailable"}People who sign in with their directory password cannot until it answers. Check the directory and the network to it.
Something needs fixingversealx-server doctor exits with 2See below. Run it every hour or so.

Alerts to an organisation’s administrators

An organisation’s administrators can be told by mail when something needs them. Each rule names what to watch, the number at which to go off, and whom to tell:

WatchingCountsGoes off at
queue-backlogThe organisation’s messages waiting to be deliveredA number from 1 to 100,000
dmarc-failuresMessages in the organisation’s name that failed DMARC, in the reports received for the last dayA number from 1 to 1,000,000,000
sign-in-failuresAccounts locked out, or collecting failed sign-ins nowA number from 1 to 100,000
dns-problemDomains whose last DNS check found something broken1
breached-passwordsAccounts whose password in use has appeared in a data breach and must be changed; see Passwords that have leakedA number from 1 to 100,000
passkey-went-backPasskeys whose signature counter went backwards in the last day, which is what a copied authenticator does; see Passkeys1
forwarding-requestsRequests to forward mail outside the organisation that wait for an administrator; see Forwarding outside the organisationA number from 1 to 100,000
sending-heldAn account’s sending put on hold for review because it sent unlike itself, each time it happens; see Sending held for review1
organisation-storageThe organisation’s mailboxes at or over its storage warning level, before its ceiling refuses mail for everybody; see The organisation’s storage1
reported-phishingA message somebody reported as phishing, or a later verdict found, and each one the server took back from everybody, each time it happens; see Reported phishing1
certificate-expiryA certificate for the organisation’s domains running out within 14 days without having been renewed, each certificate on its own, and again at 3 days, with why its last renewal failed1
blocklistedOne of the organisation’s domains on a blocklist the server looks at, each listing on its own, and again when it clears; see Blocklists1
mail-stuckThe organisation’s messages waiting more than four hours for a server that answers but will not take them, which usually means a policy block thereA number from 1 to 100,000
bounces-risingThe percent of the organisation’s messages sent out today that bounced for at least one recipient, once it has sent 20A percent from 1 to 100
approvalsEach request that starts waiting for a second administrator’s approval, and each time the server’s operator goes ahead without asking, as an emergency1
account-recoveredEach account recovered with its recovery address, when somebody forgot their password; see Forgotten passwords1
webhook-pausedEach webhook, or the security log export, paused after failing for a day1
break-glass-usedEach sign-in of a break-glass account1
below-strictThe organisation’s security controls that fall short of the Strict profile. Not among the rules an organisation starts withA number from 1 to 30
new-networksThe accounts that signed in from a network they hadn’t used in ninety days, in the last hour; see A sign-in from somewhere newA number from 1 to 100,000
not-meEach time somebody says a sign-in to their account was not them1
mail-damagedMessages the integrity scrub found lost or damaged: those it couldn’t repair, and those it restored from a backup in the last dayA number from 1 to 100,000
journal-delayedThe organisation’s journal reports that have not reached their archive after a dayA number from 1 to 100,000
lookalike-domainA newly found look-alike of the organisation’s domains, once per domain1

An organisation that has never saved its rules has thirteen, sent to its administrators: passkey-went-back, sending-held, organisation-storage, reported-phishing, approvals, account-recovered, webhook-paused, break-glass-used, not-me, journal-delayed, mail-damaged and lookalike-domain, each at 1, and new-networks at 20. An alert goes off when its number reaches the threshold, and not again until the number has fallen below it and reached it again. Each time, everybody the rule names is sent one plain-text message saying what was watched, the number now, the threshold, and where to look. Every alert is also kept in a history for 90 days.

Write the rules as a file and set them:

{
  "rules": [
    { "when": "queue-backlog", "threshold": 500, "notify": ["postmaster@example.com"] },
    { "when": "dns-problem", "threshold": 1, "notify": ["it@example.com", "on-call@pager.example.net"] }
  ]
}
vsx admin alerts set alerts.json
vsx admin alerts show
vsx admin alerts history

The server’s own alerts

The server has alerts of its own, for what only it can see. They are tenant 0’s, and only its operator sets them:

WatchingCountsGoes off at
certificate-expiryEvery certificate the server serves, its own and each domain’s1
blocklistedThe server’s sending addresses and every organisation’s domains, each listing on its own1
mail-stuckEvery organisation’s stuck messagesA number from 1 to 100,000
disk-fillingEach volume a node keeps its store or its mail on, at this percent full or moreA percent from 50 to 99
bounces-risingThe percent of every organisation’s messages sent out today that bouncedA percent from 1 to 100
provider-throttlingA receiving provider (Google, Microsoft, Yahoo, iCloud) asking the server to slow down for an hour or more1
backup-failedThe node’s daily backup failing, or none read back whole in 26 hours1
slow-scannerRuns of a scanner or milter that took more than half the filter’s time budget, in the last hourA number from 1 to 100,000
backup-test-failedThe weekly restore test of the newest backup failing1
node-missingA node of the cluster not heard from for 30 seconds, each time it goes missing; see Nodes of a cluster1
probation-keptA new organisation’s probation kept past its days because too much of its outside mail bounced, once for each organisation1

disk-filling, provider-throttling, backup-failed, backup-test-failed, node-missing, probation-kept and slow-scanner are the server’s alone: an organisation that asks for one is refused. Until the operator saves otherwise, the server has all eleven: disk-filling at 85, bounces-rising at 5, slow-scanner at 10, and the rest at 1. A node that has never been set up to take daily backups is not told about them. They tell nobody by mail until the operator names whom, and each is logged as a warning, and kept in the history, either way. Set them over the local socket, as tenant 0:

vsx admin alerts show --tenant 0
vsx admin alerts set server-alerts.json --tenant 0

The server’s alerts go every way an organisation’s go. As tenant 0, the operator can add a webhook that takes alert events, or export the security log to a SIEM. Each server alert is then sent there as alert.raised, even when it tells nobody by mail. On the console, the operator does this from the Alerts and Webhooks pages. From the command line:

vsx admin webhooks add https://pager.example.com/hooks/mail --events alert --description 'On call' --tenant 0

Blocklists

Once an hour, one node looks the server up on blocklists: each of its sending addresses (what its host name resolves to, leaving out private ones) on the lists of addresses, and each organisation’s domains, up to 1,000, on the lists of domains. The lists are the filter’s own blocklists, leaving out its allowlists. Where the filter names no list of a kind, Spamhaus ZEN is used for addresses and Spamhaus DBL for domains.

The lookups go through the node’s own resolver, as the lists’ terms require. A list asked through a public resolver refuses to answer; a refusal is never counted as a listing, and doctor says which list refused and why.

A rule may name up to 10 addresses, and an organisation may have up to 20 rules. A rule with no addresses is kept in the history only. Administrators set the rules; auditors can see them and the history. Over the API they are GET and PUT /api/v1/tenants/{tenant}/alerts, the PUT with the version the GET gave, and GET /api/v1/tenants/{tenant}/alerts/history.

What is stored

Once a day, one node counts what each organisation stores, by kind:

  • message content (each message counted once, however many people hold it)
  • message records
  • calendars and contacts
  • quarantine
  • message traces
  • the audit log
  • reports
  • change logs

The count is paced so that it never competes with mail. Each node also records how full its disks are once a day. From 30 days of those records, the server forecasts the day each disk will be full, when it is growing.

versealx-server storage status
versealx-server storage organisations
versealx-server storage kinds --tenant 3

The operator reads the same on the console’s Storage page, beside Usage, and through GET /storage, GET /storage/organisations and GET /storage/kinds. An organisation’s administrators and auditors see their own counts on the What is stored card on the Usage page, with versealx-server admin usage breakdown, or through GET /tenants/{tenant}/storage/breakdown. Counts are kept for 400 days.

When a disk fills

A node never damages mail or bounces it because a disk filled. When a volume holding the store or the blobs reaches the reserve ([storage] reserve_percent, 95% unless you change it), the node stops taking new mail. MAIL FROM on port 25, on submission and from your other premises gets this reply:

452 4.3.1 Insufficient system storage; try again later

That is a temporary refusal. Senders keep the message and try again, and mail apps keep it in their outbox. Everything else goes on. People can read, move and delete mail, and administrators can make room. As soon as the volume is back under the reserve, mail is accepted again without any action from you. The node checks its disks at most every ten seconds.

Well before the reserve, the server’s disk-filling alert tells you a volume is filling (85% unless you change it).

Logs

The server writes its log to standard output. The packaged service sends it to the systemd journal:

journalctl -u versealx-server -f
  • Lines are human-readable, with a time, a level and named fields.
  • The level is info unless you set the RUST_LOG environment variable, for example RUST_LOG=debug in a systemd drop-in.
  • Every line about a message carries its queue id, and no line ever carries what a message says. Relay passwords are never written.

Lines worth knowing

LineMeaning
mx listening, submission listening, imap listening, pop3 listening, managesieve listeningA listener is bound, with its address.
https listenerThe shared HTTPS listener is bound, with which services it carries.
local socket listeningThe admin socket is ready.
metrics and health listeningThe metrics endpoint is bound.
relay worker started, deliver worker startedThe queue workers are running.
generated a new key-encryption key; back it upFirst start: copy the key file now. See Backup and restore.
acceptedA message was accepted, with its size, recipients and role.
filteredThe filter’s verdict, the deciding stage and the score.
authenticated, authentication failedA submission sign-in, with the mechanism.
no DKIM key; sending unsignedA domain has no signing keys.
certificate obtained, could not obtain a certificate; will try again, serving the certificate another node obtainedACME.
domain proved, domain no longer proves itself; the record it was proved by has goneDomain ownership changed.
yesterday's reports sent, could not send yesterday's reportsThe daily DMARC and TLS reports.
old message traces forgottenTrace retention ran.
sieve script failed; keepingA person’s script failed; the message went to their inbox.
address is not public; skippedA destination resolved to a private address.
the serve role is on but this node runs neither the store nor the submission role; …Clients would be told to connect to services this node does not run.
shutting downThe node received a stop signal.

explain

versealx-server explain prints what a node’s configuration will do without starting it: each role, what it listens on and what it sends out of the machine; the outbound pacing; the message ceilings; where data lives; the filter pipeline; and what leaves the premises — certificate requests, metrics and whether the node may deliver to private addresses. See the Command-line reference.

doctor

doctor checks a deployment from the outside in and tells you what to do about each problem. It reads DNS, the certificate, the store and the clock, and changes nothing.

sudo -u versealx versealx-server doctor --config /etc/versealx-server/versealx-server.toml
ok   hostname: mail.example.com
ok   tls certificate: /var/lib/versealx-server/tls/certificate.pem
ok   tls key: /var/lib/versealx-server/tls/key.pem
ok   clock: reads 2026
ok   resolver: passes DNSSEC through; a name that does not exist can be proved not to and the answer kept
ok   reverse DNS: mail.example.com and its 1 address(es) confirm each other
ok   DANE: no _25._tcp.mail.example.com record, so nothing is promised about this host's key
ok   store: answers; 1 domain(s) configured
ok   example.com MX: names this host (mail.example.com)
BAD  example.com SPF: does not name the relay this node sends through: v=spf1 mx -all
         mail leaves from relay.example.net, not from this domain's MX, so every message fails SPF and then DMARC. Publish `example.com. IN TXT "v=spf1 mx include:relay.example.net -all"`
ok   example.com ownership: _versealx-verify.example.com is published, so autoconfig and MTA-STS are answered
ok   example.com DMARC: v=DMARC1; p=none; rua=mailto:dmarc@example.com
…

16 checked, 15 good, 0 to look at, 1 broken

Each line starts with ok, warn (works, but will not keep working, or is not what you meant) or BAD (mail will be lost or refused). A problem’s remedy is on the indented line below it.

Exit codes

CodeMeaning
0Everything is ok.
1Something is warn.
2Something is BAD, or doctor could not run — for example the resolver could not start: <reason>, or the configuration could not be read.

The checks

CheckWhat it looks at
hostnameA fully qualified name. warn for a single-label name or one ending in .localhost.
tls certificate, tls keyWith certificate files: both exist. BAD when one is missing.
tlsWith ACME: a certificate has been obtained and has more than 30 days left. See TLS certificates.
tls renewalFor each certificate whose last order failed: when, and what the CA or the network said. warn.
clockThe clock reads 2024 or later.
resolverWhether your resolver passes DNSSEC records through, by asking for a name that cannot exist. warn when it does not, times out, or invents an answer.
reverse DNSThe host name resolves; each of its first two addresses has a PTR; the PTR resolves back. See DNS records.
DANEAny TLSA record for the host matches the certificate in use, and the state of a key rollover. See DNS records.
storeThe store opens and answers, and how many domains it holds. BAD when it does not; the DNS checks that need no domain still run.
domainswarn when there are none.
leaked-password checkWhether the leaked-password service answers, and how quickly, or that this node does not ask (breach_check = false).
confinementWhether the kernel enforces the node’s confinement: Landlock, version <n>, and anything this kernel cannot confine. BAD when landlock = "required" and the kernel does not enforce it, so run would refuse to start.
key-encryption keyWhen the key was last rotated, or that it never has been. BAD when a rotation is half done: no node starts until the same keys rotate command is run again.
<domain> MXThe MX names this host.
<domain> SPFA record exists and agrees with whether this node sends through a relay.
<domain> ownershipThe _versealx-verify token is published.
<domain> DMARCA DMARC record exists.
<domain> DKIM <selector>Each signing key’s record is published and matches the key in the store.
<domain> MTA-STSThe advertised policy and the policy this node serves agree, and every MX host is covered.
<domain> TLS reportingA TLS reporting record exists.
blocklistThe server’s sending addresses and its domains, looked up on the blocklists as the hourly look does. BAD for each listing, with the list’s page for asking to be removed; warn for a list that did not answer.

A lookup that fails rather than finding nothing is reported as warn with check the resolver before reading anything else here.

Something unclear or out of date on this page? Tell us.