Linux infrastructure engineer. Production estates and platform migrations for banking, government and telecoms.
Thirty years in IT, more than twenty-five on Unix and Linux. Red Hat and Ubuntu at production scale, containerised platforms, high availability and disaster recovery, Ansible, and the migrations that move an estate from one to the other.
RHEL · Ubuntu · Docker · Kubernetes · Ansible · PostgreSQL · AWS · Azure · VMware · Nginx · Caddy · HAProxy · Prometheus · Grafana · ELK · TLS and PKI · SPF, DKIM and DMARC
Post-mortem
534,686 delivery attempts. Zero messages delivered. Every dashboard reported the system healthy.
A self-hosted Postal mail server, running in Docker on a VPS, had been accepting mail and queueing it normally for months. Messages went in. The queue looked healthy. Nothing came out the other side, and nobody noticed, because the layer that was broken was not the layer being watched.
Every attempt in the database carried the same failure:
Resolv::ResolvError
The obvious check is to resolve a domain from inside the container. I did, in Ruby, the same library Postal uses. It returned in 0.03 seconds.
That result was a false negative, and it cost real time. It had already produced a plausible and completely wrong theory, that name resolution was failing under concurrency and succeeding when tested by hand.
The manual lookup read /etc/resolv.conf. The worker read the resolver configured in
Postal's own config file. Two different files. The diagnostic and the failure were never
looking at the same thing.
Postal's configured resolver was 10.0.1.1. That address is the gateway of a Docker
bridge network, br-e538c46466eb. It is a routable address, so nothing errors
immediately, and it answers on no port at all. Nothing listens on 53.
Every MX lookup the mail server had ever made timed out against a gateway that was never a resolver. The config file was dated 8 May. The failures ran from then until September.
127.0.0.11, Docker's embedded DNS, which also resolves
container names, with 1.1.1.1 and 8.8.8.8 behind it.timeout:2 attempts:2, so a resolver failure fails fast instead of holding a
worker open.