Byron Coke

Linux infrastructure engineer. Production estates and platform migrations for banking, government and telecoms.

contracts@paxmentis.com ·Remote, UK ·Available at four weeks' notice

Thirty years in IT, more than twenty-five on Unix and Linux. Red Hat and Ubuntu at production scale, containerised platforms, high availability and disaster recovery, Ansible, and the migrations that move an estate from one to the other.

RHEL · Ubuntu · Docker · Kubernetes · Ansible · PostgreSQL · AWS · Azure · VMware · Nginx · Caddy · HAProxy · Prometheus · Grafana · ELK · TLS and PKI · SPF, DKIM and DMARC

A recent diagnosis, written up in full

Post-mortem

Half a million delivery attempts, nothing ever sent, and a test that passed

534,686 delivery attempts. Zero messages delivered. Every dashboard reported the system healthy.

A self-hosted Postal mail server, running in Docker on a VPS, had been accepting mail and queueing it normally for months. Messages went in. The queue looked healthy. Nothing came out the other side, and nobody noticed, because the layer that was broken was not the layer being watched.

Every attempt in the database carried the same failure:

Resolv::ResolvError

The test that made it worse

The obvious check is to resolve a domain from inside the container. I did, in Ruby, the same library Postal uses. It returned in 0.03 seconds.

That result was a false negative, and it cost real time. It had already produced a plausible and completely wrong theory, that name resolution was failing under concurrency and succeeding when tested by hand.

The manual lookup read /etc/resolv.conf. The worker read the resolver configured in Postal's own config file. Two different files. The diagnostic and the failure were never looking at the same thing.

The cause

Postal's configured resolver was 10.0.1.1. That address is the gateway of a Docker bridge network, br-e538c46466eb. It is a routable address, so nothing errors immediately, and it answers on no port at all. Nothing listens on 53.

Every MX lookup the mail server had ever made timed out against a gateway that was never a resolver. The config file was dated 8 May. The failures ran from then until September.

The fix

What I take from it