AgentSee

Hove & London Founded 18 August 2026 Issue No. 2 · Free, obviously

We built our own mail server.
Here is what it cost.

Three afternoons, sixty-five commits, and an honest accounting of what owning your own mail buys you. It is not money.

01 — The premise

Owning your own mail is a reasonable thing to do again.

Two people, building in public. The pitch involves deploying client work as infrastructure-as-code, and that claim is worth much less if your own first server was clicked together in a dashboard.

The starting position was worse than nothing. Cloudflare Email Routing forwarded everything at our domain to personal inboxes and could not send at all. Every account signup, every reply to a guest, every password reset went out from a personal Gmail. The domain published a DMARC policy of reject, which was true: nothing was authorised to send as us, because nothing could.

The design that emerged was to own the inbox and rent the reputation. Stalwart on a VPS handles inbound and stores the mail. Everything outbound goes to a relay over authenticated submission. That removes the unwinnable part of self-hosting, which is that a fresh IP address has no sending history, and no history looks identical to a spammer who rented the address an hour ago.

Working with an AI made this three afternoons of actual work, spread over a fortnight because that is how evenings go. What it did not do is make it cheaper.

02 — The decisions

Two of the three choices that mattered had nothing to do with email.

Amazon SES is cheaper per message and was the first choice. It lost on a property unrelated to mail: new accounts start in a sandbox, and leaving it needs a support request reviewed by a human, with no API and no timebox. That put an unbounded wait on the critical path of everything else. SMTP2GO's free tier is a thousand messages a month with no card, which is more than two people write.

Hetzner was the obvious host and roughly £47 a year. Vultr was on the shortlist too, mostly for the spread of datacentre locations, until we read what happened to GreatFire. Vultr took their FreeWeChat project offline after a complaint instigated by Tencent, ignored a claim-by-claim rebuttal, ignored a letter co-signed by seventeen press-freedom organisations, then terminated the account without cause.

That put provider conduct on the table as a selection criterion, and once it was there Hetzner had its own problem, having terminated a non-profit that ran censorship-evasion services.

The criterion has two sides that have to be held together. A host should refuse to carry fascists and organised harassment of marginalised people. A host should also hold the line when a state or a corporation leans on legitimate journalism. Those are not in tension, and the second is not a licence for the first.

We went to Infomaniak: a certified B Corp, employee-owned, never took outside investment, runs on renewable energy and offsets twice what it uses. It costs about £15 to £30 a year more than Hetzner and buys nothing technical. That is a values purchase, and saying so plainly seems better than dressing it up as risk mitigation. Nothing we host will ever attract a takedown at that scale.

The whole thing runs at around £60 to £80 a year, against about £37 for hosted mail from the same company we chose as our host. Self-hosting is more expensive at two users. It wins on control and on learning.

Every failure looked like one thing and was another.

03 — What went wrong

The useful content is the failures, and they share a shape.

One missing package emptied the whole box. Cloud-init installs its package list atomically. One name that does not exist in Debian took out docker, restic, fail2ban and unattended-upgrades along with it. The instance came up. SSH worked. Files were written. The only evidence was in a status command nothing prompts you to run.

The setup wizard kept reappearing. Stalwart reads its configuration from a path inside the container; ours was mounted somewhere nothing reads. So it found no configuration, entered bootstrap mode, and the wizard ran to completion against storage that did not persist. Restart, fresh wizard, fresh password, as though nothing had happened. There is no error anywhere in that sequence.

The same class of bug appeared one layer down. The volume was mounted at one path while the server defaulted its datastore to another, which existed only inside the container's ephemeral layer. The database wrote happily to somewhere that vanished on every recreate.

The mail server banned the entire internet. Docker's published ports rewrite the source address, so Stalwart saw every external connection as coming from the same bridge gateway. It did what a mail server should do with an address making repeated half-open connections, and banned it. That address was everyone. We found out because our own monitoring got blocked; the unlucky version is one spammer taking inbound mail down for the world.

Certificates, three times. Our apex domain serves a static site through Cloudflare, so the usual challenge can never validate it, and Let's Encrypt fails an entire order if any single name fails. Then: leaving the certificate's alternative names empty does not mean none, it means a default set of four hostnames that do not exist. Then, days later, editing that list destroyed the working certificate without ordering a replacement.

Silence by default. None of the above was visible until we added a console logger, because the server writes nothing to standard output out of the box. The obvious alternative writes to a directory that does not exist in the container and fails on every startup. Every problem in that phase was invisible until logging worked. If there is one habit to take from this, it is to make a system able to tell you things before you need it to.

A silent fallback made every wrong guess identical. The outbound relay would not engage. The route existed, the configuration referenced it, and mail kept going direct. An unresolvable route name falls back with no warning and no log line, so four consecutive wrong hypotheses produced the same output. Reading the object with the command line instead of the web form broke the deadlock, and the same move later exposed a hostname with one letter wrong sitting in a field the interface had displayed without complaint.

The plan that differed by machine. One file, holding every value that decides what the infrastructure looks like, was excluded from version control on an assumption that had quietly stopped being true. It existed on exactly one laptop. A plan run from the second offered to destroy the mail records of a live server.

04 — Working with a model

It was good at some things and confidently wrong about others.

The model read widely and quickly, wrote configuration with the reasoning attached, recognised the class of a problem from a log line, and recorded why a decision was made while the reason was still known. Sixty-five commit messages that explain themselves turned out to be a better artefact than the configuration they describe.

It was also wrong in ways worth naming. It claimed the server generates Apple configuration profiles, which it does not; that came from documentation for a different product built on top of it. It moved a certificate-authority record to the apex, where Cloudflare silently drops it, having taken it from the place the tool had correctly put it. It asserted without checking that a service's subdomains survive disabling it. It wrote a backup restore test that could not have passed, and that would have re-sent the mail queue if it had.

Every one of those is a confident claim about how a system behaves that had not been tested. The things that went well have the opposite shape: probes calibrated against known-good and known-bad controls, reading objects rather than forms, checking that a fix had landed before concluding it had failed.

A division of labour emerged without being designed. The model proposes and writes. The human supplies anomaly detection, the "that looks wrong" that precedes any investigation, and the authority for anything irreversible. The near-miss with the destroyed records is the clearest case: someone noticed a plan looked odd and did not explain it away, and the investigation turned up a fault nobody was hunting for.

05 — What we have now

The assets outlast the afternoon that produced them.

Infrastructure that applies cleanly, with credentials resolved from a password manager at the moment of use. Nothing is exported into a shell and there are no plaintext credential files on disk.

A configuration snapshot taken from the running server by one command that refuses to write the file if it finds a secret value in it. The clicking happened once.

A runbook transcribed from what actually happened rather than from the design. Every trap above is in it, in the file where the next person will look.

Backups at two companies, nightly, alerting separately, with restores proven to boot. That included discovering the restore test itself was broken.

And understanding. DKIM alignment, certificate challenge types, transport security policies, why an SPF record can correctly authorise nobody at all. Bought expensively and not otherwise obtainable.

06 — Whether you should

Decide what would make you stop before you start.

Worth it if you want to understand the system, or if your positioning requires the receipts. Not worth it if you want cheaper mail.

What made it defensible was writing the abandonment criteria before starting: a failed restore test not fixed within a week, two silent delivery failures in a quarter, or it stops being interesting and becomes a chore. Any of those and we move to hosted mail.

And keeping that exit costed and current. Same addresses, same domain, about an hour. An escape route you have never priced is not an escape route.

Everything described here is in the open, including the mistakes. The repository is public and the commit messages are the honest version.