On-premises infrastructure often inherits a dangerous habit: a server becomes important because it has been around a long time. It receives a familiar name, a collection of exceptions, and eventually the instruction not to touch it because too many things depend on it.
That is a pet. It is also a poor fit for an authoritative DNS service.
Our DNS architecture has moved to Version 2. The goal was straightforward: keep DNS available while we replace a server, remove a cross-region database dependency from the serving path, and make each component’s job clear enough to monitor. We replace the server that answers queries, keep writes in one defined place, and prepare the rollback before the change.
Version 1 got us operating
Version 1 solved the immediate problem. It gave us an authoritative DNS service, a management interface, and a path for DNS changes. It also carried assumptions that were reasonable during the first build and less appropriate once the service became part of the platform itself.
The main assumption was that the remote serving site could reach the writable DNS database whenever it needed to. That joined DNS availability to a dependency across regions. A slow or unavailable link did not have to become a DNS outage, but it made diagnosis and recovery more complex than they needed to be.
Version 2 separates the responsibilities:
- The writable PowerDNS primary and the DNS management application, Poweradmin, run in YUL.
- The PowerDNS secondary in YOW maintains its own local database through authenticated zone transfers.
- dnsdist sits in front of DNS service in YOW and caches responses, reducing repeated work while PowerDNS remains the authoritative source for the answer.
- HAProxy provides a stable public DNS Manager endpoint while its backend follows the application.
Each role has a clear job. Operators make changes on the primary. The secondary serves its own copy independently. dnsdist handles the high-volume query path. HAProxy makes the web application’s public address independent from the application host.
Why the management application belongs with the writer
Poweradmin changes DNS data. It therefore belongs beside the writable primary in YUL, where its database and PowerDNS API dependencies are local to the write path.
That placement removes an avoidable remote dependency from a change request. A zone creation, record update, or zone deletion reaches the writable DNS system directly. The secondary then receives the approved zone data through the normal transfer mechanism. It does not need write access to the primary database and it does not need to behave like a second control plane.
The public DNS Manager address remains stable through HAProxy. Moving the application backend does not require people to find a new bookmark or update an integration. A legacy endpoint redirects to the current service, so users land in one place while the infrastructure behind it changes safely.
A secondary should keep serving when the primary is unavailable
The secondary in YOW has its own local PowerDNS database. That is the important change.
It pulls signed zones from the primary using authenticated transfers and serves them locally. If the connection to YUL is slow or unavailable, the secondary still has the last transferred copy of each zone. It continues to answer authoritative DNS queries from that local data.
dnsdist adds a second layer in front of PowerDNS at both sites, and it is worth being precise about what it does. Each dnsdist runs a packet cache configured with stale entries disabled, so once a cached answer’s TTL expires, dnsdist stops serving it and asks PowerDNS for a fresh one. That absorbs repeat queries for the same record within its TTL and reduces backend load during normal operation. It does not carry the service through a longer disruption, because once the TTL runs out there is nothing left in the cache to hand back. The local authoritative database on the secondary, not the cache in front of it, is what keeps answering through an extended outage.
The design gives us two useful properties:
- Queries in YOW do not depend on a live database connection to YUL.
- We prepare and test a replacement server while the current system continues serving DNS.
The pattern this follows
None of this is a private invention. A dedicated writable primary, secondaries that hold their own copy of the data and keep answering on their own, changes propagated by NOTIFY and pulled with AXFR or IXFR authenticated by a shared TSIG key: these are the standard building blocks of authoritative DNS, not something we designed from scratch for two regions. RFC 2182 states the reasoning behind putting secondaries where we put ours: secondary servers belong on different networks, in different locations, so that one failure cannot take out every server able to answer for a zone.
Large DNS providers apply the same separation at a much bigger scale. Cloudflare, Akamai, and NS1 all keep a write plane, the system that accepts and validates a zone change, apart from the edge nodes that answer queries, so that trouble on one side does not become trouble on the other. Several of them also run anycast, announcing the same address from many locations so the network routes a query to whichever one is closest and healthy. We do not run anycast. Our secondary answers from its own dedicated address in its own region, and resolvers fail over between our two nameserver addresses the ordinary way, by retrying the next server in the NS set when one does not answer. It is a smaller version of the same idea: a serving node independent enough to keep answering when the write side is unreachable.
The cutover method matters as much as the target design
We did not turn off the existing DNS server and hope the replacement was ready. We built the replacement in parallel, loaded and verified its local data, checked the serving path, confirmed DNSSEC answers still validated against the zone’s trust anchor from both nodes, and kept the old server available for rollback.
The replacement did not take over the old server’s private address. It has its own. What stayed in place was the public glue address registered for that nameserver, and the firewall’s NAT rule is what we repointed, from the old private address to the new one, once the replacement passed its checks. External systems never had to learn a new route, because from outside nothing about the address changed, only which internal host answered behind it. After the replacement passed its checks, we powered off the old server and kept it available for rollback. A clear pass or fail decision and a real rollback target are more valuable than a maintenance window that assumes every step will work.
This is the cattle principle in practice. The service identity remains. The machine underneath it is replaceable.
The problems we found were part of the work
The first design pass exposed issues we might have missed if we had treated this as a simple VM replacement.
A split-brain alert that was not split brain
Our database alert reported more than one writable MySQL node. The condition was real in the metrics but wrong in meaning. It counted the independent local database used by the DNS secondary alongside the primary database cluster.
Those databases have different roles. The local database is writable because the secondary needs to store transferred zone data. It is not a competing primary for the central write path. We corrected the alert rule so it evaluates the intended primary cluster rather than every database exporter in the estate.
An alert that is precise about the wrong scope is still an operational defect. It trains people to ignore a page that should command attention.
Configuration scope must match the service role
We also found configuration that had drifted from the system it described. A control key sat inline instead of in the vault with the service’s other protected values. A deployment setting followed an old topology rather than the current DNS role.
The corrections were small, but the rule is not: protected values belong in the same controlled configuration path as comparable values, and a playbook must target the role it is meant to manage. Infrastructure automation is only repeatable when the configuration describes today’s architecture, not the last one.
A dashboard needs to answer the operator’s question
Early monitoring made the local DNS database visible, but it did so in a separate view. That creates extra work during an incident because an operator has to remember which dashboard contains which part of a service.
We integrated the local database into the existing PowerDNS and MySQL dashboards and made the site selector consistent. Operators now select YUL, YOW, or both, then compare the primary and secondary without switching between unrelated pages. Alerts and logs should identify each role and site. Operators then see what is healthy, what is independent, and what needs attention.
A monitor should follow the active service, not the retired host
The DNS Manager monitor remains enabled because the service still matters. During the move, its backend temporarily followed a retired application location and the public endpoint returned an error. We corrected the HAProxy route to the active service and verified the public path again.
That distinction matters. Muting a monitor that watches a live service hides a regression. Muting alerts for a powered-off rollback VM is appropriate for the defined rollback period. Alerting should reflect the lifecycle of each component, not serve as a way to make a dashboard quiet.
Alert fatigue
The split-brain alert and the retired-host monitor were both false positives, and a false positive is not a rounding error. It is pager noise, and pager noise trains people. The split-brain page above was accurate about the metric and wrong about what it meant, and being technically correct about something meaningless does not make it harmless. Every time an alert like that fires and turns out to mean nothing, acknowledging the next page from that source without looking closely gets a little easier.
The same failure runs the other way. Muting a monitor because a component looks retired, which is what briefly happened to the DNS Manager check during the move, removes the one signal that would have caught the regression underneath it. Alert fatigue does not only look like ignoring pages. Sometimes it looks like deciding in advance that a monitor no longer needs to be believed.
Actionable alerting means every page corresponds to something a person should actually do. We fixed the alert rule and the routing instead of muting either one, because every false page teaches the team to ignore the next real one.
What Version 2 gives us
Version 2 is deliberately more ordinary than Version 1. That is a success.
- The write path has one home in YUL.
- The YOW secondary serves from local data.
- dnsdist reduces repeated backend work without becoming a second source of truth.
- Poweradmin is colocated with the writable DNS service.
- HAProxy keeps the public management address stable.
- Monitoring, logs, dashboards, and alerts distinguish the two sites and their roles.
- The team prepares, validates, cuts over, and rolls back a replacement through a controlled change.
None of this depends on a cloud provider. The cattle model applies anywhere. A VM in a private environment is replaceable when operators manage its configuration, data, identity, and monitoring as parts of a service.
Version 3: DNS hosting as a complete platform service
We are planning Version 3 before Q2 2027, tracked on our roadmap. It will move the complete DNS server offering and its zones through a full transition, rather than treating the nameserver layer as a standalone upgrade.
Version 2 gives us the operational foundation for that work: clear writers and readers, independent serving at each site, tested replacement procedures, and observability that follows the architecture. Version 3 will build on those controls as the DNS service expands.
The lesson from this migration is not that an on-premises server needs more care because it is physical or familiar. It needs less special treatment. Give the service a clear design. Put its configuration under control. Test a replacement before it is needed. Then let the server be cattle.
Talk to us
We do this kind of work for other organizations, not only our own platform: DNS migrations onto tested infrastructure, resilience reviews of on-premises systems, replacing servers that have quietly become irreplaceable, and tuning alerting so a page means something again.
- Managed Observability & Incident Response, our alert design and pager-noise reduction work
- Private cloud build-outs, for resilience reviews and replacing pet servers on infrastructure you own
- Contact us to talk through a DNS migration or an on-premises resilience review
Related reading:
- Managed DNS on ZCP, the architecture and setup for the nameservers this migration runs on
- We Moved Our Own DNS onto Our Own Platform, the incident that put zsoftly.ca on these nameservers
- What Nobody Tells You About Migrating a VPN Server From AWS to Private Cloud, another private cloud migration built on an assumption nobody wrote down
