Gino Eising
Gino Eising
Nerd by Nature
Sep 9, 2026 22 min read

Home DNS failover, taken too far on purpose

thumbnail for this post

Cover: after Piet Mondrian, Broadway Boogie Woogie (1943) — a grid of lines carrying small pulses of colour — queries on a network — with three coloured stations on one line answering for the same address; one goes cream, the pulses flow past it to the next.

September 2026 — you would rather read this than do it. Good. I did it so you don’t have to.

Every home network has one thing that, when it breaks, makes everything look broken: DNS. The router still routes, the fibre still carries bits, the NAS still serves — and nobody in the house can open a page. If you run an ad-blocking resolver (AdGuard Home, in my case, in Kubernetes), you have also made DNS depend on the most complicated thing you own.

This is the story of a week spent on a question that does not deserve a week: how do I make home DNS failover properly? Not “add a second entry to DHCP” properly. Measured properly.

It got out of hand. There is a BGP session to a router now. There are fifteen Incus containers, a scientific benchmark suite, and five competing DNS resolver architectures running on an Orange Pi 6 Plus. Let’s go.


The starting point, and what was actually wrong with it

One AdGuard pod on a three-node cluster, exposed on 10.1.1.236 by MetalLB. DHCP hands every client 10.1.1.236, 10.1.1.1 — AdGuard first, the router’s own resolver second. A MikroTik scheduler script probes AdGuard every five minutes and, if it is dead, rewrites the DHCP option to 10.1.1.1 only.

On paper this is failover. In practice, two things break:

  1. A client with two DNS servers does not fail over; it waits. A stub resolver asks the first server, waits its timeout (2 s on most systems), then asks the second. It does that for every query while the first is down. Pages load; they load like 2003.
  2. A second DNS entry gets real traffic even when the first is fine. When I ran three AdGuards, every one of them saw queries — phones, laptops, a smart TV. The “backup” quietly serves a share of everything.

Neither is an AdGuard bug; it is how DHCP resolvers behave.


Three designs, argued into shape

Every design below went through the same treatment: write it down, write the attack on it — every way it could fail, lie, or cost more than it looks — and only build what survived.

Design 1 — the relay. Clients get one address, a stateless DNS forwarder (dnsdist). It holds an ordered list of upstreams — home AdGuard, a second AdGuard on another cluster, the router — health-checks them every two seconds with a real lookup, and sends every query to the first healthy one. It has a packet cache that keeps serving expired answers while no upstream is available (setStaleCacheEntriesTTL(600)). The DHCP script stays, demoted to one job: if the relay itself dies, hand out the router. AdGuard becomes “just software” behind a fixed address, which also let my DR tooling move it between clusters without anyone noticing.

Cost the attack found immediately: AdGuard sees the relay as the client, not your phone. dnsdist forwards the real client in EDNS Client Subnet (ECS) and AdGuard logs it, but AdGuard’s dashboard books the query to the relay — per-client rules stop working. Cost I found later, measured: AdGuard’s default per-client rate limit (100 q/s) now applies to the whole house, and its whitelist did not exempt the relay. Turn the limit off behind a relay.

Design 2 — anycast. Three AdGuards, each announcing the same service address (10.1.9.53/32) to the router over BGP, with BFD for sub-second failure detection and a small health check on each node that withdraws the route when its AdGuard stops resolving — done the boring way, by taking the interface that carries the address down, so the announcement follows the link. The router ECMPs across whoever is healthy. Clients talk to AdGuard directly — identity intact. No relay, no DHCP games, no script. This is how DNS is done at scale, shrunk to a living room.

The catch, found by measurement rather than argument: when all members are gone there is nothing — the route disappears and the address is unreachable. Every other design had the router as an implicit last resort; anycast needs it made explicit: a MikroTik netwatch that switches on a dst-nat redirect to the router’s resolver while no anycast member answers.

Design 0 — the baseline, kept as the control group.

Two things ran through all of them: adguardhome-sync keeps every AdGuard’s rules, rewrites and filters equal to the one at home, and cache_optimistic — AdGuard’s serve-stale — answers from an expired cache entry immediately and refreshes in the background, so an upstream blip is invisible for anything looked up recently.


Measuring failover: what your phone actually feels

A small load tool played the phone: a fixed stream of queries per second against whatever address the design handed out, mimicking a household mix — six popular names (cache hits), two random names under a real zone (cache misses, walking the root hierarchy), and two .lan names (the path back into the local router).

The first benchmark version had a coordinated omission bug: it sent queries sequentially, waiting for each answer. A 2-second timeout stalled the sender itself, throttling the benchmark to 0.5 q/s during outages and hiding the pain. The rewritten tool sends on a strict clock tick regardless of previous stalls.

The metrics that matter:

  • Felt outage carries the entire results table: the longest continuous stretch in which a query either failed or took more than two seconds. That is the window where a browser spinner spins, an app hangs, or a video call freezes.
  • Stalled — answered, but only after waiting out a full 2-second timeout (or longer) on the primary before falling back to the secondary DHCP entry. In real browsing, cascading DNS lookups lump together: a 2.1 s stall on one asset and a 9 s stall on another turn a page load into molasses. On paper, the baseline “answers 100% of queries”; in reality, it is completely unusable.
  • Lost — queries that got no answer at all before the client gave up. What a streaming video, voice call, or gaming session feels immediately.
  • Worst gap — the longest consecutive run of lost queries, showing how quickly the failover mechanism detects the fault.
ScenarioDesignLostStalled 2 sFelt Outage
one instance restartsbaseline019421 s
one instance restartsrelay503 s
one instance restartsanycast803 s
one instance stopped 60 sbaseline089091 s
one instance stopped 60 srelay804 s
one instance stopped 60 sanycast, container killed000
one instance stopped 60 sanycast, box frozen (20 q/s)1603 s
every instance stopped 60 sbaseline091593 s
every instance stopped 60 srelay1305 s
every instance stopped 60 sanycast, no last resort38060 s
every instance stopped 60 sanycast, router as member1103 s
flapping, 3 timesbaseline11095112 s
flapping, 3 timesrelay000
flapping, 3 timesanycast000
internet down 60 sanycast124063 s
.lan path broken on one nodeanycast000

The two big lessons from this table:

First, anycast has two fundamentally different failure modes. Killing a container loses nothing — the host kernel tears down the network namespace, closes the BGP socket, and the router receives an immediate TCP FIN/RST that withdraws the route in sub-milliseconds. BFD doesn’t even need to wake up. But a frozen node (kernel panic, power cut, severed link) sends no packet. That is where BFD’s 3 × 300 ms timers earn their keep: 16 queries lost in 0.8 seconds at 20 q/s before the router cuts the route.

Second, serve-stale (cache_optimistic / RFC 8767) is the cheapest availability you will ever buy. When the WAN was severed for 60 seconds, only the never-cached misses failed (124 out of 124). All ~1,200 cached queries answered smoothly from memory. Without serve-stale, the entire household stops loading websites the instant the upstream connection blips.

The day the anycast trio deleted itself

The strongest finding of this whole project wasn’t planned; it was an accident.

During an earlier test with three Knot Resolver nodes, my home fiber connection blipped. I looked over at the MikroTik router and watched something bizarre happen in real time: all three nodes withdrew their BGP routes simultaneously, leaving paths=0. The entire anycast address vanished.

The three resolver containers were completely healthy. Their processes were running, and their in-memory caches were packed with thousands of primed domains ready to answer. But my health-check script on each node was asking: “Can this node reach 1.1.1.1 on the public internet?”

During an ISP outage, every single node answered “no” at the exact same instant.

The failover mechanism took down the dummy interfaces, withdrew the routes, and killed DNS for the entire house at the precise moment cached resolution was needed most.

The lesson is unforgiving: an anycast health check must only test local health and the path to the router, never the public WAN. The probe must ask for a non-cached local .lan record or check the local gateway to verify that the container’s network stack is alive. Let serve-stale (RFC 8767) do its job for internet domains. Withdrawing the route during an ISP outage hands the LAN to a router that can’t reach the internet either — except now nobody can even reach the local NAS.


Taking it too far: The 15-container tournament

At this point, a reasonable person puts the anycast IP into DHCP, closes the terminal, and pours coffee. I am not that person.

With a MikroTik CCR2004 speaking BGP ECMP and BFD to an Orange Pi 6 Plus running Incus containers, the next question is inevitable: why stop at AdGuard?

Every homelab thread turns into a tribal shouting match about DNS resolvers: Pi-hole vs AdGuard vs Unbound vs Technitium vs Knot. Almost none contain real numbers measured on identical hardware under identical conditions. So I built fifteen Incus containers (Debian 12, ARM64) on the Orange Pi 6 Plus — three nodes each, announcing five distinct anycast VIPs to the router:

CompetitorAnycast VIPContainersRole / Architecture
AdGuard Home10.1.9.53dns-a, dns-b, dns-cGo all-in-one ad-blocker + forwarder, SQLite query log
Blocky + Unbound10.1.9.54dns-e, dns-f, dns-gBlocky filtering proxy (Go) + local Unbound recursive loopback (port 5335) + PostgreSQL HA logging
Unbound Standalone10.1.9.55dns-h, dns-i, dns-jPure C recursive validating resolver + RPZ blocklists, local memory cache
Technitium DNS10.1.9.56dns-k, dns-l, dns-mC# / .NET authoritative + recursive resolver, built-in apps, Web GUI, EDE support
Knot Resolver 610.1.9.57dns-n, dns-o, dns-pCZ.NIC high-throughput modular resolver (C + LuaJIT), local LMDB cache

Every single container runs FRR (bgpd + bfdd), peering directly with the CCR2004 router. Every cluster gets the exact same routing treatment, sub-second BFD timers, and health monitoring.


The scientific benchmark protocol: Meten is weten

Running five DNS servers concurrently on an 8-core ARM SoC measures thermal throttling, not DNS. To get trustworthy data, the benchmark suite enforces four rules:

  1. Strict sequential isolation. Resolvers are tested one by one with a 10-second quiet cooldown between competitors to let buffers flush and the SoC settle back to idle.
  2. Pristine cache clearing. Before every run, the target resolver services are completely restarted (systemctl restart ...). No leftover cache hits from a previous test.
  3. Identical query distribution. The test load consists of 4,000 synthetic queries mirroring a real household profile: 60% high-frequency domains (testing warm cache path), 20% distinct zones that force recursive traversal to authoritative nameservers (testing cold resolution), and 20% local .lan names.
  4. Hardware and network parity. All tests run against the respective anycast VIP from a dedicated wired gigabit Linux host across the MikroTik switch.

I evaluated six specific dimensions:

  • Cold Resolution Latency: Unprimed recursive path walking root hints down to authoritative servers.
  • Warm In-Memory Latency: Pure socket-to-cache efficiency across p50, p90, and p99.
  • Maximum Throughput (QPS): Blasting 75,000 queries at 5,000 sustained QPS with dnsperf.
  • Ad & Tracker Blocking Rate: Blasting 101 real-world tracking domains against industry-standard blocklists (HaGeZi Multi PRO, Threat Intelligence Feeds, and StevenBlack).
  • RFC 8767 Stale-Serving Resilience Drill: Priming a domain, severing authoritative upstream access, allowing TTL to expire, and querying again.
  • BFD Anycast Failover Under Flood: Streaming 500 QPS continuously while sending SIGKILL to the active container, counting dropped probes.

The Scorecard

Here are the pristine benchmark results, measured sequentially with cold starts:

MetricAdGuard Home
10.1.9.53
Blocky + Unbound
10.1.9.54
Unbound Standalone
10.1.9.55
Technitium
10.1.9.56
Knot Resolver 6
10.1.9.57
Cold Latency (p50)12.89 ms31.39 ms38.51 ms6.60 ms48.82 ms
Cold Latency (avg)17.78 ms46.27 ms62.55 ms11.11 ms48.75 ms
Cold Latency (p99)93.05 ms176.69 ms456.35 ms61.60 ms204.96 ms
Warm Latency (p50)2.04 ms11.06 ms (2.46M rules)1.17 ms1.44 ms0.99 ms
Warm Latency (p99)2.79 ms36.82 ms4.33 ms2.47 ms4.79 ms
Throughput (dnsperf)4,999 QPS (1 lost)4,998 QPS (0 lost)5,000 QPS (1 lost)1,541 QPS (300 lost)5,000 QPS (1 lost)
Block Rate (%)93.1% (94/101)96.0% (97/101)82.2% (83/101)94.1% (95/101)82.2% (83/101)
RFC 8767 Serve-StalePASS (3.63 ms)PASS (2.54 ms)PASS (1.53 ms)FAILED (3,105 ms timeout)PASS (2.46 ms)
Failover Jitter (500 QPS)1 lost (0.19%)1 lost (0.19%)1 lost (0.18%)8 lost (4.04%)0 lost (0.00%)

Dissecting the contenders: What the numbers actually mean

Knot Resolver 6: The unyielding speed demon

If your only criterion is raw speed and wire-level protocol adherence, Knot Resolver 6 is an absolute masterclass.

CZ.NIC rebuilt Knot Resolver 6 around a declarative manager and a lightweight C/LuaJIT engine. In warm cache, Knot answered queries in 0.99 ms at p50. When dnsperf pushed 5,000 QPS, Knot answered 74,999 out of 75,000 queries with zero jitter.

Its anycast failover was pure poetry: when I killed the active container under a 500 QPS flood, zero queries were dropped. The TCP socket closed, the MikroTik router rehashed the flow, and Knot resumed serving without a single dropped UDP packet.

The catch? Knot is designed for ISPs. Ad-blocking relies on text RPZ files, with no web UI, no database query log, and no temporary bypass button. You edit configs and restart.

Technitium: The high-feature paradox

Technitium DNS surprised me twice — once positively, and once catastrophically.

On the positive side, Technitium’s cold recursive resolver is astonishingly fast: 6.60 ms p50 and an average of 11.11 ms. It aggressively prefetches root zones and implements Extended DNS Errors (RFC 8914), returning structured diagnostic codes that explain why a query failed or was blocked. Its .NET web management UI is clean, polished, and comprehensive.

Then came the stress tests.

When dnsperf pushed query rates beyond 1,500 QPS, Technitium hit a hard concurrency wall. While AdGuard, Blocky, Unbound, and Knot effortlessly saturated 5,000 QPS, Technitium choked at 1,541 QPS, dropping 300 queries and throwing socket errors. On an 8-core ARM SoC, .NET’s thread pool and memory management showed clear scaling limits compared to Go and C.

Worse: it completely failed the RFC 8767 stale-serving drill. When authoritative nameservers were blocked at the firewall, Technitium ignored its expired cache and spent 3,105 ms attempting to reach the dead upstreams before timing out. And during the anycast failover drill, it dropped 8 consecutive probes (a 4% loss rate). For high-availability clustering under duress, it wasn’t ready.

Unbound Standalone: The classic anchor

Unbound is the gold standard for recursive resolution, and the benchmark shows why. Warm cache latency is 1.17 ms. Sustained throughput reached a flat 5,000 QPS with 1 dropped query. Its RFC 8767 stale cache answered in 1.53 ms — the fastest in the entire lineup.

Unbound’s limit is the control plane. Managing millions of rules via RPZ zones is clunky, achieving 82.2% block accuracy in my test. Crucially, Unbound has no database backend; tracking queries requires grepping syslog across containers. There is no web UI, no search API, and no temporary unblock. It is a brilliant recursive engine without a dashboard.

AdGuard Home: The consumer champion

AdGuard Home scored exceptionally well across almost every consumer metric. It blocked 93.1% of trackers out of the box. Its warm cache answered in 2.04 ms, throughput hit 4,999 QPS, and its cache_optimistic serve-stale engine answered expired records in 3.63 ms.

The flaw in AdGuard Home is clustering. Query logs and stats live in a local SQLite database (data/querylog.json / data/stats.db). Across three anycast containers, history is fragmented across three separate UIs. Syncing settings requires running adguardhome-sync on a cron schedule. And unblocking broken sites is an untimed manual toggle: you disable it, forget to re-enable it, and leave the house unfiltered for days.

Blocky + Unbound: The hybrid architecture

Which brings me to the hybrid pair: Blocky running as the frontend proxy on port 53, delegating recursive resolution to a local Unbound instance running on loopback (127.0.0.1:5335).

Blocky is written in Go, specifically designed as a fast, cluster-friendly ad-blocking DNS proxy. It does not attempt to be a recursive resolver; it leaves recursion, DNSSEC validation, and RFC 8767 caching to Unbound.

In the benchmarks, Blocky + Unbound delivered:

  • 96.0% ad/tracker blocking rate — stopping 97 out of 101 malicious and telemetry domains.
  • 4,998 QPS throughput with 0 dropped queries.
  • 2.54 ms RFC 8767 serve-stale latency (Unbound serving expired cache through Blocky).
  • 1 dropped probe (0.19%) during BFD anycast failover.

Its warm cache latency was 11.06 ms at p50 — higher than Unbound standalone because Blocky evaluates 2.46 million regex and wildcard blocklist rules in Go memory for every query before handing the answer back. But in real human perception, 11 ms is completely instantaneous.

More importantly, Blocky solves every operational flaw that plagues multi-instance DNS.


Why Blocky + Unbound won my homelab

Choosing what to run at home isn’t decided on an Excel sheet of microsecond latencies. It is decided at 8:30 PM on a Tuesday when your partner needs to sign a mortgage document on DocuSign, the page stalls on a tracking redirect, and the phone rings.

Here is why Blocky + Unbound earned the permanent spot:

1. Centralized, searchable query logs in PostgreSQL (CNPG HA)

In an anycast cluster, ECMP hashes queries across all three containers. A single web browsing session will spray DNS lookups across dns-e, dns-f, and dns-g.

With AdGuard or Pi-hole, you have three fragmented SQLite files. If a smart plug goes rogue and starts beaconing to a strange IP, finding its query history requires searching three different UIs.

Blocky natively supports PostgreSQL as a query log target:

queryLog:
  type: postgresql
  target: postgres://blocky:SECRET@cnpg-cluster-rw.postgres.svc:5432/blocky?sslmode=require
  logRetentionDays: 30

Because my Kubernetes cluster runs CloudNativePG (CNPG) with high availability and automated failover, all three anycast containers stream their query logs into a single, resilient database. If a container dies, no logs are lost. If I need to trace a client, I can query a month of DNS traffic across the entire house in milliseconds using standard SQL or Grafana.

2. The “DocuSign Problem” and timed web GUI unblocking

Aggressive blocklists inevitably break legitimate workflows: a DocuSign signing link in a tracking redirect (click.docusign.net), an affiliate review link, or bank authentication.

In Unbound, fixing this requires editing config files and reloading over SSH. In AdGuard, disabling protection is a global untimed toggle: you sign, grab coffee, and forget that the entire house sits unprotected for three days.

Blocky solves this via its REST API and blocky-ui. When deployed in Kubernetes and gated behind an Authentik forward-auth outpost (blocky.djieno.com), anyone in the family can open the dashboard and click “Disable blocking for 5 minutes”:

POST /api/blocking/disable?duration=5m

Blocky silences the blocklists for that client (or globally) for exactly five minutes, and then automatically re-arms itself. The document gets signed, the tracking link opens, and the firewall closes behind you without manual intervention.

3. Maximum blocking muscle without memory bloat

Blocky is designed around Go concurrency. I loaded it with the full HaGeZi Multi PRO list, the Threat Intelligence Feeds (TIF), and StevenBlack’s combined hosts — 2.46 million active rules.

Blocky compiled the entire list into in-memory radix trees and hash maps in under four seconds, consuming roughly 380 MB of RAM per container. In the tracker benchmark, it caught 96.0% of all test trackers (97 out of 101), outperforming AdGuard (93.1%) and crushing standalone RPZ (82.2%).

4. Single declarative GitOps configuration

AdGuard Home and Technitium are stateful appliances: settings are configured through web UIs into local state files. Keeping three anycast instances in sync requires external sync daemons, scheduled cron jobs, and prayer.

Blocky has zero local state. Its entire behavior is defined in a single, clean YAML file:

upstreams:
  groups:
    default:
      - 127.0.0.1:5335

blocking:
  blackLists:
    ads:
      - https://raw.githubusercontent.com/StevenBlack/hosts/master/hosts
      - https://raw.githubusercontent.com/hagezi/dns-blocklists/main/wildcard/pro-onlydomains.txt
      - https://raw.githubusercontent.com/hagezi/dns-blocklists/main/wildcard/tif-onlydomains.txt
    whiteLists:
      ads:
        - /etc/blocky/allowlist.txt
  clientGroupsBlock:
    default:
      - ads

caching:
  minTime: 2m
  maxTime: 30m
  prefetching: true
  prefetchThreshold: 3

This file is checked into Git and deployed by FluxCD. Every container mounts the exact same ConfigMap. There is no configuration drift, no sync cron job, and no database replication lag.

5. Decoupled control plane and data plane

By pairing Blocky with Unbound, each component does what it was built to do:

  • Blocky is the control plane: fast filtering, regex evaluation, PostgreSQL logging, REST API, and Authentik-secured web UI.
  • Unbound is the data plane: local loopback recursion, cryptographic DNSSEC validation, and RFC 8767 stale cache serving.

If Blocky reloads its blocklists, Unbound’s cache keeps serving queries. If an upstream ISP link fails, Unbound’s RFC 8767 stale engine delivers expired answers in 2.5 ms while Blocky logs the event.


Three side quests that turned out to matter more than topology

Three details surfaced during the benchmarking that have nothing to do with failover protocols, but everything to do with how fast a home network actually feels.

1. The SafeBrowsing tax: 12 ms on every cache miss for zero blocks

If you run AdGuard Home, you probably checked the box that says “Use AdGuard browsing security web service” (SafeBrowsing) to block phishing and malware domains.

I put a packet capture on the wire to see why cold lookups on AdGuard occasionally felt sluggish:

  • Client query arrives at t = 0.0 ms.
  • Upstream query to the recursive resolver is dispatched at t = 14.8 ms.

Inside AdGuard, every single cache miss waits ~12 to 15 ms while AdGuard sends a synchronous API query to its cloud lookup servers before it even attempts to resolve the actual domain upstream.

Over a 27,000-query household benchmark, SafeBrowsing blocked exactly 0 domains. Zero.

This is a known upstream issue (AdGuardHome #2857): the security check happens synchronously on the critical path of cold resolution rather than in parallel. Unchecking SafeBrowsing immediately claws back 12 ms on every cache miss.

2. Reverse lookups and the private PTR cache reality

A friend mentioned that reverse lookups of his own LAN addresses felt noticeably sluggish on AdGuard compared to Unbound.

I measured reverse lookups over the 57 IP addresses in my DHCP lease table, cold pass then warm pass, p50:

ResolverCached Forward LookupReverse PTR, ColdReverse PTR, Warm
Router (Source of Truth)0.2 ms0.50 ms0.51 ms
Unbound Standalone0.8 ms1.86 ms0.47 ms
Knot Resolver 60.6 ms1.65 ms0.72 ms
Blocky + Unbound11.0 ms1.91 ms1.88 ms
AdGuard Home (Same Host)1.2 ms2.41 ms2.45 ms

The sharper version now measured reveals a fundamental architectural divide: Unbound and Knot Resolver 6 cache private reverse PTR lookups, while AdGuard Home and Blocky never do.

Unbound answers warm reverse queries from memory in 0.47 ms and Knot in 0.72 ms, counting down the TTL properly.

AdGuard and Blocky treat private PTR lookups as an uncached passthrough:

  • AdGuard routes private reverse DNS (local_ptr_upstreams) through an uncached path and returns TTL 300 on every single answer, never decreasing (#6950). Every single lookup is forwarded to the router, costing ~2.45 ms every time.
  • Blocky forwards PTR lookups directly to its upstream resolver without maintaining an in-memory PTR cache (cold ≈ warm ≈ 1.9 ms). In the Blocky + Unbound hybrid, local Unbound on loopback absorbs the work, keeping latency low, but the query still traverses the loopback hop.

If you run network dashboards, Grafana panels, or home scanners that constantly reverse-lookup LAN IPs, Unbound and Knot handle them in microseconds without touching your router.

3. The rate-limiting trap behind forwarders

During the first dnsperf run, AdGuard suddenly capped out at 102 q/s with 66% query loss.

That was not a CPU bottleneck; it was AdGuard’s default per-client rate limit (ratelimit: 100). When running behind a relay or a reverse proxy, the entire house looks like a single IP address. If you place a forwarder in front of any resolver, you must explicitly disable the resolver’s internal rate limiter and enforce rate limits at the ingress layer.


Which resolver works for whom?

After testing five architectures across thousands of queries, here is the honest matrix for homelab operators:

ArchitectureBest ForKey StrengthsTrade-offs
Blocky + UnboundHA Homelabs, GitOps, Kubernetes clustersPostgreSQL HA query logging, timed web unblock (DocuSign), 96% block rate, declarative YAMLRequires two daemons per node (Blocky + Unbound loopback)
Knot Resolver 6Throughput purists, ISP networks, bare-metal routingSub-millisecond warm cache (0.99 ms), 0-loss BFD failover, minimal CPU footprintNo web GUI, no query log database, RPZ rule management via text configs
AdGuard HomeSingle-node setups, family homes, quick deploymentsExcellent out-of-the-box UI, great parental controls, fast warm cacheSQLite database tied to local disk, no native multi-node clustering, no timed unblock
Unbound StandaloneNetwork security engineers, pure recursive DNSBattle-tested C codebase, RFC 8767 stale cache in 1.5 ms, authoritative recursionNo GUI, text-only logging, lower block rate with stock RPZ lists (82%)
Technitium DNSWindows sysadmins, enterprise lab testingFastest cold resolution (6.6 ms), Extended DNS Errors (RFC 8914), rich app ecosystemConcurrency ceiling at ~1,500 QPS, failed RFC 8767 stale drill, higher memory usage

If you want to rebuild this

If you want to eliminate DNS downtime from your home network, here is the recommended roadmap:

  1. Enable serve-stale immediately. Whether you run AdGuard (cache_optimistic: true), dnsdist (setStaleCacheEntriesTTL(600)), or Unbound (serve-expired: yes), this zero-cost setting renders upstream ISP outages invisible for active domains.
  2. Hand out one address in DHCP. Stop listing two DNS servers in your router’s DHCP options. A second address is a 2-second stall penalty for your family.
  3. If you have one box: run a relay. A tiny dnsdist instance pointing at your primary resolver with a health check falling back to the router gives you instant failover with minimal complexity.
  4. If you have multiple boxes: build BGP Anycast.
    • Install FRR (bgpd and bfdd) on each node.
    • Assign the VIP (e.g. 10.1.9.54/32) to a dummy interface (lo:dns or dummy0).
    • Configure FRR to redistribute the connected /32 with a route-map.
    • Run BFD with 300 ms intervals.
    • Add a health-check script that takes down the dummy interface if DNS resolution fails.
    • On your router (MikroTik, pfSense, OPNsense, or VyOS): enable BGP ECMP and configure a fallback redirect (netwatch) so traffic routes to the router’s local resolver if all anycast nodes disappear.
  5. Hardware: Three NanoPi R3S LTS units (RK3566, 2 GB RAM, dual gigabit Ethernet, metal case) draw ~2 watts each and cost under €50. Three of them provide a physical anycast fabric that survives your main server rack going down.
  6. For GitOps homelabs: deploy Blocky + Unbound. Mount a single declarative YAML config, point Blocky’s query log at an HA PostgreSQL cluster (like CloudNativePG), expose blocky-ui behind Authentik forward-auth, and enjoy centralized search, timed bypass, and 96% ad-blocking.

What runs at home now

The anycast cluster on the Orange Pi 6 Plus is now the permanent primary DNS for the house. DHCP hands out a single address: 10.1.9.54.

Behind that address, three Incus containers run Blocky in front of Unbound, peering over BGP and BFD with the MikroTik CCR2004. Query logs flow into CloudNativePG. When someone in the living room needs to sign a mortgage agreement or click an affiliate review link, they open blocky.djieno.com, click “Pause for 5 minutes,” and the network handles the rest.

If you take one lesson from a week of DNS obsession: a second DNS server in DHCP is not failover, it is a two-second tax on every query for as long as the primary is down.

Everything else in this post was just finding the most thorough, enjoyable way to stop paying it.