Home DNS failover, taken too far on purpose

Cover: after Piet Mondrian, Broadway Boogie Woogie (1943) — a grid of lines carrying small pulses of colour — queries on a network — with three coloured stations on one line answering for the same address; one goes cream, the pulses flow past it to the next.
September 2026 — you would rather read this than do it. Good. I did it so you don’t have to.
Every home network has one thing that, when it breaks, makes everything look broken: DNS. The router still routes, the fibre still carries bits, the NAS still serves — and nobody in the house can open a page. If you run an ad-blocking resolver (AdGuard Home, in my case, in Kubernetes), you have also made DNS depend on the most complicated thing you own.
This is the story of a week spent on a question that does not deserve a week: how do I make home DNS failover properly? Not “add a second entry to DHCP” properly. Measured properly.
It got out of hand. There is a BGP session to a router now. There are fifteen Incus containers, a scientific benchmark suite, and five competing DNS resolver architectures running on an Orange Pi 6 Plus. Let’s go.
The starting point, and what was actually wrong with it
One AdGuard pod on a three-node cluster, exposed on 10.1.1.236 by MetalLB. DHCP hands every
client 10.1.1.236, 10.1.1.1 — AdGuard first, the router’s own resolver second. A MikroTik
scheduler script probes AdGuard every five minutes and, if it is dead, rewrites the DHCP
option to 10.1.1.1 only.
On paper this is failover. In practice, two things break:
- A client with two DNS servers does not fail over; it waits. A stub resolver asks the first server, waits its timeout (2 s on most systems), then asks the second. It does that for every query while the first is down. Pages load; they load like 2003.
- A second DNS entry gets real traffic even when the first is fine. When I ran three AdGuards, every one of them saw queries — phones, laptops, a smart TV. The “backup” quietly serves a share of everything.
Neither is an AdGuard bug; it is how DHCP resolvers behave.
Three designs, argued into shape
Every design below went through the same treatment: write it down, write the attack on it — every way it could fail, lie, or cost more than it looks — and only build what survived.
Design 1 — the relay. Clients get one address, a stateless DNS forwarder (dnsdist). It
holds an ordered list of upstreams — home AdGuard, a second AdGuard on another cluster, the
router — health-checks them every two seconds with a real lookup, and sends every query to
the first healthy one. It has a packet cache that keeps serving expired answers while no
upstream is available (setStaleCacheEntriesTTL(600)). The DHCP script stays, demoted to one
job: if the relay itself dies, hand out the router. AdGuard becomes “just software” behind a
fixed address, which also let my DR tooling move it between clusters without anyone noticing.
Cost the attack found immediately: AdGuard sees the relay as the client, not your phone. dnsdist forwards the real client in EDNS Client Subnet (ECS) and AdGuard logs it, but AdGuard’s dashboard books the query to the relay — per-client rules stop working. Cost I found later, measured: AdGuard’s default per-client rate limit (100 q/s) now applies to the whole house, and its whitelist did not exempt the relay. Turn the limit off behind a relay.
Design 2 — anycast. Three AdGuards, each announcing the same service address
(10.1.9.53/32) to the router over BGP, with BFD for sub-second failure detection and a
small health check on each node that withdraws the route when its AdGuard stops resolving
— done the boring way, by taking the interface that carries the address down, so the
announcement follows the link. The router ECMPs across whoever is healthy. Clients talk to
AdGuard directly — identity intact. No relay, no DHCP games, no script. This is how DNS is
done at scale, shrunk to a living room.
The catch, found by measurement rather than argument: when all members are gone there is
nothing — the route disappears and the address is unreachable. Every other design had the
router as an implicit last resort; anycast needs it made explicit: a MikroTik netwatch that
switches on a dst-nat redirect to the router’s resolver while no anycast member answers.
Design 0 — the baseline, kept as the control group.
Two things ran through all of them: adguardhome-sync keeps every AdGuard’s rules, rewrites
and filters equal to the one at home, and cache_optimistic — AdGuard’s serve-stale — answers
from an expired cache entry immediately and refreshes in the background, so an upstream blip
is invisible for anything looked up recently.
Measuring failover: what your phone actually feels
A small load tool played the phone: a fixed stream of queries per second against whatever address
the design handed out, mimicking a household mix — six popular names (cache hits), two random
names under a real zone (cache misses, walking the root hierarchy), and two .lan names (the
path back into the local router).
The first benchmark version had a coordinated omission bug: it sent queries sequentially, waiting for each answer. A 2-second timeout stalled the sender itself, throttling the benchmark to 0.5 q/s during outages and hiding the pain. The rewritten tool sends on a strict clock tick regardless of previous stalls.
The metrics that matter:
- Felt outage carries the entire results table: the longest continuous stretch in which a query either failed or took more than two seconds. That is the window where a browser spinner spins, an app hangs, or a video call freezes.
- Stalled — answered, but only after waiting out a full 2-second timeout (or longer) on the primary before falling back to the secondary DHCP entry. In real browsing, cascading DNS lookups lump together: a 2.1 s stall on one asset and a 9 s stall on another turn a page load into molasses. On paper, the baseline “answers 100% of queries”; in reality, it is completely unusable.
- Lost — queries that got no answer at all before the client gave up. What a streaming video, voice call, or gaming session feels immediately.
- Worst gap — the longest consecutive run of lost queries, showing how quickly the failover mechanism detects the fault.
| Scenario | Design | Lost | Stalled 2 s | Felt Outage |
|---|---|---|---|---|
| one instance restarts | baseline | 0 | 194 | 21 s |
| one instance restarts | relay | 5 | 0 | 3 s |
| one instance restarts | anycast | 8 | 0 | 3 s |
| one instance stopped 60 s | baseline | 0 | 890 | 91 s |
| one instance stopped 60 s | relay | 8 | 0 | 4 s |
| one instance stopped 60 s | anycast, container killed | 0 | 0 | 0 |
| one instance stopped 60 s | anycast, box frozen (20 q/s) | 16 | 0 | 3 s |
| every instance stopped 60 s | baseline | 0 | 915 | 93 s |
| every instance stopped 60 s | relay | 13 | 0 | 5 s |
| every instance stopped 60 s | anycast, no last resort | 38 | 0 | 60 s |
| every instance stopped 60 s | anycast, router as member | 11 | 0 | 3 s |
| flapping, 3 times | baseline | 1 | 1095 | 112 s |
| flapping, 3 times | relay | 0 | 0 | 0 |
| flapping, 3 times | anycast | 0 | 0 | 0 |
| internet down 60 s | anycast | 124 | 0 | 63 s |
| .lan path broken on one node | anycast | 0 | 0 | 0 |
The two big lessons from this table:
First, anycast has two fundamentally different failure modes. Killing a container loses nothing — the host kernel tears down the network namespace, closes the BGP socket, and the router receives an immediate TCP FIN/RST that withdraws the route in sub-milliseconds. BFD doesn’t even need to wake up. But a frozen node (kernel panic, power cut, severed link) sends no packet. That is where BFD’s 3 × 300 ms timers earn their keep: 16 queries lost in 0.8 seconds at 20 q/s before the router cuts the route.
Second, serve-stale (cache_optimistic / RFC 8767) is the cheapest availability you will ever buy.
When the WAN was severed for 60 seconds, only the never-cached misses failed (124 out of 124).
All ~1,200 cached queries answered smoothly from memory. Without serve-stale, the entire household
stops loading websites the instant the upstream connection blips.
The day the anycast trio deleted itself
The strongest finding of this whole project wasn’t planned; it was an accident.
During an earlier test with three Knot Resolver nodes, my home fiber connection blipped. I looked
over at the MikroTik router and watched something bizarre happen in real time: all three nodes
withdrew their BGP routes simultaneously, leaving paths=0. The entire anycast address vanished.
The three resolver containers were completely healthy. Their processes were running, and their in-memory caches were packed with thousands of primed domains ready to answer. But my health-check script on each node was asking: “Can this node reach 1.1.1.1 on the public internet?”
During an ISP outage, every single node answered “no” at the exact same instant.
The failover mechanism took down the dummy interfaces, withdrew the routes, and killed DNS for the entire house at the precise moment cached resolution was needed most.
The lesson is unforgiving: an anycast health check must only test local health and the path to
the router, never the public WAN. The probe must ask for a non-cached local .lan record or
check the local gateway to verify that the container’s network stack is alive. Let serve-stale
(RFC 8767) do its job for internet domains. Withdrawing the route during an ISP outage hands the LAN
to a router that can’t reach the internet either — except now nobody can even reach the local NAS.
Taking it too far: The 15-container tournament
At this point, a reasonable person puts the anycast IP into DHCP, closes the terminal, and pours coffee. I am not that person.
With a MikroTik CCR2004 speaking BGP ECMP and BFD to an Orange Pi 6 Plus running Incus containers, the next question is inevitable: why stop at AdGuard?
Every homelab thread turns into a tribal shouting match about DNS resolvers: Pi-hole vs AdGuard vs Unbound vs Technitium vs Knot. Almost none contain real numbers measured on identical hardware under identical conditions. So I built fifteen Incus containers (Debian 12, ARM64) on the Orange Pi 6 Plus — three nodes each, announcing five distinct anycast VIPs to the router:
| Competitor | Anycast VIP | Containers | Role / Architecture |
|---|---|---|---|
| AdGuard Home | 10.1.9.53 | dns-a, dns-b, dns-c | Go all-in-one ad-blocker + forwarder, SQLite query log |
| Blocky + Unbound | 10.1.9.54 | dns-e, dns-f, dns-g | Blocky filtering proxy (Go) + local Unbound recursive loopback (port 5335) + PostgreSQL HA logging |
| Unbound Standalone | 10.1.9.55 | dns-h, dns-i, dns-j | Pure C recursive validating resolver + RPZ blocklists, local memory cache |
| Technitium DNS | 10.1.9.56 | dns-k, dns-l, dns-m | C# / .NET authoritative + recursive resolver, built-in apps, Web GUI, EDE support |
| Knot Resolver 6 | 10.1.9.57 | dns-n, dns-o, dns-p | CZ.NIC high-throughput modular resolver (C + LuaJIT), local LMDB cache |
Every single container runs FRR (bgpd + bfdd), peering directly with the CCR2004 router.
Every cluster gets the exact same routing treatment, sub-second BFD timers, and health monitoring.
The scientific benchmark protocol: Meten is weten
Running five DNS servers concurrently on an 8-core ARM SoC measures thermal throttling, not DNS. To get trustworthy data, the benchmark suite enforces four rules:
- Strict sequential isolation. Resolvers are tested one by one with a 10-second quiet cooldown between competitors to let buffers flush and the SoC settle back to idle.
- Pristine cache clearing. Before every run, the target resolver services are completely
restarted (
systemctl restart ...). No leftover cache hits from a previous test. - Identical query distribution. The test load consists of 4,000 synthetic queries mirroring
a real household profile: 60% high-frequency domains (testing warm cache path), 20% distinct
zones that force recursive traversal to authoritative nameservers (testing cold resolution),
and 20% local
.lannames. - Hardware and network parity. All tests run against the respective anycast VIP from a dedicated wired gigabit Linux host across the MikroTik switch.
I evaluated six specific dimensions:
- Cold Resolution Latency: Unprimed recursive path walking root hints down to authoritative servers.
- Warm In-Memory Latency: Pure socket-to-cache efficiency across p50, p90, and p99.
- Maximum Throughput (QPS): Blasting 75,000 queries at 5,000 sustained QPS with
dnsperf. - Ad & Tracker Blocking Rate: Blasting 101 real-world tracking domains against industry-standard blocklists (HaGeZi Multi PRO, Threat Intelligence Feeds, and StevenBlack).
- RFC 8767 Stale-Serving Resilience Drill: Priming a domain, severing authoritative upstream access, allowing TTL to expire, and querying again.
- BFD Anycast Failover Under Flood: Streaming 500 QPS continuously while sending
SIGKILLto the active container, counting dropped probes.
The Scorecard
Here are the pristine benchmark results, measured sequentially with cold starts:
| Metric | AdGuard Home | Blocky + Unbound | Unbound Standalone | Technitium | Knot Resolver 6 |
|---|---|---|---|---|---|
| Cold Latency (p50) | 12.89 ms | 31.39 ms | 38.51 ms | 6.60 ms | 48.82 ms |
| Cold Latency (avg) | 17.78 ms | 46.27 ms | 62.55 ms | 11.11 ms | 48.75 ms |
| Cold Latency (p99) | 93.05 ms | 176.69 ms | 456.35 ms | 61.60 ms | 204.96 ms |
| Warm Latency (p50) | 2.04 ms | 11.06 ms | 1.17 ms | 1.44 ms | 0.99 ms |
| Warm Latency (p99) | 2.79 ms | 36.82 ms | 4.33 ms | 2.47 ms | 4.79 ms |
| Throughput (dnsperf) | 4,999 QPS | 4,998 QPS | 5,000 QPS | 1,541 QPS | 5,000 QPS |
| Block Rate (%) | 93.1% | 96.0% | 82.2% | 94.1% | 82.2% |
| RFC 8767 Serve-Stale | PASS | PASS | PASS | FAILED | PASS |
| Failover Jitter (500 QPS) | 1 lost | 1 lost | 1 lost | 8 lost | 0 lost |
Dissecting the contenders: What the numbers actually mean
Knot Resolver 6: The unyielding speed demon
If your only criterion is raw speed and wire-level protocol adherence, Knot Resolver 6 is an absolute masterclass.
CZ.NIC rebuilt Knot Resolver 6 around a declarative manager and a lightweight C/LuaJIT engine.
In warm cache, Knot answered queries in 0.99 ms at p50. When dnsperf pushed 5,000 QPS,
Knot answered 74,999 out of 75,000 queries with zero jitter.
Its anycast failover was pure poetry: when I killed the active container under a 500 QPS flood, zero queries were dropped. The TCP socket closed, the MikroTik router rehashed the flow, and Knot resumed serving without a single dropped UDP packet.
The catch? Knot is designed for ISPs. Ad-blocking relies on text RPZ files, with no web UI, no database query log, and no temporary bypass button. You edit configs and restart.
Technitium: The high-feature paradox
Technitium DNS surprised me twice — once positively, and once catastrophically.
On the positive side, Technitium’s cold recursive resolver is astonishingly fast: 6.60 ms p50 and an average of 11.11 ms. It aggressively prefetches root zones and implements Extended DNS Errors (RFC 8914), returning structured diagnostic codes that explain why a query failed or was blocked. Its .NET web management UI is clean, polished, and comprehensive.
Then came the stress tests.
When dnsperf pushed query rates beyond 1,500 QPS, Technitium hit a hard concurrency wall.
While AdGuard, Blocky, Unbound, and Knot effortlessly saturated 5,000 QPS, Technitium choked at
1,541 QPS, dropping 300 queries and throwing socket errors. On an 8-core ARM SoC, .NET’s thread
pool and memory management showed clear scaling limits compared to Go and C.
Worse: it completely failed the RFC 8767 stale-serving drill. When authoritative nameservers were blocked at the firewall, Technitium ignored its expired cache and spent 3,105 ms attempting to reach the dead upstreams before timing out. And during the anycast failover drill, it dropped 8 consecutive probes (a 4% loss rate). For high-availability clustering under duress, it wasn’t ready.
Unbound Standalone: The classic anchor
Unbound is the gold standard for recursive resolution, and the benchmark shows why. Warm cache latency is 1.17 ms. Sustained throughput reached a flat 5,000 QPS with 1 dropped query. Its RFC 8767 stale cache answered in 1.53 ms — the fastest in the entire lineup.
Unbound’s limit is the control plane. Managing millions of rules via RPZ zones is clunky, achieving 82.2% block accuracy in my test. Crucially, Unbound has no database backend; tracking queries requires grepping syslog across containers. There is no web UI, no search API, and no temporary unblock. It is a brilliant recursive engine without a dashboard.
AdGuard Home: The consumer champion
AdGuard Home scored exceptionally well across almost every consumer metric. It blocked 93.1% of
trackers out of the box. Its warm cache answered in 2.04 ms, throughput hit 4,999 QPS, and its
cache_optimistic serve-stale engine answered expired records in 3.63 ms.
The flaw in AdGuard Home is clustering. Query logs and stats live in a local SQLite database (data/querylog.json / data/stats.db). Across three anycast containers, history is fragmented across three separate UIs. Syncing settings requires running adguardhome-sync on a cron schedule. And unblocking broken sites is an untimed manual toggle: you disable it, forget to re-enable it, and leave the house unfiltered for days.
Blocky + Unbound: The hybrid architecture
Which brings me to the hybrid pair: Blocky running as the
frontend proxy on port 53, delegating recursive resolution to a local Unbound instance running
on loopback (127.0.0.1:5335).
Blocky is written in Go, specifically designed as a fast, cluster-friendly ad-blocking DNS proxy. It does not attempt to be a recursive resolver; it leaves recursion, DNSSEC validation, and RFC 8767 caching to Unbound.
In the benchmarks, Blocky + Unbound delivered:
- 96.0% ad/tracker blocking rate — stopping 97 out of 101 malicious and telemetry domains.
- 4,998 QPS throughput with 0 dropped queries.
- 2.54 ms RFC 8767 serve-stale latency (Unbound serving expired cache through Blocky).
- 1 dropped probe (0.19%) during BFD anycast failover.
Its warm cache latency was 11.06 ms at p50 — higher than Unbound standalone because Blocky evaluates 2.46 million regex and wildcard blocklist rules in Go memory for every query before handing the answer back. But in real human perception, 11 ms is completely instantaneous.
More importantly, Blocky solves every operational flaw that plagues multi-instance DNS.
Why Blocky + Unbound won my homelab
Choosing what to run at home isn’t decided on an Excel sheet of microsecond latencies. It is decided at 8:30 PM on a Tuesday when your partner needs to sign a mortgage document on DocuSign, the page stalls on a tracking redirect, and the phone rings.
Here is why Blocky + Unbound earned the permanent spot:
1. Centralized, searchable query logs in PostgreSQL (CNPG HA)
In an anycast cluster, ECMP hashes queries across all three containers. A single web browsing
session will spray DNS lookups across dns-e, dns-f, and dns-g.
With AdGuard or Pi-hole, you have three fragmented SQLite files. If a smart plug goes rogue and starts beaconing to a strange IP, finding its query history requires searching three different UIs.
Blocky natively supports PostgreSQL as a query log target:
queryLog:
type: postgresql
target: postgres://blocky:SECRET@cnpg-cluster-rw.postgres.svc:5432/blocky?sslmode=require
logRetentionDays: 30
Because my Kubernetes cluster runs CloudNativePG (CNPG) with high availability and automated failover, all three anycast containers stream their query logs into a single, resilient database. If a container dies, no logs are lost. If I need to trace a client, I can query a month of DNS traffic across the entire house in milliseconds using standard SQL or Grafana.
2. The “DocuSign Problem” and timed web GUI unblocking
Aggressive blocklists inevitably break legitimate workflows: a DocuSign signing link in a tracking redirect (click.docusign.net), an affiliate review link, or bank authentication.
In Unbound, fixing this requires editing config files and reloading over SSH. In AdGuard, disabling protection is a global untimed toggle: you sign, grab coffee, and forget that the entire house sits unprotected for three days.
Blocky solves this via its REST API and blocky-ui. When deployed in Kubernetes and gated behind
an Authentik forward-auth outpost (blocky.djieno.com), anyone in the family can open the dashboard
and click “Disable blocking for 5 minutes”:
POST /api/blocking/disable?duration=5m
Blocky silences the blocklists for that client (or globally) for exactly five minutes, and then automatically re-arms itself. The document gets signed, the tracking link opens, and the firewall closes behind you without manual intervention.
3. Maximum blocking muscle without memory bloat
Blocky is designed around Go concurrency. I loaded it with the full HaGeZi Multi PRO list, the Threat Intelligence Feeds (TIF), and StevenBlack’s combined hosts — 2.46 million active rules.
Blocky compiled the entire list into in-memory radix trees and hash maps in under four seconds, consuming roughly 380 MB of RAM per container. In the tracker benchmark, it caught 96.0% of all test trackers (97 out of 101), outperforming AdGuard (93.1%) and crushing standalone RPZ (82.2%).
4. Single declarative GitOps configuration
AdGuard Home and Technitium are stateful appliances: settings are configured through web UIs into local state files. Keeping three anycast instances in sync requires external sync daemons, scheduled cron jobs, and prayer.
Blocky has zero local state. Its entire behavior is defined in a single, clean YAML file:
upstreams:
groups:
default:
- 127.0.0.1:5335
blocking:
blackLists:
ads:
- https://raw.githubusercontent.com/StevenBlack/hosts/master/hosts
- https://raw.githubusercontent.com/hagezi/dns-blocklists/main/wildcard/pro-onlydomains.txt
- https://raw.githubusercontent.com/hagezi/dns-blocklists/main/wildcard/tif-onlydomains.txt
whiteLists:
ads:
- /etc/blocky/allowlist.txt
clientGroupsBlock:
default:
- ads
caching:
minTime: 2m
maxTime: 30m
prefetching: true
prefetchThreshold: 3
This file is checked into Git and deployed by FluxCD. Every container mounts the exact same ConfigMap. There is no configuration drift, no sync cron job, and no database replication lag.
5. Decoupled control plane and data plane
By pairing Blocky with Unbound, each component does what it was built to do:
- Blocky is the control plane: fast filtering, regex evaluation, PostgreSQL logging, REST API, and Authentik-secured web UI.
- Unbound is the data plane: local loopback recursion, cryptographic DNSSEC validation, and RFC 8767 stale cache serving.
If Blocky reloads its blocklists, Unbound’s cache keeps serving queries. If an upstream ISP link fails, Unbound’s RFC 8767 stale engine delivers expired answers in 2.5 ms while Blocky logs the event.
Three side quests that turned out to matter more than topology
Three details surfaced during the benchmarking that have nothing to do with failover protocols, but everything to do with how fast a home network actually feels.
1. The SafeBrowsing tax: 12 ms on every cache miss for zero blocks
If you run AdGuard Home, you probably checked the box that says “Use AdGuard browsing security web service” (SafeBrowsing) to block phishing and malware domains.
I put a packet capture on the wire to see why cold lookups on AdGuard occasionally felt sluggish:
- Client query arrives at
t = 0.0 ms. - Upstream query to the recursive resolver is dispatched at
t = 14.8 ms.
Inside AdGuard, every single cache miss waits ~12 to 15 ms while AdGuard sends a synchronous API query to its cloud lookup servers before it even attempts to resolve the actual domain upstream.
Over a 27,000-query household benchmark, SafeBrowsing blocked exactly 0 domains. Zero.
This is a known upstream issue (AdGuardHome #2857): the security check happens synchronously on the critical path of cold resolution rather than in parallel. Unchecking SafeBrowsing immediately claws back 12 ms on every cache miss.
2. Reverse lookups and the private PTR cache reality
A friend mentioned that reverse lookups of his own LAN addresses felt noticeably sluggish on AdGuard compared to Unbound.
I measured reverse lookups over the 57 IP addresses in my DHCP lease table, cold pass then warm pass, p50:
| Resolver | Cached Forward Lookup | Reverse PTR, Cold | Reverse PTR, Warm |
|---|---|---|---|
| Router (Source of Truth) | 0.2 ms | 0.50 ms | 0.51 ms |
| Unbound Standalone | 0.8 ms | 1.86 ms | 0.47 ms |
| Knot Resolver 6 | 0.6 ms | 1.65 ms | 0.72 ms |
| Blocky + Unbound | 11.0 ms | 1.91 ms | 1.88 ms |
| AdGuard Home (Same Host) | 1.2 ms | 2.41 ms | 2.45 ms |
The sharper version now measured reveals a fundamental architectural divide: Unbound and Knot Resolver 6 cache private reverse PTR lookups, while AdGuard Home and Blocky never do.
Unbound answers warm reverse queries from memory in 0.47 ms and Knot in 0.72 ms, counting down the TTL properly.
AdGuard and Blocky treat private PTR lookups as an uncached passthrough:
- AdGuard routes private reverse DNS (
local_ptr_upstreams) through an uncached path and returns TTL 300 on every single answer, never decreasing (#6950). Every single lookup is forwarded to the router, costing ~2.45 ms every time. - Blocky forwards PTR lookups directly to its upstream resolver without maintaining an in-memory PTR cache
(
cold ≈ warm ≈ 1.9 ms). In the Blocky + Unbound hybrid, local Unbound on loopback absorbs the work, keeping latency low, but the query still traverses the loopback hop.
If you run network dashboards, Grafana panels, or home scanners that constantly reverse-lookup LAN IPs, Unbound and Knot handle them in microseconds without touching your router.
3. The rate-limiting trap behind forwarders
During the first dnsperf run, AdGuard suddenly capped out at 102 q/s with 66% query loss.
That was not a CPU bottleneck; it was AdGuard’s default per-client rate limit (ratelimit: 100).
When running behind a relay or a reverse proxy, the entire house looks like a single IP address.
If you place a forwarder in front of any resolver, you must explicitly disable the resolver’s
internal rate limiter and enforce rate limits at the ingress layer.
Which resolver works for whom?
After testing five architectures across thousands of queries, here is the honest matrix for homelab operators:
| Architecture | Best For | Key Strengths | Trade-offs |
|---|---|---|---|
| Blocky + Unbound | HA Homelabs, GitOps, Kubernetes clusters | PostgreSQL HA query logging, timed web unblock (DocuSign), 96% block rate, declarative YAML | Requires two daemons per node (Blocky + Unbound loopback) |
| Knot Resolver 6 | Throughput purists, ISP networks, bare-metal routing | Sub-millisecond warm cache (0.99 ms), 0-loss BFD failover, minimal CPU footprint | No web GUI, no query log database, RPZ rule management via text configs |
| AdGuard Home | Single-node setups, family homes, quick deployments | Excellent out-of-the-box UI, great parental controls, fast warm cache | SQLite database tied to local disk, no native multi-node clustering, no timed unblock |
| Unbound Standalone | Network security engineers, pure recursive DNS | Battle-tested C codebase, RFC 8767 stale cache in 1.5 ms, authoritative recursion | No GUI, text-only logging, lower block rate with stock RPZ lists (82%) |
| Technitium DNS | Windows sysadmins, enterprise lab testing | Fastest cold resolution (6.6 ms), Extended DNS Errors (RFC 8914), rich app ecosystem | Concurrency ceiling at ~1,500 QPS, failed RFC 8767 stale drill, higher memory usage |
If you want to rebuild this
If you want to eliminate DNS downtime from your home network, here is the recommended roadmap:
- Enable serve-stale immediately. Whether you run AdGuard (
cache_optimistic: true), dnsdist (setStaleCacheEntriesTTL(600)), or Unbound (serve-expired: yes), this zero-cost setting renders upstream ISP outages invisible for active domains. - Hand out one address in DHCP. Stop listing two DNS servers in your router’s DHCP options. A second address is a 2-second stall penalty for your family.
- If you have one box: run a relay. A tiny
dnsdistinstance pointing at your primary resolver with a health check falling back to the router gives you instant failover with minimal complexity. - If you have multiple boxes: build BGP Anycast.
- Install FRR (
bgpdandbfdd) on each node. - Assign the VIP (e.g.
10.1.9.54/32) to a dummy interface (lo:dnsordummy0). - Configure FRR to redistribute the connected
/32with a route-map. - Run BFD with 300 ms intervals.
- Add a health-check script that takes down the dummy interface if DNS resolution fails.
- On your router (MikroTik, pfSense, OPNsense, or VyOS): enable BGP ECMP and configure a fallback redirect (netwatch) so traffic routes to the router’s local resolver if all anycast nodes disappear.
- Install FRR (
- Hardware: Three NanoPi R3S LTS units (RK3566, 2 GB RAM, dual gigabit Ethernet, metal case) draw ~2 watts each and cost under €50. Three of them provide a physical anycast fabric that survives your main server rack going down.
- For GitOps homelabs: deploy Blocky + Unbound. Mount a single declarative YAML config,
point Blocky’s query log at an HA PostgreSQL cluster (like CloudNativePG), expose
blocky-uibehind Authentik forward-auth, and enjoy centralized search, timed bypass, and 96% ad-blocking.
What runs at home now
The anycast cluster on the Orange Pi 6 Plus is now the permanent primary DNS for the house.
DHCP hands out a single address: 10.1.9.54.
Behind that address, three Incus containers run Blocky in front of Unbound, peering over BGP
and BFD with the MikroTik CCR2004. Query logs flow into CloudNativePG. When someone in the living
room needs to sign a mortgage agreement or click an affiliate review link, they open
blocky.djieno.com, click “Pause for 5 minutes,” and the network handles the rest.
If you take one lesson from a week of DNS obsession: a second DNS server in DHCP is not failover, it is a two-second tax on every query for as long as the primary is down.
Everything else in this post was just finding the most thorough, enjoyable way to stop paying it.



