Burnt out after weeks on a 63-device ZHA network — topology won't stabilize, looking for a fresh set of eyes

Hi everyone,

Reaching out because I think I’ve hit the edge of what I know how to try, and I’d rather get a second opinion before doing something drastic (like a full network reset).

Setup: HAOS 2026.7.4 / ZHA 2.0.0, SLZB-06Mg24U coordinator (Ethernet), 63 Zigbee devices (21 routers, 41 end devices), a good chunk of it Tuya gear.

I’ve been fighting a recurring issue for weeks: a plug drops and needs to be unplugged/replugged to come back, on roughly a 24h cycle (sometimes a different plug each time). This week I tried a big cleanup to fix it once and for all: swapped 9 router plugs for a newer model, removed a heat detector that seemed to be disrupting the network, and switched coordinators (Conbee II → SLZB-06M).

Good news: routers no longer drop completely. Bad news: the topology isn’t rebuilding like it used to. Routers that are literally in the same room barely see each other anymore, and I get occasional command failures even on devices that show as “available.”

I tried enabling source_routing (24h test, with before/after measurements) — it clearly made things worse rather than better, so I reverted.

One thing that’s nagging at me: an old plug I never replaced has stayed rock solid the whole time, while the new ones have degraded badly in link quality. I’m currently testing by adding a few old-model plugs back in as reinforcement at key spots, but honestly I’m starting to go in circles and second-guess my own conclusions.

If you’ve run into this kind of symptom before — routers that stay online but whose topology never really “heals” — on a network of comparable size, or if you see something I might be missing, I’d really appreciate it, even just a hunch. Thanks for reading this far, it helps to not be stuck on this alone.

I had (have?) the same problem. I used Claude to help me write a script that did work at least once. Since then I re-flashed the last stable firmware for my SLZB-MRW10U and haven’t really had any issues. Maybe one or the other (automation or down-flash) can help you. The automation relies on the device broadcasting a leave event, so its kind of a niche thing. It can’t hurt, though.