Matter / Thread onboarding and stability issues — observations from real-world home network environments

We have been collecting a number of real-world user reports related to Matter / Thread devices and would like to check whether the community has observed similar patterns.

This is not targeted at any specific brand or platform, but rather an attempt to identify possible common environmental factors from a networking perspective.

Observed cases (simplified)

Case 1 — Device becomes unreachable after a period of normal operation

Some Matter / Thread devices work normally during initial setup, but become unreachable after a few days of use, and cannot be reliably re-added even after a reset.


Case 2 — Matter onboarding fails during pairing stage

When onboarding via QR code:

  • Device enters pairing mode
  • Commissioning process starts (“Connecting…”)
  • Then times out
  • Repeated attempts after reset may still fail

Case 3 — Local discovery or sub-device pairing instability

In hub-based systems, sub-devices are sometimes not discoverable even when in pairing mode, or show generic “compatibility” or pairing failure errors during binding.


Environmental observation

In these cases, the underlying network connectivity is typically stable (strong WiFi signal, working internet connection, and other smart devices operating normally).

However, a common pattern is that many of these reports come from more complex home network setups, often involving managed networking systems such as UniFi-based routers, switches, or access points.

This is not intended as a conclusion, but rather as a possible shared environmental factor worth further observation.


Technical discussion points

Matter and Thread onboarding and device discovery rely on several local networking mechanisms, including:

  • mDNS / DNS-SD service discovery
  • IPv6 link-local communication
  • Multicast traffic within the local network
  • Coordination between Thread Border Routers and the IP network layer

In managed network environments, these mechanisms may be affected by configurations such as:

  • multicast forwarding / IGMP snooping behavior
  • VLAN or WiFi SSID isolation
  • firewall or security inspection features
  • cross-subnet mDNS propagation settings

Even when overall connectivity appears normal, these factors may still impact device discovery or onboarding flows, resulting in:

  • Devices entering pairing mode but not being discovered
  • Timeouts during commissioning
  • Previously working devices becoming intermittently unreachable or unable to rejoin

We are not assuming any single root cause. We are simply trying to understand:

:backhand_index_pointing_right: If you have experienced similar issues, how did you resolve or work around them?
:backhand_index_pointing_right: Are there any specific network configurations or practical experiences that made Matter / Thread devices more stable in your environment?
:backhand_index_pointing_right: When using UniFi / multi-AP / VLAN / mesh networks, what adjustments have you made to improve device discovery or onboarding success?

Any real-world troubleshooting experiences (successful or not) would be highly appreciated, as they could help us better understand potential patterns in these issues.

Matter & Thread Deep Dive

You are missing a cause in the list.

  • multiple NICs in Matter and/or Thread server host, which can cause services to flip-flop between NICs.

Besides that I run with several devices, mostly on Thread, but also some on WiFi.
I have prevented some common issues, which are also known from other mesh networks, like Zigbee and Z-wave.

  1. If there is a weak spot, do not assume a single repeatering device can fix it. RF signals are tricky to predict. Add more or even many.
  2. Make sure to use devices that the vendors support with firmware updates, especially for the repeating devices. I had issues with the Onvis S4 smart plug that had a bug, which made it stop repeating messages, but still announce itself as a repeating device. A firmware update fixed that. The Onvis was working fine the whole time, but other devices would fail until the big was patches.
  3. Avoid VLANs. I tried it in the start and I know how to set up routing, also for discovery protocols, but there are so many pitfalls that you need to know about and handle.

Those are excellent points.

You're absolutely right about multiple NICs. We've also seen cases where Matter or Thread services bind to the wrong network interface, leading to intermittent commissioning or discovery issues.

We also agree that a healthy Thread mesh requires enough powered routing devices rather than relying on a single repeater. Firmware quality on Thread Routers is another factor that's often overlooked.

Your point about avoiding unnecessary VLAN complexity is also valuable, especially for users who are just getting started with Matter.

People often overlook the fact, that Thread, Zigbee, Bluetooth and WiFi share the same 2,4 Ghz Radio Band and so they don't choose the appropriate channels to avoid interferences between those 4 Protocols. Common symptoms are failing firmware updates or devices loosing connectivity over time.
And i don't know if there are completely "failsafe" autochannel functions in the standard setups existing.

Trouble started with both matter WiFi bulbs and matter WiFi switched plugs. Using the companion app, the device onboarding process continue to saying the device is configured and configuration sent to HA. The next step, after "Done", just hangs. The device does not appear in the Matter list, nor in searching devices and entities. Running on Pi5 with all updates installed. Some devices install, but some fail. When a device is reset after a failed install, sometimes it succeeds, but mostly not. I think the following logs show the problem, but my programming days ended a couple decades ago. The router is an RT-AX86U running current Merlin and the IPV6 tests fine. Logs:

2026-06-30 18:46:09.414 INFO ConfigStorage Set config key fabricLabel to MYCTxC

2026-06-30 18:55:41.609 INFO ClientInteraction Subscription successful « @1:76•b43f⇵7918 2↔2 id: 74ff7a4a interval: 1m 9s

2026-06-30 18:56:11.834 INFO LegacyDataLoader Saved server data to 10166129605400261104.json: 21 node(s), last_node_id=118

2026-06-30 18:56:11.834 INFO LegacyDataLoader Batch update to legacy server file: added: 118

2026-06-30 18:56:45.030 INFO Session •unsecured#824383a5639da835 Session ended

2026-06-30 18:56:45.036 WARN PeerConnection @1:75 udp://[fe80::7a20:51ff:fe41:ed18%end0]:5540 Authorization rejected by peer on session resumption; clearing resumption data and retrying

2026-06-30 18:56:45.037 INFO PeerConnection @1:75•unsecured#b9fbdda8e447c866⇵7919 udp://[fe80::7a20:51ff:fe41:ed18%end0]:5540 Connecting addr #: 1 attempt #: 2 connect time: 10m 56s addr time: 10m 56s fast fallback

2026-06-30 18:56:45.060 INFO Session •unsecured#b9fbdda8e447c866 Session ended

2026-06-30 18:56:45.061 WARN PeerConnection @1:75 udp://[fe80::7a20:51ff:fe41:ed18%end0]:5540 Peer error (retry in 5m): [channel-status-response] (Failure (1) / NoSharedTrustRoots (1)) Received general error status for protocol 0 (Sigma2(Resume))

2026-06-30 19:01:45.063 INFO PeerConnection @1:75•unsecured#75a4feb545050dec⇵791a udp://[fe80::7a20:51ff:fe41:ed18%end0]:5540 Connecting addr #: 1 attempt #: 3 connect time: 15m 56s addr time: 15m 56s fast fallback

2026-06-30 19:01:45.189 INFO Session •unsecured#75a4feb545050dec Session ended

2026-06-30 19:01:45.189 WARN PeerConnection @1:75 udp://[fe80::7a20:51ff:fe41:ed18%end0]:5540 Peer error (retry in 5m): [channel-status-response] (Failure (1) / NoSharedTrustRoots (1)) Received general error status for protocol 0 (Sigma2(Resume))

Any guidance is welcome. BTW, I have been running home control systems over 20 years and currently support Z-Wave, Zigbee, Matter, and Thread. Several Thread devices work fine. Non WiFi Matter devices aslso work.

HAOS running on a mini PC.
I originally had a few Switchbot and Tapo devices that used Matter over Wifi. These have been 100% reliable for many months.
I bought a ZBT-2 and a whole bunch of Ikea Matter over Thread devices. I added the Routing devices first, then the battery devices. Took over a week to pair all the devices, the majority needing multiple attempts and factory resets. Very frustrating.
For a couple of months I would get several different devices going unavailable each day. Different devices in different rooms every time, but the GU10 light bulbs were the worst. Power-cycling fixed them, but a few hours later something else would go unavailable.
About a month ago I added my Apple TV as a second Border Router on the same Thread network, and stability has improved massively.
I still get one or two dropouts a week though - a different device each time.

I have checked which channels wifi and Thread are using, and there appears to be no overlap.

My setup is "mostly" working.

  • HA (Version: 2026.5.4) running inside a VM, running as a nixos option (services.home-assistant.enable = true)
  • thread bridge running using podman
  • matter server js latest running inside podman
  • ZBT-2

My devices lose connectivity in around 30 minutes, what solves it is if I restart matter-server things start working again and doing what they are supposed to.

To get things connected:

  1. using the iOS HA app, add the device by scanning the code
  2. goes from connecting -> setting up
  3. at this point the app report that it failed
  4. I start watching the zeroconf config page on my browser (/config/zeroconf) and after around 10-15 minutes a udp entry appears for the device
  5. I navigate to devices->matter->options->Add Manually and type in the numeric code

Then it adds it and is able to read the state (at least for my window/door sensor -- MYGGBETT)

After 30 or so minutes HA stops receiving updates,

INFO   ClientSubscription   Subscription 70669cce to peer @1:1 timed out after 30m 38s

It seems to resolve without a restart but I have not figured out yet how long it takes, I am thinking around 10 minutes. Though, after a manual restart, it fixes things right away.

I really want this to work, but this experience is not great. I'm going to try running the latest versions of everything I can next...

The biggest frustration for HA users using HA's Matter Server with the HA OTBR is Thread dataset/credential management for iOS or Android based HA Companion App. HA Companion App relies on the underlying iOS/Android Matter/Thread framework which does not provide sufficient means for managing Thread datasets stored on the respective mobile devices. As a result, when pairing a Matter over Thread device using the HA Companion App on a mobile device, the Matter over Thread device is given the incorrect Thread dataset and thus can not join the Thread network (and thus Matter can't pair it either) without any explicit/direct feedback to the user that the dataset/credentials didn't work (but may get an indirect feedback such as a TBR is required, or the Connecting....). Even worse, is that Android does not provide the means to remove or override a pre-existing dataset (without using a hack that requires all the Google Play/Store data being deleted).

1 Like

I'm new to HA and don't know the intricacies of matter/thread pairing process -- so please forgive my ignorance. How are people claiming they're using thread/matter devices natively? What are they using? What does a working setup look like? I really want to use thread/matter. I thought getting the zbt2 would make it all work out of the box, but it apparently does not. Thank you

The ZBT-2 works absolutely flawlessly and setup is easy.
The most common problems occur due to incorrect configuration of your home router/wifi, the phone used for setup and the setup of your HA system if it not installed “bare metal” on a Raspberry, MiniPC or similar but running on a VM, Proxmox, etc.
If IPv6/mDNS gets blocked or not forwarded correctly throughout this hole chain, you will likely run into (mostly solvable) problems.

I’m running HA on a Raspberry Pi5 using a ZBT-2 as Border router and i can confirm, that it runs rockstable with 46 devices (more than half of it router devices like bulbs and plugs) spread throughout the hole house.
No dropouts or unavailable devices.
As i mentioned earlier: don’t overlook the possible interferences with wifi, bluetooth and zigbee. You have to choose the different channels wisely!
My settings:
WiFi 2,4 Ghz on channel 11 (max. 20 Mhz Bandwith!)
Zigbee on channel 11 (Zigbee and WiFi use different channel numbers!)
Thread on channel 25
With these settings i can avoid overlapping/interfering radio signals.

2 Likes

Thanks, I’ll double check my channels. I’m fairly confident that my IPV6 mDNS broadcasts are working especially considering the fact that my two test devices are do work for a period of time. I was even able to upgrade their firmware. Around the 30 minute mark they just stop being updated in HA even though they are indicating that they are sending a signal. Manually rebooting my matter server appears to force a resubscription and everything starts working again. This doesn’t seem to be a radio or mDNS issue. In the logs (matter server) I see a timeout reported, then logs saying it resubscriscribed, but it no longer updates in HA. No errors are logged in HA or in the thread bridge. I can normally fix issues but the silence in the logs is killing me. I will double check the logging levels but I believe I set them to debug.

the log in matter server reporting the timeout and the resubscription:

2026-07-08 14:58:17.761 INFO   ClientSubscription   Subscription d9cd374c to peer @1:b timed out after 30m 38s
2026-07-08 14:58:17.762 INFO   ClientSubscription   Replacing subscription to @1:b due to timeout
2026-07-08 14:58:17.763 INFO   ClientInteraction    Probe » @1:b•efe4⇵6bf1
2026-07-08 14:58:28.481 INFO   ClientInteraction    Probe « @1:b•efe4⇵6bf1 (success)
2026-07-08 14:58:28.482 INFO   ClientInteraction    Subscribe » @1:b•efe4⇵6bf2 min: 0 max: 10m 35s attributes: 1 events: 1
2026-07-08 14:58:29.565 INFO   ClientInteraction    Subscription successful « @1:b•efe4⇵6bf2 2↔2 id: 2a2e0389 interval: 30m timeout: 30m 38s
2026-07-08 14:59:05.648 INFO   ClientSubscription   Subscription d629afda to peer @1:1 timed out after 30m 38s
2026-07-08 14:59:05.648 INFO   ClientSubscription   Replacing subscription to @1:1 due to timeout
2026-07-08 14:59:05.650 INFO   ClientInteraction    Probe » @1:1•efe6⇵6bf3
2026-07-08 14:59:05.961 INFO   ClientInteraction    Probe « @1:1•efe6⇵6bf3 (success)
2026-07-08 14:59:05.962 INFO   ClientInteraction    Subscribe » @1:1•efe6⇵6bf4 min: 0 max: 10m 29s attributes: 1 events: 1
2026-07-08 14:59:06.974 INFO   ClientInteraction    Subscription successful « @1:1•efe6⇵6bf4 2↔2 id: 98f754fb interval: 30m timeout: 30m 38s
2026-07-08 14:59:08.971 INFO   UdpMulticastServer   lo: send ENETUNREACH ff02::fb%lo:5353
2026-07-08 15:00:12.304 INFO   ClientInteraction    Read » @1:1•efe6⇵6bf5 attributes: 1
2026-07-08 15:00:23.064 INFO   ClientInteraction    Read « @1:1•efe6⇵6bf5 attributes: 1 events: 0
2026-07-08 15:00:49.132 INFO   UdpMulticastServer   lo: send ENETUNREACH ff02::fb%lo:5353
2026-07-08 15:02:29.302 INFO   UdpMulticastServer   lo: send ENETUNREACH ff02::fb%lo:5353
2026-07-08 15:04:08.598 INFO   UdpMulticastServer   lo: send ENETUNREACH ff02::fb%lo:5353
2026-07-08 15:05:49.350 INFO   UdpMulticastServer   lo: send ENETUNREACH ff02::fb%lo:5353
2026-07-08 15:07:29.530 INFO   UdpMulticastServer   lo: send ENETUNREACH ff02::fb%lo:5353
2026-07-08 15:09:04.324 INFO   ClientInteraction    Read » @1:1•efe6⇵6bf6 attributes: 1
2026-07-08 15:09:09.078 INFO   ClientInteraction    Read « @1:1•efe6⇵6bf6 attributes: 1 events: 0
2026-07-08 15:09:12.968 INFO   UdpMulticastServer   lo: send ENETUNREACH ff02::fb%lo:5353

How many Thread devices do you have and how many of them are routers? What does the topology overwiew show in regards to signal quality?