I was looking to retire my google minis for some time with fast local replacement, unfortunately Onju Voice wasn’t “resposnive” I’m expecting voice assistant to be able to act as fast as human would, e.g. no delay after wake word and perceived “instant” (within 1 second) response.
I had V100 8GB vGPU collecting dust on my Proxmox. It was too small for any meaningful LLM, but recently there were 2 interesting projects released:
- Confucius4-R2T2 transcript (it runs on my V100) with almost realtime transcription that is sent back to “client app”, it is much faster than whisper and understands both Polish and English without any problems
- Needle3 - used for smart home tool calls, this tiny model feels real time even on CPU
My architecture has 2 servers:
- GPU Ubuntu server running R2T2, Needle3, Kokoro
- Windows client app that handles logic and integrates with PC microphone/speakers
When command is “mistranscribed” the client app fallbacks to Groq to try figure out what the command was and feeds it back to Needle3, when Needle3 fails for the second time, it calls Groq again to try get the tool call using gpt-oss-120b.
There are some tricks to make it faster like overlap with wake word for R2T2 so you don’t need to pause after the wake word, make R2T2 transcribe as you speak with utterance detection so the transcribed text is ready as soon as you stop speaking.
It feels much faster than google home mini in most cases (when command doesn’t require multiple groq round trips). I’m planning to buy bunch of ESP32-S3 XMOS and make them work as satellites that connect to the client app.
The PoC uses OpenHab, but I’m slowly migrating to HA, didn’t check HA how it could integrate R2T2, Needle3, Kokoro yet I’m new to this ecosystem.
It also supports a single timer and weather info.
Note all of the code is vibe-coded and project is just working PoC that matches my deployment requirements running on diffrent GPUs/VM configuration might require some adjusting.