How we distilled HuBERT into 1.22 MB for wake words

JarbasAI

JarbasAI (Lead Author)

OVOS Contributor

Co-authors:

Claude (Anthropic)

Claude (Anthropic)

How we distilled HuBERT into 1.22 MB for wake words

How we distilled HuBERT into 1.22 MB for wake words

HuBERT is a self-supervised speech model from Meta AI (paper). It learned about speech from recordings with no labels, and its inner layers describe speech well enough that a small classifier on top can learn a lot from very little data. The base model, facebook/hubert-base-ls960, has about 95 million parameters and looks at whole utterances at once. A wake word engine listens all day, on small devices, to a live stream. HuBERT-base is far too large and slow for that.

So we distilled it. WakeHuBERT tiny has about 0.64 million parameters. Its int8 build is 1.46 MB, and the float32 build is 3.3 MB. It is strictly causal, so it streams. And it keeps enough of what HuBERT knows that a wake word classifier trained on its features, from synthetic speech only, detects the word in real recordings.

To put the sizes in a form everyone remembers, here they are in 3.5-inch "1.44 MB" floppy disks. Formatted with FAT12, such a disk holds 1,457,664 bytes of files (2,847 sectors of 512 bytes). In the table, 1 MB is 1,000,000 bytes.

model parameters file floppy disks
HuBERT-base about 95 million 377.6 MB 260
DistilHuBERT, ONNX float32 about 23.5 million 94.0 MB 65
DistilHuBERT, ONNX int8 about 23.5 million 50.4 MB 35
WakeHuBERT tiny, float32 about 0.64 million 3.3 MB 3
WakeHuBERT tiny, int8 about 0.64 million 1.46 MB 2
WakeHuBERT tiny, int8, gzip -9 about 0.64 million 1.22 MB 1

The int8 student misses a single floppy by 5,181 bytes. Compressed with gzip -9, it fits, with 238 KB to spare.

Yes, that is where the 1.22 MB in the title comes from. Nobody runs a gzipped model: onnxruntime loads the 1.46 MB file. We gzipped it only so it would fit on a floppy, and so the headline would have a smaller number in it. We admit the clickbait.

We were not the first to shrink HuBERT. DistilHuBERT (model) cuts it to a quarter of the size, and our first wake word experiments ran on an ONNX export of it. It works, and it is an option on a laptop with compute to spare, but it is still far too heavy for a single-board computer. That is why we distilled our own.

👉 WakeHuBERT tiny on Hugging Face (Apache-2.0)


The teacher and the student

The teacher is HuBERT-base. The student is a log-mel front end followed by a causal temporal convolutional network: a strided convolution, dilated depthwise-separable convolution blocks and a small projection. It uses only convolution, batch norm and ReLU, which quantise well. The log-mel front end is part of the ONNX graph, so the student takes the raw 16 kHz waveform and gives 128-dimensional features at 50 frames per second.

Each output frame depends only on audio that came before it, within a receptive field of 2.5 seconds. Keep 2.5 seconds of context and the streamed features are the same as the offline ones.

The student learned to predict three of the teacher's layers: 4, 8 and 12. The teacher heard clean speech. The student heard the same speech with noise, reverberation and background talkers, so it learned to describe the speech and ignore the rest.

The speech came from three open corpora, cut into 600,000 two-second crops: LibriSpeech (960 hours of English audiobooks), Multilingual LibriSpeech (seven more languages) and a language-balanced sample of Multilingual Spoken Words (41 languages). The noise came from MUSAN and AudioSet, and a quarter of the training items were non-speech sounds that teacher and student heard the same way.

What worked, and what did not

Masked distillation worked. During training, spans of the student's input were hidden, each frame masked at probability 0.065, and the student still had to reproduce the teacher. The model card reports this as the largest single gain in robustness found in the experiments.

Width helped. A wider student was better.

Lookahead did not help. Giving the student a little future audio, at the cost of latency, did not make it better.

Other teachers were no better. We also distilled students from WavLM and XEUS. Neither beat the HuBERT student. All of these students are published in the Onnx feature extractors collection, the ONNX featurizers for wake word experiments: the WakeHuBERT, WakeWav and WakeXeus families. They can all be used with wakeforge.

The last step was quantisation. The static int8 build, the 1.46 MB file, gives features that agree with float32 at a mean cosine similarity of 0.998.

You can use WakeHuBERT on its own, outside OVOS:

import numpy as np, onnxruntime as ort
from huggingface_hub import hf_hub_download

path = hf_hub_download("TigreGotico/wakehubert-tiny", "wakehubert.onnx")
sess = ort.InferenceSession(path)
feats = sess.run(None, {"waveform": np.zeros((1, 24000), np.float32)})[0]  # (1, 75, 128)

The model card has the full architecture, the training data and an evaluation.


What it lets you do: the wakeforge plugin

The new OVOS wakeforge wake word plugin runs wake word models built on WakeHuBERT. This is the ready-to-use part: it ships the featurizer in both builds and ready models, and the runtime needs only onnxruntime and numpy.

pip install --pre ovos-ww-plugin-wakeforge

The ready models (see below) are trained from synthetic speech, so a word that nobody has ever recorded is not out of reach. The featurizer and the models were made with wakeforge, our research framework for wake word experiments. It is not a one-click model maker, but it is open, and anyone who wants to experiment with their own word can use the same toolkit we did. The rest of this post explains how the models are trained, what we learned on the way, and how they score on real speech.

From 'Alexa' to OVOS' 'Hey Mycroft', check out the list of WakeHuBERT-wakewords on HuggingFace, growing as we speak.


The wake word heads

On top of WakeHuBERT sits a small GRU classifier with a hidden size of 128. One model per word is enough. We tried an ensemble of six models, and once the training data was right it barely helped.

The heads are trained with the research bench in wakeforge (scripts/research/head_bench.py). The bench is a research script, the one we used to run these experiments, not a packaged training tool. Each positive clip becomes eight augmented copies. The augmentation includes a device-response stage that imitates cheap hardware: a limited microphone band, a coloured frequency response, level changes, clipping and self-noise. Babble and noise are mixed in on top.

The checkpoint we keep is the one with the best recall at zero false accepts on held-out calibration speech. The result is exported to the plugin's format with ww_trainer-export-plugin.


The data recipe: text-to-speech plus voice cloning

This is the part we think is new. Every word starts as a grid of text-to-speech voices that covers every variant of the language: every English accent, every Portuguese voice. Any OVOS TTS plugin can supply the voices, proprietary services such as Edge and Google included, and phoonnx alone exposes thousands of models across languages. Each voice says the word at 5 speaking rates and 3 pitches, with 3 spellings that change the delivery ("jarvis", "jarvis!", "jarvis?").

The clips are then voice-cloned onto real speakers with Chatterbox, through voiceclonnx, a pure-ONNX voice cloning library. Cloning works across languages, so one pool of reference speakers serves every language.

For Catalan, Galician and Basque we are adding more voices through phoonnx: Matxa and the other Projecte AINA voices from BSC for Catalan, the Proxecto Nós voices from the University of Vigo for Galician, and the HiTZ voices for Basque.


What we learned

Sound-alike negatives hurt. It seems obvious to teach the model what the word is not, with synthetic near-homophones such as "commuter" or "come pewter" for "computer". On real speech it backfired. On real "computer" recordings, at 1 false activation per hour, recall dropped from 90.5% to 55.2% when the sound-alikes were added.

More of one voice is not more data. Adding a single multi-speaker Piper voice as extra positives also made the models slightly worse.

Voice coverage is what matters. The full voice grid made the difference. On real "jarvis" recordings, recall at 1 false activation per hour went from 85.4% for our earlier model to 99.0–99.5% (two training seeds).

The synthetic-only models never see real recordings during training. Real recordings are used only to test them.


Open data

Every word's training set is a public dataset, TigreGotico/synthetic-wakeword-<word>, under CC BY 4.0. They are grouped in collections per language:

The sound-alike negatives are published separately, as not-wake-words-soundalikes-en and not-wake-words-soundalikes-pt. Read the finding above before you use them: in our tests they cost a lot of recall on real speech.


Try it

Try-out any of the new WakeHuBERT wakewords in the online HuggingFace space, without installing anything. If you like one or more of them, install the OVOS plugin, then set it as your wake word engine in ~/.config/mycroft/mycroft.conf:

{
  "listener": {
    "wake_word": "jarvis"
  },
  "hotwords": {
    "jarvis": {
      "module": "ovos-ww-plugin-wakeforge",
      "model": "jarvis",
      "listen": true
    }
  }
}

Restart OVOS and say "jarvis". If you experiment with wakeforge and export a model of your own, point model at its ONNX file instead; the plugin README lists every option. Bugs and questions go to the issue tracker.


Built in the Open, Funded for the Commons

WakeHuBERT, wakeforge and the plugin were developed by TigreGótico for OpenVoiceOS. As we shared when OpenVoiceOS received its NGI Zero Commons Fund grant, that funding goes toward the plumbing a community project rarely has the resources to polish. A wake word engine that anyone can train for their own word and language, from open data, is that kind of plumbing.

This work is part of the OpenVoiceOS From Beta to Breakthrough milestone, funded through the NGI0 Commons Fund, a fund established by NLnet with financial support from the European Commission's Next Generation Internet programme, under the aegis of DG Communications Networks, Content and Technology under grant agreement No 101135429. Additional funding is made available by the Swiss State Secretariat for Education, Research and Innovation (SERI).


Help Us Build Voice for Everyone

OpenVoiceOS is more than software, it's a mission. If you believe voice assistants should be open, inclusive, and user-controlled, here's how you can help:

  • 💸 Donate: Help us fund development, infrastructure, and legal protection.
  • 📣 Contribute Open Data: Share voice samples and transcriptions under open licenses.
  • 🌍 Translate: Help make OVOS accessible in every language.

We're not building this for profit. We're building it for people. With your support, we can keep voice tech transparent, private, and community-owned.

👉 Support the project here

JarbasAI

JarbasAI