Wake-Word Verification Across Networked Microphones

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing network microphone devices (NMDs) are prone to false positives due to false wake word triggers, leading to inefficient operation, resource consumption, unexpected audio interruptions, and privacy concerns when offloading voice processing to a voice assistant service (VAS).

Innovation Solution

A first playback device equipped with a wake-word engine identifies a potential wake word and sends sound data to a second playback device for verification, allowing parallel processing of command determination, reducing the time to execute user commands and minimizing privacy risks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If wake word verification is performed by offloading to a voice assistant service (VAS), then wake word detection capability is improved, but response time increases and privacy risks arise

Engineering Contradiction:
Improvewake word detection accuracyVSAvoidresponse time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent divides the wake word verification process into two segments: local wake word engine processing and remote VAS verification. The local engine performs initial detection while the remote service performs final verification, allowing parallel processing and reducing overall response time while maintaining high accuracy through distributed verification

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The local wake word engine performs preliminary detection and filtering before sending data to the VAS. This preliminary action reduces the amount of data that needs to be processed remotely and allows the system to prepare verification requests in advance, minimizing latency

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If wake word verification is performed by offloading to a voice assistant service (VAS), then wake word detection accuracy is improved, but privacy risks increase

Engineering Contradiction:
Improvewake word detection accuracyVSAvoidprivacy risks
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent extracts and removes personally identifiable information (PII) from voice data before sending it to the VAS for verification. This extraction process separates sensitive personal information from the verification process, allowing accurate wake word detection while minimizing privacy exposure by only transmitting anonymized audio segments

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary processing layer that anonymizes voice data by removing PII before transmission to the VAS. This intermediary acts as a mediator that preserves verification accuracy while protecting user privacy by ensuring that personal information never reaches the remote service

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of time

If wake word verification is performed locally without remote validation, then response time is reduced, but false positive rate increases

Engineering Contradiction:
Improveresponse timeVSAvoidfalse positive rate
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The patent implements different verification qualities at different locations: the local wake word engine provides rapid initial filtering with higher sensitivity, while the remote VAS provides final verification with higher specificity. This local-quality differentiation allows fast response times while maintaining low false positive rates through hierarchical verification

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12562167B2Localized wakeword verification
Publication Date: 2026.02.24 SONOS INC
  • US12562167B2 patent drawing
  • US12562167B2 patent drawing
  • US12562167B2 patent drawing

AI summary

In one aspect, a networked microphone device is configured to (i) receive sound data, (ii) determine, via the wake-word engine, that a first portion of the sound data is representative of a wake word, (iii) determine that a second networked microphone device was added to a media playback system, (iv) transmit the first portion of the sound data to a second networked microphone device, (v) begin determining a command to be performed by the first networked microphone device, (vi) receive an indication of whether the first portion of the sound data is representative of the wake word, and (vii) output a response indicative of whether the first portion of the sound data is representative of the wake word.