Range-Based Network Microphone Activation Without Wake Words

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional wake-word engines in network microphone devices (NMDs) are prone to false positives due to false wake words and phonetically similar words, leading to resource consumption and audio playback interruptions.

Innovation Solution

Implementing a combination of physical conditions (touch or line-of-sight) with command keywords to trigger voice assistants without a pre-determined wake word, and using a local natural language unit to process voice inputs, reducing the need for cloud processing and minimizing false positives.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional wake-word engines are used in network microphone devices, then the system can trigger voice assistants through voice commands, but false positives occur due to false wake words and phonetically similar words leading to resource consumption and audio playback interruptions

Engineering Contradiction:
Improvefalse positive rateVSAvoidresource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system segments the voice processing function by introducing a local natural language unit that operates independently from the cloud-based wake-word engine. This local unit processes voice inputs first to filter out false positives before they reach the main voice assistant trigger mechanism, thereby reducing unnecessary resource consumption from false activations while maintaining reliable true positive detection

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The local natural language unit serves as an intermediary between the wake-word detection system and the voice assistant execution. It acts as a filtering layer that validates whether detected wake words represent genuine user intent, preventing false positives from triggering resource-intensive voice assistant operations and audio playback interruptions

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If cloud processing is used for all voice inputs, then comprehensive processing capability is achieved, but user privacy is compromised and response time increases

Engineering Contradiction:
Improvevoice processing accuracyVSAvoiduser privacy
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The voice processing architecture is segmented into local and cloud components. The local natural language unit handles preliminary processing of voice inputs, filtering and pre-processing data locally before selecting which inputs require cloud processing. This segmentation maintains processing accuracy for complex queries while protecting privacy by keeping sensitive local processing on-device

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The local natural language unit performs preliminary processing of voice inputs before cloud transmission. It pre-analyzes and filters voice data locally, preparing only necessary inputs for cloud processing. This preliminary action reduces the volume of data requiring cloud transmission, thereby maintaining comprehensive processing capability while minimizing privacy loss through reduced cloud exposure

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If wake-word engines continuously monitor for wake words, then voice assistant activation is enabled, but false positives from phonetically similar words cause audio playback interruptions

Engineering Contradiction:
Improvevoice assistant activationVSAvoidfalse wake word triggers
Core Design Contradiction:
Ease of operationVSObject-generated harmful factors

Solution Approach 1:

The system implements feedback mechanisms where the local natural language unit continuously monitors and evaluates wake-word detections. When phonetically similar words are detected, the local unit provides feedback to suppress false triggers by analyzing contextual relevance and user intent patterns, thereby maintaining ease of activation while filtering harmful false positives

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The local natural language unit acts as an intermediary validation layer between wake-word detection and voice assistant activation. It mediates the trigger decision by evaluating whether detected words represent genuine activation intent versus phonetically similar false positives, preventing harmful interruptions while maintaining responsive activation for valid commands

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12424220B2Network device interaction by range
Publication Date: 2025.09.23 SONOS INC
  • US12424220B2 patent drawing
  • US12424220B2 patent drawing
  • US12424220B2 patent drawing

AI summary

Examples described herein relate to triggering voice assistant(s) on a network microphone device (NMD). An NMD is a networked computing device that typically includes an arrangement of microphones, such as a microphone array, that is configured to detect sound present in the NMD's environment. Once the voice assistant is triggered, the NMD may start recording voice input as a potential voice command. Within examples, the NMD may operate in a wakewordless mode if certain conditions are met. These conditions may involve detecting user proximity in one of multiple different ranges. For instance, an example NMD may monitor for user proximity in a first range from the playback device via at least one touch-sensitive sensor and/or user line-of-sight in a second range that is further from the playback device than the first range. When either user proximity or user line-of-sight is detected, the NMD may enables the wakewordless mode.