Range-Based Network Microphone Activation Without Wake Words
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional wake-word engines in network microphone devices (NMDs) are prone to false positives due to false wake words and phonetically similar words, leading to resource consumption and audio playback interruptions.
Innovation Solution
Implementing a combination of physical conditions (touch or line-of-sight) with command keywords to trigger voice assistants without a pre-determined wake word, and using a local natural language unit to process voice inputs, reducing the need for cloud processing and minimizing false positives.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional wake-word engines are used in network microphone devices, then the system can trigger voice assistants through voice commands, but false positives occur due to false wake words and phonetically similar words leading to resource consumption and audio playback interruptions
Solution Approach 1:
The system segments the voice processing function by introducing a local natural language unit that operates independently from the cloud-based wake-word engine. This local unit processes voice inputs first to filter out false positives before they reach the main voice assistant trigger mechanism, thereby reducing unnecessary resource consumption from false activations while maintaining reliable true positive detection
Solution Approach 2:
The local natural language unit serves as an intermediary between the wake-word detection system and the voice assistant execution. It acts as a filtering layer that validates whether detected wake words represent genuine user intent, preventing false positives from triggering resource-intensive voice assistant operations and audio playback interruptions
2Measurement precision
If cloud processing is used for all voice inputs, then comprehensive processing capability is achieved, but user privacy is compromised and response time increases
Solution Approach 1:
The voice processing architecture is segmented into local and cloud components. The local natural language unit handles preliminary processing of voice inputs, filtering and pre-processing data locally before selecting which inputs require cloud processing. This segmentation maintains processing accuracy for complex queries while protecting privacy by keeping sensitive local processing on-device
Solution Approach 2:
The local natural language unit performs preliminary processing of voice inputs before cloud transmission. It pre-analyzes and filters voice data locally, preparing only necessary inputs for cloud processing. This preliminary action reduces the volume of data requiring cloud transmission, thereby maintaining comprehensive processing capability while minimizing privacy loss through reduced cloud exposure
3Ease of operation
If wake-word engines continuously monitor for wake words, then voice assistant activation is enabled, but false positives from phonetically similar words cause audio playback interruptions
Solution Approach 1:
The system implements feedback mechanisms where the local natural language unit continuously monitors and evaluates wake-word detections. When phonetically similar words are detected, the local unit provides feedback to suppress false triggers by analyzing contextual relevance and user intent patterns, thereby maintaining ease of activation while filtering harmful false positives
Solution Approach 2:
The local natural language unit acts as an intermediary validation layer between wake-word detection and voice assistant activation. It mediates the trigger decision by evaluating whether detected words represent genuine activation intent versus phonetically similar false positives, preventing harmful interruptions while maintaining responsive activation for valid commands
Data Source
AI summary
Examples described herein relate to triggering voice assistant(s) on a network microphone device (NMD). An NMD is a networked computing device that typically includes an arrangement of microphones, such as a microphone array, that is configured to detect sound present in the NMD's environment. Once the voice assistant is triggered, the NMD may start recording voice input as a potential voice command. Within examples, the NMD may operate in a wakewordless mode if certain conditions are met. These conditions may involve detecting user proximity in one of multiple different ranges. For instance, an example NMD may monitor for user proximity in a first range from the playback device via at least one touch-sensitive sensor and/or user line-of-sight in a second range that is further from the playback device than the first range. When either user proximity or user line-of-sight is detected, the NMD may enables the wakewordless mode.


