Dynamic Wake Word Thresholds for Spatially Separated Audio Sources
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Wake words are difficult to detect in noisy environments with multiple audio sources due to interference from non-voice-based audio signals, leading to incorrect voice command detection or failure to detect the wake word.
Innovation Solution
A voice-controlled device employs a fixed microphone array with multiple wake word detection engines, each using dynamically adjustable sensitivity thresholds based on spatial and temporal factors, to distinguish and accurately detect wake words amidst various audio sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the device focuses on the loudest audio source, then the detection sensitivity is improved, but the reliability deteriorates because the loudest source may not be a voice
Solution Approach 1:
The patent applies local quality by creating different detection characteristics for different spatial locations. The system divides the environment into multiple sectors and assigns different detection thresholds and sensitivity levels to each sector based on local conditions such as expected voice presence and background noise characteristics, rather than using a uniform detection approach throughout the entire environment.
Solution Approach 2:
The system dynamically changes detection parameters including sensitivity thresholds and sector assignments based on environmental conditions. The detection parameters are adjusted according to the detected audio landscape, allowing the system to adapt to varying noise levels, voice presence, and spatial distribution of audio sources in real-time.
2Reliability
If the sensitivity threshold is lowered to detect wake words in noisy environments, then the wake word detection capability is improved, but false detections from non-voice audio sources increase
Solution Approach 1:
The patent segments the detection system into multiple independent wake word detection engines, each responsible for specific spatial sectors. This segmentation allows the system to apply different sensitivity thresholds and detection strategies to different areas, reducing false detections by isolating voice detection in sectors where voices are expected while being less sensitive in sectors dominated by non-voice audio sources.
Solution Approach 2:
The system introduces spatial sector assignment as an intermediary layer between the audio input and wake word detection engines. This intermediary mechanism filters and directs audio signals to appropriate detection engines based on their spatial origin, preventing non-voice audio sources from triggering wake word detections in inappropriate sectors.
3Adaptability or versatility
If multiple detection engines are used to cover different audio sources, then the coverage is improved, but the device complexity increases
Solution Approach 1:
The patent implements multi-functionality by designing wake word detection engines that can operate across multiple sectors and handle different types of audio sources. Each detection engine is capable of performing wake word detection while also contributing to the overall environmental audio understanding, reducing the need for separate specialized components for each function.
Data Source
Figure 1
Figure 2A
Figure 2B
AI summary
Systems and methods are provided for detecting wake words. An electronic device detects an audio signal; identifies two angular sections as first and second sources of audio associated with the audio signal; processes the audio signal at two wake word detection engines (220), where each detection engine is associated with a respective angular section; determines, based on the processing at the wake word detection engines, whether the audio signal represents a wake word for the electronic device; and in accordance with a determination that the audio signal does represent a wake word, adjusts a wake word detection threshold for at least one of the wake word detection engines. Upon detecting a wake word in a first angular section, the audio signal associated with one or more of the other angular sections is used as a noise reference to perform noise cancelation in the first wake word detection engine for the first angular section.