Spatially Selective Wake-Up Word Detection Using Acoustic Zone Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech-enabled systems face challenges in accurately detecting wake-up words in noisy environments with multiple speakers, leading to user frustration and unsatisfactory responsiveness.
Innovation Solution
The system employs multiple microphones to detect acoustic information from spatially partitioned acoustic zones, using techniques like time of flight, angle of arrival, and beamforming to isolate sound sources and perform separate speech recognition on each zone, improving wake-up word detection and subsequent speech recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the system processes all acoustic input for wake-up word detection, then no speech is missed, but processing cost increases and accuracy decreases in noisy environments
Solution Approach 1:
The patent divides the acoustic environment into multiple acoustic zones using beamforming techniques. Each zone is processed independently for wake-up word detection, allowing the system to focus computational resources on specific spatial regions rather than processing all acoustic input uniformly. This segmentation improves detection accuracy in noisy environments while reducing overall processing cost.
Solution Approach 2:
The system applies different processing strategies to different acoustic zones based on local characteristics. Zones with higher probability of containing wake-up words receive more intensive processing, while quieter zones receive less processing. This local quality approach optimizes the balance between detection accuracy and processing cost by adapting resource allocation to local acoustic conditions.
2Ease of operation
If the system uses manual triggers to engage speech recognition, then processing cost is reduced, but ease of operation decreases
Solution Approach 1:
The system performs preliminary wake-up word detection in acoustic zones before engaging full speech recognition processing. This preliminary action allows the system to remain in a low-power state until a wake-up word is detected, at which point full processing is activated. This approach maintains hands-free operation while reducing processing cost during idle periods.
3Measurement precision
If the system processes speech from all directions, then no wake-up word is missed, but measurement precision decreases in noisy environments with multiple speakers
Solution Approach 1:
The patent segments the acoustic space into distinct zones and processes speech from each zone separately. This allows the system to identify which zone contains a wake-up word and focus subsequent speech recognition on that specific zone, improving accuracy even when multiple speakers are present in different locations.
Solution Approach 2:
The system uses beamforming as an intermediary technique to spatially filter and separate acoustic signals from different directions. This intermediary processing step enables the system to isolate wake-up words from specific zones before passing them to the speech recognition engine, improving precision in multi-speaker environments.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enhances wake-up word detection accuracy and robustness in noisy environments by targeting the correct acoustic zone for speech recognition, allowing users to interact effectively even amidst other sound sources.
Implementation Method 1
using techniques like time of flight, angle of arrival, and beamforming to isolate sound sources
Implementation Method 2
using techniques like time of flight, angle of arrival, and beamforming to isolate sound sources
Implementation Method 3
using techniques like time of flight, angle of arrival, and beamforming to isolate sound sources
Data Source
AI summary
According to some aspects, a system for detecting a designated wake-up word is provided, the system comprising a plurality of microphones to detect acoustic information from a physical space having a plurality of acoustic zones, at least one processor configured to receive a first acoustic signal representing the acoustic information received by the plurality of microphones, process the first acoustic signal to identify content of the first acoustic signal originating from each of the plurality of acoustic zones, provide a plurality of second acoustic signals, each of the plurality of second acoustic signals substantially corresponding to the content identified as originating from a respective one of the plurality of acoustic zones, and performing automatic speech recognition on each of the plurality of second acoustic signals to determine whether the designated wake-up word was spoken.


