Multibeam Keyword Detection Using Angle-Based Audio Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-controlled user interfaces face challenges in distinguishing user inputs from background noises and third-party speech, leading to false alarms and reduced true positive activation rates, especially in noisy environments.
Innovation Solution
A multibeam keyword detection system that uses a microphone array and beamforming technology to isolate sounds by angle of arrival, separating them into distinct beams to improve signal-to-noise ratio and reduce false alarms by processing sounds independently based on their angles of arrival, thereby enhancing keyword detection accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If increasingly strict rules are applied to keyword detection, then false alarm rate is reduced, but true positive activation rate deteriorates
Solution Approach 1:
The audio signal is segmented into multiple beams based on angle of arrival, with each beam processed by a separate keyword detector. This segmentation allows the system to apply detection rules selectively to different spatial regions, reducing false alarms from background noise while maintaining sensitivity to user-directed speech.
Solution Approach 2:
Different detection sensitivity levels are applied to different spatial regions. Beams directed toward the user receive more lenient detection rules, while other regions use stricter rules to filter noise. This local quality approach resolves the contradiction by making detection criteria spatially adaptive rather than uniformly strict.
2Reliability
If strict detection rules are applied to reduce false alarms, then reliability improves, but device complexity increases
Solution Approach 1:
The detection system is segmented into multiple independent beam processors, each handling a specific angular region. This modular segmentation distributes the computational complexity across parallel units rather than requiring a single complex detector, making the system more manageable while achieving better false alarm rejection through spatial filtering.
3Adaptability or versatility
If environmental sounds are processed together with keyword detection, then comprehensive monitoring is achieved, but false association of environmental sounds with keywords increases
Solution Approach 1:
The audio spectrum is segmented into multiple beams based on angle of arrival, allowing environmental sounds to be monitored across all directions while keyword detection is performed separately in each beam. This prevents false associations by spatially separating keyword processing from general environmental monitoring.
Solution Approach 2:
Keyword detection is extracted as a separate function from general environmental sound monitoring. The system first segments audio into beams, then applies keyword detection specifically to relevant beams while maintaining the ability to monitor all environmental sounds, thus preventing false keyword associations with background noise.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The system effectively reduces false alarm rates and improves the detection of valid user inputs by isolating noise and enhancing the signal-to-noise ratio, allowing for more accurate keyword detection and command recognition in noisy environments.
Implementation Method 1
a beamforming engine comprising a plurality of beamformers coupled to receive the composite audio signal and configured to form a plurality of beam signals from the composite audio signal, each of the beam signals having an associated beam angle and comprising audio data corresponding to sounds captured by the microphone array within a range of angles of arrival with respect to the beam angle
Data Source
AI summary
A system and method provides for multibeam keyword detection. A composite audio signal may include sound components. The system and method groups the sound components into subsets based on the angles of arrival of sound components. Keyword detectors evaluate each subset and determine whether a keyword is present.


