Keyword-Based Audio Localization with Beamforming for Noisy Speech
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Devices struggle to distinguish between a person speaking and other ambient sounds, leading to difficulty in understanding voice commands due to background noise or multiple speakers.
Innovation Solution
Implementing multiple listening zones using beamformers to detect a keyword, determine the direction and distance of the speaker, and form active acoustic beams to enhance speech recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the device listens for keywords from multiple directions using multiple listening zones, then the device can detect keywords spoken by anyone in the environment, but the device cannot distinguish between the speaker and background noise
Solution Approach 1:
The patent divides the listening environment into multiple acoustic zones (first acoustic zone, second acoustic zone, etc.) with each zone having associated beamformers that point in specific directions. This segmentation allows the system to simultaneously monitor multiple directions for keywords while also being able to identify which zone the keyword was spoken in, enabling both broad detection coverage and directional speaker identification.
Solution Approach 2:
The patent introduces a spatial dimension by using beamformers that create directional acoustic zones in three-dimensional space. Each beamformer points in a specific direction and creates an acoustic zone that can be distinguished from other zones. This adds the dimension of spatial location to the keyword detection process, allowing the system to not only detect keywords from multiple directions but also to identify the specific direction from which the keyword was spoken.
2Measurement precision
If the device uses multiple beamformers pointing in various directions to detect keywords, then the device can identify the direction of the speaker, but the device complexity increases
Solution Approach 1:
The patent makes the beamformers multi-functional by using them for both keyword detection and speaker location identification. The same beamformers that detect keywords in specific acoustic zones also provide the spatial information needed to determine speaker direction. This eliminates the need for separate systems for detection and localization, reducing overall device complexity while maintaining precision.
Solution Approach 2:
The patent performs preliminary action by pre-configuring multiple beamformers to point in various directions before keyword detection begins. These beamformers are set up in advance to create overlapping acoustic zones that cover the entire listening environment. This preliminary configuration allows the system to immediately begin both keyword detection and directional identification without additional processing complexity during operation.
3Measurement precision
If the device forms active acoustic beams directed toward the speaker after keyword detection, then the device can enhance speech recognition, but the device must switch between different listening modes
Solution Approach 1:
The patent implements dynamic switching between different listening modes: a first listening mode for keyword detection using multiple acoustic zones, and a second listening mode for enhanced speech recognition using active acoustic beams directed at the identified speaker location. The system dynamically transitions between these modes based on whether keyword detection is needed or enhanced speech recognition is required, allowing the device to adapt its behavior to different operational contexts.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enhances speech recognition by focusing on the speaker's voice, reducing background noise interference and improving command understanding.
Implementation Method 1
the device may implement multiple listening zones, such as using one or more beamformers pointing in various directions around a horizontal plane and/or a vertical plane
Implementation Method 2
form one or more active acoustic beams directed toward the person speaking
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Systems, apparatuses, and methods are described for determining a direction associated with a detected spoken keyword, forming an acoustic beam in the determined direction, and listening for subsequent speech using the acoustic beam in the determined direction.