Multi-Microphone Beamforming for Moving-User Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speech recognition performance deteriorates when a user is moving due to fixed beamforming directions, necessitating a method to adapt the electronic device's location and direction based on the user's location.
Innovation Solution
An electronic device equipped with multiple microphones that identify a designated call word, perform speech recognition, and adjust its location using phase differences and beamforming to enhance target voice signals while reducing noise and echo.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If beamforming is used to track user location, then speech recognition accuracy is improved, but device complexity increases
Solution Approach 1:
The patent divides the speech recognition process into two stages: first identifying call words using a single microphone to reduce complexity, then performing full speech recognition using multiple microphones with beamforming only when necessary. This segmentation allows the system to achieve high accuracy when needed while maintaining low complexity during normal operation.
Solution Approach 2:
The patent dynamically adjusts the number of microphones active in beamforming based on whether call words have been identified. When call words are detected, the system transitions from single-microphone mode to multi-microphone beamforming mode, adapting the device complexity to the current operational requirements and improving accuracy at the appropriate moment.
2Measurement precision
If multiple microphones are used for speech recognition, then speech recognition accuracy is improved, but computational load increases
Solution Approach 1:
The patent segments the speech recognition task into call word identification (using one microphone, low computational load) and continuous word recognition (using multiple microphones, high computational load). By separating these functions, the system performs heavy computations only when necessary, reducing overall power consumption while maintaining high accuracy when needed.
Solution Approach 2:
The system performs preliminary call word identification using a single microphone before engaging the full multi-microphone beamforming system. This preliminary action filters out unnecessary computations by identifying trigger words first, then activating the computationally intensive speech recognition only for relevant segments, thereby reducing total computational load.
3Device complexity
If beamforming direction is fixed, then device complexity is reduced, but speech recognition performance deteriorates when user moves
Solution Approach 1:
The patent implements dynamic beamforming that adjusts its direction and orientation based on detected user location and call word identification. The system transitions from a fixed beamforming configuration to an adaptive one that tracks user movement, improving reliability while managing complexity through conditional activation only when call words are detected.
Solution Approach 2:
The system uses feedback from call word identification and user location detection to dynamically adjust beamforming parameters. By continuously monitoring user position and call word detection results, the system modifies its beamforming direction and orientation in real-time, ensuring optimal performance while maintaining manageable complexity through feedback-driven adaptation.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Improves speech recognition accuracy by selectively processing voice signals from a single microphone, reducing computational load, and accurately tracking the user's location for enhanced performance.
Implementation Method 1
track a location of a user associated with the plurality of voice signals by using a phase difference between the plurality of voice signals
Implementation Method 2
identify a second voice signal, which is obtained by enhancing a target voice corresponding to an utterance of the user, from the plurality of voice signals obtained by performing beamforming toward the location of the user
Data Source
AI summary
An electronic device may include a plurality of microphones, a speaker, a processor, and a memory. The processor may be configured to obtain a plurality of voice signals by using the plurality of microphones, to identify a designated call word, by performing speech recognition on a first voice signal obtained from a first microphone, from among the plurality of voice signals according to a passage of time in obtaining the plurality of voice signals, to obtain a second time point, which precedes an utterance time corresponding to the designated call word from a first time point at which the designated call word is identified completely, and to perform speech recognition on one portion, which is obtained from the second time point, from among the plurality of voice signals.


