Azimuth Estimation Using Spatial Spectrum and Wakeup Scores
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Far-field speech recognition systems face challenges in accurately estimating the azimuth of a target voice due to environmental noise and reverberation, which affects user experience and system performance.
Innovation Solution
An azimuth estimation method that involves real-time multi-channel sampling, buffering, wakeup word detection, and spatial spectrum estimation to determine the azimuth of a target voice using the highest wakeup word score, thereby reducing noise interference and improving estimation accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If beamforming algorithm is used to improve voice signal quality, then speech recognition performance is improved, but the system becomes extremely sensitive to azimuth accuracy
Solution Approach 1:
The patent performs preliminary azimuth estimation using spatial spectrum analysis before applying beamforming. By pre-estimating the azimuth with multiple candidate values and evaluating their likelihoods, the system prepares optimal beamforming parameters in advance, reducing sensitivity to final azimuth accuracy while maintaining speech recognition performance.
Solution Approach 2:
The patent changes the parameter approach by using multiple candidate azimuth values instead of a single fixed azimuth. By evaluating the likelihood of different azimuth candidates and selecting the optimal one, the system makes the beamforming process more robust to azimuth estimation variations, thereby reducing sensitivity.
2Measurement precision
If spatial spectrum estimation is performed on all buffered signals, then azimuth estimation accuracy is improved, but computational complexity increases
Solution Approach 1:
The patent segments the azimuth estimation process by dividing it into multiple candidate azimuth evaluations. Instead of performing a single comprehensive spatial spectrum estimation, the system divides the task into evaluating multiple discrete candidate azimuths, which can be processed more efficiently and allows for selective refinement of promising candidates.
Solution Approach 2:
The patent applies partial action by performing spatial spectrum estimation selectively on candidate azimuths rather than exhaustively searching all possible azimuths. The system evaluates a limited set of candidate directions and focuses computational resources on the most promising ones, achieving good accuracy without the full computational burden of exhaustive search.
3Reliability
If multi-channel sampling is performed in real-time, then noise interference is reduced, but processing time increases
Solution Approach 1:
The patent performs preliminary processing of multi-channel sampling signals by buffering them and performing preliminary spectral analysis before final azimuth determination. This preliminary action allows the system to prepare signal representations in advance, reducing the processing time required for final noise-resistant azimuth estimation while maintaining noise reduction benefits.
Data Source
AI summary
Embodiments of this application discloses an azimuth estimation method performed at a computing device, the method including: obtaining, in real time, multi-channel sampling signals and buffering the multi-channel sampling signals; performing wakeup word detection on one or more sampling signals of the multi-channel sampling signals, and determining a wakeup word detection score for each channel of the one or more sampling signals; performing a spatial spectrum estimation on the buffered multi-channel sampling signals to obtain a spatial spectrum estimation result, when the wakeup word detection scores of the one or more sampling signals indicates that a wakeup word exists in the one or more sampling signals; and determining an azimuth of a target voice associated with the multi-channel sampling signals according to the spatial spectrum estimation result and a highest wakeup word detection score, thereby improving the accuracy of the azimuth estimation in a voice interaction process.


