Audio Source Azimuth Positioning With Echo-Canceled Time-Frequency Weighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio interaction devices face challenges in accurately determining the azimuth of a target voice due to sensitivity to the accuracy of the specified azimuth, which affects the quality and performance of voice interaction.
Innovation Solution
An audio recognition method that involves obtaining audio signals in multiple directions, performing echo cancellation, calculating weights of time-frequency points, and using a weighted covariance matrix to enhance the accuracy of sound source azimuth estimation, thereby reducing interference from noise and echoes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If beamforming algorithm is used to enhance target voice, then voice signal quality is improved, but azimuth accuracy requirement becomes more stringent
Solution Approach 1:
The patent applies preliminary action by performing echo cancellation before azimuth estimation. The echo-canceled audio signals are processed to obtain time-frequency domain expressions and calculate weights before the azimuth is determined. This preliminary processing removes echo interference that would otherwise distort the azimuth estimation, allowing the beamforming algorithm to operate with more accurate azimuth values without requiring extremely precise measurements.
Solution Approach 2:
The patent introduces an intermediary approach by using weighted covariance matrices as a mediator between the raw audio signals and the final azimuth estimation. The weights, calculated from time-frequency domain expressions of echo-canceled signals, serve as a mediator that adjusts the contribution of different signal components. This intermediary mechanism allows the system to achieve accurate azimuth estimation even in complex environments with echo and noise, resolving the contradiction between voice quality enhancement and azimuth measurement precision.
2Adaptability or versatility
If azimuth estimation is performed in complex environments with echo and noise, then audio interaction capability is maintained, but azimuth accuracy deteriorates
Solution Approach 1:
The patent applies preliminary action by performing echo cancellation before azimuth estimation. The echo-canceled audio signals are processed to obtain time-frequency domain expressions and calculate weights before the azimuth is determined. This preliminary processing removes echo interference that would otherwise distort the azimuth estimation, allowing the beamforming algorithm to operate with more accurate azimuth values without requiring extremely precise measurements.
Solution Approach 2:
The patent applies parameter changes by transforming the audio signals from the time domain to the time-frequency domain using Short-Time Fourier Transform (STFT). This transformation changes the representation parameters of the signals, allowing for more accurate extraction of time-frequency domain expressions that can be used to calculate weights. These parameter changes enable the system to distinguish between echo, noise, and target voice more effectively, maintaining azimuth accuracy in complex environments.
Data Source
AI summary
This application discloses a method for positioning a target audio signal by a computer device. The method includes: performing echo cancellation on the audio signals collected in a plurality of directions in a space, the audio signals comprising a target-audio direct signal; obtaining weights of a plurality of time-frequency points in the echo-canceled audio signals, a weight of each time-frequency point indicating a relative proportion of the target-audio direct signal in the echo-canceled audio signals at the time-frequency point; obtaining a weighted audio signal energy distribution of the audio signals in the plurality of directions by using the weights of the plurality of time-frequency points in the echo-canceled audio signals; and obtaining a sound source azimuth corresponding to the target-audio direct signal in the audio signals by using the weighted audio signal energy distribution of the audio signals in the plurality of directions.


