Robust Speaker Localization Using Relative Transfer Covariance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio signal processing systems fail to effectively isolate target audio from noise in noisy environments, especially when the target audio is weaker than the noise source, leading to poor localization and enhancement of speech signals.
Innovation Solution
The system employs a modified Time Difference of Arrival (TDOA) or Direction of Arrival (DOA) estimation method using a Relative Transfer Function (RTF) that nulls dominant noise sources, allowing for robust localization and enhancement of target audio by constructing a directional covariance matrix aligned across frequency bands and minimizing beam power subject to distortionless criteria.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional audio signal processing is used, then the system can process audio inputs, but it fails to effectively isolate target audio from noise when the target audio is weaker than the noise source
Solution Approach 1:
The system changes parameters by using relative transfer functions across multiple frequency bands to characterize audio sources, and by optimizing beamforming parameters to enhance target audio while suppressing noise. This allows the system to achieve accurate localization and enhancement even when target audio is weaker than noise sources.
Solution Approach 2:
The system moves from simple spatial filtering to a multi-dimensional approach by constructing directional covariance matrices aligned across frequency bands. This multi-frequency band analysis with directional information provides additional dimensions for distinguishing target audio from noise, improving localization accuracy in noisy environments.
2Measurement precision
If the system uses a plurality of audio input components to receive audio signals, then it can capture audio from different directions, but it struggles to determine the relative location of the audio source in noisy environments
Solution Approach 1:
The system uses relative transfer functions as key parameters to characterize the acoustic path from audio sources to microphones. By estimating these transfer functions across frequency bands and using them to construct directional covariance matrices, the system achieves accurate localization without requiring overly complex processing of raw microphone signals.
Solution Approach 2:
The system replaces complex mechanical/spatial filtering approaches with signal processing-based beamforming. By using relative transfer functions and covariance matrix operations, the system achieves sophisticated spatial filtering and source localization through mathematical operations rather than simple physical array configurations.
3Reliability
If the system processes audio signals to enhance target audio, then it can improve speech recognition, but it cannot effectively distinguish target speech from stronger noise sources when the signal-to-noise ratio is negative
Solution Approach 1:
The system changes the processing approach by using relative transfer functions estimated from the audio signals themselves, rather than relying on predefined acoustic models. By constructing directional covariance matrices aligned across frequency bands and applying beamforming with these parameters, the system can enhance target audio even in negative signal-to-noise ratio conditions, preserving speech quality for accurate recognition.
Solution Approach 2:
The system uses feedback by iteratively estimating relative transfer functions from the received audio signals and using these estimates to guide the beamforming process. The directional covariance matrices are constructed based on these estimated transfer functions, creating a feedback loop that adapts to the actual acoustic environment and improves target audio enhancement reliability.
Data Source
AI summary
Systems and methods include a plurality of audio input components configured to generate a plurality of audio input signals, and a logic device configured to receive the plurality of audio input signals, determine whether the plurality of audio signals comprise target audio associated with an audio source, estimate a relative location of the audio source with respect to the plurality of audio input components based on the plurality of audio signals and a determination of whether the plurality of audio signals comprise the target audio, and process the plurality of audio signals to generate an audio output signal by enhancing the target audio based on the estimated relative location. The logic device is further configured to use relative transfer-based covariance to construct directional covariance matrix aligned across frequency bands and find a direction that minimizes beam power subject to distortionless criteria.


