TJR Weighted Filter for Speech Enhancement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing two-microphone speech enhancement technologies, particularly those using GSC structures, face challenges in accurately estimating target sound sources due to reliance on voice activity detectors and potential impairment by non-stationary coherent interference, leading to cancellation of desired signals.
Innovation Solution
A spatially pre-processed target-to-jammer ratio (TJR) weighted filter is introduced, utilizing an FFT module, beamformer, reference generator, PSD estimator, and noise estimator to determine the TJR, which switches between optimized and new Wiener solutions based on TJR values to prevent signal cancellation and effectively eliminate noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If voice activity detector is used to determine target source presence, then estimation can be started/stopped, but performance relies on VAD accuracy which may fail in complex acoustic environments
Solution Approach 1:
The patent changes the detection parameter from binary voice activity detection to continuous target-to-jammer ratio calculation. By computing the ratio of target signal power to jammer power spectral densities, the system can adaptively determine target presence and switch between different estimation methods (VAD-based vs. TJR-based) depending on the acoustic environment, thereby improving both reliability and adaptability.
2Productivity
If noise estimation is performed without target signal presence confirmation, then estimation can be performed continuously, but desired signal may be cancelled
Solution Approach 1:
The patent implements a feedback mechanism where the target-to-jammer ratio is continuously calculated and used to control the noise estimation process. When TJR exceeds a threshold indicating target presence, the system switches to a different estimation method that preserves the desired signal. This feedback loop allows continuous operation while preventing signal cancellation through adaptive control.
Solution Approach 2:
The system dynamically switches between different noise estimation methods based on real-time TJR calculation. The estimation process is not static but adapts its behavior based on the detected acoustic scene, transitioning between VAD-based mode (when target is absent) and TJR-based mode (when target is present), thereby maintaining both productivity and reliability.
3Ease of manufacture
If conventional GSC structure is used for speech enhancement, then beamforming and null steering are achieved, but non-stationary coherent interference impairs performance
Solution Approach 1:
The patent extends the conventional static GSC structure by adding dynamic adaptation capabilities. The system continuously calculates the target-to-jammer ratio and adjusts the noise estimation method accordingly. This dynamic extension allows the system to handle non-stationary coherent interference by switching between estimation methods, while maintaining the ease of implementing the standard GSC structure for beamforming and null steering.
Data Source
AI summary
The present invention provides a spatially pre-processed target-to-jammer ratio weighted filter and a method thereof, which uses two microphones to receive audio signals. The audio signals are divided into a plurality of sinusoidal waves by a fast Fourier transform (FFT) module, and a beamformer uses the sinusoidal waves to generate beamformed signals. A reference generator generates at least one reference signal. The beamformed signals and reference signals are used to work out power spectral densities (PSD), and a target-to-jammer ratio (TJR) is worked out with the power spectral densities. TJR is used to determine whether a sound source exists. According to the determination result, a noise estimator is switched to eliminate noise from the beamformed signals and generate output signals. An inverse fast Fourier transform (IFFT) module recombines the output signals and then outputs the recombined signals.


