Selective Audio Source Enhancement via Multistage BSS
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech enhancement methods fail to provide satisfactory performance in far-field applications with limited channels and significant reverberation, leading to inadequate energy propagation and poor automatic speech recognition in noisy environments.
Innovation Solution
The implementation of a selective audio source enhancement system using Blind Source Separation (BSS) techniques, specifically a multistage processing approach involving source detection, weighted natural gradient, constrained independent component analysis (ICA), and spectral filtering, optimized for limited hardware resources, allowing for robust speech recognition and noise suppression.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional speech enhancement methods (single or multiple channel) are used, then processing simplicity is maintained, but performance is insufficient in far-field applications with limited channels and significant reverberation
Solution Approach 1:
The patent applies segmentation by dividing the speech enhancement task into multiple stages: first estimating the speech power spectral density, then estimating the noise power spectral density, and finally computing the enhanced speech signal. This multi-stage approach allows complex processing to be broken down into manageable steps that can be implemented with limited hardware resources while maintaining high performance in far-field conditions with reverberation.
2Reliability
If beam forming methods are used to enhance signals from predefined spatial directions, then directional signal enhancement is improved, but effectiveness decreases when energy propagation over steering geometrical direction is limited
Solution Approach 1:
The patent changes the fundamental parameters of the enhancement approach by moving from spatial filtering (beam forming) to spectral processing. Instead of relying on directional energy propagation, the method estimates speech and noise power spectral densities and processes signals in the frequency domain. This allows effective enhancement even when directional energy propagation is limited, as the method does not depend on strong spatial separation between signal and noise.
3Reliability
If continuous signal-to-noise ratio estimation is used in discrete time-spectral domain, then enhancement effectiveness is improved for stationary noise, but performance degrades when noise exhibits high energy variation (non-stationarity)
Solution Approach 1:
The patent implements dynamics by using time-varying estimates of speech and noise power spectral densities. The method continuously updates these estimates adaptively, allowing the enhancement process to respond to changing noise conditions. This dynamic approach enables effective handling of non-stationary noise with high energy variation, as the system can adjust its estimates in real-time rather than relying on fixed stationary assumptions.
Data Source
AI summary
A selective audio source enhancement system includes a processor and a memory, and a pre-processing unit configured to receive audio data including a target audio signal, and to perform sub-band domain decomposition of the audio data to generate buffered outputs. In addition, the system includes a target source detection unit configured to receive the buffered outputs, and to generate a target presence probability corresponding to the target audio signal, as well as a spatial filter estimation unit configured to receive the target presence probability, and to transform frames buffered in each sub-band into a higher resolution frequency-domain. The system also includes a spectral filtering unit configured to retrieve a multichannel image of the target audio signal and noise signals associated with the target audio signal, and an audio synthesis unit configured to extract an enhanced mono signal corresponding to the target audio signal from the multichannel image.


