Noise Estimation for Echo Cancellation in Multi-Microphone Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional methods for speech enhancement in noisy environments, such as beamforming and echo cancellation, are inadequate in reducing background noise and reverberation, especially in rooms with long reverberation times, leading to decreased sound quality and increased error rates in voice recognition systems.
Innovation Solution
A method for estimating the time-varying and frequency-dependent inter-microphone noise covariance matrix, which is optimal in a maximum likelihood sense, is used to improve noise reduction and speech enhancement, allowing for accurate estimation of noise levels and calculation of frequency-dependent gain weights for post-filtering.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If traditional beamforming and echo cancellation methods are used, then spatial filtering is achieved, but background noise and reverberation are not sufficiently reduced
Solution Approach 1:
A post-filter is introduced as an intermediary component between the beamformer and the output signal. This post-filter uses noise power spectral density estimates to further attenuate noise and reverberation components that escaped the beamforming stage, thereby achieving better noise reduction without compromising the spatial filtering capability
Solution Approach 2:
The patent replaces traditional mechanical/acoustic noise reduction approaches with a signal processing-based post-filter that uses statistical noise estimates. This substitution allows for more precise and adaptive noise reduction by manipulating the spectral content of the signal rather than relying solely on spatial filtering
2Measurement precision
If noise power spectral density is estimated from beamformer output, then single-channel noise tracking is achieved, but performance is suboptimal compared to multi-microphone approaches
Solution Approach 1:
The patent merges the advantages of both single-channel noise tracking and multi-microphone noise estimation by combining beamformer output with direct multi-microphone signal processing. The post-filter leverages noise PSD estimates derived from multiple microphone inputs, achieving superior noise estimation accuracy that utilizes the spatial information from the microphone array
Solution Approach 2:
The noise estimation system is designed to be multi-functional, serving both as a basis for beamforming and as input for post-filtering. The same noise PSD estimates are used to control both the beamformer weights and the post-filter gains, making the system adaptable to different operating conditions and maximizing the utility of the noise estimation
3Object-generated harmful factors
If spectral weights are applied to beamformer output, then residual echo is reduced, but target speech content must be preserved
Solution Approach 1:
The post-filter applies frequency-dependent spectral weights that are adapted to the local characteristics of each frequency bin. By using noise PSD estimates and signal-to-noise ratio calculations on a per-frequency-bin basis, the filter can aggressively attenuate echo and noise in frequencies where the target speech is weak, while preserving frequencies where the target speech dominates, thus achieving local optimization of speech preservation and echo reduction
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The application relates to a method for audio signal processing. The application further relates to a method of processing signals obtained from a multi-microphone system. The object of the present application is to reduce undesired noise sources and residual echo signals from an initial echo cancellation step. The problem is solved by receiving M communication signals in frequency subbands where M is at least two; processing the M subband communication signals in each subband with a blocking matrix (203,303,403) of M rows and N linearly independent columns in each subband, where N>=1 and N<M, to obtain N target-cancelled signals in each subband; processing the M subband communication signals and the N target-cancelled signals in each subband with a set of beamformer coefficients (204,304,404) to obtain a beamformer output signal in each subband; processing the communication signals with a target absence detector (309) to obtain a target absence signal in each subband; using the target absence signal to obtain an inverse target-cancelled covariance matrix of order N (310,410) in each band; processing the N target-cancelled signals in each subband with the inverse target-cancelled covariance matrix in a quadratic form (312, 412) to yield a real-valued noise correction factor in each subband; using the target absence signal to obtain an initial estimate (311, 411) of the noise power in the beamformer output signal averaged over recent frames with target absence in each subband; multiplying the initial noise estimate with the noise correction factor to obtain a refined estimate (417) of the power of the beamformer output noise signal component in each subband; processing the refined estimate of the power of the beamformer output noise signal component with the magnitude of the beamformer output to obtain a postfilter gain value in each subband; processing the beamformer output signal with the postfilter gain value (206,306,406) to obtain a postfilter output signal in each subband; processing the postfilter output subband signals through a synthesis filterbank (207,307,407) to obtain an enhanced beamformed output signal where the target signal is enhanced by attenuation of noise signal components. This has the advantage of providing improved sound quality and reduction of undesired signal components such as the late reverberant part of an acoustic echo signal. The invention may e.g. be used for headsets, hearing aids, active ear protection systems, mobile telephones, teleconferencing systems, karaoke systems, public address systems, mobile communication devices, hands-free communication devices, voice control systems, car audio systems, navigation systems, audio capture, video cameras, and video telephony.