Multi-microphone Speech Intelligibility via Spectral Variance Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Hearing aid users face difficulties in understanding speech in reverberant environments due to the lack of effective signal processing algorithms that can efficiently separate direct sound from reverberation and noise, leading to reduced speech intelligibility.
Innovation Solution
A method and system for processing noisy audio signals using two or more microphones to estimate the spectral variances of target and reverberant signal components dynamically, employing Maximum Likelihood Estimation (MLE) and spatial filtering, which models the reverberant sound field as isotropic and assumes a known speaker direction, allowing for joint estimation of spectral variances of target and interference.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multi-microphone systems and linear prediction models are used to eliminate reverberation, then speech intelligibility is improved, but the device complexity increases
Solution Approach 1:
The patent replaces complex mechanical signal processing approaches with statistical signal processing methods. Specifically, it uses Maximum Likelihood Estimation (MLE) to estimate spectral variances of target and noise components, substituting traditional linear prediction and beamforming methods with a statistically optimal approach that is computationally more efficient and easier to implement in hearing aids.
Solution Approach 2:
The patent changes the fundamental parameters used for reverberation suppression by estimating spectral variances in the frequency domain rather than processing time-domain signals through complex filters. This parameter transformation from time-domain to frequency-domain processing, combined with MLE-based variance estimation, simplifies the overall system while maintaining effectiveness.
2Reliability
If adaptive beamforming and postfiltering are applied to reduce noise components, then speech intelligibility is improved, but the difficulty of detecting and measuring increases
Solution Approach 1:
The patent substitutes adaptive beamforming and postfiltering with a Maximum Likelihood Estimation approach that directly estimates spectral variances. This statistical method replaces the multi-stage adaptive processing with a unified estimation framework that is more straightforward to implement and tune, reducing the difficulty of detection and measurement while maintaining noise reduction effectiveness.
3Reliability
If isotropic noise suppression is applied by symmetric microphone arrays, then speech intelligibility is improved, but the device complexity increases
Solution Approach 1:
The patent creates a universal noise suppression method based on MLE that can handle both isotropic and non-isotropic noise fields. By formulating the problem in terms of spectral variance estimation rather than relying on specific array geometries or noise assumptions, the solution becomes universally applicable to different microphone configurations and acoustic environments, reducing device complexity.
Data Source
AI summary
The application relates to an audio processing system and a method of processing a noisy (e.g. reverberant) signal comprising first (v) and optionally second (w) noise signal components and a target signal component (x), the method comprising a) Providing or receiving a time-frequency representation Yi(k,m) of a noisy audio signal yi at an ith input unit, i=1, 2, . . . , M, where M≧2; b) Providing (e.g. predefined spatial) characteristics of said target signal component and said noise signal component(s); and c) Estimating spectral variances or scaled versions thereof λV, λX of said first noise signal component v (representing reverberation) and said target signal component x, respectively, said estimates of λV and λX being jointly optimal in maximum likelihood sense, based on the statistical assumptions that a) the time-frequency representations Yi(k,m), Xi(k,m), and Vi(k,m) (and Wi(k,m)) of respective signals yi(n), and signal components xi, and vi (and wi) are zero-mean, complex-valued Gaussian distributed, b) that each of them are statistically independent across time m and frequency k, and c) that Xi(k,m) and Vi(k,m) (and Wi(k,m)) are uncorrelated. An advantage of the invention is that it provides the basis for an improved intelligibility of an input speech signal. The invention may e.g. be used for hearing assistance devices, e.g. hearing aids.


