Maximum Likelihood Beamforming for Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing beamforming methods, such as MVDR, struggle to accurately estimate noise signals in environments where a target sound is present, leading to deteriorated speech recognition performance due to noise distortion.
Innovation Solution
A beamforming method using maximum likelihood estimation is employed, where a filter is estimated to maximize the log likelihood of a probability density function under the assumption that the beamforming output signal satisfies a complex generalized Gaussian or complex gamma distribution, with the goal of improving sound quality and speech recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If MVDR beamforming algorithm is used to minimize beamforming output signal, then noise removal capability is improved, but speech recognition performance deteriorates when target sound is present
Solution Approach 1:
The patent changes the fundamental parameter of filter estimation from minimizing output power (MVDR) to maximizing log-likelihood under complex generalized Gaussian distribution assumption. This parameter change allows accurate noise estimation even when target sound is present, resolving the contradiction between noise removal and speech recognition performance
Solution Approach 2:
The patent replaces the traditional MVDR mechanical optimization approach with a statistical maximum likelihood estimation approach based on complex generalized Gaussian distribution. This substitution enables more accurate noise signal estimation in the presence of target sound, improving speech recognition reliability
2Object-affected harmful factors
If filter estimation minimizes beamforming output signal power, then noise components are removed, but log likelihood of probability distribution is not maximized
Solution Approach 1:
The patent fundamentally changes the estimation parameter from output power minimization to log-likelihood maximization under complex generalized Gaussian distribution. This parameter change ensures that the filter estimation accurately reflects the statistical properties of the noise signal, improving measurement precision
Solution Approach 2:
The patent employs an iterative algorithm where the filter estimation and variance estimation mutually refine each other. The algorithm uses the estimated filter to compute variance, then uses the variance to improve filter estimation, creating a self-improving system that maximizes log-likelihood
Data Source
AI summary
Provided is a method for beamforming by using maximum likelihood estimation in a speech recognition apparatus, including: (a) receiving an input signal (Xn,k) at a time frame n and a frequency k where noise is mixed: (b) determining a probability density function for a target signal (Yn,k) obtained by removing the noise from the input signal satisfies a complex generalized Guassian distribution or a complex gamma distribution where an average value is zero in a time-frequency domain; (c) estimating a variance (λn,k) of the target signal so as to maximize log likelihood for the probability density function; (d) estimating a filter (wk) maximizing a cost function so as to maximize the log likelihood for the probability density function; and (e) repeatedly performing the estimation of the steps (c) and (d) until the filter (wk) coverages, and finally acquiring a final filter (wk).

