Speech Noise Reduction Using Bark Domain Likelihood Ratio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech noise reduction technologies face challenges in accurately distinguishing between frames containing speech and noise, leading to low recognition rates due to reliance on priori signal-to-noise ratios alone.
Innovation Solution
A computer-implemented method that estimates a posteriori signal-to-noise ratio and priori signal-to-noise ratio, determines a speech/noise likelihood ratio in the Bark domain, and calculates a gain for converting noisy speech signals into pure speech signals, using a frequency domain transfer function.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a priori speech existence probability is estimated using only priori signal-to-noise ratios on all frequency points, then the method is simple to implement, but it cannot well distinguish frames containing both speech and noise from frames containing only noise
Solution Approach 1:
The patent segments the frequency domain into multiple frequency points and introduces frame-level speech existence probability estimation. Instead of using a single global priori probability, the system estimates speech existence probability separately for each frequency point and then combines them at the frame level. This segmentation allows the system to distinguish between frames with speech and frames with noise more accurately by analyzing the distribution of speech existence probabilities across different frequency points.
Solution Approach 2:
The patent adds a temporal dimension to the frequency-domain analysis by introducing frame-level speech existence probability. The system transitions from estimating speech existence probability solely in the frequency domain to incorporating temporal information through frame-level analysis. This dimensional change enables the system to capture the temporal characteristics of speech signals, improving the ability to distinguish speech frames from noise frames.
2Reliability
If Wiener gain fluctuation is kept small in time and frequency, then speech recognition rate improves, but noise suppression effectiveness decreases
Solution Approach 1:
The patent introduces dynamic frame-level speech existence probability estimation that adapts to the specific characteristics of each frame. Instead of using a static, uniform Wiener gain, the system dynamically adjusts the gain for each frequency point based on the estimated speech existence probability. This dynamic approach allows the system to maintain small Wiener gain fluctuations for stable speech recognition while introducing appropriate gain variations to suppress noise effectively.
Solution Approach 2:
The patent applies local quality by estimating speech existence probability separately for each frequency point and using this information to adjust the Wiener gain locally. The system calculates a local Wiener gain for each frequency point based on the estimated speech existence probability at that frequency, rather than applying a uniform gain across all frequency points. This local adjustment enables the system to suppress noise at frequency points where speech is unlikely while preserving speech content where speech probability is high.
Data Source
AI summary
This application discloses a speech noise reduction method performed by a computing device. The method includes: obtaining a noisy speech signal, the noisy speech signal including a pure speech signal and a noise signal; estimating a posteriori signal-to-noise ratio and a priori signal-to-noise ratio of the noisy speech signal; determining a speech/noise likelihood ratio in a Bark domain based on the estimated posteriori signal-to-noise ratio and the estimated priori signal-to-noise ratio; estimating a priori speech existence probability based on the determined speech/noise likelihood ratio; determining a gain based on the estimated posteriori signal-to-noise ratio, the estimated priori signal-to-noise ratio, and the estimated priori speech existence probability, the gain being a frequency domain transfer function used for converting the noisy speech signal into an estimation of the pure speech signal; and exporting the estimation of the pure speech signal from the noisy speech signal based on the gain.


