Speech Presence Probability Calculation for Dual-Microphone Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech presence probability calculation methods for dual-microphone speech enhancement systems are computationally intensive and sensitive to parameter fluctuations, failing to ensure that the speech presence probability approaches zero in inactive segments, leading to inaccuracies in noise estimation and speech enhancement quality.
Innovation Solution
A method involving the calculation of two metric parameters, signal-to-noise ratio (SNR) and signal power level difference (PLD), followed by normalization and non-linear transformation, is used to determine the speech presence probability, employing a formula that fits the product term and power term of these parameters to normalize the fitting coefficient, thereby reducing computational complexity and improving robustness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If traditional VAD algorithms with binary decision methods are used, then the system is simple to implement, but misjudgment occurs frequently affecting noise estimation accuracy
Solution Approach 1:
The patent transforms the binary decision parameter into a continuous speech presence probability parameter. Instead of using fixed thresholds for binary decisions, the system calculates a probability value that can take any value between 0 and 1, allowing for more precise noise estimation while maintaining implementation simplicity through the use of standardized probability calculations.
2Measurement precision
If soft decision VAD technology is used to calculate speech presence probability, then noise estimation accuracy improves, but computational complexity increases significantly
Solution Approach 1:
The patent changes the calculation parameters to use only the signal-to-noise ratio (SNR) and signal power level difference (PLD), which are already computed in the speech enhancement system. This avoids the need for complex multi-parameter calculations while maintaining accurate speech presence probability estimation.
Solution Approach 2:
The patent extracts only the essential features (SNR and PLD) from the complex signal processing pipeline and uses these simplified parameters to calculate the speech presence probability. This extraction approach reduces computational complexity by eliminating unnecessary intermediate calculations while preserving the accuracy needed for effective noise estimation.
3Measurement precision
If existing speech presence probability methods are used, then some noise estimation can be performed, but the method is sensitive to parameter fluctuations reducing robustness
Solution Approach 1:
The patent changes the calculation to use a normalized probability formula that inherently compensates for parameter fluctuations. By using the standardized probability calculation based on SNR and PLD ratios, the system achieves greater robustness while maintaining noise estimation capability.
4Productivity
If traditional speech presence probability calculation is used, then processing can be performed, but the speech presence probability does not approach zero in inactive segments
Solution Approach 1:
The patent changes the output parameter to be a true probability value that approaches zero for inactive segments. The formula ensures that when SNR and PLD indicate clear inactive conditions, the calculated probability naturally approaches zero, providing accurate discrimination between active and inactive speech segments.
Data Source
AI summary
A method and apparatus for determining a speech presence probability and an electronic device are provided. According to present disclosure, a metric parameter of a signal to noise ratio of a signal of a first channel and a metric parameter of a signal power level difference between the first channel and the second channel are introduced in determining the speech presence probability, the normalization and non-linear transformation processing is performed on the above-mentioned metric parameters, and the speech presence probability is obtained by fitting the product term and a first power term of a power exponent of the above-mentioned parameters. Therefore, the calculation amount of calculating the speech presence probability is reduced, the calculation result has good robustness to parameter fluctuations, and the disclosure can be widely applied to various application scenarios of dual-microphone speech enhancement systems.


