Speech Enhancement Source Identification Using Confidence Levels
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for speech signal enhancement in multisource environments, such as those used in Automatic Speech Recognition (ASR), are computationally demanding due to the iterative nature of Gaussian Mixture Model (GMM) fitting and often misidentify sources, especially in noisy conditions with multiple interfering signals.
Innovation Solution
The proposed solution employs a wrapped Gaussian mixture model and a Bayes decision process to identify and isolate sources, using confidence levels from wake-up-word detection to determine which sources participate in a dialogue phase, while excluding interfering sources and adapting to source movement, thereby improving source separation and reducing computational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If iterative Gaussian Mixture Model fitting is used for source classification, then source identification accuracy may be improved, but computational complexity increases significantly
Solution Approach 1:
The patent applies preliminary action by performing GMM fitting less frequently (e.g., once per dialogue phase or at selected frames) rather than in every frame. The system maintains classification models in advance and reuses them, avoiding repeated iterative fitting while still achieving accurate source identification when needed.
Solution Approach 2:
The patent uses partial action by selectively applying full GMM fitting only when necessary (e.g., when source conditions change or at dialogue phase boundaries), while using simpler classification methods for routine frames. This partial application of the complex algorithm reduces overall computational burden while maintaining accuracy where critical.
2Measurement precision
If GMM fitting restarts in every frame to track source movement, then source tracking accuracy is improved, but processing time increases
Solution Approach 1:
The system performs GMM fitting in advance at dialogue phase boundaries or when source conditions change, rather than restarting in every frame. The pre-computed models are then reused for tracking, maintaining accuracy while significantly reducing processing time during continuous operation.
Solution Approach 2:
The patent implements periodic action by updating GMM models at specific intervals (e.g., once per dialogue phase or at selected frames) rather than continuously in every frame. This periodic updating maintains source tracking accuracy while reducing the overall computational burden and processing time.
3Adaptability or versatility
If multiple sources are classified in noisy environments, then source separation capability is improved, but misidentification rate increases
Solution Approach 1:
The patent uses feedback by incorporating confidence scores from wake-up-word detection and ASR systems to validate and correct source classifications. When confidence levels indicate uncertainty or potential misidentification, the system can adjust classifications or request re-evaluation, improving reliability in noisy multisource environments.
Solution Approach 2:
The patent replaces purely algorithmic GMM-based classification with a hybrid approach that incorporates neural network-based wake-up-word detection and ASR confidence scoring. This substitution of mechanical/classical methods with intelligent/AI methods improves source identification reliability by leveraging pattern recognition and confidence-based decision making.
Data Source
AI summary
A method, computer program product, and computer system for receiving, by a computing device, a first signal emitted from one or more sources. A second signal may be received emitted from the one or more sources. A first confidence level that the wake-up-word is included in the first signal may be determined. A second confidence level that the wake-up-word is included in the second signal may be determined. It may be identified that the wake-up-word originated from a first source of the one or more sources based upon, at least in part, the first and second confidence levels. The first source may be enabled to participate in a dialog phase.


