Microphone Array Speech Recognition Dynamic Noise Masking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing microphone-array-based speech recognition systems face challenges in maintaining consistent noise cancellation across different environments, leading to varying speech recognition accuracy due to mismatched training and test conditions, and lack of consideration for noise interference information.
Innovation Solution
A microphone-array-based speech recognition system that incorporates a noise masking module, a confidence measure score computation module, and a threshold adjustment module to optimize noise cancellation by using a speech model and a filler model, adjusting noise masking thresholds to achieve maximum confidence measure scores, thereby improving speech recognition rates in noisy environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If noise cancellation is performed using fixed thresholds in traditional microphone-array-based speech recognition systems, then the system structure remains simple, but speech recognition accuracy varies across different environments due to mismatched training and test conditions
Solution Approach 1:
The patent implements dynamic threshold adjustment by introducing a confidence measure score computation module that evaluates speech model confidence and a threshold adjustment module that adapts noise masking thresholds in real-time based on environmental conditions. This transforms the static threshold system into a dynamic one that automatically adapts to different noisy environments, resolving the contradiction between maintaining simple system structure and achieving environmental adaptability.
Solution Approach 2:
The patent establishes a feedback loop where the confidence measure score computation module continuously evaluates the confidence of speech recognition results, and this feedback is used by the threshold adjustment module to dynamically adjust noise masking thresholds. This feedback mechanism enables the system to self-optimize for different environments, improving speech recognition accuracy without requiring manual reconfiguration.
2Measurement precision
If noise masking parameters are adjusted to improve speech recognition accuracy in noisy environments, then speech recognition accuracy improves, but the system complexity increases due to additional modules
Solution Approach 1:
The patent achieves multi-functionality by integrating the confidence measure score computation and threshold adjustment capabilities within the existing speech recognition framework. The confidence measure score computation module leverages existing speech model outputs, and the threshold adjustment module builds upon the existing noise masking module, allowing the system to perform both traditional speech recognition and adaptive noise cancellation using the same core components, thereby minimizing additional complexity.
Solution Approach 2:
The patent focuses on adjusting a single critical parameter (noise masking threshold) based on confidence measure scores rather than redesigning the entire noise cancellation system. By changing this key parameter dynamically, the system achieves improved speech recognition accuracy in noisy environments without introducing complex structural modifications, thus resolving the contradiction between improving accuracy and maintaining system simplicity.
3Ease of manufacture
If traditional speech models are used without considering noise interference information, then the training process remains simple, but speech recognition performance degrades in noisy environments
Solution Approach 1:
The patent introduces the confidence measure score as an intermediary that bridges the gap between traditional speech models and noisy environment performance. Rather than modifying the speech models themselves, the confidence measure score computation module evaluates model confidence and uses this information to adjust noise masking thresholds, allowing traditional speech models to maintain their training simplicity while achieving improved reliability in noisy environments through the intermediary confidence evaluation mechanism.
Data Source
AI summary
A microphone-array-based speech recognition system combines a noise cancelling technique for cancelling noise of input speech signals from an array of microphones, according to at least an inputted threshold. The system receives noise-cancelled speech signals outputted by a noise masking module through at least a speech model and at least a filler model, then computes a confidence measure score with the at least a speech model and the at least a filler model for each threshold and each noise-cancelled speech signal, and adjusts the threshold to continue the noise cancelling for achieving a maximum confidence measure score, thereby outputting a speech recognition result related to the maximum confidence measure score.


