Microphone Array Speech Recognition Dynamic Noise Masking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing microphone-array-based speech recognition systems face challenges in maintaining consistent noise cancellation across different environments, leading to varying speech recognition accuracy due to mismatched training and test conditions, and lack of consideration for noise interference information.

Innovation Solution

A microphone-array-based speech recognition system that incorporates a noise masking module, a confidence measure score computation module, and a threshold adjustment module to optimize noise cancellation by using a speech model and a filler model, adjusting noise masking thresholds to achieve maximum confidence measure scores, thereby improving speech recognition rates in noisy environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If noise cancellation is performed using fixed thresholds in traditional microphone-array-based speech recognition systems, then the system structure remains simple, but speech recognition accuracy varies across different environments due to mismatched training and test conditions

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidenvironmental adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic threshold adjustment by introducing a confidence measure score computation module that evaluates speech model confidence and a threshold adjustment module that adapts noise masking thresholds in real-time based on environmental conditions. This transforms the static threshold system into a dynamic one that automatically adapts to different noisy environments, resolving the contradiction between maintaining simple system structure and achieving environmental adaptability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent establishes a feedback loop where the confidence measure score computation module continuously evaluates the confidence of speech recognition results, and this feedback is used by the threshold adjustment module to dynamically adjust noise masking thresholds. This feedback mechanism enables the system to self-optimize for different environments, improving speech recognition accuracy without requiring manual reconfiguration.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If noise masking parameters are adjusted to improve speech recognition accuracy in noisy environments, then speech recognition accuracy improves, but the system complexity increases due to additional modules

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent achieves multi-functionality by integrating the confidence measure score computation and threshold adjustment capabilities within the existing speech recognition framework. The confidence measure score computation module leverages existing speech model outputs, and the threshold adjustment module builds upon the existing noise masking module, allowing the system to perform both traditional speech recognition and adaptive noise cancellation using the same core components, thereby minimizing additional complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent focuses on adjusting a single critical parameter (noise masking threshold) based on confidence measure scores rather than redesigning the entire noise cancellation system. By changing this key parameter dynamically, the system achieves improved speech recognition accuracy in noisy environments without introducing complex structural modifications, thus resolving the contradiction between improving accuracy and maintaining system simplicity.

Inventive Principle:
Principle #35Parameter changes

3Ease of manufacture

If traditional speech models are used without considering noise interference information, then the training process remains simple, but speech recognition performance degrades in noisy environments

Engineering Contradiction:
Improvetraining simplicityVSAvoidspeech recognition reliability
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent introduces the confidence measure score as an intermediary that bridges the gap between traditional speech models and noisy environment performance. Rather than modifying the speech models themselves, the confidence measure score computation module evaluates model confidence and uses this information to adjust noise masking thresholds, allowing traditional speech models to maintain their training simplicity while achieving improved reliability in noisy environments through the intermediary confidence evaluation mechanism.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS8744849B2Microphone-array-based speech recognition system and method
Publication Date: 2014.06.03 IND TECH RES INST
  • US8744849B2 patent drawing
  • US8744849B2 patent drawing
  • US8744849B2 patent drawing

AI summary

A microphone-array-based speech recognition system combines a noise cancelling technique for cancelling noise of input speech signals from an array of microphones, according to at least an inputted threshold. The system receives noise-cancelled speech signals outputted by a noise masking module through at least a speech model and at least a filler model, then computes a confidence measure score with the at least a speech model and the at least a filler model for each threshold and each noise-cancelled speech signal, and adjusts the threshold to continue the noise cancelling for achieving a maximum confidence measure score, thereby outputting a speech recognition result related to the maximum confidence measure score.