Dynamic Hyperparameter Adjustment for Automatic Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automatic speech recognition systems face limitations due to the use of fixed hyper parameters, which do not account for variations in acoustic conditions and application domains, leading to suboptimal recognition performance, especially in outlier conditions, speakers, and channels.

Innovation Solution

A dynamic and adaptive approach is introduced, where a machine learning model trained on audio data and metadata estimates optimal hyper parameters for automatic speech recognition, allowing for real-time or batch mode adjustments of parameters such as word insertion penalty, language model scale, and beam pruning width, to improve recognition accuracy across different environments and speakers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If fixed hyper parameters are used in automatic speech recognition systems, then system complexity is reduced and ease of operation is improved, but recognition accuracy deteriorates in outlier conditions and varied acoustic environments

Engineering Contradiction:
Improveease of operationVSAvoidrecognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent implements dynamic hyperparameter adjustment by training a machine learning model to predict optimal hyperparameters based on acoustic environment characteristics. The system transitions from static fixed parameters to dynamic parameters that adapt to different acoustic conditions, speakers, and channels, resolving the contradiction between ease of operation and recognition accuracy.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the state of hyperparameters from fixed to variable by introducing a machine learning model that outputs different hyperparameter values based on input acoustic characteristics. This allows the system to maintain ease of operation while improving recognition accuracy through parameter adaptation to different conditions.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If dynamic hyper parameter adjustment via machine learning is implemented, then recognition accuracy is improved in varied conditions, but device complexity increases

Engineering Contradiction:
Improverecognition accuracyVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a machine learning model as an intermediary component that bridges the gap between fixed hyperparameters and dynamic adaptation needs. This intermediary model processes acoustic characteristics and generates appropriate hyperparameter adjustments, enabling accuracy improvement while containing complexity through a dedicated specialized component.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies preliminary action by pre-training the machine learning model on acoustic characteristics and hyperparameter combinations before deployment. This allows the system to have dynamic adaptation capability ready in advance, reducing the complexity of real-time decision-making and enabling accurate hyperparameter selection without excessive computational burden during operation.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If hyper parameters are tuned on sample audio data with fixed settings, then development time is reduced and productivity is improved, but adaptability to different acoustic conditions and speakers deteriorates

Engineering Contradiction:
ImproveproductivityVSAvoidadaptability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent makes the hyperparameter system dynamic by training a machine learning model that adapts to different acoustic conditions, speakers, and channels. This dynamic approach maintains productivity through automated model-based selection while significantly improving adaptability across varied environments, resolving the contradiction between development efficiency and system versatility.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent creates a universal hyperparameter selection system through the machine learning model that can handle multiple acoustic conditions, speaker types, and channel characteristics with a single trained model. This multi-functional approach maintains productivity while enabling the system to adapt to diverse conditions that would otherwise require separate tuning processes.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11972753B2System and method for performing automatic speech recognition system parameter adjustment via machine learning
Publication Date: 2024.04.30 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11972753B2 patent drawing
  • US11972753B2 patent drawing
  • US11972753B2 patent drawing

AI summary

A system, method and computer-readable storage device provides an improved speech processing approach in which hyper parameters used for speech recognition are modified dynamically or in batch mode rather than fixed statically. The method includes estimating, via a model trained on audio data and/or metadata, a set of parameters useful for performing automatic speech recognition, receiving speech at an automatic speech recognition system, applying, by the automatic speech recognition system, the set of parameters to processing the speech to yield text and outputting the text from the automatic speech recognition system.