ASR Error Correction Function Mapping Parameters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional automatic speech recognition systems face challenges in reducing transcription errors when adapting general acoustic models to specific speakers, as existing adaptation techniques are complex and costly.

Innovation Solution

An enhanced automatic speech recognition system generates an error correction function that maps supervised parameters to unsupervised parameters, allowing for improved model or feature space adaptation, thereby reducing transcription errors by learning from correct and incorrect transcriptions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional model space adaptation or feature space adaptation techniques are used to tailor a general acoustic model to a specific speaker, then transcription accuracy is improved, but system complexity and implementation cost increase

Engineering Contradiction:
Improvetranscription accuracyVSAvoidadaptation technique complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates a simplified copy of the adaptation process by using pre-computed transformation matrices from unsupervised training that can be directly applied during runtime without complex optimization procedures. This copying approach maintains the essential adaptation function while eliminating computational complexity

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs adaptation-related computations in advance during unsupervised training, pre-computing transformation matrices that capture speaker-specific characteristics. These pre-computed matrices are then directly applied during runtime without requiring complex real-time adaptation algorithms

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If conventional adaptation techniques are implemented to reduce transcription errors, then recognition accuracy improves, but implementation cost increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidimplementation cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent uses a simplified copied version of the adaptation process that leverages pre-computed transformation matrices from unsupervised training, eliminating the need for expensive complex adaptation implementations while maintaining accuracy benefits

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent employs lightweight transformation matrices that are computationally inexpensive to store and apply during runtime, replacing costly complex adaptation algorithms with simple, efficient matrix operations that have minimal computational overhead

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Measurement precision

If speaker adaptation is performed using training data from a particular speaker, then transcription errors for that speaker are reduced, but processing time and computational resources increase

Engineering Contradiction:
Improvetranscription error reductionVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs all speaker-specific adaptation computations during the unsupervised training phase, pre-computing transformation matrices that encode speaker characteristics. During runtime, only simple matrix application is required, dramatically reducing processing time while maintaining error reduction benefits

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a simplified runtime version of the adaptation process by copying and applying pre-computed transformation matrices, eliminating the need for time-consuming real-time adaptation computations while preserving the error reduction effect

Inventive Principle:
Principle #26Copying

Data Source

PatentUS8306819B2Enhanced automatic speech recognition using mapping between unsupervised and supervised speech model parameters trained on same acoustic training data
Publication Date: 2012.11.06 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8306819B2 patent drawing
  • US8306819B2 patent drawing
  • US8306819B2 patent drawing

AI summary

Techniques for enhanced automatic speech recognition are described. An enhanced ASR system may be operative to generate an error correction function. The error correction function may represent a mapping between a supervised set of parameters and an unsupervised training set of parameters generated using a same set of acoustic training data, and apply the error correction function to an unsupervised testing set of parameters to form a corrected set of parameters used to perform speaker adaptation. Other embodiments are described and claimed.