ASR Error Correction Function Mapping Parameters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional automatic speech recognition systems face challenges in reducing transcription errors when adapting general acoustic models to specific speakers, as existing adaptation techniques are complex and costly.
Innovation Solution
An enhanced automatic speech recognition system generates an error correction function that maps supervised parameters to unsupervised parameters, allowing for improved model or feature space adaptation, thereby reducing transcription errors by learning from correct and incorrect transcriptions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional model space adaptation or feature space adaptation techniques are used to tailor a general acoustic model to a specific speaker, then transcription accuracy is improved, but system complexity and implementation cost increase
Solution Approach 1:
The patent creates a simplified copy of the adaptation process by using pre-computed transformation matrices from unsupervised training that can be directly applied during runtime without complex optimization procedures. This copying approach maintains the essential adaptation function while eliminating computational complexity
Solution Approach 2:
The patent performs adaptation-related computations in advance during unsupervised training, pre-computing transformation matrices that capture speaker-specific characteristics. These pre-computed matrices are then directly applied during runtime without requiring complex real-time adaptation algorithms
2Measurement precision
If conventional adaptation techniques are implemented to reduce transcription errors, then recognition accuracy improves, but implementation cost increases
Solution Approach 1:
The patent uses a simplified copied version of the adaptation process that leverages pre-computed transformation matrices from unsupervised training, eliminating the need for expensive complex adaptation implementations while maintaining accuracy benefits
Solution Approach 2:
The patent employs lightweight transformation matrices that are computationally inexpensive to store and apply during runtime, replacing costly complex adaptation algorithms with simple, efficient matrix operations that have minimal computational overhead
3Measurement precision
If speaker adaptation is performed using training data from a particular speaker, then transcription errors for that speaker are reduced, but processing time and computational resources increase
Solution Approach 1:
The patent performs all speaker-specific adaptation computations during the unsupervised training phase, pre-computing transformation matrices that encode speaker characteristics. During runtime, only simple matrix application is required, dramatically reducing processing time while maintaining error reduction benefits
Solution Approach 2:
The patent creates a simplified runtime version of the adaptation process by copying and applying pre-computed transformation matrices, eliminating the need for time-consuming real-time adaptation computations while preserving the error reduction effect
Data Source
AI summary
Techniques for enhanced automatic speech recognition are described. An enhanced ASR system may be operative to generate an error correction function. The error correction function may represent a mapping between a supervised set of parameters and an unsupervised training set of parameters generated using a same set of acoustic training data, and apply the error correction function to an unsupervised testing set of parameters to form a corrected set of parameters used to perform speaker adaptation. Other embodiments are described and claimed.


