Acoustic Model Training from User-Corrected Speech Terms
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems often misinterpret user utterances, leading to incorrect transcriptions that require manual correction, which is inefficient and can lead to inaccurate training of acoustic models.
Innovation Solution
A method where user corrections to incorrect transcriptions are used to train an acoustic model, isolating audio data corresponding to the corrected terms to improve recognition of specific terms with various pronunciations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If the speech recognition system uses automated speech recognition to transcribe user utterances, then the system can automatically process voice inputs, but the transcription accuracy deteriorates leading to incorrect transcriptions that require manual correction
Solution Approach 1:
The system implements feedback by collecting user corrections to transcriptions and using them to retrain the acoustic model. When users correct incorrect transcriptions, these corrections are fed back into the training process, allowing the system to learn from its errors and improve future transcription accuracy automatically
Solution Approach 2:
The system performs self-improvement by automatically training its acoustic model using user correction data. The speech recognition system serves itself by utilizing user feedback to retrain and enhance its own transcription capabilities without requiring external intervention beyond the initial user corrections
2Measurement precision
If the system collects user corrections to improve transcription accuracy, then the acoustic model can learn from real user data, but the system complexity increases due to additional data processing and model training requirements
Solution Approach 1:
The system extracts only the necessary components from user interactions - specifically isolating the correction data and corresponding audio segments - to use for model training. This extraction approach allows the system to leverage user feedback without being overwhelmed by the complexity of processing entire conversation contexts
Solution Approach 2:
The system performs preliminary actions by collecting and organizing user correction data during normal operation, preparing training datasets in advance. This allows the acoustic model to be trained on real user correction data without requiring complex real-time processing during the actual transcription task
Data Source
AI summary
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for speech recognition. One of the methods includes receiving first audio data corresponding to an utterance; obtaining a first transcription of the first audio data; receiving data indicating (i) a selection of one or more terms of the first transcription and (ii) one or more of replacement terms; determining that one or more of the replacement terms are classified as a correction of one or more of the selected terms; in response to determining that the one or more of the replacement terms are classified as a correction of the one or more of the selected terms, obtaining a first portion of the first audio data that corresponds to one or more terms of the first transcription; and using the first portion of the first audio data that is associated with the one or more terms of the first transcription to train an acoustic model for recognizing the one or more of the replacement terms.


