Unsupervised Language Model Adaptation for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automated speech assessment systems face challenges in achieving high recognition accuracy due to the time-consuming and costly process of obtaining human transcriptions for large-scale language tests, making it impractical to use ordinary supervised training for language model adaptation.

Innovation Solution

The use of unsupervised and semi-supervised language model adaptation methods, which involve using untranscribed speech samples and Internet data to adapt language models, reducing the need for human transcription and improving recognition accuracy through techniques like active learning and directed manual transcription.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If ordinary supervised training is used for language model adaptation, then recognition accuracy can be improved, but the process becomes time-consuming and costly due to the need for human transcriptions

Engineering Contradiction:
Improverecognition accuracyVSAvoidtime for obtaining human transcriptions
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system uses automatically generated transcriptions from an ASR system to adapt the language model, rather than relying on human transcriptions. The ASR system serves itself by providing its own output data for model adaptation, eliminating the need for external human transcription services and significantly reducing time and cost while maintaining improved recognition accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces an intermediary ASR system that generates transcriptions which are then used for language model adaptation. This intermediary layer allows the system to bridge the gap between raw speech data and adapted language models without requiring direct human transcription involvement, thus resolving the contradiction between accuracy improvement and time/cost efficiency

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If human-generated transcript data is used for language model adaptation, then recognition accuracy improves, but the cost and time requirements increase significantly

Engineering Contradiction:
Improverecognition accuracyVSAvoidcost of obtaining transcriptions
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The system eliminates the need for expensive human transcription services by using the ASR system's own automatic transcriptions for language model adaptation. This self-service approach maintains the benefit of improved recognition accuracy while dramatically reducing the cost factor associated with human transcription labor

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent uses automatically generated transcription data that can be quickly produced and discarded or updated as needed, replacing the expensive and time-consuming human-generated transcripts. These automatic transcriptions serve as sufficient training data for language model adaptation without requiring the high-cost human transcription process

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Data Source

PatentUS9224383B2Unsupervised language model adaptation for automated speech scoring
Publication Date: 2015.12.29 EDUCATIONAL TESTING SERVICE
  • US9224383B2 patent drawing
  • US9224383B2 patent drawing
  • US9224383B2 patent drawing

AI summary

Systems and methods are provided for generating a transcript of a speech sample response to a test question. The speech sample response to the test question is provided to a language model, where the language model is configured to perform an automated speech recognition function. The language model is adapted to the test question to improve the automated speech recognition function by providing to the language model automated speech recognition data related to the test question, Internet data related to the test question, or human-generated transcript data related to the test question. The transcript of the speech sample is generated using the adapted language model.