ASR Training Data Sampling via Benchmark Classification Distribution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automatic Speech Recognition (ASR) systems face performance variability due to differences in training corpus characteristics, leading to inaccurate transcriptions when used across different language domains, such as voice instant messaging versus news broadcasts.
Innovation Solution
A method involving benchmark text string classification to determine a classification distribution, which is used to sample a training corpus that matches the target domain's characteristics, ensuring the ASR system is trained with data that approximates the domain-specific topic distribution, thereby improving recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If a large corpus of training text strings is used to train the ASR system, then the system can learn language syntax, but the performance varies based on the characteristics of the corpus leading to inaccurate transcriptions in different language domains
Solution Approach 1:
The patent applies local quality by filtering and selecting training data based on domain-specific characteristics. Instead of treating all training data uniformly, the system identifies and emphasizes data from specific domains (e.g., news broadcasts vs. instant messaging) to match the target application's linguistic patterns, thereby improving transcription accuracy for that specific domain while maintaining overall corpus utilization.
Solution Approach 2:
The system changes the parameter of data selection criteria by using domain classifiers to evaluate and filter training text strings based on their domain characteristics. This parameter change enables the system to adjust the composition of the training corpus to match the target domain's linguistic properties, resolving the contradiction between corpus size and domain-specific accuracy.
2Adaptability or versatility
If the ASR system is trained on general language data, then it can handle diverse language syntax, but it performs inaccurately when applied to specific language domains such as voice instant messaging or news broadcasts
Solution Approach 1:
The patent segments the training data into different domains using classifiers that identify specific language characteristics. By dividing the general corpus into domain-specific subsets (e.g., formal speech, casual conversation, news), the system can then select and train on appropriate segments that match the target domain, maintaining both versatility and precision.
Solution Approach 2:
The system performs preliminary domain classification and filtering of training data before the actual training process. This preliminary action ensures that the training corpus is pre-adjusted to match the target domain's characteristics, allowing the ASR system to achieve high accuracy for specific domains while maintaining the ability to handle multiple language types.
Data Source
AI summary
A set of benchmark text strings may be classified to provide a set of benchmark classifications. The benchmark text strings in the set may correspond to a benchmark corpus of benchmark utterances in a particular language. A benchmark classification distribution of the set of benchmark classifications may be determined. A respective classification for each text string in a corpus of text strings may also be determined. Text strings from the corpus of text strings may be sampled to form a training corpus of training text strings such that the classifications of the training text strings have a training text string classification distribution that is based on the benchmark classification distribution. The training corpus of training text strings may be used to train an automatic speech recognition (ASR) system.


