ASR Pronunciation Prediction via Co-emitted Text Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing Automatic Speech Recognition (ASR) systems face challenges in accurately predicting the pronunciation of text entities, especially for named entities, words in different languages, or unknown words, due to variability in human speech and limited pronunciation data.
Innovation Solution
The method involves generating an encoding of allowable pronunciations for a text sample, selecting predicted text samples corresponding to an audio sample, outputting the text sample, and updating the encoding based on the pronunciations of co-emitted text samples using processing circuitry.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If ASR systems use limited pronunciation data for text entities, then the system complexity is reduced, but the pronunciation prediction accuracy deteriorates
Solution Approach 1:
The patent combines pronunciation information from multiple co-emitted text samples to create a more comprehensive pronunciation model for the target text entity. By merging pronunciation data from several related text samples that were recognized together in the ASR process, the system accumulates sufficient pronunciation information without requiring external data sources, thereby improving prediction accuracy while maintaining system simplicity
Solution Approach 2:
The patent uses co-emitted text samples as intermediary elements to transfer pronunciation information to the target text entity. These co-emitted samples serve as mediators that bridge the gap between limited available data and the need for accurate pronunciation prediction, allowing the system to infer pronunciations through intermediate representations without directly requiring extensive pronunciation databases
2Adaptability or versatility
If ASR systems incorporate more pronunciation variations from co-emitted samples, then the adaptability to user-specific speech patterns improves, but the device complexity increases
Solution Approach 1:
The patent implements a feedback mechanism where pronunciation information from co-emitted text samples is continuously incorporated into the pronunciation model for the target text entity. The ASR system uses the recognition results from multiple text samples to refine and update pronunciation predictions, creating a self-improving loop that adapts to user-specific speech patterns through iterative feedback without requiring complex external adaptation modules
Solution Approach 2:
The system performs self-service by automatically extracting and utilizing pronunciation information from its own ASR recognition outputs. Rather than requiring separate adaptation systems or external data sources, the ASR system serves its own pronunciation needs by leveraging the co-emitted text samples it already generates during normal operation, thereby improving adaptability while avoiding additional system complexity
Data Source
AI summary
A method, device, and computer-readable storage medium for predicting pronunciation of a text sample, including generating an encoding of allowable pronunciations of the text sample, selecting predicted text samples corresponding to an audio sample, the predicted text samples including the text sample and one or more co-emitted text samples, outputting the text sample, and updating the encoding of allowable pronunciations of the text sample based on pronunciations of the one or more co-emitted text samples.


