ASR Pronunciation Prediction via Co-emitted Text Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing Automatic Speech Recognition (ASR) systems face challenges in accurately predicting the pronunciation of text entities, especially for named entities, words in different languages, or unknown words, due to variability in human speech and limited pronunciation data.

Innovation Solution

The method involves generating an encoding of allowable pronunciations for a text sample, selecting predicted text samples corresponding to an audio sample, outputting the text sample, and updating the encoding based on the pronunciations of co-emitted text samples using processing circuitry.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If ASR systems use limited pronunciation data for text entities, then the system complexity is reduced, but the pronunciation prediction accuracy deteriorates

Engineering Contradiction:
Improvepronunciation prediction accuracyVSAvoidpronunciation data availability
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent combines pronunciation information from multiple co-emitted text samples to create a more comprehensive pronunciation model for the target text entity. By merging pronunciation data from several related text samples that were recognized together in the ASR process, the system accumulates sufficient pronunciation information without requiring external data sources, thereby improving prediction accuracy while maintaining system simplicity

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent uses co-emitted text samples as intermediary elements to transfer pronunciation information to the target text entity. These co-emitted samples serve as mediators that bridge the gap between limited available data and the need for accurate pronunciation prediction, allowing the system to infer pronunciations through intermediate representations without directly requiring extensive pronunciation databases

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If ASR systems incorporate more pronunciation variations from co-emitted samples, then the adaptability to user-specific speech patterns improves, but the device complexity increases

Engineering Contradiction:
Improveadaptation to user-specific speech patternsVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a feedback mechanism where pronunciation information from co-emitted text samples is continuously incorporated into the pronunciation model for the target text entity. The ASR system uses the recognition results from multiple text samples to refine and update pronunciation predictions, creating a self-improving loop that adapts to user-specific speech patterns through iterative feedback without requiring complex external adaptation modules

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs self-service by automatically extracting and utilizing pronunciation information from its own ASR recognition outputs. Rather than requiring separate adaptation systems or external data sources, the ASR system serves its own pronunciation needs by leveraging the co-emitted text samples it already generates during normal operation, thereby improving adaptability while avoiding additional system complexity

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250131910A1Automated prediction of pronunciation of text entities based on co-emitted speech recognition predictions
Publication Date: 2025.04.24 GOOGLE LLC
  • US20250131910A1 patent drawing
  • US20250131910A1 patent drawing
  • US20250131910A1 patent drawing

AI summary

A method, device, and computer-readable storage medium for predicting pronunciation of a text sample, including generating an encoding of allowable pronunciations of the text sample, selecting predicted text samples corresponding to an audio sample, the predicted text samples including the text sample and one or more co-emitted text samples, outputting the text sample, and updating the encoding of allowable pronunciations of the text sample based on pronunciations of the one or more co-emitted text samples.