Acoustic Pronunciation Learning via Conditional Probability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for generating pronunciation dictionaries are labor-intensive and costly, particularly for complex words like names and surnames, as they rely on letter-to-phone engines that struggle with varied pronunciations and require linguist verification, which is slow and expensive, and are not effective without acoustic samples.
Innovation Solution
A computerized method using conditional probability techniques to generate pronunciations by graphing initial pronunciations, determining the highest-scoring sets, substituting lowest-probability phones with unique substitutes, and weighting with linguistic probabilities, incorporating acoustic data and feedback to refine matches.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If letter-to-phone engines are used to generate pronunciations, then the process is automated, but the accuracy deteriorates for complex words like names and surnames
Solution Approach 1:
The system uses acoustic feedback from actual speech samples to correct and refine automatically generated pronunciations. The feedback mechanism compares letter-to-phone engine outputs with real acoustic data, allowing iterative improvement of pronunciation accuracy for complex words while maintaining automation.
Solution Approach 2:
The system changes the parameter approach by incorporating acoustic probability scores and linguistic context information alongside traditional letter-to-phone mapping. This multi-parameter approach enables the system to handle complex words like names and surnames more accurately by considering phonetic probability distributions rather than fixed mappings.
2Measurement precision
If linguists are employed to verify and adjust pronunciations, then the accuracy improves, but the cost and time increase
Solution Approach 1:
The system performs self-verification by automatically comparing generated pronunciations against acoustic samples and using probabilistic models to assess quality. The conditional probability framework enables the system to self-correct pronunciation errors without requiring linguist intervention, maintaining high accuracy while eliminating manual verification time.
Solution Approach 2:
The system replaces the mechanical process of manual linguist verification with an automated computational framework using conditional probability and acoustic modeling. This substitution eliminates human labor while maintaining or improving verification accuracy through algorithmic assessment of pronunciation quality against acoustic data.
3Reliability
If more acoustic samples are collected to improve pronunciation accuracy, then the reliability improves, but the complexity of the system increases
Solution Approach 1:
The system achieves universality by using a single conditional probability framework to handle multiple functions: generating pronunciations, evaluating accuracy, and determining reliability. This multi-functional approach allows the system to process diverse acoustic samples and word types through a unified model, improving reliability without proportionally increasing complexity.
Solution Approach 2:
The system manages complexity by dynamically adjusting probability thresholds and acoustic modeling parameters based on the availability and quality of acoustic samples. When more samples are available, the system automatically refines its probability distributions, improving reliability while keeping the underlying computational structure manageable through adaptive parameter tuning.
Data Source
AI summary
A computerized method is provided for generating pronunciations for words and storing the pronunciations in a pronunciation dictionary. The method includes graphing sets of initial pronunciations; thereafter in an ASR subsystem determining a highest-scoring set of initial pronunciations; generating sets of alternate pronunciations, wherein each set of alternate pronunciations includes the highest-scoring set of initial pronunciations with a lowest-probability phone of the highest-scoring initial pronunciation substituted with a unique-substitute phone; graphing the sets of alternate pronunciations; determining in the ASR subsystem a highest-scoring set of alternate pronunciations; and adding to a pronunciation dictionary the highest-scoring set of alternate pronunciations.


