Acoustic Pronunciation Learning via Conditional Probability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for generating pronunciation dictionaries are labor-intensive and costly, particularly for complex words like names and surnames, as they rely on letter-to-phone engines that struggle with varied pronunciations and require linguist verification, which is slow and expensive, and are not effective without acoustic samples.

Innovation Solution

A computerized method using conditional probability techniques to generate pronunciations by graphing initial pronunciations, determining the highest-scoring sets, substituting lowest-probability phones with unique substitutes, and weighting with linguistic probabilities, incorporating acoustic data and feedback to refine matches.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If letter-to-phone engines are used to generate pronunciations, then the process is automated, but the accuracy deteriorates for complex words like names and surnames

Engineering Contradiction:
Improveautomation of pronunciation generationVSAvoidpronunciation accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The system uses acoustic feedback from actual speech samples to correct and refine automatically generated pronunciations. The feedback mechanism compares letter-to-phone engine outputs with real acoustic data, allowing iterative improvement of pronunciation accuracy for complex words while maintaining automation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system changes the parameter approach by incorporating acoustic probability scores and linguistic context information alongside traditional letter-to-phone mapping. This multi-parameter approach enables the system to handle complex words like names and surnames more accurately by considering phonetic probability distributions rather than fixed mappings.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If linguists are employed to verify and adjust pronunciations, then the accuracy improves, but the cost and time increase

Engineering Contradiction:
Improvepronunciation accuracyVSAvoidtime for verification
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs self-verification by automatically comparing generated pronunciations against acoustic samples and using probabilistic models to assess quality. The conditional probability framework enables the system to self-correct pronunciation errors without requiring linguist intervention, maintaining high accuracy while eliminating manual verification time.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system replaces the mechanical process of manual linguist verification with an automated computational framework using conditional probability and acoustic modeling. This substitution eliminates human labor while maintaining or improving verification accuracy through algorithmic assessment of pronunciation quality against acoustic data.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If more acoustic samples are collected to improve pronunciation accuracy, then the reliability improves, but the complexity of the system increases

Engineering Contradiction:
Improvepronunciation reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system achieves universality by using a single conditional probability framework to handle multiple functions: generating pronunciations, evaluating accuracy, and determining reliability. This multi-functional approach allows the system to process diverse acoustic samples and word types through a unified model, improving reliability without proportionally increasing complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system manages complexity by dynamically adjusting probability thresholds and acoustic modeling parameters based on the availability and quality of acoustic samples. When more samples are available, the system automatically refines its probability distributions, improving reliability while keeping the underlying computational structure manageable through adaptive parameter tuning.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS7280963B1Method for learning linguistically valid word pronunciations from acoustic data
Publication Date: 2007.10.09 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7280963B1 patent drawing
  • US7280963B1 patent drawing
  • US7280963B1 patent drawing

AI summary

A computerized method is provided for generating pronunciations for words and storing the pronunciations in a pronunciation dictionary. The method includes graphing sets of initial pronunciations; thereafter in an ASR subsystem determining a highest-scoring set of initial pronunciations; generating sets of alternate pronunciations, wherein each set of alternate pronunciations includes the highest-scoring set of initial pronunciations with a lowest-probability phone of the highest-scoring initial pronunciation substituted with a unique-substitute phone; graphing the sets of alternate pronunciations; determining in the ASR subsystem a highest-scoring set of alternate pronunciations; and adding to a pronunciation dictionary the highest-scoring set of alternate pronunciations.