Pronunciation Model Dialect Separation via Phoneme Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current pronunciation modeling techniques struggle to effectively separate dialectal variations from other acoustic differences, leading to confusion and inefficiency, as they often prioritize non-dialectal variations and require expensive and time-consuming manual processes.

Innovation Solution

A system that generates a pronunciation model by identifying a generic model of speech, labeling interchangeable phonemic alternatives as the same phoneme, and substituting these alternatives to create dialect-dependent dictionaries, allowing for improved dialectal modeling and recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional pronunciation modeling techniques are used, then the system can handle general speech recognition, but it cannot effectively separate dialectal variations from other acoustic differences

Engineering Contradiction:
Improvedialectal variation separation accuracyVSAvoidmodeling complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the acoustic space by separating dialectal variations from non-dialectal variations through clustering. It divides the phonetic data into dialect-specific clusters while maintaining a unified phoneme representation, allowing the system to handle dialectal differences without increasing overall modeling complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary layer between the acoustic signal and the phoneme representation. This intermediary consists of dialect-specific acoustic models that map acoustic features to phonemes while accounting for dialectal variations, enabling precise separation of dialectal from non-dialectal characteristics.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If manual pronunciation dictionary creation is used, then dialectal variations can be captured, but the process is expensive and time-consuming

Engineering Contradiction:
Improvedialectal pronunciation accuracyVSAvoidmodel creation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements self-service through automatic clustering algorithms that autonomously identify dialectal variations and create pronunciation models without human intervention. The system automatically analyzes acoustic data, clusters speakers by dialect, and generates pronunciation dictionaries, eliminating the need for manual linguist involvement.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the approach from manual parameter specification to automatic parameter extraction. Instead of requiring linguists to manually define phonetic rules and dictionaries, the system automatically extracts dialectal parameters through acoustic clustering and statistical analysis, significantly reducing creation time and cost.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If alternative phoneme symbols are used for dialectal variations, then different dialects can be represented, but confusion and disparity arise when modeling various speech dialects

Engineering Contradiction:
Improvedialect representation capabilityVSAvoidpronunciation modeling consistency
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent creates a universal phoneme representation that serves multiple dialects simultaneously. Each phoneme has a core representation that is common across all dialects, while dialect-specific acoustic models provide the necessary variations. This universal approach allows the system to represent multiple dialects without creating separate phoneme symbols for each, eliminating confusion and maintaining consistency.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8862470B2System and method for pronunciation modeling
Publication Date: 2014.10.14 INTERACTIONS LLC (US)
  • US8862470B2 patent drawing
  • US8862470B2 patent drawing
  • US8862470B2 patent drawing

AI summary

Systems, computer-implemented methods, and tangible computer-readable media for generating a pronunciation model. The method includes identifying a generic model of speech composed of phonemes, identifying a family of interchangeable phonemic alternatives for a phoneme in the generic model of speech, labeling the family of interchangeable phonemic alternatives as referring to the same phoneme, and generating a pronunciation model which substitutes each family for each respective phoneme. In one aspect, the generic model of speech is a vocal tract length normalized acoustic model. Interchangeable phonemic alternatives can represent a same phoneme for different dialectal classes. An interchangeable phonemic alternative can include a string of phonemes.