Reduced Complexity Acoustic Model for Speech Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automatic speech recognition systems face poor performance and misrecognitions due to their rich structure models with many Gaussian distributions, which lead to minimal feature changes and poor adaptation to new speakers, especially when speakers are outliers.

Innovation Solution

The approach involves using a reduced complexity model with fewer distributions for adaptation, gradually increasing model size through recursive training to move features closer to the overall centroid, resulting in higher recognition accuracy and more homogeneous models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a rich structure model with many Gaussian distributions is used for speech recognition, then the model can represent speech sounds accurately, but the adaptation to new speakers becomes ineffective and feature changes become minimal

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidadaptation to new speakers
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the adaptation process into two distinct phases: (1) a global adaptation phase using a reduced complexity model with fewer Gaussian distributions to capture overall speaker characteristics, and (2) a local adaptation phase using the full complexity model to capture speaker-specific details. This segmentation allows the system to first establish a baseline adaptation that moves features toward the overall centroid, then refine with speaker-specific adjustments, thereby resolving the contradiction between model richness and adaptability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent dynamically adjusts the complexity of the adaptation model based on the adaptation stage. During initial adaptation, a reduced complexity model with fewer distributions is used to enable larger feature changes and better capture of outlier speakers. As adaptation progresses, the full complexity model is employed for finer adjustments. This dynamic approach allows the system to optimize between adaptation effectiveness and recognition accuracy at different stages.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If local mixture components are used for adaptation to match nearest distributions, then speaker-specific adaptation is achieved, but features move away from the overall centroid leading to poor performance

Engineering Contradiction:
Improvespeaker-specific adaptationVSAvoidoverall speech recognition performance
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent applies preliminary global adaptation using a reduced complexity model before performing local speaker-specific adaptation. This preliminary action ensures that features are first adjusted toward the overall centroid and establish a solid baseline representation. Only after this preliminary global adaptation is the local mixture component adaptation applied, ensuring that speaker-specific adjustments are built upon a stable foundation rather than potentially moving features away from the centroid.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If a reduced complexity model with fewer distributions is used for adaptation, then feature changes increase and adaptation to new speakers improves, but model representation accuracy may decrease

Engineering Contradiction:
Improveadaptation to new speakersVSAvoidspeech sound representation accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the model into two versions: a reduced complexity model with fewer Gaussian distributions optimized for adaptation, and a full complexity model with many distributions optimized for accurate speech sound representation. The reduced model is used exclusively for the adaptation phase to maximize feature changes and speaker adaptability, while the full model is used for the recognition phase to ensure accurate representation of speech sounds. This segmentation allows each model to be optimized for its specific purpose without compromise.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8600749B2System and method for training adaptation-specific acoustic models for automatic speech recognition
Publication Date: 2013.12.03 INTERACTIONS LLC (US)
  • US8600749B2 patent drawing
  • US8600749B2 patent drawing
  • US8600749B2 patent drawing

AI summary

Disclosed herein are systems, methods, and computer-readable storage media for training adaptation-specific acoustic models. A system practicing the method receives speech and generates a full size model and a reduced size model, the reduced size model starting with a single distribution for each speech sound in the received speech. The system finds speech segment boundaries in the speech using the full size model and adapts features of the speech data using the reduced size model based on the speech segment boundaries and an overall centroid for each speech sound. The system then recognizes speech using the adapted features of the speech. The model can be a Hidden Markov Model (HMM). The reduced size model can also be of a reduced complexity, such as having fewer mixture components than a model of full complexity. Adapting features of speech can include moving the features closer to an overall feature distribution center.