Acoustic Model Training with Reduced Feature Space Variation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing acoustic model training methods do not adequately differentiate between frequently and infrequently occurring text elements, leading to reduced accuracy in speech recognition engines due to feature space variation.

Innovation Solution

A system and method for training an acoustic model using a combined phoneme set, dictionary, and transcription set that renames specific phonemes and text elements, focusing on the most frequently occurring elements to reduce feature space variation and improve recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a single acoustic model is trained to cover a wide range of applications using general training data, then the system versatility is improved, but the recognition accuracy for individual applications deteriorates

Engineering Contradiction:
Improveapplication coverageVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the training process into two distinct phases: a general training phase using diverse training data to establish broad acoustic patterns, and a fine-tuning phase using application-specific data to optimize for particular domains. This segmentation allows the model to first achieve versatility across applications, then specialize in individual applications without compromising the general capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by conducting general training before application-specific fine-tuning. The general training establishes a foundational acoustic model that can handle various applications, and only after this preliminary step is completed does the system proceed to application-specific optimization. This ensures that versatility is built first, then precision is enhanced for specific domains.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If feature space is reduced to improve recognition accuracy, then measurement precision is improved, but the ability to handle diverse phonetic variations deteriorates

Engineering Contradiction:
Improverecognition accuracyVSAvoidphonetic variation handling
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies local quality by differentiating between general phonetic features that should be preserved across all applications and application-specific phonetic variations that should be optimized locally. The system maintains a core set of acoustic features for universal recognition while allowing application-specific fine-tuning of phonetic characteristics, enabling both accuracy and adaptability to diverse phonetic patterns.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements dynamics by making the acoustic model adaptable through sequential training. The model starts with a static, general acoustic representation and dynamically adjusts its feature space through application-specific fine-tuning. This dynamic adjustment allows the model to optimize feature extraction for specific applications while retaining the capability to handle phonetic variations through the general training foundation.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS8301446B2System and method for training an acoustic model with reduced feature space variation
Publication Date: 2012.10.30 ADACEL SYST
  • US8301446B2 patent drawing
  • US8301446B2 patent drawing
  • US8301446B2 patent drawing

AI summary

Feature space variation associated with specific text elements is reduced by training an acoustic model with a phoneme set, dictionary and transcription set configured to better distinguish the specific text elements and at least some specific phonemes associated therewith. The specific text elements can include the most frequently occurring text elements from a text data set, which can include text data beyond the transcriptions of a training data set. The specific text elements can be identified using a text element distribution table sorted by occurrence within the text data set. Specific phonemes can be limited to consonant phonemes to improve speed and accuracy.