Vocal Tract Length Normalization Codebooks for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems are hindered by sensitivity to variable background environments, accents, dialects, speaker characteristics, and recording conditions, requiring a minimum of 10 to 20 seconds to normalize speech, making them impractical for applications with limited speech samples.

Innovation Solution

The method involves generating codebooks with precomputed speaker normalization factors by computing vocal tract lengths and clustering speech vectors, allowing for rapid selection and application of the appropriate normalization factor to normalize speech samples for improved recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional speech normalization methods are used to estimate vocal tract length, then speech recognition accuracy is improved, but computational time increases to 10-20 seconds

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidnormalization computational time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent pre-computes and stores vocal tract length normalization factors in codebooks during an offline training phase. During runtime, the system only needs to perform fast codebook lookup and selection based on acoustic distance, avoiding the computationally intensive 10-20 second normalization process while maintaining recognition accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates simplified representations (codebooks) that copy the essential normalization characteristics of speakers without storing complete speech samples. Each codebook contains clustered speech vectors and associated normalization factors, enabling fast approximation of vocal tract length estimation through acoustic distance calculation rather than full speech analysis.

Inventive Principle:
Principle #26Copying

2Productivity

If codebook-based normalization is used instead of reference acoustic models, then computational time is reduced, but adaptability to different speakers may be affected

Engineering Contradiction:
Improvenormalization speedVSAvoidspeaker adaptation capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent transforms the continuous vocal tract length parameter into discrete normalization factors stored in codebooks. By clustering speech vectors and computing representative normalization factors for each cluster, the system maintains adaptability to different speakers while enabling fast codebook lookup instead of continuous computation.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments the speech space into clustered regions, with each cluster represented by a codebook entry containing normalized speech vectors and associated normalization factors. This segmentation allows the system to adapt to different speakers by selecting the appropriate codebook segment based on acoustic distance, maintaining versatility while improving speed.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS7797158B2System and method for improving robustness of speech recognition using vocal tract length normalization codebooks
Publication Date: 2010.09.14 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7797158B2 patent drawing
  • US7797158B2 patent drawing
  • US7797158B2 patent drawing

AI summary

Disclosed are systems, methods, and computer readable media for performing speech recognition. The method embodiment comprises selecting a codebook from a plurality of codebooks with a minimal acoustic distance to a received speech sample, the plurality of codebooks generated by a process of (a) computing a vocal tract length for a each of a plurality of speakers, (b) for each of the plurality of speakers, clustering speech vectors, and (c) creating a codebook for each speaker, the codebook containing entries for the respective speaker's vocal tract length, speech vectors, and an optional vector weight for each speech vector, (2) applying the respective vocal tract length associated with the selected codebook to normalize the received speech sample for use in speech recognition, and (3) recognizing the received speech sample based on the respective vocal tract length associated with the selected codebook.