Extended Feature Vectors for Wideband Acoustic Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face challenges in training effective wideband acoustic models using narrowband telephone data, as the latter lacks essential frequency components, leading to suboptimal performance due to the mixing and truncation of features in the cepstral domain.

Innovation Solution

A method is developed to generate extended feature vectors by estimating missing components in narrowband data, allowing for the training of wideband acoustic models in the cepstral domain using a combination of wideband and narrowband data through an iterative algorithm that converts model parameters between the spectral and cepstral domains, ensuring full rank covariance matrices and equal contribution from all dimensions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If narrowband telephone data is used for training, then data collection cost is reduced, but model performance deteriorates due to missing frequency components

Engineering Contradiction:
Improvedata collection costVSAvoidmodel performance
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent introduces an intermediary process of bandwidth extension that acts as a mediator between narrowband data and wideband model requirements. The system estimates missing high-frequency components by treating them as hidden variables and using the observed low-frequency components to infer their values, thereby bridging the gap between limited input data and comprehensive model training requirements

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameter representation by transforming the problem from directly using narrowband spectral data to estimating extended spectral parameters. By formulating missing frequency components as hidden variables and using probabilistic modeling, the system transforms insufficient observed parameters into comprehensive extended feature vectors that include both observed and estimated components

Inventive Principle:
Principle #35Parameter changes

2Reliability

If wideband speech data is used for training, then model performance improves due to more frequency components, but data collection cost increases

Engineering Contradiction:
Improvemodel performanceVSAvoiddata collection cost
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The patent creates a virtual copy of wideband training data by estimating missing high-frequency components from available narrowband data. Instead of requiring actual wideband recordings, the system generates synthetic wideband-like feature vectors through probabilistic estimation, effectively copying the beneficial properties of wideband data without the associated collection costs

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary bandwidth extension during the training phase by pre-estimating missing components and creating extended feature vectors before model training begins. This preliminary action prepares the data in advance, allowing the model to learn from enriched features without requiring actual wideband recordings during the training process

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If cepstral domain training is used, then speech recognition accuracy improves, but training from narrowband data becomes difficult due to feature mixing and truncation

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidtraining complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent inverts the conventional approach by not directly training in the cepstral domain from narrowband data, but instead first estimating missing spectral components in the spectral domain where frequency components are separable, then transforming the extended spectral features to the cepstral domain for training. This inversion avoids the mixing problem by working in the domain where components can be independently estimated

Inventive Principle:
Principle #13The other way round (Inversion)

4Ease of operation

If narrowband speech is sampled through telephone network, then data collection becomes easier, but information content is reduced due to low sampling rate

Engineering Contradiction:
Improvedata collection easeVSAvoidfrequency components
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent applies a counterweight approach by introducing estimated high-frequency components that compensate for the information lost during telephone network transmission. The estimation process generates virtual high-frequency content that balances and offsets the missing information, allowing the system to counteract the information loss inherent in narrowband sampling

Inventive Principle:
Principle #8Anti-weight (Counterweight)

Data Source

PatentUS7454338B2Training wideband acoustic models in the cepstral domain using mixed-bandwidth training data and extended vectors for speech recognition
Publication Date: 2008.11.18 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7454338B2 patent drawing
  • US7454338B2 patent drawing
  • US7454338B2 patent drawing

AI summary

A method and apparatus are provided that generate values for a first set of dimensions of a feature vector from a speech signal. The values of the first set of dimensions are used to estimate values for a second set of dimensions of the feature vector to form an extended feature vector. The extended feature vector is then used to train an acoustic model.