Extended Feature Vectors for Wideband Acoustic Model Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face challenges in training effective wideband acoustic models using narrowband telephone data, as the latter lacks essential frequency components, leading to suboptimal performance due to the mixing and truncation of features in the cepstral domain.
Innovation Solution
A method is developed to generate extended feature vectors by estimating missing components in narrowband data, allowing for the training of wideband acoustic models in the cepstral domain using a combination of wideband and narrowband data through an iterative algorithm that converts model parameters between the spectral and cepstral domains, ensuring full rank covariance matrices and equal contribution from all dimensions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If narrowband telephone data is used for training, then data collection cost is reduced, but model performance deteriorates due to missing frequency components
Solution Approach 1:
The patent introduces an intermediary process of bandwidth extension that acts as a mediator between narrowband data and wideband model requirements. The system estimates missing high-frequency components by treating them as hidden variables and using the observed low-frequency components to infer their values, thereby bridging the gap between limited input data and comprehensive model training requirements
Solution Approach 2:
The patent changes the parameter representation by transforming the problem from directly using narrowband spectral data to estimating extended spectral parameters. By formulating missing frequency components as hidden variables and using probabilistic modeling, the system transforms insufficient observed parameters into comprehensive extended feature vectors that include both observed and estimated components
2Reliability
If wideband speech data is used for training, then model performance improves due to more frequency components, but data collection cost increases
Solution Approach 1:
The patent creates a virtual copy of wideband training data by estimating missing high-frequency components from available narrowband data. Instead of requiring actual wideband recordings, the system generates synthetic wideband-like feature vectors through probabilistic estimation, effectively copying the beneficial properties of wideband data without the associated collection costs
Solution Approach 2:
The patent performs preliminary bandwidth extension during the training phase by pre-estimating missing components and creating extended feature vectors before model training begins. This preliminary action prepares the data in advance, allowing the model to learn from enriched features without requiring actual wideband recordings during the training process
3Measurement precision
If cepstral domain training is used, then speech recognition accuracy improves, but training from narrowband data becomes difficult due to feature mixing and truncation
Solution Approach 1:
The patent inverts the conventional approach by not directly training in the cepstral domain from narrowband data, but instead first estimating missing spectral components in the spectral domain where frequency components are separable, then transforming the extended spectral features to the cepstral domain for training. This inversion avoids the mixing problem by working in the domain where components can be independently estimated
4Ease of operation
If narrowband speech is sampled through telephone network, then data collection becomes easier, but information content is reduced due to low sampling rate
Solution Approach 1:
The patent applies a counterweight approach by introducing estimated high-frequency components that compensate for the information lost during telephone network transmission. The estimation process generates virtual high-frequency content that balances and offsets the missing information, allowing the system to counteract the information loss inherent in narrowband sampling
Data Source
AI summary
A method and apparatus are provided that generate values for a first set of dimensions of a feature vector from a speech signal. The values of the first set of dimensions are used to estimate values for a second set of dimensions of the feature vector to form an extended feature vector. The extended feature vector is then used to train an acoustic model.


