Sparse MAP Acoustic Model Adaptation for Storage Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face challenges in creating user-specific acoustic models due to large storage requirements and inefficiencies in adapting acoustic models with limited training data, leading to reduced accuracy and increased storage needs.
Innovation Solution
The implementation of a Maximum A Posteriori (MAP) adaptation process with sparseness constraints to generate acoustic parameter adaptation data, which identifies changes in a small fraction of acoustic parameters from a baseline model, resulting in user-specific models that require significantly less storage space while maintaining high accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If conventional linear regression methods are used to adapt baseline acoustic models to specific speakers, then the adaptation can be performed with small amounts of training data, but the storage space requirements remain huge because all acoustic parameters must be stored
Solution Approach 1:
The patent extracts only the necessary acoustic parameters from the baseline model for adaptation. Instead of storing or processing all acoustic parameters, the system identifies and extracts only those parameters that need to be adapted to specific speakers, significantly reducing storage requirements while maintaining adaptation capability.
Solution Approach 2:
The patent applies local quality by treating different acoustic parameters differently based on their adaptability. Some parameters are kept fixed from the baseline model, while only specific parameters are adapted to individual speakers. This selective adaptation approach reduces storage needs by focusing only on the parameters that require customization.
2Measurement precision
If distinct acoustic models are created for each user to improve recognition accuracy, then user-specific accuracy increases, but storage requirements increase tremendously
Solution Approach 1:
The patent creates a universal baseline acoustic model that serves all users, combined with user-specific adaptations stored in a compact form. This multi-functional approach allows the system to provide user-specific recognition accuracy while maintaining a shared baseline model that reduces overall storage requirements compared to storing complete separate models for each user.
Solution Approach 2:
The patent implements a nested structure where user-specific adapted parameters are nested within the context of the baseline acoustic model. The adapted parameters for each user are stored as modifications to the baseline model rather than as complete separate models, creating a compact nested representation that reduces storage requirements.
3Measurement precision
If all acoustic parameters are adapted for each user to maximize accuracy, then user-specific performance improves, but the complexity of model management and storage increases
Solution Approach 1:
The patent changes the parameter representation by storing only the adapted parameters for each user rather than complete models. This parameter change approach simplifies model management by reducing the number of parameters that need to be stored and managed for each user, while still achieving user-specific recognition accuracy through the adapted parameters.
Data Source
AI summary
Techniques disclosed herein include using a Maximum A Posteriori (MAP) adaptation process that imposes sparseness constraints to generate acoustic parameter adaptation data for specific users based on a relatively small set of training data. The resulting acoustic parameter adaptation data identifies changes for a relatively small fraction of acoustic parameters from a baseline acoustic speech model instead of changes to all acoustic parameters. This results in user-specific acoustic parameter adaptation data that is several orders of magnitude smaller than storage amounts otherwise required for a complete acoustic model. This provides customized acoustic speech models that increase recognition accuracy at a fraction of expected data storage requirements.


