Reduced Complexity Acoustic Model for Speech Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automatic speech recognition systems face poor performance and misrecognitions due to their rich structure models with many Gaussian distributions, which lead to minimal feature changes and poor adaptation to new speakers, especially when speakers are outliers.
Innovation Solution
The approach involves using a reduced complexity model with fewer distributions for adaptation, gradually increasing model size through recursive training to move features closer to the overall centroid, resulting in higher recognition accuracy and more homogeneous models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a rich structure model with many Gaussian distributions is used for speech recognition, then the model can represent speech sounds accurately, but the adaptation to new speakers becomes ineffective and feature changes become minimal
Solution Approach 1:
The patent segments the adaptation process into two distinct phases: (1) a global adaptation phase using a reduced complexity model with fewer Gaussian distributions to capture overall speaker characteristics, and (2) a local adaptation phase using the full complexity model to capture speaker-specific details. This segmentation allows the system to first establish a baseline adaptation that moves features toward the overall centroid, then refine with speaker-specific adjustments, thereby resolving the contradiction between model richness and adaptability.
Solution Approach 2:
The patent dynamically adjusts the complexity of the adaptation model based on the adaptation stage. During initial adaptation, a reduced complexity model with fewer distributions is used to enable larger feature changes and better capture of outlier speakers. As adaptation progresses, the full complexity model is employed for finer adjustments. This dynamic approach allows the system to optimize between adaptation effectiveness and recognition accuracy at different stages.
2Adaptability or versatility
If local mixture components are used for adaptation to match nearest distributions, then speaker-specific adaptation is achieved, but features move away from the overall centroid leading to poor performance
Solution Approach 1:
The patent applies preliminary global adaptation using a reduced complexity model before performing local speaker-specific adaptation. This preliminary action ensures that features are first adjusted toward the overall centroid and establish a solid baseline representation. Only after this preliminary global adaptation is the local mixture component adaptation applied, ensuring that speaker-specific adjustments are built upon a stable foundation rather than potentially moving features away from the centroid.
3Adaptability or versatility
If a reduced complexity model with fewer distributions is used for adaptation, then feature changes increase and adaptation to new speakers improves, but model representation accuracy may decrease
Solution Approach 1:
The patent segments the model into two versions: a reduced complexity model with fewer Gaussian distributions optimized for adaptation, and a full complexity model with many distributions optimized for accurate speech sound representation. The reduced model is used exclusively for the adaptation phase to maximize feature changes and speaker adaptability, while the full model is used for the recognition phase to ensure accurate representation of speech sounds. This segmentation allows each model to be optimized for its specific purpose without compromise.
Data Source
AI summary
Disclosed herein are systems, methods, and computer-readable storage media for training adaptation-specific acoustic models. A system practicing the method receives speech and generates a full size model and a reduced size model, the reduced size model starting with a single distribution for each speech sound in the received speech. The system finds speech segment boundaries in the speech using the full size model and adapts features of the speech data using the reduced size model based on the speech segment boundaries and an overall centroid for each speech sound. The system then recognizes speech using the adapted features of the speech. The model can be a Hidden Markov Model (HMM). The reduced size model can also be of a reduced complexity, such as having fewer mixture components than a model of full complexity. Adapting features of speech can include moving the features closer to an overall feature distribution center.


