Mixed Normalization Factor for Acoustic Model Robustness
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional acoustic models for speech recognition are not sufficiently robust due to limited diversity of audio data, which affects their performance in recognizing voices with different cepstral means and variances, especially in multi-speaker conversations.
Innovation Solution
The method involves obtaining audio data from multiple sources, calculating normalization factors for each source, mixing these factors to create a mixed normalization factor, and using it to normalize the audio data, thereby enhancing the training data for acoustic models to improve robustness and adaptability in recognizing diverse voices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If audio data from multiple sources is normalized using separate normalization factors, then each source maintains its original characteristics, but the acoustic model lacks robustness for multi-speaker conversations
Solution Approach 1:
The patent combines normalization factors from multiple audio sources into a single mixed normalization factor. This is achieved by calculating normalization factors for each source separately, then mixing them through weighted averaging or other combination techniques to create a unified normalization factor that represents the diversity of all sources while enabling the acoustic model to generalize across different speakers and conditions.
Solution Approach 2:
The patent changes the normalization parameter from source-specific to mixed/combined. By transforming the normalization approach from individual source normalization to mixed normalization, the system adapts the acoustic model to handle variations in cepstral means and variances across different speakers, thereby improving robustness for multi-speaker conversations.
2Reliability
If more diverse audio data is used for training, then robustness improves, but computational resource requirements increase
Solution Approach 1:
The patent performs preliminary normalization of audio data from multiple sources using a mixed normalization factor before training the acoustic model. By pre-processing the data to account for source diversity through mixed normalization, the system reduces the computational burden during model training and inference, as the model doesn't need to learn normalization patterns from scratch during training.
3Ease of manufacture
If conventional normalization is used for each audio source, then processing is simpler, but recognition accuracy for diverse voices deteriorates
Solution Approach 1:
The patent creates a universal mixed normalization factor that serves multiple audio sources simultaneously. This single normalization factor is derived from and applicable to multiple sources with different characteristics, making the normalization process multi-functional. The mixed normalization factor universally handles diverse voice characteristics while maintaining a unified processing approach.
Data Source
AI summary
A method, computer system, and a computer program product for audio data augmentation are provided. Sets of audio data from different sources may be obtained. A respective normalization factor for at least two sources of the different sources may be calculated. The normalization factors from the at least two sources may be mixed to determine a mixed normalization factor. A first set of the sets may be normalized by using the mixed normalization factor and to obtain training data for training an acoustic model.


