Mixed Normalization Factor for Acoustic Model Robustness

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional acoustic models for speech recognition are not sufficiently robust due to limited diversity of audio data, which affects their performance in recognizing voices with different cepstral means and variances, especially in multi-speaker conversations.

Innovation Solution

The method involves obtaining audio data from multiple sources, calculating normalization factors for each source, mixing these factors to create a mixed normalization factor, and using it to normalize the audio data, thereby enhancing the training data for acoustic models to improve robustness and adaptability in recognizing diverse voices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If audio data from multiple sources is normalized using separate normalization factors, then each source maintains its original characteristics, but the acoustic model lacks robustness for multi-speaker conversations

Engineering Contradiction:
Improverobustness of acoustic modelVSAvoidadaptability to different cepstral means and variances
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent combines normalization factors from multiple audio sources into a single mixed normalization factor. This is achieved by calculating normalization factors for each source separately, then mixing them through weighted averaging or other combination techniques to create a unified normalization factor that represents the diversity of all sources while enabling the acoustic model to generalize across different speakers and conditions.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent changes the normalization parameter from source-specific to mixed/combined. By transforming the normalization approach from individual source normalization to mixed normalization, the system adapts the acoustic model to handle variations in cepstral means and variances across different speakers, thereby improving robustness for multi-speaker conversations.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If more diverse audio data is used for training, then robustness improves, but computational resource requirements increase

Engineering Contradiction:
Improverobustness of acoustic modelVSAvoidcomputational resource requirements
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary normalization of audio data from multiple sources using a mixed normalization factor before training the acoustic model. By pre-processing the data to account for source diversity through mixed normalization, the system reduces the computational burden during model training and inference, as the model doesn't need to learn normalization patterns from scratch during training.

Inventive Principle:
Principle #10Preliminary action

3Ease of manufacture

If conventional normalization is used for each audio source, then processing is simpler, but recognition accuracy for diverse voices deteriorates

Engineering Contradiction:
Improvesimplicity of processingVSAvoidrecognition accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent creates a universal mixed normalization factor that serves multiple audio sources simultaneously. This single normalization factor is derived from and applicable to multiple sources with different characteristics, making the normalization process multi-functional. The mixed normalization factor universally handles diverse voice characteristics while maintaining a unified processing approach.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12112767B2Acoustic data augmentation with mixed normalization factors
Publication Date: 2024.10.08 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12112767B2 patent drawing
  • US12112767B2 patent drawing
  • US12112767B2 patent drawing

AI summary

A method, computer system, and a computer program product for audio data augmentation are provided. Sets of audio data from different sources may be obtained. A respective normalization factor for at least two sources of the different sources may be calculated. The normalization factors from the at least two sources may be mixed to determine a mixed normalization factor. A first set of the sets may be normalized by using the mixed normalization factor and to obtain training data for training an acoustic model.