Neural Network Training Apparatus for Noise-Robust Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition technologies face performance degradation due to noise mismatches between training and actual environments, leading to inefficient noise removal.
Innovation Solution
A neural network training apparatus and method that employs a primary trainer for clean data training and a secondary trainer for noise-robust training using a probability distribution of the output class calculated during primary training, with objective functions that combine clean and noisy data training.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If training is performed using only clean training data, then the neural network model achieves high accuracy on clean data, but the speech recognition performance degrades in noisy real-world environments
Solution Approach 1:
The patent applies preliminary action by performing primary training on clean data first to establish a solid foundation, then subsequently performing secondary training on noisy data to prepare the model for real-world conditions. This sequential preparation allows the model to first learn accurate patterns from clean data and then adapt to handle noise, resolving the contradiction between accuracy and noise robustness.
Solution Approach 2:
The patent converts the harmful effect of noise into a beneficial training opportunity by introducing noisy training data in the secondary training phase. The noise that would normally degrade performance is instead utilized to teach the model to distinguish between relevant speech information and irrelevant background noise, thereby improving real-world performance.
2Adaptability or versatility
If training is performed using only noisy training data, then the model becomes robust to noise, but the training efficiency and convergence speed decrease
Solution Approach 1:
The patent segments the training process into two distinct phases: primary training using clean data and secondary training using noisy data. This segmentation allows the model to first efficiently learn from clean data without noise interference, then subsequently adapt to noisy conditions. By dividing the training into manageable stages, the patent maintains training efficiency while achieving noise robustness.
Solution Approach 2:
The primary training phase serves as a preliminary action that establishes the model's basic capabilities on clean data before introducing noise in the secondary phase. This preliminary foundation ensures that the model has learned accurate patterns efficiently, and the subsequent noisy training builds upon this foundation rather than starting from scratch, thereby maintaining overall training efficiency.
3Productivity
If a single training phase is used, then the training process is simple and fast, but the model cannot handle both clean and noisy data effectively
Solution Approach 1:
The patent segments the training process into two distinct phases: primary training for clean data and secondary training for noisy data. Each phase is optimized for its specific condition, allowing the model to learn effective representations from both clean and noisy data. This segmentation enables the model to achieve multi-condition performance while maintaining reasonable training speed through structured progression.
Solution Approach 2:
The patent introduces dynamic adaptation by switching from static clean data training to dynamic noisy data training in the second phase. This dynamic approach allows the model to adapt its learning process to different data conditions, improving its ability to handle both clean and noisy data effectively while maintaining training efficiency through purposeful progression.
Data Source
AI summary
A neural network training apparatus includes a primary trainer configured to perform a primary training of a neural network model based on clean training data and target data corresponding to the clean training data; and a secondary trainer configured to perform a secondary training of the neural network model on which the primary training has been performed based on noisy training data and an output probability distribution of an output class for the clean training data calculated during the primary training of the neural network model.


