Generative Noise Model for Robust ASR Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current Automatic Speech Recognition (ASR) systems are not robust enough to handle a wide range of real-world noise and acoustic distortions, particularly in unseen conditions, due to limitations in multi-conditioned training datasets that often repeat noise types and fail to capture all possible noise scenarios.
Innovation Solution
A method and system for generating synthetic multi-conditioned data sets using a generative noise model that creates unique noise signals for additive and channel distortions, spanning the entire noise space, and applying constraints to simulate real-world effects, allowing for robust training of ASR models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a large number of noise signals are used during training, then robustness against additive noise is improved, but the possibility of noise signal repetition increases, effectively reducing the strength of training dataset
Solution Approach 1:
The patent uses Room Impulse Responses (RIRs) as templates to generate multiple synthetic noise signals through linear combinations. Instead of using many actual recorded noise signals that may repeat, the system creates diverse noise variations by combining a small set of basis RIRs with different parameters, thereby maintaining noise signal diversity while achieving robustness.
Solution Approach 2:
The patent varies parameters such as Signal-to-Noise Ratio (SNR), linear combination coefficients, and RIR selection to generate diverse noise conditions from a limited set of basis signals. By changing these parameters systematically, the training dataset achieves coverage of various noise scenarios without requiring a large number of distinct recorded noise signals.
2Reliability
If existing multi-conditioned training approaches are used, then performance in seen conditions is improved, but performance in unseen degradation conditions lags behind
Solution Approach 1:
The patent performs preliminary action by generating a comprehensive set of synthetic noise conditions during training that cover both seen and unseen degradation scenarios. By anticipating potential unseen conditions and incorporating them into the training dataset through parametric variations of basis RIRs, the system prepares the ASR model to generalize better to novel noise conditions while maintaining performance on known conditions.
3Adaptability or versatility
If noise basis signals spanning entire noise space are used, then coverage of noise types is improved, but complexity of generating unique noise signals increases
Solution Approach 1:
The patent segments the complex task of generating diverse noise signals into manageable components: a small set of basis RIRs representing different acoustic environments, which are then combined through linear operations. This segmentation allows systematic coverage of noise space while keeping the generation process computationally tractable through structured combination of basis elements.
Data Source
AI summary
Performance of Automatic Speech Recognition (ASR) for robustness against real world noises and channel distortions is critical. Embodiments herein provide method and system for generating synthetic multi-conditioned data sets for additive noise and channel distortion for training multi-conditioned acoustic models for robust ASR. The method provides a generative noise model generating plurality of types of noise signals for additive noise based on weighted linear combination of plurality of noise basis signals and channel distortion based on estimated channel responses. The generative noise model is a parametric model, wherein basis function selection, number of basis functions to be combined linearly and weightages to be applied to the combinations is tunable, thereby enabling generation of wide variety of noise signals. Further, the noise signals are added to set of training speech utterances under set of constraints providing the multi-conditioned data sets, imitating real world effects.


