Audio Source Separation for Stationary and Non-Stationary Noise Remixing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing neural network models for audio signal processing struggle to effectively separate and suppress noise components that differ from the predefined noise types used during training, leading to suboptimal performance in enhancing speech intelligibility and maintaining desired audio components in various environments.
Innovation Solution
A method for audio processing that separates speech, stationary noise, and non-stationary noise by using trained models to determine weighting factors, allowing for precise manipulation of these components to enhance speech intelligibility while preserving ambient noise, using a combination of speech isolator and stationary noise isolator models, and optionally applying bandpass filters to non-stationary noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a neural network model is trained to remove a specific type of predetermined noise, then the noise suppression performance for that noise type is improved, but the performance decreases when applied to remove noise defined differently from the training noise
Solution Approach 1:
The patent segments the audio signal into multiple distinct components: speech content, stationary noise content, and non-stationary noise content. This segmentation is achieved through separate neural network models trained for each component type, allowing each model to specialize in removing its specific noise type while maintaining high performance across diverse noise conditions
Solution Approach 2:
The system creates a universal audio processing framework that handles multiple noise types through a combination of specialized models. The speech isolator model, stationary noise isolator model, and non-stationary noise isolator model work together to provide multi-functional noise suppression capability, making the system adaptable to various noise environments without requiring retraining
2Measurement precision
If an aggressive speech separation model is used to treat all non-speech components as noise, then speech intelligibility is enhanced, but desired audio components such as birdsong and rattling leaves are suppressed
Solution Approach 1:
The patent applies different processing qualities to different audio components. The speech isolator model aggressively enhances speech intelligibility, while the stationary and non-stationary noise isolator models provide more selective suppression that preserves desired ambient sounds like birdsong and nature noises, allowing each component to be treated according to its specific characteristics
Solution Approach 2:
The system changes the parameters and characteristics of noise suppression by using separate models with different training objectives. The speech isolator focuses on speech enhancement, while the noise isolators are trained to preserve certain non-speech components, thereby maintaining speech intelligibility without losing desired audio information
3Measurement precision
If a less aggressive speech separation model trained to remove only stationary background noise is used, then speech intelligibility is maintained, but unwanted non-stationary noise such as passing airplanes is not suppressed
Solution Approach 1:
The patent segments noise suppression into two distinct functions: stationary noise removal by the stationary noise isolator model and non-stationary noise removal by the non-stationary noise isolator model. This segmentation allows each model to specialize in its respective noise type, enabling comprehensive noise suppression that includes both stationary background noise and transient non-stationary disturbances like passing airplanes
Solution Approach 2:
The system maintains continuous noise suppression across different noise types by chaining multiple isolator models. The stationary noise isolator continuously removes background noise, while the non-stationary noise isolator continuously suppresses transient disturbances, ensuring ongoing speech intelligibility enhancement without interruption or loss of effectiveness
Data Source
Figure 1a~2
Figure 3a~3b
Figure 3c~4
AI summary
The present disclosure relates to a method and audio processing system (1) for performing source separation. The method comprises obtaining (S1) an audio signal (Sin) including a mixture of speech content and noise content, determining (S2a, S2b, S2c), from the audio signal, speech content (formula A), stationary noise content (formula C) and non-speech content (formula B). The stationary noise content (formula C) is a true subset of the non-speech content (formula B) and the method further comprises determining (S3), based on a difference between the stationary noise content (formula C) and the non-speech content (formula B) a non-stationary noise content formula D), obtaining (S5) a set of weighting factors and forming (S6) a processed audio signal based on a combination of the speech content (formula A), the stationary noise content (formula C), and the non-stationary noise content (formula D) weighted with their respective weighting factor.