Audio Source Separation for Stationary and Non-Stationary Noise Remixing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing neural network models for audio signal processing struggle to effectively separate and suppress noise components that differ from the predefined noise types used during training, leading to suboptimal performance in enhancing speech intelligibility and maintaining desired audio components in various environments.

Innovation Solution

A method for audio processing that separates speech, stationary noise, and non-stationary noise by using trained models to determine weighting factors, allowing for precise manipulation of these components to enhance speech intelligibility while preserving ambient noise, using a combination of speech isolator and stationary noise isolator models, and optionally applying bandpass filters to non-stationary noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a neural network model is trained to remove a specific type of predetermined noise, then the noise suppression performance for that noise type is improved, but the performance decreases when applied to remove noise defined differently from the training noise

Engineering Contradiction:
Improvenoise suppression performanceVSAvoidadaptability to different noise types
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the audio signal into multiple distinct components: speech content, stationary noise content, and non-stationary noise content. This segmentation is achieved through separate neural network models trained for each component type, allowing each model to specialize in removing its specific noise type while maintaining high performance across diverse noise conditions

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system creates a universal audio processing framework that handles multiple noise types through a combination of specialized models. The speech isolator model, stationary noise isolator model, and non-stationary noise isolator model work together to provide multi-functional noise suppression capability, making the system adaptable to various noise environments without requiring retraining

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If an aggressive speech separation model is used to treat all non-speech components as noise, then speech intelligibility is enhanced, but desired audio components such as birdsong and rattling leaves are suppressed

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidloss of desired audio components
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent applies different processing qualities to different audio components. The speech isolator model aggressively enhances speech intelligibility, while the stationary and non-stationary noise isolator models provide more selective suppression that preserves desired ambient sounds like birdsong and nature noises, allowing each component to be treated according to its specific characteristics

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system changes the parameters and characteristics of noise suppression by using separate models with different training objectives. The speech isolator focuses on speech enhancement, while the noise isolators are trained to preserve certain non-speech components, thereby maintaining speech intelligibility without losing desired audio information

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If a less aggressive speech separation model trained to remove only stationary background noise is used, then speech intelligibility is maintained, but unwanted non-stationary noise such as passing airplanes is not suppressed

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidunwanted non-stationary noise
Core Design Contradiction:
Measurement precisionVSObject-generated harmful factors

Solution Approach 1:

The patent segments noise suppression into two distinct functions: stationary noise removal by the stationary noise isolator model and non-stationary noise removal by the non-stationary noise isolator model. This segmentation allows each model to specialize in its respective noise type, enabling comprehensive noise suppression that includes both stationary background noise and transient non-stationary disturbances like passing airplanes

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system maintains continuous noise suppression across different noise types by chaining multiple isolator models. The stationary noise isolator continuously removes background noise, while the non-stationary noise isolator continuously suppresses transient disturbances, ensuring ongoing speech intelligibility enhancement without interruption or loss of effectiveness

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentEP4434032B1Source separation and remixing in signal processing
Publication Date: 2025.07.30 DOLBY LABORATORIES LICENSING CORP
  • EP4434032B1 patent drawingFigure 1a~2
  • EP4434032B1 patent drawingFigure 3a~3b
  • EP4434032B1 patent drawingFigure 3c~4

AI summary

The present disclosure relates to a method and audio processing system (1) for performing source separation. The method comprises obtaining (S1) an audio signal (Sin) including a mixture of speech content and noise content, determining (S2a, S2b, S2c), from the audio signal, speech content (formula A), stationary noise content (formula C) and non-speech content (formula B). The stationary noise content (formula C) is a true subset of the non-speech content (formula B) and the method further comprises determining (S3), based on a difference between the stationary noise content (formula C) and the non-speech content (formula B) a non-stationary noise content formula D), obtaining (S5) a set of weighting factors and forming (S6) a processed audio signal based on a combination of the speech content (formula A), the stationary noise content (formula C), and the non-stationary noise content (formula D) weighted with their respective weighting factor.