Audio Source Separation Using Spatial and Neural Cues
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing systems for source separation either ignore spatial cues or source cues, leading to suboptimal performance when combining different source separation processes.
Innovation Solution
A method and system that combine spatial cue based separation and source cue based separation by first processing the input audio signal with a spatial cue based separation module to determine mixing parameters, and then using a neural network based source cue based separation module to further process the intermediate audio signal and reduce noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If source cue based separation is used to remove noise and background audio, then audio quality is improved, but spatial information is lost
Solution Approach 1:
The system segments the source separation task into two distinct stages: first spatial cue-based separation to preserve spatial information and extract intermediate signals, then source cue-based separation to remove noise and background audio. This segmentation allows each module to specialize in one aspect without compromising the other.
Solution Approach 2:
The patent introduces an intermediate audio signal as a mediator between the spatial cue-based separation module and the source cue-based separation module. This intermediate signal carries both spatial information from the first stage and serves as input for noise removal in the second stage, preventing direct loss of spatial information.
2Loss of information
If spatial cue based separation is used to preserve spatial information, then spatial cues are maintained, but noise and background audio are not sufficiently removed
Solution Approach 1:
The system divides the source separation process into two specialized modules: spatial cue-based separation handles spatial information preservation, while source cue-based separation handles noise and background audio removal. This segmentation allows each module to optimize for its specific function without compromise.
Solution Approach 2:
The patent creates a continuous processing pipeline where the output of the spatial cue-based separation module feeds into the source cue-based separation module. This continuous action ensures that spatial information is preserved in the intermediate signal while subsequently enabling effective noise removal without interrupting the useful spatial cues.
3Adaptability or versatility
If multiple source separation processes are combined, then comprehensive audio processing is achieved, but system complexity increases
Solution Approach 1:
The system segments the complex source separation task into two manageable modules with distinct functions: spatial cue-based separation and source cue-based separation. Each module has a specific role and processes specific types of information, making the overall system easier to design, implement, and maintain despite handling multiple separation objectives.
Solution Approach 2:
The intermediate audio signal acts as a mediator that simplifies the interface between the two separation modules. The first module outputs spatially-separated intermediate signals with standard characteristics, and the second module accepts these standardized inputs for noise removal, reducing the complexity of coordinating multiple processes.
Data Source
AI summary
The present disclosure relates to a method and system for processing audio for source separation. The method comprises obtaining an input audio signal (A) comprising at least two channels and processing the input audio signal (A) with a spatial cue based separation module (10) to obtain an intermediate audio signal (B). The spatial cue based separation module (10) is configured to determine a mixing parameter of the at least two channels of the input audio signal (A) and modify the channels, based on the mixing parameter, to obtain the intermediate audio signal (B). The method further comprises processing the intermediate audio signal (B) with a source cue based separation module (20) to generate an output audio signal (C), wherein the source cue based separation module (20) is configured to implement a neural network trained to predict a noise reduced output audio signal (C) given the intermediate audio signal (B).


