Audio Source Separation via Joint Additive and Independent Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio source separation methods face challenges in achieving stable and rapid convergence, particularly in real-world applications, due to issues like estimation instabilities, permutation indeterminacy, and mismatch between training data and actual audio properties, especially when prior information is unavailable.
Innovation Solution
A method and system that jointly utilize additive and independent/uncorrelated source modeling to determine spatial parameters based on linear combination and orthogonality characteristics, enabling perceptually natural audio source separation with stable and rapid convergence, even for highly non-stationary sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If blind source separation is performed without prior information, then adaptability to real-world applications is improved, but estimation stability and convergence reliability deteriorate
Solution Approach 1:
The patent combines additive source modeling (linear combination characteristic) and independent/uncorrelated source modeling (orthogonality characteristic) into a unified joint determination framework. This merging allows the system to leverage the strengths of both approaches: additive modeling provides flexibility for real-world adaptability while independent/uncorrelated modeling provides orthogonality constraints that ensure estimation stability and prevent permutation indeterminacy.
Solution Approach 2:
The patent dynamically adjusts the weighting between linear combination characteristic and orthogonality characteristic during the separation process. By changing the parameter weights adaptively, the system can maintain stability during convergence while preserving adaptability to different audio scenarios, resolving the contradiction between reliability and adaptability.
2Adaptability or versatility
If additive source modeling is used alone, then adaptability to non-stationary sources is improved, but estimation stability deteriorates due to permutation indeterminacy
Solution Approach 1:
The patent merges additive source modeling with independent/uncorrelated source modeling to create a hybrid approach. The additive component handles non-stationary sources effectively through linear combination, while the independent/uncorrelated component provides orthogonality constraints that eliminate permutation indeterminacy and ensure stable estimation.
3Stability of the object's composition
If independent/uncorrelated source modeling is used alone, then convergence stability is improved, but perceptual naturalness and adaptability deteriorate
Solution Approach 1:
The patent combines independent/uncorrelated source modeling with additive source modeling. The independent/uncorrelated approach provides stable convergence through orthogonality constraints, while the additive approach enhances perceptual naturalness by allowing flexible linear combinations that better represent real audio sources.
Solution Approach 2:
The patent dynamically adjusts the balance between orthogonality constraint strength and linear combination flexibility. By modifying parameters during processing, the system maintains convergence stability while improving perceptual quality and adaptability to different audio characteristics.
4Speed
If training data is used for source separation, then initial convergence speed is improved, but mismatch between training data and actual audio properties reduces reliability
Solution Approach 1:
The patent implements a self-adjusting mechanism where the system automatically adapts the weighting between linear combination and orthogonality characteristics based on the actual audio input properties. This self-service approach eliminates dependency on pre-trained data while maintaining fast convergence through the inherent mathematical constraints of the joint modeling framework.
Data Source
Figure 1~2
Figure 3
Figure 4
AI summary
A method of audio source separation from audio content is disclosed. The method includes determining a spatial parameter of an audio source based on a linear combination characteristic of the audio source and an orthogonality characteristic of two or more audio sources to be separated in the audio content. The method also includes separating the audio source from the audio content based on the spatial parameter. Corresponding system and computer program product are also disclosed.