Reverberant Audio Source Separation Using Learned Mixing Parameters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Source separation in reverberant environments is challenging due to increased spatial spread of sources caused by echoes, and existing methods require prior knowledge of source positions and room characteristics, which is often impractical in real-world applications.
Innovation Solution
A semi-supervised method for source separation that learns mixing parameters and spectral bases during a training phase with only one source active, and applies these parameters to estimate a reconstruction model during a testing phase with all sources active, allowing for accurate source separation without prior information about the recording environment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If prior knowledge of source positions and room characteristics is used for source separation, then separation accuracy is improved, but the complexity of the system increases and practical applicability decreases
Solution Approach 1:
The patent applies preliminary action by performing a training phase before the actual source separation task. During this training phase, the system learns mixing parameters and spectral bases from recorded audio data without requiring knowledge of source positions or room characteristics. These learned parameters are then stored and reused during the testing phase, enabling accurate source separation without repeatedly solving the complex inverse problem of estimating mixing parameters from scratch.
Solution Approach 2:
The patent uses copying by creating a statistical model (reconstruction model) that copies the essential characteristics of the mixing process. Instead of directly analyzing and solving the complex reverberant mixing equations each time, the system creates a simplified statistical representation during training that captures the acoustic environment's properties. This copied model can then be applied repeatedly for source separation without re-analyzing the original complex system.
2Adaptability or versatility
If existing source separation methods are applied in reverberant environments, then source separation can be attempted, but the spatial spread of sources due to echoes degrades separation performance
Solution Approach 1:
The patent applies parameter changes by transforming the source separation problem from the time domain to the frequency domain using Short-Time Fourier Transform (STFT). This parameter transformation allows the system to analyze and separate sources in the frequency domain where the mixing model becomes simpler and more tractable. The system learns frequency-domain mixing parameters during training and uses these to guide separation in the frequency domain, which can then be transformed back to time domain signals.
Solution Approach 2:
The patent substitutes the mechanical/acoustic analysis approach with a statistical learning approach. Instead of directly analyzing the physical acoustic equations and reverberation patterns, the system uses statistical models to learn the mixing characteristics from data. This substitution replaces complex physical modeling with data-driven statistical inference, making the system more robust to reverberation without requiring explicit knowledge of room acoustics.
3Adaptability or versatility
If training data with multiple sources is used, then the model can learn more complex mixing patterns, but it becomes impossible to isolate individual source characteristics
Solution Approach 1:
The patent applies segmentation by dividing the learning process into two distinct phases: training phase and testing phase. During the training phase, the system processes recordings where individual sources are isolated, learning their specific spectral bases and mixing parameters separately. This segmentation allows the system to capture pure source characteristics without contamination from other sources. The learned parameters are then combined and applied during the testing phase to handle complex multi-source mixing scenarios.
Data Source
AI summary
Embodiments of source separation for reverberant environment are disclosed. According to a method, first microphone signals for each individual one of at least one source are captured respectively by at least two microphones for a period during which only the individual one produces sounds. Mixing parameters for modeling acoustic paths between the at least one source and the at least two microphones are learned by a processor based on the first microphone signals. Second microphone signals are captured respectively by the at least two microphones for a period during which all of the at least one source produce sounds. The reconstruction model is estimated by the processor based on the mixing parameters and second microphone signals. The processor performs the source separation by applying the reconstruction model.


