Audio Source Separation Using Time-Domain Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for audio source separation in teleconferencing scenarios with stereo microphones face challenges in efficiently separating signals from multiple sources due to crosstalk and propagation delays, particularly in environments with convolutive mixtures, where existing approaches often result in permutation issues and increased computational complexity.
Innovation Solution
A time-domain approach using fractional delay allpass filters and attenuation factors to model relative transfer functions between microphones, combined with a novel optimization method called 'random directions' to minimize the negative Kullback-Leibler divergence objective function, effectively separating audio sources by iteratively refining delay and attenuation values.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If frequency-domain ICA methods are used for audio source separation, then separation capability is improved, but computational complexity and permutation issues increase
Solution Approach 1:
The patent replaces frequency-domain processing with time-domain processing. Instead of using FFT-based ICA methods that require complex frequency bin processing and permutation resolution, the invention directly processes audio signals in the time domain using simple linear combinations of microphone signals with time-varying coefficients, eliminating the need for frequency-domain transformations and associated computational complexity
Solution Approach 2:
The patent segments the audio processing into independent time frames, where separation coefficients are adapted frame-by-frame. This allows the system to process each time frame independently without requiring global frequency-domain analysis, reducing computational complexity while maintaining separation capability through temporal localization
2Object-affected harmful factors
If stereo microphone separation is used, then crosstalk reduction is improved, but system delay increases
Solution Approach 1:
The patent implements adaptive coefficient adjustment where the separation system automatically adapts to changing acoustic environments in real-time. The coefficients are updated using simple recursive algorithms that process current and past signal values without requiring lookahead buffers or complex optimization, achieving crosstalk reduction with minimal system delay
Solution Approach 2:
The patent uses time-varying separation coefficients that adapt to changing acoustic conditions. Instead of fixed frequency-domain filters that introduce phase delay, the system dynamically adjusts time-domain coefficients based on instantaneous signal statistics, reducing group delay while maintaining separation performance
3Measurement precision
If complex optimization algorithms are used, then separation accuracy is improved, but processing speed decreases
Solution Approach 1:
The patent uses simple, computationally inexpensive optimization algorithms that can be executed rapidly. Instead of employing complex gradient-based or evolutionary optimization methods, the invention uses straightforward recursive least squares or normalized least squares algorithms that provide sufficient separation accuracy with minimal computational burden, enabling real-time processing
Solution Approach 2:
The patent applies optimization only to the necessary extent for achieving adequate separation. Rather than pursuing perfect separation through exhaustive optimization, the system uses simplified cost functions and termination criteria that provide sufficient separation accuracy for teleconferencing applications while maintaining high processing speed
Data Source
AI summary
Methods, apparatus and techniques for acquiring output signals associated with different sources (such as audio sources) are presented. A first input signal is combined with a delayed and scaled version of a second input signal, to acquire a first output signal. A second input signal is combined with a delayed and scaled version of the first input signal, to acquire a second output signal. Using a random direction optimization, scaling values (forming a candidates' vector) are determined by iteratively modifying the candidates' vector. A Kullback-Leibler divergence is measured. The first output signal and the second output signal are selected to be those measurements associated with the candidate parameters associated with Kullback-Leibler divergence which indicates lowest similarity.


