Audio Source Separation Using Time-Domain Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for audio source separation in teleconferencing scenarios with stereo microphones face challenges in efficiently separating signals from multiple sources due to crosstalk and propagation delays, particularly in environments with convolutive mixtures, where existing approaches often result in permutation issues and increased computational complexity.

Innovation Solution

A time-domain approach using fractional delay allpass filters and attenuation factors to model relative transfer functions between microphones, combined with a novel optimization method called 'random directions' to minimize the negative Kullback-Leibler divergence objective function, effectively separating audio sources by iteratively refining delay and attenuation values.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If frequency-domain ICA methods are used for audio source separation, then separation capability is improved, but computational complexity and permutation issues increase

Engineering Contradiction:
Improveseparation capabilityVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces frequency-domain processing with time-domain processing. Instead of using FFT-based ICA methods that require complex frequency bin processing and permutation resolution, the invention directly processes audio signals in the time domain using simple linear combinations of microphone signals with time-varying coefficients, eliminating the need for frequency-domain transformations and associated computational complexity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent segments the audio processing into independent time frames, where separation coefficients are adapted frame-by-frame. This allows the system to process each time frame independently without requiring global frequency-domain analysis, reducing computational complexity while maintaining separation capability through temporal localization

Inventive Principle:
Principle #1Segmentation

2Object-affected harmful factors

If stereo microphone separation is used, then crosstalk reduction is improved, but system delay increases

Engineering Contradiction:
ImprovecrosstalkVSAvoidsystem delay
Core Design Contradiction:
Object-affected harmful factorsVSLoss of time

Solution Approach 1:

The patent implements adaptive coefficient adjustment where the separation system automatically adapts to changing acoustic environments in real-time. The coefficients are updated using simple recursive algorithms that process current and past signal values without requiring lookahead buffers or complex optimization, achieving crosstalk reduction with minimal system delay

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent uses time-varying separation coefficients that adapt to changing acoustic conditions. Instead of fixed frequency-domain filters that introduce phase delay, the system dynamically adjusts time-domain coefficients based on instantaneous signal statistics, reducing group delay while maintaining separation performance

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If complex optimization algorithms are used, then separation accuracy is improved, but processing speed decreases

Engineering Contradiction:
Improveseparation accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent uses simple, computationally inexpensive optimization algorithms that can be executed rapidly. Instead of employing complex gradient-based or evolutionary optimization methods, the invention uses straightforward recursive least squares or normalized least squares algorithms that provide sufficient separation accuracy with minimal computational burden, enabling real-time processing

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Solution Approach 2:

The patent applies optimization only to the necessary extent for achieving adequate separation. Rather than pursuing perfect separation through exhaustive optimization, the system uses simplified cost functions and termination criteria that provide sufficient separation accuracy for teleconferencing applications while maintaining high processing speed

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12190899B2Apparatus and method for acquiring a plurality of audio signals associated with different sound sources
Publication Date: 2025.01.07 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US12190899B2 patent drawing
  • US12190899B2 patent drawing
  • US12190899B2 patent drawing

AI summary

Methods, apparatus and techniques for acquiring output signals associated with different sources (such as audio sources) are presented. A first input signal is combined with a delayed and scaled version of a second input signal, to acquire a first output signal. A second input signal is combined with a delayed and scaled version of the first input signal, to acquire a second output signal. Using a random direction optimization, scaling values (forming a candidates' vector) are determined by iteratively modifying the candidates' vector. A Kullback-Leibler divergence is measured. The first output signal and the second output signal are selected to be those measurements associated with the candidate parameters associated with Kullback-Leibler divergence which indicates lowest similarity.