Audio Signal Processing via Source Separation and Coherence Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multiple-microphone speech enhancement algorithms face challenges in reverberant environments and diffuse noise, failing to provide artefact-free enhancement and intelligibility improvement, with existing methods either relying on complex model parameters or lacking in noise reduction in time-frequency domains with speech contributions.

Innovation Solution

A composite method combining source separation and coherence-based envelope filtering in the Bark domain, where source separation is used in coherent sound fields and coherence-based filtering in diffuse fields, with sound field diffuseness detection to optimize signal enhancement and reconstruct a virtual stereophonic sound field.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If source separation algorithms are used in reverberant environments, then speech and noise can be separated in coherent sound fields, but the method fails to provide artefact-free enhancement in diffuse noise fields

Engineering Contradiction:
Improvespeech enhancement qualityVSAvoidperformance across different acoustic environments
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system dynamically adapts its processing strategy based on the detected sound field characteristics. When coherent sound fields are detected, source separation algorithms are applied; when diffuse noise fields are detected, coherence-based envelope filtering is applied. This dynamic adaptation resolves the contradiction by making the system versatile across different acoustic environments while maintaining reliable speech enhancement in each specific condition.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes its processing parameters based on the acoustic environment. The diffuseness parameter is used to determine which enhancement strategy to apply, allowing the system to optimize its performance for either coherent or diffuse sound fields. This parameter-based adaptation enables the system to maintain high speech enhancement quality across varying acoustic conditions.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If complex model parameters are used for source separation, then speech and noise can be separated in coherent fields, but the device complexity increases

Engineering Contradiction:
Improvesource separation accuracyVSAvoidmodel parameter complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

Instead of always applying complex source separation algorithms, the system applies them only partially - specifically when coherent sound fields are detected. In diffuse noise fields, simpler coherence-based filtering is used. This partial application of complex methods reduces device complexity while maintaining source separation accuracy when it is most beneficial.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system extracts and applies only the necessary processing complexity for each acoustic condition. Complex source separation model parameters are extracted and applied only when needed (in coherent fields), while simpler methods are used otherwise. This extraction approach maintains separation accuracy when required while reducing overall device complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

3Manufacturing precision

If single-channel speech enhancement algorithms are used, then signal quality can be improved, but speech intelligibility remains unchanged

Engineering Contradiction:
Improvesignal qualityVSAvoidspeech intelligibility
Core Design Contradiction:
Manufacturing precisionVSMeasurement precision

Solution Approach 1:

The system transitions from single-channel processing to multi-channel spatial processing by utilizing microphone arrays and analyzing spatial characteristics of sound fields. This dimensional change from temporal-only processing to spatio-temporal processing enables improvement in both signal quality and speech intelligibility through spatial filtering and beamforming techniques.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system introduces spatial coherence and sound field diffuseness as intermediary parameters that bridge signal quality improvement and intelligibility enhancement. By analyzing these intermediary spatial characteristics, the system can apply appropriate enhancement strategies that simultaneously improve both signal quality and speech intelligibility.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS7761291B2Method for processing audio-signals
Publication Date: 2010.07.20 OTICON
  • US7761291B2 patent drawing
  • US7761291B2 patent drawing
  • US7761291B2 patent drawing

AI summary

The invention regards a method for processing audio-signals whereby audio signals are captured at two spaced apart locations and subject to a transformation in the perceptual domain (Bar or Mel), whereupon: a) a (blind or supervised) source separation process is performed to give a first estimate of the wanted signal parts and the noise parts of the microphone signals and b) a coherence based separation process is performed to give a second estimate of the wanted signal parts and the noise parts of the microphone signals, and where further a sound field diffuseness detection is performed on the at least two signals, whereby further the sound field diffuseness detections is used to mix the output from the blind source separation and the coherence based separation process in order to achieve the best possible signal. The transfer functions calculated from the source separation are used to reconstruct a virtual stereophonic sound field in restore the spatial information about the source position in the enhanced signals.