Selective Audio Source Enhancement via Multistage BSS

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech enhancement methods fail to provide satisfactory performance in far-field applications with limited channels and significant reverberation, leading to inadequate energy propagation and poor automatic speech recognition in noisy environments.

Innovation Solution

The implementation of a selective audio source enhancement system using Blind Source Separation (BSS) techniques, specifically a multistage processing approach involving source detection, weighted natural gradient, constrained independent component analysis (ICA), and spectral filtering, optimized for limited hardware resources, allowing for robust speech recognition and noise suppression.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional speech enhancement methods (single or multiple channel) are used, then processing simplicity is maintained, but performance is insufficient in far-field applications with limited channels and significant reverberation

Engineering Contradiction:
Improvespeech enhancement performanceVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the speech enhancement task into multiple stages: first estimating the speech power spectral density, then estimating the noise power spectral density, and finally computing the enhanced speech signal. This multi-stage approach allows complex processing to be broken down into manageable steps that can be implemented with limited hardware resources while maintaining high performance in far-field conditions with reverberation.

Inventive Principle:
Principle #1Segmentation

2Reliability

If beam forming methods are used to enhance signals from predefined spatial directions, then directional signal enhancement is improved, but effectiveness decreases when energy propagation over steering geometrical direction is limited

Engineering Contradiction:
Improvedirectional signal enhancementVSAvoidenergy propagation limitation
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent changes the fundamental parameters of the enhancement approach by moving from spatial filtering (beam forming) to spectral processing. Instead of relying on directional energy propagation, the method estimates speech and noise power spectral densities and processes signals in the frequency domain. This allows effective enhancement even when directional energy propagation is limited, as the method does not depend on strong spatial separation between signal and noise.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If continuous signal-to-noise ratio estimation is used in discrete time-spectral domain, then enhancement effectiveness is improved for stationary noise, but performance degrades when noise exhibits high energy variation (non-stationarity)

Engineering Contradiction:
Improveenhancement effectivenessVSAvoidnoise stationarity adaptability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamics by using time-varying estimates of speech and noise power spectral densities. The method continuously updates these estimates adaptively, allowing the enhancement process to respond to changing noise conditions. This dynamic approach enables effective handling of non-stationary noise with high energy variation, as the system can adjust its estimates in real-time rather than relying on fixed stationary assumptions.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10123113B2Selective audio source enhancement
Publication Date: 2018.11.06 SYNAPTICS INC
  • US10123113B2 patent drawing
  • US10123113B2 patent drawing
  • US10123113B2 patent drawing

AI summary

A selective audio source enhancement system includes a processor and a memory, and a pre-processing unit configured to receive audio data including a target audio signal, and to perform sub-band domain decomposition of the audio data to generate buffered outputs. In addition, the system includes a target source detection unit configured to receive the buffered outputs, and to generate a target presence probability corresponding to the target audio signal, as well as a spatial filter estimation unit configured to receive the target presence probability, and to transform frames buffered in each sub-band into a higher resolution frequency-domain. The system also includes a spectral filtering unit configured to retrieve a multichannel image of the target audio signal and noise signals associated with the target audio signal, and an audio synthesis unit configured to extract an enhanced mono signal corresponding to the target audio signal from the multichannel image.