Neural Network Feature Combiner for Speech Enhancement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech enhancement technologies face challenges in accurately estimating the noise spectrum, especially in non-stationary environments, leading to residual noise and distortions, particularly when using methods that require multiple spectrogram processing steps or are sensitive to noise mismatches.

Innovation Solution

A neural network-based feature combiner processes band-wise spectral shape features to estimate noise and signal-to-noise ratios, using a supervised learning method that learns characteristics of speech to provide accurate noise estimation without inherent delays, suitable for non-stationary noise conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If spectral weighting methods are used to reduce noise energy, then noise reduction is achieved, but residual noise and distortions occur due to inaccurate noise spectrum estimation in non-stationary environments

Engineering Contradiction:
Improvenoise energyVSAvoidnoise spectrum estimation accuracy
Core Design Contradiction:
Object-affected harmful factorsVSMeasurement precision

Solution Approach 1:

The patent implements dynamic noise spectrum estimation by continuously updating the noise spectrum based on current signal characteristics rather than using static or slowly adapting estimates. This allows the system to track non-stationary noise environments and maintain accurate noise spectrum estimation despite changing conditions, thereby reducing residual noise and distortions while preserving speech quality.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If multiple spectrogram processing steps are used for noise estimation, then processing thoroughness is improved, but system delays increase

Engineering Contradiction:
Improvenoise estimation accuracyVSAvoidsystem delays
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts essential noise estimation functionality from complex multi-step spectrogram processing by directly estimating the noise spectrum from the input signal in a streamlined manner. This extraction of the core estimation function eliminates unnecessary processing steps while maintaining accurate noise spectrum estimation, thereby reducing system delays without sacrificing noise estimation precision.

Inventive Principle:
Principle #2Taking out (Extraction)

3Object-affected harmful factors

If noise reduction schemes are applied in low SNR situations, then noise attenuation is improved, but processing artifacts are introduced

Engineering Contradiction:
Improvenoise levelVSAvoidprocessing artifacts
Core Design Contradiction:
Object-affected harmful factorsVSObject-generated harmful factors

Solution Approach 1:

The patent employs feedback mechanisms by continuously monitoring the estimated noise spectrum and signal characteristics to adaptively control the noise reduction process. This feedback allows the system to adjust processing intensity based on current conditions, preventing excessive noise attenuation that would cause artifacts in low SNR situations while still achieving effective noise reduction when appropriate.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP2151822B8Apparatus and method for processing an audio signal for speech enhancement using a feature extraction
Publication Date: 2018.10.24 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV

AI summary

An apparatus for processing an audio signal to obtain control information for a speech enhancement filter comprises a feature extractor for extracting at least one feature per frequency band of a plurality of frequency bands of a short-time spectral representation of a plurality of short-time spectral representations, where the at least one feature represents a spectral shape of the short-time spectral representation in the frequency band. The apparatus additionally comprises a feature combiner for combining the at least one feature for each frequency band using combination parameters to obtain the control information for the speech enhancement filter for a time portion of the audio signal. The feature combiner can use a neural network regression method, which is based on combination parameters determined in a training phase for the neural network.