Neural Network Feature Combiner for Speech Enhancement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech enhancement technologies face challenges in accurately estimating the noise spectrum, especially in non-stationary environments, leading to residual noise and distortions, particularly when using methods that require multiple spectrogram processing steps or are sensitive to noise mismatches.
Innovation Solution
A neural network-based feature combiner processes band-wise spectral shape features to estimate noise and signal-to-noise ratios, using a supervised learning method that learns characteristics of speech to provide accurate noise estimation without inherent delays, suitable for non-stationary noise conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If spectral weighting methods are used to reduce noise energy, then noise reduction is achieved, but residual noise and distortions occur due to inaccurate noise spectrum estimation in non-stationary environments
Solution Approach 1:
The patent implements dynamic noise spectrum estimation by continuously updating the noise spectrum based on current signal characteristics rather than using static or slowly adapting estimates. This allows the system to track non-stationary noise environments and maintain accurate noise spectrum estimation despite changing conditions, thereby reducing residual noise and distortions while preserving speech quality.
2Measurement precision
If multiple spectrogram processing steps are used for noise estimation, then processing thoroughness is improved, but system delays increase
Solution Approach 1:
The patent extracts essential noise estimation functionality from complex multi-step spectrogram processing by directly estimating the noise spectrum from the input signal in a streamlined manner. This extraction of the core estimation function eliminates unnecessary processing steps while maintaining accurate noise spectrum estimation, thereby reducing system delays without sacrificing noise estimation precision.
3Object-affected harmful factors
If noise reduction schemes are applied in low SNR situations, then noise attenuation is improved, but processing artifacts are introduced
Solution Approach 1:
The patent employs feedback mechanisms by continuously monitoring the estimated noise spectrum and signal characteristics to adaptively control the noise reduction process. This feedback allows the system to adjust processing intensity based on current conditions, preventing excessive noise attenuation that would cause artifacts in low SNR situations while still achieving effective noise reduction when appropriate.
Data Source
AI summary
An apparatus for processing an audio signal to obtain control information for a speech enhancement filter comprises a feature extractor for extracting at least one feature per frequency band of a plurality of frequency bands of a short-time spectral representation of a plurality of short-time spectral representations, where the at least one feature represents a spectral shape of the short-time spectral representation in the frequency band. The apparatus additionally comprises a feature combiner for combining the at least one feature for each frequency band using combination parameters to obtain the control information for the speech enhancement filter for a time portion of the audio signal. The feature combiner can use a neural network regression method, which is based on combination parameters determined in a training phase for the neural network.