Signal Enhancer Using Weighted Filter Merging for Speech Intelligibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech enhancement techniques are not robust enough under high noise levels, as they either prioritize noise reduction over auditory perception or introduce complexity with spectrum holes and undesirable processing steps, leading to poor performance in automatic speech recognition systems.

Innovation Solution

A signal enhancer that generates multiple filters, including noise-reduction and noise-masking filters, and merges them using weighted sums based on detected speech probability, allowing for a compromise between SNR improvement and intelligibility enhancement, with adaptive weight determination and iterative filter adaptation to minimize differences between filtered and target signals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If noise reduction filtering is applied to maximize SNR, then noise is reduced, but speech component is attenuated and intelligibility deteriorates

Engineering Contradiction:
Improvenoise levelVSAvoidspeech intelligibility
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The patent combines multiple filtering approaches (spectral subtraction, Wiener filtering, and noise masking) into a unified signal enhancement system. By merging these different filtering methods and adaptively selecting their outputs based on speech presence probability, the system achieves both noise reduction and speech preservation, resolving the contradiction between SNR improvement and intelligibility maintenance.

Inventive Principle:
Principle #5Merging (Combining)

2Device complexity

If simple binary mask technique is used for noise masking, then processing complexity is reduced, but spectrum holes are created and performance deteriorates

Engineering Contradiction:
Improveprocessing complexityVSAvoidenhancement robustness
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent moves beyond the simple binary mask approach by introducing continuous weighting factors and speech presence probability parameters. Instead of binary selection, the system uses parameter-based blending where filtered signals are combined with weighted sums based on detected speech probability, creating smoother transitions and avoiding spectrum holes while maintaining robustness.

Inventive Principle:
Principle #35Parameter changes

3Object-affected harmful factors

If strong noise reduction processing is applied, then SNR is improved, but speech component is attenuated resulting in poor ASR performance

Engineering Contradiction:
Improvesignal-to-noise ratioVSAvoidautomatic speech recognition accuracy
Core Design Contradiction:
Object-affected harmful factorsVSMeasurement precision

Solution Approach 1:

The patent implements dynamic adaptation of filtering strength based on real-time speech presence detection. The system continuously adjusts the filtering parameters and weight values according to the detected speech probability in each frame, making the noise reduction processing dynamic rather than static. This ensures that speech components are preserved when present while maintaining noise reduction when speech is absent, thereby improving ASR performance.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP3692529B1An apparatus and a method for signal enhancement
Publication Date: 2023.05.24 HUAWEI TECH CO LTD
  • EP3692529B1 patent drawingFigure 1
  • EP3692529B1 patent drawingFigure 2
  • EP3692529B1 patent drawingFigure 3

AI summary

A signal enhancer comprises an input configured to receive an audio signal. It also comprises a processor that is configured to generate at least two different filters based on the audio signal X of a current frame. The processor is also configured to generate at least two filtered signals by applying each of the at least two filters to the audio signal of the current frame respectively. The processor is further configured to generate an enhanced audio signal Y for the current frame by merging the n filtered signals. This improves the robustness of the signal enhancement.