Multi-Talker Babble Noise Reduction via Sub-Band Wavelet Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional cochlear implants and communication devices struggle to effectively improve speech intelligibility in noisy environments, particularly in the presence of multi-talker babble noise, due to the difficulty in accurately estimating background noise spectra and the introduction of distortion and artifacts from noise removal methods.
Innovation Solution
A method and system that classify input audio signals into noise-dominant and speech-dominant categories using principle component analysis, followed by parallel wavelet de-noising of sub-band components with Tunable Q-Factor Wavelet Transforms, adjusting thresholds based on previous noise-dominant frames, to aggressively de-noise noise-dominant signals while preserving speech clarity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If conventional noise removal methods are applied to noisy signals, then noise reduction is achieved, but distortion and artifacts are introduced that degrade speech intelligibility
Solution Approach 1:
The patent segments the noisy signal into multiple sub-bands using filter banks, allowing different processing strategies to be applied to different frequency regions. This segmentation enables selective noise removal in noise-dominant sub-bands while preserving speech content in speech-dominant sub-bands, thereby reducing overall distortion and artifacts.
Solution Approach 2:
The patent applies local quality by classifying each sub-band as either noise-dominant or speech-dominant and applying different de-noising strengths accordingly. Noise-dominant sub-bands receive aggressive de-noising while speech-dominant sub-bands receive minimal processing, optimizing the balance between noise removal and speech preservation at local frequency regions.
2Loss of information
If aggressive noise removal is applied to improve speech intelligibility, then noise reduction increases, but signal distortion increases leading to musical noise artifacts
Solution Approach 1:
The patent implements dynamic processing by adaptively adjusting the de-noising strength for each sub-band based on real-time classification of noise vs. speech dominance. The system dynamically switches between aggressive de-noising for noise-dominant sub-bands and gentle processing for speech-dominant sub-bands, optimizing the trade-off between noise removal and signal fidelity in each local region.
Solution Approach 2:
The patent changes processing parameters (de-noising strength) based on the classification of each sub-band. By adjusting the de-noising parameter dynamically according to whether a sub-band is noise-dominant or speech-dominant, the system achieves effective noise removal while maintaining signal fidelity and avoiding musical noise artifacts.
3Device complexity
If single-channel noise reduction algorithms are used, then processing complexity is reduced, but speech understanding in the presence of competing talkers remains difficult
Solution Approach 1:
The patent transitions from single-channel processing to multi-channel processing by dividing the audio signal into multiple frequency sub-bands. This dimensional expansion allows the system to exploit spectral information across different frequency regions, classifying each sub-band independently and applying targeted de-noising, thereby significantly improving speech intelligibility in babble noise while maintaining manageable processing complexity.
Data Source
AI summary
A system and method for improving intelligibility of speech is provided. The system and method may include obtaining an input audio signal frame, classifying the input audio signal frame into a first category or a second category, wherein the first category corresponds to the noise being stronger than the speech signal, and the second category corresponds to the speech signal being stronger than the noise, decomposing the input audio signal frame into a plurality of sub-band components; de-noising each sub-band component of the input audio signal frame in parallel by applying a first wavelet de-noising method including a first wavelet transform and a predetermined threshold for the sub-band component, and a second wavelet de-noising method including a second wavelet transform and the predetermined threshold for the sub-band component, wherein the predetermined threshold for each sub-band component is based on at least one previous noise-dominant signal frame received by the receiving arrangement.


