Adaptive De-Esser for In-Car Speech Sibilance Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In-Car Communication (ICC) systems face challenges in handling sibilant sounds, which are speaker-dependent and become dominant due to noise suppression, leading to an overemphasis of higher frequency bands, making them annoying for listeners.
Innovation Solution
A dynamic deesser method that employs spectral envelope and phoneme-dependent attenuation, using psychoacoustic measures like sharpness to adaptively suppress sibilant sounds, considering various speakers and acoustic scenarios, and integrating this method into ICC systems for improved speech signal processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If noise suppression is applied to improve speech clarity in ICC systems, then speech intelligibility is improved, but sibilant sounds become over-emphasized and more annoying
Solution Approach 1:
The patent applies local quality by implementing frequency-selective attenuation that targets only the specific frequency ranges where sibilant sounds occur (typically 2-8 kHz) while preserving other frequency components. The deesser dynamically adjusts attenuation levels based on the detected sibilant energy in different frequency bands, thereby locally modifying the spectral characteristics to reduce annoyance without affecting overall speech intelligibility.
Solution Approach 2:
The patent employs parameter changes by dynamically adjusting the attenuation parameter based on the detected sibilant characteristics. The system continuously monitors the spectral content and adapts the attenuation level, frequency range, and time constant parameters in real-time to match the speaker's specific sibilant profile and the acoustic environment, thereby resolving the contradiction between noise suppression benefits and sibilant annoyance.
2Device complexity
If a static equalizer setting is used to simplify system configuration, then system complexity is reduced, but the system cannot adapt to different speakers' sibilant characteristics
Solution Approach 1:
The patent applies dynamics by transforming the static equalizer setting into a dynamic adaptive system. The deesser continuously analyzes the input speech signal to detect sibilant characteristics and automatically adjusts the attenuation parameters in real-time. This dynamic adaptation allows the system to accommodate different speakers' sibilant profiles without requiring manual reconfiguration, thereby resolving the contradiction between system simplicity and adaptability.
Solution Approach 2:
The patent employs self-service by enabling the system to automatically detect and adapt to each speaker's unique sibilant characteristics without external intervention. The deesser performs self-adjustment by monitoring the spectral content and modifying its attenuation parameters based on the detected sibilant energy, thereby providing speaker-specific optimization automatically and eliminating the need for manual tuning.
3Measurement precision
If a deesser method optimized for known speakers is used to improve sibilant suppression, then suppression effectiveness is improved, but the method cannot work robustly for unknown speakers and various acoustic scenarios
Solution Approach 1:
The patent applies universality by designing a deesser algorithm that functions effectively across diverse speakers and acoustic scenarios without requiring speaker-specific optimization. The system uses universal sibilant detection criteria based on spectral characteristics that are common to all speakers, combined with adaptive parameter adjustment that works in various acoustic environments (car interior, different noise levels, different driving conditions), thereby achieving both precision and robustness.
Solution Approach 2:
The patent employs feedback by continuously monitoring the output signal's spectral content and using this information to adjust the attenuation parameters in real-time. The system measures the effectiveness of sibilant suppression and dynamically modifies its operation based on the detected sibilant energy levels and spectral distribution, ensuring effective performance across unknown speakers and varying acoustic scenarios through closed-loop control.
Data Source
AI summary
Methods and systems for deessing of speech signals are described. A deesser of a speech processing system includes an analyzer configured to receive a full spectral envelope for each time frame of a speech signal presented to the speech processing system, and to analyze the full spectral envelope to identify frequency content for deessing. The deesser also includes a compressor configured to receive results from the analyzer and to spectrally weight the speech signal as a function of results of the analyzer. The analyzer can be configured to calculate a psychoacoustic measure from the full spectral envelope, and may be further configured to detect sibilant sounds of the speech signal using the psychoacoustic measure. The psychoacoustic measure can include, for example, a measure of sharpness, and the analyzer may be further configured to calculate deesser weights based on the measure of sharpness. An example application includes in-car communications.


