Adaptive In-Car De-Esser for Sibilant Speech Suppression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In-Car Communication (ICC) systems face challenges in handling sibilant sounds, which are speaker-dependent and become dominant due to noise suppression, leading to over-emphasis of higher frequency bands, and existing deesser methods fail to adapt to various speakers and acoustic scenarios effectively.
Innovation Solution
A dynamic deesser method that employs spectral envelope and phoneme-dependent attenuation, using psychoacoustic measures like sharpness to adaptively control sibilant sound suppression, optimizing the process for robust performance across different speakers and acoustic scenarios, including idle car, town traffic, and highway conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If noise suppression is applied to improve speech clarity in ICC systems, then speech clarity is improved, but sibilant sounds become over-emphasized and more dominant
Solution Approach 1:
The deesser applies localized frequency-specific attenuation targeting only the sibilant frequency range (typically 4-12 kHz) while preserving other frequency bands. This is achieved through spectral analysis that identifies sibilant energy distribution and applies targeted gain reduction only where needed, maintaining speech clarity while reducing annoying sibilant dominance.
Solution Approach 2:
The system dynamically adjusts the attenuation parameters based on the measured sibilant energy levels in the signal. The deesser continuously monitors the spectral content and adapts the attenuation amount in real-time, changing the parameter settings according to the speaker characteristics and acoustic scenario to optimally reduce sibilants without affecting overall speech quality.
2Device complexity
If static equalizer settings are used to simplify system configuration, then system complexity is reduced, but the system cannot adapt to different speakers and acoustic scenarios
Solution Approach 1:
The deesser implements dynamic adaptation by continuously analyzing the input signal's spectral envelope and adjusting attenuation parameters in real-time based on detected sibilant characteristics. This allows the system to automatically adapt to different speakers, speaking styles, and acoustic scenarios without requiring manual reconfiguration, while maintaining relatively simple system architecture through algorithmic adaptability.
Solution Approach 2:
The system performs self-adjustment by automatically detecting sibilant sounds and configuring appropriate attenuation levels without external intervention. The deesser monitors its own input signal, identifies when sibilant reduction is needed, and applies appropriate processing parameters autonomously, eliminating the need for manual tuning across different usage scenarios.
3Measurement precision
If deesser methods are optimized for known speakers and controlled scenarios, then deessing performance is improved, but the method fails to work robustly for unknown speakers and varied acoustic scenarios
Solution Approach 1:
The deesser employs dynamic parameter adjustment that adapts to any speaker and acoustic scenario in real-time through spectral analysis. Rather than relying on pre-configured settings for specific speakers, the system continuously analyzes the spectral envelope, identifies sibilant characteristics, and adjusts attenuation parameters dynamically, ensuring robust performance across unknown speakers and varied scenarios like idle car, town traffic, and highway conditions.
Data Source
AI summary
Methods and systems for deessing of speech signals are described. A deesser of a speech processing system includes an analyzer configured to receive a full spectral envelope for each time frame of a speech signal presented to the speech processing system, and to analyze the full spectral envelope to identify frequency content for deessing. The deesser also includes a compressor configured to receive results from the analyzer and to spectrally weight the speech signal as a function of results of the analyzer. The analyzer can be configured to calculate a psychoacoustic measure from the full spectral envelope, and may be further configured to detect sibilant sounds of the speech signal using the psychoacoustic measure. The psychoacoustic measure can include, for example, a measure of sharpness, and the analyzer may be further configured to calculate deesser weights based on the measure of sharpness. An example application includes in-car communications.


