Adaptive De-Esser for In-Car Speech Sibilance Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In-Car Communication (ICC) systems face challenges in handling sibilant sounds, which are speaker-dependent and become dominant due to noise suppression, leading to an overemphasis of higher frequency bands, making them annoying for listeners.

Innovation Solution

A dynamic deesser method that employs spectral envelope and phoneme-dependent attenuation, using psychoacoustic measures like sharpness to adaptively suppress sibilant sounds, considering various speakers and acoustic scenarios, and integrating this method into ICC systems for improved speech signal processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If noise suppression is applied to improve speech clarity in ICC systems, then speech intelligibility is improved, but sibilant sounds become over-emphasized and more annoying

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidsibilant sound annoyance
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent applies local quality by implementing frequency-selective attenuation that targets only the specific frequency ranges where sibilant sounds occur (typically 2-8 kHz) while preserving other frequency components. The deesser dynamically adjusts attenuation levels based on the detected sibilant energy in different frequency bands, thereby locally modifying the spectral characteristics to reduce annoyance without affecting overall speech intelligibility.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent employs parameter changes by dynamically adjusting the attenuation parameter based on the detected sibilant characteristics. The system continuously monitors the spectral content and adapts the attenuation level, frequency range, and time constant parameters in real-time to match the speaker's specific sibilant profile and the acoustic environment, thereby resolving the contradiction between noise suppression benefits and sibilant annoyance.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If a static equalizer setting is used to simplify system configuration, then system complexity is reduced, but the system cannot adapt to different speakers' sibilant characteristics

Engineering Contradiction:
Improvesystem configurationVSAvoidspeaker adaptation
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent applies dynamics by transforming the static equalizer setting into a dynamic adaptive system. The deesser continuously analyzes the input speech signal to detect sibilant characteristics and automatically adjusts the attenuation parameters in real-time. This dynamic adaptation allows the system to accommodate different speakers' sibilant profiles without requiring manual reconfiguration, thereby resolving the contradiction between system simplicity and adaptability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent employs self-service by enabling the system to automatically detect and adapt to each speaker's unique sibilant characteristics without external intervention. The deesser performs self-adjustment by monitoring the spectral content and modifying its attenuation parameters based on the detected sibilant energy, thereby providing speaker-specific optimization automatically and eliminating the need for manual tuning.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If a deesser method optimized for known speakers is used to improve sibilant suppression, then suppression effectiveness is improved, but the method cannot work robustly for unknown speakers and various acoustic scenarios

Engineering Contradiction:
Improvesibilant suppression effectivenessVSAvoidrobustness across speakers and scenarios
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies universality by designing a deesser algorithm that functions effectively across diverse speakers and acoustic scenarios without requiring speaker-specific optimization. The system uses universal sibilant detection criteria based on spectral characteristics that are common to all speakers, combined with adaptive parameter adjustment that works in various acoustic environments (car interior, different noise levels, different driving conditions), thereby achieving both precision and robustness.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent employs feedback by continuously monitoring the output signal's spectral content and using this information to adjust the attenuation parameters in real-time. The system measures the effectiveness of sibilant suppression and dynamically modifies its operation based on the detected sibilant energy levels and spectral distribution, ensuring effective performance across unknown speakers and varying acoustic scenarios through closed-loop control.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20240062770A1Enhanced de-esser for in-car communications systems
Publication Date: 2024.02.22 CERENCE OPERATING CO
  • US20240062770A1 patent drawing
  • US20240062770A1 patent drawing
  • US20240062770A1 patent drawing

AI summary

Methods and systems for deessing of speech signals are described. A deesser of a speech processing system includes an analyzer configured to receive a full spectral envelope for each time frame of a speech signal presented to the speech processing system, and to analyze the full spectral envelope to identify frequency content for deessing. The deesser also includes a compressor configured to receive results from the analyzer and to spectrally weight the speech signal as a function of results of the analyzer. The analyzer can be configured to calculate a psychoacoustic measure from the full spectral envelope, and may be further configured to detect sibilant sounds of the speech signal using the psychoacoustic measure. The psychoacoustic measure can include, for example, a measure of sharpness, and the analyzer may be further configured to calculate deesser weights based on the measure of sharpness. An example application includes in-car communications.