Automatic De-Esser Using Relative Sibilance Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional de-essers require manual parameter setting and are ineffective in reducing sibilance across varying signal levels and singer genders, often failing to address soft sibilance and being inconsistent with changes in overall signal levels.
Innovation Solution
An automatic de-esser that processes audio signals by transforming them into the frequency domain, using a multi-band compressor that applies attenuation based on energy level comparisons and zero-crossing rates, independent of absolute signal levels, to effectively reduce sibilance in both loud and soft parts of a performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If conventional de-essers use manual parameter setting with absolute thresholds, then they can be configured for specific sessions, but they fail to adapt to varying signal levels and singer genders, requiring repeated manual intervention
Solution Approach 1:
The de-esser uses dynamic threshold adjustment based on signal level detection. The threshold is not fixed but adapts automatically to the current signal level, allowing the device to handle varying signal levels and different singer genders without manual intervention. The system dynamically calculates the threshold as a function of the detected signal level, making the de-essing process adaptive rather than static.
Solution Approach 2:
The system implements feedback by continuously monitoring the signal level and adjusting the threshold accordingly. The output of the signal level detection feeds back to the threshold calculation mechanism, creating a closed-loop system that automatically adapts to changing conditions in real-time, eliminating the need for manual parameter changes between sessions.
2Reliability
If conventional de-essers use fixed absolute thresholds, then the device structure remains simple, but the de-essing effectiveness varies with overall signal level changes
Solution Approach 1:
The system changes the parameter from a fixed absolute threshold to a dynamic threshold that is a function of signal level. By expressing the threshold as a relative parameter rather than an absolute value, the de-esser maintains consistent performance across different signal levels. The threshold is calculated as a percentage or ratio of the detected signal level, ensuring that de-essing effectiveness remains reliable regardless of overall volume changes.
3Reliability
If conventional de-essers are configured for loud sibilance, then they reduce prominent sibilance, but they either miss soft sibilance or require manual threshold adjustment over time
Solution Approach 1:
The system performs self-service by automatically detecting and adapting to the presence of sibilance at any level. The signal level detection mechanism enables the de-esser to serve itself by automatically adjusting its parameters based on the actual input signal characteristics, eliminating the need for operators to manually tweak thresholds to catch soft sibilance.
4Adaptability or versatility
If de-essing is applied independently of signal level, then soft sibilance can be detected, but the system must process a wider dynamic range of signals
Solution Approach 1:
The system uses dynamic threshold calculation where the threshold is expressed as a function of signal level rather than a fixed value. This dynamic approach allows the de-esser to process both loud and soft sibilance effectively. The threshold automatically scales with the signal level, maintaining appropriate de-essing pressure across the full dynamic range without requiring complex multi-threshold systems.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Methods, systems, and computer program products of automatic de-essing are disclosed. An automatic de-esser can be used without manually setting parameters and can perform reliable sibilance detection and reduction regardless of absolute signal level, singer gender and other extraneous factors. An audio processing device divides input audio signals into buffers each containing a number of samples, the buffers overlapping one another. The audio processing device transforms each buffer from the time domain into the frequency domain and implements de-essing as a multi‑band compressor that only acts on a designated sibilance band. The audio processing device determines an amount of attenuation in the sibilance band based on comparison of energy level in sibilance band of a buffer to broadband energy level in a previous buffer. The amount of attenuation is also determined based on a zero‑crossing rate, as well as a slope and onset of a compression curve.