Content-Aware De-Essing Using Phoneme Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional de-essers often compress non-sibilant sounds, leading to undesirable artifacts and require multiple units to handle different sibilant sounds, resulting in a complex signal chain.
Innovation Solution
A pseudo real-time content-aware auditory cleansing method that uses a machine learning-based phoneme classifier to predict sibilance in audio samples, allowing for targeted gain control and compression to mitigate sibilant sounds without affecting other audio components.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If conventional de-essers compress only sibilant frequency bands, then sibilant sounds are reduced, but non-sibilant sounds are also compressed causing undesirable artifacts
Solution Approach 1:
The patent uses a feedback-based approach where the audio signal is analyzed in real-time to detect sibilant sounds, and compression is applied dynamically based on this detection. The system continuously monitors the input signal and adjusts compression parameters accordingly, ensuring that only actual sibilant sounds are compressed while preserving other audio content.
Solution Approach 2:
The patent dynamically changes compression parameters (threshold, ratio, attack, release) based on the detected sibilant characteristics. By adjusting these parameters in real-time according to the specific sibilant event being detected, the system achieves precise control over which sounds are compressed and which are preserved, eliminating the need for fixed frequency-based compression.
2Adaptability or versatility
If multiple de-essers are used to handle different sibilant sounds, then coverage of various sibilant types improves, but signal chain complexity increases
Solution Approach 1:
The patent implements a universal de-esser that can handle all types of sibilant sounds through a single device. The system uses machine learning models trained to recognize multiple sibilant patterns (s, sh, t, ch, etc.), allowing one de-esser to perform the function of multiple specialized de-essers would otherwise be needed.
Solution Approach 2:
The patent employs an automated detection and classification system that identifies and processes different sibilant types without requiring manual configuration or multiple separate devices. The system self-adjusts its parameters and processing approach based on the detected sibilant characteristics, eliminating the need for complex manual signal chain arrangements.
Data Source
Figure 1~2
Figure 3
Figure 4A
AI summary
Techniques for pseudo real-time content-aware auditory cleansing are described herein. A method for pseudo real-time content-aware auditory cleansing may include introducing a delay in the delivery of an audio file by an audio delivery system, collecting a sample of the audio file, predicting the presence of a sibilance and other unwanted sounds in the sample, cleansing the sample of the sibilance and other unwanted sounds, and outputting a cleansed version of the sample. The delay corresponds to a length of the sample. Cleansing the sample may include one or more of a gain reduction, an equalization adjustment, a replacement, and removal of the sibilance and other unwanted sounds.