Content-Aware Audio De-Essing With Phoneme-Based Sibilance Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional de-essers are frequency-dependent and lack context awareness, often compressing non-sibilant sounds and resulting in undesirable artifacts, leading to complex signal chains when multiple de-essers are used to handle different sibilant sounds.
Innovation Solution
A pseudo real-time content-aware auditory cleansing method that uses a machine learning-based phoneme classifier to predict the presence of sibilance in an audio sample, followed by cleansing and outputting a cleansed version of the sample, incorporating a look-ahead delay and functional modules for gain control and compression.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If conventional de-essers are used to compress sibilant sounds, then sibilant reduction is achieved, but non-sibilant sounds are also compressed resulting in undesirable artifacts
Solution Approach 1:
The system performs preliminary classification of audio content using a machine learning phoneme classifier before applying compression. The classifier identifies whether sibilant sounds are present in the incoming audio signal, and only then is compression applied. This preliminary action prevents unnecessary compression of non-sibilant sounds, eliminating the generation of compression artifacts while still effectively reducing sibilant harshness when needed.
Solution Approach 2:
The system dynamically changes the compression parameter (gain reduction) based on the detected sibilant content. When sibilants are detected, compression is applied with appropriate gain reduction; when they are not present, compression is bypassed or reduced. This parameter change approach ensures that compression artifacts are only generated when necessary for sibilant control, not continuously on all audio content.
2Adaptability or versatility
If multiple de-essers are used to handle different sibilant sounds, then comprehensive sibilant coverage is achieved, but signal chain complexity increases
Solution Approach 1:
The system employs a single de-esser unit that performs multiple functions through the integration of a machine learning phoneme classifier. This classifier can identify different types of sibilant sounds (such as 's', 'sh', 't' sounds) and route them to appropriate processing paths within the same device. This multi-functional approach provides comprehensive sibilant coverage equivalent to multiple separate de-essers while maintaining a simplified, unified signal chain.
Solution Approach 2:
The system merges the classification and compression functions into a single integrated processing stage. The machine learning classifier and compression engine work together in one unified system rather than requiring separate de-esser units for different sibilant types. This merging reduces the overall signal chain complexity while maintaining the ability to handle various sibilant sounds effectively.
Data Source
AI summary
Techniques for pseudo real-time content-aware auditory cleansing are described herein. A method for pseudo real-time content-aware auditory cleansing may include introducing a delay in the delivery of an audio file by an audio delivery system, collecting a sample of the audio file, predicting the presence of a sibilance and other unwanted sounds in the sample, cleansing the sample of the sibilance and other unwanted sounds, and outputting a cleansed version of the sample. The delay corresponds to a length of the sample. Cleansing the sample may include one or more of a gain reduction, an equalization adjustment, a replacement, and removal of the sibilance and other unwanted sounds.


