Content-Aware Audio De-Essing With Phoneme-Based Sibilance Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional de-essers are frequency-dependent and lack context awareness, often compressing non-sibilant sounds and resulting in undesirable artifacts, leading to complex signal chains when multiple de-essers are used to handle different sibilant sounds.

Innovation Solution

A pseudo real-time content-aware auditory cleansing method that uses a machine learning-based phoneme classifier to predict the presence of sibilance in an audio sample, followed by cleansing and outputting a cleansed version of the sample, incorporating a look-ahead delay and functional modules for gain control and compression.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If conventional de-essers are used to compress sibilant sounds, then sibilant reduction is achieved, but non-sibilant sounds are also compressed resulting in undesirable artifacts

Engineering Contradiction:
Improvesibilant harshnessVSAvoidcompression artifacts
Core Design Contradiction:
Object-affected harmful factorsVSObject-generated harmful factors

Solution Approach 1:

The system performs preliminary classification of audio content using a machine learning phoneme classifier before applying compression. The classifier identifies whether sibilant sounds are present in the incoming audio signal, and only then is compression applied. This preliminary action prevents unnecessary compression of non-sibilant sounds, eliminating the generation of compression artifacts while still effectively reducing sibilant harshness when needed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically changes the compression parameter (gain reduction) based on the detected sibilant content. When sibilants are detected, compression is applied with appropriate gain reduction; when they are not present, compression is bypassed or reduced. This parameter change approach ensures that compression artifacts are only generated when necessary for sibilant control, not continuously on all audio content.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If multiple de-essers are used to handle different sibilant sounds, then comprehensive sibilant coverage is achieved, but signal chain complexity increases

Engineering Contradiction:
Improvesibilant sound coverageVSAvoidsignal chain length
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system employs a single de-esser unit that performs multiple functions through the integration of a machine learning phoneme classifier. This classifier can identify different types of sibilant sounds (such as 's', 'sh', 't' sounds) and route them to appropriate processing paths within the same device. This multi-functional approach provides comprehensive sibilant coverage equivalent to multiple separate de-essers while maintaining a simplified, unified signal chain.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system merges the classification and compression functions into a single integrated processing stage. The machine learning classifier and compression engine work together in one unified system rather than requiring separate de-esser units for different sibilant types. This merging reduces the overall signal chain complexity while maintaining the ability to handle various sibilant sounds effectively.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250118324A1Pseudo Real-time Content-Aware Auditory Cleansing
Publication Date: 2025.04.10 ANTARES AUDIO STRATEGIES LLC
  • US20250118324A1 patent drawing
  • US20250118324A1 patent drawing
  • US20250118324A1 patent drawing

AI summary

Techniques for pseudo real-time content-aware auditory cleansing are described herein. A method for pseudo real-time content-aware auditory cleansing may include introducing a delay in the delivery of an audio file by an audio delivery system, collecting a sample of the audio file, predicting the presence of a sibilance and other unwanted sounds in the sample, cleansing the sample of the sibilance and other unwanted sounds, and outputting a cleansed version of the sample. The delay corresponds to a length of the sample. Cleansing the sample may include one or more of a gain reduction, an equalization adjustment, a replacement, and removal of the sibilance and other unwanted sounds.