Selective Audio Dereverberation for Speech Intelligibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio technologies face challenges in effectively suppressing reverberation in mixed audio content, such as user-generated content like podcasts, which can lead to reduced audio quality and compromised speech intelligibility when applying dereverberation techniques.

Innovation Solution

The method involves classifying input audio signals into types like speech, music, or speech over music, and selectively performing dereverberation based on the classification, using techniques such as spatial component separation, neural networks, and reverberation time analysis to determine the need for dereverberation, thereby improving speech intelligibility while preserving audio quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If dereverberation is applied to all audio content, then speech intelligibility is improved, but audio quality is degraded for non-speech content

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidaudio quality degradation
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The system applies different processing quality to different parts of the audio content by classifying media types and selectively applying dereverberation only to speech content, while preserving music and speech-over-music content without dereverberation processing

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The audio content is segmented into different media types (speech, music, speech over music) through classification, allowing differential processing where dereverberation is applied only to the speech segment while leaving other segments unchanged

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If dereverberation is applied to mixed media content, then speech intelligibility is improved, but the complexity of the processing system increases

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The processing system is segmented into distinct functional modules: a media type classifier that categorizes audio content, and a selective dereverberation processor that applies processing only to identified speech content. This modular architecture manages complexity by separating classification logic from processing logic

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If dereverberation is applied to user-generated content, then speech intelligibility is improved, but loss of original audio characteristics occurs

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidaudio quality
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system preserves original audio characteristics by applying dereverberation selectively only to speech portions identified through media type classification, while leaving music and mixed content portions untouched, thus maintaining their original audio quality and characteristics

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP4305620B1Dereverberation based on media type
Publication Date: 2025.01.15 DOLBY LABORATORIES LICENSING CORP
  • EP4305620B1 patent drawingFigure 1A
  • EP4305620B1 patent drawingFigure 1B
  • EP4305620B1 patent drawingFigure 2

AI summary

A method for reverberation suppression may involve receiving an input audio signal. The method may involve classifying a media type of the input audio signal as one of a group comprising at least: 1) speech; 2) music; or 3) speech over music. The method may involve determining whether to perform dereverberation on the input audio signal based at least on a determination that the media type of the input audio signal has been classified as speech. The method may involve generating an output audio signal by performing dereverberation on the input audio signal in response to determining that dereverberation is to be performed on the input audio signal.