Dialogue Component Compression for Clearer Audio Codec Speech

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio codec systems provide subtle dialogue enhancement, and existing solutions are not directly applicable or effective for improving dialogue clarity in audio signals.

Innovation Solution

The proposed solution involves additional processing of estimated dialogue components using compression and equalization, where dialogue enhancement parameters are used to estimate and enhance the dialogue component separately from the audio signal, increasing its signal-to-noise ratio and making it more audible by applying a compressor and user-determined gain.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If dialogue enhancement parameters are applied as linear gains directly to the audio signal, then the dialogue enhancement is subtle, but the processing complexity is low

Engineering Contradiction:
Improvedialogue clarityVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The audio signal is segmented into dialogue components and non-dialogue components using dialogue enhancement parameters. This segmentation allows separate processing of the dialogue portion, enabling application of compression and equalization specifically to enhance dialogue clarity without processing the entire audio signal, thus resolving the contradiction between dialogue clarity and processing complexity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The dialogue components are estimated and separated from the audio signal before applying enhancement processing. This preliminary action of component separation enables subsequent compression and equalization to be applied more effectively to the dialogue portion, improving dialogue clarity while maintaining reasonable processing complexity through targeted processing

Inventive Principle:
Principle #10Preliminary action

2Power

If compression is applied to the estimated dialogue component, then the average power of dialogue increases, but the processing complexity increases

Engineering Contradiction:
Improveaverage power of dialogueVSAvoidprocessing complexity
Core Design Contradiction:
PowerVSDevice complexity

Solution Approach 1:

Compression is applied locally only to the estimated dialogue component rather than the entire audio signal. This local processing approach increases the average power of dialogue components specifically, while limiting the overall processing complexity by avoiding global processing of all audio content

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

Compression parameters are applied to the dialogue component to change its power characteristics. By modifying the compression ratio, threshold, and attack/release parameters specifically for the dialogue portion, the average power of dialogue is increased while the processing complexity remains manageable through parameter optimization

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3956886B1Dialogue enhancement in audio codec
Publication Date: 2024.05.29 DOLBY INTERNATIONAL AB
  • EP3956886B1 patent drawingFigure 1~3
  • EP3956886B1 patent drawingFigure 4~6
  • EP3956886B1 patent drawingFigure 7a~7b

AI summary

Dialogue enhancement of an audio signal, comprising obtaining a set of time-varying parameters configured to estimate a dialogue component present in said audio signal, estimating the dialogue component from the audio signal, applying a compressor only to the estimated dialogue component, to generate a processed dialogue component, applying a user-determined gain to the processed dialogue component, to provide an enhanced dialogue component. The processing of the estimated dialogue may be performed on the decoder side or encoder side. The invention enables an improved dialogue enhancement.