Audio Decoder Dialog Enhancement via Object Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio systems face increased computational complexity when attempting to enhance dialog in object-based audio scenes due to the need to manage a large number of audio objects, leading to potential mixing issues between dialog and other objects.

Innovation Solution

A method and apparatus that modify coefficients for reconstructing audio objects representing dialog within the decoder, using an enhancement parameter to reduce complexity by selectively enhancing dialog objects without full reconstruction of all audio objects, thereby reducing computational load and maintaining the mutual ratio between coefficients.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If dialog enhancement is implemented in an object-based audio system with many audio objects, then dialog enhancement capability is improved, but decoder computational complexity increases

Engineering Contradiction:
Improvedialog enhancement capabilityVSAvoiddecoder computational complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the audio objects into two distinct groups: dialog objects and non-dialog objects. This segmentation allows the decoder to apply different processing strategies to each group, specifically applying complex enhancement operations only to dialog objects while using simpler operations for non-dialog objects, thereby reducing overall computational complexity while maintaining dialog enhancement capability

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by selectively performing full reconstruction and enhancement operations only on dialog objects rather than all audio objects. For non-dialog objects, a simplified processing path is used, which performs fewer computational operations while still maintaining acceptable audio quality, thus reducing the total computational load on the decoder

Inventive Principle:
Principle #16Partial or excessive action

2Device complexity

If the number of audio objects is reduced through object clustering, then complexity and data requirements are reduced, but mixing between dialog and other objects occurs

Engineering Contradiction:
Improveaudio scene complexityVSAvoiddialog separation quality
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent applies local quality by providing different processing treatments to different types of audio objects based on their content. Dialog objects receive enhanced processing with full reconstruction and selective enhancement operations, while non-dialog objects receive standard processing. This differentiated approach ensures that dialog objects maintain their separation quality even when other objects are clustered, as the dialog-specific processing path preserves their integrity

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the audio object processing into distinct pathways: one for dialog objects and another for non-dialog objects. This segmentation allows dialog objects to be extracted and processed separately from clustered non-dialog objects, preventing unwanted mixing while still allowing the overall number of processed objects to be reduced through clustering of non-dialog content

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10163446B2Audio encoder and decoder
Publication Date: 2018.12.25 DOLBY INTERNATIONAL AB
  • US10163446B2 patent drawing
  • US10163446B2 patent drawing
  • US10163446B2 patent drawing

AI summary

This disclosure falls into the field of audio coding, in particular it is related to the field of spatial audio coding, where the audio information is represented by multiple audio objects including at least one dialog object. In particular the disclosure provides a method and apparatus for enhancing dialog in a decoder in an audio system. Furthermore, this disclosure provides a method and apparatus for encoding such audio objects for allowing dialog to be enhanced by the decoder in the audio system.