Audio Decoder Dialog Enhancement via Object Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio systems face increased computational complexity when attempting to enhance dialog in object-based audio scenes due to the need to manage a large number of audio objects, leading to potential mixing issues between dialog and other objects.
Innovation Solution
A method and apparatus that modify coefficients for reconstructing audio objects representing dialog within the decoder, using an enhancement parameter to reduce complexity by selectively enhancing dialog objects without full reconstruction of all audio objects, thereby reducing computational load and maintaining the mutual ratio between coefficients.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If dialog enhancement is implemented in an object-based audio system with many audio objects, then dialog enhancement capability is improved, but decoder computational complexity increases
Solution Approach 1:
The patent segments the audio objects into two distinct groups: dialog objects and non-dialog objects. This segmentation allows the decoder to apply different processing strategies to each group, specifically applying complex enhancement operations only to dialog objects while using simpler operations for non-dialog objects, thereby reducing overall computational complexity while maintaining dialog enhancement capability
Solution Approach 2:
The patent applies partial action by selectively performing full reconstruction and enhancement operations only on dialog objects rather than all audio objects. For non-dialog objects, a simplified processing path is used, which performs fewer computational operations while still maintaining acceptable audio quality, thus reducing the total computational load on the decoder
2Device complexity
If the number of audio objects is reduced through object clustering, then complexity and data requirements are reduced, but mixing between dialog and other objects occurs
Solution Approach 1:
The patent applies local quality by providing different processing treatments to different types of audio objects based on their content. Dialog objects receive enhanced processing with full reconstruction and selective enhancement operations, while non-dialog objects receive standard processing. This differentiated approach ensures that dialog objects maintain their separation quality even when other objects are clustered, as the dialog-specific processing path preserves their integrity
Solution Approach 2:
The patent segments the audio object processing into distinct pathways: one for dialog objects and another for non-dialog objects. This segmentation allows dialog objects to be extracted and processed separately from clustered non-dialog objects, preventing unwanted mixing while still allowing the overall number of processed objects to be reduced through clustering of non-dialog content
Data Source
AI summary
This disclosure falls into the field of audio coding, in particular it is related to the field of spatial audio coding, where the audio information is represented by multiple audio objects including at least one dialog object. In particular the disclosure provides a method and apparatus for enhancing dialog in a decoder in an audio system. Furthermore, this disclosure provides a method and apparatus for encoding such audio objects for allowing dialog to be enhanced by the decoder in the audio system.


