Immersive Audio Metadata Redundancy Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing technologies for encoding and decoding of immersive audio face challenges in reducing the data rate of metadata, particularly in object-based audio systems, which leads to increased complexity and data rate due to the provision of multiple sets of metadata.

Innovation Solution

The method involves identifying redundant data elements or structures within different sets of metadata and encoding them by reference to external metadata, reducing redundant data transmission by using flags or differential encoding to indicate dependencies between metadata sets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple sets of metadata are provided for different decoder types, then flexibility and compatibility are improved, but data rate increases substantially

Engineering Contradiction:
Improvedecoder compatibilityVSAvoiddata rate
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent merges multiple metadata sets (side information, object audio metadata, and SimpleRendererInfo) into a unified metadata structure. This consolidation eliminates redundant data elements while preserving all necessary information for different decoder types, thereby reducing the overall data rate while maintaining decoder compatibility.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The unified metadata structure is designed to serve multiple functions simultaneously: it provides side information for parameterized upmix, object audio metadata for object-based rendering, and SimpleRendererInfo for backward-compatible downmix. This multi-functional design eliminates the need for separate metadata sets for different decoder types.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If object-based audio is used instead of channel-based audio, then spatial playback quality is improved, but data rate and complexity increase

Engineering Contradiction:
Improvespatial playback accuracyVSAvoidencoding complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the audio scene into independent audio objects, each with its own metadata describing spatial properties. This segmentation enables precise spatial playback control while allowing the encoder to process and encode each object separately, thereby managing complexity through modular processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms audio objects into a unified metadata representation that captures essential spatial parameters (position, gain, width) in a compact format. This parameter-based representation maintains high spatial playback accuracy while reducing the complexity of encoding and transmitting object-based audio information.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS9955278B2Exploiting metadata redundancy in immersive audio metadata
Publication Date: 2018.04.24 DOLBY INTERNATIONAL AB
  • US9955278B2 patent drawing
  • US9955278B2 patent drawing
  • US9955278B2 patent drawing

AI summary

The present document relates to the field of encoding and decoding of audio. In particular, the present document relates to encoding and decoding of an audio scene comprising audio objects. A method (400) for encoding metadata relating to a plurality of audio objects (106a) of an audio scene (102) is described. The metadata comprises a first set (114, 314) of metadata and a second set (104) of metadata. The first and second sets (104, 114, 314) of metadata comprise one or more data elements which are indicative of a property of an audio object (106a) from the plurality of audio objects (106a) and/or of a downmix signal (112) derived from the plurality of audio objects (106a). The method (400) comprises identifying (401) a redundant data element which is common to the first and second sets (104, 114, 314) of metadata. Furthermore, the method comprises encoding (402) the redundant data element of the first set (114, 314) of metadata by referring to a redundant data element of a set (104) of metadata external for the first set (114, 314) of metadata.