Immersive Audio Metadata Redundancy Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing technologies for encoding and decoding of immersive audio face challenges in reducing the data rate of metadata, particularly in object-based audio systems, which leads to increased complexity and data rate due to the provision of multiple sets of metadata.
Innovation Solution
The method involves identifying redundant data elements or structures within different sets of metadata and encoding them by reference to external metadata, reducing redundant data transmission by using flags or differential encoding to indicate dependencies between metadata sets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple sets of metadata are provided for different decoder types, then flexibility and compatibility are improved, but data rate increases substantially
Solution Approach 1:
The patent merges multiple metadata sets (side information, object audio metadata, and SimpleRendererInfo) into a unified metadata structure. This consolidation eliminates redundant data elements while preserving all necessary information for different decoder types, thereby reducing the overall data rate while maintaining decoder compatibility.
Solution Approach 2:
The unified metadata structure is designed to serve multiple functions simultaneously: it provides side information for parameterized upmix, object audio metadata for object-based rendering, and SimpleRendererInfo for backward-compatible downmix. This multi-functional design eliminates the need for separate metadata sets for different decoder types.
2Measurement precision
If object-based audio is used instead of channel-based audio, then spatial playback quality is improved, but data rate and complexity increase
Solution Approach 1:
The patent segments the audio scene into independent audio objects, each with its own metadata describing spatial properties. This segmentation enables precise spatial playback control while allowing the encoder to process and encode each object separately, thereby managing complexity through modular processing.
Solution Approach 2:
The patent transforms audio objects into a unified metadata representation that captures essential spatial parameters (position, gain, width) in a compact format. This parameter-based representation maintains high spatial playback accuracy while reducing the complexity of encoding and transmitting object-based audio information.
Data Source
AI summary
The present document relates to the field of encoding and decoding of audio. In particular, the present document relates to encoding and decoding of an audio scene comprising audio objects. A method (400) for encoding metadata relating to a plurality of audio objects (106a) of an audio scene (102) is described. The metadata comprises a first set (114, 314) of metadata and a second set (104) of metadata. The first and second sets (104, 114, 314) of metadata comprise one or more data elements which are indicative of a property of an audio object (106a) from the plurality of audio objects (106a) and/or of a downmix signal (112) derived from the plurality of audio objects (106a). The method (400) comprises identifying (401) a redundant data element which is common to the first and second sets (104, 114, 314) of metadata. Furthermore, the method comprises encoding (402) the redundant data element of the first set (114, 314) of metadata by referring to a redundant data element of a set (104) of metadata external for the first set (114, 314) of metadata.


