Metadata Decoder for Low Delay Spatial Audio Object Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing metadata coding concepts for spatial audio object coding are inefficient, particularly in terms of data compression, and do not support real-time operation or limited system delay.
Innovation Solution
The proposed solution involves an apparatus and method for efficient object metadata coding, which includes a metadata decoder that generates reconstructed metadata signals from processed metadata signals based on a control signal, and an audio channel generator that produces audio channels using the audio object signals and reconstructed metadata signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If text-based metadata formats (SpatDIF, ASDF) are used to describe audio objects, then the metadata can be structured and transmitted, but the data compression efficiency is poor and system delay increases
Solution Approach 1:
The patent transforms metadata from text-based representation to binary format, changing the parameter of data representation. This binary format enables efficient compression algorithms to be applied, reducing the amount of data that needs to be transmitted while maintaining full information content about audio object positions and trajectories.
Solution Approach 2:
The patent replaces the mechanical/text-based metadata structure (SpatDIF, ASDF) with a signal-processing approach using processed metadata signals. Instead of transmitting structured text data, the system transmits compressed binary metadata signals that can be efficiently decoded, substituting a data-processing mechanism for a document-structure mechanism.
2Loss of information
If AudioBIFS format is used for metadata encoding, then binary format provides better compression, but random access capability and real-time operation are not supported
Solution Approach 1:
The patent segments the metadata stream into individual processed metadata signals for each audio object, where each signal contains position information over time. This segmentation enables random access to specific object metadata without requiring decoding of the entire stream, while maintaining the compression efficiency of binary format. Each segmented signal can be independently processed and accessed.
3Quantity of substance
If existing metadata coding concepts are applied, then audio object positions can be transmitted, but bit rate efficiency for pure azimuth changes is poor
Solution Approach 1:
The patent extracts and separates the positional parameters (azimuth, elevation, radius) into distinct processed metadata signals. By taking out the azimuth information as a separate signal component, the system can apply targeted compression strategies specific to azimuth changes, improving bit rate efficiency for this particular parameter while maintaining position information accuracy through selective processing.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An apparatus (100) for generating one or more audio channels is provided. The apparatus comprises a metadata decoder (110) for generating one or more reconstructed metadata signals (x1',...,xN') from one or more processed metadata signals (z1,...,zN) depending on a control signal (b), wherein each of the one or more reconstructed metadata signals (x1',...,xN') indicates information associated with an audio object signal of one or more audio object signals, wherein the metadata decoder (110) is configured to generate the one or more reconstructed metadata signals (x1',...,xN') by determining a plurality of reconstructed metadata samples (x1'(n),...,xN'(n)) for each of the one or more reconstructed metadata signals (x1',...,xN'). Moreover, the apparatus comprises an audio channel generator (120) for generating the one or more audio channels depending on the one or more audio object signals and depending on the one or more reconstructed metadata signals (x1',...,xN'). The metadata decoder (110) is configured to receive a plurality of processed metadata samples (z1(n),...,zN(n)) of each of the one or more processed metadata signals (z1,...,zN). Moreover, the metadata decoder (110) is configured to receive the control signal (b). Furthermore, the metadata decoder (110) is configured to determine each reconstructed metadata sample (xi'(n)) of the plurality of reconstructed metadata samples (xi'(1).... xi'(n-1), xi'(n)) of each reconstructed metadata signal (xi') of the one or more reconstructed metadata signals (x1',...,xN'), so that, when the control signal (b) indicates a first state (b(n)=0), said reconstructed metadata sample (xi'(n)) is a sum of one of the processed metadata samples (zi(n)) of one of the one or more processed metadata signals (zi) and of another already generated reconstructed metadata sample (xi'(n-1)) of said reconstructed metadata signal (xi'), and so that, when the control signal indicates a second state (b(n)=1) being different from the first state, said reconstructed metadata sample (xi'(n)) is said one (zi(n)) of the processed metadata samples (zi(1),...,zi(n)) of said one (zi) of the one or more processed metadata signals (z1,...,zN). Moreover, an apparatus (250) for generating encoded audio information is provided.