Metadata Decoder for Low Delay Spatial Audio Object Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing metadata coding concepts for spatial audio object coding are inefficient, particularly in terms of data compression, and do not support real-time operation or limited system delay.

Innovation Solution

The proposed solution involves an apparatus and method for efficient object metadata coding, which includes a metadata decoder that generates reconstructed metadata signals from processed metadata signals based on a control signal, and an audio channel generator that produces audio channels using the audio object signals and reconstructed metadata signals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If text-based metadata formats (SpatDIF, ASDF) are used to describe audio objects, then the metadata can be structured and transmitted, but the data compression efficiency is poor and system delay increases

Engineering Contradiction:
Improvemetadata compression efficiencyVSAvoidsystem delay
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent transforms metadata from text-based representation to binary format, changing the parameter of data representation. This binary format enables efficient compression algorithms to be applied, reducing the amount of data that needs to be transmitted while maintaining full information content about audio object positions and trajectories.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the mechanical/text-based metadata structure (SpatDIF, ASDF) with a signal-processing approach using processed metadata signals. Instead of transmitting structured text data, the system transmits compressed binary metadata signals that can be efficiently decoded, substituting a data-processing mechanism for a document-structure mechanism.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If AudioBIFS format is used for metadata encoding, then binary format provides better compression, but random access capability and real-time operation are not supported

Engineering Contradiction:
Improvedata compression rateVSAvoidrandom access capability
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent segments the metadata stream into individual processed metadata signals for each audio object, where each signal contains position information over time. This segmentation enables random access to specific object metadata without requiring decoding of the entire stream, while maintaining the compression efficiency of binary format. Each segmented signal can be independently processed and accessed.

Inventive Principle:
Principle #1Segmentation

3Quantity of substance

If existing metadata coding concepts are applied, then audio object positions can be transmitted, but bit rate efficiency for pure azimuth changes is poor

Engineering Contradiction:
Improvebit rate efficiencyVSAvoidposition information accuracy
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent extracts and separates the positional parameters (azimuth, elevation, radius) into distinct processed metadata signals. By taking out the azimuth information as a separate signal component, the system can apply targeted compression strategies specific to azimuth changes, improving bit rate efficiency for this particular parameter while maintaining position information accuracy through selective processing.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentEP4542544A1Apparatus and method for low delay object metadata coding
Publication Date: 2025.04.23 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • EP4542544A1 patent drawingFigure 1
  • EP4542544A1 patent drawingFigure 2
  • EP4542544A1 patent drawingFigure 3

AI summary

An apparatus (100) for generating one or more audio channels is provided. The apparatus comprises a metadata decoder (110) for generating one or more reconstructed metadata signals (x1',...,xN') from one or more processed metadata signals (z1,...,zN) depending on a control signal (b), wherein each of the one or more reconstructed metadata signals (x1',...,xN') indicates information associated with an audio object signal of one or more audio object signals, wherein the metadata decoder (110) is configured to generate the one or more reconstructed metadata signals (x1',...,xN') by determining a plurality of reconstructed metadata samples (x1'(n),...,xN'(n)) for each of the one or more reconstructed metadata signals (x1',...,xN'). Moreover, the apparatus comprises an audio channel generator (120) for generating the one or more audio channels depending on the one or more audio object signals and depending on the one or more reconstructed metadata signals (x1',...,xN'). The metadata decoder (110) is configured to receive a plurality of processed metadata samples (z1(n),...,zN(n)) of each of the one or more processed metadata signals (z1,...,zN). Moreover, the metadata decoder (110) is configured to receive the control signal (b). Furthermore, the metadata decoder (110) is configured to determine each reconstructed metadata sample (xi'(n)) of the plurality of reconstructed metadata samples (xi'(1).... xi'(n-1), xi'(n)) of each reconstructed metadata signal (xi') of the one or more reconstructed metadata signals (x1',...,xN'), so that, when the control signal (b) indicates a first state (b(n)=0), said reconstructed metadata sample (xi'(n)) is a sum of one of the processed metadata samples (zi(n)) of one of the one or more processed metadata signals (zi) and of another already generated reconstructed metadata sample (xi'(n-1)) of said reconstructed metadata signal (xi'), and so that, when the control signal indicates a second state (b(n)=1) being different from the first state, said reconstructed metadata sample (xi'(n)) is said one (zi(n)) of the processed metadata samples (zi(1),...,zi(n)) of said one (zi) of the one or more processed metadata signals (z1,...,zN). Moreover, an apparatus (250) for generating encoded audio information is provided.