Audio Object Reconstruction via Time-Variable Side Information

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for encoding and decoding audio objects in audio scenes, particularly in multichannel configurations like 5.1, face challenges in reconstructing audio objects accurately, leading to audible artifacts due to insufficient downmix reconstruction and high computational complexity.

Innovation Solution

An encoder and decoder method that calculates time-variable side information and downmix signals, including transition data for interpolation, to facilitate efficient and improved reconstruction of audio objects, reducing computational complexity and allowing for resampling without affecting playback quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If audio objects are combined into a multichannel downmix, then compatibility with legacy decoders is improved, but reconstruction fidelity of audio objects deteriorates

Engineering Contradiction:
Improvecompatibility with legacy decodersVSAvoidreconstruction fidelity
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent segments the audio signal into two independent parts: a multichannel downmix signal for legacy compatibility and object-based side information for high-quality reconstruction. The downmix is created by combining audio objects according to a downmix matrix, while the side information preserves individual object characteristics, allowing both legacy compatibility and high fidelity to coexist.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces object-based side information as an intermediary element that bridges the gap between the downmix signal and the original audio objects. This side information contains parameters that enable accurate reconstruction of individual audio objects from the downmix, effectively mediating between the need for compatibility and the desire for high reconstruction quality.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of manufacture

If traditional side information formats are used, then encoding simplicity is maintained, but computational complexity of reconstruction increases

Engineering Contradiction:
Improveencoding simplicityVSAvoidcomputational complexity of reconstruction
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The patent changes the parameters of the side information format to include time-variable parameters that describe the spatial and spectral characteristics of audio objects. By using parameter-based representation instead of full signal data, the patent reduces the computational complexity required for reconstruction while maintaining encoding simplicity.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If audio objects with similar horizontal positions are combined into the same channel, then downmix compactness is improved, but reconstruction accuracy deteriorates

Engineering Contradiction:
Improvedownmix compactnessVSAvoidreconstruction accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent applies local quality by preserving distinct vertical position information for audio objects even when they share the same horizontal position. The side information includes parameters that capture local spatial characteristics (vertical position, depth) of each audio object, allowing accurate reconstruction while maintaining compact downmix representation.

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP3127109B1Efficient coding of audio scenes comprising audio objects
Publication Date: 2018.03.14 DOLBY INTERNATIONAL AB
  • EP3127109B1 patent drawingFigure 1~2
  • EP3127109B1 patent drawingFigure 3~4
  • EP3127109B1 patent drawingFigure 5

AI summary

There is provided encoding and decoding methods for encoding and decoding of object based audio. An exemplary decoding method described is for reconstructing audio objects based on a data stream, wherein the data stream corresponds to a plurality of time frames, wherein the data stream comprises a plurality of side information instances, wherein the data stream further comprises, for each side information instance, transition data including two independently assignable portions which in combination define a point in time to begin a transition from a current reconstruction setting to a desired reconstruction setting specified by the side information instance, and a point in time to complete the transition.