Audio Object Reconstruction via Time-Variable Side Information
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for encoding and decoding audio objects in audio scenes, particularly in multichannel configurations like 5.1, face challenges in reconstructing audio objects accurately, leading to audible artifacts due to insufficient downmix reconstruction and high computational complexity.
Innovation Solution
An encoder and decoder method that calculates time-variable side information and downmix signals, including transition data for interpolation, to facilitate efficient and improved reconstruction of audio objects, reducing computational complexity and allowing for resampling without affecting playback quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If audio objects are combined into a multichannel downmix, then compatibility with legacy decoders is improved, but reconstruction fidelity of audio objects deteriorates
Solution Approach 1:
The patent segments the audio signal into two independent parts: a multichannel downmix signal for legacy compatibility and object-based side information for high-quality reconstruction. The downmix is created by combining audio objects according to a downmix matrix, while the side information preserves individual object characteristics, allowing both legacy compatibility and high fidelity to coexist.
Solution Approach 2:
The patent introduces object-based side information as an intermediary element that bridges the gap between the downmix signal and the original audio objects. This side information contains parameters that enable accurate reconstruction of individual audio objects from the downmix, effectively mediating between the need for compatibility and the desire for high reconstruction quality.
2Ease of manufacture
If traditional side information formats are used, then encoding simplicity is maintained, but computational complexity of reconstruction increases
Solution Approach 1:
The patent changes the parameters of the side information format to include time-variable parameters that describe the spatial and spectral characteristics of audio objects. By using parameter-based representation instead of full signal data, the patent reduces the computational complexity required for reconstruction while maintaining encoding simplicity.
3Quantity of substance
If audio objects with similar horizontal positions are combined into the same channel, then downmix compactness is improved, but reconstruction accuracy deteriorates
Solution Approach 1:
The patent applies local quality by preserving distinct vertical position information for audio objects even when they share the same horizontal position. The side information includes parameters that capture local spatial characteristics (vertical position, depth) of each audio object, allowing accurate reconstruction while maintaining compact downmix representation.
Data Source
Figure 1~2
Figure 3~4
Figure 5
AI summary
There is provided encoding and decoding methods for encoding and decoding of object based audio. An exemplary decoding method described is for reconstructing audio objects based on a data stream, wherein the data stream corresponds to a plurality of time frames, wherein the data stream comprises a plurality of side information instances, wherein the data stream further comprises, for each side information instance, transition data including two independently assignable portions which in combination define a point in time to begin a transition from a current reconstruction setting to a desired reconstruction setting specified by the side information instance, and a point in time to complete the transition.