Audio Object Reconstruction via Spatial Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio encoding methods fail to efficiently reconstruct audio objects, particularly when multiple objects share similar spatial positions, leading to imperfect reconstruction and audible artifacts, and lack flexibility in adapting to dynamic audio scenes.
Innovation Solution
An encoder and decoder system that calculates adaptive downmix signals and side information independently of loudspeaker configurations, allowing for perfect reconstruction of audio objects by clustering audio objects based on spatial proximity and importance, and including time-varying metadata for dynamic audio scenes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If audio objects are combined into a multichannel downmix for legacy decoder compatibility, then legacy playback systems can use the downmix directly, but the reconstruction of audio objects at the decoder side becomes insufficient and may lead to audible artifacts
Solution Approach 1:
The patent segments the audio scene into multiple clusters based on spatial proximity and importance, where each cluster is represented by a separate audio object. This segmentation allows the decoder to reconstruct multiple distinct audio objects from the downmix, improving reconstruction accuracy while maintaining compatibility with legacy systems through the downmix signal.
Solution Approach 2:
The patent introduces a new dimension of side information (metadata) that describes the spatial positions and characteristics of audio objects beyond the traditional multichannel downmix. This additional dimensional information enables accurate reconstruction of audio objects in three-dimensional space while preserving backward compatibility with legacy decoders that only process the downmix.
2Loss of information
If the number of audio objects is very large (hundreds of objects), then the audio scene can be represented comprehensively, but the coding and reconstruction complexity increases significantly
Solution Approach 1:
The patent merges multiple audio objects into clusters based on their spatial proximity and importance, representing each cluster as a single audio object. This merging reduces the total number of audio objects from hundreds to a manageable number, significantly decreasing coding and reconstruction complexity while preserving the essential spatial and perceptual characteristics of the original audio scene.
Solution Approach 2:
The patent changes the parameter representation by using cluster-based spatial positions and importance weights instead of individual object parameters for every audio source. This parameter transformation reduces the data dimensionality and computational load while maintaining faithful representation of the audio scene's spatial structure.
3Device complexity
If audio objects with the same horizontal position but different vertical positions are combined into the same downmix channel, then the downmix can be simplified, but the decoder cannot ensure perfect reconstruction and may produce audible artifacts
Solution Approach 1:
The patent applies local quality by assigning different spatial characteristics to different regions of the audio scene. Audio objects at different vertical positions are assigned distinct spatial parameters in the side information, allowing the decoder to reconstruct them with appropriate vertical localization. This local differentiation preserves reconstruction fidelity while allowing simplification in regions where objects are truly indistinguishable.
Data Source
AI summary
There is provided encoding and decoding methods for encoding and decoding of object based audio. An exemplary encoding method includes inter alia calculating M downmix signals by forming combinations of N audio objects, wherein M≦N, and calculating parameters which allow reconstruction of a set of audio objects formed on basis of the N audio objects from the M downmix signals. The calculation of the M downmix signals is made according to a criterion which is independent of any loudspeaker configuration.


