Audio Object Clustering for Legacy Decoder Compatibility
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio encoding methods fail to efficiently reconstruct audio objects, particularly when multiple objects share similar spatial positions, leading to imperfect reconstruction and audible artifacts, and require high computational complexity.
Innovation Solution
An encoder and decoder system that calculates adaptive downmix signals and side information independently of loudspeaker configurations, allowing for flexible combination of audio objects and reducing computational complexity by clustering and using time-varying metadata for improved fidelity and separation of audio objects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If audio objects are combined into a multichannel downmix, then compatibility with legacy decoders is improved, but reconstruction fidelity of audio objects deteriorates
Solution Approach 1:
The patent segments the audio objects into spatial clusters based on their position characteristics. Audio objects with similar spatial positions are grouped together, allowing the system to maintain separate spatial information while creating a compatible downmix. This segmentation enables legacy decoders to process the downmix while preserving reconstruction fidelity through the clustering approach.
Solution Approach 2:
The patent introduces spatial clustering as an intermediary mechanism between the audio objects and the multichannel downmix. This intermediary structure allows the system to transform audio objects into a downmix format compatible with legacy decoders while maintaining spatial information through the clustering parameters, thus resolving the contradiction between compatibility and fidelity.
2Device complexity
If audio objects with similar horizontal positions are combined into the same channel, then device complexity is reduced, but reconstruction precision of individual audio objects deteriorates
Solution Approach 1:
The patent extends the spatial clustering from two-dimensional horizontal positioning to three-dimensional spatial positioning by incorporating vertical position information. This dimensional expansion allows the system to differentiate audio objects based on vertical positions even when they share horizontal positions, thereby maintaining reconstruction precision without significantly increasing channel assignment complexity.
Solution Approach 2:
The patent applies local quality by creating spatial clusters that group audio objects with similar spatial characteristics. Within each cluster, audio objects are treated as a unit for downmix purposes, while the cluster itself preserves the spatial differentiation information. This local clustering approach maintains precision for individual audio objects while simplifying the overall channel assignment process.
3Device complexity
If parametric reconstruction is used, then computational complexity is reduced, but reconstruction fidelity deteriorates
Solution Approach 1:
The patent performs preliminary spatial clustering of audio objects before creating the downmix. This preliminary action organizes the audio objects into spatially coherent groups, which simplifies the subsequent parametric reconstruction process. By pre-grouping the audio objects, the system reduces the computational complexity required for reconstruction while maintaining fidelity through the preserved spatial cluster information.
Solution Approach 2:
The patent changes the parameters used for audio object representation by introducing spatial cluster identifiers alongside traditional audio parameters. This parameter transformation allows the system to use simpler parametric reconstruction methods while maintaining high fidelity, as the spatial cluster information provides the necessary contextual data for accurate reconstruction without requiring complex computations.
Data Source
AI summary
There is provided encoding and decoding methods for encoding and decoding of object based audio. An exemplary encoding method includes inter alia calculating M downmix signals by forming combinations of N audio objects, wherein M≤N, and calculating parameters which allow reconstruction of a set of audio objects formed on basis of the N audio objects from the M downmix signals. The calculation of the M downmix signals is made according to a criterion which is independent of any loudspeaker configuration.


