Audio Object Clustering for Legacy Decoder Compatibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio encoding methods fail to efficiently reconstruct audio objects, particularly when multiple objects share similar spatial positions, leading to imperfect reconstruction and audible artifacts, and require high computational complexity.

Innovation Solution

An encoder and decoder system that calculates adaptive downmix signals and side information independently of loudspeaker configurations, allowing for flexible combination of audio objects and reducing computational complexity by clustering and using time-varying metadata for improved fidelity and separation of audio objects.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If audio objects are combined into a multichannel downmix, then compatibility with legacy decoders is improved, but reconstruction fidelity of audio objects deteriorates

Engineering Contradiction:
Improvecompatibility with legacy decodersVSAvoidreconstruction fidelity
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent segments the audio objects into spatial clusters based on their position characteristics. Audio objects with similar spatial positions are grouped together, allowing the system to maintain separate spatial information while creating a compatible downmix. This segmentation enables legacy decoders to process the downmix while preserving reconstruction fidelity through the clustering approach.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces spatial clustering as an intermediary mechanism between the audio objects and the multichannel downmix. This intermediary structure allows the system to transform audio objects into a downmix format compatible with legacy decoders while maintaining spatial information through the clustering parameters, thus resolving the contradiction between compatibility and fidelity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If audio objects with similar horizontal positions are combined into the same channel, then device complexity is reduced, but reconstruction precision of individual audio objects deteriorates

Engineering Contradiction:
Improvechannel assignment complexityVSAvoidvertical position differentiation
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent extends the spatial clustering from two-dimensional horizontal positioning to three-dimensional spatial positioning by incorporating vertical position information. This dimensional expansion allows the system to differentiate audio objects based on vertical positions even when they share horizontal positions, thereby maintaining reconstruction precision without significantly increasing channel assignment complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent applies local quality by creating spatial clusters that group audio objects with similar spatial characteristics. Within each cluster, audio objects are treated as a unit for downmix purposes, while the cluster itself preserves the spatial differentiation information. This local clustering approach maintains precision for individual audio objects while simplifying the overall channel assignment process.

Inventive Principle:
Principle #3Local quality

3Device complexity

If parametric reconstruction is used, then computational complexity is reduced, but reconstruction fidelity deteriorates

Engineering Contradiction:
Improvecomputational complexityVSAvoidreconstruction fidelity
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent performs preliminary spatial clustering of audio objects before creating the downmix. This preliminary action organizes the audio objects into spatially coherent groups, which simplifies the subsequent parametric reconstruction process. By pre-grouping the audio objects, the system reduces the computational complexity required for reconstruction while maintaining fidelity through the preserved spatial cluster information.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the parameters used for audio object representation by introducing spatial cluster identifiers alongside traditional audio parameters. This parameter transformation allows the system to use simpler parametric reconstruction methods while maintaining high fidelity, as the spatial cluster information provides the necessary contextual data for accurate reconstruction without requiring complex computations.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11705139B2Efficient coding of audio scenes comprising audio objects
Publication Date: 2023.07.18 DOLBY INTERNATIONAL AB
  • US11705139B2 patent drawing
  • US11705139B2 patent drawing
  • US11705139B2 patent drawing

AI summary

There is provided encoding and decoding methods for encoding and decoding of object based audio. An exemplary encoding method includes inter alia calculating M downmix signals by forming combinations of N audio objects, wherein M≤N, and calculating parameters which allow reconstruction of a set of audio objects formed on basis of the N audio objects from the M downmix signals. The calculation of the M downmix signals is made according to a criterion which is independent of any loudspeaker configuration.