Audio Object Clustering via Renderer-Aware Perceptual Difference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional audio object clustering methods based on spatial proximity fail to accurately combine audio objects in object-based audio systems, leading to undesirable rendering results due to mismatches between spatial and rendering differences, especially in scenarios like speaker and headphone playback systems.

Innovation Solution

The proposed method introduces a renderer-aware perceptual difference as a new factor for audio object clustering, determining the rendering difference between audio objects with respect to a renderer and using this information to control the clustering process, either replacing or combining it with spatial differences to improve clustering accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If audio objects are clustered based on spatial proximity, then the clustering process is simple, but the rendering accuracy deteriorates due to mismatches between spatial and rendering differences

Engineering Contradiction:
Improveclustering process simplicityVSAvoidrendering accuracy
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent changes the clustering parameter from spatial proximity to renderer-aware perceptual difference. Instead of using simple spatial distance metrics, the system calculates perceptual differences by simulating how audio objects would be rendered to speakers, using parameters such as object-to-speaker gains and renderer configuration. This parameter transformation resolves the contradiction by maintaining clustering simplicity while dramatically improving rendering accuracy.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If the number of audio objects is reduced through clustering, then transmission bandwidth and computational complexity are reduced, but audio quality may deteriorate

Engineering Contradiction:
Improvetransmission efficiencyVSAvoidaudio quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent implements feedback by using renderer simulation to evaluate the perceptual impact of clustering decisions. The system simulates how clustered audio objects would be rendered to speakers and compares this against the original unclustered rendering. This feedback mechanism allows the system to optimize clustering to minimize perceptual differences, ensuring that bandwidth reduction does not compromise audio quality.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs preliminary renderer simulation and perceptual difference calculation before finalizing clustering decisions. By pre-evaluating how different clustering configurations will affect rendering quality, the system can make informed decisions about which objects to cluster together, ensuring that quality is preserved while achieving the desired reduction in transmission bandwidth.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If renderer-aware perceptual difference is used for clustering, then rendering accuracy is improved, but computational complexity increases

Engineering Contradiction:
Improverendering accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies partial action by selectively calculating perceptual differences only for audio object pairs that are candidates for clustering, rather than computing all possible pairwise differences. The system identifies relevant objects based on spatial proximity or other heuristics first, then performs the more computationally intensive renderer simulation only on these subsets. This approach maintains high rendering accuracy while significantly reducing overall computational complexity.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP3488623B1Audio object clustering based on renderer-aware perceptual difference
Publication Date: 2020.12.02 DOLBY LABORATORIES LICENSING CORP
  • EP3488623B1 patent drawingFigure 1A~1B
  • EP3488623B1 patent drawingFigure 2~3
  • EP3488623B1 patent drawingFigure 4~5

AI summary

Example embodiments disclosed herein relate to audio object clustering based on renderer-aware perceptual difference. A method of processing audio objects is provided. The method includes obtaining renderer-related information indicating a configuration of a renderer. The method also includes determining, based on the obtained renderer-related information, a rendering difference between a first audio object and a second audio object among the audio objects with respect to the renderer. The method further includes clustering the audio objects at least in part based on the rendering difference. Corresponding system, device, and computer program product are also disclosed.