Audio Object Clustering for Arbitrary Speaker Layouts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

As audio playback systems become increasingly complex with more channels and three-dimensional speaker layouts, existing methods struggle to efficiently process and render audio objects in a way that maintains sound quality and reduces data complexity.

Innovation Solution

A method involving audio object clustering, where N audio objects are grouped into M clusters by determining cluster centroid positions and gain contributions based on cost functions that consider the difference between audio object and cluster or speaker positions, allowing for efficient data reduction and rendering in adaptive audio playback systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If audio objects are clustered to reduce data complexity, then device complexity and data transmission requirements are reduced, but sound quality and spatial accuracy may deteriorate

Engineering Contradiction:
Improveaudio data complexityVSAvoidspatial accuracy
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent combines multiple audio objects into clusters, where each cluster is represented by a single audio object with aggregated properties. This merging reduces the number of individual audio objects that need to be processed and transmitted, thereby reducing device complexity and data transmission requirements while maintaining the overall spatial characteristics of the original audio scene through careful selection of cluster representatives and gain calculations.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transforms the audio scene by changing parameters such as position, gain, and spatial coordinates. By calculating new positions based on weighted averages of original object positions and determining appropriate gain values, the system maintains spatial accuracy despite the reduction in object count. The parameter transformations ensure that the clustered representation preserves the perceptual spatial characteristics of the original multi-object scene.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If the number of speakers increases to create immersive audio experiences, then audio quality and immersion are improved, but device complexity and processing requirements increase

Engineering Contradiction:
Improveaudio qualityVSAvoidspeaker configuration complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent creates a universal audio rendering system that can adapt to various speaker configurations (stereo, 5.1, 7.1, immersive setups) using the same core algorithms. The system calculates gain values and spatial positions that are applicable across different speaker layouts, allowing a single implementation to support multiple speaker configurations without requiring separate processing pipelines for each configuration type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements dynamic adaptation to different speaker configurations by calculating optimal gain values and spatial positions based on the specific playback environment. The system can dynamically adjust the rendering parameters according to the number and arrangement of speakers available, enabling the same audio content to be optimally reproduced across static (2D) and dynamic (3D immersive) speaker layouts.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If audio objects are rendered to arbitrary speaker layouts, then adaptability to different playback environments is improved, but processing complexity increases

Engineering Contradiction:
Improveplayback environment adaptabilityVSAvoidrendering process complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the audio rendering process into distinct stages: clustering audio objects into groups, calculating cluster representative positions, determining gain values for each speaker, and finally rendering to the target speaker layout. This segmentation allows each stage to be processed independently and efficiently, reducing the overall processing complexity while maintaining adaptability to arbitrary speaker configurations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces cluster representatives as intermediary elements between the original audio objects and the final speaker output. These intermediaries simplify the rendering process by reducing the number of direct object-to-speaker mappings that need to be calculated, thereby reducing processing complexity while still preserving the spatial characteristics needed for adaptability to different playback environments.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP3028476B1Panning of audio objects to arbitrary speaker layouts
Publication Date: 2019.03.13 DOLBY INTERNATIONAL AB
  • EP3028476B1 patent drawingFigure 1
  • EP3028476B1 patent drawingFigure 2
  • EP3028476B1 patent drawingFigure 3A~3B

AI summary

A gain contribution of the audio signal for each of the N audio objects to at least one of M speakers may be determined. Determining the gain contribution may involve determining a center of loudness position that is a function of speaker (or cluster) positions and gains assigned to each speaker (or cluster). Determining the gain contribution also may involve determining a minimum value of a cost function. A first term of the cost function may represent a difference between the center of loudness position and an audio object position.