Composite Spatial Audio Object Rendering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for rendering object-based audio scenes with many audio objects face high computational load and loss of spatial information, limiting creative freedom and efficiency, especially in dynamic and multi-user scenarios.

Innovation Solution

The method transforms clusters of audio objects into composite spatial audio objects, preserving spatial information independent of listening positions, reducing rendering complexity and enabling efficient adaptation to changes in user and object positions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If direct binaural rendering of each individual audio object is performed, then spatial accuracy is improved, but computational load increases significantly

Engineering Contradiction:
Improvespatial accuracyVSAvoidcomputational load
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent merges multiple individual audio objects into a cluster represented by a single composite audio object with aggregated spatial characteristics. This combining approach maintains the overall spatial impression of the cluster while dramatically reducing the number of separate rendering operations required, thus lowering computational load while preserving perceptual spatial accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent segments the audio scene into clusters of objects that can be rendered together as a unit. By dividing the scene into manageable clusters rather than rendering every individual object separately, the system reduces computational complexity while maintaining spatial fidelity at the cluster level.

Inventive Principle:
Principle #1Segmentation

2Power

If clustering of audio objects is performed to reduce rendering complexity, then computational load is reduced, but spatial information is lost

Engineering Contradiction:
Improvecomputational loadVSAvoidspatial information
Core Design Contradiction:
PowerVSLoss of information

Solution Approach 1:

The patent transforms the spatial representation parameters by computing aggregate spatial characteristics (such as mean position, spatial covariance) for each cluster. This parameter transformation allows the cluster to represent multiple objects with different spatial positions while maintaining essential spatial information in a compressed form that can be rendered efficiently.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If indirect binaural rendering through pre-rendering to virtual loudspeaker configuration is used, then rendering complexity is reduced, but adaptability to dynamic positions decreases

Engineering Contradiction:
Improverendering complexityVSAvoidadaptability to dynamic positions
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic cluster formation and spatial parameter computation that adapts to changing object positions and user orientation. Rather than using fixed pre-rendered configurations, the system continuously updates cluster spatial characteristics based on current scene state, enabling real-time adaptation to dynamic positions while maintaining reduced rendering complexity.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20230088922A1Representation and rendering of audio objects
Publication Date: 2023.03.23 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • US20230088922A1 patent drawing
  • US20230088922A1 patent drawing
  • US20230088922A1 patent drawing

AI summary

A method (900) for representing a cluster of audio objects (202). The method includes obtaining (s902) reference position information identifying a reference position within the cluster of audio objects; and using the obtained reference position information, transforming (s904) the cluster of audio objects into a composite spatial audio object. The transforming of the cluster of audio objects into the composite spatial audio object is performed independently of any listening position of any user.