Composite Spatial Audio Object Rendering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for rendering object-based audio scenes with many audio objects face high computational load and loss of spatial information, limiting creative freedom and efficiency, especially in dynamic and multi-user scenarios.
Innovation Solution
The method transforms clusters of audio objects into composite spatial audio objects, preserving spatial information independent of listening positions, reducing rendering complexity and enabling efficient adaptation to changes in user and object positions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If direct binaural rendering of each individual audio object is performed, then spatial accuracy is improved, but computational load increases significantly
Solution Approach 1:
The patent merges multiple individual audio objects into a cluster represented by a single composite audio object with aggregated spatial characteristics. This combining approach maintains the overall spatial impression of the cluster while dramatically reducing the number of separate rendering operations required, thus lowering computational load while preserving perceptual spatial accuracy.
Solution Approach 2:
The patent segments the audio scene into clusters of objects that can be rendered together as a unit. By dividing the scene into manageable clusters rather than rendering every individual object separately, the system reduces computational complexity while maintaining spatial fidelity at the cluster level.
2Power
If clustering of audio objects is performed to reduce rendering complexity, then computational load is reduced, but spatial information is lost
Solution Approach 1:
The patent transforms the spatial representation parameters by computing aggregate spatial characteristics (such as mean position, spatial covariance) for each cluster. This parameter transformation allows the cluster to represent multiple objects with different spatial positions while maintaining essential spatial information in a compressed form that can be rendered efficiently.
3Device complexity
If indirect binaural rendering through pre-rendering to virtual loudspeaker configuration is used, then rendering complexity is reduced, but adaptability to dynamic positions decreases
Solution Approach 1:
The patent implements dynamic cluster formation and spatial parameter computation that adapts to changing object positions and user orientation. Rather than using fixed pre-rendered configurations, the system continuously updates cluster spatial characteristics based on current scene state, enabling real-time adaptation to dynamic positions while maintaining reduced rendering complexity.
Data Source
AI summary
A method (900) for representing a cluster of audio objects (202). The method includes obtaining (s902) reference position information identifying a reference position within the cluster of audio objects; and using the obtained reference position information, transforming (s904) the cluster of audio objects into a composite spatial audio object. The transforming of the cluster of audio objects into the composite spatial audio object is performed independently of any listening position of any user.


