Perceptual Distance Metric for Spatial Audio Object Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio reproduction systems face challenges in efficiently reducing the number of audio objects in object-based immersive sound scenes while maintaining high perceptual quality, as they lack a computationally efficient method to estimate the perceptual impact of spatial property changes in real-time applications like virtual reality.
Innovation Solution
A perceptual distance metric is introduced to represent perceptual differences in spatial properties of audio scenes, utilizing a perceptual coordinate system, 3D directional loudness map, and spatial masking model to optimize object clustering and generate audio output signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If the number of audio objects is reduced through clustering, then transmission efficiency and storage efficiency are improved, but the perceptual quality of the audio scene deteriorates
Solution Approach 1:
The patent applies parameter changes by transforming the clustering criterion from physical distance to perceptual distance. The perceptual distance metric incorporates psychoacoustic parameters such as directional loudness maps and spatial masking models to determine the perceptual similarity between audio objects. This allows the system to cluster objects based on their perceptual characteristics rather than their physical positions, thereby maintaining perceptual quality while reducing the number of audio objects for transmission and storage.
2Device complexity
If traditional clustering algorithms are used, then computational requirements are reduced, but the perceptual accuracy of spatial property estimation deteriorates
Solution Approach 1:
The patent replaces traditional mechanical distance calculation with a perceptual-based measurement system. Instead of using simple Euclidean distance, the system employs directional loudness maps and spatial masking models to estimate perceptual differences in spatial properties. This substitution enables more accurate perception-based clustering while maintaining computational efficiency through optimized algorithms that leverage the structure of the perceptual space.
3Speed
If real-time processing is implemented, then responsiveness is improved, but computational complexity increases
Solution Approach 1:
The patent applies preliminary action by pre-computing and storing directional loudness maps and spatial masking models before real-time processing is needed. These pre-computed perceptual characteristics can be quickly queried during real-time audio object clustering, significantly reducing the computational complexity of real-time operations. The system prepares the perceptual measurement framework in advance, enabling fast real-time processing without requiring complex computations during critical time windows.
Data Source
AI summary
An apparatus according to an embodiment is provided. The apparatus comprises an input interface for receiving a plurality of audio objects of an audio sound scene. Moreover, the apparatus comprises a processor. Each of the plurality of audio objects represents a sound source being different from any other sound source being represented by any other audio object of the plurality of audio objects; or at least two of the plurality of audio objects represent a same sound source at different locations. The processor is configured to obtain information on a perceptual difference between two audio objects of the plurality of audio objects depending on a distance metric, wherein the distance metric represents perceptual differences in spatial properties of the audio sound scene. And/or, the processor is configured to process the plurality of audio objects to obtain a plurality of audio object clusters or a plurality of processed audio objects depending on the distance metric.


