Perceptual Distance Metric for Spatial Audio Object Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio reproduction systems face challenges in efficiently reducing the number of audio objects in object-based immersive sound scenes while maintaining high perceptual quality, as they lack a computationally efficient method to estimate the perceptual impact of spatial property changes in real-time applications like virtual reality.

Innovation Solution

A perceptual distance metric is introduced to represent perceptual differences in spatial properties of audio scenes, utilizing a perceptual coordinate system, 3D directional loudness map, and spatial masking model to optimize object clustering and generate audio output signals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If the number of audio objects is reduced through clustering, then transmission efficiency and storage efficiency are improved, but the perceptual quality of the audio scene deteriorates

Engineering Contradiction:
Improvetransmission efficiencyVSAvoidperceptual quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies parameter changes by transforming the clustering criterion from physical distance to perceptual distance. The perceptual distance metric incorporates psychoacoustic parameters such as directional loudness maps and spatial masking models to determine the perceptual similarity between audio objects. This allows the system to cluster objects based on their perceptual characteristics rather than their physical positions, thereby maintaining perceptual quality while reducing the number of audio objects for transmission and storage.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If traditional clustering algorithms are used, then computational requirements are reduced, but the perceptual accuracy of spatial property estimation deteriorates

Engineering Contradiction:
Improvecomputational requirementsVSAvoidperceptual accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent replaces traditional mechanical distance calculation with a perceptual-based measurement system. Instead of using simple Euclidean distance, the system employs directional loudness maps and spatial masking models to estimate perceptual differences in spatial properties. This substitution enables more accurate perception-based clustering while maintaining computational efficiency through optimized algorithms that leverage the structure of the perceptual space.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Speed

If real-time processing is implemented, then responsiveness is improved, but computational complexity increases

Engineering Contradiction:
Improveprocessing speedVSAvoidcomputational complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-computing and storing directional loudness maps and spatial masking models before real-time processing is needed. These pre-computed perceptual characteristics can be quickly queried during real-time audio object clustering, significantly reducing the computational complexity of real-time operations. The system prepares the perceptual measurement framework in advance, enabling fast real-time processing without requiring complex computations during critical time windows.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250287170A1Apparatus and method employing a perception-based distance metric for spatial audio
Publication Date: 2025.09.11 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US20250287170A1 patent drawing
  • US20250287170A1 patent drawing
  • US20250287170A1 patent drawing

AI summary

An apparatus according to an embodiment is provided. The apparatus comprises an input interface for receiving a plurality of audio objects of an audio sound scene. Moreover, the apparatus comprises a processor. Each of the plurality of audio objects represents a sound source being different from any other sound source being represented by any other audio object of the plurality of audio objects; or at least two of the plurality of audio objects represent a same sound source at different locations. The processor is configured to obtain information on a perceptual difference between two audio objects of the plurality of audio objects depending on a distance metric, wherein the distance metric represents perceptual differences in spatial properties of the audio sound scene. And/or, the processor is configured to process the plurality of audio objects to obtain a plurality of audio object clusters or a plurality of processed audio objects depending on the distance metric.