Perception-Based Audio Clustering for Object Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing object-based audio clustering algorithms fail to consider perceptual properties relative to the listener, leading to inefficient reduction of audio objects while maintaining high perceptual quality.

Innovation Solution

Implement perception-based clustering algorithms that group audio objects into clusters using models such as Gaussian mixture models (GMM) and hierarchical clustering, considering perceptual distance metrics, directional loudness maps, and spatial masking to reduce the number of audio objects while preserving quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If existing object-based audio clustering algorithms are used to reduce the number of audio objects, then transmission and storage efficiency is improved, but the algorithms fail to consider perceptual properties relative to the listener resulting in suboptimal reduction while maintaining quality

Engineering Contradiction:
Improvenumber of audio objectsVSAvoidperceptual quality
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent transforms the clustering approach by changing from purely spatial parameters to perceptual parameters. It introduces a perception model that maps physical audio object positions to perceptual space, where clustering is performed based on perceptual distance rather than Euclidean distance. This parameter transformation enables the system to group audio objects according to how they are perceived by the listener, achieving better quality preservation while reducing the number of objects.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces a perception model as an intermediary between the physical audio scene and the clustering process. This intermediary layer translates physical positions into perceptual representations, allowing the clustering algorithm to operate on perceptually relevant features rather than raw spatial coordinates. The perception model acts as a mediator that bridges the gap between physical reality and perceptual experience, enabling quality-aware clustering.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If audio objects are clustered to reduce computational requirements for real-time rendering, then processing complexity is reduced, but existing algorithms do not optimize for perceptual quality

Engineering Contradiction:
Improvecomputational complexityVSAvoidperceptual quality
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent changes the clustering parameters from spatial coordinates to perceptual attributes derived from a perception model. This transformation allows the system to perform clustering based on perceptual distance metrics that account for human hearing characteristics, such as directional loudness maps and spatial masking. By operating in perceptual space, the algorithm achieves more efficient rendering with better quality preservation.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces traditional spatial-based clustering mechanics with perception-based clustering. Instead of using straightforward geometric distance calculations, the system substitutes a perception model that incorporates auditory masking, localization accuracy limits, and directional loudness characteristics. This substitution transforms the clustering process into a perceptually-driven mechanism that optimizes for human hearing properties rather than purely computational efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Ease of manufacture

If spatial properties of audio objects are considered for clustering, then clustering can be performed, but location dependency in spatial localization accuracy is not considered

Engineering Contradiction:
Improveclustering capabilityVSAvoidlocalization accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent transforms the clustering approach by changing from fixed spatial coordinates to perceptual parameters that account for location-dependent localization accuracy. It introduces a perception model that maps physical positions to perceptual space, where the transformation varies with location. This parameter change enables the system to capture the non-uniform nature of human localization accuracy across different spatial regions, improving measurement precision while maintaining clustering capability.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies local quality by making the perception model location-dependent. Different regions of the audio space are treated differently according to their perceptual characteristics. The perception model incorporates directional loudness maps and spatial masking that vary with position, allowing the clustering algorithm to adapt to local perceptual properties. This local differentiation improves localization accuracy measurement while preserving the overall clustering function.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20250287169A1Apparatus and method for perception-based clustering of object-based audio scenes
Publication Date: 2025.09.11 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US20250287169A1 patent drawing
  • US20250287169A1 patent drawing
  • US20250287169A1 patent drawing

AI summary

An apparatus according to an embodiment is provided The apparatus comprises an input interface for receiving information on three or more audio objects. Moreover, the apparatus comprises a cluster generator for generating two or more audio object clusters by associating each of the three or more audio objects with at least one of the two or more audio object clusters, such that, for each of the two or more audio object clusters, at least one of the three or more audio objects is associated to said audio object cluster, and such that, for each of at least one of the two or more audio object clusters, at least two of the three or more audio objects are associated with said audio object cluster. The cluster generator is configured to generate the two or more audio object clusters depending on a perception-based model.