Perception-Based Audio Clustering for Object Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing object-based audio clustering algorithms fail to consider perceptual properties relative to the listener, leading to inefficient reduction of audio objects while maintaining high perceptual quality.
Innovation Solution
Implement perception-based clustering algorithms that group audio objects into clusters using models such as Gaussian mixture models (GMM) and hierarchical clustering, considering perceptual distance metrics, directional loudness maps, and spatial masking to reduce the number of audio objects while preserving quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If existing object-based audio clustering algorithms are used to reduce the number of audio objects, then transmission and storage efficiency is improved, but the algorithms fail to consider perceptual properties relative to the listener resulting in suboptimal reduction while maintaining quality
Solution Approach 1:
The patent transforms the clustering approach by changing from purely spatial parameters to perceptual parameters. It introduces a perception model that maps physical audio object positions to perceptual space, where clustering is performed based on perceptual distance rather than Euclidean distance. This parameter transformation enables the system to group audio objects according to how they are perceived by the listener, achieving better quality preservation while reducing the number of objects.
Solution Approach 2:
The patent introduces a perception model as an intermediary between the physical audio scene and the clustering process. This intermediary layer translates physical positions into perceptual representations, allowing the clustering algorithm to operate on perceptually relevant features rather than raw spatial coordinates. The perception model acts as a mediator that bridges the gap between physical reality and perceptual experience, enabling quality-aware clustering.
2Device complexity
If audio objects are clustered to reduce computational requirements for real-time rendering, then processing complexity is reduced, but existing algorithms do not optimize for perceptual quality
Solution Approach 1:
The patent changes the clustering parameters from spatial coordinates to perceptual attributes derived from a perception model. This transformation allows the system to perform clustering based on perceptual distance metrics that account for human hearing characteristics, such as directional loudness maps and spatial masking. By operating in perceptual space, the algorithm achieves more efficient rendering with better quality preservation.
Solution Approach 2:
The patent replaces traditional spatial-based clustering mechanics with perception-based clustering. Instead of using straightforward geometric distance calculations, the system substitutes a perception model that incorporates auditory masking, localization accuracy limits, and directional loudness characteristics. This substitution transforms the clustering process into a perceptually-driven mechanism that optimizes for human hearing properties rather than purely computational efficiency.
3Ease of manufacture
If spatial properties of audio objects are considered for clustering, then clustering can be performed, but location dependency in spatial localization accuracy is not considered
Solution Approach 1:
The patent transforms the clustering approach by changing from fixed spatial coordinates to perceptual parameters that account for location-dependent localization accuracy. It introduces a perception model that maps physical positions to perceptual space, where the transformation varies with location. This parameter change enables the system to capture the non-uniform nature of human localization accuracy across different spatial regions, improving measurement precision while maintaining clustering capability.
Solution Approach 2:
The patent applies local quality by making the perception model location-dependent. Different regions of the audio space are treated differently according to their perceptual characteristics. The perception model incorporates directional loudness maps and spatial masking that vary with position, allowing the clustering algorithm to adapt to local perceptual properties. This local differentiation improves localization accuracy measurement while preserving the overall clustering function.
Data Source
AI summary
An apparatus according to an embodiment is provided The apparatus comprises an input interface for receiving information on three or more audio objects. Moreover, the apparatus comprises a cluster generator for generating two or more audio object clusters by associating each of the three or more audio objects with at least one of the two or more audio object clusters, such that, for each of the two or more audio object clusters, at least one of the three or more audio objects is associated to said audio object cluster, and such that, for each of at least one of the two or more audio object clusters, at least two of the three or more audio objects are associated with said audio object cluster. The cluster generator is configured to generate the two or more audio object clusters depending on a perception-based model.


