Audio Object Clustering Using Spatial Error Metrics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing technologies face challenges in determining spatial error metrics and audio quality degradation during audio object clustering, leading to inefficiencies in format conversion, rendering, and transmission, particularly in reducing the number of audio objects while maintaining spatial fidelity.
Innovation Solution
The development of a method and system that computes spatial error metrics and audio quality degradation by analyzing intra-frame and inter-frame spatial errors, using importance-weighted and normalized error metrics, and predicting subjective audio quality through correlation with user surveys, to optimize the audio object clustering process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If the number of audio objects is reduced through clustering, then bandwidth and processing requirements are decreased, but spatial fidelity and audio quality may deteriorate
Solution Approach 1:
The patent implements feedback by computing spatial error metrics (such as azimuth angle error, elevation angle error, and distance error) between original audio objects and clustered audio objects. These error metrics are used to evaluate and optimize the clustering process, ensuring that spatial fidelity is maintained while achieving compression goals. The feedback loop allows iterative refinement of clustering parameters to balance efficiency and quality.
Solution Approach 2:
The patent employs parameter changes by adjusting clustering thresholds, error metric weights, and spatial tolerance levels to optimize the balance between compression ratio and spatial fidelity. Different parameter sets can be applied depending on the specific application requirements, allowing flexible control over the trade-off between encoding efficiency and audio quality.
2Measurement precision
If spatial error metrics are computed for all audio objects, then audio quality assessment is improved, but processing time and computational complexity increase
Solution Approach 1:
The patent applies local quality by computing spatial error metrics selectively for audio objects that meet specific criteria (such as those with significant spatial characteristics or those that are likely to impact perceived audio quality). This selective approach focuses computational resources on critical audio objects while reducing overall processing time for less important objects.
Solution Approach 2:
The patent implements partial action by computing a subset of spatial error metrics (such as only azimuth angle error or only elevation angle error) depending on the specific application needs, rather than computing all possible metrics. This partial computation approach reduces processing time while still providing sufficient quality assessment for the given application.
3Adaptability or versatility
If audio content is adapted for multiple distribution settings, then versatility and adaptability are improved, but device complexity and processing requirements increase
Solution Approach 1:
The patent implements universality by creating a unified audio object-based representation that can be adapted to multiple distribution settings (such as immersive audio, stereo, mono, different channel configurations) through a single source. The spatial error metric computation framework serves multiple functions by evaluating quality across different rendering scenarios and providing optimization guidance for various distribution formats simultaneously.
Data Source
Figure 1~2
Figure 3A
Figure 3B
AI summary
Audio objects that are present in input audio content in one or more frames are determined. Output clusters that are present in output audio content in the one or more frames are also determined. Here, the audio objects in the input audio content are converted to the output clusters in the output audio content. One or more spatial error metrics are computed based at least in part on positional metadata of the audio objects and positional metadata of the output clusters.