Audio Object Extraction Using Projection Spaces for Adaptive Playback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Channel-based audio content formats are inefficient in adapting to various playback configurations, leading to degraded listening experiences due to mismatched playback settings and limitations in binaural rendering, particularly in separating and positioning audio objects.
Innovation Solution
A method and system for extracting audio objects from channel-based content by identifying projection spaces and determining correlations between channels to separate and re-render objects, using source separation techniques to produce clean multi-channel or mono representations adaptable to different playback settings.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If channel-based audio content formats are used, then the audio content can be created and stored with predefined physical locations, but the listening experience degrades when played back with different playback settings due to mismatch between playback settings and predefined channel positions
Solution Approach 1:
The patent segments the audio content into independent audio objects with their own positional metadata, separating them from the channel-based structure. This allows each object to be independently positioned and rendered according to the actual playback configuration, enabling adaptation to different speaker setups while maintaining accurate spatial positioning and listening experience quality.
Solution Approach 2:
The patent implements dynamic rendering where audio objects are re-positioned and re-rendered in real-time based on the detected playback configuration. Instead of static channel assignments, the system dynamically adjusts the positioning and rendering of audio objects to match the actual speaker arrangement, ensuring optimal listening experience across various playback settings.
2Adaptability or versatility
If channel-based formats are used, then the audio content can be distributed with fixed speaker assignments, but the binaural rendering is limited to a small number of head-related transfer functions (HRTFs) specific to speaker positions, degrading the binaural listening experience for other positions
Solution Approach 1:
The patent makes the audio rendering system universal by extracting audio objects with positional metadata that can be rendered with any HRTF set corresponding to any speaker configuration. Instead of being limited to specific channel-HRTF mappings, the system can universally adapt to any speaker arrangement by selecting appropriate HRTFs for each audio object's position, achieving accurate binaural rendering across diverse playback settings.
3Measurement precision
If source separation techniques are used to separate objects from multi-channel mixtures, then clean audio objects can be extracted for accurate position estimation, but the device complexity increases
Solution Approach 1:
The patent applies source separation techniques during the audio content creation and encoding phase, before distribution. This preliminary action extracts and identifies audio objects with their positional information, storing this metadata with the audio content. During playback, the system simply retrieves and uses this pre-computed positional information without requiring complex real-time separation, thus achieving accurate position estimation while minimizing playback system complexity.
Data Source
Figure 1~2
Figure 3~4
AI summary
A method is disclosed for audio object extraction from an audio content which includes identifying a first set of projection spaces including a first subset for a first channel and a second subset for a second channel of the plurality of channels. The method may further include determining a first set of correlations between the first and second channels, each of the first set of correlations corresponding to one of the first subset of projection spaces and one of the second subset of projection spaces. Still further, the method may include extracting an audio object from an audio signal of the first channel at least in part based on a first correlation among the first set of correlations and the projection space from the first subset corresponding to the first correlation, the first correlation being greater than a first predefined threshold. Corresponding system and computer program products are also disclosed.