Audio Object Extraction via Spectral Similarity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies lack an effective method for accurately and efficiently extracting audio objects from traditional channel-based audio content, which limits the ability to provide immersive experiences similar to object-based audio formats.
Innovation Solution
A method and system for audio object extraction that applies frame-level audio object extraction based on frequency spectral similarities among channels, followed by audio object composition across frames to generate complete tracks of audio objects, utilizing techniques such as hierarchical clustering and probability matrix generation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If frame-level audio object extraction is applied to extract audio objects from channel-based content, then the ability to provide immersive experiences is improved, but the complexity of the extraction process increases
Solution Approach 1:
The extraction process is divided into two distinct stages: frame-level audio object extraction and audio object composition across frames. This segmentation allows the system to handle complex extraction tasks in manageable steps, first identifying objects within individual frames using frequency spectral similarities, then composing these objects across multiple frames to maintain temporal consistency and provide immersive experiences.
Solution Approach 2:
The system performs preliminary audio object extraction on individual frames before composing objects across frames. By pre-identifying audio objects in each frame based on frequency spectral similarities among channels, the system prepares the necessary data structures and object identifiers that will be used in the subsequent composition stage, reducing the overall computational complexity.
2Measurement precision
If audio object extraction is performed on individual frames based on frequency spectral similarities, then the accuracy of object identification is improved, but the computational time increases
Solution Approach 1:
The system applies audio object extraction only on individual frames where it is necessary to identify audio objects, rather than processing every frame with the full extraction algorithm. By selectively applying extraction based on frequency spectral similarities, the system achieves accurate object identification while reducing unnecessary computational overhead.
Solution Approach 2:
The system extracts only the essential information needed for audio object identification from each frame, specifically focusing on frequency spectral similarities among channels. By taking out only the critical features required for object identification rather than processing all audio data, the system improves accuracy while minimizing computational time requirements.
3Reliability
If audio object composition is performed across frames, then the completeness of audio object tracks is improved, but the processing complexity increases
Solution Approach 1:
The system merges audio object information across multiple frames by composing extracted objects temporally. By combining the results from individual frame extractions and maintaining temporal consistency through composition operations, the system produces complete and reliable audio object tracks that span across the entire audio content duration.
Solution Approach 2:
The audio object composition process maintains continuous tracking of audio objects across frames, ensuring that objects are followed and tracked throughout the temporal sequence. This continuity ensures complete audio object tracks are generated, with the system maintaining useful action by consistently applying composition operations across all relevant frames.
Data Source
AI summary
Embodiments of the present invention relate to audio object extraction. A method for audio object extraction from audio content of a format based on a plurality of channels is disclosed. The method comprises applying audio object extraction on individual frames of the audio content at least partially based on frequency spectral similarities among the plurality of channels. The method further comprises performing audio object composition across the frames of the audio content, based on the audio object extraction on the individual frames, to generate a track of at least one audio object. Corresponding system and computer program product are also disclosed.


