Video-Assisted Audio Object Extraction for Immersive Channel-Based Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional channel-based audio content formats lack the capability to provide immersive experiences similar to object-based audio, necessitating the extraction of audio objects from channel-based content to align with video content for enhanced immersion.
Innovation Solution
A method and system for video content-assisted audio object extraction, where video object-based information is used to assist in extracting audio objects from channel-based audio content, utilizing techniques like video object extraction, audio template generation, and metadata estimation to improve the precision and alignment of audio object extraction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If audio objects are extracted from channel-based audio content without video assistance, then the extraction process is simpler, but the precision and alignment of audio object extraction deteriorates
Solution Approach 1:
The patent uses video content as an intermediary to assist audio object extraction. Video objects and their metadata (position, velocity, size) serve as mediating information that guides the extraction of corresponding audio objects from channel-based audio, improving extraction precision without requiring complex audio-only processing
Solution Approach 2:
The patent transitions from one-dimensional audio signal processing to two-dimensional video-audio joint processing. By incorporating video spatial information (x, y coordinates) and temporal information, the system achieves more precise audio object localization and extraction than audio-only methods
2Adaptability or versatility
If traditional channel-based audio formats are used, then compatibility with existing systems is maintained, but the immersive listening experience deteriorates
Solution Approach 1:
The patent segments the audio signal into distinct audio objects with associated metadata (position, velocity, size, type), separating them from the background audio bed. This segmentation enables independent control and spatial positioning of individual sound sources, creating immersive 3D audio experiences while maintaining compatibility with existing channel-based playback systems
Solution Approach 2:
The extracted audio objects with metadata can be universally applied to various playback configurations (stereo, surround 5.1, 7.1, etc.). The same audio object data structure works across different speaker arrangements and formats, providing versatile immersive experience support without requiring format-specific processing
3Measurement precision
If video content is used to assist audio object extraction, then the alignment between audio and video improves, but the processing time and computational resources increase
Solution Approach 1:
The patent performs preliminary video object detection and metadata extraction before audio object extraction. By pre-processing the video content to identify objects, their positions, velocities, and sizes, the system prepares reference information that accelerates the subsequent audio object extraction and alignment process
Solution Approach 2:
The system uses video object metadata as feedback to guide and constrain audio object extraction. The known video object positions and characteristics provide feedback signals that help identify corresponding audio objects more quickly and accurately, reducing the computational search space and processing time
Data Source
Figure 1~2
Figure 3~4
Figure 5~6
AI summary
Embodiments of the present invention relate to video content assisted audio object extraction. A method of audio object extraction from channel-based audio content is disclosed. The method comprises extracting at least one video object from video content associated with the channel-based audio content, and determining information about the at least one video object. The method further comprises extracting from the channel-based audio content an audio object to be rendered as an upmixed audio signal based on the determined information. Corresponding system and computer program product are also disclosed.