Merging-Friendly Video File Format for Spatial Subset Streaming
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing file formats for coded video data, such as ISO/IEC 14496-12 and ISO/IEC 23090-2, incur significant overhead and complexity in handling spatial subsets of coded videos, particularly in 360-degree video streaming, due to separate extractor tracks for each viewing direction, leading to inefficient resource usage and delayed data processing.
Innovation Solution
A 'merging friendly' file format that groups source tracks into mergeable sets, uses templates for configurable parameter sets and SEI messages, and aligns random access points, allowing efficient merging and decoding of spatial video subsets without requiring complete data download.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If separate extractor tracks are used for each viewing direction in 360-degree video streaming, then spatial subset extraction is enabled, but file format overhead and complexity increase significantly
Solution Approach 1:
The patent merges multiple extractor tracks into a single unified track structure. Instead of having separate extractor tracks for each viewing direction, the invention uses a single track that contains all extractor NAL units, eliminating the need for multiple separate track structures and reducing overall file format complexity while maintaining the ability to extract spatial subsets for different viewing directions.
Solution Approach 2:
The single track structure serves multiple functions by containing extractor NAL units for different viewing directions. This universal track design enables the system to handle various spatial subset extraction requirements without requiring separate dedicated tracks for each function, thereby reducing overhead and simplifying the file format.
2Adaptability or versatility
If separate extractor tracks are used for each viewing direction, then spatial subset extraction is enabled, but processing time is delayed
Solution Approach 1:
The patent implements preliminary action by pre-organizing all extractor NAL units within a single track structure during the encoding phase. This pre-organization allows the decoder to efficiently locate and process the required extractor units without needing to search through multiple separate tracks, thereby reducing processing time while maintaining the capability to extract spatial subsets for different viewing directions.
3Adaptability or versatility
If separate extractor tracks are used for each viewing direction, then spatial subset extraction is enabled, but computational resources are wasted
Solution Approach 1:
The patent merges multiple extractor tracks into a single unified track structure. Instead of having separate extractor tracks for each viewing direction, the invention uses a single track that contains all extractor NAL units, eliminating the need for multiple separate track structures and reducing overall file format complexity while maintaining the ability to extract spatial subsets for different viewing directions.
Solution Approach 2:
The invention extracts only the necessary extractor NAL units from the single track when processing different viewing directions. By using a unified track structure, the system can selectively extract and process only the required portions without the overhead of managing multiple separate tracks, thereby improving computational efficiency and reducing resource waste.
4Measurement precision
If complete data is downloaded for all tiles, then spatial subset extraction is accurate, but bandwidth consumption increases
Solution Approach 1:
The patent applies partial action by enabling the system to download and process only the necessary extractor NAL units for the required spatial subsets rather than requiring complete data for all tiles. The single track structure with organized extractor units allows selective extraction of only the needed portions, reducing bandwidth consumption while maintaining extraction accuracy for the specific viewing directions required.
Data Source
AI summary
Video data for deriving a spatially variable section of a scene therefrom as well as corresponding methods and apparatuses for creating video data for deriving a spatially variable section of a scene therefrom and for deriving a spatially variable section of a scene from video data. The video data comprises a set of source tracks comprising coded video data representing spatial portions of a video showing the scene and is formatted in a specific file format and supports the merging of different spatial portions into a joint bitstream through compressed-domain processing.


