Partitioned Video Encapsulation via Spatial Tracks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing media data encapsulation and streaming technologies face challenges in adaptability to new devices and usage scenarios, particularly in efficiently managing and transmitting partitioned video data with low description overhead and dynamic track references.
Innovation Solution
A method for encapsulating partitioned video data by creating spatial tracks and a base track with reconstruction instructions, allowing flexible reconstruction of bit-streams from partitioned frames, enabling dynamic change of track references and low metadata overhead, conforming to standards like MPEG Versatile Video Coding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If partitioned video data is encapsulated using traditional ISOBMFF format, then the media content can be streamed over HTTP, but the description overhead is high and adaptability to new devices is limited
Solution Approach 1:
The video data is divided into spatial partitions (tiles, slices, or subpictures) where each partition is independently encoded and encapsulated as a separate track. This segmentation enables selective transmission of specific spatial regions, reducing the amount of metadata required compared to traditional formats while improving adaptability to different device capabilities and network conditions.
Solution Approach 2:
The patent introduces a new dimensional organization of video data by spatial partitions, moving beyond traditional temporal frame-based organization. Each spatial partition is treated as an independent entity that can be selectively accessed and transmitted, enabling efficient region-of-interest extraction and adaptive streaming with reduced description overhead.
2Adaptability or versatility
If traditional encapsulation methods are used, then media content can be delivered, but dynamic track references and flexible reconstruction are not supported
Solution Approach 1:
The patent implements dynamic track references where the base track can reference different spatial track combinations at different time points. This dynamic referencing mechanism allows flexible reconstruction of video streams based on changing device capabilities or network conditions, while the standardized ISOBMFF container keeps the overall complexity manageable.
Solution Approach 2:
The base track acts as an intermediary that coordinates between the client and multiple spatial tracks. It contains reconstruction instructions that guide the client on how to combine spatial partitions, providing a standardized interface that simplifies the complexity of managing multiple independent video partitions.
3Adaptability or versatility
If spatial partitions are independently encoded, then region of interest extraction is enabled, but metadata overhead increases
Solution Approach 1:
Video content is segmented into independent spatial partitions (tiles, slices, or subpictures) that can be selectively transmitted. By only encoding and transmitting partitions containing regions of interest, the system reduces overall metadata overhead while maintaining the capability for selective region extraction.
Solution Approach 2:
Different spatial partitions can have different levels of encoding quality or be selectively included in the stream based on their importance. This local quality differentiation allows the system to focus metadata resources on critical regions while using minimal or no metadata for less important areas.
Data Source
AI summary
According to embodiments, the invention provides a method for encapsulating partitioned timed media data comprising timed samples, comprising in turn subsamples, the timed samples being grouped into groups, the method comprising: obtaining spatial tracks, each spatial track comprising at least one subsample of a first timed sample and one corresponding subsample of the other timed samples, the corresponding subsamples being located at the same spatial position in its own timed sample as the at least one subsample; creating a base track referencing at least some of the spatial tracks, the base track comprising reconstruction instructions, each of the reconstruction instructions being associated with a group of timed samples, enabling generating a portion of a bit-stream from subsamples of spatial tracks, that belong to a same group of timed samples; and independently encapsulating each of the tracks in a least one media file.


