HEVC Layered Media Encapsulation Metadata Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for encapsulating and streaming video data, particularly in the context of HTTP and RTP protocols, face inefficiencies when dealing with multi-layer HEVC video streams, especially in accessing and combining spatial tiles, leading to complex metadata organization and potential duplication of data.
Innovation Solution
A method for encapsulating multi-layer partitioned timed media data that simplifies the process by creating tracks with descriptive metadata organized into main descriptive boxes, optionally including sub-boxes for layer organization, and using specific parameters to configure decoding devices, ensuring efficient parsing and streaming of spatial tiles.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If multi-layer HEVC video streams are encapsulated using existing methods, then the video data can be transmitted and stored, but the metadata organization becomes complex and data duplication occurs
Solution Approach 1:
The patent segments the metadata organization by creating separate main descriptive boxes for each layer, with sub-boxes containing layer-specific information. This segmentation prevents metadata mixing and reduces organizational complexity when dealing with multi-layer HEVC video streams
Solution Approach 2:
The patent extracts layer-specific descriptive information into separate sub-boxes within the main descriptive boxes. This extraction method isolates redundant metadata for each layer, preventing duplication while maintaining organized structure for efficient parsing
2Productivity
If spatial tiles are accessed and combined in multi-layer HEVC streams, then user-selected regions of interest can be streamed efficiently, but the parsing and combination process becomes complex
Solution Approach 1:
The patent segments spatial video data into independently decodable tiles within each layer, with clear metadata boundaries. This segmentation enables client devices to parse and combine only the specific tiles needed for regions of interest, improving streaming efficiency while maintaining manageable parsing complexity through structured organization
Solution Approach 2:
The patent introduces standardized main descriptive boxes as intermediaries between the encoded video data and the client device parsing process. These boxes provide a structured interface that simplifies the combination of spatial tiles from different layers by establishing consistent metadata formats that reduce parsing complexity
3Quantity of substance
If complete video frames are transmitted, then full video quality is maintained, but unnecessary data is transmitted when only specific regions are needed
Solution Approach 1:
The patent segments video frames into spatial tiles that can be independently decoded and transmitted. This allows the server to transmit only the specific tiles corresponding to user-selected regions of interest rather than complete frames, reducing data transmission volume while maintaining video quality for the transmitted regions through the structured metadata organization
Data Source
AI summary
The invention relates to a method for encapsulating multi-layer partitioned timed media data in a server, the multi-layer partitioned timed media data comprising timed samples, each timed sample being encoded into a first layer and at least one second layer, at least one timed sample comprising at least one subsample, each subsample being encoded into the first layer or the at least one second layer. The method comprises: obtaining at least one subsample from at least one of the timed samples; creating a track comprising the at least one obtained subsample; and generating a descriptive metadata associated with the created track, the descriptive metadata being organized into one main descriptive box per track, the descriptive information about the organization of the different layers being included into one or more sub-boxes, wherein at most one main descriptive box comprises the one or more sub-boxes.


