HEVC Layered Media Encapsulation Metadata Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for encapsulating and streaming video data, particularly in the context of HTTP and RTP protocols, face inefficiencies when dealing with multi-layer HEVC video streams, especially in accessing and combining spatial tiles, leading to complex metadata organization and potential duplication of data.

Innovation Solution

A method for encapsulating multi-layer partitioned timed media data that simplifies the process by creating tracks with descriptive metadata organized into main descriptive boxes, optionally including sub-boxes for layer organization, and using specific parameters to configure decoding devices, ensuring efficient parsing and streaming of spatial tiles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If multi-layer HEVC video streams are encapsulated using existing methods, then the video data can be transmitted and stored, but the metadata organization becomes complex and data duplication occurs

Engineering Contradiction:
Improveencapsulation process simplicityVSAvoidmetadata organization complexity
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The patent segments the metadata organization by creating separate main descriptive boxes for each layer, with sub-boxes containing layer-specific information. This segmentation prevents metadata mixing and reduces organizational complexity when dealing with multi-layer HEVC video streams

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts layer-specific descriptive information into separate sub-boxes within the main descriptive boxes. This extraction method isolates redundant metadata for each layer, preventing duplication while maintaining organized structure for efficient parsing

Inventive Principle:
Principle #2Taking out (Extraction)

2Productivity

If spatial tiles are accessed and combined in multi-layer HEVC streams, then user-selected regions of interest can be streamed efficiently, but the parsing and combination process becomes complex

Engineering Contradiction:
Improvestreaming efficiencyVSAvoidparsing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments spatial video data into independently decodable tiles within each layer, with clear metadata boundaries. This segmentation enables client devices to parse and combine only the specific tiles needed for regions of interest, improving streaming efficiency while maintaining manageable parsing complexity through structured organization

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces standardized main descriptive boxes as intermediaries between the encoded video data and the client device parsing process. These boxes provide a structured interface that simplifies the combination of spatial tiles from different layers by establishing consistent metadata formats that reduce parsing complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

3Quantity of substance

If complete video frames are transmitted, then full video quality is maintained, but unnecessary data is transmitted when only specific regions are needed

Engineering Contradiction:
Improvedata transmission volumeVSAvoidvideo quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent segments video frames into spatial tiles that can be independently decoded and transmitted. This allows the server to transmit only the specific tiles corresponding to user-selected regions of interest rather than complete frames, reducing data transmission volume while maintaining video quality for the transmitted regions through the structured metadata organization

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11005904B2Method, device, and computer program for encapsulating HEVC layered media data
Publication Date: 2021.05.11 CANON KK
  • US11005904B2 patent drawing
  • US11005904B2 patent drawing
  • US11005904B2 patent drawing

AI summary

The invention relates to a method for encapsulating multi-layer partitioned timed media data in a server, the multi-layer partitioned timed media data comprising timed samples, each timed sample being encoded into a first layer and at least one second layer, at least one timed sample comprising at least one subsample, each subsample being encoded into the first layer or the at least one second layer. The method comprises: obtaining at least one subsample from at least one of the timed samples; creating a track comprising the at least one obtained subsample; and generating a descriptive metadata associated with the created track, the descriptive metadata being organized into one main descriptive box per track, the descriptive information about the organization of the different layers being included into one or more sub-boxes, wherein at most one main descriptive box comprises the one or more sub-boxes.