Sub-Picture Layer Mapping for Region-of-Interest Video Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding standards face complexity in mapping encoded elements to spatial subdivision structures like slices, tiles, and sub-pictures, leading to increased decoder complexity and inefficiency in accessing spatial regions.
Innovation Solution
The implementation of sub-pictures as logical units within media coding, utilizing metadata to define and map spatial regions to coding layers, allowing for region-of-interest scalability and efficient decoding by indicating sub-pictures and their relationships to layers, even in systems that do not natively support sub-picture definitions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If spatial subdivision structures (slices, tiles, sub-pictures) are specified in high-level syntax with low-level mapping, then coding flexibility and error resiliency are improved, but decoder complexity and difficulty in accessing spatial regions increase
Solution Approach 1:
The patent introduces an intermediary mapping structure that connects high-level spatial subdivision definitions with low-level coded data. This intermediary layer provides a systematic translation mechanism that simplifies decoder access to spatial regions while preserving the flexibility of high-level syntax specifications for slices, tiles, and sub-pictures.
Solution Approach 2:
The patent segments the complex mapping process into distinct hierarchical levels: high-level spatial subdivision structure definition, intermediary mapping relationship specification, and low-level coded data organization. This segmentation allows each level to be processed independently, reducing overall decoder complexity while maintaining coding flexibility.
2Loss of information
If all coded video data is decoded, then complete video content is available, but bitrate and processing resources are wasted when only certain spatial regions are needed
Solution Approach 1:
The patent enables extraction of only the necessary spatial regions (sub-pictures, tiles, or slices) from the coded video bitstream based on decoder needs. By organizing coded data with explicit spatial region identifiers and mapping relationships, decoders can selectively decode and extract only the required portions, avoiding unnecessary processing power consumption and reducing effective bitrate for partial content retrieval.
3Productivity
If traditional slicing and tiling methods are used, then parallel decoding and error resiliency are improved, but access efficiency to specific spatial regions deteriorates
Solution Approach 1:
The patent adds a new dimensional layer to traditional slicing and tiling by introducing explicit sub-picture concepts with high-level syntax definitions. This additional dimension organizes spatial regions in a more accessible hierarchy, allowing decoders to efficiently access specific spatial regions while maintaining the parallel decoding capabilities provided by slices and tiles. The new dimension provides direct pointers to spatial regions without requiring complex low-level mapping calculations.
Data Source
AI summary
The present disclosure relates to leveraging media coding processes and syntax elements, which ordinarily would support layer-based coding, to representations of sub-pictures within video. A sub-picture relates to a spatial region of a video that is organized into a logical unit separate from other region(s) of the videos' content. Sub-pictures can be used, for example, to support region of interest scalability with each sub-picture corresponding to a different region of interest. Techniques for indicating, in a coded video bitstream, that spatial layers are used as sub-pictures and providing metadata to specify and interpret the relationships between sub-pictures and their correspondence to a final reconstructed picture are also specified. Moreover, metadata may be revised before delivery to consuming terminals based on information developed about the consuming terminal's capabilities and processing environments.


