Sub-Picture Layer Mapping for Region-of-Interest Video Scalability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding standards face complexity in accessing spatial regions due to the mapping of encoded elements to spatial subdivision structures, which complicates decoding processes, especially in immersive and 3D video applications.
Innovation Solution
The implementation of sub-pictures as logical units within video frames, with metadata to specify relationships between sub-pictures and layers, allowing for efficient indication and interpretation in the coded video bitstream, and high-level syntax signaling for sub-picture support.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If spatial subdivision structures (slices, tiles, subpictures) are specified in high-level syntax with low-level mapping, then spatial independence and parallel processing are improved, but decoder complexity and parsing difficulty increase
Solution Approach 1:
The patent segments the video frame into independent subpictures that can be decoded in parallel. Each subpicture is treated as a separate decoding unit with its own parameter set, enabling independent processing while maintaining overall frame coherence. This segmentation resolves the contradiction by allowing parallel decoding (improving productivity) while keeping each subpicture's parsing simple (managing device complexity).
Solution Approach 2:
The patent introduces a new organizational dimension by mapping subpictures to layer structures in the bitstream. Instead of only spatial subdivision, it adds a hierarchical layer dimension where subpictures can be grouped and accessed systematically. This dimensional addition simplifies decoder access patterns while preserving parallel processing capabilities.
2Adaptability or versatility
If layer-based coding structures are used for spatial and temporal scalability, then video quality and adaptability are improved, but syntax complexity and processing overhead increase
Solution Approach 1:
The patent makes the layer structure multi-functional by using it simultaneously for spatial subdivision (subpictures), temporal scalability (different frame rates), and quality differentiation (different resolutions). This universal layer structure reduces the need for separate syntax mechanisms for each function, thereby improving adaptability while managing syntax complexity through consolidation.
Solution Approach 2:
The patent enables scalable parameter changes across layers, where base layers contain essential video information and enhancement layers add incremental quality or resolution. Decoders can selectively process layers based on available bandwidth and device capabilities, achieving adaptability while the progressive parameter changes keep syntax processing manageable through hierarchical refinement.
Data Source
Figure 1
Figure 2(a)~2(d)
Figure 3~4
AI summary
The present disclosure relates to leveraging media coding processes and syntax elements, which ordinarily would support layer-based coding, to representations of sub-pictures within video. A sub-picture relates to a spatial region of a video that is organized into a logical unit separate from other region(s) of the videos' content. Sub-pictures can be used, for example, to support region of interest scalability with each sub-picture corresponding to a different region of interest. Techniques for indicating, in a coded video bitstream, that spatial layers are used as sub-pictures and providing metadata to specify and interpret the relationships between sub-pictures and their correspondence to a final reconstructed picture are also specified. Moreover, metadata may be revised before delivery to consuming terminals based on information developed about the consuming terminal's capabilities and processing environments.