Sub-Picture Layer Mapping for Region-of-Interest Video Scalability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video coding standards face complexity in accessing spatial regions due to the mapping of encoded elements to spatial subdivision structures, which complicates decoding processes, especially in immersive and 3D video applications.

Innovation Solution

The implementation of sub-pictures as logical units within video frames, with metadata to specify relationships between sub-pictures and layers, allowing for efficient indication and interpretation in the coded video bitstream, and high-level syntax signaling for sub-picture support.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If spatial subdivision structures (slices, tiles, subpictures) are specified in high-level syntax with low-level mapping, then spatial independence and parallel processing are improved, but decoder complexity and parsing difficulty increase

Engineering Contradiction:
Improveparallel decoding efficiencyVSAvoiddecoder parsing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the video frame into independent subpictures that can be decoded in parallel. Each subpicture is treated as a separate decoding unit with its own parameter set, enabling independent processing while maintaining overall frame coherence. This segmentation resolves the contradiction by allowing parallel decoding (improving productivity) while keeping each subpicture's parsing simple (managing device complexity).

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new organizational dimension by mapping subpictures to layer structures in the bitstream. Instead of only spatial subdivision, it adds a hierarchical layer dimension where subpictures can be grouped and accessed systematically. This dimensional addition simplifies decoder access patterns while preserving parallel processing capabilities.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If layer-based coding structures are used for spatial and temporal scalability, then video quality and adaptability are improved, but syntax complexity and processing overhead increase

Engineering Contradiction:
Improvespatial and temporal scalabilityVSAvoidsyntax processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent makes the layer structure multi-functional by using it simultaneously for spatial subdivision (subpictures), temporal scalability (different frame rates), and quality differentiation (different resolutions). This universal layer structure reduces the need for separate syntax mechanisms for each function, thereby improving adaptability while managing syntax complexity through consolidation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent enables scalable parameter changes across layers, where base layers contain essential video information and enhancement layers add incremental quality or resolution. Decoders can selectively process layers based on available bandwidth and device capabilities, achieving adaptability while the progressive parameter changes keep syntax processing manageable through hierarchical refinement.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4668750A1Layer based methods for sub-picture support and region of interest scalability
Publication Date: 2025.12.24 APPLE INC
  • EP4668750A1 patent drawingFigure 1
  • EP4668750A1 patent drawingFigure 2(a)~2(d)
  • EP4668750A1 patent drawingFigure 3~4

AI summary

The present disclosure relates to leveraging media coding processes and syntax elements, which ordinarily would support layer-based coding, to representations of sub-pictures within video. A sub-picture relates to a spatial region of a video that is organized into a logical unit separate from other region(s) of the videos' content. Sub-pictures can be used, for example, to support region of interest scalability with each sub-picture corresponding to a different region of interest. Techniques for indicating, in a coded video bitstream, that spatial layers are used as sub-pictures and providing metadata to specify and interpret the relationships between sub-pictures and their correspondence to a final reconstructed picture are also specified. Moreover, metadata may be revised before delivery to consuming terminals based on information developed about the consuming terminal's capabilities and processing environments.