Canvas Size Scalable Video Coding for ROI-Based HDR Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video coding standards struggle to efficiently support scalable distribution of high dynamic range (HDR) content across different resolutions, such as 4K and 8K, while maintaining compatibility with existing playback devices and minimizing coding efficiency loss.
Innovation Solution
Implement canvas size scalability by enhancing HEVC tile concepts with layer-adaptive slice addressing, cross-boundary prediction and entropy coding, and post-filtering to enable independent decoding of regions of interest (ROIs) within video frames, using SEI messaging for in-loop filtering adjustments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If video content is distributed in multiple resolutions (4K and 8K) to support future playback devices, then forward compatibility and adaptability are improved, but device complexity and coding complexity increase
Solution Approach 1:
The video content is segmented into multiple resolution versions (4K and 8K) with different canvas sizes, allowing playback devices to select the appropriate resolution. The encoder processes each resolution separately with optimized coding parameters, reducing the overall coding complexity compared to processing all resolutions uniformly.
Solution Approach 2:
The patent introduces canvas size as an additional dimension for scalability beyond traditional resolution scaling. By defining different canvas sizes (e.g., 3840×2160 for 4K, 7680×4320 for 8K) and using layer-adaptive slice addressing, the system enables efficient multi-resolution distribution without proportionally increasing coding complexity across all dimensions.
2Productivity
If independent decoding of regions of interest is enabled through tile concepts and cross-boundary prediction, then decoding efficiency and adaptability are improved, but device complexity increases
Solution Approach 1:
The video frame is divided into tiles with independent decoding capabilities. Each tile can be decoded independently or in combination with adjacent tiles, allowing flexible region-of-interest decoding. This segmentation enables playback devices to decode only necessary regions, improving decoding efficiency while managing device complexity through modular processing.
Solution Approach 2:
The tile-based structure with cross-boundary prediction enables multiple decoding modes (independent tile decoding, combined tile decoding, selective ROI decoding) within a single framework. This multi-functional approach improves decoding efficiency for different scenarios without requiring separate decoding systems, thereby controlling device complexity.
3Adaptability or versatility
If canvas size scalability is implemented with layer-adaptive slice addressing, then adaptability across resolutions is improved, but coding complexity increases
Solution Approach 1:
The patent implements layer-adaptive slice addressing where slice boundaries and addressing modes are dynamically adjusted based on the target resolution layer. For 4K content, one addressing scheme is used, while for 8K content, the addressing is adapted to the larger canvas size. This dynamic adaptation enables scalable distribution without requiring completely separate coding systems for each resolution, thereby controlling coding complexity.
Solution Approach 2:
The system changes key parameters (canvas size, slice boundaries, addressing offsets) based on the target resolution layer. By systematically adjusting these parameters rather than redesigning the entire coding structure, the patent achieves canvas size scalability while managing coding complexity through parameter-based adaptation.
4Manufacturing precision
If post-filtering and in-loop filtering adjustments are applied to maintain visual quality, then image quality is improved, but coding complexity and processing time increase
Solution Approach 1:
Filtering parameters and post-filtering adjustments are determined and configured in advance during the encoding stage for each resolution layer. This preliminary configuration allows playback devices to apply appropriate filtering without real-time complex calculations, maintaining visual quality while reducing processing complexity during decoding.
Solution Approach 2:
The patent uses filtering as an intermediary process between decoding and final image output. By introducing configurable post-filtering and in-loop filtering adjustments, the system mediates between the decoded image data and the final visual output, improving visual quality while allowing flexible control over processing complexity based on device capabilities.
Data Source
AI summary
Methods and systems for canvas size scalability across the same or different bitstream layers of a video coded bitstream are described. Offset parameters for a conformance window, a reference region of interest (ROI) in a reference layer, and a current ROI in a current layer are received. The width and height of a current ROI and a reference ROI are computed based on the offset parameters and they are used to generate a width and height scaling factor to be used by a reference picture resampling unit to generate an output picture based on the current ROI and the reference ROI.


