Tree-Based Block Division for 360-Degree Image Encoding and Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image processing systems struggle with the massive data generated by 360-degree images for virtual and augmented reality, necessitating improved performance in image encoding and decoding, particularly for high-resolution and high-quality images.
Innovation Solution
A method for decoding 360-degree images involves generating a predicted image using syntax information, combining it with a residual image, and reconstructing the image in a projection format, utilizing projection formats like ERP, CMP, OHP, and ISP, and performing image expansion based on partitioning units to enhance compression performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If 360-degree images are captured with multiple cameras for virtual reality and augmented reality, then image quality and resolution are improved, but the amount of data generated increases massively
Solution Approach 1:
The patent divides the 360-degree image into multiple projection formats (ERP, CMP, OHP, ISP) and processes different regions with different encoding strategies. The image is segmented into face regions, edge regions, and corner regions, each handled differently to optimize compression while maintaining quality.
Solution Approach 2:
Different encoding parameters and compression ratios are applied to different regions of the 360-degree image. High-quality encoding is applied to face regions where visual quality is critical, while lower-quality encoding is applied to edge and corner regions where quality is less critical, thus reducing overall data amount while maintaining perceived quality.
2Measurement precision
If high-resolution and high-quality images are processed, then image quality is improved, but the performance of the image processing system becomes insufficient
Solution Approach 1:
The encoding process is divided into multiple independent stages: projection format selection, region classification, differential encoding, and reconstruction. This segmentation allows parallel processing of different regions and formats, improving overall processing performance.
Solution Approach 2:
The patent applies full high-quality encoding only to critical face regions, while using compressed encoding for less critical edge and corner regions. This partial application of high-quality encoding maintains visual quality where needed while improving processing performance overall.
3Ease of manufacture
If traditional image encoding methods are used for 360-degree images, then implementation is simple, but compression performance is insufficient
Solution Approach 1:
The patent creates a universal encoding framework that handles multiple projection formats (ERP, CMP, OHP, ISP) and multiple region types within a single encoding process. This multi-functional approach improves compression performance without requiring separate encoding systems for each format.
Solution Approach 2:
The patent dynamically changes encoding parameters based on the projection format and region type. Different quantization parameters, transformation blocks, and prediction modes are applied depending on the specific format and region, optimizing compression performance for each case.
Data Source
AI summary
Disclosed are methods and apparatuses for image data encoding/decoding. A method of decoding an image includes receiving a bitstream in which the image is encoded; obtaining index information for specifying a block division type of a current block in the image; and determining the block division type of the current block from a candidate group pre-defined in the decoding apparatus. The candidate group includes a plurality of candidate division types, including at least one of a non-division, a first quad-division, a second quad-division, a binary-division or a triple-division. The method also includes dividing the current block into a plurality of sub-blocks; and decoding each of the sub-blocks with reference to syntax information obtained from the bitstream.


