Image Encoding And Decoding With Projection-Aware Block Partitioning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image processing systems struggle with the massive data generated for 360-degree images in virtual and augmented reality, necessitating improved performance in image encoding and decoding, particularly for 360-degree images.
Innovation Solution
A method for decoding 360-degree images involves generating a predicted image using syntax information, combining it with a residual image, and reconstructing the image in a specific projection format, utilizing projection formats like ERP, CMP, OHP, and ISP, and performing image expansion based on partitioning units to enhance compression performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If 360-degree images are captured and processed for virtual reality and augmented reality, then the realism and quality of media service are improved, but the amount of data generated increases massively and the performance of image processing systems becomes insufficient
Solution Approach 1:
The patent applies segmentation by dividing the 360-degree image into multiple projection formats (ERP, CMP, OHP, ISP) and processing different regions with different encoding strategies. The image is segmented into face regions and edge regions, with different prediction and transformation methods applied to each segment, enabling efficient processing of the massive data while maintaining service quality.
Solution Approach 2:
The patent transforms the 360-degree image from a spherical coordinate system to multiple 2D projection formats (equi-rectangular, cubic, octahedral, icosahedral projections). This dimensionality change allows the system to process the immersive image data using conventional 2D image processing techniques, improving processing performance while preserving the full 360-degree visual experience.
2Loss of substance
If image encoding and decoding is performed on 360-degree images, then compression is achieved, but the complexity of processing multiple projection formats increases system complexity
Solution Approach 1:
The patent implements a universal encoding framework that handles multiple projection formats (ERP, CMP, OHP, ISP) through a single integrated system. The same encoding and decoding apparatus can process any of these projection formats by selecting the appropriate projection type, reducing the need for separate specialized systems for each format while achieving effective compression.
Solution Approach 2:
The patent performs preliminary classification of the 360-degree image into different projection formats and identifies suitable reference regions before the main encoding process. By pre-determining the projection type and selecting appropriate reference pictures in advance, the system avoids complex real-time decisions during encoding, thereby reducing processing complexity while maintaining compression efficiency.
3Loss of substance
If image expansion is performed on partitioning units to generate predicted images, then compression performance is enhanced, but the processing time and computational load increase
Solution Approach 1:
The patent applies local quality by performing image expansion and prediction only on specific partitioning units where it provides the most benefit. Different regions of the 360-degree image are processed with different levels of expansion and prediction complexity based on their local characteristics, enhancing compression performance in critical areas while minimizing unnecessary processing elsewhere, thus reducing overall processing time.
Data Source
AI summary
Disclosed are methods and apparatuses for image data encoding/decoding. A method of decoding an image includes receiving a bitstream in which the image is encoded; obtaining index information for specifying a block division type of a current block in the image; and determining the block division type of the current block from a candidate group pre-defined in the decoding apparatus. The candidate group includes a plurality of candidate division types, including at least one of a non-division, a first quad-division, a second quad-division, a binary-division or a triple-division. The method also includes dividing the current block into a plurality of sub-blocks; and decoding each of the sub-blocks with reference to syntax information obtained from the bitstream.


