360-Degree Image Reconstruction Across Multi-Projection Formats
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image processing systems struggle with the massive data generated for 360-degree images in virtual and augmented reality, requiring improved performance in encoding and decoding methods.
Innovation Solution
A method for encoding and decoding 360-degree images that includes generating a predicted image using syntax information, combining it with a residual image, and reconstructing the image in various projection formats, such as Equi-Rectangular, CubeMap, OctaHedron, and IcoSahedral, to enhance compression performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multi-view images captured with multiple cameras are processed for 360-degree images, then the visual quality and immersion for virtual reality and augmented reality are improved, but the amount of data generated increases massively and the image processing system performance becomes insufficient
Solution Approach 1:
The patent divides the 360-degree image into multiple projection formats (ERP, CMP, OHP, ISP) and processes different regions with different encoding strategies. The image is segmented into face regions, boundary regions, and transition regions, each handled separately to optimize compression while maintaining visual quality in critical areas.
Solution Approach 2:
Different encoding parameters and compression ratios are applied to different regions of the 360-degree image based on their importance. Face regions (high visual importance) use higher quality encoding, while boundary and transition regions use lower quality encoding, optimizing the overall data-to-quality ratio.
2Ease of manufacture
If conventional image encoding methods are used for 360-degree images, then the encoding process is simple, but the compression performance is insufficient for large data volumes
Solution Approach 1:
The patent changes multiple encoding parameters including projection format selection, region-wise packing configuration, and quantization parameters based on the importance of different image regions. These parameter changes enable better compression performance while maintaining a systematic encoding approach.
Solution Approach 2:
The patent introduces region-wise packing in the spatial domain and combines it with frequency-domain transformation (DCT). This multi-dimensional approach to encoding allows for more efficient compression by exploiting both spatial and frequency characteristics of different image regions.
3Adaptability or versatility
If the 360-degree image is projected into various projection formats, then the adaptability for different VR/AR applications is improved, but the complexity of the encoding and decoding process increases
Solution Approach 1:
The patent creates a universal encoding framework that can output multiple projection formats (ERP, CMP, OHP, ISP) from a single encoding process. The region-wise packing methodology is format-agnostic and can be applied to any projection type, making the system versatile while maintaining a consistent processing approach.
Solution Approach 2:
The patent performs region identification and classification before the actual encoding process. By pre-segmenting the image into face, boundary, and transition regions and determining their importance, the subsequent encoding process becomes more straightforward and less complex.
Data Source
AI summary
Disclosed are methods and apparatuses for decoding an image. A method includes receiving a bitstream obtained by encoding the image; dividing a first coding block into a plurality of second coding blocks; generating a prediction block of a second coding block based on syntax information obtained from the bitstream; and reconstructing the second coding block based on the prediction block and a residual block of the second coding block, the residual block being obtained by performing a dequantization and an inverse-transform on quantized transform coefficients from the bitstream. The first coding block has a recursive division structure. The first coding block is divided based on at least one of a quad tree division, a binary tree division or a triple tree division.


