Video Coding Region Resampling for 360-Degree Panoramic Content
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Video coding systems face inefficiencies in encoding and decoding 360-degree panoramic content due to the equirectangular projection, which results in increased encoding and decoding complexity and decreased rate-distortion performance, particularly in areas near the nadir and zenith where pixels are stretched disproportionately.
Innovation Solution
A method and apparatus for video coding that involves decoding and reconstructing regions of a picture, including resampling and rearranging preliminary reconstructed regions for improved prediction, using a processor and memory to manage bitstream data and apply sampling ratios for efficient encoding and decoding, particularly in the vertical field of view regions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If equirectangular projection is used to map 360-degree panoramic content, then the full field-of-view can be represented in a rectangular format, but the nadir and zenith areas are stretched resulting in unnecessarily large number of pixels and increased encoding complexity
Solution Approach 1:
The panoramic image is divided into multiple regions with different sampling ratios. Regions near the nadir and zenith poles are identified and assigned a first sampling ratio, while other regions are assigned a second sampling ratio. This segmentation allows differential processing of areas with different distortion characteristics.
Solution Approach 2:
Different sampling ratios are applied to different regions of the panoramic image based on their local characteristics. The nadir and zenith regions, which suffer from excessive stretching in equirectangular projection, receive a lower sampling ratio to reduce pixel density, while other regions maintain the standard sampling ratio. This local adaptation optimizes the balance between coverage and complexity.
2Area of stationary object
If equirectangular projection is used to map 360-degree panoramic content, then the full field-of-view can be represented in a rectangular format, but the nadir and zenith areas are stretched resulting in unnecessarily large number of pixels and decreased rate-distortion performance
Solution Approach 1:
The panoramic image is divided into multiple regions with different sampling ratios. Regions near the nadir and zenith poles are identified and assigned a first sampling ratio, while other regions are assigned a second sampling ratio. This segmentation allows differential processing of areas with different distortion characteristics.
Solution Approach 2:
Different sampling ratios are applied to different regions of the panoramic image based on their local characteristics. The nadir and zenith regions, which suffer from excessive stretching in equirectangular projection, receive a lower sampling ratio to reduce pixel density, while other regions maintain the standard sampling ratio. This local adaptation optimizes the balance between coverage and complexity.
3Ease of manufacture
If uniform sampling is applied to all regions of the panoramic image, then the encoding process is simple, but the rate-distortion performance decreases due to excessive pixels in stretched areas
Solution Approach 1:
The panoramic image is divided into multiple regions with different sampling ratios. Regions near the nadir and zenith poles are identified and assigned a first sampling ratio, while other regions are assigned a second sampling ratio. This segmentation allows differential processing of areas with different distortion characteristics.
Solution Approach 2:
Different sampling ratios are applied to different regions of the panoramic image based on their local characteristics. The nadir and zenith regions, which suffer from excessive stretching in equirectangular projection, receive a lower sampling ratio to reduce pixel density, while other regions maintain the standard sampling ratio. This local adaptation optimizes the balance between coverage and complexity.
Data Source
AI summary
A method comprising: decoding, from a bitstream, a first encoded region of first picture into a first preliminary reconstructed region; forming a first reconstructed region from the first preliminary reconstructed region, wherein the forming comprises resampling and/or rearranging the first preliminary reconstructed region, wherein the rearranging comprises relocating, rotating and/or mirroring; and decoding at least a second region, wherein the first reconstructed region is used as a reference for prediction in decoding the at least second region and the second region either belongs to a second picture and is spatially collocated with the first reconstructed region or belongs to the first picture.


