Homographic Matrix Inter-View Prediction for 360 Video
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current encoding techniques for 360° videos, such as those used in virtual reality headsets, are limited to 3 degrees of freedom, allowing only rotational movements, and fail to accurately render translational movements, leading to user discomfort due to parasitic pixel display issues.
Innovation Solution
A method that improves inter-view prediction by using a homographic matrix to compensate for geometric distortions between neighboring views in 360° video encoding and decoding, allowing for the creation of a new reference image that enhances prediction accuracy in overlap zones, thereby improving compression performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional 2D video encoding is used for 360° videos, then the encoding process is simple, but the rendering is limited to 3 degrees of freedom (rotational movements only) and cannot accurately render translational movements
Solution Approach 1:
The patent divides the 360° video content into multiple overlapping views captured by different cameras. Each view is processed separately with dedicated homographic transformation, allowing independent optimization of each view's rendering capabilities while maintaining overall system functionality.
Solution Approach 2:
The patent introduces homographic matrices as intermediary mathematical transformations between different camera views. These matrices serve as mediators that enable accurate translational movement rendering by mapping pixels from one view to another, bridging the gap between simple 2D encoding and complex 6DoF rendering requirements.
2Productivity
If multi-view encoders like MV-HEVC or 3D-HEVC are used, then inter-view prediction can be exploited, but the encoders are designed for converging views with camera centers outside the scene, making them suboptimal for divergent 360° camera views
Solution Approach 1:
The patent changes the fundamental parameter of camera geometry modeling from converging views (standard multi-view) to divergent views (360° panorama). By modifying the geometric model and using homographic transformations specifically tailored for divergent camera arrangements, the system achieves both compression efficiency and compatibility with 360° camera configurations.
Solution Approach 2:
The patent implements dynamic homographic transformation that adapts to each specific pair of neighboring views. Rather than using a fixed transformation model, the system dynamically calculates and applies appropriate homographic matrices based on the specific geometric relationship between adjacent camera views, optimizing compression for each local region.
3Productivity
If overlapping areas between neighboring views are used for inter-view prediction, then compression can be improved, but the pixels in overlapping areas have undergone geometric transformations, making simple pixel copying inefficient
Solution Approach 1:
The patent uses homographic matrices as intermediary transformations to accurately map pixels from source views to target views in overlapping regions. This intermediary step corrects for geometric distortions and perspective differences, enabling high-precision inter-view prediction that maintains both compression efficiency and prediction accuracy.
Solution Approach 2:
The patent replaces simple mechanical pixel copying with mathematical homographic transformation. Instead of directly copying pixels from one view to another (which fails due to geometric differences), the system substitutes a mathematical transformation model that accounts for perspective, rotation, and scaling differences between divergent camera views.
Data Source
Figure 1
Figure 2~3
Figure 4
AI summary
The invention relates to a method and device for decoding an encoded data signal representing a multi-view video sequence representative of an omnidirectional video, the multi-view video sequence comprising at least a first view and a second view. Parameters allowing a homographic matrix to be obtained (61), representing the transformation of a plane of the second view into a plane of the second view, are read (60) out from the signal. An image of the second view comprises a so-called active zone comprising pixels which, when they are projected via the homographic matrix onto an image of the first view, are included in the image of the first view. An image of the second view is decoded (62) by generation (620) of a reference image comprising pixel values determined from previously reconstructed pixels of an image of the first view and from the homographic matrix and, for at least one block of the image of the second view, the reference image generated is included in the list of reference images when the block belongs (622) to the active zone. The block is reconstructed (625) from a reference image indicated by an index read out (621) from the data signal.