Motion Vector Transformation for Omnidirectional Video Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for encoding and decoding large field of view videos, such as those with 360° content, face challenges due to geometry distortions and unsuitable motion vector prediction when projecting 3D surfaces onto 2D pictures, leading to increased encoding costs and inefficiencies.
Innovation Solution
A method and apparatus that determine and transform motion vectors based on a projection function to adapt to the 2D representation, allowing for effective encoding and decoding of large field of view videos by clipping and transforming motion vectors to improve prediction accuracy and reduce encoding costs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If motion vectors are predicted using classical video coding methods from neighboring blocks, then encoding process is simple, but prediction accuracy deteriorates due to geometry distortions from projection
Solution Approach 1:
The patent transforms motion vectors by applying the inverse of the projection function to compensate for geometric distortions. This parameter transformation adjusts the motion vector components based on the projection geometry, improving prediction accuracy without fundamentally changing the encoding architecture
Solution Approach 2:
The patent introduces a new dimension of processing by transforming motion vectors from the 2D projected space back to the 3D spherical space for prediction, then transforming back. This dimensional transition allows prediction to occur in a space where motion is more uniform, resolving the contradiction between accuracy and complexity
2Adaptability or versatility
If 3D surface is projected onto 2D picture using standard projection functions, then content can be displayed on conventional devices, but geometry distortions occur causing non-uniform motion representation
Solution Approach 1:
The patent applies parameter transformations to motion vectors based on the projection function used. By transforming motion vectors according to the specific projection geometry (e.g., equirectangular, cubic), the patent compensates for distortion effects while maintaining compatibility with standard display devices
Solution Approach 2:
The patent introduces an intermediary transformation step that converts motion vectors between the projected 2D space and the original 3D space. This intermediary process allows the system to maintain both compatibility with conventional devices and uniformity in motion representation by operating in the appropriate coordinate system for each task
3Productivity
If motion vectors are transformed based on projection function to improve prediction accuracy, then encoding efficiency improves, but computational complexity increases
Solution Approach 1:
The patent transforms motion vector parameters using the projection function, which improves prediction accuracy and encoding efficiency. The transformation is applied selectively based on the projection type and block characteristics, balancing the computational overhead with the encoding benefits
Solution Approach 2:
The patent applies motion vector transformation selectively rather than uniformly across all blocks. By transforming only where necessary (e.g., for inter-predicted blocks with significant distortion), the system achieves encoding efficiency improvements while limiting the increase in computational complexity
Data Source
AI summary
A method for decoding a large field of view video is disclosed. At least one picture of said large field of view video is represented as a 3D surface projected onto at least one 2D picture using a projection function. The method comprises, for at least one current block of said 2D picture: —determining whether an absolute value of at least one component of a motion vector (d V) associated with another block of said 2D picture satisfies a condition; —transforming, based on said determining, said motion vector (d V) into a current motion vector (d P) associated with said current block responsive to said projection function; and —decoding said current block using said current motion vector (d P).


