Immersive Video Formatting for 6DoF Motion Parallax
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current omnidirectional video technologies only support motion parallax for seated viewers and lack immersion in six degrees of freedom, failing to provide natural video experiences for left/right and up/down movements in virtual reality environments.
Innovation Solution
The method involves acquiring basic and multiple view videos, generating residual video plus depth (RVD) and packed video plus depth (PVD) videos, and using metadata to decode and synthesize immersive videos that support motion parallax by packing non-overlapping video regions, allowing for six degrees of freedom in viewer movement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If omnidirectional video is used for 3DoF, then rotation movement is supported, but translation movement (left/right and up/down) is not supported
Solution Approach 1:
The video content is segmented into multiple view videos captured from different positions. Each view video covers a specific spatial region, and by combining these segmented views with their corresponding depth information, the system reconstructs complete 6DoF motion parallax that includes both rotation and translation movements.
Solution Approach 2:
The patent transitions from 3DoF (three rotational degrees of freedom) to 6DoF by adding three translational degrees of freedom. This is achieved by capturing videos at multiple positions along the translation axes and using depth information to synthesize views that maintain proper motion parallax for both rotational and translational movements.
2Adaptability or versatility
If multiple view videos are captured to support 6DoF, then motion parallax for translation is improved, but data amount and processing complexity increase
Solution Approach 1:
Depth information is pre-calculated and stored for each view video during the encoding phase. This preliminary computation of depth maps allows the decoding device to efficiently perform view synthesis and warping operations without requiring complex real-time calculations, thereby reducing processing complexity at playback.
Solution Approach 2:
Instead of capturing and transmitting all possible views for every potential viewer position, the system creates synthesized copies of video content through view synthesis algorithms. These synthesized views are generated by warping and combining the pre-captured multiple view videos according to the viewer's specific position and orientation, reducing the amount of data that needs to be transmitted and processed.
3Adaptability or versatility
If basic video and multiple view videos are combined to generate RVD and PVD, then 6DoF immersion is improved, but encoding and transmission complexity increases
Solution Approach 1:
The video data is segmented into residual components (RVD) that represent only the differences between the basic video and multiple view videos. This segmentation allows the system to transmit only the essential additional information needed for 6DoF reconstruction, rather than transmitting complete multiple view videos, thereby simplifying encoding and reducing data transmission requirements.
Data Source
AI summary
Disclosed herein is an immersive video formatting method and apparatus for supporting motion parallax, The immersive video formatting method includes acquiring a basic video at a basic position, acquiring a multiple view video at at least one position different from the basic position, acquiring at least one residual video plus depth (RVD) video using the basic video and the multiple view video, and generating at least one of a packed video plus depth (PVD) video or predetermined metadata using the acquired basic video and the at least one RVD video.


