Multi-Viewpoint Image Reprojection for Occlusion-Free Immersive Video
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for generating immersive video content in virtual reality (VR) face challenges such as time-consuming manual editing to fill in occluded parts of subjects, impractical multiple takes, and computational overheads, leading to gaps and reduced immersion when viewers change viewpoints.
Innovation Solution
Capture multiple images of a subject from different viewpoints using a primary and additional cameras, re-project and combine these images to generate a composite image that fills in occluded parts without requiring complete 3D reconstruction or excessive rendering, using a simple mesh and stereoscopic texture projection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a stereoscopic camera is used to capture video of a performer, then the captured video can be projected onto a mesh for immersive viewing, but occluded parts of the performer appear as gaps when viewers change viewpoints
Solution Approach 1:
Multiple images are captured in advance from different camera positions during the filming process. These pre-captured images from alternative viewpoints are stored and later used to fill in occluded regions when the viewer's viewpoint changes, eliminating the need for real-time generation or manual editing.
Solution Approach 2:
Image data from alternative viewpoints is copied and re-projected to fill occluded regions. Instead of generating new image data or manually creating content, the system replicates existing image data from different camera angles and integrates it into the immersive video stream based on the viewer's current viewpoint.
2Manufacturing precision
If manual painting is used to fill occluded parts, then the performer appears more realistic, but the process is time-consuming and reduces productivity
Solution Approach 1:
Instead of manual painting, the system automatically copies image data from pre-captured alternative viewpoint images to fill occluded regions. This automated copying process maintains high image realism while eliminating the time-consuming manual editing process entirely.
Solution Approach 2:
The system performs self-service by automatically selecting and integrating appropriate image data from multiple captured viewpoints based on the viewer's current position. The computational system autonomously determines which pre-captured images to use and blends them seamlessly without human intervention.
3Loss of information
If multiple takes are captured to remove occluding objects, then occluded parts can be supplemented, but this interferes with the performer's performance and is impractical
Solution Approach 1:
Multiple images from different angles are captured in a single continuous take during normal performance. This preliminary capture of diverse viewpoint data eliminates the need for multiple takes or performance interruptions, as all necessary image data is available for later integration based on viewer position.
4Adaptability or versatility
If complete 3D reconstruction is performed to fill occluded regions, then viewpoint flexibility is improved, but computational overhead becomes excessive
Solution Approach 1:
Instead of performing computationally intensive complete 3D reconstruction, the system copies and re-projects existing 2D image data from alternative viewpoints. This approach provides sufficient viewpoint flexibility for immersive viewing while avoiding the excessive computational overhead of full 3D modeling and rendering.
Data Source
Figure 1
Figure 2~3
Figure 4~5B
AI summary
A method of generating an image of a subject (702) in a scene comprises obtaining a first image and a second image of a subject in a scene, each image corresponding to a different respective viewpoint of the subject, each image being captured by a different respective camera (708A, 704A), wherein at least some of the subject is occluded in the first image and not the second image by virtue of the different viewpoints, obtaining camera pose data indicating a pose of a camera for each image, re-projecting (710), based on the difference in camera poses associated with each image, at least a portion of the second image to correspond to the viewpoint from which the first image was captured, and combining the re-projected portion of the second image with at least some of the first image so as to generate a composite image of the subject from the viewpoint of the first image, the re-projected portion of the second image providing image data for at least some of the occluded part or parts of the subject in the first image.