Neural Network Stereo Pair Generation from Single 2D Video
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for generating 3D stereoscopic effects from 2D videos require explicit 3D reconstruction of individuals, which is challenging and inefficient, especially in real-time applications like virtual and augmented reality.
Innovation Solution
A method using multiple neural networks to process 2D videos of humans in motion, generating stereo pairs of images without explicit 3D reconstruction by creating a mapping between RGB pixels and a 3D surface-based model, refining the model, and applying it to generate complete textures and views from different virtual viewpoints.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If explicit 3D reconstruction is used to generate stereoscopic effects from 2D videos, then the 3D effect quality can be improved, but the processing time and computational complexity increase significantly
Solution Approach 1:
The patent extracts and removes the explicit 3D reconstruction step from the traditional pipeline. Instead of reconstructing 3D models first and then generating stereoscopic views, the method directly transforms 2D video frames into stereoscopic pairs using neural networks, eliminating the time-consuming 3D reconstruction phase while maintaining 3D effect quality
Solution Approach 2:
The patent introduces neural network models as intermediary components that directly map 2D images to stereoscopic pairs. These neural networks learn the transformation relationship between 2D views and 3D stereoscopic outputs, serving as a mediator that bypasses the need for explicit 3D reconstruction while preserving depth information and 3D effect quality
2Manufacturing precision
If explicit 3D reconstruction is performed, then accurate 3D representation can be achieved, but the system complexity and computational resources required increase
Solution Approach 1:
The patent removes the complex explicit 3D reconstruction module from the system. By directly transforming 2D images to stereoscopic pairs through neural networks, the system achieves accurate 3D representation without the intermediate complex 3D model reconstruction steps, thereby reducing overall system complexity and computational resource requirements
Solution Approach 2:
The patent replaces the traditional mechanical 3D reconstruction pipeline with a neural network-based direct transformation approach. Instead of using complex geometric algorithms and 3D modeling techniques, the system employs trained neural networks to directly generate stereoscopic pairs from 2D images, simplifying the system while maintaining 3D representation accuracy
3Manufacturing precision
If traditional stereoscopic methods are used, then 3D effect can be generated, but the process requires multiple offset images to be processed separately
Solution Approach 1:
The patent merges the processing of multiple offset images into a single unified neural network transformation process. Instead of separately processing left and right offset images through multiple independent pipelines, the system uses a single neural network model that takes one 2D image and directly outputs the complete stereoscopic pair, significantly improving processing efficiency and productivity
Data Source
AI summary
In one embodiment, one or more computing systems may receive an image comprising pixels corresponding to a person captured by a camera from a camera viewpoint. The one or more computing systems may generate, based on the image, (1) a first body-surface mapping associated with the camera viewpoint, the first body-surface mapping indicates, for each of the pixels corresponding to the person, a corresponding location on a surface of a human body, and (2) a second body-surface mapping associated with a first virtual viewpoint different from the camera viewpoint. The one or more computing systems may generate a partial texture of the person by warping the pixels corresponding to the person based on the first body-surface mapping, the partial texture having incomplete texel information. The one or more computing systems may generate, based on the partial texture, a full texture of the person, the full texture having complete texel information.


