Neural Network Novel View Synthesis from Sparse Images
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for presenting objects from multiple viewing angles and lighting conditions are expensive and inefficient, especially for complex objects, as they require extensive data capture and storage, and are limited by the specific lighting conditions of sparse image reconstructions.
Innovation Solution
A neural network system that generates images from sparse input images, allowing for novel viewpoints and lighting directions by creating a 3D volume of data and processing it to produce accurate predictions, including depth probabilities, resulting in a relit image or depth map corresponding to the desired view and lighting.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a light stage with many cameras and lights is used to capture images from all viewing angles and lighting conditions, then the completeness and quality of image data is improved, but the cost and device complexity increase significantly
Solution Approach 1:
The patent creates a virtual copy of the physical light stage setup through a neural network model. Instead of using actual multiple cameras and lights, the system learns from a sparse set of real images and generates synthetic images that simulate what would be captured by a full light stage, thereby replacing expensive physical equipment with a computational model
Solution Approach 2:
The patent changes the parameter of image acquisition from physical capture using multiple devices to computational generation through neural networks. The system transforms the approach from capturing all possible views physically to synthesizing views computationally, fundamentally changing how image data is obtained
2Adaptability or versatility
If a light stage is used to capture images from all viewing angles and lighting conditions, then the versatility of image retrieval is improved, but the data storage and organization requirements increase significantly
Solution Approach 1:
The patent generates virtual copies of images for all possible viewing angles and lighting conditions through neural network synthesis. Instead of storing thousands of actual captured images, the system stores a compact neural network model that can generate any required view on-demand, dramatically reducing storage requirements while maintaining full versatility
Solution Approach 2:
The patent transforms the static storage of pre-captured images into dynamic generation of images. The neural network model can adaptively generate images for any requested viewpoint and lighting condition in real-time, replacing the need to store and retrieve from large static datasets
3Device complexity
If sparse image reconstruction is used to generate 3D models for rendering different viewpoints, then the cost and data requirements are reduced, but the accuracy deteriorates for complex objects
Solution Approach 1:
The patent introduces photometric stereo feature maps as an intermediary representation that bridges sparse images and accurate 3D understanding. These feature maps encode surface normal and reflectance information, serving as a mediator that enables accurate reconstruction of complex objects from sparse views without requiring full 360-degree image capture
Solution Approach 2:
The patent changes the approach from direct geometric reconstruction to physics-based rendering with learned material properties. By incorporating photometric stereo parameters and learned surface reflectance models, the system achieves accurate rendering of complex objects with varied materials and geometries from sparse input
4Loss of time
If sparse photographs are used for 3D model reconstruction, then the cost and time requirements are reduced, but the lighting condition flexibility is limited
Solution Approach 1:
The patent makes the lighting conditions dynamic and adjustable through neural network synthesis. Instead of being fixed to the lighting conditions during capture, the system can generate images under any requested lighting condition by manipulating the neural network model, enabling full flexibility without additional physical photography
Solution Approach 2:
The patent changes the lighting parameters from fixed physical conditions to adjustable computational parameters. The neural network model accepts lighting direction and intensity as input parameters, allowing dynamic control of lighting conditions in generated images independent of the original capture conditions
Data Source
AI summary
Embodiments are generally directed to generating novel images of an object having a novel viewpoint and a novel lighting direction based on sparse images of the object. A neural network is trained with training images rendered from a 3D model. Utilizing the 3D model, training images, ground truth predictive images from particular viewpoint(s), and ground truth predictive depth maps of the ground truth predictive images, can be easily generated and fed back through the neural network for training. Once trained, the neural network can receive a sparse plurality of images of an object, a novel viewpoint, and a novel lighting direction. The neural network can generate a plane sweep volume based on the sparse plurality of images, and calculate depth probabilities for each pixel in the plane sweep volume. A predictive output image of the object, having the novel viewpoint and novel lighting direction, can be generated and output.


