Neural Network Novel View Synthesis from Sparse Images

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for presenting objects from multiple viewing angles and lighting conditions are expensive and inefficient, especially for complex objects, as they require extensive data capture and storage, and are limited by the specific lighting conditions of sparse image reconstructions.

Innovation Solution

A neural network system that generates images from sparse input images, allowing for novel viewpoints and lighting directions by creating a 3D volume of data and processing it to produce accurate predictions, including depth probabilities, resulting in a relit image or depth map corresponding to the desired view and lighting.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a light stage with many cameras and lights is used to capture images from all viewing angles and lighting conditions, then the completeness and quality of image data is improved, but the cost and device complexity increase significantly

Engineering Contradiction:
Improvecompleteness of image dataVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates a virtual copy of the physical light stage setup through a neural network model. Instead of using actual multiple cameras and lights, the system learns from a sparse set of real images and generates synthetic images that simulate what would be captured by a full light stage, thereby replacing expensive physical equipment with a computational model

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent changes the parameter of image acquisition from physical capture using multiple devices to computational generation through neural networks. The system transforms the approach from capturing all possible views physically to synthesizing views computationally, fundamentally changing how image data is obtained

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If a light stage is used to capture images from all viewing angles and lighting conditions, then the versatility of image retrieval is improved, but the data storage and organization requirements increase significantly

Engineering Contradiction:
Improveversatility of image retrievalVSAvoiddata storage requirements
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent generates virtual copies of images for all possible viewing angles and lighting conditions through neural network synthesis. Instead of storing thousands of actual captured images, the system stores a compact neural network model that can generate any required view on-demand, dramatically reducing storage requirements while maintaining full versatility

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms the static storage of pre-captured images into dynamic generation of images. The neural network model can adaptively generate images for any requested viewpoint and lighting condition in real-time, replacing the need to store and retrieve from large static datasets

Inventive Principle:
Principle #15Dynamics

3Device complexity

If sparse image reconstruction is used to generate 3D models for rendering different viewpoints, then the cost and data requirements are reduced, but the accuracy deteriorates for complex objects

Engineering Contradiction:
Improvedevice complexityVSAvoidreconstruction accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent introduces photometric stereo feature maps as an intermediary representation that bridges sparse images and accurate 3D understanding. These feature maps encode surface normal and reflectance information, serving as a mediator that enables accurate reconstruction of complex objects from sparse views without requiring full 360-degree image capture

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the approach from direct geometric reconstruction to physics-based rendering with learned material properties. By incorporating photometric stereo parameters and learned surface reflectance models, the system achieves accurate rendering of complex objects with varied materials and geometries from sparse input

Inventive Principle:
Principle #35Parameter changes

4Loss of time

If sparse photographs are used for 3D model reconstruction, then the cost and time requirements are reduced, but the lighting condition flexibility is limited

Engineering Contradiction:
Improvetime for image captureVSAvoidlighting condition flexibility
Core Design Contradiction:
Loss of timeVSAdaptability or versatility

Solution Approach 1:

The patent makes the lighting conditions dynamic and adjustable through neural network synthesis. Instead of being fixed to the lighting conditions during capture, the system can generate images under any requested lighting condition by manipulating the neural network model, enabling full flexibility without additional physical photography

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the lighting parameters from fixed physical conditions to adjustable computational parameters. The neural network model accepts lighting direction and intensity as input parameters, allowing dynamic control of lighting conditions in generated images independent of the original capture conditions

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10950037B2Deep novel view and lighting synthesis from sparse images
Publication Date: 2021.03.16 ADOBE INC
  • US10950037B2 patent drawing
  • US10950037B2 patent drawing
  • US10950037B2 patent drawing

AI summary

Embodiments are generally directed to generating novel images of an object having a novel viewpoint and a novel lighting direction based on sparse images of the object. A neural network is trained with training images rendered from a 3D model. Utilizing the 3D model, training images, ground truth predictive images from particular viewpoint(s), and ground truth predictive depth maps of the ground truth predictive images, can be easily generated and fed back through the neural network for training. Once trained, the neural network can receive a sparse plurality of images of an object, a novel viewpoint, and a novel lighting direction. The neural network can generate a plane sweep volume based on the sparse plurality of images, and calculate depth probabilities for each pixel in the plane sweep volume. A predictive output image of the object, having the novel viewpoint and novel lighting direction, can be generated and output.