Relightable 3D Portrait Generation Using Neural Point Clouds
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies face challenges in generating photorealistic, relightable 3D models of human heads or upper bodies from conventional camera videos, especially with limited resources, and often require complex and costly equipment like light stages.
Innovation Solution
The method employs a deep neural network that uses neural point-based graphics and 3D point clouds to generate relightable 3D portraits. It processes interleaved flash and no-flash images from smartphone cameras, disentangling albedo and lighting factors using face-specific priors, and renders 3D portraits from arbitrary viewpoints with interactive speed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If specialized equipment like light stages is used to acquire relightable models, then the quality and realism of 3D models is improved, but the device complexity and cost increase significantly
Solution Approach 1:
The patent uses a camera system to capture images that copy the appearance of the object under different lighting conditions, then uses neural networks to synthesize the relightable model from these 2D images rather than requiring specialized light stage equipment to directly capture 3D relightable data
Solution Approach 2:
The patent replaces complex mechanical light stage equipment with a combination of standard camera imaging and computational neural network processing to achieve the same relightable model acquisition goal
2Ease of operation
If conventional camera videos are used to generate 3D portraits, then the ease of operation and accessibility are improved, but the manufacturing precision and photorealism of the models deteriorate
Solution Approach 1:
The patent performs preliminary capture of interleaved flash and no-flash images with conventional cameras, then uses pre-trained neural networks to process these images and generate the relightable model, making the process accessible while maintaining quality
Solution Approach 2:
The patent changes the parameter representation by using latent descriptors in a neural network to encode appearance information, allowing high-quality model generation from conventional camera inputs through learned parameter transformations
3Productivity
If deep neural networks are used to process images in real-time, then the productivity and interactive speed are improved, but the use of energy and computational resources increase
Solution Approach 1:
The patent performs the computationally intensive training of deep neural networks in advance during an offline stage, then uses the trained model for fast real-time inference during actual operation, shifting the energy burden to a preliminary phase
Solution Approach 2:
The patent uses a trained neural network model that copies the learned processing patterns, allowing real-time rendering without repeating the full computational training process, thus reducing real-time energy consumption
Data Source
AI summary
The disclosure provides a method for generating relightable 3D portrait using a deep neural network and a computing device implementing the method. A possibility of obtaining, in real time and on computing devices having limited processing resources, realistically relighted 3D portraits having quality higher or at least comparable to quality achieved by prior art solutions, but without utilizing complex and costly equipment is provided. A method for rendering a relighted 3D portrait of a person, the method including: receiving an input defining a camera viewpoint and lighting conditions, rasterizing latent descriptors of a 3D point cloud at different resolutions based on the camera viewpoint to obtain rasterized images, wherein the 3D point cloud is generated based on a sequence of images captured by a camera with a blinking flash while moving the camera at least partly around an upper body, the sequence of images comprising a set of flash images and a set of no-flash images, processing the rasterized images with a deep neural network to predict albedo, normals, environmental shadow maps, and segmentation mask for the received camera viewpoint, and fusing the predicted albedo, normals, environmental shadow maps, and segmentation mask into the relighted 3D portrait based on the lighting conditions.


