Few-shot Talking Head Synthesis via 3D Mesh and Neural Textures
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems face challenges in generating high-quality, novel views of talking heads and torsos using few input images, particularly in accurately representing 3D views and preserving user identity, as they often compress appearance information into a single latent vector, leading to loss of visual identity and high-frequency details.
Innovation Solution
The system employs a 3D mesh proxy and learned neural textures to represent user appearance, using a combination of face mesh and planar proxies to generate accurate and realistic images, with techniques like inverse rendering and attention mechanisms to fuse information from input views and preserve user identity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If appearance information is compressed into a single latent vector, then the dimensionality is reduced and processing is simplified, but visual identity and high-frequency details are lost
Solution Approach 1:
The patent segments appearance information into multiple separate latent vectors organized in a hierarchical structure, rather than compressing all information into a single vector. This segmentation preserves high-frequency details and visual identity while still achieving dimensionality reduction through organized decomposition of appearance features across multiple vectors at different hierarchical levels.
2Loss of time
If only a few input images are used, then the data requirement is reduced and processing time is shortened, but the quality and accuracy of novel views deteriorate
Solution Approach 1:
The patent introduces a hierarchical dimension to the latent space organization, arranging latent vectors in multiple levels where higher levels capture global appearance and lower levels preserve fine details. This hierarchical structure enables the system to generate high-quality novel views from few input images by systematically reconstructing appearance information across hierarchical levels rather than relying on large datasets.
3Device complexity
If conventional rendering is used, then the system is simpler to implement, but accurate 3D views and user identity preservation are compromised
Solution Approach 1:
The patent introduces a hierarchical latent space as an intermediary representation between input images and novel view synthesis. This hierarchical latent space acts as a mediator that systematically organizes appearance information across multiple levels, enabling accurate 3D view generation and identity preservation while maintaining relative system simplicity through structured intermediate representation.
Data Source
AI summary
Systems and methods are described for utilizing an image processing system with at least one processing device to perform operations including receiving a plurality of input images of a user, generating a three-dimensional mesh proxy based on a first set of features extracted from the plurality of input images and a second set of features extracted from the plurality of input images. The method may further include generating a neural texture based on a three-dimensional mesh proxy and the plurality of input images, generating a representation of the user including at least a neural texture, and sampling at least one portion of the neural texture from the three-dimensional mesh proxy. In response to providing the at least one sampled portion to a neural renderer, the method may include receiving, from the neural renderer, a synthesized image of the user that is previously not captured by the image processing system.


