Few-shot Talking Head Synthesis via 3D Mesh and Neural Textures

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems face challenges in generating high-quality, novel views of talking heads and torsos using few input images, particularly in accurately representing 3D views and preserving user identity, as they often compress appearance information into a single latent vector, leading to loss of visual identity and high-frequency details.

Innovation Solution

The system employs a 3D mesh proxy and learned neural textures to represent user appearance, using a combination of face mesh and planar proxies to generate accurate and realistic images, with techniques like inverse rendering and attention mechanisms to fuse information from input views and preserve user identity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If appearance information is compressed into a single latent vector, then the dimensionality is reduced and processing is simplified, but visual identity and high-frequency details are lost

Engineering Contradiction:
ImprovedimensionalityVSAvoidvisual identity
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent segments appearance information into multiple separate latent vectors organized in a hierarchical structure, rather than compressing all information into a single vector. This segmentation preserves high-frequency details and visual identity while still achieving dimensionality reduction through organized decomposition of appearance features across multiple vectors at different hierarchical levels.

Inventive Principle:
Principle #1Segmentation

2Loss of time

If only a few input images are used, then the data requirement is reduced and processing time is shortened, but the quality and accuracy of novel views deteriorate

Engineering Contradiction:
Improveprocessing timeVSAvoidquality of novel views
Core Design Contradiction:
Loss of timeVSManufacturing precision

Solution Approach 1:

The patent introduces a hierarchical dimension to the latent space organization, arranging latent vectors in multiple levels where higher levels capture global appearance and lower levels preserve fine details. This hierarchical structure enables the system to generate high-quality novel views from few input images by systematically reconstructing appearance information across hierarchical levels rather than relying on large datasets.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Device complexity

If conventional rendering is used, then the system is simpler to implement, but accurate 3D views and user identity preservation are compromised

Engineering Contradiction:
Improvesystem implementationVSAvoidaccuracy of 3D views
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent introduces a hierarchical latent space as an intermediary representation between input images and novel view synthesis. This hierarchical latent space acts as a mediator that systematically organizes appearance information across multiple levels, enabling accurate 3D view generation and identity preservation while maintaining relative system simplicity through structured intermediate representation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12026833B2Few-shot synthesis of talking heads
Publication Date: 2024.07.02 GOOGLE LLC
  • US12026833B2 patent drawing
  • US12026833B2 patent drawing
  • US12026833B2 patent drawing

AI summary

Systems and methods are described for utilizing an image processing system with at least one processing device to perform operations including receiving a plurality of input images of a user, generating a three-dimensional mesh proxy based on a first set of features extracted from the plurality of input images and a second set of features extracted from the plurality of input images. The method may further include generating a neural texture based on a three-dimensional mesh proxy and the plurality of input images, generating a representation of the user including at least a neural texture, and sampling at least one portion of the neural texture from the three-dimensional mesh proxy. In response to providing the at least one sampled portion to a neural renderer, the method may include receiving, from the neural renderer, a synthesized image of the user that is previously not captured by the image processing system.