Neural Network Talking Head Generation with Few-Shot Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies for generating video sequences with talking heads face limitations, such as unrealistic movement and rotation in warping-based methods and the need for extensive data and training time in GAN-based approaches, especially when using a small number of images.

Innovation Solution

A method involving a neural network model that performs first learning using meta-learning on multiple users' video sequences and second learning through few-shot learning for fine-tuning with a limited number of images from a new user, incorporating landmark information to generate realistic video sequences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If warping-based technology is used to generate video sequences with talking heads, then the amount of learning data required is reduced, but the realism of head movement and rotation deteriorates

Engineering Contradiction:
Improveamount of learning dataVSAvoidrealism of head movement
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent segments the learning process into two distinct phases: first learning (pre-training) on a large dataset of multiple users, and second learning (fine-tuning) on a small dataset of the target user. This segmentation allows the model to first learn general talking head generation capabilities broadly, then specialize for the specific user with minimal data, thereby achieving both data efficiency and realism.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by performing first learning (pre-training) on a large dataset of multiple users before conducting second learning (fine-tuning) on the target user's small dataset. This preliminary training establishes a robust foundation that enables the model to achieve realistic head movements with minimal user-specific data, resolving the contradiction between data quantity and realism.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If GAN-based technology is used to generate video sequences with talking heads, then the degree of reality is improved, but the training time and data requirements increase

Engineering Contradiction:
Improvedegree of realityVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the training process into first learning (pre-training) on a large dataset and second learning (fine-tuning) on a small dataset. This segmentation reduces total training time by leveraging transfer learning, where the model acquires general capabilities during pre-training and only requires brief fine-tuning to achieve user-specific realism, thereby maintaining high degree of reality while reducing training time.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary action by conducting first learning (pre-training) on a large dataset of multiple users before fine-tuning on the target user. This preliminary training establishes robust general capabilities that reduce the time and data needed for subsequent user-specific training, resolving the contradiction between achieving high realism and minimizing training time.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If a neural network model is trained on multiple users' video sequences, then the adaptability to different users is improved, but the complexity of the training process increases

Engineering Contradiction:
Improveadaptability to different usersVSAvoidcomplexity of training process
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the training process into two distinct stages: first learning (pre-training) on multiple users' video sequences to learn general talking head generation capabilities, and second learning (fine-tuning) on the target user's small dataset. This segmentation improves adaptability to different users while managing training complexity by reusing the pre-trained model for each new user, avoiding the need to train from scratch.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies universality by training a single neural network model on multiple users' data during first learning, creating a universal model that can generate talking heads for any user. This universal model is then fine-tuned for specific users during second learning, achieving high adaptability while avoiding the complexity of maintaining separate models for each user.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12039456B2Electronic device and controlling method thereof
Publication Date: 2024.07.16 SAMSUNG ELECTRONICS CO LTD
  • US12039456B2 patent drawing
  • US12039456B2 patent drawing
  • US12039456B2 patent drawing

AI summary

An electronic device and a controlling method thereof are provided. A controlling method of an electronic device according to the disclosure includes: performing first learning for a neural network model for acquiring a video sequence including a talking head of a random user based on a plurality of learning video sequences including talking heads of a plurality of users, performing second learning for fine-tuning the neural network model based on at least one image including a talking head of a first user different from the plurality of users and first landmark information included in the at least one image, and acquiring a first video sequence including the talking head of the first user based on the at least one image and pre-stored second landmark information using the neural network model for which the first learning and the second learning were performed.