Neural Network Talking Head Generation with Few-Shot Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies for generating video sequences with talking heads face limitations, such as unrealistic movement and rotation in warping-based methods and the need for extensive data and training time in GAN-based approaches, especially when using a small number of images.
Innovation Solution
A method involving a neural network model that performs first learning using meta-learning on multiple users' video sequences and second learning through few-shot learning for fine-tuning with a limited number of images from a new user, incorporating landmark information to generate realistic video sequences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If warping-based technology is used to generate video sequences with talking heads, then the amount of learning data required is reduced, but the realism of head movement and rotation deteriorates
Solution Approach 1:
The patent segments the learning process into two distinct phases: first learning (pre-training) on a large dataset of multiple users, and second learning (fine-tuning) on a small dataset of the target user. This segmentation allows the model to first learn general talking head generation capabilities broadly, then specialize for the specific user with minimal data, thereby achieving both data efficiency and realism.
Solution Approach 2:
The patent applies preliminary action by performing first learning (pre-training) on a large dataset of multiple users before conducting second learning (fine-tuning) on the target user's small dataset. This preliminary training establishes a robust foundation that enables the model to achieve realistic head movements with minimal user-specific data, resolving the contradiction between data quantity and realism.
2Reliability
If GAN-based technology is used to generate video sequences with talking heads, then the degree of reality is improved, but the training time and data requirements increase
Solution Approach 1:
The patent segments the training process into first learning (pre-training) on a large dataset and second learning (fine-tuning) on a small dataset. This segmentation reduces total training time by leveraging transfer learning, where the model acquires general capabilities during pre-training and only requires brief fine-tuning to achieve user-specific realism, thereby maintaining high degree of reality while reducing training time.
Solution Approach 2:
The patent performs preliminary action by conducting first learning (pre-training) on a large dataset of multiple users before fine-tuning on the target user. This preliminary training establishes robust general capabilities that reduce the time and data needed for subsequent user-specific training, resolving the contradiction between achieving high realism and minimizing training time.
3Adaptability or versatility
If a neural network model is trained on multiple users' video sequences, then the adaptability to different users is improved, but the complexity of the training process increases
Solution Approach 1:
The patent segments the training process into two distinct stages: first learning (pre-training) on multiple users' video sequences to learn general talking head generation capabilities, and second learning (fine-tuning) on the target user's small dataset. This segmentation improves adaptability to different users while managing training complexity by reusing the pre-trained model for each new user, avoiding the need to train from scratch.
Solution Approach 2:
The patent applies universality by training a single neural network model on multiple users' data during first learning, creating a universal model that can generate talking heads for any user. This universal model is then fine-tuned for specific users during second learning, achieving high adaptability while avoiding the complexity of maintaining separate models for each user.
Data Source
AI summary
An electronic device and a controlling method thereof are provided. A controlling method of an electronic device according to the disclosure includes: performing first learning for a neural network model for acquiring a video sequence including a talking head of a random user based on a plurality of learning video sequences including talking heads of a plurality of users, performing second learning for fine-tuning the neural network model based on at least one image including a talking head of a first user different from the plurality of users and first landmark information included in the at least one image, and acquiring a first video sequence including the talking head of the first user based on the at least one image and pre-stored second landmark information using the neural network model for which the first learning and the second learning were performed.


