Personalized portrait face video reconstruction method based on meta-learning
Through a meta-learning-based method, the source image and driver video are divided into T tasks. The improved ResNet-50 and FLAME models are used, combined with UV mapping and LBS algorithm, and the problems of low computing efficiency and poor generalization capabilities in the existing technology are solved, and efficient personalized facial video reconstruction is achieved to generate realistic personalized facial videos.
Patent Information
- Application Number
- CN202510379935.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-08-19
AI Technical Summary
The existing video reconstruction methods have low computational efficiency and poor generalization capabilities in personalized facial reconstruction, making it difficult to quickly and accurately generate highly personalized reconstruction results, and their performance is significantly reduced when facing diversified data or cross-domain tasks.
Using a meta-learning-based method, the source image and driver video are divided into T non-overlapping tasks. The improved ResNet-50 encoder and FLAME three-dimensional face model are used, combined with UV mapping technology and linear hybrid skin LBS algorithm, and the model parameters are optimized through internal and external cycles, and a fine three-dimensional face model is generated and dynamically deformed.
It significantly improves the generalization ability and personalized adaptability of the model, and can quickly adapt to new tasks under the conditions of few samples, generate realistic personalized facial videos, especially in the simulation of teeth and lip movements, with more realistic texture details, adapting to real-time expressions and posture changes.
Smart Images

Figure CN120510237A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video reconstruction, and in particular to a personalized portrait facial video reconstruction method based on meta-learning. Background Art
[0002] Traditional video compression methods, such as H.264 / AVC and H.265 / HEVC, primarily reduce data volume by optimizing spatial, temporal, perceptual, and information entropy redundancy. However, when dealing with large data volumes and high quality requirements, these methods fail to fully exploit the content correlation between image frames, resulting in limited compression ratio improvements and a technical bottleneck. To address this, video reconstruction technology has emerged. To simplify the problem, it is optimized for conversational videos with simple backgrounds, which primarily focus on accurately reconstructing portraits. Existing 3D reconstruction methods, such as the FLAME model-based method proposed by Baert et al., can more accurately model faces and improve the reproduction of facial details; the emotion-driven method proposed by Filntisis et al. significantly enhances the realism of facial reconstruction by capturing emotion-driven facial expressions; and the 3D Facial Expressions method proposed by Retsinas et al., based on a 3D mesh model, accurately simulates facial dynamics and improves the reconstruction of facial expressions.
[0003] Existing models typically employ a global optimization strategy during training, optimizing the entire task based on a unified set of parameters. While this approach achieves good results on some standard tasks, it ignores the personalized needs of different users and tasks, resulting in poor generalization and slow convergence when faced with limited data and new tasks. Meta-learning, as an advanced learning paradigm, was first applied to the field of classification and recognition due to its remarkable ability to quickly adapt to new tasks with minimal data. For example, Siamese Networks, by learning a similarity metric between pairs of input samples, can quickly identify new categories with only a small amount of labeled data, demonstrating strong generalization capabilities. MAML, on the other hand, significantly improves the model's adaptability to new tasks by optimizing the model's initial parameters, enabling it to quickly converge to a better solution on new tasks with a small number of gradient updates. Based on its principles, meta-learning is also applicable to the field of video reconstruction, where it can improve the model's performance and generalization capabilities on new tasks through personalized adaptation strategies.
[0004] Existing 3D reconstruction methods, such as the FLAME model, emotion-driven methods, and 3D Facial Expressions, can simulate basic facial structure and dynamic expression changes well, but they have obvious shortcomings in processing individual details and texture information. For example, the FLAME model performs poorly in simulating tooth details and lip opening and closing movements; while the emotion-driven method can generate emotion-related expressions, it is not perfect when integrating facial details (such as wrinkles and moles) and texture information; and the 3D Facial Expressions method has low computational efficiency when processing complex expressions and posture changes, and it also has difficulty effectively integrating individual details and texture information. These shortcomings result in a large gap between the realism and accuracy of the reconstruction results and actual needs, highlighting the limitations of existing methods in expressing personalized details.
[0005] Furthermore, existing methods face numerous challenges in practical training and application. On the one hand, traditional deep learning algorithms are computationally expensive and inefficient, making them difficult to meet real-time processing requirements. On the other hand, after training, these algorithms typically only provide optimal weight configurations on average, limiting the model's ability to generalize to new tasks. Consequently, when deployed in practice, current models are limited by both computational efficiency and generalization capabilities, making it difficult to quickly and accurately generate highly personalized reconstruction results.
[0006] Furthermore, most existing research directly uses public or self-produced datasets for training, causing the models to focus more on the overall characteristics of the dataset while ignoring the subtle differences of specific users or specific environments. Furthermore, these methods are often based on the independent and identically distributed (IID) assumption, resulting in the models performing well only in situations similar to the training data. Performance degrades significantly when faced with diverse data or cross-domain tasks in real-world applications. This limitation further hinders the development and practical application of personalized facial reconstruction technology. Summary of the Invention
[0007] In view of the shortcomings of the existing technology, the present invention aims to propose a personalized portrait facial video reconstruction method based on meta-learning, comprising:
[0008] Step 1: Acquire multiple source images of T users, wherein the multiple source images of each user include images of the user under different lighting conditions, hairstyle characteristics, and facial expressions, and acquire a driving video corresponding to each source image, wherein the driving video is used to make the source image move;
[0009] Step 2: Process each source image and driving video using a face detection algorithm to obtain a processed source image and a processed driving video;
[0010] Step 3: Divide the processed source image and its corresponding driving video into T tasks. Each task contains a source image and its corresponding driving video processed by a user under different lighting, hairstyle characteristics, and facial expressions. Divide the T tasks into training and test sets according to a preset ratio. The processed source image and its corresponding driving video are used as a set of data. In the training and test sets, all the group data of each task are divided according to a preset ratio to obtain the support set and query set.
[0011] Step 4: Obtain a MetaFace model, which includes an encoder module, a decoder module, and a rendering module. Initialize the parameters in the encoder module, the decoder module, and the rendering module. The initialized parameters form a local parameter vector. The initialized local parameter vector is used as the current parameter vector. Set the first and second iteration times. The initial values of the first and second iteration times are both 0.
[0012] Step 5: For each task in the training set, the processed source image and its corresponding processed driving video in the support set of the task are input into the MetaFace model based on the current parameter vector. After passing through the encoder module, decoder module and rendering module, the predicted 2D video frame is output.
[0013] Step 6: Based on the output predicted 2D video frame, the source image input to the MetaFace model, and the real 2D video frame corresponding to the processed driving video, calculate the facial key point loss and detail consistency loss. Based on the facial key point loss and detail consistency loss, update the current parameter vector through k-step stochastic gradient descent or Adam optimizer. The first iteration number is increased by one, and it is determined whether the first iteration number reaches the preset threshold. If the first iteration number does not reach the preset threshold, the updated parameter vector is used as the new current parameter vector, and the process returns to step 5. If the first iteration number reaches the preset threshold, the final parameter vector of the task is obtained, and then the final parameter vectors of all tasks are obtained.
[0014] Step 7: Aggregate the final parameter vectors of all tasks into the current global parameter vector. Update the current global parameter vector through the Adam optimizer to obtain the updated global parameter vector.
[0015] Step 8: The second iteration number is increased by one, and it is determined whether the second iteration number reaches the preset threshold. If the second iteration number does not reach the preset threshold, the parameter vector of each task in the updated global parameter vector is used as the current parameter vector, and the process returns to step 5. If the second iteration number reaches the preset threshold, the final global parameter vector is obtained, and then the final MetaFace model is obtained.
[0016] Optionally, step 2 specifically includes:
[0017] Performing face region detection on each source image using a face detection algorithm, calculating the proportion of the face region in the source image, and if the proportion of the face region in the source image is greater than or equal to a preset threshold, not processing the source image and using the source image as the processed source image; and if the proportion of the face region in the source image is less than the preset threshold, cropping the source image to obtain a processed source image.
[0018] Similarly, for each frame image in the driving video, face area detection is performed on each frame image, and the proportion of the face area in the image is calculated. When the proportion of the face area in the image is greater than or equal to the preset threshold, the image is not processed. When the proportion of the face area in the image is less than the preset threshold, the image is cropped. After all frame images of all driving videos are processed as above, the processed driving video is obtained.
[0019] Optionally, step 5 specifically includes:
[0020] Step 5.1: In the encoder module based on the current parameter vector, for each task in the training set, the processed source image in the support set of the task is input into the encoder to obtain semantic features, where the semantic features include content parameters and texture features, and the content parameters include shape features, expression features, and posture features;
[0021] The encoder is obtained by modifying the output dimension of the fully connected layer in the ResNet-50 model;
[0022] Step 5.2: In the decoder module based on the current parameter vector, generate the final 3D face model based on the semantic features and the general FLAME 3D face model;
[0023] Step 5.3: In the rendering module, the final 3D face model is rendered based on the geometric structure, texture information, and ambient lighting conditions, and the predicted 2D video frame is output.
[0024] Optionally, step 5.2 specifically includes:
[0025] Step 5.2.1: Based on the content parameters, adjust the general FLAME 3D face model to obtain an initial 3D face model. Based on the linear blend skinning (LBS) algorithm, adjust the vertices of three target areas in the initial 3D face model to obtain three target areas after vertex adjustment. The three target areas include the area where the teeth are located, the area where the nose is located, and the area where the cheeks are located. For each target area after vertex adjustment, connect the adjusted vertices through patches to obtain a rough 3D face model.
[0026] Step 5.2.2: Using UV mapping technology, based on texture features, the rough 3D face model is processed to obtain a refined 3D face model;
[0027] Specifically, through the texture decoder T d , the texture features are processed to obtain the UV displacement map D, which is specifically achieved through the following formula:
[0028]
[0029] Where δ is the texture parameter, is the expression parameter, θ jaw is the mandibular posture parameter;
[0030] Based on the UV displacement map D, the vertices of the rough 3D face model are adjusted to obtain the final position of the vertices in the UV space. This is achieved by the following formula:
[0031] M' uv =M uv +D⊙N uv ;
[0032] Among them, M′ uv is the final position of the vertex in UV space, M uv is the vertex position of the rough 3D face model in UV space, N uv M uv The corresponding surface normal, ⊙ represents element-wise multiplication;
[0033] The rough 3D face model is adjusted based on the final positions of the vertices in the UV space to obtain a fine 3D face model;
[0034] Step 5.2.3: For the source image input to the encoder in step 5, obtain the corresponding processed driving video. For each frame of the image, identify the current posture θ of the user in the image and obtain the vertex position v′ of the user's face area. Based on the linear blending skinning (LBS) algorithm, adjust the vertex position of the user's face area to obtain the LBS deformed vertex position v", thereby obtaining the LBS deformed vertex position v" corresponding to each frame of the image; perform expression deformation on each frame of the image using the Blendshapes technology to obtain the expression deformed vertex position v expr , thus obtaining the vertex position v after expression deformation corresponding to each frame image expr ;
[0035] Based on the processed driving video, the vertex position v" after LBS deformation and the vertex position v after expression deformation corresponding to each frame image expr , adjust the refined three-dimensional face model to obtain the final three-dimensional face model.
[0036] Optionally, in step 5.2.3, the vertex positions of the user's face area are adjusted based on the linear blend skinning LBS algorithm to obtain the vertex position v" after LBS deformation, which is specifically achieved by the following formula:
[0037]
[0038] Among them, K b is the total number of bones, ω i is the weight of vertex v′ relative to the i-th bone, T i (θ) is the transformation matrix of the current posture θ.
[0039] Optionally, in step 5.2.3, each frame of the image is deformed by the Blendshapes technology to obtain the vertex position v after the deformation of the expression expr , which is specifically achieved through the following formula:
[0040]
[0041] in, is the average face shape, K e is the total number of Blendshape models, φ j is the expression coefficient corresponding to each frame image, B j is the j-th blendshape model.
[0042] Optionally, in step 7, the current global parameter vector is updated using the Adam optimizer to obtain an updated global parameter vector, which is specifically implemented using the following formula:
[0043]
[0044] Among them, φ′ is the updated global parameter vector, φ n is the final parameter vector of the nth task, φ is the current global parameter vector, and β is the outer loop learning rate.
[0045] The beneficial effects of adopting the above technical solution are:
[0046] Compared to existing technologies, the present invention divides the processed source image and its corresponding processed driving video into T non-overlapping tasks, each corresponding to a subset of characters with a specific identity. This innovative design significantly improves the model's generalization and personalized adaptability. Secondly, the encoder uses a modified fully connected ResNet-50 network as the backbone network to output semantic features. These features are then combined with the universal FLAME 3D face model to construct a coarse 3D face model. This allows for accurate simulation of basic facial structure and dynamic expression, particularly in tooth modeling and lip movement simulation, significantly enhancing the realism and detail of the reconstruction. Furthermore, the present invention leverages UV mapping technology to provide detailed texture restoration capabilities for the coarse 3D face model, rendering features such as skin texture, pores, and moles more realistically. This results in a refined 3D face model, further enhancing the model's ability to express personalized appearance. Furthermore, the present invention dynamically deforms the refined 3D face model based on the motion information of the driving video. The rendering module converts the final 3D face model output by the decoder module into 2D video frames, completing the efficient 3D to 2D conversion. Finally, the present invention employs a meta-learning module to optimize model performance through an inner and outer loop. The inner loop fine-tunes parameters for a single task, while the outer loop integrates parameter update trajectories across tasks and uses meta-gradient optimization to generate universal initialization parameters. This enables the model to quickly adapt to new tasks with few samples, significantly improving generalization and transfer efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 Schematic diagram of a process for personalized portrait facial video reconstruction based on meta-learning in an embodiment of the present invention;
[0048] Figure 2 A schematic diagram of constructing a refined three-dimensional face model in an embodiment of the present invention;
[0049] Figure 3: This is a comparison diagram of the implementation results of the present invention and other models in the embodiment of the present invention, wherein, Figure (a1) is the source image of person 1, (a2) is the driving video of person 1, (a3) is the image of person 1 obtained based on the FOMM model, (a4) is the image of person 1 obtained based on the TPSMM model, (a5) is the image of person 1 obtained based on the DaGAN model, (a6) is the image of person 1 obtained based on the MCNET model, and (a7) is the image of person 1 obtained under the model of the present invention; Figure (b1) is the source image of person 2, (b2) is the driving video of person 2, (b3) is the image of person 2 obtained based on the FOMM model, (b4) is the image of person 2 obtained based on the TPSMM model, (b5) is the image of person 2 obtained based on the DaGAN model, (b6) is the image of person 2 obtained based on the MCNET model, and (b7) is the image of person 2. 2 is the image obtained under the model of the present invention; Figure (c1) is the source image of person 3, (c2) is the driving video of person 3, (c3) is the image of person 3 obtained based on the FOMM model, (c4) is the image of person 3 obtained based on the TPSMM model, (c5) is the image of person 3 obtained based on the DaGAN model, (c6) is the image of person 3 obtained based on the MCNET model, and (c7) is the image of person 3 obtained under the model of the present invention; Figure (d1) is the source image of person 4, (d2) is the driving video of person 4, (d3) is the image of person 4 obtained based on the FOMM model, (d4) is the image of person 4 obtained based on the TPSMM model, (d5) is the image of person 4 obtained based on the DaGAN model, (d6) is the image of person 4 obtained based on the MCNET model, and (d7) is the image of person 4 obtained under the model of the present invention. DETAILED DESCRIPTION
[0050] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.
[0051] In view of the problems existing in the prior art, the present invention provides a personalized portrait facial video reconstruction method based on meta-learning, combining Figure 1 , which may include the following steps:
[0052] Step 1: Acquire multiple source images of T users, where each source image includes images of the user under different lighting conditions, hairstyle characteristics, and facial expressions. Acquire a driving video corresponding to each source image, where the driving video is used to cause the source image to move and can provide motion information.
[0053] The source images and their corresponding driving videos can be obtained from the YouTube Faces dataset.
[0054] Step 2: Process each source image and driving video using a face detection algorithm to obtain a processed source image and a processed driving video;
[0055] Performing face region detection on each source image using a face detection algorithm, calculating the proportion of the face region in the source image, and if the proportion of the face region in the source image is greater than or equal to a preset threshold, not processing the source image and using the source image as the processed source image; and if the proportion of the face region in the source image is less than the preset threshold, cropping the source image to obtain a processed source image.
[0056] Similarly, for each frame image in the driving video, face area detection is performed on each frame image, and the proportion of the face area in the image is calculated. When the proportion of the face area in the image is greater than or equal to the preset threshold, the image is not processed. When the proportion of the face area in the image is less than the preset threshold, the image is cropped. After all frame images of all driving videos are processed as above, the processed driving video is obtained.
[0057] Step 3: Divide the processed source image and its corresponding driving video to obtain T tasks, where one task includes a source image processed by a user under different lighting, hairstyle characteristics and facial expressions and its corresponding driving video. The T tasks are divided into training sets and test sets according to a preset ratio, wherein the training set is used for model training and the test set is used to evaluate the effect of task adaptability. The present invention adopts the N-Way K-Shot algorithm to realize the division into training sets and test sets according to a preset ratio. The processed source image and its corresponding driving video are used as a group of data. In the training set and test set, all group data in each task are divided according to a preset ratio to obtain a support set and a query set. In the test set, it is also necessary to divide the support set and the query set according to the preset ratio.
[0058] Step 4: Obtain a MetaFace model, which includes an encoder module, a decoder module, and a rendering module. Initialize the parameters in the encoder module, the decoder module, and the rendering module. The initialized parameters form a local parameter vector. The initialized local parameter vector is used as the current parameter vector. Set the first and second iteration times. The initial values of the first and second iteration times are both 0.
[0059] Step 5: For each task in the training set, the processed source image and its corresponding processed driving video in the support set of the task are input into the MetaFace model based on the current parameter vector. After passing through the encoder module, decoder module and rendering module, the predicted 2D video frame is output.
[0060] Step 5.1: In the encoder module based on the current parameter vector, for each task in the training set, the processed source image in the support set of the task is input to the encoder (i.e. Figure 2 In the content encoder in ), semantic features are obtained, wherein the semantic features include content parameters and texture features, and the content parameters include shape features, expression features, and posture features;
[0061] The encoder is obtained by modifying the output dimension of the fully connected layer in the ResNet-50 model;
[0062] Step 5.2: In the decoder module based on the current parameter vector, generate the final 3D face model based on the semantic features and the general FLAME 3D face model;
[0063] Step 5.2.1: Binding Figure 2 Based on the content parameters, the general FLAME 3D face model is adjusted to obtain the initial 3D face model. Based on the linear blend skinning LBS algorithm, the vertices of the three target areas in the initial 3D face model are adjusted to obtain the three target areas after adjusting the vertices. The three target areas include the area where the teeth are located, the area where the nose is located, and the area where the cheeks are located. For each target area after adjusting the vertices, the adjusted vertices are connected through the patch to obtain a rough 3D face model, that is, Figure 2 The rough face modeling part.
[0064] Taking the area where the teeth are located as an example, 216 parameterizable vertices are introduced for the 14 visible teeth to accurately control the shape of individual teeth (such as the protrusion of the incisors and the width of the tooth gaps), and the occlusal surface curvature and gingival transition structure are constructed through triangular patches, realizing high-freedom modeling of tooth geometric details, while ensuring the physical fit between the lip muscles and teeth in facial animations. For example, when smiling, the upper lip naturally rises to expose the complete tooth surface.
[0065] Step 5.2.2: Using UV mapping technology, based on texture features, the rough 3D face model is processed to obtain a refined 3D face model;
[0066] Specifically, through the texture decoder T d , the texture features are processed to obtain the UV displacement map D, which is specifically achieved through the following formula:
[0067]
[0068] Where δ is the texture parameter, is the expression parameter, θ jaw is the mandibular posture parameter;
[0069] Based on the UV displacement map D, the vertices of the rough 3D face model are adjusted to obtain the final position of the vertices in the UV space (i.e. Figure 2 The facial details in the image are obtained by the following formula:
[0070] M' uv =M uv +D⊙N uv ;
[0071] Among them, M′ uv is the final position of the vertex in UV space, M uv is the vertex position of the rough 3D face model in UV space, N uv M uv The corresponding surface normal, ⊙ represents element-wise multiplication;
[0072] The rough 3D face model is adjusted based on the final position of the vertices in the UV space to obtain a fine 3D face model, i.e. Figure 2 Medium-fine face modeling part.
[0073] Step 5.2.3: For the source image input to the encoder in step 5, obtain the corresponding processed driving video. For each frame of the image, identify the current posture θ of the user in the image and obtain the vertex position v′ of the user's face area. Based on the linear blend skinning algorithm, adjust the vertex position of the user's face area to obtain the vertex position v' after LBS deformation. This is achieved by the following formula:
[0074]
[0075] Among them, K b is the total number of bones, ω i is the weight of vertex v′ relative to the i-th bone, T i (θ) is the transformation matrix of the current posture θ.
[0076] Thus, the vertex position v" after LBS deformation corresponding to each frame image is obtained;
[0077] The expression of each frame image is deformed by Blendshapes technology to obtain the vertex position v after expression deformation expr , which is specifically achieved through the following formula:
[0078]
[0079] in, is the average face shape, K e is the total number of Blendshape models, φ j is the expression coefficient corresponding to each frame image, B jis the j-th blendshape model.
[0080] Thus, the vertex position v after expression deformation corresponding to each frame image is obtained expr ;
[0081] Therefore, the present invention uses linear blend skinning (LBS) and Blendshapes technology to identify changes in the processed driving video and ensure that the MetaFace model can adapt to real-time expression and posture changes.
[0082] Based on the processed driving video, the vertex position v" after LBS deformation and the vertex position v after expression deformation corresponding to each frame image expr , adjust the refined three-dimensional face model to obtain the final three-dimensional face model.
[0083] Step 5.3: In the rendering module, the final 3D face model is rendered based on the geometric structure, texture information, and ambient lighting conditions, and the predicted 2D video frame is output.
[0084] Next, the present invention combines a meta-learning module to optimize the initialization weights through inner and outer loops. The inner loop performs gradient descent to optimize local parameters for the same subject within a single task, then passes the updated parameters to the outer loop; the outer loop then summarizes common knowledge across multiple tasks to generate the final global initialization parameters. For details, refer to steps 6 through 8.
[0085] Step 6: Based on the output predicted 2D video frame, the source image input to the MetaFace model, and the real 2D video frame corresponding to the processed driving video, calculate the facial key point loss and detail consistency loss. Based on the facial key point loss and detail consistency loss, update the current parameter vector through k-step stochastic gradient descent or Adam optimizer. The first iteration number is increased by one, and it is determined whether the first iteration number reaches the preset threshold. If the first iteration number does not reach the preset threshold, the updated parameter vector is used as the new current parameter vector, and the process returns to step 5. If the first iteration number reaches the preset threshold, the final parameter vector of the task is obtained, and then the final parameter vectors of all tasks are obtained.
[0086] Step 7: Aggregate the final parameter vectors of all tasks into the current global parameter vector. Update the current global parameter vector through the Adam optimizer to obtain the updated global parameter vector. This is achieved by the following formula:
[0087]
[0088] Among them, φ′ is the updated global parameter vector, φ nis the final parameter vector of the nth task, φ is the current global parameter vector, and β is the outer loop learning rate.
[0089] Step 8: The second iteration number is increased by one, and it is determined whether the second iteration number reaches the preset threshold. If the second iteration number does not reach the preset threshold, the parameter vector of each task in the updated global parameter vector is used as the current parameter vector, and the process returns to step 5. If the second iteration number reaches the preset threshold, the final global parameter vector is obtained, and then the final MetaFace model is obtained.
[0090] In a specific implementation of the present invention, the preset threshold of the first iteration number is set to 24, the preset threshold of the second iteration number is set to 300, and the value of T is 532.
[0091] Based on this, the key technical points of the present invention are:
[0092] 1. Based on an improved FLAME model, this invention combines coarse and fine facial modeling modules with a deformation module to accurately simulate basic facial structure and dynamic facial expressions, while effectively restoring individual details. This technology is particularly accurate in simulating dental details and lip movements, significantly enhancing the realism and detail of the reconstruction.
[0093] 2. Based on UV mapping technology, this invention restores texture features of 3D face models, such as skin texture, pores, moles, etc., making the reconstruction effect more delicate and realistic, and further improving the reconstruction quality.
[0094] 3. This paper performs multi-task partitioning on the YouTube Faces dataset, dividing it into T non-overlapping meta-tasks, each corresponding to a subset of characters with a specific identity. By simulating the adaptation of new faces to a scene, the model captures commonalities across characters while retaining individual specificity, thereby improving the efficiency of subsequent reconstruction.
[0095] 4. This paper employs a meta-learning module, implementing an inner and outer loop. The inner loop fine-tunes parameters on a single task, while the outer loop integrates parameter update trajectories across tasks and generates universal initialization parameters using meta-gradient optimization. This module allows the model to quickly adapt to new tasks with a small number of samples, significantly improving generalization and transfer efficiency.
[0096] Compared to existing technologies, the present invention divides the processed source image and its corresponding processed driving video into T non-overlapping tasks, each corresponding to a subset of characters with a specific identity. This innovative design significantly improves the model's generalization and personalized adaptability. Secondly, the encoder uses a modified fully connected ResNet-50 network as the backbone network to output semantic features. These features are then combined with the universal FLAME 3D face model to construct a coarse 3D face model. This allows for accurate simulation of basic facial structure and dynamic expression, particularly in tooth modeling and lip movement simulation, significantly enhancing the realism and detail of the reconstruction. Furthermore, the present invention leverages UV mapping technology to provide detailed texture restoration capabilities for the coarse 3D face model, rendering features such as skin texture, pores, and moles more realistically. This results in a refined 3D face model, further enhancing the model's ability to express personalized appearance. Furthermore, the present invention dynamically deforms the refined 3D face model based on the motion information of the driving video. The rendering module converts the final 3D face model output by the decoder module into 2D video frames, completing the efficient 3D to 2D conversion. Finally, the present invention employs a meta-learning module to optimize model performance through an inner and outer loop. The inner loop fine-tunes parameters for a single task, while the outer loop integrates parameter update trajectories across tasks and utilizes meta-gradient optimization to generate universal initialization parameters. This enables the model to quickly adapt to new tasks with few samples, significantly improving generalization and transfer efficiency. In summary, this method not only quickly and accurately adapts to new tasks and achieves high-precision portrait facial reconstruction, but also demonstrates significant research value and practical application potential, and is expected to be widely used in fields such as virtual reality, game development, and film and television special effects.
[0097] Figure 3This is a comparison diagram of the results achieved by the present invention and other models, wherein Figure (a1) is the source image of person 1, (a2) is the driving video of person 1, (a3) is the image of person 1 obtained based on the FOMM model, (a4) is the image of person 1 obtained based on the TPSMM model, (a5) is the image of person 1 obtained based on the DaGAN model, (a6) is the image of person 1 obtained based on the MCNET model, and (a7) is the image of person 1 obtained under the model of the present invention; Figure (b1) is the source image of person 2, (b2) is the driving video of person 2, (b3) is the image of person 2 obtained based on the FOMM model, (b4) is the image of person 2 obtained based on the TPSMM model, (b5) is the image of person 2 obtained based on the DaGAN model, (b6) is the image of person 2 obtained based on the MCNET model, and (b7) is the image of person 2 in this Images obtained under the model of the invention; Figure (c1) is the source image of person 3, (c2) is the driving video of person 3, (c3) is the image of person 3 obtained based on the FOMM model, (c4) is the image of person 3 obtained based on the TPSMM model, (c5) is the image of person 3 obtained based on the DaGAN model, (c6) is the image of person 3 obtained based on the MCNET model, and (c7) is the image of person 3 obtained under the model of the present invention; Figure (d1) is the source image of person 4, (d2) is the driving video of person 4, (d3) is the image of person 4 obtained based on the FOMM model, (d4) is the image of person 4 obtained based on the TPSMM model, (d5) is the image of person 4 obtained based on the DaGAN model, (d6) is the image of person 4 obtained based on the MCNET model, and (d7) is the image of person 4 obtained under the model of the present invention;
[0098] Depend on Figure 3 It can be seen that compared with other models, the present invention has more advantages when the face is occluded or has large movements. For example, the opening range of the mouth in the last row clearly shows that our model is more consistent with the motion trajectory of the driving video. For example, in the third row, when the character in the driving video is in the eye-closed state, only our model and MCNET have their eyes closed.
[0099] The above description is merely a preferred embodiment of the present disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above-mentioned technical features, but should also encompass other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned inventive concept. For example, a technical solution formed by mutually replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A personalized portrait facial video reconstruction method based on meta-learning, characterized by: include: Step 1: Acquire multiple source images of T users, wherein the multiple source images of each user include images of the user under different lighting conditions, hairstyle characteristics, and facial expressions, and acquire a driving video corresponding to each source image, wherein the driving video is used to make the source image move; Step 2: Process each source image and driving video using a face detection algorithm to obtain a processed source image and a processed driving video; Step 3: Divide the processed source image and its corresponding driving video into T tasks. Each task contains a source image and its corresponding driving video processed by a user under different lighting, hairstyle characteristics, and facial expressions. Divide the T tasks into training and test sets according to a preset ratio. The processed source image and its corresponding driving video are used as a set of data. In the training and test sets, all the group data of each task are divided according to a preset ratio to obtain the support set and query set. Step 4: Obtain a MetaFace model, which includes an encoder module, a decoder module, and a rendering module. Initialize the parameters in the encoder module, the decoder module, and the rendering module. The initialized parameters form a local parameter vector. The initialized local parameter vector is used as the current parameter vector. Set the first and second iteration times. The initial values of the first and second iteration times are both 0. Step 5: For each task in the training set, the processed source image and its corresponding processed driving video in the support set of the task are input into the MetaFace model based on the current parameter vector. After passing through the encoder module, decoder module and rendering module, the predicted 2D video frame is output. Step 6: Based on the output predicted 2D video frame, the source image input to the MetaFace model, and the real 2D video frame corresponding to the processed driving video, calculate the facial key point loss and detail consistency loss. Based on the facial key point loss and detail consistency loss, update the current parameter vector through k-step stochastic gradient descent or Adam optimizer. The first iteration number is increased by one, and it is determined whether the first iteration number reaches the preset threshold. If the first iteration number does not reach the preset threshold, the updated parameter vector is used as the new current parameter vector, and the process returns to step 5. If the first iteration number reaches the preset threshold, the final parameter vector of the task is obtained, and then the final parameter vectors of all tasks are obtained. Step 7: Aggregate the final parameter vectors of all tasks into the current global parameter vector. Update the current global parameter vector through the Adam optimizer to obtain the updated global parameter vector. Step 8: The second iteration number is increased by one, and it is determined whether the second iteration number reaches the preset threshold. If the second iteration number does not reach the preset threshold, the parameter vector of each task in the updated global parameter vector is used as the current parameter vector, and the process returns to step 5. If the second iteration number reaches the preset threshold, the final global parameter vector is obtained, and then the final MetaFace model is obtained.
2. The personalized portrait facial video reconstruction method based on meta-learning according to claim 1, characterized in that Step 2 specifically includes: Performing face region detection on each source image using a face detection algorithm, calculating the proportion of the face region in the source image, and if the proportion of the face region in the source image is greater than or equal to a preset threshold, not processing the source image and using the source image as the processed source image; and if the proportion of the face region in the source image is less than the preset threshold, cropping the source image to obtain a processed source image. Similarly, for each frame image in the driving video, face area detection is performed on each frame image, and the proportion of the face area in the image is calculated. When the proportion of the face area in the image is greater than or equal to the preset threshold, the image is not processed. When the proportion of the face area in the image is less than the preset threshold, the image is cropped. After all frame images of all driving videos are processed as above, the processed driving video is obtained.
3. The personalized portrait facial video reconstruction method based on meta-learning according to claim 1, characterized in that Step 5 specifically includes: Step 5.1: In the encoder module based on the current parameter vector, for each task in the training set, the processed source image in the support set of the task is input into the encoder to obtain semantic features, where the semantic features include content parameters and texture features, and the content parameters include shape features, expression features, and posture features; The encoder is obtained by modifying the output dimension of the fully connected layer in the ResNet-50 model; Step 5.2: In the decoder module based on the current parameter vector, generate the final 3D face model based on the semantic features and the general FLAME 3D face model; Step 5.3: In the rendering module, the final 3D face model is rendered based on the geometric structure, texture information, and ambient lighting conditions, and the predicted 2D video frame is output.
4. The personalized portrait facial video reconstruction method based on meta-learning according to claim 3, characterized in that: Step 5.2 specifically includes: Step 5.2.1: Based on the content parameters, adjust the general FLAME 3D face model to obtain an initial 3D face model. Based on the linear blend skinning (LBS) algorithm, adjust the vertices of three target areas in the initial 3D face model to obtain three target areas after vertex adjustment. The three target areas include the area where the teeth are located, the area where the nose is located, and the area where the cheeks are located. For each target area after vertex adjustment, connect the adjusted vertices through patches to obtain a rough 3D face model. Step 5.2.2: Using UV mapping technology, based on texture features, the rough 3D face model is processed to obtain a refined 3D face model; Specifically, through the texture decoder T d , the texture features are processed to obtain the UV displacement map D, which is specifically achieved through the following formula: Where δ is the texture parameter, is the expression parameter, θ jaw is the mandibular posture parameter; Based on the UV displacement map D, the vertices of the rough 3D face model are adjusted to obtain the final position of the vertices in the UV space. This is achieved by the following formula: M' uv =M uv +D⊙N uv ; Among them, M' uv is the final position of the vertex in UV space, M uv is the vertex position of the rough 3D face model in UV space, N uv M uv The corresponding surface normal, ⊙ represents element-wise multiplication; The rough 3D face model is adjusted based on the final positions of the vertices in the UV space to obtain a fine 3D face model; Step 5.2.3: For the source image input to the encoder in step 5, obtain the corresponding processed driving video. For each frame of the image, identify the current posture θ of the user in the image and obtain the vertex position v′ of the user's face area. Based on the linear blending skinning (LBS) algorithm, adjust the vertex position of the user's face area to obtain the LBS deformed vertex position v", thereby obtaining the LBS deformed vertex position v" corresponding to each frame of the image; perform expression deformation on each frame of the image using the Blendshapes technology to obtain the expression deformed vertex position v expr , thus obtaining the vertex position v after expression deformation corresponding to each frame image expr ; Based on the processed driving video, the vertex position v" after LBS deformation and the vertex position v after expression deformation corresponding to each frame image expr , adjust the refined three-dimensional face model to obtain the final three-dimensional face model.
5. The personalized portrait facial video reconstruction method based on meta-learning according to claim 4, characterized in that: In step 5.2.3, the vertex positions of the user's face area are adjusted based on the linear blend skinning LBS algorithm to obtain the vertex position v" after LBS deformation. This is specifically achieved through the following formula: Among them, K b is the total number of bones, ω i is the weight of vertex v′ relative to the i-th bone, T i (θ) is the transformation matrix of the current posture θ.
6. The personalized portrait facial video reconstruction method based on meta-learning according to claim 4, characterized in that: In step 5.2.3, the expression of each frame image is deformed by Blendshapes technology to obtain the vertex position v after expression deformation expr , which is specifically achieved through the following formula: in, is the average face shape, K e is the total number of Blendshape models, φ j is the expression coefficient corresponding to each frame image, B j is the j-th blendshape model.
7. The personalized portrait facial video reconstruction method based on meta-learning according to claim 1, characterized in that: In step 7, the current global parameter vector is updated through the Adam optimizer to obtain the updated global parameter vector, which is specifically achieved through the following formula: Among them, φ′ is the updated global parameter vector, φ n is the final parameter vector of the nth task, φ is the current global parameter vector, and β is the outer loop learning rate.