Face high-fidelity and drivable reconstruction method based on three-dimensional gaussian splatting

By combining the three-dimensional Gaussian splash model and the FLAME framework, the facial identity and expression base are decoupled, which solves the problems of low reconstruction efficiency and poor driving effect in the existing technology, and achieves high-fidelity face synthesis and drivable reconstruction.

CN118736108BActive Publication Date: 2025-10-10ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410723951.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-05
Publication Date
2025-10-10
Estimated Expiration
2044-06-05

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively decoupling identity and expression in high-fidelity three-dimensional reconstruction of faces, resulting in low reconstruction efficiency and poor driving effect.

Method used

Based on the three-dimensional Gaussian splash model, combined with the FLAME framework and the pre-trained face identity information extraction network, high-fidelity face reconstruction is achieved by decoupling the expression basis and identity basis. The FLAME parameters and camera parameters are optimized through supervised training to improve the accuracy of the driving.

Benefits of technology

It achieves complete decoupling of facial expression and identity, improves reconstruction effect and driving accuracy, enhances generation quality and geometric details, and has significantly improved efficiency compared to previous methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118736108B_ABST
    Figure CN118736108B_ABST
Patent Text Reader

Abstract

The application discloses a high-fidelity and drivable reconstruction method of a face based on three-dimensional Gaussian splashing. The application introduces a three-dimensional Gaussian splashing model, provides a high-fidelity and drivable reconstruction method of a face based on three-dimensional Gaussian splashing, realizes high-fidelity reconstruction of a face, guarantees the availability of a novel expression driving to the maximum, avoids the coupling problem of an expression base and an identity base, and guarantees the accuracy of expression driving.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of computer vision, and particularly relates to a high-fidelity and drivable face reconstruction method based on three-dimensional Gaussian splatting. BACKGROUND

[0002] In the past few decades, with the development of virtual reality technology and digital human technology, high-fidelity three-dimensional face reconstruction has been more widely applied, such as 3D content generation, games, virtual reality, etc. However, up to now, high-fidelity three-dimensional face reconstruction is still a major problem in computer vision. High-fidelity three-dimensional face reconstruction is roughly divided into the following steps: first, data acquisition, then preprocessing, then face expression and identity information extraction, and finally three-dimensional face reconstruction according to the relevant information obtained in the previous steps.

[0003] The human head shape space can be effectively divided into three parts of identity, expression and appearance, on this premise, a three-dimensional deformable model (3DMM) is proposed, which is a face deformable reconstruction technology, which realizes the face deformable driving by parameterizing the face expression and structure, but this method cannot well fit the texture of the face surface, and the performance of the expression fitting is also not ideal, in addition, its efficiency for face reconstruction is also not high. Subsequent research has extended this grid-based parametric head model by developing a multilinear model or a nonlinear model, and equipped with a hinge model that corrects the mixed shape, enhancing its performance.

[0004] In order to accurately capture the complex deformation related to facial expressions, recent advanced techniques have made progress by adding additional displacement maps that can respond to image input. In addition, machine learning-driven generative models such as GANs and StyleGAN have been integrated into current frameworks to improve the accuracy of facial texture and geometry modeling. However, despite these advances, the parameter models currently used are still mostly limited to capturing the geometry and appearance of facial regions at a fairly basic level through explicit mesh models. This limitation affects the realism quality of reconstruction and animation based on these models.

[0005] Three-dimensional avatar synthesis can be divided into explicit representation and implicit representation. Explicit representation is mainly based on mesh models, which have been developed for many years. In recent years, some researchers have created photo-realistic portraits using two-dimensional neural rendering technology, but these methods often ignore non-facial areas, or face the challenge of temporal inconsistency due to the lack of tight integration with three-dimensional geometry. Other methods focus on learning vertex offsets to more accurately capture head geometry details.

[0006] However, due to the limitations of mesh model representation and the challenges of differentiable rendering, these methods still suffer from geometric and texture artifacts in complex areas such as hair, eyes, and mouths. PointAvatar introduces a novel, deformable point-based approach that overcomes some of the limitations of mesh models, but at the cost of requiring too many points and a lot of training time.

[0007] On the other hand, implicit models utilize neural networks to create digital avatars. While extensive research has focused on achieving high fidelity, this often comes at the expense of training and inference efficiency. To improve efficiency and reduce computational overhead, innovative approaches such as volumetric primitives and local feature grids have been proposed. Furthermore, Flashavatar proposes a UV sampling strategy to enhance image realism.

[0008] Advances in geometric reconstruction have significantly improved the accuracy of geometric synthesis. Neural Radiance Fields (NeRF) is the most representative work in this field, which has demonstrated strong capabilities in handling complex objects and has achieved higher quality results. Some methods generate photo-realistic human portraits by optimizing additional continuous volumetric deformation fields, while others combine traditional 3DMMs methods and have the ability to generalize to new deformations. However, volume rendering methods rely on heavy sampling and alpha synthesis, which limits the inference speed.

[0009] Therefore, the recent 3D Gaussian splatting (3DGS) method uses a set of 3D Gaussian points to describe 3D real-world scenes, endowing these points with variable properties, demonstrating the feasibility and efficiency of photo-realistic novel view synthesis. Regarding animated avatar synthesis, some current methods use multi-layer perceptrons to create a deformation field from a canonical space to a deformation space. While these methods have made significant progress in photo-realistic avatar synthesis, they are unable to separate identity from avatar.

[0010] High-fidelity 3D reconstruction of the face requires not only 3D reconstruction of the face, but also video and voice drive. Therefore, how to decouple the identity and expression of the person is also an important issue. During the reconstruction process, the identity information of the person should be fixed as much as possible and the expression information should be optimized. If these two parts are optimized at the same time, it may lead to the coupling of expression and identity, which will have a negative impact on the driving effect. Summary of the Invention

[0011] To address these technical issues, the present invention introduces a 3D Gaussian splatter model to provide a high-fidelity and drivable method for facial reconstruction based on 3D Gaussian splatter. This method decouples the facial expression basis from the identity basis, improving the accuracy of the driven expression while maintaining good reconstruction quality.

[0012] The technical solutions specifically adopted in the present invention are as follows:

[0013] In a first aspect, the present invention provides a high-fidelity and drivable reconstruction method for a human face based on three-dimensional Gaussian splattering, comprising the following steps:

[0014] S1. Obtain a facial action video captured by a first user using a monocular camera, wherein the first frame of the facial action video needs to keep the face expressionless, while the remaining frames need to make different facial expressions;

[0015] S2, performing image processing on the facial action video of the first photographer, extracting a face portrait frame with the background removed and a face mask image from each video frame;

[0016] S3, extracting the identity vector of the first photographer in the FLAME framework from the first frame of the facial action video of the first photographer using a pre-trained face identity information extraction network;

[0017] S4, using the facial key points in each facial portrait frame in the facial action video of the first photographer as supervision labels, and performing supervised learning optimization on the FLAME parameters and camera parameters in the facial expression tracking network based on the FLAME framework, to obtain the optimized FLAME parameters and camera parameters corresponding to each facial portrait frame; the facial expression tracking network based on the FLAME framework uses the photographer's FLAME parameters and camera parameters as model inputs, the FLAME parameters including an identity vector, a facial expression vector, an eye posture vector, a chin posture vector, and an eye closure vector, and uses the facial key points of the facial portrait frame as model outputs; and in the supervised learning process, the identity vector in the FLAME parameters remains fixed, and the remaining FLAME parameters and camera parameters are optimized, thereby achieving decoupling of the facial expression basis and the identity basis vector;

[0018] S5, using the FLAME parameters and camera parameters of each facial portrait frame obtained by the first photographer through supervised learning optimization as driving parameter inputs of the three-dimensional Gaussian splash model, and using the corresponding facial portrait frame and facial mask image as supervision labels to perform supervised training on the three-dimensional Gaussian splash model;

[0019] S6. Replace the first photographer with the second photographer, and re-obtain the FLAME parameters and camera parameters of each facial portrait frame in the facial action video corresponding to the second photographer according to S1 to S4. Then, input the identity vector corresponding to the first photographer and the facial expression vector, eye posture vector, chin posture vector, eye closure vector, and camera parameters of each facial portrait frame corresponding to the second photographer as driving parameters into the three-dimensional Gaussian splash model trained in S5 to obtain the virtual human rendering effect of the first photographer driven by the second photographer.

[0020] Preferably, the monocular camera needs to keep a fixed position when photographing the first photographer and the second photographer.

[0021] Preferably, in S2, the image processing includes image cropping and background segmentation and removal operations.

[0022] Preferably, in S3, the face identity information extraction network is composed of an ArcFace face recognition model and a decoder. The face portrait frame is first subjected to the ArcFace face recognition model to extract facial features, and then the facial features are input into the decoder for 3DMM coefficient extraction to obtain the identity vector of the photographer.

[0023] Preferably, in S4, in the facial expression tracking network based on the FLAME framework, the identity vector, facial expression vector, eye posture vector, chin posture vector, and eye closure vector of the photographer in the model input are first converted into a transformation matrix through a linear mixed skinning function, and then the standard Mesh head model is spatially transformed through the transformation matrix to obtain the photographer Mesh head model corresponding to the current model input, and then the three-dimensional facial key points are extracted from the photographer Mesh head model and the three-dimensional facial key points are mapped to two dimensions using external camera parameters, the output two-dimensional facial key points and the true value label calculate the loss and reversely optimize other FLAME parameters and camera parameters in the model input except the identity vector based on the loss.

[0024] Preferably, in S5, when the three-dimensional Gaussian splash model is rendered, a Mesh head model of the photographer is constructed according to the FLAME parameters of the input facial portrait frame, and a Gaussian kernel is assigned to the center point of each triangular face of the Mesh head model. At the same time, a Gaussian kernel is assigned to each of the three lines connecting the center point of the triangular face to the three vertices of the triangular face, so that there are four Gaussian kernels on each triangular face, and the attributes of each Gaussian kernel include the position attribute, opacity attribute, rotation attribute, size attribute, and color attribute based on the viewing direction of the Gaussian point; then, the rendering viewing angle is determined using the input camera parameters, and the three-dimensional Gaussian kernel on the Mesh head model is mapped to a two-dimensional plane. Then, a tile-based rasterizer renderer is used to perform portrait rendering on the two-dimensional plane to obtain a facial portrait frame and a corresponding facial mask map.

[0025] Preferably, the positions of the four Gaussian kernels allocated to each triangle are as follows:

[0026]

[0027] Where: represents the position coordinates of the center point of the triangle, whose position is determined by the average value of the position coordinates x1, x2, and x3 of the three vertices of the triangle; x′1, x′2, and x′3 represent the position coordinates of the three Gaussian kernels located on the three connecting lines; n is a learnable parameter within the preset upper and lower limits.

[0028] As a preference, the loss function used for supervised training of the 3D Gaussian splatter model is for:

[0029]

[0030] Where:

[0031]

[0032] in is a mask vector, if the Gaussian kernel is visible at the current rendering view, is 0, otherwise Take 1; s is the Gaussian kernel size; ξ scaling ,λ α ,λ ssim ,λ scale ,λ invis All are hyperparameters; A render and A gt Render the face mask image and its true value label respectively; is the structural similarity index loss term; I render and I gt The facial portrait frames and their true value labels are rendered respectively.

[0033] In a second aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for high-fidelity and drivable reconstruction of a face based on three-dimensional Gaussian splattering as described in any of the schemes of the first aspect above is implemented.

[0034] In a third aspect, the present invention provides a computer electronic device, characterized in that it includes a memory and a processor;

[0035] The memory is used to store computer programs;

[0036] The processor is configured to implement the high-fidelity and drivable reconstruction method of a human face based on three-dimensional Gaussian splattering as described in any one of the solutions of the first aspect when executing the computer program.

[0037] Compared with the existing high-fidelity 3D face reconstruction and driving methods, the present invention has the following advantages:

[0038] Beneficial effects:

[0039] (1) The present invention adopts an expression-based and identity-based extraction method based on the FLAME framework, and decouples the two information during preprocessing, which improves the reconstruction effect.

[0040] (2) The present invention adopts a model based on the three-dimensional Gaussian splashing method and proposes a fully drivable high-fidelity portrait synthesis. This is a new portrait synthesis method that completely decouples identity and expression, can provide fully controllable driving and achieve high-fidelity face synthesis, and its efficiency is significantly improved compared to previous methods.

[0041] (3) The present invention adopts a point-based random initialization field, which greatly improves the generation quality and geometric details.

[0042] (4) The present invention adopts a redesigned model structure and additional loss terms, such as mask map loss terms and other regularization losses, to adapt to different FLAME parameter models. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The technical solution of this invention is further illustrated below using embodiments and specific implementation methods. In addition, some drawings are used in the process of illustrating the technical solution. Those skilled in the art can also derive other drawings and the intent of the present invention based on these drawings without making any creative efforts.

[0044] Figure 1 : Schematic diagram of the overall process structure of the high-fidelity and drivable reconstruction method of the face based on three-dimensional Gaussian splattering;

[0045] Figure 2 :Schematic diagram of the extracted identity base process structure based on high-fidelity and drivable reconstruction method of human face based on 3D Gaussian splattering;

[0046] Figure 3 : Schematic diagram of the structure of extracting expression base and posture parameters for high-fidelity and drivable reconstruction of human faces based on 3D Gaussian splattering;

[0047] Figure 4 :Schematic diagram of the 3D Gaussian splash model training structure for high-fidelity and drivable reconstruction method of human faces based on 3D Gaussian splash.

[0048] Figure 5 Some examples of driving results are given. The first column in the figure shows the image of photographer B, and the second column shows the face image of virtual person A. This example uses the expression of photographer B to drive the image of virtual person A. DETAILED DESCRIPTION

[0049] The technical solutions of the embodiments of the present invention are further described in detail below with reference to the accompanying drawings. It should be noted that the embodiments described herein are only for explaining the present invention and are not intended to limit the present invention in any way. Those skilled in the art may make various changes to the embodiments without departing from the spirit of the present invention.

[0050] The embodiments of this specification disclose a high-fidelity and drivable reconstruction method for human faces based on three-dimensional Gaussian splattering. This invention introduces a three-dimensional Gaussian splattering model and provides a high-fidelity and drivable reconstruction method for human faces based on three-dimensional Gaussian splattering. This method achieves high-fidelity three-dimensional facial reconstruction while maximizing the usability of novel expression-driven driving. It also avoids the coupling issue between the expression base and the identity base, while ensuring the accuracy of the driven expression representation.

[0051] In view of this, Figure 1 A schematic flow chart of a high-fidelity and drivable reconstruction method for a human face based on three-dimensional Gaussian splattering provided in a preferred embodiment of the present invention is shown. The specific implementation process of the method is described in detail below.

[0052] S1. Obtain a facial action video shot by a first shooter using a monocular camera, wherein the first frame of the facial action video needs to keep the face expressionless, while the remaining frames need to make different facial expressions.

[0053] In the embodiment of the present invention, when using a monocular camera to capture the face of a first subject, the camera position should be kept fixed, and the first subject should initially maintain a calm expression to ensure the accuracy of subsequent identity extraction from the first frame. The subject should then be able to express as many different expressions as possible to improve robustness during driving.

[0054] S2. Perform image processing on the facial action video of the first photographer, and extract a face portrait frame with the background removed and a face mask image from each video frame.

[0055] The image processing of the present invention is based on the ability to extract the corresponding face portrait frame and the face mask, and generally includes image cropping and background segmentation and removal operations. In an embodiment of the present invention, it can be achieved as follows:

[0056] Step 2.1: Crop the original video V. Usually, due to the high resolution of the camera, the original image will contain other areas besides the face. Through the cropping method, the final cropped image I crop It mainly contains information about the face of the person.

[0057] Step 2.2: Use the pre-trained portrait segmentation model to complete the cropped image I cropDo segmentation, remove the background of the portrait image, and get the face portrait frame as the true value label I for training gt , and also need to output the mask map of the cropped face portrait frame as another part of the training truth value label A gt .

[0058] S3. Extracting the identity vector of the first photographer under the FLAME framework from the first frame of the facial action video of the first photographer through a pre-trained face identity information extraction network.

[0059] In the present invention, the face feature extraction network and the 3DMM coefficient extraction module can be used to extract the face features from the face. gt Extract the identity feature latent vector based on the FLAME framework. It should be noted that FLAME (Faces Learned with an Articulated Model and Expressions) is a 3D face modeling framework that creates detailed and expressive face models by learning thousands of 3D scan models. Its main purpose is to achieve parametric description of the shape, expression, and posture of the human head. The FLAME framework can be the FLAME2020 framework. The identity feature latent vector is mainly the identity vector representation. In an embodiment of the present invention, Figure 2 As shown, the facial identity information extraction network employed can be composed of an ArcFace face recognition model and a decoder. The first facial portrait frame of a facial action video is first input into the ArcFace face recognition model to extract facial features. These facial features are then input into the decoder for 3DMM coefficient extraction to obtain the subject's identity vector. The decoder uses the 3DMM coefficient extraction algorithm to decode these facial features into a 300-dimensional identity vector based on the FLAME framework. This 3DMM coefficient extraction algorithm is state-of-the-art and can be implemented using the MICA module (code available at https: / / github.com / Zielon / MICA).

[0060] S4. The facial key points in each facial portrait frame in the facial action video of the first photographer are used as supervision labels, and the FLAME parameters and camera parameters in the facial expression tracking network based on the FLAME framework are optimized in a supervised manner to obtain the optimized FLAME parameters and camera parameters corresponding to each facial portrait frame; the facial expression tracking network based on the FLAME framework uses the photographer's FLAME parameters and camera parameters as model inputs, and the FLAME parameters include an identity vector, a facial expression vector, an eye posture vector, a chin posture vector and an eye closure vector, and uses the facial key points of the facial portrait frame as model outputs; and in the supervised learning process, the identity vector in the FLAME parameters remains fixed, and the remaining FLAME parameters and camera parameters are optimized, thereby achieving decoupling of the facial expression basis and the identity basis vector.

[0061] In an embodiment of the present invention, Figure 3 As shown in the figure, the specific process in the facial expression tracking network based on the FLAME framework is as follows: first, the identity vector, facial expression vector, eye posture vector, chin posture vector, and eye closure vector of the photographer in the model input are converted into a transformation matrix through a linear mixed skinning function, and then the standard Mesh head model is spatially transformed through the transformation matrix to obtain the photographer Mesh head model corresponding to the current model input, and then the three-dimensional facial key points are extracted from the photographer Mesh head model and the camera external parameters are used to map the three-dimensional facial key points to two dimensions, the output two-dimensional facial key points and the true value label calculate the loss and reversely optimize the other FLAME parameters and camera parameters in the model input except the identity vector based on the loss.

[0062] It should be noted that during the optimization process of this facial expression tracking network, each facial portrait frame in the facial action video of the first photographer is optimized frame by frame, that is, each facial portrait frame needs to use its own facial key points as supervision labels to optimize the input of the facial expression tracking network. A loss is made between the predicted value and the true value label of the facial key points to supervise the fitting optimization of the FLAME parameters, so that the facial key points of this facial portrait frame can be accurately obtained. The optimized model input can be used as a driving parameter in the subsequent three-dimensional Gaussian splash model.

[0063] The FLAME parameters of the three-dimensional Gaussian splash model finally input by the present invention contain a total of five vectors. The FLAME parameters can be expressed as:

[0064]

[0065] In addition to the aforementioned identity vector representation In addition, it also contains expression vector representation Eye pose vector representation Chin pose vector representation Eyelid closure vector representation However, it should be noted that during the supervised learning optimization of FLAME parameters and camera parameters in the facial expression tracking network based on the FLAME framework, the identity vector representation that has been extracted previously Fixed, only optimizing other FLAME parameters. This maintains the fundamental structure of the face during the iterative optimization process, decoupling the facial expression and identity bases and improving reconstruction quality. Furthermore, in addition to FLAME parameters, the optimization process also iteratively optimizes camera parameters, including but not limited to camera intrinsics and extrinsics. The output camera parameters determine the subsequent rendering perspective of the 3D Gaussian splatter model.

[0066] In addition, in the above embodiment of the present invention, the facial expression tracking network based on the FLAME framework can be constructed by itself, or it can be directly implemented using the Metrical Photometric Tracker, whose code can be found at https: / / github.com / Zielon / metrical-tracker.

[0067] S5. The FLAME parameters and camera parameters of each facial portrait frame obtained by the first photographer through supervised learning optimization are used as driving parameter inputs of the three-dimensional Gaussian splash model, and the corresponding facial portrait frames and facial mask images are used as supervision labels to perform supervised training on the three-dimensional Gaussian splash model.

[0068] It should be noted that the 3D Gaussian splatting model itself belongs to the existing technology. The embodiment of the present invention directly adopts the existing model code during implementation. For specific code, please refer to https: / / github.com / graphdeco-inria / gaussian-splatting.

[0069] In an embodiment of the present invention, Figure 4As shown, the Gaussian kernel in the three-dimensional Gaussian splash model is encrypted. The specific method is as follows: when the three-dimensional Gaussian splash model is rendered, the photographer's Mesh head model is constructed according to the FLAME parameters of the input facial portrait frame. The method of constructing the Mesh head model is consistent with the aforementioned S4, that is, the identity vector, facial expression vector, eye posture vector, chin posture vector, and eye closure vector are converted into a transformation matrix through a linear mixed skinning function, and then the standard Mesh head model is spatially transformed through the transformation matrix to convert it into a Mesh structure head model without texture. After the Mesh head model is obtained, a Gaussian kernel is assigned to the center point of each triangle of the Mesh head model, and a Gaussian kernel is assigned to each of the three lines from the center point of the triangle to the three vertices of the triangle, so that there are four Gaussian kernels on each triangle. After a total of four Gaussian kernels are assigned to a triangle, the density of the Gaussian kernel is greatly increased, thereby improving the fidelity of the final reconstruction.

[0070] The positions of the four Gaussian kernels assigned to each triangle are as follows:

[0071]

[0072]

[0073] Where: represents the position coordinates of the center point of the triangle, and its position is determined by the position coordinates x1, x2, and x3 of the three vertices of the triangle by calculating the average value; x′1, x′2, and x′3 represent the position coordinates of the three Gaussian kernels located on the three connecting lines; n is a parameter that can be learned within the preset upper and lower limits, and its data range is [n min ,n max ]. min ,n max There are two hyperparameters

[0074] In addition, in this three-dimensional Gaussian splash model, each Gaussian kernel contains the following attributes in is the location attribute of the Gaussian point, is the opacity property of the Gaussian point, is the rotation property of the Gaussian point, is the size property of the Gaussian points, is the color attribute of the Gaussian point based on the viewing direction.

[0075] After determining the Gaussian kernel on each triangular face of the Mesh head model, the input camera parameters are used to determine the rendering perspective, and the three-dimensional Gaussian kernel on the Mesh head model is mapped to a two-dimensional plane. Then, a tile-based rasterizer is used to efficiently render the portrait on the two-dimensional plane to obtain a high-fidelity face portrait frame I render , can be obtained through I gt In addition, the renderer will also output the rendered face mask image A render , through A gt Supervision is performed to enhance the final geometric reconstruction results.

[0076] Furthermore, the present invention uses the loss function when conducting supervised training on the three-dimensional Gaussian splash model. Each loss term in can be reasonably designed. In the embodiment of the present invention, the loss function used is for:

[0077]

[0078] Where:

[0079]

[0080] in is the regularization term for the Gaussian kernel that is invisible at the current rendering perspective, is a mask vector, if the Gaussian kernel is visible at the current rendering view, is 0, otherwise Take 1; s is the Gaussian kernel size; are the other four loss terms, ξ scaling ,λ α ,λ ssim ,λ scale ,λ invis All are hyperparameters; A render and A gt Render the face mask image and its true value label respectively; is the structural similarity index loss term; I render and I gt The facial portrait frames and their true value labels are rendered respectively.

[0081] S6. Replace the first photographer with the second photographer, and re-obtain the FLAME parameters and camera parameters of each facial portrait frame in the visual facial motion frequency corresponding to the second photographer according to S1 to S4. Then, input the identity vector corresponding to the first photographer and the facial expression vector, eye posture vector, chin posture vector, eye closure vector, and camera parameters of each facial portrait frame corresponding to the second photographer as driving parameters into the three-dimensional Gaussian splash model trained in S5 to obtain the virtual human rendering effect of the first photographer driven by the second photographer.

[0082] It's important to note that the monocular camera must remain fixed in the same position when capturing both the first and second subjects to avoid distorting the reconstruction results due to different shooting angles. During the S6's virtual human rendering process, the FLAME parameters and camera parameters of each facial portrait frame of the second subject can be replaced with the identity vector corresponding to the first subject. This allows the 3D Gaussian splatter model to render the first subject's virtual human, driven by the second subject's FLAME parameters and camera parameters.

[0083] In order to better demonstrate the effect of the algorithm implemented in steps S1 to S6 of the present invention on face driving, Figure 5 Some driving examples are given. The first column in the figure shows the image of photographer B, and the second column shows the face image of avatar A. This example uses the expression of photographer B to drive the image of avatar A.

[0084] Similarly, based on the same inventive concept, the present invention also provides a computer electronic device corresponding to the high-fidelity and drivable reconstruction method of a human face based on three-dimensional Gaussian splattering provided in the above embodiment, which includes a memory and a processor;

[0085] The memory is used to store computer programs;

[0086] The processor is configured to implement the high-fidelity and drivable reconstruction method of a human face based on three-dimensional Gaussian splattering as described in the aforementioned embodiment when executing the computer program.

[0087] Furthermore, the logic instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention.

[0088] Therefore, based on the same inventive concept, the present invention provides a computer-readable storage medium corresponding to a high-fidelity and drivable reconstruction method for a human face based on three-dimensional Gaussian splattering, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it can implement the high-fidelity and drivable reconstruction method for a human face based on three-dimensional Gaussian splattering as described in the aforementioned embodiment.

[0089] Therefore, based on the same inventive concept, the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, can implement the high-fidelity and drivable reconstruction method of a human face based on three-dimensional Gaussian splattering as described in the aforementioned embodiment.

[0090] It is understood that the storage medium may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Furthermore, the storage medium may be any medium capable of storing program code, such as a USB flash drive, a mobile hard drive, a magnetic disk, or an optical disk.

[0091] It is understandable that the above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0092] It should also be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the system described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here. In the various embodiments provided in this application, the division of steps or modules in the system and method is only a logical function division. In actual implementation, there may be other division methods, for example, multiple modules or steps can be combined or integrated together, and a module or step can also be split.

[0093] The embodiment described above is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Persons skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent substitution or equivalent transformation falls within the scope of protection of the present invention.

Claims

1. A high-fidelity and drivable reconstruction method for human faces based on three-dimensional Gaussian splattering, characterized by: The following steps are involved: S1. Obtain a facial action video captured by a first user using a monocular camera, wherein the first frame of the facial action video needs to keep the face expressionless, while the remaining frames need to make different facial expressions; S2, performing image processing on the facial action video of the first photographer, extracting a face portrait frame with the background removed and a face mask image from each video frame; S3, extracting the identity vector of the first photographer in the FLAME framework from the first frame of the facial action video of the first photographer using a pre-trained face identity information extraction network; S4, using the facial key points in each facial portrait frame in the facial action video of the first photographer as supervision labels, and performing supervised learning optimization on the FLAME parameters and camera parameters in the facial expression tracking network based on the FLAME framework, to obtain the optimized FLAME parameters and camera parameters corresponding to each facial portrait frame; the facial expression tracking network based on the FLAME framework uses the photographer's FLAME parameters and camera parameters as model inputs, the FLAME parameters including an identity vector, a facial expression vector, an eye posture vector, a chin posture vector, and an eye closure vector, and uses the facial key points of the facial portrait frame as model outputs; and in the supervised learning process, the identity vector in the FLAME parameters remains fixed, and the remaining FLAME parameters and camera parameters are optimized, thereby achieving decoupling of the facial expression basis and the identity basis vector; S5, using the FLAME parameters and camera parameters of each facial portrait frame obtained by the first photographer through supervised learning optimization as driving parameter inputs of the three-dimensional Gaussian splash model, and using the corresponding facial portrait frame and facial mask image as supervision labels to perform supervised training on the three-dimensional Gaussian splash model; S6. Replace the first photographer with the second photographer, and re-obtain the FLAME parameters and camera parameters of each facial portrait frame in the facial action video corresponding to the second photographer according to S1 to S4. Then, input the identity vector corresponding to the first photographer and the facial expression vector, eye posture vector, chin posture vector, eye closure vector, and camera parameters of each facial portrait frame corresponding to the second photographer as driving parameters into the three-dimensional Gaussian splash model trained in S5 to obtain the virtual human rendering effect of the first photographer driven by the second photographer.

2. The high-fidelity and drivable reconstruction method of a human face based on three-dimensional Gaussian splattering according to claim 1, characterized in that: The monocular camera needs to keep a fixed position when photographing the first photographer and the second photographer.

3. The high-fidelity and drivable reconstruction method of a human face based on three-dimensional Gaussian splattering according to claim 1, characterized in that: In S2, the image processing includes image cropping and background segmentation and removal operations.

4. The high-fidelity and drivable reconstruction method of a human face based on three-dimensional Gaussian splattering according to claim 1, characterized in that: In S3, the face identity information extraction network is composed of an ArcFace face recognition model and a decoder. The face portrait frame is firstly subjected to the ArcFace face recognition model to extract facial features, and then the facial features are input into the decoder to extract 3DMM coefficients to obtain the identity vector of the photographer.

5. The high-fidelity and drivable reconstruction method of a human face based on three-dimensional Gaussian splattering according to claim 1, characterized in that: In the S4, in the facial expression tracking network based on the FLAME framework, the identity vector, facial expression vector, eye posture vector, chin posture vector, and eye closure vector of the photographer in the model input are first converted into a transformation matrix through a linear mixed skinning function, and then the standard Mesh head model is spatially transformed through the transformation matrix to obtain the photographer Mesh head model corresponding to the current model input, and then the three-dimensional facial key points are extracted from the photographer Mesh head model and the three-dimensional facial key points are mapped to two dimensions using the camera external parameters. The output two-dimensional facial key points and the true value label calculate the loss and reversely optimize the other FLAME parameters and camera parameters in the model input except the identity vector based on the loss.

6. The high-fidelity and drivable reconstruction method of a human face based on three-dimensional Gaussian splattering according to claim 1, characterized in that: In the S5, when the three-dimensional Gaussian splash model is rendered, a Mesh head model of the photographer is constructed according to the FLAME parameters of the input facial portrait frame, and a Gaussian kernel is assigned to the center point of each triangular face of the Mesh head model. At the same time, a Gaussian kernel is assigned to each of the three lines connecting the center point of the triangular face to the three vertices of the triangular face, so that there are four Gaussian kernels on each triangular face, and the attributes of each Gaussian kernel include the position attribute, opacity attribute, rotation attribute, size attribute, and color attribute based on the viewing direction of the Gaussian point; then, the rendering viewing angle is determined using the input camera parameters, and the three-dimensional Gaussian kernel on the Mesh head model is mapped to a two-dimensional plane. Then, a tile-based rasterizer renderer is used to perform portrait rendering on the two-dimensional plane to obtain a facial portrait frame and a corresponding facial mask map.

7. The high-fidelity and drivable reconstruction method of a human face based on three-dimensional Gaussian splattering according to claim 6, characterized in that: The positions of the four Gaussian kernels assigned to each triangle are as follows: Where: represents the position coordinates of the center point of the triangle, whose position is determined by the average value of the position coordinates x1, x2, and x3 of the three vertices of the triangle; x′1, x′2, and x′3 represent the position coordinates of the three Gaussian kernels located on the three connecting lines; n is a learnable parameter within the preset upper and lower limits.

8. The high-fidelity and drivable reconstruction method of a human face based on three-dimensional Gaussian splattering according to claim 6, characterized in that: The loss function used for supervised training of the 3D Gaussian splatter model is for: Where: in is a mask vector, if the Gaussian kernel is visible at the current rendering view, is 0, otherwise Take 1; s is the Gaussian kernel size; ξ scaling ,λ α ,λ ssim ,λ scale ,λ invis All are hyperparameters; A render and A gt Render the face mask image and its true value label respectively; is the structural similarity index loss term; I render and I gt The facial portrait frames and their true value labels are rendered respectively.

9. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and when the computer program is executed by the processor, the high-fidelity and drivable reconstruction method of a human face based on three-dimensional Gaussian splattering according to any one of claims 1 to 8 is implemented.

10. A computer electronic device, characterized in that: including memory and processor; The memory is used to store computer programs; The processor is configured to implement the high-fidelity and drivable reconstruction method of a human face based on three-dimensional Gaussian splattering according to any one of claims 1 to 8 when executing the computer program.

Citation Information

Patent Citations

  • Three-dimensional face reconstruction method and device, equipment, medium and product

    CN115330947A

  • High-fidelity three-dimensional face reconstruction and generation method based on implicit neural function

    CN116071494A