Monocular image-driven efficient three-dimensional face modeling and rendering method

By using generative models and 3D Gaussian reconstruction models, and leveraging monocular portrait images and their annotation information, combined with generative adversarial loss and pseudo-supervised loss functions, the problems of high data acquisition cost and poor generalization in monocular image-driven 3D face reconstruction are solved, achieving efficient and stable 3D face modeling and rendering.

CN121074255APending Publication Date: 2025-12-05BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511194836.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing technologies rely on large-scale 3D datasets or video data for monocular image-driven 3D face reconstruction, resulting in high data acquisition costs and poor generalization, making it difficult to effectively utilize monocular images for efficient 3D face modeling and rendering.

Method used

By establishing a generative model and a 3D Gaussian reconstruction model, training is performed using monocular human portrait images and their annotation information. Combined with generative adversarial loss and pseudo-supervised loss functions, end-to-end 3D face reconstruction is achieved.

Benefits of technology

It reduces the difficulty and cost of data collection, improves the generalization and reconstruction quality of 3D face reconstruction, and achieves fast and efficient 3D face rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121074255A_ABST
    Figure CN121074255A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision and computer graphics, in particular to a monocular image-driven efficient three-dimensional face modeling and rendering method, which comprises the following steps of: acquiring a plurality of monocular portrait images and annotation information acquired by a camera, and establishing a portrait data set; establishing a generative model, wherein the generative model generates a portrait image from the Gaussian noise; pre-training the generated model to obtain a pre-trained 3D decoder; establishing a three-dimensional Gaussian reconstruction model, wherein the three-dimensional Gaussian reconstruction model generates a portrait image from the input image; performing joint fine tuning on the three-dimensional Gaussian reconstruction model to obtain a trained three-dimensional Gaussian reconstruction model; inputting a to-be-reconstructed monocular image into the trained three-dimensional Gaussian reconstruction model, and performing reasoning and rendering to obtain a reconstructed portrait image; according to the method, the training data collection difficulty can be reduced, and the three-dimensional face reconstruction generalization is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and computer graphics, specifically to a monocular image-driven, efficient method for 3D face modeling and rendering. Background Technology

[0002] Reconstructing 3D human figures from monocular images has long been a focal point in computer vision and computer graphics, particularly important for numerous applications such as Augmented Reality (AR) and Virtual Reality (VR). Traditional techniques, along with recent advancements like Neural Radiance Field (NeRF) and 3D Gaussian Sputtering (3DGS), have made significant progress in reconstructing static, non-facial 3D objects using multi-view image information. 3DGS, in particular, excels in both the clarity of the reconstruction results and its fast rendering speed. However, extending 3DGS-based methods to 3D reconstruction from monocular images remains challenging. Compared to multi-view images, monocular images lack complete object information, leading to ambiguity in the reconstruction results for invisible areas, resulting in blurred reconstruction outcomes.

[0003] Recent studies have explored the use of 3D Morphable Models (3DMMs) to model the coarse geometric information of 3D human images. GAGAvatar, proposed by Xuangeng Chu et al., is a representative example. It uses multiple neural networks to predict the 3D attributes of the 3D Morphable Model, trained on a monocular video dataset. While this approach utilizes the 3D Morphable Model as a geometric prior, it requires a large amount of monocular video data and frame-by-frame 3D face parameter annotations for computation and driving. Furthermore, the use of a neural rendering module reduces the accuracy of geometric reconstruction. SynShot, proposed by Wojciech Zielonka et al., is another representative example. It constructs a 3D Gaussian generative model based on the 3D Morphable Model on a multi-view dataset, then uses this generative model as a prior to fit the input monocular human image, obtaining the latent vector representation of the generative model. Although this method uses the generative model as a prior to resolve ambiguity in the reconstruction results, it requires a large amount of high-quality 3D face data to train the generative model.

[0004] Current methods rely on large-scale 3D datasets or video data; however, collecting such data is costly and limited in scale. For example, the commonly used 3D dataset Objaverse-XL contains approximately 10 million 3D objects, while the 2D image dataset LAION contains over 5 billion images—a difference of 500 times. For facial recognition data, this disparity in data size is even more pronounced due to sensitive issues such as identity and ethnicity. Summary of the Invention

[0005] In view of the above problems, the present invention provides a highly efficient 3D face modeling and rendering method driven by monocular images, which solves the technical problems of difficulty in collecting training data and poor generalization of 3D face reconstruction in the prior art.

[0006] This invention provides a highly efficient 3D face modeling and rendering method driven by a monocular image, comprising the following steps:

[0007] Step S1: Acquire multiple monocular human portrait images captured by the camera and their corresponding annotation information, and use the human portrait images and their corresponding annotation information as samples of the dataset to establish a human portrait dataset; the annotation information includes camera intrinsic and extrinsic parameters, three-dimensional deformation parameters and segmentation mask;

[0008] Step S2: Establish a generative model. The generative model processes Gaussian noise sequentially through a mapping network, a 3D decoder, and a 3D Gaussian splashing process to generate a portrait image. The generative model is pre-trained based on the portrait dataset and generative adversarial loss to obtain a pre-trained 3D decoder.

[0009] Step S3: Establish a 3D Gaussian reconstruction model. The 3D Gaussian reconstruction model sequentially processes the input image through a pre-trained image encoder, a bridging network, a pre-trained 3D decoder, and 3D Gaussian splash rendering to obtain a portrait image. Based on the portrait dataset and the total loss function, the 3D Gaussian reconstruction model is jointly fine-tuned to obtain a trained 3D Gaussian reconstruction model.

[0010] The total loss function includes a loss used to measure the correctness of the generated portrait image from the new perspective;

[0011] Step S4: Input the monocular image to be reconstructed into the trained 3D Gaussian reconstruction model and perform inference and rendering to obtain the reconstructed portrait image.

[0012] Preferably, step S1 specifically includes:

[0013] Step S1-1: Select a large-scale 2D face image dataset to obtain multiple monocular portrait images;

[0014] Steps S1-2: Use a 3D face reconstruction network to obtain the camera intrinsic parameters, extrinsic parameters, and 3D deformation parameters for each monocular portrait image; use a face segmentation model to obtain the segmentation mask for the face region;

[0015] Steps S1-3: Use the portrait images and corresponding annotation information as samples to create a portrait dataset.

[0016] Preferably, in step S2, the mapping network is composed of a multilayer perceptron, and the 3D decoder includes a convolutional layer, an upsampling layer, and an adaptive normalization layer;

[0017] The steps for generating portrait images using the 3D Gaussian sputtering process specifically include: using 3D deformation parameters as geometric information and static PBR materials as texture information; and rendering the 3D Gaussian primitives obtained from the 3D decoder through 3D Gaussian sputtering rendering based on the camera's intrinsic and extrinsic parameters to obtain a portrait image from the corresponding viewpoint.

[0018] Preferably, in step S2, the generative model further includes a discriminator, which is composed of a multi-layer convolutional network;

[0019] The steps for pre-training the generative model based on the human image dataset and generative adversarial loss specifically include:

[0020] Step 21: Generate a portrait image by 3D Gaussian splashing, input the generated image and the portrait images in the portrait dataset into the discriminator, and perform error backpropagation based on the difference between the discriminator's prediction result and the actual result to adjust the network weights of the discriminator;

[0021] Step 22: Input the generated image into the discriminator, perform backpropagation of the error based on the discriminator output value, and adjust the network weights of the mapping network and the 3D decoder.

[0022] Step 23: Return to step 21 until the preset number of training iterations are completed, and the pre-trained 3D decoder is obtained.

[0023] Preferably, in step S3, the pre-trained image encoder is pre-trained on a large-scale 2D image dataset, and the pre-trained image encoder converts the input image into a feature map;

[0024] The bridging network includes a multi-layer transformer network; the bridging network uses a self-attention mechanism to obtain image features and latent vectors from feature maps and learnable tokens;

[0025] The pre-trained 3D decoder converts latent vectors into 3D Gaussian primitives, and then renders the 3D Gaussian primitives using 3D Gaussian sputtering rendering to obtain a portrait image.

[0026] Preferably, in step S3, the three-dimensional Gaussian reconstruction model further includes an EMA version of the bridging network and a region mask prediction head;

[0027] The EMA version of the bridging network is a copy obtained by exponentially weighted averaging of the parameters of each layer of the bridging network; the region mask prediction head can process the image features to obtain the predicted segmentation mask.

[0028] Preferably, in step S3, the step of jointly fine-tuning the 3D Gaussian reconstruction model based on the portrait dataset and the total loss function includes:

[0029] Step 31: Obtain the mask loss of the region mask prediction head;

[0030] Step 32: Input the portrait images from the portrait dataset into the 3D Gaussian reconstruction model to obtain the rendered portrait image and the predicted segmentation mask, and calculate the reconstruction loss and image quality loss;

[0031] Step 33: Randomly generate new camera extrinsic parameters. The 3D Gaussian reconstruction model obtains portrait images from the portrait images in the portrait dataset under the new perspective based on the new camera extrinsic parameters, and calculates the pseudo-supervision loss of the new perspective.

[0032] Step 34: Based on the reconstruction loss, mask loss, image quality loss and new perspective pseudo-supervision loss, obtain the total loss function and perform error backpropagation to adjust the network weights of each network in the 3D Gaussian reconstruction model, thus completing the joint fine-tuning of the 3D Gaussian reconstruction model.

[0033] Preferably, step 32 specifically includes:

[0034] Calculate the real human portrait image x in the portrait dataset gt The final rendered portrait image x recon L1 distance and true segmentation mask m gt The reconstruction loss is obtained from the L1 distance between the predicted segmentation mask m and the target segmentation mask m.

[0035] Calculate x gt With x recon The SSIM loss between them serves as the image quality loss.

[0036] Preferably, step 33 specifically includes:

[0037] New camera extrinsic parameters are randomly generated. The new camera extrinsic parameters are used to perform 3D Gaussian splashing on the 3D Gaussian primitives obtained by the pre-trained 3D decoder to render a new perspective image and a segmentation mask.

[0038] The new perspective image is processed by a pre-trained image encoder and then by the EMA version of the bridging network to obtain the latent vector and image features of the new perspective image. The image features are then input into the region mask prediction head to obtain the predicted segmentation mask of the new perspective image processed by the EMA version of the bridging network.

[0039] The pseudo-supervised loss for the new viewpoint is obtained from the latent vectors of the new viewpoint image processed by the EMA version of the bridging network and the predicted segmentation mask, and is expressed as follows:

[0040]

[0041] in, Let m represent the pseudo-supervised loss from the new perspective, ||·||2 represent the L2 norm, and m represent the L2 norm. n w represents the segmentation mask of the 3D Gaussian primitive obtained by the bridging network in the new view and the latent vector of the original view, while m′ and w′ represent the predicted segmentation mask and latent vector of the new view obtained by the EMA version of the bridging network.

[0042] Compared with the prior art, the present invention has at least the following beneficial effects:

[0043] (1) This invention uses only monocular human images and establishes corresponding annotation information, without relying on the three-dimensional or multi-view images required by traditional methods. It can directly utilize large-scale and diverse two-dimensional image datasets for training, effectively reducing the difficulty and cost of data collection.

[0044] (2) This invention employs a head-aware guidance mechanism, effectively extracting and utilizing key features in 2D portrait images through segmentation masks and bridging networks, thus suppressing the negative impact of irrelevant perturbations such as background on 3D reconstruction. This invention provides a novel perspective pseudo-supervised loss function calculation, enabling the model to directly improve the quality of the synthesized portrait images from the novel perspective during training, thereby ensuring the realism and detail consistency of the novel perspective rendering. The entire framework is optimized end-to-end with reconstruction quality as the goal, further improving the stability and efficiency of training.

[0045] (3) This invention combines the powerful capabilities of end-to-end neural networks in feature extraction and representation with the high efficiency of three-dimensional Gaussian splashing technology in real-time three-dimensional rendering, thereby achieving fast, efficient and stable reconstruction of three-dimensional human figures. Attached Figure Description

[0046] The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of the invention.

[0047] Figure 1 The flowchart of the efficient 3D face modeling and rendering method driven by monocular image provided by the present invention is shown.

[0048] Figure 2 The network structure diagram of the efficient 3D face modeling and rendering method driven by monocular image provided by the present invention is shown.

[0049] Figure 3 An example image of the three-dimensional human portrait reconstruction result provided by the present invention. Detailed Implementation

[0050] To better understand the above-described objectives, features, and advantages of the present invention, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other. Furthermore, the present invention can be implemented in other ways different from those described herein; therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.

[0051] The purpose of this invention is to deeply integrate an end-to-end neural network with a 3D Gaussian model based on a 3D face model, and to design a training strategy and pseudo-supervised loss function to achieve training on a large-scale image dataset. First, this invention establishes a dataset of monocular images and labeled information, provides a pre-training strategy based on a 3D perceptual generative adversarial network, and builds a generative model capable of establishing a high-quality latent space representing 3D human images. Then, on the same dataset, this invention fine-tunes a large-scale bridging network and a 3D decoder to achieve end-to-end reconstruction from image features to 3D human images. During the fine-tuning stage, to ensure the viewpoint generalization of the end-to-end network reconstruction results, this invention utilizes a bridging network to obtain pseudo-supervised signals from new viewpoint images to supervise the rendering results of the reconstructed 3D face from the new viewpoint. This invention reduces the dependence on difficult-to-collect data and provides a highly generalizable 3D face reconstruction model.

[0052] like Figure 1 , Figure 2 As shown, this invention discloses an efficient 3D face modeling and rendering method driven by a monocular image. The specific implementation steps are as follows:

[0053] Step S1: Acquire multiple monocular human portrait images captured by the camera and their corresponding annotation information, and use the human portrait images and their corresponding annotation information as samples of the dataset to establish a human portrait dataset; the annotation information includes camera intrinsic and extrinsic parameters, three-dimensional deformation parameters and segmentation mask;

[0054] This invention uses monocular human portrait images acquired by a camera as a basis, and obtains annotation information through preprocessing to establish a dataset.

[0055] In some embodiments, a large-scale 2D face image dataset of real-world portraits can be selected, such as the FFHQ dataset (containing 140,000 face images).

[0056] Large-scale 2D face image datasets only provide information such as face images and 2D key points, without providing annotation information such as camera intrinsic and extrinsic parameters, 3D deformation parameters, and segmentation masks to represent the 3D shooting state. This invention obtains the annotation information through preprocessing, thereby obtaining large-scale usable data.

[0057] The preprocessing process includes: predicting the camera intrinsic and extrinsic parameters and 3D deformation parameters for each image using a 3D face reconstruction method; and obtaining the segmentation mask for the face region using a face segmentation model. The camera intrinsic and extrinsic parameters, 3D deformation parameters, and segmentation mask are then used together as the annotation information.

[0058] The 3D face reconstruction method can be an existing deep learning-based 3D face reconstruction network, such as DECA or EMOCA. Based on the input face image, the 3D face reconstruction method automatically predicts and outputs camera intrinsic parameters, camera extrinsic parameters, and 3D deformation parameters for 3D reconstruction. The face segmentation model can employ currently common image semantic segmentation networks, such as BiSeNet or FaRL. This segmentation model processes the input face image, automatically generating segmentation masks that annotate the face regions, thereby effectively distinguishing the face from the background region and providing accurate region information for subsequent processing.

[0059] The three-dimensional deformation parameters include shape, expression, and pose parameters, which are used to describe the geometric structure, facial expression, and position orientation information of the human face, respectively.

[0060] Finally, a portrait dataset is established by using portrait images and corresponding annotation information as samples. This invention can also divide the dataset samples into a training set and a test set at a 9:1 ratio. The training set is used for training to learn the model parameters, while the test set is used to evaluate the model's generalization performance and recognition accuracy.

[0061] Step S2: Establish a generative model. The generative model processes Gaussian noise sequentially through a mapping network, a 3D decoder, and a 3D Gaussian splashing process to generate a portrait image. The generative model is pre-trained based on the portrait dataset and generative adversarial loss to obtain a pre-trained 3D decoder.

[0062] This invention provides a pre-training strategy based on a 3D perceptual generative adversarial network to establish a generative model that can build a high-quality latent space representing the latent vectors of a 3D human image through a mapping network and a 3D decoder, for subsequent joint fine-tuning.

[0063] The generative model of this invention is as follows Figure 2As shown, a vector is randomly sampled from a standard normal distribution (Gaussian distribution). The random noise vector is used as input, and the mapping network maps the random noise vector to the latent space to generate a latent vector. The 3D decoder decodes the latent vector into a three-dimensional Gaussian primitive with geometric details. Finally, based on the camera's intrinsic and extrinsic parameters and the three-dimensional deformation parameters, the three-dimensional Gaussian primitive is rendered by three-dimensional Gaussian sputtering to obtain the portrait image under the viewpoint corresponding to the camera's intrinsic and extrinsic parameters.

[0064] In some embodiments, the mapping network may be a multilayer perceptron (MLP) used to map random noise vectors to latent vectors.

[0065] In some embodiments, the 3D decoder may include convolutional layers, upsampling layers, and adaptive normalization layers for decoding latent vectors into three-dimensional Gaussian primitive parameters containing geometric details.

[0066] In some embodiments, the rendering process of 3D Gaussian sputtering includes: using 3D deformation parameters as geometric information, using static PBR materials as texture information, and using predefined 3D Gaussian primitive mask information to obtain the rendered image and segmentation mask under the corresponding viewpoint through 3D Gaussian sputtering rendering based on camera intrinsic and extrinsic parameters.

[0067] The generative model of the present invention also includes a discriminator, which can be composed of a multi-layer convolutional network and a prediction head, used to determine the probability that the portrait image and segmentation mask generated in the previous steps are true.

[0068] The present invention pre-trains the generative model based on a human portrait dataset, as described in detail below.

[0069] The process of generating a human portrait image by sequentially passing Gaussian noise through a mapping network, a 3D decoder, and a 3D Gaussian splashing process is called the generator image generation process.

[0070] The pre-training process includes alternating training of the discriminator and the generator. The discriminator training process includes: inputting images generated by the generator and portrait images from the portrait dataset into the discriminator, and performing error backpropagation based on the difference between the discriminator's prediction results and the actual results to adjust the discriminator's network weights.

[0071] The process of training the generator includes: inputting the image generated by the generator into the discriminator, backpropagating the error of the generator based on the output value of the discriminator, and adjusting the network weights of the mapping network and the 3D decoder in the generator.

[0072] After multiple alternating training processes of the discriminator and generator, the training of the generative model is finally completed. The resulting mapping network and 3D decoder in the generator can represent the high-quality latent space of the three-dimensional human image.

[0073] The above training process can be represented as:

[0074]

[0075] Where G(·) represents the generator, D(·) represents the discriminator, w represents the latent vector, x represents the ground truth image, and E x [·] represents the mathematical expectation of the real image. This represents a latent vector that follows a latent space distribution formed by processing Gaussian noise using a multilayer perceptron. Let ln represent the mathematical expectation of the latent vector, and let ln denote the natural logarithm. This represents minimizing the generator's loss and maximizing the discriminator's loss. 3D face parameters and camera information are omitted.

[0076] Through this training process, the generative model of this invention can realize the mapping from random Gaussian noise to a three-dimensional Gaussian human image, and finally obtain a pre-trained 3D decoder.

[0077] Step S3: Establish a 3D Gaussian reconstruction model. The 3D Gaussian reconstruction model sequentially processes the input image through a pre-trained image encoder, a bridging network, a pre-trained 3D decoder, and a 3D Gaussian splashing process to generate a portrait image. Based on the portrait dataset and the total loss function, the 3D Gaussian reconstruction model is jointly fine-tuned to obtain a trained 3D Gaussian reconstruction model.

[0078] The total loss function includes a loss used to measure the correctness of the generated portrait image from the new perspective.

[0079] The previous steps have yielded a model for generating random portraits from noise. In this step, to generate a portrait of a specified person from other angles, a 2D image encoder and bridging network are connected before the pre-trained 3D decoder to construct a 3D Gaussian reconstruction model. This model uses a self-supervised pre-trained image encoder to obtain feature maps. These feature maps and learnable tokens are then input into a multi-layer transformer network. A self-attention mechanism is used to obtain the latent vectors input to the pre-trained 3D decoder, which in turn generates the reconstruction result. A detailed description follows.

[0080] The 3D Gaussian reconstruction model sequentially processes the input image through a pre-trained image encoder, a bridging network, a pre-trained 3D decoder, and a 3D Gaussian splashing process to generate a portrait image.

[0081] The pre-trained image encoder is pre-trained on a large-scale 2D image dataset. In some embodiments, the pre-trained image encoder can be a VAE, DINO image encoder, etc. The pre-trained image encoder processes the 2D image to obtain a feature map representing the features of the 2D image.

[0082] The bridging network receives the feature map output by the pre-trained image encoder and maps the feature map to latent vectors, obtaining latent vectors and image features. The bridging network of this invention employs a multi-layer transformer structure, which can use a self-attention mechanism to obtain latent vectors and image features from the feature map and learnable tokens.

[0083] The 3D Gaussian reconstruction model of this invention also includes an Exponential Moving Average (EMA) version of the bridging network, which is a "smoothed" copy obtained by exponentially weighting the parameters of each layer of the bridging network during training. The EMA version of the bridging network can not only provide pseudo-supervision signals for new perspective images, but also effectively reduce the fluctuation of network parameters, improve the generalization ability of the model in the validation and inference stages, and achieve more stable reconstruction results by using the EMA version of the bridging network for inference and result generation.

[0084] The pre-trained 3D decoder is obtained in step S2 and is used to generate three-dimensional Gaussian primitive parameters from the latent vectors.

[0085] The mapping process from monocular image to latent vector in the above steps can be represented as follows:

[0086] [f, w] = B([E(x), v])

[0087] Where f represents image features, B(·) represents bridging network, E(·) represents pre-trained image encoder, v represents learnable token, w represents latent vector, and x represents real image.

[0088] The 3D Gaussian reconstruction model of this invention includes a region mask prediction head. Image features obtained from the bridging network are input into the region mask prediction head to obtain a predicted segmentation mask. By setting the region mask prediction head, the bridging network can focus on the human figure region in the image features, improving the quality of new perspective synthesis.

[0089] This invention includes a step of jointly fine-tuning a 3D Gaussian reconstruction model based on a human portrait dataset and a total loss function. It adopts a new perspective strategy for generating human images, which can improve the quality of the final generated image. The specific description is as follows.

[0090] (1) Obtain the mask loss of the region mask prediction head.

[0091] The portrait images in the dataset are input into a pre-trained image encoder, which then passes through a bridging network to obtain image features. These image features are then input into a region mask prediction head to obtain segmentation mask predictions. The mask loss is determined by the ground truth segmentation mask values ​​in the dataset, expressed as:

[0092] m = M(f)

[0093]

[0094] in, M represents the masking loss, M(·) represents the region masking prediction head, and m is the segmented masking prediction value. gt The true labeled values ​​of the segmentation mask in the dataset are denoted by ||·||1, which is the L1 norm. The L1 norm is used to avoid gradient explosion and gradient vanishing.

[0095] (2) Generate an image from the original viewpoint and calculate the loss.

[0096] The portrait images in the dataset are input into the 3D Gaussian reconstruction model for inference. This includes inputting a pre-trained image encoder, passing it through a bridging network and a 3D decoder, and then performing 3D Gaussian splashing on the camera intrinsic and extrinsic parameters and 3D deformation parameters corresponding to the portrait images in the dataset to render the portrait images. Finally, the model passes through a bridging network and a region mask prediction head to obtain a predicted segmentation map.

[0097] The loss of the rendered portrait image is calculated, including the reconstruction loss, masking loss, and image quality loss.

[0098] The reconstruction loss Used to calculate the real image x gt The final rendered portrait image x recon The L1 distance between them.

[0099] The image quality loss Used to calculate the real image x gt The rendered portrait image x recon The SSIM loss between the two primarily penalizes the structural similarity of the images.

[0100] (3) Generate a new perspective and calculate the pseudo-supervised loss of the new perspective.

[0101] The feature map is input into a bridging network to obtain a latent vector w. New camera extrinsic parameters are randomly generated, representing a new perspective different from the current input portrait image. The latent vector is then input into a 3D decoder, and the obtained 3D Gaussian primitives are subjected to 3D Gaussian splashing using the new camera extrinsic parameters to render an image with the new perspective.

[0102] The new perspective image is processed by a pre-trained image encoder, and then by the EMA version of the bridging network to obtain the latent vector and image features of the new perspective image processed by the EMA version of the bridging network; the image features are input into the region mask prediction head to obtain the predicted segmentation mask of the new perspective image processed by the EMA version of the bridging network.

[0103] Finally, the pseudo-supervised loss of the new perspective is calculated by the difference between the segmentation mask and the latent vector obtained from the input image x and the EMA version of the bridging network.

[0104] The expression for the above process is:

[0105] m n x n =G(w)

[0106] [f′,w′]=B ema ([E(x n ), v])

[0107] m′=M(f′)

[0108]

[0109] Where, m n and x n Let G(w) be the segmentation mask and the portrait image rendered from the input image x under the new viewpoint, respectively. Let G(w) represent the generator's processing of the latent vector w, and B... ema (·) represents the EMA version of the bridging network, f′ and w′ are the image features and latent vectors of the new viewpoint obtained by the EMA version of the bridging network, respectively, and m′ represents the segmentation mask of the new viewpoint obtained by the EMA version of the bridging network. Let ||·||2 represent the pseudo-supervised loss from the new perspective, and let ||·||2 represent the L2 norm. The L2 norm is used to calculate the loss to ensure the smoothness of the optimization.

[0110] Through the above steps, this invention establishes the reconstruction loss, image quality loss, and new perspective pseudo-supervision loss, and the final total loss function used in the 3D Gaussian reconstruction model. The expression is:

[0111]

[0112] Where, λ mask , λ SSIM and λ pseudo Adjustment parameters for mask loss, image quality loss, and pseudo-supervision loss from new perspectives.

[0113] During training, based on the total loss function Error backpropagation is performed to adjust the network weights of each network in the 3D Gaussian reconstruction model, thus completing the joint fine-tuning of the 3D Gaussian reconstruction model.

[0114] In some embodiments, the present invention employs a fixed number of 3D Gaussian primitives during pre-training and joint fine-tuning, and skips the densification and pruning steps of traditional 3D Gaussian methods to maintain the simplicity of training.

[0115] Step S4: Input the monocular image to be reconstructed into the trained 3D Gaussian reconstruction model and perform inference and rendering to obtain the reconstructed portrait image.

[0116] In this step, the monocular human image to be reconstructed is first input into a pre-trained 3D Gaussian reconstruction model. The model estimates and generates the 3D structure of the input image through forward inference. Then, the generated 3D Gaussian primitives are used for rendering, outputting the final reconstructed human image.

[0117] To illustrate the effectiveness of the method proposed in this invention, the following detailed description of the above technical solution of this invention is provided through a specific embodiment.

[0118] Example 1

[0119] This embodiment provides an efficient 3D face modeling and rendering method for the FFHQ dataset, including the following steps.

[0120] 1. Preparation of real-world portrait video dataset. Complete dataset selection, data preprocessing, and dataset partitioning.

[0121] 1.1 The dataset selected includes monocular human face images acquired by a camera. Specifically, to verify the generalization ability of the model, this implementation selects the FFHQ dataset (containing 140,000 face images) as the real dataset.

[0122] 1.2 Dataset preprocessing includes cropping face regions and predicting camera intrinsic and extrinsic parameters, as well as 3D face deformation model parameters, for each image using a 3D face reconstruction method. Then, a joint optimization approach is used to obtain more accurate parameters. Additionally, a face segmentation model is used to obtain face region masks.

[0123] 1.3 The dataset is divided into training and test sets according to a 9:1 ratio.

[0124] 2. Design model training constraints. During the model pre-training phase, the loss is the GAN loss. During the joint fine-tuning phase, the total model loss includes reconstruction loss, image quality loss, and pseudo-supervision loss. Specific constraint designs have been discussed in the invention description and will not be repeated here. The trade-off parameters and related hyperparameters for each sub-constraint of the total model loss are set as follows:

[0125] 3. Training the Gaussian reconstruction model. The backpropagation algorithm is used to update and optimize the network parameter weights until the model loss converges. In this example, the training and evaluation of the Gaussian reconstruction model of this invention are both completed on the PyTorch platform. The model of this invention is trained on four NVIDIA RTX 3090 GPUs (24GB each). This invention uses the Adam optimizer to optimize the parameters of the bridging network and the 3D decoder network. The model batch size is 32 during training, and iterative optimization is performed 1,000,000 times.

[0126] 4. After model training is complete, perform 3D portrait reconstruction. Input the monocular portrait image to be reconstructed into the trained 3D Gaussian reconstruction model and render it. The rendering result is the reconstructed 3D portrait. Figure 3 As shown, the present invention can effectively complete the task of 3D human portrait reconstruction and has good generalization on different data.

[0127] While the specific embodiments of the present invention depict actions or steps in a particular order, this should be understood as requiring such actions or steps to be performed in the specific order shown or in sequential order, or requiring all illustrated actions or steps to be performed to achieve the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations.

[0128] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A monocular image driven efficient 3D face modeling and rendering method, characterized in that, The method comprises the following steps: Step S1, obtaining a plurality of monocular portrait images collected by a camera and corresponding label information, and establishing a portrait data set by taking the portrait images and the corresponding label information as samples of the data set; the label information comprises camera internal and external parameters, three-dimensional deformation parameters and a segmentation mask; Step S2, establishing a generation model, wherein the generation model sequentially passes Gaussian noise through a mapping network, a 3D decoder and three-dimensional Gaussian splashing processing to generate a portrait image; the generation model is pre-trained based on the portrait data set and a generative adversarial loss, and a pre-trained 3D decoder is obtained; Step S3, establishing a three-dimensional Gaussian reconstruction model, wherein the three-dimensional Gaussian reconstruction model sequentially passes an input image through a pre-trained image encoder, a bridging network, the pre-trained 3D decoder and three-dimensional Gaussian splashing rendering to obtain a portrait image; the three-dimensional Gaussian reconstruction model is jointly fine-tuned based on the portrait data set and a total loss function, and a trained three-dimensional Gaussian reconstruction model is obtained; The total loss function comprises a loss for measuring the correctness of the generated portrait image under a new view angle; Step S4, inputting a to-be-reconstructed monocular image into the trained three-dimensional Gaussian reconstruction model and performing inference and rendering to obtain a reconstructed portrait image.

2. The monocular image driven high-efficiency three-dimensional face modeling and rendering method according to claim 1, characterized in that, Step S1 specifically comprises: Step S1-1, selecting a 2D large-scale face picture data set to obtain a plurality of monocular portrait images; Step S1-2, obtaining camera internal and external parameters and three-dimensional deformation parameters of each monocular portrait image by using a three-dimensional face reconstruction network; and obtaining a segmentation mask of a face region by using a face segmentation model; Step S1-3, establishing a portrait data set by taking the portrait images and the corresponding label information as samples of the data set.

3. The monocular image driven high-efficiency three-dimensional face modeling and rendering method according to claim 2, characterized in that, In step S2, the mapping network is composed of multiple layers of perception mechanisms, and the 3D decoder comprises a convolution layer, an up-sampling layer and an adaptive normalization layer; The step of generating a portrait image by the three-dimensional Gaussian splashing processing specifically comprises: taking the three-dimensional deformation parameters as geometric information and taking a static PBR material as texture information, rendering a three-dimensional Gaussian primitive obtained by the 3D decoder to obtain a portrait image under a corresponding view angle by three-dimensional Gaussian splashing rendering according to the camera internal and external parameters.

4. The monocular image driven high-efficiency three-dimensional face modeling and rendering method according to claim 3, characterized in that, In step S2, the generation model further comprises a discriminator, and the discriminator is composed of multiple layers of convolution networks; The step of pre-training the generation model based on the portrait data set and the generative adversarial loss specifically comprises: Step 21, generating a portrait image by three-dimensional Gaussian splashing processing, inputting the generated image and the portrait images in the portrait data set into the discriminator, and performing error back propagation according to the difference between the prediction result and the true result of the discriminator to adjust the network weight of the discriminator; Step 22, inputting the generated image into the discriminator, performing error back propagation based on the output value of the discriminator, and adjusting the network weight of the mapping network and the 3D decoder; Step 23, returning to step 21 until a preset number of training times is completed, and a pre-trained 3D decoder is obtained.

5. The monocular image driven high-efficiency three-dimensional face modeling and rendering method according to claim 2, characterized in that, In step S3, the pre-trained image encoder is obtained by pre-training on a large-scale 2D image data set, and the pre-trained image encoder converts an input image into a feature map; The bridge network comprises a multi-layer transformer network; the bridge network passes the feature map and the learnable token through a self-attention mechanism to obtain an image feature and a latent vector; The pre-trained 3D decoder converts the latent vector into a three-dimensional Gaussian primitive, and renders the three-dimensional Gaussian primitive to obtain a portrait image through three-dimensional Gaussian sputtering rendering.

6. The monocular image driven high-efficiency three-dimensional face modeling and rendering method according to claim 5, characterized in that, In step S3, the three-dimensional Gaussian reconstruction model further comprises an EMA version of the bridge network and a region mask prediction head; The EMA version of the bridge network is a copy obtained by exponentially weighted averaging the parameters of each layer of the bridge network; and the region mask prediction head can process the image feature to obtain a predicted segmentation mask.

7. The monocular image driven high-efficiency three-dimensional face modeling and rendering method according to claim 6, characterized in that, In step S3, the step of jointly fine-tuning the three-dimensional Gaussian reconstruction model based on the portrait dataset and the total loss function comprises: Step 31, obtaining a mask loss of the region mask prediction head; Step 32, inputting a portrait image in the portrait dataset into the three-dimensional Gaussian reconstruction model to obtain a rendered portrait image and a predicted segmentation mask, and calculating a reconstruction loss and an image quality loss; Step 33, randomly generating a new camera extrinsic parameter, and based on the new camera extrinsic parameter, the three-dimensional Gaussian reconstruction model obtains a portrait image in a new view angle from the portrait image in the portrait dataset, and calculates a new view angle pseudo-supervised loss; Step 34, based on the reconstruction loss, the mask loss, the image quality loss and the new view angle pseudo-supervised loss, the total loss function is obtained to perform error back propagation, adjust the network weights of each network in the three-dimensional Gaussian reconstruction model, and complete the joint fine-tuning of the three-dimensional Gaussian reconstruction model.

8. The monocular image driven high-efficiency three-dimensional face modeling and rendering method according to claim 7, characterized in that, The step 32 specifically comprises: computing a real portrait image x in a portrait dataset gt the L1 distance between the final rendered portrait image x recon and the real segmentation mask m gt the L1 distance between the predicted segmentation mask m and the real segmentation mask m, resulting in the reconstruction loss Compute x gt between x recon as the image quality loss 9. The monocular image-driven high-efficiency three-dimensional face modeling and rendering method according to claim 8, characterized in that, The step 33 specifically comprises: Randomly generating a new camera extrinsic parameter, and performing three-dimensional Gaussian sputtering on the three-dimensional Gaussian primitive obtained by the pre-trained 3D decoder with the new camera extrinsic parameter to render a new view angle image and a segmentation mask; Processing the new view angle image through the pre-trained image encoder, and then processing the new view angle image through the EMA version of the bridge network to obtain a latent vector and an image feature of the new view angle image; inputting the image feature into the region mask prediction head to obtain a predicted segmentation mask of the new view angle image processed by the EMA version of the bridge network; The new view angle pseudo-supervised loss is obtained from the latent vector and the predicted segmentation mask of the new view angle image processed by the EMA version of the bridge network, and the expression is: wherein, represents the new-view pseudo-supervised loss, ||·||2represents the L2 norm, m n and w represent the segmentation mask of the three-dimensional Gaussian super-pixel at the new view and the hidden vector at the original view obtained by the bridge network, and m', w' represent the predicted segmentation mask at the new view and the hidden vector obtained by the EMA version of the bridge network.