Head model reconstruction method based on Gaussian splash and generative model

By combining Gaussian splattering and generative models, the problem of missing side and back structures of head model reconstruction in the prior art is solved, and the head model reconstruction effect with higher texture accuracy and structural integrity is achieved.

CN120219625APending Publication Date: 2025-06-27XIDIAN UNIV

Patent Information

Application Number
CN202510295770.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, due to the lack of side and back structures, the reconstruction quality of the head model is low, especially in terms of texture details and structural integrity.

Method used

The head model reconstruction method based on Gaussian splattering and generative models is adopted. The FLAME head model is fitted with shape parameters, expression and pose parameters optimization, and the 3D Gaussian splattering algorithm is used to train 3D Gaussian splattering algorithm. Combined with the generative model, a multi-view appearance image surrounding the head is generated by further training the 3D Gaussian splattering on the side and back to achieve a more complete head model reconstruction.

Benefits of technology

The texture accuracy and structural integrity of the head model were improved. The experimental results showed that the PSNR, SSIM and LPIPS indicators were superior to the existing technology, which significantly improved the detail fidelity and structural integrity of the reconstruction model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219625A_ABST
    Figure CN120219625A_ABST
Patent Text Reader

Abstract

The invention provides a head model reconstruction method based on Gaussian splash and a generative model. The method comprises the following steps: fitting a shot face of a person; the expression parameters and the posture parameters of the FLAME head model are optimized; training the 3D Gaussian ball based on a Gaussian splashing algorithm; and obtaining a complete head model reconstruction result based on the generative model. According to the method, the head appearance is subjected to multi-view synthesis through the generative model, a multi-view image which has consistent texture information and surrounds the head by 360 degrees can be inferred, and information loss of the input image in a non-front area is made up; and meanwhile, the multi-view image can also adopt a Gaussian splash algorithm to train the side and back areas of the head model, so that the head model with higher texture precision and more complete structure is reconstructed, and experimental results show that the detail fidelity of the reconstructed model is improved while the completeness of the head structure is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and relates to a method for reconstructing a head model, specifically to a method for reconstructing a head model based on Gaussian splatting and generative model, which can be applied to fields such as virtual reality and facial animation. Background Technique

[0002] Model reconstruction refers to the process of using computer vision algorithms to restore the three-dimensional structure of an object from a two-dimensional planar image. Head model reconstruction specifically aims at the complex and dynamically changing part of the human head, and through fusing multi-view images and depth data, a fine and realistic three-dimensional head model is restored. This technology aims to overcome the reconstruction errors and loss of model details caused by factors such as occlusion, dynamic expressions, lighting changes, and data noise. Compared with three-dimensional scene reconstruction, head model reconstruction not only requires higher texture details, but also poses more stringent requirements on the integrity of the model structure.

[0003] Head model reconstruction methods are divided into two categories: reconstruction methods based on three-dimensional deformable face models (3D Morphable Face Model, 3DMM) and traditional multi-view geometry reconstruction methods. Among them, the most representative in 3DMM is the FLAME model (Faces Learned with an Articulated Model and Expressions). The FLAME model parameterizes the shape, expression, and pose of the head by learning thousands of 3D scanned head models. However, this model cannot well restore the texture details, thus affecting the reconstruction quality of the head model. The Gaussian Splatting algorithm effectively solves the problem of low model texture details. As an emerging three-dimensional reconstruction algorithm, it has achieved remarkable results in the fields of novel view synthesis and model reconstruction. This method realizes high-quality reconstruction of the target by distributing a large number of 3D Gaussian spheres with their respective attributes (such as position, scale, rotation, opacity, and spherical harmonic coefficients, etc.) in space. Although the Gaussian Splatting algorithm can achieve high-quality reconstruction according to the training data, it is difficult to effectively reconstruct the areas outside the training data. With the development of generative models, the generative model based on triplane representation has reached an extremely realistic level in image generation. This model embeds two-dimensional feature maps on three orthogonal planes, projects and samples features for any point in space, thereby realizing efficient three-dimensional information expression and image generation.

[0004] To obtain better head model reconstruction quality, for example, the patent application with the publication number CN118736108A and the title "High-Fidelity and Drivable Facial Reconstruction Method Based on 3D Gaussian Splashing" discloses a head model reconstruction method based on 3D Gaussian splashing. This invention introduces a 3D Gaussian splashing model, which realizes high-fidelity 3D facial reconstruction while maximizing the usability of novel expression driving and can avoid the coupling problem between expression basis and identity basis. However, due to the limited head information collected by a monocular camera, there are missing structures on the side and back of the reconstructed head model, which affects the further improvement of the head model reconstruction quality. Summary of the Invention

[0005] The object of the present invention is to overcome the defects of the above-mentioned existing technologies, and propose a head model reconstruction method based on Gaussian splashing and generative model, which is used to solve the technical problem of low reconstruction quality caused by the missing structures on the side and back in the existing technologies.

[0006] To achieve the above object, the technical solutions adopted by the present invention include the following steps:

[0007] (1) Fit the captured human face:

[0008] Decode the facial expression video of a person captured by a monocular RGB camera, and perform shape parameter fitting of the FLAME head model on a neutral expression image in the decoded image sequence based on a face shape prediction model to obtain a FLAME head model with a neutral expression.

[0009] (2) Optimize the expression parameters and pose parameters of the FLAME head model:

[0010] Detect the positions of the facial key points of each image in the image sequence through a facial feature point detection model, and optimize the expression parameter θ and pose parameter δ of the FLAME head model through the detected positions of multiple key points to obtain the optimized expression parameter θ * and pose parameter δ * ;

[0011] (3) Train the 3D Gaussian sphere based on the Gaussian splashing algorithm:

[0012] According to the optimized expression parameter θ * and pose parameter δ * perform linear deformation on the FLAME head model, and use the Gaussian splashing algorithm to train the 3D Gaussian sphere on the surface of the initialized FLAME head model according to the image sequence and the deformed FLAME head model to obtain a frontal head model;

[0013] (4) Obtain the reconstructed result of the complete head model based on the generative model:

[0014] Generate multi-view appearance images that surround the head by 360 degrees for the frontal head model through the generative model, and use the same method as in step (3) to retrain the 3D Gaussian spheres on the sides and back of the frontal head model through this multi-view appearance image to obtain the complete head model.

[0015] Compared with the prior art, the present invention has the following advantages:

[0016] Through the generative model, the present invention performs multi-view synthesis on the head appearance, and can infer multi-view images that surround the head by 360 degrees with consistent texture information, making up for the lack of information in the non-frontal area of the input image; at the same time, the multi-view images can also use the Gaussian splashing algorithm to train the side and back areas of the head model, thereby reconstructing a head model with higher texture accuracy and more complete structure. Experimental results show that while ensuring the integrity of the head structure, the present invention improves the detail fidelity of the reconstructed model. Description of the Drawings

[0017] Figure 1 It is a flowchart for the implementation of the present invention.

[0018] Figure 2 It is a visualization result diagram for reconstructing the head model through different data sets of the present invention. Detailed Embodiment

[0019] The present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0020] Refer to Figure 1 , the present invention includes the following steps:

[0021] Step 1) Fit the captured human face:

[0022] Step 1a) Use a monocular RGB camera to record a facial expression video of a person's face facing forward for a period of time and decode it to obtain an image sequence. In this embodiment, the recording frame rate is 25 FPS, the recording duration is 60 s, and FFMPEG is used to decode the video to obtain the image sequence;

[0023] Step 1b) Based on the face shape prediction model, fit the shape parameters of the FLAME head model for a neutral expression image in the decoded image sequence to obtain the FLAME head model of the neutral expression;

[0024] The face shape prediction model is an open-source pre-trained model, including an identity encoding module and a geometry decoding module. The identity encoding module is stacked based on ResNet 100, encodes the face information in neutral expression images, and is not affected by lighting, expression, rotation, and occlusion. The geometry decoding module consists of multiple fully connected layers and ReLU activation functions, and can convert a set of shape parameters of implicit facial geometry and identity features output by the identity encoder into shape parameters suitable for the FLAME head model, obtaining the FLAME head model with neutral expression.

[0025] The FLAME head model is a three-dimensional deformable modeling framework jointly controlled by shape parameters, expression parameters, and pose parameters, used to describe the complete structure of the face and head. This model can jointly model facial geometry, muscle expression changes, and head pose, and has a complete topological structure covering the ears, top of the head, and back of the head, providing a structural basis for the fine reconstruction and dynamic driving of the head model. Among them, the shape parameters consist of 300 floating-point type values, the expression parameters consist of 100 floating-point type values, and the pose parameters consist of 6 floating-point type values.

[0026] Step 2) Optimize the expression parameters and pose parameters of the FLAME head model:

[0027] Step 2a) Detect the positions of face key points in each image of the image sequence through a facial feature point detection model, and obtain the detected face key points f.

[0028] The facial feature point detection model includes a feature encoding module and a key point decoding module. The feature encoding module extracts features from each image. The key point decoding module decodes the extracted high-dimensional facial features to obtain the 2D coordinates of 68 face key points. Among them, the 68 face key points include 17 key points identifying the lower jaw contour, 5 key points each identifying the left and right eyebrow regions and the nose tip region, 4 key points identifying the nose bridge region, 6 key points each identifying the left and right eye contours, and 20 key points identifying the lip contour.

[0029] Step 2b) Initialize the iteration number as s, the maximum optimization number as S, S≥10. In the s-th iteration, the expression parameters and pose parameters of the FLAME head model for optimization are respectively represented as θ s 、δ s , let s = 1. In this embodiment, the maximum optimization number S = 10;

[0030] Step 2c) Perform linear deformation on the FLAME head model with neutral expression according to the expression parameters θ s and pose parameters δ s in the s-th iteration to obtain the deformed FLAME head model.

[0031] The FLAME head model deforms linearly according to the expression parameter θ s on the neutral-expression FLAME head model, and then uses the pose parameter δ s to perform translation and rotation operations on the linearly deformed FLAME head model to obtain the adjusted FLAME head model. The deformation formula of the FLAME head model is:

[0032]

[0033] where represents the shape parameter, represents the expression parameter, represents the pose parameter, the function LBS is the linear blend skinning function, represents rotating the vertex around the joint and performing linear smoothing through the blending weight w, the specific definition of the facial vertex represented by

[0034]

[0035] where represents the standard FLAME model without individual differences, no expression, and no pose, is the shape basis function that controls the static shape deviation by the shape parameter, is the pose deformation basis function that controls the non-rigid pose deformation by the pose parameter, is the expression deformation basis function that controls the facial muscle movement deformation by the expression parameter; the shape basis function in represents the a-th shape basis, β a is the weight corresponding to the a-th shape basis, is the number of shape bases; the pose deformation basis function in represents the rotation matrix of the b-th joint in the current pose, represents the rotation matrix of the b-th joint without pose, represents the b-th pose deformation basis, B represents the number of rotating joints; the expression deformation basis function in e represents the e-th expression deformation basis, is the weight corresponding to the e-th expression deformation basis, is the number of expression deformation bases;

[0036] Step 2d) Calculate the loss value L through the face key points ν of the deformed FLAME head model and the detected face key points f corresponding to it s and through Ls Optimize the expression parameters θ s and the pose parameters δ s where the loss value L s The calculation formula and the optimization formula are respectively:

[0037]

[0038] where, ∑· represents the summation operation, ||·||2 represents the operation of calculating the L2 norm, and f i represents the 2D coordinates of the i-th facial key point detected, and ν i represents the 3D coordinates of the i-th facial key point of the FLAME head model. Φ(·) represents the FLAME head model to perform a deformation operation according to the expression parameter θ s and the pose parameter δ s The projection operation from 3D coordinates to 2D coordinates is represented by Π(·). θ s ', δ s ' respectively represent the optimization results of θ s , δ s , and η θ , η δ respectively represent the learning rates of the expression parameter θ and the pose parameter δ. respectively represent the partial derivative operations of L s with respect to θ s , δ s ;

[0039] Step 2e) Judge whether s > S holds. If so, obtain the optimized expression parameter θ of the FLAME head model * and the pose parameter δ * , otherwise, set s = s + 1 and execute step (2c).

[0040] Step 3) Train the 3D Gaussian sphere based on the Gaussian splash algorithm:

[0041] Step 3a) Perform a linear deformation on the FLAME head model according to the optimized expression parameter θ of the FLAME head model * and the pose parameter δ * ;

[0042] Step 3b) Initialize the 3D Gaussian sphere on the surface of the FLAME head model. The attributes of the 3D Gaussian sphere are defined as:

[0043] G = {p, r, z, o, SH}

[0044] where p, r, z, o, and SH respectively represent the position, rotation, scaling coefficient, opacity, and spherical harmonic coefficient of the 3D Gaussian sphere;

[0045] Step 3c) Initialize the number of training times as t, the maximum number of iterations as T, where T≥6000, and the 3D Gaussian sphere property parameters used for optimization in the t-th training are denoted as G t , and let t = 1. In this embodiment, the maximum number of iterations T = 6000;

[0046] Step 3d) Adopt the Gaussian splashing algorithm and render the image I through the 3D Gaussian spheres on the surface of the FLAME head model pred ;

[0047] The projection formula of the 3D Gaussian sphere onto the 2D image plane is:

[0048] Σ′=JWΣW T J T

[0049] where J represents the affine approximation Jacobian matrix of the projective transformation, which can convert a 3D vector into a 2D vector, W is the viewing transformation matrix, Σ is the covariance matrix of the 3D Gaussian sphere, and J T , W T are the transposes of J and W respectively, and Σ′ is the 2D covariance matrix projected onto the image plane;

[0050] The center position and color of the projected 2D Gaussian sphere can be directly obtained from the parameters of the 3D Gaussian sphere, while the opacity of the 2D Gaussian sphere needs to be calculated according to the opacity and covariance matrix of the 3D Gaussian sphere. The calculation formula is:

[0051]

[0052] where o′ is the opacity of the 2D Gaussian sphere, o is the opacity of the 3D Gaussian sphere, Σ is the covariance matrix of the 3D Gaussian sphere. In addition, the projected 2D Gaussian spheres are sorted according to their depth values to correctly handle the occlusion relationship during the subsequent rendering process;

[0053] By fusing the colors and opacities of all overlapping 2D Gaussian spheres at the same position, the rendered image I is obtained pred , and the rendering formula is:

[0054]

[0055] where c x represents the color after projection of the x-th 2D Gaussian sphere, o′ x represents the opacity of the x-th 2D Gaussian sphere, o′ y represents the opacity of each 2D Gaussian sphere before the x-th 2D Gaussian sphere, represents the set of all Gaussian spheres sorted by depth at the current pixel position, and Π· represents the product operation. The color of the 2D Gaussian sphere at the x-th position is determined by calculating the cumulative visibility occlusion;

[0056] Step 3e) Calculate the loss value of the 3D Gaussian sphere on the surface of the FLAME head model by rendering the image I pred and the corresponding ground truth image I in the image sequence gt Calculate the loss value of the 3D Gaussian sphere on the surface of the FLAME head model and optimize the 3D Gaussian sphere attribute parameters G through the loss value The calculation formula of the loss value L and the formula for optimizing the 3D Gaussian sphere attribute parameters G t are as follows respectively: t

[0057]

[0058] where is the mean absolute error is the structural similarity error, λ represents the loss term weight, U is the total number of pixels of I pred I pred (u), I gt (u) respectively represent the pixel values of the u-th pixel of I pred and I gt |·| represents the absolute value operation, SSIM represents calculating the structural similarity between two images, G t ' represents the optimization result of G t η G represents the learning rate of each attribute of the 3D Gaussian sphere represents the partial derivative operation of G t ;

[0059] Step 3f) Judge whether t > T holds. If so, obtain the trained frontal head model; otherwise, set t = t + 1 and execute step (3d).

[0060] Step 4) Obtain the reconstruction result of the complete head model based on the generative model:

[0061] Step 4a) Generate multi-view appearance images that surround the head by 360 degrees for the frontal head model through the generative model. In this embodiment, the generative model generates a total of 8 multi-view appearance images that surround the head by 360 degrees, and one image is generated every 45 degrees;

[0062] ​The generative model is trained based on a 3D perception generative adversarial network, including a view encoding module and an appearance generation module. The view encoding module is mainly used to extract high-dimensional semantic feature information from the frontal head image. First, its appearance features are extracted through a multi-layer convolutional neural network, and then the features are fused with the target viewing angle information specified by the user, so that the model has the ability to infer the head appearance under the target view. After receiving the above fusion features, the appearance generation module converts them into a head image corresponding to the target view through a generative adversarial network based on the tri-planar representation, which can be used for subsequent training and completion of the side and back regions of the head model.

[0063] Step 4b) uses the same method as step 3) to train the 3D Gaussian spheres on the side and back of the frontal head model again through this multi-view appearance image. In this embodiment, the number of iterations F for training the 3D Gaussian spheres on the side and back of the frontal head model is F≥2000, and a complete head model is obtained.

[0064] Next, combined with the experimental results, the technical effects of the present invention will be further described:

[0065] 1. Experimental conditions and content:

[0066] The hardware platform for the experiment is: the processor is an Intel(R) Core i7-12700KF CPU with a main frequency of 3.6 GHz, the memory is 32 GB, and the graphics card is an NVIDIA GeForce RTX 3090. The experimental test platform is: Ubuntu 20.04 operating system, Python version is 3.8.20, Pytorch version is 2.1.0, and CUDA version is 11.8.

[0067] Simulation 1, simulating the effect of the head model reconstruction of the present invention on the INSTA public dataset and the self-built dataset, and the results are as Figure 2 (a) and Figure 2 (b) shown;

[0068] Simulation 2, comparing and simulating the peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), and learned perceptual image patch similarity (LPIPS) of the present invention and the existing 3D Gaussian splash-based head model reconstruction method through the INSTA public dataset and the self-built dataset, and the results are shown in Tables 1 and 2.

[0069] Table 1

[0070] PSNR↑ SSIM↑ LPIPS↓ Prior art 29.14 0.936 0.0672 The present invention 30.28 0.948 0.0594

[0071] Table 2

[0072] PSNR↑ SSIM↑ LPIPS↓ Prior art 27.96 0.915 0.0807 The present invention 29.12 0.927 0.0691

[0073] 2. Analysis of experimental results:

[0074] Referring to Figure 2 , in which Figure 2 (a) and Figure 2 (b) are respectively the simulation experiment results on the INSTA public dataset and the self-built dataset. The appearances of the reconstructed head models from the front, right, back, and left views are shown in sequence from left to right. It can be seen from the figure that the present invention can effectively solve the problem of incomplete modeling in the non-front area of the prior art, and the generated head models have significant improvements in both structural integrity and texture consistency.

[0075] Referring to Table 1 and Table 2, it can be seen from Table 1 that in terms of the PSNR index, the result of the present invention is 30.28, higher than 29.14 of the prior art, indicating a better performance in the reconstruction accuracy of the head model; in terms of the SSIM index, the result of the present invention is 0.948, showing stronger structural similarity compared with 0.936 of the prior art; in terms of the LPIPS index, the result of the present invention is 0.0594, better than 0.0672 of the prior art, indicating that it is closer to the real image in terms of perceptual quality.

[0076] It can be seen from Table 2 that the present invention also maintains a performance comparable to that of the public dataset on this dataset and is superior to the prior methods in all indexes. Among them, the PSNR index is 29.12, showing an obvious improvement compared with 27.96 of the prior art; the SSIM index is 0.927, higher than 0.915 of the prior art; the LPIPS index is 0.0691, better than 0.0807 of the prior art, further verifying the comprehensive advantages of the present invention in model reconstruction accuracy, structural similarity, and perceptual quality.

[0077] Based on the above analysis of the experimental results, the head model reconstruction method based on Gaussian splash and generative model proposed by the present invention leads the prior art solutions in the PSNR, SSIM, and LPIPS index tests, and this method is applicable to various datasets, which can improve the detail fidelity of the reconstructed model while ensuring the structural integrity of the head model.

Claims

1. A head model reconstruction method based on Gaussian splashing and generative model, characterized in that: The steps include: (1) Fitting the photographed face: Decode the facial expression video of the person shot by the monocular RGB camera, and fit the shape parameters of the FLAME head model to a neutral expression image in the decoded image sequence based on the face shape prediction model to obtain the FLAME head model with neutral expression; (2) Optimize the expression parameters and posture parameters of the FLAME head model: The facial key point position of each image in the image sequence is detected by the facial feature point detection model, and the FLAME head model expression parameter θ and posture parameter δ are optimized through the detected multiple key point positions to obtain the optimized FLAME head model expression parameter θ * and attitude parameter δ * ; (3) Training 3D Gaussian sphere based on Gaussian splash algorithm: According to the optimized FLAME head model expression parameters θ * and attitude parameter δ * The FLAME head model is linearly deformed, and the Gaussian splash algorithm is used to train the 3D Gaussian sphere on the surface of the initialized FLAME head model according to the image sequence and the deformed FLAME head model to obtain the frontal head model. (4) Obtain the complete head model reconstruction result based on the generative model: A generative model is used to generate a 360-degree multi-view appearance image around the head for the front head model, and the 3D Gaussian sphere of the side and back of the front head model is trained again using the multi-view appearance image to obtain a complete head model.

2. The method according to claim 1, characterized in that The step (1) is to fit the shape parameters of the FLAME head model to a neutral expression image in the decoded image sequence based on the face shape prediction model, wherein the face shape prediction model includes an identity encoding module and a geometry decoding module, and the fitting is implemented in the following steps: The identity encoding module encodes the facial information in the neutral expression image, and the geometry decoding module converts a set of shape parameters of implicit facial geometry and identity features obtained by encoding into shape parameters suitable for the FLAME head model, thereby obtaining the FLAME head model with neutral expression.

3. The method according to claim 1, characterized in that The facial feature point detection model described in step (2) detects the facial key point position of each image in the image sequence, wherein the facial feature point detection model includes a feature encoding module and a key point decoding module, and the detection is implemented in the following steps: The feature encoding module extracts features from each image; the key point decoding module decodes the extracted high-dimensional facial features to obtain the 2D coordinates of 68 facial key points, including 17 key points for identifying the jaw contour, 5 key points for identifying the left and right eyebrow areas and the nose tip area, 4 key points for identifying the nose bridge area, 6 key points for identifying the left and right eye contours, and 20 key points for identifying the lip contour.

4. The method according to claim 1, characterized in that The expression parameter θ and the posture parameter δ of the FLAME head model described in step (2) are optimized, and the optimization steps are: (2a) The number of initialization iterations is s, the maximum number of optimizations is S, S ≥ 10, and the expression parameters and posture parameters of the FLAME head model used for optimization in the sth iteration are denoted as θ s , δ s , let s = 1; (2b) According to the expression parameter θ of the sth iteration s and attitude parameter δ s Perform linear deformation on the FLAME head model with neutral expression to obtain the deformed FLAME head model; (2c) The loss value L is calculated by the deformed face key point ν of the FLAME head model and its corresponding detected face key point f s , and through L s For expression parameter θ s , attitude parameter δ s Optimize, where the loss value L s The calculation formula and optimization formula are: Among them, ∑· represents the sum operation, ||·||2 represents the L2 norm operation, and f i represents the 2D coordinates of the detected i-th facial key point, ν i represents the 3D coordinates of the i-th facial key point of the FLAME head model, Φ(·) represents the FLAME head model according to the expression parameter θ s and the attitude parameter δ s Perform deformation operation, Π(·) represents the projection operation from 3D coordinates to 2D coordinates, θ s ',δ s 'represents θ s , δ s The optimization result, η θ , η δ denote the learning rates of expression parameter θ and posture parameter δ respectively, Respectively represent L s For θ s , δ s The partial derivative operation of ; (2d) Determine whether s>S holds. If so, obtain the optimized FLAME head model expression parameter θ * and attitude parameter δ * , otherwise, let s=s+1 and execute step (2b).

5. The method according to claim 4, characterized in that The expression parameter θ according to the s-th iteration described in step (2b) s and attitude parameter δ s The FLAME head model with neutral expression is linearly deformed, where the FLAME head model includes a posture deformation base module and an expression deformation base module. The implementation steps of the linear deformation are as follows: The table situation rebase module according to the expression parameter θ s The facial area of ​​the FLAME head model is adjusted; the posture deformation base module is based on the posture parameter δ s The posture of the FLAME head model is adjusted to obtain the deformed FLAME head model.

6. The method according to claim 1, characterized in that The properties of the 3D Gaussian sphere of the surface of the FLAME head model initialized in step (3) are defined as: G={p,r,z,o,SH} Among them, p, r, z, o, and SH represent the position, rotation, scaling factor, opacity, and spherical harmonic coefficients of the 3D Gaussian sphere, respectively.

7. The method according to claim 1, characterized in that The Gaussian splash algorithm described in step (3) is used to train the 3D Gaussian sphere on the surface of the initialized FLAME head model according to the image sequence and the deformed FLAME head model. The implementation steps are: (3a) The number of initial training times is t, the maximum number of iterations is T, T ≥ 6000, and the attribute parameters of the 3D Gaussian ball used for optimization in the t-th training are represented by G t , let t = 1; (3b) Using the Gaussian splash algorithm, the 3D Gaussian sphere on the surface of the FLAME head model is used to render the image I pred ; (3c) By rendering image I pred The true value image I corresponding to the image sequence gt Calculate the loss value of the 3D Gaussian sphere on the surface of the FLAME head model And through the loss value For the 3D Gaussian sphere attribute parameter G t Optimize (3d) Determine whether t>T holds. If so, obtain the trained frontal head model. Otherwise, set t=t+1 and execute step (3b).

8. The method according to claim 7, characterized in that The loss value described in step (3c) The calculation formula for the 3D Gaussian sphere attribute parameter G t The optimization formulas are: in, is the mean absolute error, is the structural similarity error, λ represents the loss term weight, and U is I pred Total number of pixels, I pred (u), I gt (u) respectively represent I pred ,I gt The pixel value of the u-th pixel, Σ· represents the summation operation, |·| represents the absolute value operation, SSIM represents the calculation of the structural similarity between two images, G t ' indicates G t The optimization result, η G Represents the learning rate of each attribute of the 3D Gaussian ball, express For G t The partial derivative operation.

9. The method according to claim 1, characterized in that: The front head model described in step (4) generates a multi-view appearance image of 360 degrees around the head through a generative model, wherein the generative model includes a view encoding module and an appearance generation module, and the generation is implemented by the following steps: The view encoding module extracts features from the frontal head image and fuses the given observation angle information with the extracted features. The appearance generation module generates the head appearance at the corresponding angle using the encoded features to obtain a 360-degree multi-view appearance image surrounding the head.

Citation Information

Patent Citations

  • Face high-fidelity and drivable reconstruction method based on three-dimensional Gaussian splashing

    CN118736108A

Cited By

  • Head reconstruction method and system based on diffusion model and three-dimensional Gaussian sputtering

    CN121414991A