Three-dimensional human head reconstruction method based on geometric prior network and related device

By employing a 3D head reconstruction method based on geometric prior networks, combined with deformation field networks and 3D Gaussian sputtering technology, and binding 3D Gaussian vertices to FLAME meshes, the problem of geometric distortion in virtual digital head reconstruction is solved, achieving high-precision and detailed 3D head reconstruction.

CN121482228BActive Publication Date: 2026-03-24NANCHANG HANGKONG UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-09
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing virtual digital human head reconstruction technologies struggle to accurately fit the complex geometry of real human heads, particularly in areas with intricate structures such as hair and glasses. Furthermore, geometric distortions easily occur during expression and posture-driven processes, making it difficult to achieve high-precision reconstruction.

Method used

A 3D head reconstruction method based on geometric prior networks is adopted. By extracting the pose and expression parameters of monocular head videos, and combining the deformation field network model and 3D Gaussian sputtering technology, the vertices of the 3D Gaussian mesh are bound to the FLAME mesh, and the Gaussian initialization and offset prediction are optimized to achieve high-precision geometric structure reconstruction.

Benefits of technology

It improves the detail and accuracy of human head reconstruction, maintains the consistency of geometric structure when facial expressions and postures change, and achieves high-precision 3D human head reconstruction, which is suitable for rendering virtual digital heads and real-time interactive applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482228B_ABST
    Figure CN121482228B_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional human head reconstruction method based on a geometric prior network and related devices, and the method comprises the following steps: extracting a plurality of posture parameters and a plurality of expression parameters of each frame of picture in a preprocessed monocular human head video; performing linear mixed skinning changes on a plurality of initial mesh vertex positions of each frame of picture to obtain a plurality of first mesh vertex positions of each frame of picture, and assigning a reference three-dimensional Gaussian to each first mesh vertex position; initializing each reference three-dimensional Gaussian to obtain a first attribute set; determining a geometric offset of each reference three-dimensional Gaussian according to the plurality of posture parameters, the plurality of expression parameters, hidden coding and the plurality of initial mesh vertex positions; determining a second attribute set of each reference three-dimensional Gaussian according to the first attribute set and the geometric offset; and determining a human head reconstruction result according to the second attribute set of each reference three-dimensional Gaussian. The application realizes high-precision restoration of human head reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer vision and virtual reality technology, and in particular to a method and related apparatus for three-dimensional head reconstruction based on geometric prior networks. Background Technology

[0002] In recent years, virtual digital humans have shown great application potential in media production, virtual avatars, online education, and other fields. Users hope to reconstruct realistic digital heads that match the identity of the target person from monocular speaking head videos shot with simple devices such as mobile phones, and render virtual digital heads with controllable expressions from a new perspective.

[0003] However, current virtual digital human head reconstruction technology still faces multiple bottlenecks. For example, it cannot accurately fit the complex geometry of a real human head, such as performing poorly in detailed structural areas like hair and glasses; and in the process of expression and posture driving, geometric distortion often occurs, making it difficult to reproduce the deformation characteristics of a real head when switching expressions and adjusting postures, which is significantly different from users' high-precision requirements for head reconstruction. Summary of the Invention

[0004] This application provides a three-dimensional head reconstruction method and related apparatus based on geometric prior networks to improve the detail reproduction of head reconstruction and achieve high-precision head reconstruction.

[0005] In a first aspect, embodiments of this application provide a three-dimensional head reconstruction method based on a geometric prior network, including:

[0006] Extract multiple pose parameters and multiple facial expression parameters from each frame of the preprocessed monocular head video;

[0007] A linear blending skin transformation is performed on multiple initial mesh vertex positions of each frame to obtain multiple first mesh vertex positions of each frame, and a reference 3D Gaussian is assigned to each first mesh vertex position.

[0008] Each reference 3D Gaussian is initialized to obtain a first attribute set, which includes at least the first mesh vertex position, latent code, and vertex index. The latent code is used to characterize the multimodal feature fusion of a single first mesh vertex position.

[0009] The geometric offset of each reference 3D Gaussian is determined based on the plurality of pose parameters, the plurality of expression parameters, the hidden code, and the plurality of initial mesh vertex positions;

[0010] Based on the first attribute set and the geometric offset, a second attribute set is determined for each reference 3D Gaussian. The second attribute set includes at least the second mesh vertex position, the implicit code, and the vertex index. The second mesh vertex position is bound to the vertex index. The second mesh vertex position is determined based on the first mesh vertex position and the geometric offset.

[0011] The head reconstruction result is determined based on the second set of attributes of each reference 3D Gaussian.

[0012] The first attribute set further includes a first rotation direction, a first scaling factor, opacity, and color. The step of determining the geometric offset of each reference 3D Gaussian based on the plurality of pose parameters, the plurality of expression parameters, the implicit coding, and the plurality of initial mesh vertex positions includes:

[0013] The multiple posture parameters and the multiple facial expression parameters are fused to obtain fused parameters;

[0014] The position encoding module in the pre-trained deformation field network model is used to encode the position of each initial grid vertex to obtain the encoded features.

[0015] The offset prediction module in the deformation field network model predicts the offset of each reference 3D Gaussian based on the fusion parameters, the encoding features, and the hidden encoding, thereby obtaining the first offset of each initial mesh vertex position, the second offset of the first rotation direction, and the third offset of the first scaling scale.

[0016] The geometric offset of each reference 3D Gaussian is obtained based on the first offset, the second offset, and the third offset.

[0017] The step of determining the second attribute set of each reference 3D Gaussian based on the first attribute set and the geometric offset includes:

[0018] The position of the second grid vertex is determined based on the first offset and the position of each first grid vertex;

[0019] The second grid vertex position and the vertex index are bound together, and the second grid vertex position and the vertex index are fixed attributes;

[0020] The second rotation direction is determined based on the first rotation direction and the second offset.

[0021] The second scaling scale is determined based on the first scaling scale and the third offset.

[0022] The first attribute set is updated based on the second grid vertex position, the second rotation direction, and the second scaling scale to obtain the second attribute set.

[0023] The step of determining the head reconstruction result based on the second attribute set of each reference 3D Gaussian includes:

[0024] Determine the view space for each reference 3D Gaussian;

[0025] Determine the position gradient of the view space;

[0026] If the position gradient is greater than the first threshold, then the second attribute set of the newly added three-dimensional Gaussian is determined according to the second attribute set of the reference three-dimensional Gaussian corresponding to the position gradient.

[0027] Obtain the opacity of each reference 3D Gaussian and the newly added 3D Gaussian;

[0028] If the opacity is lower than the second threshold, then in each reference 3D Gaussian and the newly added 3D Gaussian, the 3D Gaussian corresponding to the opacity is deleted to obtain the remaining 3D Gaussian.

[0029] The head reconstruction result is determined based on the second set of attributes of the remaining three-dimensional Gaussian.

[0030] The step of determining the second attribute set of the newly added three-dimensional Gaussian based on the second attribute set of the reference three-dimensional Gaussian corresponding to the position gradient includes:

[0031] The second attribute set of the reference three-dimensional Gaussian corresponding to the position gradient is used as the reference attribute set of the newly added three-dimensional Gaussian.

[0032] Based on the dimensions of the reference 3D Gaussian corresponding to the position gradient, the dimensions of the newly added 3D Gaussian are adjusted to obtain the target dimension;

[0033] The hidden code of the reference 3D Gaussian corresponding to the position gradient is randomly perturbed to obtain the hidden code of the newly added 3D Gaussian.

[0034] Based on the target size and the implicit encoding of the newly added 3D Gaussian, the reference attribute set is updated to obtain the second attribute set of the newly added 3D Gaussian.

[0035] The step of determining the head reconstruction result based on the second attribute set of the remaining three-dimensional Gaussian includes:

[0036] Through projection operations, each remaining three-dimensional Gaussian is mapped to a local pixel contribution region in a two-dimensional image space according to the second attribute set of the remaining three-dimensional Gaussian, resulting in a set of pixel feature maps. Each pixel feature map in the set of pixel feature maps is a two-dimensional projection feature representation of a single remaining reference three-dimensional Gaussian.

[0037] Each pixel feature map is subjected to feature fusion processing to obtain a fused feature map;

[0038] The fused feature map is rendered to obtain the head reconstruction result.

[0039] Wherein, after determining the head reconstruction result based on the second attribute set of each reference 3D Gaussian, the method further includes:

[0040] The pixel-level difference between the head reconstruction result and the target image is determined to obtain the first loss, and the target image is used to characterize a single frame in the monocular head video corresponding to the head reconstruction result;

[0041] Extract the head reconstruction results and the semantic features associated with identity from the target image;

[0042] The similarity between the semantic features of the reconstructed head and the semantic features of the target image is determined to obtain the second loss.

[0043] Based on the first loss and the second loss, the learnable properties of the reference 3D Gaussian and the parameters of the deformation field network model are optimized to obtain the optimized learnable properties and the parameter configuration of the deformation field network model. The learnable properties include the first rotation direction, the first scaling scale, the opacity, the color, and the hidden encoding.

[0044] Based on the optimized learnable attributes and the parameter configuration of the deformation field network model, head reconstruction is performed on the next frame of the monocular head video.

[0045] Secondly, embodiments of this application provide a three-dimensional head reconstruction device based on a geometric prior network, comprising:

[0046] The extraction unit is used to extract multiple pose parameters and multiple facial expression parameters from each frame of the preprocessed monocular head video.

[0047] The deformation unit is used to perform linear blending skinning transformation on multiple initial mesh vertex positions of each frame to obtain multiple first mesh vertex positions of each frame, and assign a reference three-dimensional Gaussian to each first mesh vertex position.

[0048] An initialization unit is used to initialize each reference 3D Gaussian to obtain a first attribute set. The first attribute set includes at least the first mesh vertex position, latent code, and vertex index. The latent code is used to characterize the multimodal feature fusion of a single first mesh vertex position.

[0049] The first determining unit is used to determine the geometric offset of each reference 3D Gaussian based on the plurality of pose parameters, the plurality of expression parameters, the hidden code, and the plurality of initial mesh vertex positions;

[0050] The second determining unit is configured to determine a second attribute set for each reference 3D Gaussian based on the first attribute set and the geometric offset. The second attribute set includes at least the second mesh vertex position, the implicit code, and the vertex index. The second mesh vertex position is bound to the vertex index. The second mesh vertex position is determined based on the first mesh vertex position and the geometric offset.

[0051] The third determining unit is used to determine the head reconstruction result based on the second attribute set of each reference three-dimensional Gaussian.

[0052] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and executable program code stored in the memory and executable on the processor, wherein the processor executes the executable program code and performs the steps of the method described in the first aspect.

[0053] Fourthly, embodiments of this application provide a computer-readable storage medium storing executable program code, the executable program code including execution instructions for performing the steps of the method as described in the first aspect.

[0054] Fifthly, embodiments of this application provide a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps described in the first aspect of embodiments of this application. The computer program product may be a software installation package.

[0055] As can be seen, in this embodiment, firstly, multiple pose parameters and multiple expression parameters are extracted from each frame of the preprocessed monocular head video; then, linear blending skinning transformation is performed on multiple initial mesh vertex positions of each frame to obtain multiple first mesh vertex positions of each frame, and a reference 3D Gaussian is assigned to each first mesh vertex position; then, each reference 3D Gaussian is initialized to obtain a first attribute set, which includes at least the first mesh vertex position, the latent code, and the vertex index, wherein the latent code is used to characterize the multimodal feature fusion of a single first mesh vertex position; next, the geometric offset of each reference 3D Gaussian is determined based on the multiple pose parameters, the multiple expression parameters, the latent code, and the multiple initial mesh vertex positions; then, a second attribute set of each reference 3D Gaussian is determined based on the first attribute set and the geometric offset, wherein the second attribute set includes at least the second mesh vertex position, the latent code, and the vertex index, the second mesh vertex position is bound to the vertex index, and the second mesh vertex position is determined based on the first mesh vertex position and the geometric offset; finally, the head reconstruction result is determined based on the second attribute set of each reference 3D Gaussian.

[0056] This application optimizes the Gaussian initialization accuracy and constrains the Gaussian spatial distribution by binding each 3D Gaussian vertices to a mesh vertex, which helps ensure the consistency of the overall geometric structure during head reconstruction. Simultaneously, latent coding attributes are introduced during the 3D Gaussian initialization stage, providing a unique identifier for each 3D Gaussian vertices. This latent coding is applied to the prediction of geometric offsets, filling the gaps in detailed structures such as hair and glasses areas within the mesh, thus improving the detail reproduction of the head reconstruction. By using pose parameters, expression parameters, initial mesh vertex positions, and latent coding, the geometric offsets of each 3D Gaussian vertices are predicted, fully considering the dynamic changes in Gaussian attributes such as vertex positions under different facial expressions and poses, achieving accurate capture of non-rigid deformation features. Finally, head reconstruction is completed based on the second attribute set of the 3D Gaussians, ensuring that the output head reconstruction result accurately matches facial expressions and poses, achieving high-precision head reconstruction. Attached Figure Description

[0057] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0058] Figure 1 This application provides a system architecture diagram of a three-dimensional head reconstruction system.

[0059] Figure 2 This is a flowchart illustrating a three-dimensional head reconstruction method based on a geometric prior network provided in an embodiment of this application.

[0060] Figure 3 This is a schematic diagram of a three-dimensional Gaussian sputtering head reconstruction model based on geometric mesh binding provided in an embodiment of this application;

[0061] Figure 4 This is a schematic diagram illustrating a qualitative comparison of head reconstruction effects provided in an embodiment of this application;

[0062] Figure 5 This is a schematic diagram illustrating a quantitative comparison of head reconstruction effects provided in an embodiment of this application;

[0063] Figure 6 This is a schematic diagram illustrating a multi-view head reconstruction result provided in an embodiment of this application;

[0064] Figure 7 This is a schematic diagram illustrating a cross-identity transfer facial expression and pose generation head reconstruction result provided in an embodiment of this application;

[0065] Figure 8 This is a schematic diagram of an ablation experiment for verifying a binding strategy provided in an embodiment of this application;

[0066] Figure 9 This is a functional unit block diagram of a three-dimensional head reconstruction device based on a geometric prior network provided in an embodiment of this application;

[0067] Figure 10 This is a functional unit block diagram of another three-dimensional head reconstruction device based on a geometric prior network provided in this application embodiment;

[0068] Figure 11 This is a schematic diagram of the structure of an electronic device proposed in an embodiment of this application. Detailed Implementation

[0069] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.

[0070] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0071] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0072] In recent years, virtual digital humans have shown great application potential in media production, virtual image creation, online education and other fields. Users expect to be able to quickly reconstruct a three-dimensional virtual digital head that is consistent with the identity of the target person, with controllable expressions and freely switchable perspectives from a monocular head video shot with simple devices such as mobile phones.

[0073] Among them, 3D Gaussian sputtering is an efficient and real-time 3D scene representation and rendering technology. It represents the scene through an explicit 3D Gaussian distribution and combines it with differentiable rasterization to achieve high-quality, real-time rendering. Compared with the traditional neural radiation field method, it has significant improvements in both speed and rendering quality. Currently, the head virtual image modeling method based on 3D Gaussian sputtering has become the main technical route for achieving multi-view consistency and fast rendering.

[0074] However, 3D Gaussian sputtering uses a randomly initialized 3D Gaussian distribution and learns the offset through a deformation field network built using a multilayer perceptron. Specifically, in the initialization stage of Gaussian sputtering, a sparse set of 3D Gaussian point clouds is usually randomly generated as the initial model, which is difficult to accurately cover the features of detailed areas such as hair and glasses, resulting in a large deviation between the reconstructed head and the geometric structure of the real head. Furthermore, the dynamic adjustment of the number of 3D Gaussian points fails to effectively integrate the geometric prior information of the Faces Learned with an Articulated Model and Expressions (FLAME) mesh, leading to an excessive increase in the number of 3D Gaussian points to achieve satisfactory rendering quality. This results in hundreds of thousands of 3D Gaussian points, and the large number of redundant Gaussian points significantly reduces rendering efficiency, making it difficult to meet the needs of real-time interactive applications.

[0075] In the deformation field learning offset stage, when dealing with non-rigid deformation caused by changes in expression and posture, the dynamic changes of the three-dimensional Gaussian properties under different expressions and postures are not fully captured. As a result, when switching expressions or adjusting postures, it is difficult to reproduce the deformation characteristics of the real head when switching expressions and adjusting postures. This easily leads to geometric distortion problems such as facial contour distortion and blurred expression details, which seriously affects the authenticity and accuracy of the virtual digital head reconstruction results.

[0076] To address the aforementioned problems, this application provides a method and related apparatus for three-dimensional head reconstruction based on geometric prior networks. The embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0077] Please see Figure 1 , Figure 1 This is a system architecture diagram of a three-dimensional head reconstruction system provided in an embodiment of this application. Figure 1 As shown, the 3D head reconstruction system 100 includes a data acquisition module 101, a model building module 102, a model training module 103, and a result generation module 104. The data acquisition module 101, model building module 102, model training module 103, and result generation module 104 are interconnected.

[0078] The data acquisition module 101 is used to acquire and preprocess the monocular head video dataset, and to extract FLAME grid parameters from the monocular head video dataset.

[0079] The model building module 102 is used to construct a geometrically mesh-bound three-dimensional Gaussian sputtering head reconstruction model based on the three-dimensional Gaussian sputtering technology. The three-dimensional Gaussian sputtering head reconstruction model includes a Gaussian initialization model and a deformation field network model.

[0080] The process involves using a Gaussian initialization model to bind the center position of a 3D Gaussian distribution to each FLAME vertex, generating an initial Gaussian distribution set. The deformation field network model, using facial and pose parameters as conditions, predicts higher-order geometric offsets, thereby precisely correcting and dynamically adjusting various attributes of the initial 3D Gaussian distribution. Finally, the sputtering module of the Gaussian initialization model projects the corrected 3D Gaussian distribution onto a 2D image space, renders it, and outputs the reconstructed head image.

[0081] The model training module 103 is used to iteratively train and optimize the model built by the model building module 102 based on the training dataset, and to complete the model performance verification and generalization ability evaluation based on the test dataset, and finally output a model that meets the accuracy requirements of human head 3D reconstruction.

[0082] The result generation module 104 is used to input the preprocessed monocular head video into the trained three-dimensional Gaussian sputtering head reconstruction model and output the head reconstruction image of each frame in the video.

[0083] Based on this, this application provides a three-dimensional head reconstruction method and related apparatus based on geometric prior networks. The following is a detailed description of this application with reference to the accompanying drawings.

[0084] Please see Figure 2 , Figure 2 This is a flowchart illustrating a three-dimensional head reconstruction method based on a geometric prior network provided in an embodiment of this application, as shown below. Figure 2 As shown, the method includes the following steps:

[0085] S210: Extract multiple pose parameters and multiple facial expression parameters from each frame of the preprocessed monocular head video.

[0086] This involves using simple shooting devices such as mobile phones to capture monocular head videos in different scenes and with different expressions and postures. The captured monocular head video data is then preprocessed to obtain preprocessed monocular head videos. Specifically, this may include head region cropping, frame rate adjustment, pixel adjustment, foreground segmentation, and mouth region recognition.

[0087] Among them, the head bounding box, including the forehead, chin, ears, etc., can be located by face detection algorithm. Then, the bounding box is expanded outward by 5% to 10% of pixels to avoid cropping the head edge, and finally the area containing only the head is cropped out.

[0088] Next, the frame rate and pixel values ​​of the cropped area containing only the head are adjusted to unify the temporal sampling rate and spatial resolution. For example, the frame rate can be unified to 25 FPS and the pixel resolution to 512×512.

[0089] Furthermore, foreground segmentation is performed on the preprocessed monocular head video to obtain a head foreground mask and a pure head region image. For example, this can be achieved using Robust Video Matting (RVM). Then, mouth analysis is performed within the segmented pure head region image to obtain a pixel-level mouth segmentation mask, preventing similar textures in the background from being misidentified as mouths. For example, the existing face analysis framework Bilateral SeNet can be used to identify the mouth region, or other algorithms with high-precision face region segmentation capabilities can be selected.

[0090] Subsequently, the MICA face tracker is used to compare the head image in the input video with the rendering result of the FLAME model. The FLAME parameters are continuously adjusted through optimization algorithms to make the model output match the input image as closely as possible, thereby obtaining the FLAME mesh parameters with temporal continuity. The FLAME mesh parameters include shape parameters, pose parameters, and expression parameters.

[0091] Among them, multiple posture parameters include eye posture and chin posture, and multiple expression parameters include expression coefficient and eyelid coefficient.

[0092] S220, perform linear blending skinning transformation on multiple initial mesh vertex positions of each frame to obtain multiple first mesh vertex positions of each frame, and assign a reference 3D Gaussian to each first mesh vertex position.

[0093] The initial mesh vertex position is the reference mesh vertex under neutral expression and pose. The first mesh vertex position is the frame-specific, real-time deformed FLAME mesh vertex coordinates obtained by inputting multiple pose parameters and multiple expression parameters into the FLAME model and performing linear blending skin deformation for the current frame.

[0094] For each frame of the video, the 3D coordinates of all FLAME mesh vertices after deformation based on pose and expression parameters are first calculated, i.e., the position of the first mesh vertex. Then, a 3D Gaussian sphere is assigned to each deformed vertex, establishing a binding relationship between the FLAME mesh vertex and the 3D Gaussian sphere. When the FLAME mesh changes, the changed coordinates of the corresponding mesh vertex can be obtained as the reference position of the 3D Gaussian sphere based on the binding relationship between the 3D Gaussian sphere and the corresponding mesh vertex.

[0095] S230 initializes each reference 3D Gaussian to obtain the first attribute set.

[0096] The first attribute set includes at least the first grid vertex position, latent coding, and vertex index, wherein the latent coding is used to characterize the multimodal feature fusion of a single first grid vertex position.

[0097] Specifically, the set of properties of a 3D Gaussian includes center position, implicit encoding, vertex index, rotation direction, scaling scale, opacity, and color.

[0098] Among them, the center position is used to locate the specific coordinates of the 3D Gaussian in 3D space; the implicit code is a unique identifier for each 3D Gaussian, which helps the deformation field network to accurately learn and predict its geometric offset, and fill in detailed structures such as hair and glasses; the vertex index is used to establish the binding relationship between the 3D Gaussian and the mesh vertices; the rotation direction is used to control the spatial orientation of the 3D Gaussian; the scaling scale is used to adjust the spatial volume of the 3D Gaussian; the opacity is used to control the transparency of the 3D Gaussian during rendering; and the color is used to define the color information of the 3D Gaussian.

[0099] In this embodiment, an implicit coding attribute is extended on top of the original 3D Gaussian to characterize multimodal features such as 2D visual features and 3D geometric features of the bound vertices. The implicit coding attribute is a learnable low-dimensional vector, with one implicit code corresponding to each 3D Gaussian, used to identify the uniqueness of the Gaussian. This implicit code is input into subsequent offset predictions to enhance the detail representation of non-mesh regions.

[0100] The initialization process relies on the FLAME human head prior geometric model to determine the fixed properties of each reference 3D Gaussian and set reasonable initial values ​​for the learnable properties. Finally, all properties are integrated to obtain the first property set.

[0101] Specifically, the native indices of the FLAME mesh vertices corresponding to the reference 3D Gaussian are used as the initial values ​​for the vertex indices of the reference 3D Gaussian, where the vertex indices are fixed properties. The position of the first mesh vertex is used as the initial value for the center position of the reference 3D Gaussian, and after offset correction by the deformation field network model output, it remains fixed and is also a fixed property.

[0102] Among them, the implicit encoding, rotation direction, scaling scale, opacity, and color are learnable attributes. Only initial values ​​are set, and subsequent updates can be made through training or optimization iterations to adapt to the detailed features of the human head.

[0103] Specifically, each reference 3D Gaussian is initialized using the initialization module in the trained Gaussian initialization model, resulting in the first grid vertex position, implicit code, vertex index, first rotation direction, first scaling scale, opacity, and color of each reference 3D Gaussian. Note that the first grid vertex position is not necessarily the center position of the reference 3D Gaussian; this center position needs to be determined in conjunction with the deformation field network model.

[0104] S240, determine the geometric offset of each reference 3D Gaussian based on the plurality of pose parameters, the plurality of expression parameters, the hidden code, and the plurality of initial mesh vertex positions.

[0105] In one possible embodiment, the first attribute set further includes a first rotation direction, a first scaling factor, opacity, and color. Determining the geometric offset of each reference 3D Gaussian based on the plurality of pose parameters, the plurality of expression parameters, the implicit encoding, and the plurality of initial mesh vertex positions includes: fusing the plurality of pose parameters and the plurality of expression parameters to obtain fused parameters; encoding the position of each initial mesh vertex position using a position encoding module in a pre-trained deformation field network model to obtain encoded features; predicting the offset of each reference 3D Gaussian attribute based on the fused parameters, the encoded features, and the implicit encoding using an offset prediction module in the deformation field network model to obtain a first offset of each initial mesh vertex position, a second offset of the first rotation direction, and a third offset of the first scaling factor; and obtaining the geometric offset of each reference 3D Gaussian based on the first offset, the second offset, and the third offset.

[0106] Specifically, the facial expression coefficients, eye poses, chin poses, and eyelid coefficients in the FLAME mesh parameters are concatenated as the fusion parameters of the deformation field network model.

[0107] In this model, the initial mesh vertex positions are in three-dimensional spatial coordinates. Low-dimensional features struggle to capture subtle geometric changes in the local area of ​​a face. To improve the modeling capability for high-frequency details, multiple sets of frequency components can be set in the position encoding module to map the initial mesh vertex positions to a high-dimensional space, generating encoded features with high-frequency information. Preferably, the position encoding module includes six sets of frequency components.

[0108] The migration prediction module can employ a 6-layer, 256-dimensional multilayer perceptron model. The input is the fusion parameters. Implicit coding of each reference 3D Gaussian and the encoded features obtained through positional encoding The output is the offset of the 3D Gaussian geometric properties, including the offset of the center position, rotation direction, and size scaling. The specific formula is as follows:

[0109] ,in, These are the first offset, the second offset, and the third offset. It is the set of learnable parameters for a multilayer perceptron model, including parameters such as weights and biases for each layer of the model.

[0110] Introducing skip connections and combining them with positional encoding into the multilayer perceptron model can avoid activation saturation and gradient vanishing problems, while strengthening the ability to model high-frequency details. This allows the deformation field network to capture the fine geometric features of the human head more accurately and stably, while ensuring the training efficiency and stability of the deep network. Ultimately, this improves the detail reproduction and the naturalness of pose and expression transfer in the 3D Gaussian sputtering head reconstruction.

[0111] S250, determine the second set of attributes for each reference 3D Gaussian based on the first set of attributes and the geometric offset.

[0112] The second attribute set includes at least the second mesh vertex position, the implicit code, and the vertex index. The second mesh vertex position is bound to the vertex index and is determined based on the first mesh vertex position and the geometric offset.

[0113] In one possible embodiment, determining the second attribute set of each reference 3D Gaussian based on the first attribute set and the geometric offset includes: determining the second mesh vertex position based on the first offset and the position of each first mesh vertex; binding the second mesh vertex position and the vertex index, wherein the second mesh vertex position and the vertex index are fixed attributes; determining a second rotation direction based on the first rotation direction and the second offset; determining a second scaling scale based on the first scaling scale and the third offset; and updating the first attribute set based on the second mesh vertex position, the second rotation direction, and the second scaling scale to obtain the second attribute set.

[0114] Specifically, the initial Gaussian set output by the Gaussian initialization model and the geometric offset output by the deformation field network model are concatenated to generate the final Gaussian distribution set.

[0115] Specifically, the first offset and the first mesh vertex position are concatenated to obtain the second mesh vertex position, which is the center position of the final reference 3D Gaussian. The second offset and the first rotation direction are concatenated to obtain the second rotation direction; and the third offset and the first scaling factor are concatenated to obtain the second scaling factor. The specific formulas are shown below:

[0116] ,in, They represent the fusion parameters respectively. The reference 3D Gaussian's center position, rotation direction, and scaling scale. The position of the first grid vertex. The first direction of rotation, This is the first scaling scale.

[0117] After obtaining the center position of the reference 3D Gaussian, it is bound to the vertex indices of the reference 3D Gaussian. The center position of the 3D Gaussian and the corresponding mesh vertex indices are fixed values ​​that remain unchanged after binding; the remaining attributes are learnable.

[0118] Specifically, the second grid vertex position, the second rotation direction, and the second scaling scale are replaced with the first grid vertex position, the first rotation direction, and the first scaling scale in the first attribute set, respectively, to obtain the second attribute set of each reference 3D Gaussian.

[0119] As can be seen, in this embodiment, the high-order geometric offset is predicted by using the deformation field network model with expression parameters and pose parameters as conditions, thereby effectively decoupling expression-related features from identity-independent features, fully considering the dynamic changes of vertex positions under different expression poses, and enhancing the modeling ability for non-rigid deformations; by using implicitly encoded attributes to compensate for the structural deficiencies of the mesh in areas such as hair and glasses, the detail restoration and viewpoint consistency of the head reconstruction are significantly improved.

[0120] Furthermore, the dynamic changes of the human head are handled jointly by the deformation of the FLAME mesh and the learning offset of the multilayer perceptron, avoiding the need for a deeper network layer that would otherwise require the multilayer perceptron to learn dynamic scene changes alone, thus achieving a rendering speed of up to 300 FPS.

[0121] S260, determine the head reconstruction result based on the second attribute set of each reference three-dimensional Gaussian.

[0122] Among them, the head reconstruction result refers to the three-dimensional head reconstruction image corresponding to the input video or image frame, which has a real physical space geometric structure and visual appearance details. It not only restores the overall outline of the head, the precise shape of the facial features and the geometric features such as hair growth distribution, but also restores the appearance details such as skin texture, hair texture, and facial light and shadow changes. It can be applied to application scenarios such as virtual digital humans, real-time interaction, AR / VR.

[0123] The head reconstruction results include head reconstruction images from each frame of the monocular head video.

[0124] In this model, each reference 3D Gaussian collectively constitutes the 3D spatial distribution of the human head. The sputtering module in the trained Gaussian initialization model performs 3D spatial distribution optimization and feature enhancement processing. Then, through differentiable rasterization, the Gaussian set in the 3D space is accurately projected onto the 2D image space to construct a mapping relationship corresponding to the input monocular video frame. Finally, the renderer outputs the 3D reconstructed human head image based on this mapping relationship.

[0125] As can be seen, in this embodiment, the rendering stage relies on only one differentiable rasterization operation to generate a high-quality image, significantly reducing computational complexity and increasing the rendering frame rate, thereby achieving millisecond-level response while ensuring image realism.

[0126] In one possible embodiment, determining the head reconstruction result based on the second attribute set of each reference 3D Gaussian includes: determining the view space of each reference 3D Gaussian; determining the position gradient of the view space; if the position gradient is greater than a first threshold, determining the second attribute set of a new 3D Gaussian based on the second attribute set of the reference 3D Gaussian corresponding to the position gradient; obtaining the opacity of each reference 3D Gaussian and the new 3D Gaussian; if the opacity is lower than a second threshold, deleting the 3D Gaussian corresponding to the opacity from each reference 3D Gaussian and the new 3D Gaussian to obtain the remaining 3D Gaussian; and determining the head reconstruction result based on the second attribute set of the remaining 3D Gaussian.

[0127] A three-dimensional Gaussian array with only the same number of vertices as the FLAME mesh struggles to achieve high-quality rendering, particularly in capturing detailed skin textures and expressive movements. Using a dynamic adjustment strategy for the original three-dimensional Gaussian array fails to leverage the geometric priors of the FLAME mesh, and frequent increases and decreases in the number of Gaussian elements due to drastic gradient changes result in numerous floating artifacts from new perspectives. Therefore, this application proposes an inherited, binding-based adaptive Gaussian element addition and deletion strategy that minimizes the number of Gaussian elements while maintaining rendering quality and supports cross-identity pose and expression transfer.

[0128] Specifically, by using view-dependent gradient feedback and opacity filtering, the number of 3D Gaussians is adaptively and dynamically adjusted, which ensures the rendering accuracy of the geometric details and textures of the human head, while relying on the geometric prior of the FLAME mesh to avoid redundant Gaussians and artifacts.

[0129] The view space refers to a three-dimensional coordinate system constructed based on the observation perspective of the current monocular video frame. Its origin coincides with the optical center of the camera, and the coordinate axes are aligned with the camera's imaging plane.

[0130] This process involves transforming a reference 3D Gaussian from the FLAME model space to the view space of the current observation perspective. During this transformation, a projection matrix is ​​calculated using camera intrinsic and extrinsic parameters to map the 3D coordinates of each reference 3D Gaussian to its corresponding position in the view space, while preserving its binding relationship with the FLAME mesh vertices.

[0131] The position gradient refers to the partial derivative of the three-dimensional position parameters of each reference 3D Gaussian with respect to the rendering loss in the current view space. The larger the gradient value, the greater the imaging error of the grid where the Gaussian is located, and the Gaussian at the current position cannot fully capture the geometric details of the corresponding grid.

[0132] The first threshold is a gradient determination threshold obtained through statistical optimization of training data. When the gradient of a reference 3D Gaussian exceeds the first threshold, it indicates that a higher density Gaussian distribution is needed in the region where the Gaussian is located to fit the details. In this case, a new 3D Gaussian is added in the region where the Gaussian is located. Based on the reference 3D Gaussian that triggered the gradient threshold, attribute inheritance and local adjustment are performed to obtain the second attribute set of the new 3D Gaussian.

[0133] The second threshold is a quantification standard used to distinguish between effective and redundant Gaussians. It can also be determined through training and optimization to ensure that only 3D Gaussians that do not substantially contribute to the rendering result are removed. When the opacity of a Gaussian is lower than the second threshold, it indicates that the 3D Gaussian contributes very little to the rendering result. Moreover, such 3D Gaussians usually appear as artifacts, affecting the rendering quality from new perspectives and increasing computational overhead, thus leading to their removal.

[0134] Since both the reference Gaussian and the newly added Gaussian are bound to the topology of the FLAME mesh, the deletion operation only targets redundant Gaussians and will not destroy the core geometric region of the human head defined in advance by FLAME. This avoids the problem of geometric structure disorder caused by frequent addition and subtraction of Gaussians, so that the remaining three-dimensional Gaussians obtained in the end not only cover all effective geometric and texture regions of the human head, but also minimize the number of Gaussians.

[0135] In one possible embodiment, determining the second attribute set of the newly added 3D Gaussian based on the second attribute set of the reference 3D Gaussian corresponding to the position gradient includes: using the second attribute set of the reference 3D Gaussian corresponding to the position gradient as the reference attribute set of the newly added 3D Gaussian; adjusting the size of the newly added 3D Gaussian according to the size of the reference 3D Gaussian corresponding to the position gradient to obtain a target size; randomly perturbing the implicit code of the reference 3D Gaussian corresponding to the position gradient to obtain the implicit code of the newly added 3D Gaussian; and updating the reference attribute set according to the target size and the implicit code of the newly added 3D Gaussian to obtain the second attribute set of the newly added 3D Gaussian.

[0136] Among them, the reference three-dimensional Gaussian with a position gradient exceeding the second threshold is used as the parent Gaussian, so that the newly added three-dimensional Gaussian inherits its second attribute set, ensuring that the newly added three-dimensional Gaussian fully utilizes the geometric prior of FLAME from the initial stage.

[0137] For a reference 3D Gaussian whose position gradient is greater than the first threshold, different processing methods are used according to its size. If its size is large, such as greater than the third threshold, it is split into two smaller 3D Gaussians. The sizes of these two 3D Gaussians can be the same or different. If its size is small, such as less than the fourth threshold, a 3D Gaussian of the same size is cloned.

[0138] Meanwhile, the newly added 3D Gaussian is bound to the same FLAME mesh vertex as the original 3D Gaussian. In order to maintain the uniqueness of each 3D Gaussian and make up for the lack of expression parameters in the FLAME mesh, the latent code is randomly perturbed after copying. On the one hand, this can avoid complete redundancy between the features of the newly added Gaussian and the original Gaussian, ensuring the uniqueness of each Gaussian. On the other hand, it can make up for the limitations of the expression parameterization of the FLAME mesh, allowing the newly added Gaussian to capture more subtle expression or texture details beyond the preset range of FLAME, while not deviating from the semantic features of the original Gaussian.

[0139] In this process, after obtaining the target size and the perturbation-encoded implicit code, the original size and implicit code in the reference attribute set are replaced with them, while the remaining inherited attributes remain unchanged, ultimately forming the second attribute set of the newly added three-dimensional Gaussian.

[0140] As can be seen, in this embodiment, a dynamic addition and subtraction strategy is introduced, using the view space position gradient and opacity as the basis for judgment, to split, clone or delete Gaussians, add Gaussians that inherit the vertex indices of the original Gaussians, and randomly perturb the implicit encoding. This provides the expressive power of Gaussians while maintaining the mesh binding relationship, effectively utilizing the geometric prior of explicit modeling, thereby avoiding the problems of low rendering quality due to too few 3D Gaussians or rendering speed affected by too many.

[0141] In one possible embodiment, determining the head reconstruction result based on the second attribute set of the remaining three-dimensional Gaussian includes: mapping each remaining three-dimensional Gaussian to a local pixel contribution region in a two-dimensional image space through a projection operation, based on the second attribute set of the remaining three-dimensional Gaussian, to obtain a set of pixel feature maps, wherein each pixel feature map in the set is a two-dimensional projection feature representation of a single remaining reference three-dimensional Gaussian; performing feature fusion processing on each pixel feature map to obtain a fused feature map; and rendering the fused feature map to obtain the head reconstruction result.

[0142] The second attribute set of the remaining 3D Gaussians has been optimized through gradient-driven addition and opacity-driven deletion, and its distribution density matches the head detail requirements of the current viewpoint. The second attribute set of the remaining 3D Gaussians can be input into the sputtering module of the Gaussian initialization model for fine-tuning of spatial distribution. Then, through projection operations, each remaining 3D Gaussian is precisely mapped to a local pixel contribution region in the 2D image space based on its second attribute set. Based on the spatial distribution characteristics of the Gaussian kernel, the Gaussian diffuses outwards from its projection point in the image space to surrounding pixels. The contribution weight to the pixel is determined by combining the Gaussian's color, opacity, and other attributes, ultimately yielding a pixel feature map for each remaining 3D Gaussian. Each pixel feature map represents the geometric and appearance information corresponding to that Gaussian.

[0143] In this process, each pixel feature map is fused, and the contribution weights of all remaining three-dimensional Gaussian features to each pixel and color information are superimposed to construct the mapping relationship between the remaining three-dimensional Gaussian attributes and image pixels, resulting in a fused feature map. The fused feature map is used to characterize the detailed features and structural information of the human head.

[0144] The process involves rendering the fused feature map. The renderer, based on the aforementioned mapping relationship and combined with lighting models such as ambient light and diffuse reflection, simulates realistic lighting effects to generate a reconstructed head image with three-dimensional geometric structure and realistic texture details.

[0145] As can be seen, in this embodiment, the reconstructed head image captures details such as skin lines and facial wrinkles due to the high-density distribution of Gaussians, while the removal of redundant Gaussians ensures rendering efficiency. Furthermore, it leverages FLAME priors to achieve cross-identity pose and expression transfer capabilities. This solves the problems of fixed Gaussian quantity, insufficient detail modeling, and low rendering efficiency in existing methods, achieving fast rendering speed, high image quality, and strong identity consistency.

[0146] In one possible embodiment, after determining the head reconstruction result based on the second attribute set of each reference 3D Gaussian, the method further includes: determining the pixel-level difference between the head reconstruction result and the target image to obtain a first loss, wherein the target image is used to characterize a single frame in the monocular head video corresponding to the head reconstruction result; extracting identity-related semantic features from the head reconstruction result and the target image; determining the similarity between the semantic features of the head reconstruction result and the semantic features of the target image to obtain a second loss; optimizing the learnable attributes of the reference 3D Gaussian and the parameters of the deformation field network model based on the first loss and the second loss to obtain optimized learnable attributes and the parameter configuration of the deformation field network model, wherein the learnable attributes include the first rotation direction, the first scaling scale, the opacity, the color, and the implicit coding; and reconstructing the head in the next frame of the monocular head video based on the optimized learnable attributes and the parameter configuration of the deformation field network model.

[0147] In this embodiment, multiple loss functions can be set to ensure that the generated image is as close as possible to the real image.

[0148] Specifically, to achieve dual optimization of pixel-level accuracy and identity semantic consistency in head reconstruction results, this application designs a multi-loss fusion model optimization framework. First, the pixel-level difference between the reconstructed head image and the target image is calculated and defined as the first loss, which is the mean squared error loss. By calculating the squared difference between the reconstructed head image and the target image at each pixel position and taking the global average, a global image quality evaluation index is constructed.

[0149] The formula for calculating the first loss is as follows: ,in, Reconstructing images of human heads, The target image.

[0150] Secondly, to ensure a high degree of consistency between the reconstruction result and the target image in terms of identity features, high-level semantic features associated with identity are extracted from both. The second loss is obtained by calculating the similarity of these semantic features, which is a perceptual loss based on identity features.

[0151] The second loss can utilize a pre-trained deep convolutional neural network VGG to extract high-level semantic features of the image, perform similarity measurement in the feature space, and then accurately capture core identity features such as facial structure, skin color, and hairstyle, thereby effectively reducing the blur and distortion of the reconstructed image, strengthening the ability to preserve identity features, and enhancing the ability to preserve identity features.

[0152] The formula for calculating the second loss is as follows: .

[0153] The fusion loss is obtained by weighting and summing the first loss and the second loss based on the first weight of the first loss and the second weight of the second loss output during the training process.

[0154] The formula for calculating the fusion loss is as follows:

[0155] ,in, As the first weight, It is the second weight.

[0156] Based on this fusion loss, the learnable attributes of the reference 3D Gaussian and the parameters of the deformation field network model are jointly optimized to obtain the optimal parameter configuration that balances pixel accuracy and identity consistency. Finally, based on the optimized attributes and parameters, efficient and high-precision head reconstruction can be performed on the next frame of a monocular head video, ensuring the continuity and consistency of the reconstructed head images in the video sequence.

[0157] In one possible embodiment, a Gaussian initialization model and a deformation field network model are constructed. The Gaussian initialization model includes an initialization module, an attribute binding module, and a sputtering module. The deformation field network model includes a position encoding module and an offset prediction module.

[0158] The process begins by acquiring a monocular head video dataset, followed by preprocessing. Each preprocessed monocular head video can be between 1 and 3 minutes long, with the last 400-600 frames of each video used as the test set and the remainder as the training set. Preferably, the last 500 frames are used as the test set, and the remainder as the training set.

[0159] The Gaussian initialization model and the deformation field network model are iteratively trained and their parameters are optimized using the training set. The model performance is verified and its generalization ability is evaluated based on the test set. Finally, a model that meets the accuracy requirements for 3D head reconstruction is output.

[0160] Specifically, the initialization module takes the three-dimensional coordinates of all vertices of the FLAME mesh and their corresponding vertex index information, the preprocessed frame-level feature data of the monocular head video, and the initialization prior parameters of the three-dimensional Gaussian attributes as inputs, and outputs the preliminary assignments of the three-dimensional Gaussian attributes to obtain the initial Gaussian distribution set.

[0161] The input data of the deformation field network model includes the stitched facial expression and pose parameters, the latent coding attributes in the initial Gaussian distribution set, and the reference mesh vertices under neutral facial expression and pose. First, the reference mesh vertices are position-encoded by the position encoding module to obtain the high-dimensional vertex coordinates. Then, the high-dimensional vertex coordinates, the stitched facial expression and pose parameters, and the latent coding attributes in the initial Gaussian distribution set are input into the offset prediction module, which outputs the offset of the three-dimensional Gaussian geometric attributes.

[0162] Specifically, the Gaussian properties generated during initialization are corrected and dynamically adjusted by offset to obtain the corrected three-dimensional Gaussian properties, and thus the final Gaussian distribution set is obtained.

[0163] Specifically, the attribute binding module in the Gaussian initialization model binds the center position attribute and vertex index attribute of each 3D Gaussian.

[0164] Finally, the final Gaussian distribution set is input into the sputtering module of the Gaussian initialization model to output the reconstructed head image.

[0165] During model training, the number of 3D Gaussians is dynamically adjusted based on the position gradient in the view space and the opacity of each 3D Gaussian.

[0166] During training, mean squared error loss and perceptual loss are combined, and pixel-level error and high-level semantic feature error are used as optimization targets. In the early stage of training, the weight of mean squared error loss is increased to prioritize the alignment of the reconstructed image with the real image in terms of overall structure, contour and color. In the later stage of training, the weight of perceptual loss is increased to ensure the consistency of identity features, facial details and pose restoration of the reconstruction results.

[0167] Specifically, in facial expression generation tasks, the mean squared error loss method, through pixel-by-pixel comparison, pays particular attention to preserving the detailed features of key facial regions, such as the eyes and mouth, which are sensitive to facial expressions. This pixel-level alignment-based optimization mechanism allows the neural network to gradually correct local distortions and global biases in the output image during training.

[0168] During training, mean squared error loss can not only effectively improve the structural integrity of the generated image, but also significantly improve the realism of texture details. Ultimately, it enables the synthesized facial expressions to maintain the consistency of identity features while achieving lighting effects and muscle movement patterns that are difficult to distinguish from real images.

[0169] Furthermore, by dynamically adjusting the weight coefficients of the mean squared error loss, the model can focus more on optimizing subtle facial features in the later stages of training, thereby surpassing the performance of traditional image generation methods in both the dimensions of facial expression intensity and naturalness.

[0170] Specifically, increasing the weight of the mean squared error loss can enhance the pixel-level accuracy of the generated image; increasing the weight of the perceptual loss can better maintain the semantic consistency of identity features. Furthermore, the optimal ratio of different loss terms can be achieved through grid search or automatic parameter tuning strategies based on the validation set.

[0171] For example, embodiments of this application may use the Adam optimizer to train a model, whose parameters... The learning rate for the 3D Gaussian parameters can be 1e-5, while the learning rate for the offset prediction module can be 2e-6. The perceptual loss weight is set to 0 for the first 15,000 training samples, and then to 0.05 thereafter. For each iteration during training, this embodiment can randomly sample 2,000 frames from the training dataset for training.

[0172] The training process involves continuous iteration until the loss function value converges to a preset threshold, or the reconstruction performance index on the test dataset reaches a preset standard. Finally, the trained 3D Gaussian sputtering head reconstruction model is output, enabling real-time rendering of head images from new perspectives, with new expressions and poses, and supporting cross-identity expression and pose transfer.

[0173] Among them, the model can be trained on a test platform with 11GB of memory and an NVIDIA RTX 2080 Ti, using PyTorch as the backend.

[0174] After model training, peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), learned perceptual image patch similarity (LPIPS), and frame rate (Frames Per Second, FPS) can be used as evaluation metrics to evaluate the head reconstruction results.

[0175] PSNR is used to measure the pixel-level distortion between the reconstructed image and the real image. It is derived by calculating the mean square error between the two images. The higher the value, the smaller the pixel error and the stronger the detail restoration and color fidelity of the reconstruction result.

[0176] SSIM is used to calculate image similarity from three dimensions: brightness, contrast, and structure. The closer the value is to 1, the more the reconstructed image matches the real image in terms of overall structure and local texture layout.

[0177] Among them, LPIPS is a learning-based perceptual similarity measurement method that is more in line with human perception. It is used to evaluate the high-level semantic and perceptual consistency between reconstructed images and real images. Specifically, it extracts deep semantic features of images through a pre-trained deep network and calculates the distance between the two in the feature space. The lower the value, the smaller the perceptual difference.

[0178] FPS is used to measure the smoothness of a real-time rendering system, and it is defined as the number of frames that can be continuously rendered per second. Different frame rate standards correspond to different application scenarios. For example, 30 FPS can meet basic real-time needs such as virtual try-on and simple interaction, while 100 FPS and above can meet the smoothness requirements of professional-grade real-time animation, high-fidelity virtual digital human live streaming and other scenarios.

[0179] In one possible embodiment, please refer to Figure 3 , Figure 3 This is a schematic diagram of a three-dimensional Gaussian sputtering head reconstruction model based on geometric mesh binding provided in an embodiment of this application, as shown below. Figure 3 As shown, the target object head to be reconstructed, input from the left, is used to obtain the standard mesh vertex positions through a face tracker. Each mesh is assigned a 3D Gaussian symmetric generator, which is then bound to each mesh vertex. A unique implicit code is also assigned to each 3D Gaussian symmetric generator. Then, based on facial expression pose... For standard mesh vertex positions Perform a linear blending skin transformation to obtain the vertex positions of the mesh after the linear blending skin transformation. .

[0180] Simultaneously, facial expression and posture Standard mesh vertex positions With implicit coding Input the deformation field network, which learns and outputs the position offset and other attribute offsets of a 3D Gaussian. The position offset is then compared with... The center position of the 3D Gaussian is obtained by concatenating the other attribute offsets with the other initial attributes after initialization, resulting in the final attribute set of each 3D Gaussian.

[0181] The number of 3D Gaussians is then dynamically increased or decreased to reduce the number while ensuring rendering quality. At the same time, it supports pose and expression transfer across identities. Finally, the model is iteratively optimized through the loss function, and the final output is the head reconstruction results under different poses and expressions, as well as the head reconstruction results under multiple views.

[0182] In one possible embodiment, the proposed implementation method is evaluated by combining visual results with quantitative metrics, and compared with three other representative neural rendering methods.

[0183] First, the embodiments of this application are qualitatively compared with three other methods. Among them, the Instant Volumetric Head Avatars (INSTA) method is based on real-time virtual digital head modeling technology using dynamic neural radiation fields. It achieves rapid reconstruction of high-fidelity digital humans from monocular color videos by embedding Neural Graphics Primitives (NGPs) around a parameterized face model in the form of multi-resolution hash codes.

[0184] Among them, the PointAvatar method represents the dynamic scene as a set of points in a normalized space and learns a conditional deformation field. It deforms the point cloud to the target space according to the pose and expression parameters of the FLAME model. At the same time, it adopts a coarse-to-fine point cloud growth strategy to gradually increase the point density to improve the details and achieve high-fidelity animation control.

[0185] Among them, the fast virtual avatar method, FlashAvatar, is a 3D Gaussian sputtering method based on UV parameter initialization. By unfolding the FLAME mesh into a 2D UV space to uniformly sample Gaussian points, and then mapping it back to the 3D surface, it achieves geometry-aware Gaussian distribution optimization, significantly reducing redundant computation. Here, U and V correspond to the horizontal and vertical axes of the 2D plane, respectively.

[0186] Please refer to Figure 4 , Figure 4 This is a schematic diagram illustrating a qualitative comparison of head reconstruction effects provided in an embodiment of this application, as shown below. Figure 4 As shown, the first column includes real images of 5 different objects, the second column includes reconstructed head images of 5 different objects output by the PointAvatar method, the third column includes reconstructed head images of 5 different objects output by the INSTA method, the fourth column includes reconstructed head images of 5 different objects output by the FlashAvatar method, and the fifth column includes reconstructed head images of 5 different objects output by the method of this application.

[0187] Among them, according to Figure 4 As shown in the green and red boxes, while the PointAvatar method can generate smooth volumetric effects, it struggles to render sharp images, such as blurred reconstructed teeth and the inability to reproduce fine hairs. The INSTA method, based on the nearest triangle deformation point, exhibits misalignment between the FLAME mesh and the target geometry, significantly reducing rendering quality, especially noticeable noise artifacts in the mouth, hair, and eye areas. The FlashAvatar method and the method described in this application, while achieving fast rendering, also provide high-quality rendering of elements like hair and glasses that are lacking in the geometric prior model.

[0188] In this application, all real images involved in the embodiments have been explicitly authorized and agreed upon by the relevant parties, or are synthetic images generated by artificial intelligence technology, or are virtual images constructed based on 3D modeling tools; and all image data are used only for model training and verification of the technical solution of this application.

[0189] In one possible embodiment, in order to further measure the performance difference between the FlashAvatar method and the method of this application, a comparison strategy is designed. During the training process, the average PSNR value of the most recent 50 generated images is calculated. When both can render high-quality images and are at the same PSNR level, the training is stopped, and the number of three-dimensional Gaussians used by the model at this time is counted.

[0190] Please refer to Figure 5 , Figure 5 This is a schematic diagram illustrating a quantitative comparison of head reconstruction effects provided in an embodiment of this application, as shown below. Figure 5 As shown, the image includes head reconstruction from two perspectives. The first column contains real images of three different objects from the first perspective. The second column contains head reconstruction images of three different objects output by the FlashAvatar method from the first perspective. The third column contains head reconstruction images of three different objects output by the method of this application from the first perspective. The fourth column contains real images of three different objects from the second perspective. The fifth column contains head reconstruction images of three different objects output by the FlashAvatar method from the second perspective. The sixth column contains head reconstruction images of three different objects output by the method of this application from the second perspective.

[0191] The FlashAvatar with a UV mapping resolution of 128x128 has 13k Gaussians, while this application can achieve the same quality rendering effect with only 10k Gaussians.

[0192] In one possible embodiment, for quantitative evaluation, this application uses three metrics—PSNR, SSIM, and LPIPS—to measure the degree of matching between the generated result and the real image and the quality of the image, and records its rendering speed in FPS.

[0193] Please refer to Table 1, which shows that this application demonstrates significant advantages in multiple metrics. Specifically, in the PSNR metric, which measures image fidelity, this application significantly outperforms the PointAvatar, INSTA, and FlashAvatar methods. Particularly noteworthy is that this application achieved the lowest LPIPS score, directly reflecting higher visual realism and a closer resemblance to a real portrait. Table 1 is as follows:

[0194] Table 1. Quantitative comparison results of different methods

[0195]

[0196] Since the deformation of the FLAME mesh and the offset of the three-dimensional Gaussian are both based on the facial expression and pose encoding decoupled from the identity space, this application can easily realize the facial expression replay task with ultra-high-speed rendering performance, achieving a rendering speed of up to 331 FPS.

[0197] Furthermore, in this embodiment, the virtual head is represented by a three-dimensional Gaussian with multi-view consistency, allowing for free adjustment of the global camera pose. Please refer to [link / reference]. Figure 6 , Figure 6 This is a schematic diagram illustrating a multi-view head reconstruction result provided in an embodiment of this application, as shown below. Figure 6 As shown, the first column is the original view image, which includes images of the same object from two different initial angles; the second, third and fourth columns are all composite images from new viewpoints, but the second, third and fourth columns are composite images from different viewpoints generated based on the original view images at the corresponding positions in the first column.

[0198] Among them, from Figure 6 As can be seen, based on the original images from different initial angles in the first column, new perspective content with a unified style and consistent details is generated in the second to fourth columns. These synthesized images not only accurately reproduce the facial features, skin texture and expression details of the objects in the original images, but also naturally adapt to different viewing angles, realizing high-quality rendering extension from a single original perspective to any target perspective. This not only solves the problem of high cost of acquiring multi-view content, but also ensures the realism and feature consistency of the synthesized results.

[0199] In one possible embodiment, replicating shape appearance changes caused by different expressions and poses is crucial for virtual digital head evaluation. This application embodiment uses the FLAME expression and pose parameters of the source character to drive the reconstruction of the target character.

[0200] Please see Figure 7 , Figure 7 This is a schematic diagram illustrating a cross-identity transfer facial expression and pose generation head reconstruction result provided in an embodiment of this application, as shown below. Figure 7 As shown, the first column contains real images of three different objects and their corresponding local details, namely close-ups of lip expressions; the second column contains cross-identity expression and pose transfer results generated by the PointAvatar method based on the images in the first column; the third column contains cross-identity transfer and reconstruction results generated by the INSTA method based on the images in the first column; the fourth column contains cross-identity transfer and reconstruction results generated by the FlashAvatar method based on the images in the first column; and the fifth column contains cross-identity expression and pose transfer results generated by the method of this application based on the images in the first column.

[0201] Among them, from Figure 7 As can be seen, different methods all drive the reconstruction of the target character based on the FLAME parameters of the source character. However, the results of methods such as PointAvatar, INSTA, and FlashAvatar have certain deviations in details such as lip shape matching and facial texture naturalness. In contrast, the results of this application in the fifth column not only accurately align the expression and pose corresponding to the source parameters, but also restore the personalized facial details of the target character, such as the dynamic shape of the lips and the realism of facial texture, achieving a higher fidelity cross-identity transfer rendering effect.

[0202] In summary, this application maintains real-time rendering performance while still providing superior image quality, surpassing existing methods in multiple dimensions. Therefore, considering both qualitative and quantitative results, this application outperforms the aforementioned existing methods.

[0203] In one possible embodiment, an ablation experiment was conducted to verify the effectiveness of the embodiments of this application, and the analysis was mainly performed qualitatively and / or quantitatively.

[0204] Ablation experiments were conducted regarding the binding strategy. In this embodiment, a 3D Gaussian binding mesh vertices is used to provide prior geometric features. To evaluate the effectiveness of this scheme, a 3D Gaussian model without the binding strategy was trained as a control. Quantitative analysis is shown in the second and last rows of Table 2. This application exhibits high PSNR and SSIM, and low LPIPS. Qualitative analysis is provided in [reference needed]. Figure 8 , Figure 8 This is a schematic diagram of an ablation experiment for verifying a binding strategy provided in an embodiment of this application, as shown below. Figure 8 As shown, high PSNR and SSIM, and low LPIPS, not only indicate a decrease in image quality but also suggest that the 3D Gaussian distribution cannot be constrained to the vicinity of the head, resulting in noticeable floating artifacts. This may be due to the excessive offset of the 3D Gaussian distribution from the head without binding the mesh vertices, leading to inadequate constraint during gradient optimization. Furthermore, directly learning the dynamic changes in position using the 3D Gaussian distribution without leveraging the geometric deformation capabilities of the personality mesh makes it difficult for the neural network to converge. Therefore, this sufficiently demonstrates the effectiveness of the proposed strategy regarding binding geometric meshes. Table 2 is as follows:

[0205] Table 2 Quantitative comparison results of ablation experiments

[0206]

[0207] The ablation experiment included dynamic addition and deletion. This application proposes a strategy to dynamically add and delete 3D Gaussian elements during 3D Gaussian training, and to inherit the vertex index attributes of the cloned Gaussian elements from the newly added ones. To evaluate the effectiveness of this scheme, a fixed number of 10,000 3D Gaussian elements is used for comparison. The quantitative analysis is shown in the third and last rows of Table 2. It can be observed that the image quality after removing the dynamic addition and deletion strategy is not high, which fully demonstrates the necessity and rationality of the above strategy.

[0208] The invention includes an ablation experiment on the latent coding properties of 3D Gaussians. This embodiment extends the basic properties of 3D Gaussians, introducing latent coding properties that serve two purposes. First, latent coding identifies the uniqueness of a Gaussian, allowing a multilayer perceptron to distinguish between different 3D Gaussians. Second, the geometric prior features of the FLAME mesh used in this embodiment are inevitably affected by its inability to accurately fit real human heads with complex geometry, such as poor performance in detailed structural areas like hair and glasses. To compensate for this deficiency, latent coding is used to learn these features and apply them to the multilayer perceptron to compensate for the lack of geometric structure. To evaluate the effectiveness of this approach, this embodiment trains a model using positional coding instead of this approach for comparison. Quantitative analysis is shown in the fourth and last rows of Table 2. It can be noted that this application enhances image rendering quality, demonstrating its effectiveness.

[0209] As can be seen, in this embodiment of the application, by introducing additional three-dimensional parameterized facial expression pose parameters of the human head as conditions, high-quality rendering with multi-view consistency can be achieved for the target character, and a training speed of rapid convergence and a rendering speed of up to 300 FPS can be achieved within tens of minutes.

[0210] For examples consistent with the above embodiments, please refer to... Figure 9 , Figure 9 This is a functional unit block diagram of a three-dimensional head reconstruction device based on a geometric prior network provided in an embodiment of this application, as shown below. Figure 9As shown, the 3D head reconstruction device 90 based on a geometric prior network includes: an extraction unit 91, used to extract multiple pose parameters and multiple expression parameters from each frame of a preprocessed monocular head video; a deformation unit 92, used to perform linear hybrid skinning transformation on multiple initial mesh vertex positions of each frame to obtain multiple first mesh vertex positions of each frame, and assign a reference 3D Gaussian to each first mesh vertex position; and an initialization unit 93, used to initialize each reference 3D Gaussian to obtain a first attribute set, the first attribute set including at least the first mesh vertex position, latent code, and vertex index, wherein the latent code is used to characterize the multimodal feature fusion of a single first mesh vertex position; the first... A determining unit 94 is configured to determine the geometric offset of each reference 3D Gaussian based on the plurality of pose parameters, the plurality of expression parameters, the implicit coding, and the plurality of initial mesh vertex positions; a second determining unit 95 is configured to determine a second attribute set of each reference 3D Gaussian based on the first attribute set and the geometric offset, wherein the second attribute set includes at least the second mesh vertex position, the implicit coding, and the vertex index, wherein the second mesh vertex position is bound to the vertex index, and the second mesh vertex position is determined based on the first mesh vertex position and the geometric offset; a third determining unit 96 is configured to determine the head reconstruction result based on the second attribute set of each reference 3D Gaussian.

[0211] In one possible embodiment, the first attribute set further includes a first rotation direction, a first scaling factor, opacity, and color. Regarding determining the geometric offset of each reference 3D Gaussian based on the plurality of pose parameters, the plurality of expression parameters, the implicit encoding, and the plurality of initial mesh vertex positions, the first determining unit 94 is specifically configured to: fuse the plurality of pose parameters and the plurality of expression parameters to obtain fused parameters; perform position encoding on each initial mesh vertex position using a position encoding module in a pre-trained deformation field network model to obtain encoded features; perform offset prediction on the attributes of each reference 3D Gaussian based on the fused parameters, the encoded features, and the implicit encoding using an offset prediction module in the deformation field network model to obtain a first offset of each initial mesh vertex position, a second offset of the first rotation direction, and a third offset of the first scaling factor; and obtain the geometric offset of each reference 3D Gaussian based on the first offset, the second offset, and the third offset.

[0212] In one possible embodiment, in determining the second attribute set of each reference 3D Gaussian based on the first attribute set and the geometric offset, the second determining unit 95 is specifically configured to: determine the second mesh vertex position based on the first offset and the position of each first mesh vertex; bind the second mesh vertex position and the vertex index, wherein the second mesh vertex position and the vertex index are fixed attributes; determine a second rotation direction based on the first rotation direction and the second offset; determine a second scaling scale based on the first scaling scale and the third offset; and update the first attribute set based on the second mesh vertex position, the second rotation direction, and the second scaling scale to obtain the second attribute set.

[0213] In one possible embodiment, in determining the head reconstruction result based on the second attribute set of each reference 3D Gaussian, the third determining unit 96 is specifically configured to: determine the view space of each reference 3D Gaussian; determine the position gradient of the view space; if the position gradient is greater than a first threshold, determine the second attribute set of a newly added 3D Gaussian based on the second attribute set of the reference 3D Gaussian corresponding to the position gradient; obtain the opacity of each reference 3D Gaussian and the newly added 3D Gaussian; if the opacity is lower than a second threshold, delete the 3D Gaussian corresponding to the opacity from each reference 3D Gaussian and the newly added 3D Gaussian to obtain the remaining 3D Gaussian; and determine the head reconstruction result based on the second attribute set of the remaining 3D Gaussian.

[0214] In one possible embodiment, in determining the second attribute set of the newly added three-dimensional Gaussian based on the second attribute set of the reference three-dimensional Gaussian corresponding to the position gradient, the third determining unit 96 is further configured to: use the second attribute set of the reference three-dimensional Gaussian corresponding to the position gradient as the reference attribute set of the newly added three-dimensional Gaussian; adjust the size of the newly added three-dimensional Gaussian according to the size of the reference three-dimensional Gaussian corresponding to the position gradient to obtain a target size; randomly perturb the implicit code of the reference three-dimensional Gaussian corresponding to the position gradient to obtain the implicit code of the newly added three-dimensional Gaussian; and update the reference attribute set according to the target size and the implicit code of the newly added three-dimensional Gaussian to obtain the second attribute set of the newly added three-dimensional Gaussian.

[0215] In one possible embodiment, in determining the head reconstruction result based on the second attribute set of the remaining three-dimensional Gaussian, the third determining unit 96 is further configured to: through a projection operation, map each remaining three-dimensional Gaussian to a local pixel contribution region in a two-dimensional image space based on the second attribute set of the remaining three-dimensional Gaussian, to obtain a set of pixel feature maps, wherein each pixel feature map in the set is a two-dimensional projection feature representation of a single remaining reference three-dimensional Gaussian; perform feature fusion processing on each pixel feature map to obtain a fused feature map; and render the fused feature map to obtain the head reconstruction result.

[0216] In one possible embodiment, after determining the head reconstruction result based on the second attribute set of each reference 3D Gaussian, the 3D head reconstruction device 90 based on the geometric prior network is further configured to: determine the pixel-level difference between the head reconstruction result and the target image to obtain a first loss, wherein the target image is used to characterize a single frame in the monocular head video corresponding to the head reconstruction result; extract semantic features associated with identity in the head reconstruction result and the target image; determine the similarity between the semantic features of the head reconstruction result and the semantic features of the target image to obtain a second loss; optimize the learnable attributes of the reference 3D Gaussian and the parameters of the deformation field network model based on the first loss and the second loss to obtain optimized learnable attributes and the parameter configuration of the deformation field network model, wherein the learnable attributes include the first rotation direction, the first scaling scale, the opacity, the color, and the implicit coding; and reconstruct the head in the next frame of the monocular head video based on the optimized learnable attributes and the parameter configuration of the deformation field network model.

[0217] It is understood that since the method embodiments and the device embodiments are different presentations of the same technical concept, the content of the method embodiment section in this application should be adapted to the device embodiment section in a synchronous manner, and will not be repeated here.

[0218] In the case of using integrated units, please refer to Figure 10 , Figure 10 This is a functional unit block diagram of another three-dimensional head reconstruction device based on a geometric prior network provided in this application embodiment, such as... Figure 10As shown, the 3D head reconstruction device 90 based on a geometric prior network includes a processing module 902 and a communication module 901. The processing module 902 controls and manages the actions of the 3D head reconstruction device 90 based on the geometric prior network, for example, executing the steps of the extraction unit 91, deformation unit 92, initialization unit 93, first determination unit 94, second determination unit 95, and third determination unit 96, and / or other processes of the technology described herein. The communication module 901 is used for interaction between the 3D head reconstruction device 90 based on the geometric prior network and other devices.

[0219] Among them, such as Figure 10 As shown, the three-dimensional head reconstruction device 90 based on geometric prior networks may further include a storage module 903, which is used to store the program code and data of the three-dimensional head reconstruction device 90 based on geometric prior networks.

[0220] The processing module 902 can be a processor or controller, such as a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0221] The communication module 901 can be a transceiver, RF circuit, or communication interface, etc. The storage module 903 can be a memory.

[0222] All relevant content in each scenario involved in the above method embodiments can be referenced from the functional descriptions of the corresponding functional modules, and will not be repeated here. The above-mentioned three-dimensional head reconstruction device 90 based on geometric prior networks can perform the above-mentioned... Figure 2 The method for 3D head reconstruction based on geometric prior networks is shown.

[0223] Please see Figure 11 , Figure 11 This is a schematic diagram of the structure of an electronic device proposed in an embodiment of this application, such as... Figure 11As shown, the electronic device 1100 includes a processor 1110, a memory 1120, a communication interface 1130, and one or more programs 1121. The one or more programs 1121 are stored in the memory and configured to be executed by the processor. When the program is executed, it includes some or all of the steps of any three-dimensional head reconstruction method based on geometric prior networks described in the above method embodiments. The processor, memory, and communication interface are interconnected and complete communication between them.

[0224] The memory can be volatile memory such as Dynamic Random Access Memory (DRAM) or non-volatile memory such as a hard disk drive. The memory stores a set of executable program code, and the processor calls the executable program code stored in the memory to execute some or all of the steps of any of the geometric prior network-based 3D head reconstruction methods described in the above embodiments of the geometric prior network-based 3D head reconstruction method.

[0225] As can be seen, the electronic device 1100 described in this application embodiment first extracts multiple pose parameters and multiple expression parameters from each frame of the preprocessed monocular head video; then, it performs linear blending skinning transformation on multiple initial mesh vertex positions of each frame to obtain multiple first mesh vertex positions of each frame, and assigns a reference 3D Gaussian to each first mesh vertex position; then, it initializes each reference 3D Gaussian to obtain a first attribute set, which includes at least the first mesh vertex position, the latent code, and the vertex index, wherein the latent code is used to characterize the multimodal feature fusion of a single first mesh vertex position; then, it determines the geometric offset of each reference 3D Gaussian based on the multiple pose parameters, the multiple expression parameters, the latent code, and the multiple initial mesh vertex positions; then, it determines a second attribute set of each reference 3D Gaussian based on the first attribute set and the geometric offset, wherein the second attribute set includes at least the second mesh vertex position, the latent code, and the vertex index, wherein the second mesh vertex position is bound to the vertex index, and the second mesh vertex position is determined based on the first mesh vertex position and the geometric offset; finally, it determines the head reconstruction result based on the second attribute set of each reference 3D Gaussian.

[0226] This application optimizes the Gaussian initialization accuracy and constrains the Gaussian spatial distribution by binding each 3D Gaussian vertices to a mesh vertex, which helps ensure the consistency of the overall geometric structure during head reconstruction. Simultaneously, latent coding attributes are introduced during the 3D Gaussian initialization stage, providing a unique identifier for each 3D Gaussian vertices. This latent coding is applied to the prediction of geometric offsets, filling the gaps in detailed structures such as hair and glasses areas within the mesh, thus improving the detail reproduction of the head reconstruction. By using pose parameters, expression parameters, initial mesh vertex positions, and latent coding, the geometric offsets of each 3D Gaussian vertices are predicted, fully considering the dynamic changes in Gaussian attributes such as vertex positions under different facial expressions and poses, achieving accurate capture of non-rigid deformation features. Finally, head reconstruction is completed based on the second attribute set of the 3D Gaussians, ensuring that the output head reconstruction result accurately matches facial expressions and poses, achieving high-precision head reconstruction.

[0227] This application also provides a computer storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of any of the methods described in the above method embodiments, wherein the computer includes an electronic device.

[0228] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer may include an electronic device.

[0229] It should be noted that, for the sake of simplicity, the aforementioned methods are described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are optional, and the actions and modules involved are not necessarily essential to this application.

[0230] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0231] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.

[0232] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0233] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software program module.

[0234] If the integrated unit is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0235] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage device, which may include: a flash drive, a read-only memory, a random access memory, a magnetic disk, or an optical disk, etc.

[0236] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The above description of the embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for three-dimensional head reconstruction based on geometric prior networks, characterized in that, include: Extract multiple pose parameters and multiple facial expression parameters from each frame of the preprocessed monocular head video; A linear blending skin transformation is performed on multiple initial mesh vertex positions of each frame to obtain multiple first mesh vertex positions of each frame, and a reference 3D Gaussian is assigned to each first mesh vertex position. Each reference 3D Gaussian is initialized to obtain a first attribute set, which includes at least the first mesh vertex position, latent code, vertex index, first rotation direction, first scaling scale, opacity, and color. The latent code is used to characterize the multimodal feature fusion of a single first mesh vertex position. Based on the multiple pose parameters, multiple expression parameters, the hidden encoding, and the multiple initial mesh vertex positions, the geometric offset of each reference 3D Gaussian is determined; wherein, the multiple pose parameters and multiple expression parameters are fused to obtain fusion parameters; the position encoding module in the pre-trained deformation field network model encodes the position of each initial mesh vertex to obtain encoded features; the offset prediction module in the deformation field network model predicts the offset of each reference 3D Gaussian's attributes based on the fusion parameters, the encoded features, and the hidden encoding to obtain a first offset of each initial mesh vertex position, a second offset of the first rotation direction, and a third offset of the first scaling scale; based on the first offset, the second offset, and the third offset, the geometric offset of each reference 3D Gaussian is obtained. Based on the first attribute set and the geometric offset, a second attribute set is determined for each reference 3D Gaussian. The second attribute set includes at least the second mesh vertex position, the implicit code, and the vertex index. The second mesh vertex position is bound to the vertex index and is determined based on the first mesh vertex position and the geometric offset. Specifically, the second mesh vertex position is determined based on the first offset and each first mesh vertex position; the second mesh vertex position and the vertex index are bound, and these are fixed attributes; a second rotation direction is determined based on the first rotation direction and the second offset; a second scaling scale is determined based on the first scaling scale and the third offset; and the first attribute set is updated based on the second mesh vertex position, the second rotation direction, and the second scaling scale to obtain the second attribute set. The head reconstruction result is determined based on the second set of attributes of each reference 3D Gaussian.

2. The method according to claim 1, characterized in that, The step of determining the head reconstruction result based on the second attribute set of each reference 3D Gaussian includes: Determine the view space for each reference 3D Gaussian; Determine the position gradient of the view space; If the position gradient is greater than the first threshold, then the second attribute set of the newly added three-dimensional Gaussian is determined according to the second attribute set of the reference three-dimensional Gaussian corresponding to the position gradient. Obtain the opacity of each reference 3D Gaussian and the newly added 3D Gaussian; If the opacity is lower than the second threshold, then in each reference 3D Gaussian and the newly added 3D Gaussian, the 3D Gaussian corresponding to the opacity is deleted to obtain the remaining 3D Gaussian. The head reconstruction result is determined based on the second set of attributes of the remaining three-dimensional Gaussian.

3. The method according to claim 2, characterized in that, The step of determining the second attribute set of the newly added three-dimensional Gaussian based on the second attribute set of the reference three-dimensional Gaussian corresponding to the position gradient includes: The second attribute set of the reference three-dimensional Gaussian corresponding to the position gradient is used as the reference attribute set of the newly added three-dimensional Gaussian. Based on the dimensions of the reference 3D Gaussian corresponding to the position gradient, the dimensions of the newly added 3D Gaussian are adjusted to obtain the target dimension; The hidden code of the reference 3D Gaussian corresponding to the position gradient is randomly perturbed to obtain the hidden code of the newly added 3D Gaussian. Based on the target size and the implicit encoding of the newly added 3D Gaussian, the reference attribute set is updated to obtain the second attribute set of the newly added 3D Gaussian.

4. The method according to claim 2, characterized in that, Determining the head reconstruction result based on the second attribute set of the remaining three-dimensional Gaussian includes: Through projection operations, each remaining three-dimensional Gaussian is mapped to a local pixel contribution region in a two-dimensional image space according to the second attribute set of the remaining three-dimensional Gaussian, resulting in a set of pixel feature maps. Each pixel feature map in the set of pixel feature maps is a two-dimensional projection feature representation of a single remaining reference three-dimensional Gaussian. Each pixel feature map is subjected to feature fusion processing to obtain a fused feature map; The fused feature map is rendered to obtain the head reconstruction result.

5. The method according to claim 1, characterized in that, After determining the head reconstruction result based on the second attribute set of each reference 3D Gaussian, the method further includes: The pixel-level difference between the head reconstruction result and the target image is determined to obtain the first loss, and the target image is used to characterize a single frame in the monocular head video corresponding to the head reconstruction result; Extract the head reconstruction results and the semantic features associated with identity from the target image; The similarity between the semantic features of the reconstructed head and the semantic features of the target image is determined to obtain the second loss. Based on the first loss and the second loss, the learnable properties of the reference 3D Gaussian and the parameters of the deformation field network model are optimized to obtain the optimized learnable properties and the parameter configuration of the deformation field network model. The learnable properties include the first rotation direction, the first scaling scale, the opacity, the color, and the hidden encoding. Based on the optimized learnable attributes and the parameter configuration of the deformation field network model, head reconstruction is performed on the next frame of the monocular head video.

6. A three-dimensional head reconstruction device based on a geometric prior network, characterized in that, include: The extraction unit is used to extract multiple pose parameters and multiple facial expression parameters from each frame of the preprocessed monocular head video. The deformation unit is used to perform linear blending skinning transformation on multiple initial mesh vertex positions of each frame to obtain multiple first mesh vertex positions of each frame, and assign a reference three-dimensional Gaussian to each first mesh vertex position. An initialization unit is used to initialize each reference 3D Gaussian to obtain a first attribute set. The first attribute set includes at least the first mesh vertex position, latent code, vertex index, first rotation direction, first scaling scale, opacity, and color. The latent code is used to characterize the multimodal feature fusion of a single first mesh vertex position. The first determining unit is configured to determine the geometric offset of each reference 3D Gaussian based on the plurality of pose parameters, the plurality of expression parameters, the implicit encoding, and the plurality of initial mesh vertex positions; specifically, it is configured to fuse the plurality of pose parameters and the plurality of expression parameters to obtain fusion parameters; perform position encoding on the position of each initial mesh vertex using the position encoding module in the pre-trained deformation field network model to obtain encoded features; perform offset prediction on the attributes of each reference 3D Gaussian based on the fusion parameters, the encoded features, and the implicit encoding using the offset prediction module in the deformation field network model to obtain a first offset of each initial mesh vertex position, a second offset of the first rotation direction, and a third offset of the first scaling scale; and obtain the geometric offset of each reference 3D Gaussian based on the first offset, the second offset, and the third offset. The second determining unit is configured to determine a second attribute set for each reference 3D Gaussian based on the first attribute set and the geometric offset. The second attribute set includes at least the second mesh vertex position, the implicit code, and the vertex index. The second mesh vertex position is bound to the vertex index, and the second mesh vertex position is determined based on the first mesh vertex position and the geometric offset. Specifically, the unit is configured to determine the second mesh vertex position based on the first offset and each first mesh vertex position; bind the second mesh vertex position and the vertex index, which are fixed attributes; determine a second rotation direction based on the first rotation direction and the second offset; determine a second scaling scale based on the first scaling scale and the third offset; and update the first attribute set based on the second mesh vertex position, the second rotation direction, and the second scaling scale to obtain the second attribute set. The third determining unit is used to determine the head reconstruction result based on the second attribute set of each reference three-dimensional Gaussian.

7. An electronic device, characterized in that, The device includes: The device includes a memory, a processor, and executable program code stored in the memory and executable on the processor, wherein the processor executes the executable program code to perform the steps of the three-dimensional head reconstruction method based on a geometric prior network as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores executable program code, which includes execution instructions for performing the steps of the three-dimensional head reconstruction method based on geometric prior networks as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Emotion-controllable facial animation generation method and device, equipment and medium

    CN118691725A

  • Expression-editable voice-driven face reconstruction method based on three-dimensional Gaussian sputtering technology

    CN118762133A