A digital human body reconstruction method focusing on the face

Through a digital human body reconstruction method that focuses on the face, using FLAME model and multi-layer perceptron technology, the problems of blurred and unnatural animation of digital human faces are solved, and high-precision and smooth rendering of human body animations are achieved.

CN119888028BActive Publication Date: 2025-06-03CHENGDU UNIV OF INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510361546.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-03
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

Existing digital human reconstruction methods are blurry or unnatural when reconstructing digital human faces, and it is difficult to achieve accurate face animation in fast-moving frames and subtle frames.

Method used

A digital human body reconstruction method focusing on the face is proposed. By obtaining the RGB image frame sequence of the dynamic human body, the shape parameters, expression parameters and head camera position of the FLAME model are predicted, the head grid is constructed, and the Gaussky primitive parameters are optimized through a multi-layer perceptron to achieve high-precision face reconstruction and animation rendering.

Benefits of technology

It realizes high-precision digital human facial reconstruction and animation rendering, presenting smooth and realistic human animation effects, and does not require multi-view training. It only requires monocular video to achieve high-definition face and full body rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888028B_ABST
    Figure CN119888028B_ABST
Patent Text Reader

Abstract

The present invention discloses a digital human body reconstruction method focusing on the face, which relates to the technical field of digital human reconstruction. It includes successively performing super-resolution processing and semantic segmentation on the face area recognized by face recognition to obtain the current face RGB image frame; based on the initialization parameters and the expression parameters of the current RGB image frame, obtaining new parameters of the Gaussian basis elements of the head mesh corresponding to the current RGB image frame through a multi-layer perceptron; through a training strategy of minimizing the loss, optimizing the parameters of the Gaussian basis elements of the head mesh corresponding to the current RGB image frame, iteratively processing other RGB image frames, and combining all the optimized Gaussian basis elements to obtain a reconstructed digital head; rendering a human body image focusing on the face according to the reconstructed digital head and the reconstructed digital human body. The present invention not only focuses on the overall reconstruction of the human body, but also strengthens the high-precision reconstruction of the face, and can present a smooth and realistic human body animation effect in subsequent rendering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital human reconstruction, and particularly to a digital human body reconstruction method focusing on the face. Background Art

[0002] Traditional digital human reconstruction methods usually focus on the reconstruction of the body while ignoring facial details, which may result in a relatively blurred or unnatural face of the reconstructed digital human. On the contrary, only focusing on the reconstruction of the human face while ignoring the reconstruction of the human body. To achieve more natural human body animations and more realistic facial rendering effects, a high-precision digital human reconstruction model needs to be constructed to make the rendering effect of the digital human in the virtual environment more realistic and delicate.

[0003] Traditional digital human reconstruction methods mainly include the following: Existing methods for full-body human reconstruction and fine facial reconstruction mainly include the following aspects: The method of digital human reconstruction based on Nerf has the problem of low quality of both the whole body and the face; The method of reconstructing digital humans using a learned dynamic neural network has the problem that accurate facial animations cannot be achieved in fast-moving frames and subtle frames, and there are also significant artifacts for unseen poses; The digital human reconstruction method based on Gaussian still has obvious deficiencies in the reconstruction effect of the face area; There are also some methods that separately reconstruct the head using a monocular video. For virtual reality, it is obviously unreasonable that the interactive object is only the head; The SMPL-X model requires multi-view training and has high requirements for hardware devices; ExAvatar and GaussianAvatar have incorrect expressions and artifacts.

[0004] Therefore, a digital human body reconstruction method focusing on the face is developed to solve the above problems. Summary of the Invention

[0005] The present invention proposes a digital human body reconstruction method focusing on the face to solve the problem that the face of the digital human reconstructed by the existing digital human reconstruction methods is relatively blurred or unnatural.

[0006] The present invention achieves the above object through the following technical solutions:

[0007] A digital human body reconstruction method focusing on the face according to the present invention includes:

[0008] Obtaining first information, where the first information includes a sequence of RGB image frames of a dynamic human body;

[0009] Obtain second information and third information according to the first information, where the second information includes the shape parameters, expression parameters, and head camera poses of the FLAME model corresponding to each predicted RGB image frame, and the third information includes the shape parameters, pose parameters, and body camera poses of the digital human reconstruction model corresponding to each predicted RGB image frame;

[0010] Construct a head mesh corresponding to each RGB image frame based on the preset FLAME model according to the second information;

[0011] Uniformly sample the UV texture map of the head mesh corresponding to each RGB image frame to obtain corresponding sampling results;

[0012] Initialize the Gaussian field for the corresponding head mesh according to the corresponding sampling results to obtain the initialization parameters of the Gaussian basis elements of the corresponding head mesh;

[0013] Enter the processing step of the current RGB image frame:

[0014] Perform face recognition on the current RGB image frame, and perform super-resolution processing and semantic segmentation on the recognized face region in sequence to obtain the current face RGB image frame;

[0015] Based on the initialization parameters and the expression parameters of the current RGB image frame, obtain the new parameters of the Gaussian basis elements of the head mesh corresponding to the current RGB image frame based on a multi-layer perceptron;

[0016] Render a head rendering picture according to the new parameters of the Gaussian basis elements of the head mesh corresponding to the current RGB image frame and the head camera pose, calculate the loss between the head rendering picture and the current face RGB image frame, and optimize the parameters of the Gaussian basis elements of the head mesh corresponding to the current RGB image frame through a training strategy of minimizing the loss, and then end the processing step of the current RGB image frame;

[0017] Iteratively process other RGB image frames according to the processing step of the current RGB image frame until the parameters of the Gaussian basis elements of the head meshes corresponding to all the RGB image frames are optimized, and combine all the optimized Gaussian basis elements to obtain a reconstructed digital head;

[0018] Input the third information into a preset digital human reconstruction model to output a reconstructed digital human;

[0019] Render a human image with a focused face according to the reconstructed digital head and the reconstructed digital human.

[0020] Further, obtaining the second information and the third information according to the first information includes:

[0021] Obtain the RGB image frame sequence of a dynamic human body captured by a monocular camera;

[0022] Input the RGB image frame sequence of the dynamic human body captured by the monocular camera into a pre-trained human head parameter prediction model, and output the second information;

[0023] Input the RGB image frame sequence of the dynamic human body captured by the monocular camera into a pre-trained human parameter prediction model, and output the third information.

[0024] Further, the initialization parameter G is expressed as:

[0025]

[0026] Among them, μ represents the texture coordinate, and a fixed point on the head mesh can be located through the texture coordinate μ , and this point is the position of the Gaussian basis element, represents the rotation coefficient, represents the scaling coefficient, α represents the opacity, h represents the spherical harmonic function, is initialized as the quaternion (1, 0, 0, 0), is initialized as the distance to the Gaussian basis element with the closest distance, and the calculation formula is as follows:

[0027]

[0028] Among them, represents the center position of the Gaussian basis element with the closest distance to the point , and the opacity α of each Gaussian basis element is initialized to 0.1, and the spherical harmonic function h is initialized as a tensor of dimension, where N is the number of Gaussian basis elements, the spherical harmonic function order d = 3, the value of the tensor at the position (N, 3, 0) is 0.5, and the rest are 0, and the dist function is used to calculate the Euclidean distance between two coordinate points.

[0029] Further, perform face recognition on the current RGB image frame, and perform super-resolution processing and semantic segmentation on the recognized face area in sequence to obtain the current face RGB image frame, including:

[0030] Perform face recognition on the current RGB image frame based on the face recognition algorithm to obtain the face area;

[0031] Perform super-resolution processing on the face area to obtain a clear face picture;

[0032] Perform semantic segmentation on the clear face picture to obtain the current face RGB image frame.

[0033] Further, based on the initialization parameters and the expression parameters of the current RGB image frame, new parameters of the Gaussian basis elements of the head mesh corresponding to the current RGB image frame are obtained based on a multi-layer perceptron, including:

[0034] Encode the texture coordinates;

[0035] Input the encoded result and the expression parameters of the current RGB image frame into the multi-layer perceptron, and output the offset, rotation amount, and scaling amount of the position of the Gaussian basis elements of the head mesh;

[0036] Add the offset to the position of the Gaussian basis element to obtain the new position of the Gaussian basis element of the head mesh corresponding to the current RGB image frame;

[0037] Multiply the rotation amount by the rotation coefficient to obtain the new rotation coefficient of the Gaussian basis element of the head mesh corresponding to the current RGB image frame;

[0038] Multiply the scaling amount by the scaling coefficient to obtain the new scaling coefficient of the Gaussian basis element of the head mesh corresponding to the current RGB image frame.

[0039] Further, the new parameters of the Gaussian basis elements of the head mesh corresponding to the current RGB image frame are expressed as:

[0040]

[0041]

[0042] Among them, , , represent the offset, rotation amount, and scaling amount in sequence, D represents the multi-layer perceptron, θ i represents the expression parameters of the current RGB image frame, represents the encoded result, , , represent the new position, new rotation coefficient, and new scaling coefficient of the Gaussian basis element of the head mesh corresponding to the current RGB image frame in sequence.

[0043] Further, the calculation steps of the loss between the head rendering picture and the current face RGB image frame are as follows:

[0044] Render the head rendering picture through the new position, new rotation coefficient, new scaling coefficient, and opacity of the Gaussian basis element of the head mesh corresponding to the final current RGB image frame using spherical harmonics ;

[0045] The rendered head rendering picture Perform loss descent optimization with the current face RGB image frame, and the formula is as follows:

[0046]

[0047]

[0048] where is the dynamic human body image captured by the camera, is the position of the face recognition frame, crop is image cropping, super is image super-resolution, is the result of the super-resolution processing, is the result of the semantic segmentation, is the total loss function, and loss is the loss.

[0049] Furthermore, the total loss function is:

[0050]

[0051] where λ 1 and λ 2 are constants, and the Huber loss is:

[0052]

[0053] where represents the image captured by the camera, represents the rendered image, ;

[0054] Its perceptual similarity loss is:

[0055]

[0056] where represents the rendered image at the i-th layer feature map in the VGG deep network, where represents the image captured by the camera at the i-th layer feature map in the VGG deep network, N is the number of feature layers for calculating the loss, represents calculating of the L2 norm and then squaring.

[0057] Furthermore, render a human body image with a focused face based on the reconstructed digital head and the reconstructed digital human body, including:

[0058] Render a dynamic human body rendering based on the reconstructed digital human body and the predicted human camera pose ;

[0059] Render a dynamic human head rendering based on the reconstructed digital head, the predicted head camera pose, and the predicted expression parameters ;

[0060] Based on the mask result of the semantic segmentation corresponding to the reconstructed digital head Segment the dynamic human head rendering to obtain a facial image ;

[0061] Fuse the dynamic human rendering and the facial image to obtain the human image with the focused face, the formula is as follows: , the formula is as follows:

[0062] .

[0063] Furthermore, the digital human reconstruction model is the SMPL model.

[0064] The beneficial effects of the present invention are as follows:

[0065] A digital human reconstruction method with a focused face proposed by the present invention not only focuses on the overall reconstruction of the human body, but also strengthens the high-precision reconstruction of the face, and can present a smooth and realistic human animation effect in subsequent rendering. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 is a flowchart of a method for reconstructing a digital human with a focused face according to the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0067] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.

[0068] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0069] Terms such as "first", "second", etc. are only used for descriptive distinction and cannot be construed as indicating or implying relative importance.

[0070] The following will combine the accompanying drawings to provide a detailed description of the specific embodiments of the present invention.

[0071] As Figure 1 shown, this embodiment provides a digital human reconstruction method focusing on the face of the present invention, including:

[0072] Obtain first information, where the first information includes an RGB image frame sequence of a dynamic human body;

[0073] Obtain second information and third information according to the first information. The second information includes the shape parameters, expression parameters, and head camera poses of the FLAME model corresponding to each predicted RGB image frame. The third information includes the shape parameters, pose parameters, and human camera poses of the digital human reconstruction model corresponding to each predicted RGB image frame;

[0074] Construct a head mesh corresponding to each RGB image frame based on the preset FLAME model according to the second information;

[0075] Uniformly sample the UV texture map of the head mesh corresponding to each RGB image frame to obtain corresponding sampling results;

[0076] Perform Gaussian field initialization on the corresponding head mesh according to the corresponding sampling results to obtain the initialization parameters of the Gaussian basis elements of the corresponding head mesh;

[0077] Enter the processing step of the current RGB image frame:

[0078] Perform face recognition on the current RGB image frame, and perform super-resolution processing and semantic segmentation on the recognized face area in sequence to obtain the current face RGB image frame;

[0079] Based on the initialization parameters and the expression parameters of the current RGB image frame, obtain new parameters of the Gaussian basis elements of the head mesh corresponding to the current RGB image frame based on a multi-layer perceptron;

[0080] Render a head rendering picture according to the new parameters of the Gaussian basis elements of the head mesh corresponding to the current RGB image frame and the head camera pose, calculate the loss between the head rendering picture and the current face RGB image frame, and optimize the parameters of the Gaussian basis elements of the head mesh corresponding to the current RGB image frame through a training strategy of minimizing the loss, and then end the processing step of the current RGB image frame;

[0081] Iteratively process other RGB image frames according to the processing step of the current RGB image frame until the parameters of the Gaussian basis elements of the head meshes corresponding to all the RGB image frames are optimized, and combine all the optimized Gaussian basis elements to obtain a reconstructed digital head;

[0082] Input the third information into a preset digital human body reconstruction model to output a reconstructed digital human body;

[0083] Render a human body image focusing on the face based on the reconstructed digital head and the reconstructed digital human body.

[0084] In some embodiments, obtaining the second information and the third information according to the first information includes:

[0085] Obtain a sequence of RGB image frames of a dynamic human body captured by a monocular camera;

[0086] Input the sequence of RGB image frames of the dynamic human body captured by the monocular camera into a pre-trained human head parameter prediction model to output the second information;

[0087] Input the sequence of RGB image frames of the dynamic human body captured by the monocular camera into a pre-trained human body parameter prediction model to output the third information.

[0088] In some embodiments, the initialization parameter G is expressed as:

[0089]

[0090] where μ represents texture coordinates, and a fixed point on the head mesh can be located through the texture coordinates μ , and this point is the position of the Gaussian basis element, represents the rotation coefficient, represents the scaling coefficient, α represents the opacity, h represents the spherical harmonic function, is initialized as a quaternion (1, 0, 0, 0), is initialized as the distance to the Gaussian basis element with the closest distance, and the calculation formula is as follows:

[0091]

[0092] where represents the center position of the Gaussian basis element with the closest distance to the point , and the opacity α of each Gaussian basis element is initialized to 0.1, and the spherical harmonic function h is initialized as a tensor of dimension, where is the number of Gaussian basis elements, the spherical harmonic function order d = 3, the value of the tensor at the position (N, 3, 0) is 0.5, and the rest are 0, and the dist function is used to calculate the Euclidean distance between two coordinate points.

[0093] In some embodiments, face recognition is performed on the current RGB image frame, and super-resolution processing and semantic segmentation are sequentially performed on the face region recognized by face recognition to obtain the current face RGB image frame, including:

[0094] Perform face recognition on the current RGB image frame based on a face recognition algorithm to obtain a face region;

[0095] Perform super-resolution processing on the face region to obtain a clear face picture;

[0096] Perform semantic segmentation on the clear face picture to obtain the current face RGB image frame.

[0097] Further, based on the initialization parameters and the expression parameters of the current RGB image frame, new parameters of the Gaussian basis elements of the head mesh corresponding to the current RGB image frame are obtained based on a multi-layer perceptron, including:

[0098] Encode the texture coordinates;

[0099] Input the encoded result and the expression parameters of the current RGB image frame into the multi-layer perceptron, and output the offset, rotation amount, and scaling amount of the Gaussian basis element position of the head mesh;

[0100] Add the offset to the position of the Gaussian basis element to obtain the new position of the Gaussian basis element of the head mesh corresponding to the current RGB image frame;

[0101] Multiply the rotation amount by the rotation coefficient to obtain the new rotation coefficient of the Gaussian basis element of the head mesh corresponding to the current RGB image frame;

[0102] Multiply the scaling amount by the scaling coefficient to obtain the new scaling coefficient of the Gaussian basis element of the head mesh corresponding to the current RGB image frame.

[0103] In some embodiments, the new parameters of the Gaussian basis elements of the head mesh corresponding to the current RGB image frame are represented as:

[0104]

[0105]

[0106] Among them, , , represent the offset, rotation amount, and scaling amount in sequence, D represents the multi-layer perceptron, θ i represents the expression parameters of the current RGB image frame, represents the encoded result, , , The new position, new rotation coefficient and new scaling coefficient of the Gaussian basis element of the head grid corresponding to the current RGB image frame are sequentially represented.

[0107] In some embodiments, the calculation steps of the loss between the head rendering picture and the current face RGB image frame are as follows:

[0108] The head rendering image is obtained by spherical harmonic rendering through the new position, new rotation coefficient, new scaling coefficient, and opacity of the Gaussian primitive corresponding to the head mesh of the current RGB image frame. ;

[0109] Render the head image The loss reduction optimization is performed with the current face RGB image frame, and the formula is as follows:

[0110]

[0111]

[0112] in Dynamic human body images captured by the camera, is the position of the face recognition frame, crop is the image cropping, super is the image super-resolution, is the result of the super-resolution processing, is the result of the semantic segmentation, is the total loss function, and loss is the loss.

[0113] In some embodiments, the total loss function for:

[0114]

[0115] where λ 1 , 2 is a constant, where the Huber loss for:

[0116]

[0117] in, Indicates that the camera takes an image. Represents a rendered image, ;

[0118] Its perceptual similarity loss for:

[0119]

[0120] in Represents a rendered image The feature map of the i-th layer in the VGG deep network, where represents the image captured by the camera The feature map of the i-th layer in the VGG deep network, and N is the number of feature layers for calculating the loss. represents the calculation of the L2 norm of, and then square it.

[0121] As Figure 1 shown, in some embodiments, a human body image with a focused face is rendered according to the reconstructed digital head and the reconstructed digital human body, including:

[0122] Rendering a dynamic human body rendering diagram according to the reconstructed digital human body and the predicted human body camera pose ;

[0123] Rendering a dynamic human body head rendering diagram according to the reconstructed digital head, the predicted head camera pose, and the predicted expression parameters ;

[0124] According to the mask result of the semantic segmentation corresponding to the reconstructed digital head Segmenting the dynamic human body head rendering diagram to obtain a face image ;

[0125] Fusing the dynamic human body rendering diagram and the face image to obtain the human body image with the focused face , and the formula is as follows:

[0126] .

[0127] In some embodiments, the digital human body reconstruction model is an SMPL model.

[0128] The advantages of the present invention compared with the prior art are as follows:

[0129] 1. On the basis of the traditional digital human body reconstruction model, the FLAME model is combined to construct a human body reconstruction method with high-precision face detail expression, which not only pays attention to the overall reconstruction of the human body, but also strengthens the high-precision reconstruction of the face. Based on the reconstructed human body model, real-time rendering of human body animation is realized, the dynamic effect of natural facial expression changes is generated, and the coordination between actions and expressions is ensured, presenting a smooth and realistic human body animation effect;

[0130] 2. During the training process, the loss function is designed by considering the color, normal map, depth map, joint point data and facial features of the pixels in the corresponding perspective image to optimize the human body reconstruction parameters, and at the same time, the 3D points reconstructed in the facial area are fused in real time;

[0131] 3. During the human body animation rendering stage, a monocular camera is used to capture the human body posture in real time, predict the 3D posture data of the human body in the picture and the camera parameters, so as to drive the digital human to change to the human body animation in the new posture. Specifically, according to the triangular patches corresponding to the human body mesh model in the current posture and the human body mesh model in the standard posture, the rotation difference and the scaling difference are calculated, and based on this, the final primitive parameter values in the Gaussian animation field are calculated. Rendering the Gaussian animation field in real time according to the predicted camera parameters can generate the human body animation in the current posture. For the face area in the image, real-time recognition and tracking are carried out, the face posture and the camera pose are extracted, and a face animation Gaussian field is generated, which is used to update the face area in the human body Gaussian animation field, so as to render the human body animation with rich facial expressions;

[0132] 4. It does not require training with multi-views of traditional methods. This method only uses monocular videos. A high-fidelity digital human is reconstructed to achieve high-definition face and full-body rendering;

[0133] 5. The rendering speed is fast. This method is based on the Gaussian splashing technology and has a faster rendering speed than the traditional Nerf-based method;

[0134] 6. The face rendering quality is high. Compared with the monocular video reconstruction method based on the SMPL-X model, this method can achieve higher accuracy and animation details by reconstructing the face segmentation separately.

[0135] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and retouches can still be made, and these improvements and retouches should also be regarded as the protection scope of the present invention.

Claims

1. A digital human body reconstruction method focusing on the face, characterized in that: include: Acquire first information, wherein the first information includes a sequence of RGB image frames of a dynamic human body; Acquire second information and third information according to the first information, wherein the second information includes shape parameters, expression parameters, and head camera pose of the FLAME model corresponding to each predicted RGB image frame, and the third information includes shape parameters, posture parameters, and human camera pose of the digital human body reconstruction model corresponding to each predicted RGB image frame; Constructing a head grid corresponding to each RGB image frame based on a preset FLAME model according to the second information; Uniformly sampling the UV texture map of the head mesh corresponding to each RGB image frame to obtain a corresponding sampling result; Performing Gaussian field initialization on the corresponding head grid according to the corresponding sampling result to obtain initialization parameters of the Gaussian basis element of the corresponding head grid; Enter the current RGB image frame processing steps: Perform face recognition on the current RGB image frame, perform super-resolution processing and semantic segmentation on the face area recognized by the face, and obtain the current face RGB image frame; According to the initialization parameters and the expression parameters of the current RGB image frame, new parameters of the Gaussian basis element of the head grid corresponding to the current RGB image frame are obtained based on a multi-layer perceptron; Obtain a head rendering image according to the new parameters of the Gaussian primitives of the head mesh corresponding to the current RGB image frame and the head camera pose rendering, calculate the loss between the head rendering image and the current face RGB image frame, optimize the parameters of the Gaussian primitives of the head mesh corresponding to the current RGB image frame by a training strategy that minimizes the loss, and then end the current RGB image frame processing step; Iteratively process other RGB image frames according to the current RGB image frame processing step until the parameters of the Gaussian primitives corresponding to the head grids of all the RGB image frames are optimized, and combine all the optimized Gaussian primitives to obtain a reconstructed digital head; Inputting the third information into a preset digital human body reconstruction model, and outputting a reconstructed digital human body; A human body image with a focused face is obtained by rendering according to the reconstructed digital head and the reconstructed digital human body.

2. The method for reconstructing a digital human body focusing on the face according to claim 1, characterized in that: Acquiring second information and third information according to the first information includes: Get the RGB image frame sequence of the dynamic human body captured by the monocular camera; Inputting the RGB image frame sequence of the dynamic human body captured by the monocular camera into a pre-trained human head parameter prediction model, and outputting second information; The RGB image frame sequence of the dynamic human body captured by the monocular camera is input into a pre-trained human body parameter prediction model to output third information.

3. The method for reconstructing a digital human body focusing on the face according to claim 1, characterized in that: The initialization parameter G is expressed as: , Wherein, μ represents the texture coordinate, and the texture coordinate μ can be used to locate a fixed point on the head mesh. , the point is the position of the Gaussian basis element, represents the rotation coefficient, represents the scaling factor, α represents the opacity, h represents the spherical harmonic function, Initialized to quaternion (1,0,0,0), Initialized to the distance between the nearest Gaussian primitive, the calculation formula is as follows: , in, Indicates the departure point The center position of the nearest Gaussian primitive, the opacity α of each Gaussian primitive is initialized to 0.1, and the spherical harmonic function h is initialized to A tensor of dimension N, where N is the number of Gaussian basis elements, the order of spherical harmonics is d=3, the value of the tensor at (N,3,0) is 0.5, and the rest are 0. The dist function is used to calculate the Euclidean distance between two coordinate points.

4. A face-focused digital human reconstruction method according to claim 1 or 3, characterized in that: Perform face recognition on the current RGB image frame, perform super-resolution processing and semantic segmentation on the face area identified by face recognition, and obtain the current face RGB image frame, including: Perform face recognition on the current RGB image frame based on the face recognition algorithm to obtain the face area; Performing super-resolution processing on the face area to obtain a clear face image; Perform semantic segmentation on the clear face image to obtain a current face RGB image frame.

5. The method for reconstructing a digital human body focused on the face according to claim 3, characterized in that: According to the initialization parameters and the expression parameters of the current RGB image frame, new parameters of the Gaussian basis element of the head grid corresponding to the current RGB image frame are obtained based on a multi-layer perceptron, including: encoding the texture coordinates; Input the encoding result and the expression parameters of the current RGB image frame into a multi-layer perceptron, and output the offset, rotation and scaling of the Gaussian primitive position of the head grid; Adding the offset to the position of the Gaussian primitive to obtain a new position of the Gaussian primitive corresponding to the head grid of the current RGB image frame; Multiplying the rotation amount by the rotation coefficient to obtain a new rotation coefficient of the Gaussian basis element of the head grid corresponding to the current RGB image frame; The scaling amount is multiplied by the scaling factor to obtain a new scaling factor of the Gaussian basis element of the head grid corresponding to the current RGB image frame.

6. The method for reconstructing a digital human body focusing on the face according to claim 5, characterized in that: The new parameters of the Gaussian primitives of the head grid corresponding to the current RGB image frame are expressed as: , , in, , , represents the offset, rotation and scaling respectively, D represents the multilayer perceptron, θ i represents the expression parameters of the current RGB image frame, represents the result of the encoding, , , The new position, new rotation coefficient and new scaling coefficient of the Gaussian basis element of the head grid corresponding to the current RGB image frame are sequentially represented.

7. The method for reconstructing a digital human body focused on the face according to claim 6, characterized in that: The calculation steps of the loss between the head rendering picture and the current face RGB image frame are as follows: The head rendering image is obtained by spherical harmonic rendering through the new position, new rotation coefficient, new scaling coefficient, and opacity of the Gaussian primitive corresponding to the head mesh of the current RGB image frame. ; Render the head image The loss reduction optimization is performed with the current face RGB image frame, and the formula is as follows: , , in Dynamic human body images captured by the camera, is the position of the face recognition frame, crop is the image cropping, super is the image super-resolution, is the result of the super-resolution processing, is the result of the semantic segmentation, is the total loss function, and loss is the loss.

8. The method for reconstructing a digital human body focused on the face according to claim 7, characterized in that: The total loss function for: , Where λ1 and λ2 are constants, and Huber loss for: , in, Indicates that the camera takes an image. Represents a rendered image, ; Its perceptual similarity loss for: , in Represents a rendered image The feature map of layer i in the VGG deep network is Indicates that the camera captures an image In the VGG deep network, the i-th layer feature map, N is the number of feature layers for calculating the loss, Representation calculation The L2 norm of is then squared.

9. The method for reconstructing a digital human body focused on the face according to claim 8, characterized in that: Rendering a human body image with a focused face according to the reconstructed digital head and the reconstructed digital human body includes: A dynamic human body rendering image is obtained according to the reconstructed digital human body and the predicted human body camera pose rendering ; A dynamic human head rendering image is obtained by rendering according to the reconstructed digital head, the predicted head camera pose and the predicted expression parameters. ; The mask result of the semantic segmentation corresponding to the reconstructed digital head Rendering of the dynamic human head Segmentation to get the face image ; The dynamic human body rendering and the facial image Fusion is performed to obtain the human body image of the focused face , the formula is as follows: 。 10. The method for reconstructing a digital human body focused on the face according to claim 1, characterized in that: The digital human body reconstruction model is an SMPL model.

Citation Information

Patent Citations

  • Living body recognition method and device, electronic equipment and storage medium

    CN115705758A

  • Drivable human body modeling and rendering method and system based on Gaussian point cloud

    CN118196285A