Three-dimensional head portrait model construction method, system and device and storage medium
Through multi-view face information extraction and anchor point adjustment technology, the problem that the existing technology is difficult to capture high-frequency details of the face is solved, and efficient three-dimensional avatar model is realized to generate animated virtual three-dimensional avatar with high visual detail quality.
Patent Information
- Application Number
- CN202510179154.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-02-18
AI Technical Summary
The prior art is difficult to efficiently capture and restore complex expressions and posture changes of the face, especially in the animation process, it is difficult to accurately capture high-frequency dynamic details of the face, resulting in unnatural expression changes and distorted animation effects.
By obtaining the multi-view face information of the target character, extracting shape parameters, posture parameters and expression parameters, adjusting the anchor point information on the face template based on these parameters, activate the corresponding three-dimensional Gaussian points, accurately capturing the high-frequency details of the face, and generating animated virtual three-dimensional avatar with high visual detail quality.
It achieves highly restored complex facial expressions and posture changes, accurately capture high-frequency details such as subtle muscle movements, and generates animated virtual three-dimensional avatars with high visual detail quality.
Smart Images

Figure CN120236005A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular, to a method, system, device and storage medium for constructing a three-dimensional avatar model. Background Art
[0002] Creating personalized and animatable three-dimensional avatars has attracted increasing attention in applications such as augmented reality / virtual reality (AR / VR), immersive telepresence, film production, and gaming.
[0003] In the early stage, three-dimensional deformable models (3DMMs) were mainly used. It explored the diversity of specific identities and expressions in the low-dimensional space through principal component analysis (PCA), and had a certain effect in fitting shapes and transforming the expressions of given individuals. However, due to the linear interpolation characteristics of PCA and the fixed topological structure, it was limited in restoring details (such as wrinkles) and modeling facial accessories (such as complex hairstyles or glasses).
[0004] In recent years, the neural radiance field (NeRF) technology has emerged and promoted the development of 3D avatar generation technology. It uses a deep neural network to learn and capture the internal patterns of three-dimensional shapes, and has advantages in depicting facial features and dynamic information in detail.
[0005] 3D Gaussian Splatting (3DGS) uses anisotropic and discrete three-dimensional Gaussian basis elements, which is more efficient than the neural radiance field (NeRF) in reconstructing head avatars. It directly uses Gaussian points for scene representation. When animating three-dimensional avatars, it is difficult to accurately capture the high-frequency dynamic details of the face, and the transition of facial expressions between different frames is not natural, and the animation effect is prone to distortion. Summary of the Invention
[0006] The main purpose of the embodiments of the present disclosure is to propose a method, system, device and medium for constructing a three-dimensional avatar model, which can highly restore the complex expressions and pose changes of the face, accurately capture high-frequency details such as subtle muscle movements, and generate an animatable virtual three-dimensional avatar with high visual detail quality.
[0007] To achieve the above object, on the one hand, an embodiment of the present application proposes a method for constructing a three-dimensional avatar model, including the following steps:
[0008] Obtain multi-view face information of a target person;
[0009] Extract features from the multi-view face information to obtain shape parameters, first pose parameters, and first expression parameters;
[0010] Determine a plurality of anchor point information on a face template according to the shape parameters, where each anchor point information is used to manage the attribute data of three-dimensional Gaussian points associated with a plurality of associated face templates;
[0011] Adjust the anchor point information according to the first pose parameter and the first expression parameter to obtain the first attribute data of the three-dimensional Gaussian point corresponding to the anchor point information, wherein the first attribute data of the three-dimensional Gaussian point is used to obtain the second attribute data of the three-dimensional Gaussian point;
[0012] Obtain a target three-dimensional head portrait according to the second attribute data of the three-dimensional Gaussian point and the face template.
[0013] In some embodiments, the determining a plurality of anchor point information on the face template according to the shape parameter includes the following steps:
[0014] Constrain the face template according to the shape parameter to obtain an initial three-dimensional head portrait;
[0015] Sample the initial three-dimensional head portrait to obtain a plurality of the anchor point information.
[0016] In some embodiments, the adjusting the anchor point information according to the first pose parameter and the first expression parameter to obtain the first attribute data of the three-dimensional Gaussian point corresponding to the anchor point information includes the following steps:
[0017] Input the first pose parameter into a pose prediction network to obtain a predicted pose basis, and input the first expression parameter into an expression prediction network to obtain a predicted expression basis;
[0018] Adjust the anchor point information according to the predicted pose basis and the predicted expression basis;
[0019] Decode the adjusted anchor point information through a multi-layer perceptron to obtain the first attribute data of the three-dimensional Gaussian point corresponding to the anchor point information, wherein the attribute data includes opacity information, rotation information, scale information, and color information.
[0020] In some embodiments, the adjusting the anchor point information according to the predicted pose basis and the predicted expression basis includes the following steps:
[0021] Predict a linear skinning weight according to the unadjusted anchor point information, the first pose parameter, and the first expression parameter;
[0022] Adjust the anchor point information according to the predicted pose basis, the predicted expression basis, and the linear skinning weight.
[0023] In some embodiments, after adjusting the anchor point information according to the first pose parameter and the first expression parameter to obtain the first attribute data of the three-dimensional Gaussian point corresponding to the anchor point information, the following steps are further included:
[0024] Based on the attribute data of the three-dimensional Gaussian points and the face template, a first three-dimensional avatar is obtained;
[0025] Through the attribute adapter, based on the adjusted anchor point information and the first three-dimensional avatar, second attribute data of the adjusted three-dimensional Gaussian points is obtained.
[0026] In some embodiments, the obtaining of the second attribute data of the adjusted three-dimensional Gaussian points through the attribute adapter based on the first three-dimensional avatar includes the following steps:
[0027] Feature extraction is performed on the first three-dimensional avatar to obtain second pose parameters and second expression parameters;
[0028] Through the first module of the attribute adapter, based on the adjusted anchor point information, the second pose parameters and the second expression parameters, second attribute data of the initial three-dimensional Gaussian points is obtained;
[0029] Through the second module of the attribute adapter, based on the adjusted anchor point information, the second pose parameters and the second expression parameters, an offset of the second attribute data is obtained;
[0030] Based on the offset and the second attribute data of the initial three-dimensional Gaussian points, second attribute data of the adjusted three-dimensional Gaussian points is obtained.
[0031] In some embodiments, the obtaining of the target three-dimensional avatar based on the second attribute data of the three-dimensional Gaussian points and the face template includes the following steps:
[0032] Based on the second attribute data of the three-dimensional Gaussian points and the face template, an initialized three-dimensional avatar model is obtained;
[0033] Based on the projection of the three-dimensional avatar model under the target perspective onto a two-dimensional plane, a predicted rendering is obtained;
[0034] Based on the predicted rendering and the real rendering under the target perspective, a loss function of the three-dimensional avatar model is established;
[0035] By minimizing the loss function, the target three-dimensional avatar is obtained.
[0036] On the other hand, an embodiment of the present invention proposes a three-dimensional avatar model construction system, including:
[0037] A first module for obtaining multi-perspective face information of a target person;
[0038] A second module for performing feature extraction on the multi-perspective face information to obtain shape parameters, first pose parameters and first expression parameters;
[0039] A third module, configured to determine a plurality of anchor point information on a face template according to the shape parameters, wherein each piece of the anchor point information is used to manage the attribute data of three-dimensional Gaussian points on a plurality of associated face templates;
[0040] A fourth module, configured to adjust the anchor point information according to the first pose parameter and the first expression parameter to obtain first attribute data of three-dimensional Gaussian points corresponding to the anchor point information, wherein the first attribute data of the three-dimensional Gaussian points is used to obtain second attribute data of the three-dimensional Gaussian points;
[0041] A fifth module, configured to obtain a target three-dimensional head portrait according to the second attribute data of the three-dimensional Gaussian points and the face template.
[0042] On the other hand, an embodiment of the present invention provides an electronic device, including:
[0043] At least one processor;
[0044] At least one memory, configured to store at least one program;
[0045] When the at least one program is executed by the at least one processor, the at least one processor implements the three-dimensional head portrait model construction method as described in the previous embodiment.
[0046] On the other hand, an embodiment of the present invention further provides a computer-readable storage medium, where the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause a computer to execute the three-dimensional head portrait model construction method as described in the previous embodiment.
[0047] At least one of the above technical solutions of the present invention has at least the following advantages or beneficial effects: Based on anchor point-guided three-dimensional Gaussian modeling, this application first responds to changes in expression and pose through anchor points, and then activates corresponding Gaussian points based on the precise positioning and constraints of the anchor points, which can highly restore complex facial expression and pose changes, accurately capture high-frequency details such as subtle muscle movements, and generate an animatable virtual three-dimensional head portrait with high visual detail quality. Description of the Drawings
[0048] Figure 1 is a flowchart of the three-dimensional head portrait model construction method provided by an embodiment of the present application;
[0049] Figure 2 is a schematic input diagram of the three-dimensional head portrait model provided by an embodiment of the present application;
[0050] Figure 3 is a schematic output diagram of the three-dimensional head portrait model provided by an embodiment of the present application;
[0051] Figure 4 It is a schematic diagram of the three-dimensional avatar model training process provided by an embodiment of the present application;
[0052] Figure 5 It is a schematic diagram of the face template structure provided by an embodiment of the present application;
[0053] Figure 6 It is a schematic diagram of the EPAA operation provided by an embodiment of the present application;
[0054] Figure 7 It is a qualitative experimental effect diagram based on the NeRSemble dataset provided by an embodiment of the present application;
[0055] Figure 8 It is a qualitative experimental effect diagram based on the INSTA dataset provided by an embodiment of the present application;
[0056] Figure 9 It is a qualitative ablation experimental effect diagram based on the NeRSemble dataset provided by an embodiment of the present application;
[0057] Figure 10 It is a disabled EPAA ablation experimental effect diagram provided by an embodiment of the present application;
[0058] Figure 11 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0059] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0060] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. The terms "first", "second", etc. in the specification, claims and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence.
[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application, and are not intended to limit the present application.
[0062] First, several nouns involved in the present application are analyzed:
[0063] Three-dimensional Gaussian Splatting (3DGS) is a technique for real-time radiance field rendering that can generate high-quality images in the Novel View Synthesis (NVS) task. This method represents the scene by optimizing a three-dimensional Gaussian distribution and uses anisotropic splatting technology for efficient rendering.
[0064] The FLAME model (Faces Learned with an Articulated Model and Expressions) is a generative model for 3D face modeling that combines parametric modeling of human shape, joint movement, and facial expressions, enabling the generation of highly realistic and controllable 3D faces and head animations. The model integrates articulated chin, neck, and eye models and uses pose correction and global expression blend shapes to accurately capture facial geometry, expressions, and dynamic changes.
[0065] Creating personalized and animatable 3D avatars has attracted increasing attention in applications such as augmented reality / virtual reality (AR / VR), immersive telepresence, film production, and gaming.
[0066] In the early stage, 3D Morphable Models (3DMMs) explored the diversity of specific identities and expressions in a low-dimensional space through Principal Component Analysis (PCA), showing significant effects in fitting shapes and transforming the expressions of given individuals. However, due to the linear interpolation property of PCA and the fixed topology, simply using grid-based 3DMMs limits their ability to recover details (such as wrinkles) and model facial accessories (such as complex hairstyles or glasses). Recent research has successfully reconstructed 3D avatars by combining Neural Radiance Fields (NeRF) technology, leveraging its excellent capabilities in novel view synthesis. Hybrid use with 3DMMs has also achieved remarkable results, performing well in maintaining details and achieving high-quality rendering. Despite numerous efforts in various NeRF variants (such as improved methods of multi-resolution hash encoding like InstantNGP) to accelerate the training and rendering processes, these methods still have trade-offs because the volume rendering mechanism still requires a large amount of training and rendering time.
[0067] 3D Gaussian Sputtering (3DGS) uses anisotropic and discrete three-dimensional Gaussian basis elements, which is more efficient than Neural Radiance Fields (NeRF) when reconstructing head avatars. It directly uses Gaussian points for scene representation. When animating three-dimensional avatars, it is difficult to accurately capture the high-frequency dynamic details of the face, and the transition of facial expression changes between different frames is not natural, and the animation effect is prone to distortion.
[0068] Based on this, the embodiments of the present disclosure provide a method, system, device and medium for constructing a three-dimensional avatar model, which can highly restore the complex expressions and pose changes of the face, accurately capture high-frequency details such as subtle muscle movements, and generate an animatable virtual three-dimensional avatar with high visual detail quality.
[0069] Refer to Figure 1 as shown Figure 1 is an optional flowchart of the method for constructing a three-dimensional avatar model provided by some embodiments of the present application. A method for constructing a three-dimensional avatar model according to an embodiment of the present invention includes but is not limited to steps S100 to S500.
[0070] Step S100, obtaining multi-view face information of a target person;
[0071] Step S200, extracting features from the multi-view face information to obtain shape parameters, first pose parameters and first expression parameters;
[0072] Step S300, determining a plurality of anchor point information on a face template according to the shape parameters, wherein each anchor point information is used to manage the attribute data of a plurality of three-dimensional Gaussian points associated on the face template;
[0073] Step S400, adjusting the anchor point information according to the first pose parameters and the first expression parameters to obtain first attribute data of the three-dimensional Gaussian points corresponding to the anchor point information, wherein the first attribute data of the three-dimensional Gaussian points is used to obtain second attribute data of the three-dimensional Gaussian points;
[0074] Step S500, obtaining a target three-dimensional avatar according to the second attribute data of the three-dimensional Gaussian points and the face template.
[0075] In step S100 of some embodiments, the multi-view face information, as the basic data for constructing the three-dimensional avatar model, plays a key role in enhancing the realism and interaction experience of the virtual image. Refer to Figure 2 , Figure 2Displays multi - perspective images of the target person with different expressions as the input for model training. The multi - perspective face information can be collected with the help of multiple cameras, which can capture images from different perspectives such as the front, side, etc. Usually, multiple high - definition cameras are installed at positions such as directly in front, on the left and right sides, and obliquely on the sides to photograph the target person from multiple angles. Obtain the two - dimensional image information of the face, providing richer data support for subsequent 3D reconstruction. During the shooting process, the target person can be guided to show various expressions, including common expressions such as smiling, frowning, and being surprised, so as to obtain diverse expression data. These data help the model learn features such as the geometric shape changes, texture details, and lighting effects of the face under different perspectives and expressions. The 3D avatar model constructed based on these data can provide users with a more realistic and vivid virtual image, such as Figure 3 shown.
[0076] In step S200 of some embodiments, referring to Figure 4 , the facial key points in the image can be extracted by using 2D face detection algorithms (such as MTCNN, OpenPose). Through an optimization process, the 2D facial key point information is matched with the 3D facial structure in the FLAME model. Through a fitting process, for example, using the non - linear least - squares method for fitting, an objective function containing 2D facial key points, FLAME model parameters, and related constraint conditions is constructed. By continuously adjusting the parameters to minimize the value of the objective function, integrating multi - perspective face information, ensuring the parameter consistency under different perspectives, so that the model can present reasonable facial features at each angle. Finally, the shape, expression, and pose parameters of the FLAME model are determined, which can be represented as shape parameters, first - pose parameters, and first - expression parameters. These parameters can be directly applied to the FLAME model to generate a 3D face model that matches the input multi - perspective face information. In subsequent steps of 3D avatar model training, such as determining the anchor point information on the face template and adjusting the attribute data of 3D Gaussian points, these parameters will play an important role in ensuring that the final generated target 3D avatar can accurately reflect the facial features, posture, and expression of the target person.
[0077] In step S300 of some embodiments, such as Figure 4For the reconstructed part of the head geometry, the FLAME model is selected as the face template. The FLAME model can be regarded as essentially composed of multiple small facial meshes. These facial meshes each carry specific geometric information, and are interconnected and interact with each other to jointly form a three-dimensional model that can accurately describe the complex structure and dynamic changes of the human face. Each mesh corresponds to a specific area on the human face. By adjusting and controlling the attributes such as the shape, position, and texture of these meshes, the simulation of different facial expressions, poses, and individual characteristics of the human face can be achieved. Constraining the FLAME model according to the shape parameters, simulating the geometric shape of the target person's head, and initializing the scene voxelization by sampling the FLAME mesh to anchor the geometric shape of the neutral head of the subject, multiple anchor points can be obtained, and the anchor point information can describe the characteristic information such as the position of the anchor points.
[0078] In this embodiment, please refer to Figure 5 , the template of the FLAME model in the standard space adopted is the "mouth open" state, rather than the "mouth closed" state, and each mesh of the model corresponds to multiple anchor points and Gaussian points.
[0079] In some embodiments, step S300 may include but is not limited to steps S310 to S330:
[0080] Step S310, according to the shape parameters, constrain the face template to obtain an initial three-dimensional head portrait;
[0081] Step S320, sample the initial three-dimensional head portrait to obtain multiple anchor point information.
[0082] In steps S310 to S320 of some embodiments, in the initial stage, according to the shape parameters obtained in step S200, the attributes of each facial mesh in the FLAME model are modified, so that the model gradually approaches the real face shape of the target person, and thus a predefined neutral face template is obtained. The predefined neutral face template can be represented as an initial three-dimensional head portrait. The predefined neutral face template is a three-dimensional head portrait without expression. Then, Poisson disk sampling is performed on the predefined neutral face template to generate uniformly distributed sampling points as the initial anchor point positions. These anchor points are close to the head surface, forming a support structure similar to a bone framework, ensuring that the model can quickly converge to an accurate geometric shape in the early stage of model training. In addition, these anchor points not only provide the initial layout for Gaussian points, but also enable the model to more effectively capture the detailed features and complex expression changes of the human face by defining clear geometric constraints, laying a solid foundation for subsequent expression and pose modeling.
[0083] In step S400 of some embodiments, as Figure 4For the animated part of the virtual human head, the predefined neutral face template is deformed to generate specific poses and expressions. Different from the FLAME method based on template meshes, an anchor - based approach is adopted, which maps the anchors in the standard space to the deformed space that combines poses and expressions. In step S300, a face template in the "mouth open" state in the standard space is used. This processing helps to more effectively capture the deformation basis in the mouth area, especially when dealing with dynamics such as lip opening and closing. In this way, the problem that the MLP model cannot accurately learn the change due to the drastic change of the LBS weights between the lips during continuous deformation can be avoided. Specifically, the open - mouth pose provides a larger range of shape changes, which helps to improve the expression ability of facial dynamics and avoid unnatural deformations caused by overly extreme weight distributions.
[0084] By adopting an anchor - based approach to learn the underlying deformation logic of the face parameter template, the model can more flexibly capture the details of the face under the changes of dynamic expressions and poses. This mapping process is similar to the expression method of FLAME. By integrating factors such as poses and expressions into the deformation basis, natural and realistic dynamic expression effects can be generated without relying on the mesh structure. The expression of this process is shown in Equation (1).
[0085]
[0086] Among them, x c represents the coordinates of the anchor in the standard space, x d represents the coordinates of the anchor in the deformed space, θ represents the pose parameter, ψ represents the expression parameter, is the pose basis to be learned, ε is the expression basis to be learned, represents the mixing weight; B P represents the animation offset obtained by linearly combining the pose basis and the driving signal θ corresponding to the pose parameter, B E represents the animation offset obtained by linearly combining the expression basis ε and the driving signal θ corresponding to the expression parameter.
[0087] In some embodiments, step S400 may include but is not limited to steps S410 to S430:
[0088] Step S410, input the first pose parameter into the pose prediction network to obtain the predicted pose basis, and input the first expression parameter into the expression prediction network to obtain the predicted expression basis;
[0089] Step S420, adjust the anchor information according to the predicted pose basis and the predicted expression basis;
[0090] Step S430: Decode the adjusted anchor information through a multi-layer perceptron to obtain the first attribute data of the three-dimensional Gaussian points corresponding to the anchor information, where the attribute data includes opacity information, rotation information, scale information, and color information.
[0091] In steps S410 to S430 in some embodiments, to deform the predefined neutral face template in the standard space and make the predefined neutral face template obtain the target pose and expression, it is necessary to ensure that the anchor points can accurately respond to the changes in pose and expression. Spread the fixed number of template vertex deformation bases to the anchor points that need to grow, so as to learn high-quality pose bases, expression bases, and skinning weights. This diffusion process endows the anchor points with flexible deformation ability, making the model more natural in showing complex dynamic features during the generation process. To obtain these deformation bases (pose bases and expression bases), a tri-plane designed to extract the regional features of the anchor points The features obtained through this tri-plane Then by the prediction network Output the deformation bases of each anchor point. In particular, encode the first pose parameter and the first expression parameter and send them to independent branches to predict the pose base and the expression base respectively. This branch processing method not only enhances the generalization ability of the model to unseen poses and expressions, but also ensures accurate reproduction under different expression and pose changes.
[0092] The anchor points have simplified properties similar to those of the mesh vertices because they mainly focus on the changes in position. During the deformation process, the anchor points act as the skeleton of the geometric structure. After they move to the predetermined positions, the subsequent detailed carving is handed over to the lightweight multi-layer perceptron (MLP). The MLP activates the corresponding Gaussian points according to the scene requirements to capture and depict fine facial features. This division of labor not only significantly reduces the computational burden, but also ensures the efficient operation of each module, each performing its own duties, and finally achieving a detailed and realistic dynamic facial performance.
[0093] The k neural Gaussian points managed by the anchor points have multiple attributes (the first attribute data of the three-dimensional Gaussian points corresponding to the anchor information), including position μ, opacity α, rotation r and scale S related to the covariance, and color c. The activation of these Gaussian points is related to the position of their anchor points. Only when the anchor points appear within the input frustum, the anchor points will adaptively activate their linked k Gaussian points. This process is only carried out during the rendering stage. Mathematically, the central position μ of the Gaussian points k Can be expressed as shown in Equation (2):
[0094] μ k = x + offset k ·s anchor , (2)
[0095] where x is the position of the anchor point, and offset k is a learnable offset, and offset k ∈R k×3 , and S anchor is the scaling parameter related to the anchor point. Other attributes depend on the optimizable abstract features of the anchor point itself and, combined with the perspective information, respectively decode the corresponding attribute values through multiple small multi-layer perceptrons independently, so that the attributes of the Gaussian points can be dynamically adjusted according to the perspective information, thereby rendering the facial details more precisely. The neural Gaussian attributes are decoded from the features f anchor of the anchor points they manage and the perspective information d, as shown in Equations (3) to (5).
[0096]
[0097] During the process of carving details, the head of the target person is embedded in the voxel grid because the anchor points are densified based on the voxel grid. This process identifies important regions through spatial quantization and gradient evaluation, and adds new anchor points in areas with insufficient initial coverage. To control the expansion, multi-resolution voxel grids and random selection are used to ensure balanced and efficient placement of anchor points. For anchor points with lower contributions, that is, those whose generated Gaussian points' opacity cannot reach the ideal threshold, they will be removed.
[0098] In some embodiments, step S420 may include but is not limited to steps S421 to S422:
[0099] Step S421, predicting the linear skinning weights based on the unadjusted anchor point information, the first pose parameter, and the first expression parameter;
[0100] Step S422, adjusting the anchor point information according to the predicted pose basis, the predicted expression basis, and the linear skinning weights.
[0101] In steps S421 to S422 in some embodiments, first, feature extraction is performed on the unadjusted anchor point information, and the regional features of the anchor points can be extracted through tri-planes to obtain the regional features of the anchor points Then, combined with the first pose parameter and the first expression parameter, the predicted network outputs the linear skinning (LBS) weights, and the specific expression is as shown in Equation (6).
[0102]
[0103] Among them, through the regional features of the anchor points the first pose parameter and the first expression parameter are input into the prediction network to obtain the predicted pose basis Predictive expression basis ε and linear skinning weights
[0104] The prediction network is a neural network that integrates and processes these input data, and predicts pose bases through different branches Expression basis ε and linear skinning weights such as pose parameter θ, and combines the regional features of the anchor points Through the internal calculations and learning of the network, the predicted output pose bases are obtained The expression prediction network, in a similar manner, based on the expression coefficient ψ and the regional features of the anchor points obtains the predicted expression basis ε. The outputs of these two sub-networks, together with other information, participate in the prediction network's calculation of the linear skinning weights for calculation.
[0105] Using the obtained predicted pose bases Predictive expression basis ε and linear skinning weights By changing the anchor information through Equation (1), it can be expressed as changing the position information of the anchor points, realizing the position coordinate change of the anchor points from the standard space to the deformed space. After the anchor points move to the predetermined positions, the corresponding Gaussian points are activated by a multi-layer perceptron (MLP) according to the scene requirements, which are used to capture and depict fine facial features.
[0106] Combining the first pose parameter and the first expression parameter, after blendshape and linear skinning, the adaptability of the avatar model in details and the expression performance can be enhanced, and finally a virtual human head representation with accurate geometric structure and animatable is obtained.
[0107] By combining the driving signals, the predicted deformation bases, and the linear skinning (LBS) operations, the model can fully present the basic movements of the head. However, due to its linear interpolation characteristics, the linear skinning (LBS) is difficult to accurately reproduce details such as wrinkle changes and mouth opening and closing under extreme expressions and poses. Therefore, it is difficult to reproduce some exaggerated dynamic behaviors. Therefore, this problem is solved through the following steps S610 to S620, and the ability to capture high-frequency features is improved.
[0108] In some embodiments, step S400 may further include but is not limited to steps S610 to S620:
[0109] Step S610, obtaining the first three-dimensional head based on the first attribute data of the three-dimensional Gaussian points and the face template;
[0110] Step S620, through the attribute adapter, obtaining the second attribute data of the three-dimensional Gaussian points according to the adjusted anchor information and the first three-dimensional head.
[0111] In steps S610 to S620 in some embodiments, such as Figure 4 In the behavior perception optimization part, the first attribute data of the three-dimensional Gaussian points is combined with the face template, and by adjusting the template (such as geometric deformation, expression change, etc.), the first three-dimensional head is obtained. The generated first three-dimensional head is a preliminary static three-dimensional facial model, which already contains the basic facial geometry, pose and expression information of the target person.
[0112] Through this combination, the obtained first three-dimensional head is a basic version, which can reflect the facial geometric features of the target person, but may lack high-frequency details in dynamic changes. Therefore, an Expression and Pose-dependent Attribute Adapter (EPAA) is designed to make up for the deficiencies of LBS in dealing with complex dynamics. The attribute adapter dynamically fine-tunes the position and shape of the Gaussian points by injecting driving parameters into the attributes of each Gaussian point, and obtains the second attribute data of the three-dimensional Gaussian points. This method can take into account the offset effects of pose and expression on the facial structure.
[0113] In some embodiments, step S620 may include but is not limited to steps S621 to S624:
[0114] Step S621, extracting features from the first three-dimensional head to obtain the second pose parameter and the second expression parameter;
[0115] Step S622, through the first module of the attribute adapter, according to the adjusted anchor point information, the second pose parameter and the second expression parameter, obtain the second attribute data of the initial three-dimensional Gaussian points;
[0116] Step S623, through the second module of the attribute adapter, according to the adjusted anchor point information, the second pose parameter and the second expression parameter, obtain the offset of the second attribute data;
[0117] Step S624, according to the offset and the second attribute data of the initial three-dimensional Gaussian points, obtain the second attribute data of the three-dimensional Gaussian points.
[0118] In steps S621 to S625 of some embodiments, the adjusted anchor information is obtained based on step S400. The adjusted anchor information will be used as the input for subsequent steps. The first 3D avatar obtained in step S610 includes the geometric shape, pose, and expression changes of the target person's face. Feature extraction is performed on the first 3D avatar to obtain the second pose parameters and second expression parameters by detecting features such as facial key points, angles, and textures. The second pose parameters and second expression parameters are used as the input for the attribute adapter in subsequent steps to further adjust the attributes and details of the avatar.
[0119] Use (The first module of the attribute adapter) injects the dynamic parameters (the second pose parameters and second expression parameters) into the abstract features of the anchor points, which are then decoded by the image multi-layer perceptron (MLP) specialized for each Gaussian attribute. This process is shown in Equations (8) and (9), as Figure 6 shown, by injecting the current expression and pose parameters, the anchor point feature f is enhanced to f′, thereby generating Gaussian attributes (the second attribute data of the initial 3D Gaussian points) that are sensitive to head movement.
[0120] To further optimize these attributes, apply (The second module of the attribute adapter) to combine the given signal and the anchor point position to generate the spatial residuals of the Gaussian attributes (the offsets of the second attribute data): the position δμ, rotation δr, and scale δs. The above process can be expressed as shown in Equation (7).
[0121]
[0122] where γ represents the position information encoding of the anchor point coordinates, embedded as a high-dimensional sequence, θ i represents the pose parameters (the second pose parameters) of the current model, and ψ i represents the expression parameters (the second expression parameters) of the current model.
[0123] The attributes of the Gaussian points are affected by head movement. By introducing non-linear elements, the linear interpolation mode of LBS is broken, thus avoiding unnatural expressions. With the transformation of the anchor points and the enhancement of the abstract features, the Gaussian neural point attribute positions μ d , rotation r d , scale s d , opacity α d and color c d , the above Gaussian neural point attributes are the second attribute data and can be obtained through Formulas (9) and (10).
[0124]
[0125] Among them, in formula (10) is represented as MLP and is used to decode the enhanced anchor features. In formula (11), offset k is a learnable offset, offset k ∈R k×3 , s anchor is a scaling parameter related to the anchor, and x d is the corresponding anchor position.
[0126] Finally, these two network structures work together to input the additional offset into the refined attributes, that is, add the offset to the second attribute data of the initial three-dimensional Gaussian points, so as to generate the finally deformed Gaussian primitive, which can be expressed as the second attribute data of the three-dimensional Gaussian points, as shown in formula (11).
[0127] G d ={μ d +δμ,r d +δr,s d +δs,α d , c d}, (11)
[0128] In some embodiments, step S500 may further include but is not limited to steps S510 to S540:
[0129] Step S510, obtaining an initialized three-dimensional avatar model according to the second attribute data of the three-dimensional Gaussian points and the face template;
[0130] Step S520, projecting the three-dimensional avatar model under the target perspective onto a two-dimensional plane to obtain a predicted rendering;
[0131] Step S530, establishing a loss function of the three-dimensional avatar model according to the predicted rendering and the true rendering under the target perspective;
[0132] Step S540, obtaining the target three-dimensional avatar by minimizing the loss function.
[0133] In steps S510 to S540 in some embodiments, with reference to Figure 4For the loss and rendering part, the second attribute data of the 3D Gaussian points is obtained from step S624. Combining with the face template, an initial 3D avatar model is obtained, which is a preliminarily formed 3D avatar model. In order to further evaluate and optimize this 3D avatar model to make it closer to the real face image, it is necessary to compare and analyze it with the real rendering image. The initial 3D avatar needs to be projected onto a 2D plane to obtain the predicted rendering image. To ensure the accuracy and effectiveness of the comparison, the predicted rendering image and the real rendering image must be consistent. They need to have the same perspective, that is, present the avatar from the same viewing angle, and the poses and expressions shown by both must also be the same. Meeting these conditions can provide a basis for establishing the loss function of the 3D avatar model in the follow-up. The loss function established in this way can more accurately measure the gap between the model and the real situation, and then guide the optimization and improvement of the 3D avatar model to make it more vividly restore the features and appearance of the real face.
[0134] The specific projection process is as follows:
[0135] The 3D Gaussian represents the elements in the scene, which is defined by the center point μ ∈ R 3 and the covariance matrix ∑. For any position p in the 3D scene, each discrete Gaussian primitive is represented by the following formula (12).
[0136]
[0137] To ensure the stable progress of the optimization process, the covariance matrix must be kept positive semi - definite. Therefore, the matrix combines the scale and rotation transformations, and its calculation formula is as shown in equation (13).
[0138]
[0139] where R and S are the rotation and scaling matrices of the 3D Gaussian points respectively.
[0140] In the rendering stage, the 3D Gaussian is projected onto the 2D plane through a differentiable rasterizer, thus realizing the smooth transition from the 3D covariance matrix to the 2D Jacobian matrix. Mathematically speaking, with the help of a specific transformation matrix, the formula is as shown in equation (14).
[0141]
[0142] where J is the Jacobian matrix, which approximately represents the linear transformation in a small local area. W is the view transformation matrix.
[0143] In addition, the final color C of each pixel is determined by the cumulative contribution of N depth - sorted Gaussian points, and is calculated by α - blending as shown in equation (15).
[0144]
[0145] Among them, c j represents the color of the j-th Gauss point, and α k is obtained by multiplying N projected Gaussians with their respective transparencies.
[0146] The final output consists of a series of rendered images, which are compared with the ground truth, using L1 loss, SSIM loss L SSIM and perceptual loss L LPIPS for supervision. The RGB loss is shown in Equation (16)
[0147] L rgb =(1 - λ ssim )L1 + λ ssim L SSIM + λ lpips L LPIPS , (16)
[0148] where λ ssim = 0.2, λ lpips = 0.1.
[0149] During the training process, as the anchors expand, their range far exceeds the vertex set of the initial template. Therefore, it is crucial to guide these anchors, by establishing benchmarks representing poses and expressions to guide them, allowing the network to output corresponding correction benchmarks. The sampling formula of the loss function is shown in Equation (17).
[0150]
[0151] Among them, respectively represent the expression benchmarks of the predicted and pseudo-ground truth labels, respectively represent the pose benchmarks of the predicted and pseudo-ground truth labels, respectively represent the LBS (Linear Blend Skin) weights of the predicted and pseudo-ground truth labels, In addition, a squared L2 norm loss L reg is imposed on the 3DGS correction offsets (including position, scale, and rotation).
[0152] The main objective of training is to minimize the overall loss function, and the overall loss function is shown in Equation (18):
[0153] L total = L rgb + L flame + L reg , (18)
[0154] When the loss function reaches a relatively stable minimum value, it indicates that the model has been optimized as much as possible under the current training conditions. Based on the optimized model parameters and the attribute data of the three-dimensional Gaussian points, combined with the face template, the finally generated three-dimensional avatar can highly approximate a real human face in terms of shape, expression, pose, and texture details.
[0155] In some embodiments, the method is initialized by sampling the FLAME mesh and voxelizing the scene, thereby anchoring the geometry of the subject's neutral head. The deformation process combines expression and pose bases obtained by inputting tri-planar features into a deformation network and using corresponding coefficients, and can effectively distort the anchor points into the dynamic space. To optimize the consistency of facial textures, an Expression and Pose-dependent Attribute Adapter (EPAA) is introduced to refine the abstract features of the anchor points, which are decoded by their respective small multi-layer perceptrons to derive motion-related attributes. The activated Gaussian point spattering generates head renderings with arbitrary expressions and poses, and is optimized under the supervision of the ground truth.
[0156] By guiding the three-dimensional Gaussian distribution through anchor points, high-fidelity and controllable avatar reconstruction is achieved, where the animation is driven by explicit expression and pose parameters. The basic deformation structure of the parameterized facial model is extended to individual points, endowing each point with stronger robustness to cope with expressive changes. In addition, an expression- and pose-based attribute adapter is introduced to optimize the Gaussian attributes bound to the fixed points, customized for specific actions at the current viewing point.
[0157] In some embodiments, the above method can be implemented using the PyTorch framework and the Adam optimization algorithm. During the construction of the geometric head, the number of Gaussian points bound to each anchor point is set to k = 10, and the small MLP network used to activate the attributes of the three-dimensional Gaussian points (3DGS) is configured with 2 layers, with 32 hidden units in each layer. In the training of the deformation network, an exponential decay strategy similar to 3DGS is adopted to learn the basic facial expression features. At the beginning of the training, the learning rate is set to 1×10 -4 and gradually decreased to 8×10 -6 at the end of the training. The overall training process consists of 600,000 iterations, with gradient accumulation for the Gaussian points every 5,000 steps. Anchor point densification starts from the 10,000th step and is performed every 2,000 steps until the end of the training. The entire task can be completed on a single NVIDIA RTX 3090 GPU.
[0158] In some embodiments, to evaluate the performance of the 3D avatar model, two challenging datasets were selected: the NeRSemble dataset processed by GaussianAvatars and the INSTA dataset. The NeRSemble dataset includes video recordings from 9 subjects with 16 shooting angles, including front and side views. Each subject recorded 10 video sequences, and each sequence contains approximately 200 time steps. The resolution of all images was normalized to 802×550. Nine out of 10 available expression sequences and 15 out of 16 cameras were selected for training. This configuration enables the quantitative evaluation of the method's ability in new expression and view synthesis. The INSTA dataset contains 10 subjects, and the last 350 frames of each subject were used for testing. The resolution of all images is 512×512.
[0159] To determine the baseline methods, several state-of-the-art (SOTA) methods were selected as baselines for head avatar creation for comparison. AvatarMAV uses a motion-aware voxel grid to model expression motifs and predicts the required deformation offsets by combining correlation coefficients. INSTA embeds the neural radiance field into a multi-resolution hash grid and deforms based on adjacent triangles. PointAvatar is a point cloud-based head geometry construction method that uses a coarse-to-fine sampling strategy to control the scale of points. GaussianAvatars is based on 3DGS and combines a parametric shape model to create animatable head avatars. SplattingAvatar also embeds Gaussian points into the grid and designs an optimization method to control their movement between triangles. FlashAvatar attaches 3DGS to the facial UV map and predicts deformable offsets to achieve high-speed avatar rendering.
[0160] By comparing with the ground truth images, it is measured using Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS). A higher PSNR indicates less distortion, a higher SSIM indicates better structural quality, and a lower LPIPS score better reflects the consistency with human vision.
[0161] The comparison results are shown in Tables 1 and 2. Table 1 shows the quantitative comparison with the state-of-the-art methods in new view synthesis and self-reproduction on the NeRSemble dataset. Table 2 shows the quantitative comparison with the state-of-the-art methods in self-reproduction on the INSTA dataset.
[0162] The best results are marked in bold, and the second-best results are underlined.
[0163] Table 1 Quantitative comparison experimental results of the NeRSemble dataset in terms of novel view synthesis and self-reproduction
[0164]
[0165] Table 2 Quantitative comparison experimental results of the INSTA dataset in terms of self-reproduction
[0166]
[0167] To evaluate novel view synthesis and self-reproduction, the NeRSemble dataset was chosen to measure the effect of novel view synthesis because the INSTA dataset is limited to monocular images. In the evaluation of novel expression synthesis, the results from the INSTA dataset and the NeRSemble dataset were compared.
[0168] From the quantitative comparison, Table 1 highlights that this model significantly outperforms recent baseline methods in terms of PSNR, SSIM, and LPIPS metrics. Especially in novel view synthesis, this advantage is mainly attributed to the model's excellent ability to capture continuous high-frequency dynamic details (such as wrinkles, blinks, and lip movements), thus achieving more realistic deformations. Additionally, compared with the original 3DGS, the anchor-guided 3DGS introduces the view direction in the anchor features, improving the sensitivity to view changes. Therefore, through these view-dependent, motion-aware Gaussian attributes, high-fidelity novel view renderings are generated. This model also performs outstandingly in self-reproduction. To further verify its superiority in detail retention and reproduction of unseen expressions, experiments were conducted on the INSTA dataset. As shown in Table 2, this method constructs an accurate head geometry by anchoring the skeleton and enhances this process through the deformation supervision of the FLAME template mesh. This not only ensures high-quality rendering effects but also achieves efficient deformations, obtaining impressive performances in various metrics.
[0169] For qualitative comparison, focus was placed on the visual effects of this method and the baseline method in the reconstruction of 3D head avatars, such as Figure 7 and Figure 8 shown. PointAvatar uses fixed shapes and isotropic points as the original elements for avatar construction, which limits the clear presentation of facial structures and usually produces dot-like artifacts (such as Figure 5as shown in the first column). Although Gaussian Avatars can easily synthesize new expressions using their inherent topology, it is this very characteristic that produces rough deformations and fails to capture fine facial details, complex hairstyles, and facial accessories (such as Figure 7 and Figure 8 as shown in the second column). This duality is both an advantage and a limitation. Splatting Avatar allows Gaussian points to move freely between triangles; however, the lack of explicit attachment fixed points results in interrupted deformations and produces a blurring effect. Obvious noise can also be seen in the neck and shoulder regions (such as Figure 7 and Figure 8 as shown in the third column). In contrast, the present method accurately reconstructs the geometry of the subject by anchoring the template mesh, laying the foundation for this. By using an anchor-based deformation benchmark to model dynamic changes and applying EPAA adjustment, this model better preserves complex details such as wrinkles, hair, teeth, and blinking, and the final rendered effects (such as Figure 7 and Figure 8 as shown in the fourth column) are closer to real images.
[0170] In some embodiments, ablation experiments are conducted on the various components of the present method to evaluate the contributions of each part, and the experimental results are shown in Table 3 and Figure 9 as shown.
[0171] Table 3 Results of Ablation Experiments
[0172]
[0173] Disable the expression and pose-dependent attribute adapter (w / o EPAA): The expression and pose-dependent attribute adapter is disabled, which is used to optimize Gaussian attributes, and only blend shapes and linear skinning are used for animating 3D head avatars. As Figure 10 shown, after disabling the EPAA, although certain deformations can be performed, due to the limitations of linear transformation, high-frequency details and dynamic facial textures cannot be accurately restored. In addition, optimizing Gaussian attributes only in the standard space results in insufficient retention of facial changes, thus reducing the overall realism. It can be seen from this that through the expression and pose-dependent attribute adapter (EPAA), high-frequency changes during the animation process can be captured more precisely, thus achieving a more realistic performance.
[0174] Original 3D Gaussian sampling (w / o anchor-guided): Directly use the original 3DGS for head geometry modeling. As Figure 7 shown, in the early stage of training, the 3DGS did not reconstruct the subject shape as clearly as the anchor-guided strategy. However, the anchor-guided 3DGS achieved faster and more accurate convergence, laying a solid foundation for the subsequent rendering of facial textures.
[0175] Disable Triplane (w / o Triplane): Replace the positional encoding with triplane feature extraction via anchor points. Although this simplification reduces complexity, relying solely on positional encoding fails to fully capture the complete spatial features, especially in the facial regions that are prone to occlusion, which may lead to unrealistic deformations.
[0176] Disable Linear Blend Skinning function (w / o LBS): Remove the LBS function from the complete model, and it is found that the evaluation metrics of self-replay decrease significantly. This is because LBS represents skin stretching through bone transformation and skinning weights. Without this part, the model shows stiffness and inaccuracy when dealing with new expressions.
[0177] The core of this method lies in using anchor points to construct the head geometry, enabling the representation of heads with additional facial accessories and complex hairstyles. By combining the Expression and Pose-dependent Attribute Adapter (EPAA) with the Blend Shape-based Deformation Network, the deformation module can effectively retain exaggerated expressions and delicate facial details, generating realistic animatable head renderings. Superior performance compared to the baseline method is demonstrated on a multi-view dataset, aiming to create animatable head avatars in a way that balances computational efficiency and rendering quality.
[0178] An embodiment of the present invention also provides a three-dimensional avatar model construction system, including:
[0179] The first module is used to obtain multi-view face information of a target person;
[0180] The second module is used to perform feature extraction on the multi-view face information to obtain shape parameters, first pose parameters, and first expression parameters;
[0181] The third module is used to determine multiple anchor point information on the face template according to the shape parameters, where each anchor point information is used to manage the attribute data of the three-dimensional Gaussian points associated with several face templates;
[0182] The fourth module is used to adjust the anchor point information according to the first pose parameters and the first expression parameters to obtain the first attribute data of the three-dimensional Gaussian points corresponding to the anchor point information, where the first attribute data of the three-dimensional Gaussian points is used to obtain the second attribute data of the three-dimensional Gaussian points;
[0183] The fifth module is used to obtain the target three-dimensional avatar according to the second attribute data of the three-dimensional Gaussian points and the face template.
[0184] It can be understood that the content in the above embodiments of the three-dimensional avatar model construction method is applicable to the embodiments of this system. The functions specifically implemented by the embodiments of this system are the same as those of the above embodiments of the three-dimensional avatar model construction method, and the beneficial effects achieved are also the same as those of the above embodiments of the three-dimensional avatar model construction method.
[0185] Next, in conjunction with Figure 11 the electronic device of the embodiments of this application will be introduced in detail.
[0186] As Figure 11 , Figure 11 schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0187] A processor 1100, which can be implemented by using a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of this disclosure;
[0188] A memory 1200, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1200 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1200 and are called by the processor 1100 to execute the three-dimensional avatar model construction method of the embodiments of this disclosure;
[0189] An input / output interface 1300, which is used to implement information input and output;
[0190] A communication interface 1400, which is used to implement communication interaction between this device and other devices, and can implement communication through a wired method (such as USB, network cable, etc.) or through a wireless method (such as mobile network, WIFI, Bluetooth, etc.);
[0191] A bus 1500, which transmits information between various components of the device (such as the processor 1100, the memory 1200, the input / output interface 1300, and the communication interface 1400);
[0192] Among them, the processor 1100, the memory 1200, the input / output interface 1300, and the communication interface 1400 are communicatively connected to each other inside the device through the bus 1500.
[0193] An embodiment of the present disclosure also provides a storage medium, which is a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the above-described three-dimensional avatar model construction method.
[0194] As a non-transitory computer-readable storage medium, a memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0195] The embodiments described in the embodiments of the present disclosure are for more clearly illustrating the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.
[0196] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present disclosure, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.
[0197] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0198] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0199] In the description of the present application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0200] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0201] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.
[0202] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0203] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0204] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0205] The preferred embodiments of the present disclosure have been described above with reference to the accompanying drawings, and thus do not limit the scope of rights of the present disclosure. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the present disclosure shall be within the scope of rights of the present disclosure.
Claims
1. A method for constructing a three-dimensional avatar model, characterized in that: The following steps are involved: Obtain multi-view facial information of the target person; Extracting features from the multi-view facial information to obtain shape parameters, first posture parameters, and first expression parameters; Determine multiple anchor point information on the face template according to the shape parameter, wherein each anchor point information is used to manage attribute data of three-dimensional Gaussian points on several associated face templates; Adjust the anchor point information according to the first posture parameter and the first expression parameter to obtain first attribute data of the three-dimensional Gaussian point corresponding to the anchor point information, wherein the first attribute data of the three-dimensional Gaussian point is used to obtain second attribute data of the three-dimensional Gaussian point; A target three-dimensional avatar is obtained according to the second attribute data of the three-dimensional Gaussian points and the face template.
2. The method for constructing a three-dimensional portrait model according to claim 1, characterized in that: Determining multiple anchor point information on the face template according to the shape parameters includes the following steps: According to the shape parameters, constraining the face template to obtain an initial three-dimensional head portrait; The initial three-dimensional head portrait is sampled to obtain a plurality of anchor point information.
3. The method for constructing a three-dimensional portrait model according to claim 1, characterized in that: The step of adjusting the anchor point information according to the first posture parameter and the first expression parameter to obtain first attribute data of the three-dimensional Gaussian point corresponding to the anchor point information comprises the following steps: Inputting the first posture parameter into a posture prediction network to obtain a predicted posture base, and inputting the first expression parameter into an expression prediction network to obtain a predicted expression base; adjusting the anchor point information according to the predicted posture base and the predicted expression base; The adjusted anchor point information is decoded through a multi-layer perceptron to obtain first attribute data of the three-dimensional Gaussian point corresponding to the anchor point information, wherein the attribute data includes opacity information, rotation information, scale information and color information.
4. The method for constructing a three-dimensional head portrait model according to claim 3, characterized in that: The step of adjusting the anchor point information according to the predicted posture base and the predicted expression base comprises the following steps: Predicting a linear skinning weight according to the unadjusted anchor point information, the first posture parameter, and the first expression parameter; The anchor point information is adjusted according to the predicted pose basis, the predicted expression basis and the linear skin weight.
5. The method for constructing a three-dimensional head portrait model according to claim 1, characterized in that: After adjusting the anchor point information according to the first posture parameter and the first expression parameter to obtain the first attribute data of the three-dimensional Gaussian point corresponding to the anchor point information, the following steps are also included: Obtaining a first three-dimensional avatar according to the first attribute data of the three-dimensional Gaussian points and the face template; The second attribute data of the three-dimensional Gaussian point is obtained through an attribute adapter according to the adjusted anchor point information and the first three-dimensional avatar.
6. The method for constructing a three-dimensional head portrait model according to claim 5, characterized in that: The method of obtaining the second attribute data of the three-dimensional Gaussian point according to the first three-dimensional avatar through the attribute adapter includes the following steps: Extracting features of the first three-dimensional avatar to obtain second posture parameters and second expression parameters; Obtaining, by the first module of the attribute adapter, the initial second attribute data of the three-dimensional Gaussian point according to the adjusted anchor point information, the second posture parameter and the second expression parameter; Obtaining, by the second module of the attribute adapter, an offset of the second attribute data according to the adjusted anchor point information, the second posture parameter and the second expression parameter; The second attribute data of the three-dimensional Gaussian point is obtained according to the offset and the initial second attribute data of the three-dimensional Gaussian point.
7. The method for constructing a three-dimensional head portrait model according to claim 1, characterized in that: The step of obtaining a target three-dimensional head portrait according to the second attribute data of the three-dimensional Gaussian points and the face template comprises the following steps: Obtaining an initialized three-dimensional head portrait model according to the second attribute data of the three-dimensional Gaussian points and the face template; Project the three-dimensional head portrait model under the target perspective onto a two-dimensional plane to obtain a predicted rendering image; Establishing a loss function of the three-dimensional avatar model according to the predicted rendering and the real rendering under the target perspective; The target three-dimensional head portrait is obtained by minimizing the loss function.
8. A three-dimensional head portrait model construction system, characterized in that: include: The first module is used to obtain multi-view facial information of the target person; The second module is used to extract features from the multi-view facial information to obtain shape parameters, first posture parameters and first expression parameters; A third module is used to determine multiple anchor point information on the face template according to the shape parameters, wherein each anchor point information is used to manage attribute data of three-dimensional Gaussian points on several associated face templates; A fourth module is used to adjust the anchor point information according to the first posture parameter and the first expression parameter to obtain first attribute data of the three-dimensional Gaussian point corresponding to the anchor point information, wherein the first attribute data of the three-dimensional Gaussian point is used to obtain second attribute data of the three-dimensional Gaussian point; The fifth module is used to obtain a target three-dimensional head portrait according to the second attribute data of the three-dimensional Gaussian points and the face template.
9. An electronic device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the three-dimensional avatar model construction method as described in any one of claims 1-7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The program executable by the processor is used to implement the three-dimensional avatar model construction method as described in any one of claims 1 to 7 when executed by the processor.
Citation Information
Patent Citations
Scene-adaptive fine three-dimensional face reconstruction method and system, and electronic equipment
CN113269862A
Dynamic human body modeling method based on three-dimensional Gaussian
CN117671108A
Gaussian mixture shape method suitable for dynamic modeling of human head
CN118135655A
Low-bit-rate free viewpoint video generation method for efficient streaming based on 3D Gaussian, computer equipment, readable storage medium and program product
CN118573978A
Expression-editable voice-driven face reconstruction method based on three-dimensional Gaussian sputtering technology
CN118762133A
Cited By
Scene semantic occupancy prediction method based on semantic-distance adaptive Gaussian
CN120673417A
Three-dimensional head reconstruction method based on geometric prior network and related device
CN121482228A