Digital human 3D face reconstruction method, computer equipment and readable storage medium
By obtaining the ARKit expression model and 3D modeling software, the target 3D head model was created, and combined with the neural network training data set, the problem that ARKit cannot fully express real-person expressions is solved, and a richer 3D face reconstruction of head shapes and expressions is achieved, improving the reconstruction accuracy and speed.
Patent Information
- Application Number
- CN202510594532.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-07-29
AI Technical Summary
The existing ARKit-based 3D face reconstruction technology cannot fully express various expressions of real people, and lacks shape parameters, resulting in poor reconstruction results.
By obtaining multiple expression models under the ARKit framework, combining 3D modeling software to create basic 3D head model, fit expression parameters and shape parameters, generate target 3D head model, and use target 3D head model to perform 3D face reconstruction, further real-time generation of 3D faces through the neural network model training data set.
It achieves richer expressions of head shapes and expressions, improves the accuracy and speed of 3D face reconstruction, and can generate 3D face models that are close to real people in real time.
Smart Images

Figure CN120388114A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method for 3D face reconstruction of digital humans, a computer device, and a readable storage medium. Background Art
[0002] 3D (3 Dimensions, three-dimensional) face reconstruction refers to reconstructing a 3D model of a human face from one or more 2D (2 Dimensions, two-dimensional) images. With the rise of VR, AR, and the metaverse, more and more computer application software needs to have the function of 3D face reconstruction, which requires mapping the real human face image into the virtual world, generating a virtual digital human with a real human image, and driving the expression of the virtual human through real human face expressions.
[0003] ARKit is an augmented reality (AR) development framework developed by Apple for iOS devices. Currently, the basic expression widely used in the industry for 3D face reconstruction is the face expression model under the ARKit framework. This expression model only includes the eye sockets, nose, and mouth, and only the facial part, and cannot clearly express various human expressions in all directions. Moreover, the face shape provided by the face expression model under the ARKit framework cannot be changed and there is only one shape.
[0004] Therefore, the existing 3D face reconstruction based on ARKit has problems such as poor 3D face reconstruction effect due to the lack of shape parameters and the inability to fully express various real human expressions. Summary of the Invention
[0005] In order to overcome the above defects, the present invention proposes a method for 3D face reconstruction of digital humans, a computer device, and a readable storage medium that can make the reconstructed human face richer and closer to real humans in terms of shape and expression.
[0006] In a first aspect, the present invention provides a method for 3D face reconstruction of digital humans, including:
[0007] Obtaining a plurality of different expression models under the ARKit framework, and creating a basic 3D head model based on 3D modeling software;
[0008] Using the basic 3D head model to fit each expression model respectively to determine multiple groups of expression parameters and shape parameters, and creating a target 3D head model based on the multiple groups of expression parameters, shape parameters, and the basic 3D head model;
[0009] Performing 3D face reconstruction on a face image based on the target 3D head model.
[0010] In some embodiments, the obtaining a plurality of different expression models under the ARKit framework includes:
[0011] For 52 sets of motion factors under the ARKit framework, 52 expression models are obtained by setting the expression parameters corresponding to one set of motion factors to 1 each time and setting the expression parameters corresponding to the other sets of motion factors to 0 to obtain an expression model.
[0012] In some embodiments, creating a basic 3D head model based on 3D modeling software includes:
[0013] Drawing a basic head structure model based on 3D modeling software, wherein the shape parameters of the basic head structure model are configured as parameter values representing a neutral head shape;
[0014] Obtaining human head bone data, and binding the head bones to the basic head structure model based on the human head bone data to obtain the basic 3D head model. Among them, the shape parameters of the basic 3D head model are the same as those of the basic head structure model, and the expression parameters of the basic 3D head model are configured as parameter values representing no expression.
[0015] In some embodiments, the process of the basic 3D head model fitting each expression model includes:
[0016] Determining a first residual function according to the distance between the vertices in the grid corresponding to the basic 3D head model and the faces in the grid corresponding to the expression model;
[0017] Updating the model parameters of the basic 3D head model based on the first residual function and the gradient descent method until the model converges, and determining the expression parameters and shape parameters based on the model parameters when the model converges.
[0018] Further, the determining a first residual function according to the distance between the vertices in the grid corresponding to the basic 3D head model and the faces in the grid corresponding to the expression model includes:
[0019] Obtaining all the vertices of the grid corresponding to the basic 3D head model, determining the face closest to each vertex from all the constituent faces of the grid corresponding to the expression model for each vertex, and determining the first residual function according to the distance calculation formula between each vertex and the determined closest face.
[0020] In some embodiments, the 3D face reconstruction of a face image based on the target 3D head model includes:
[0021] Extracting 2D face key points based on the face image;
[0022] Obtaining target shape parameters and target expression parameters matched to the face image based on the 2D face key points and the target 3D head model;
[0023] Perform 3D face reconstruction based on the target shape parameters and the target expression parameters.
[0024] Further, the obtaining of the target shape parameters and the target expression parameters for matching the face image based on the 2D face key points and the target 3D head model includes:
[0025] Initialize the model parameters of the target 3D head model, where the model parameters include shape parameters and expression parameters;
[0026] Based on the 2D face key points and the target 3D head model, obtain 3D face key points corresponding to the 2D face key points, and obtain the projected key points by projecting the 3D face key points onto the face image;
[0027] Construct a function for calculating the Euclidean distance between the projected key points and the 2D face key points as the second residual function;
[0028] Based on the second residual function and the gradient descent method, update the model parameters of the target 3D head model until the model converges, and determine the target shape parameters and the target expression parameters according to the model parameters at the time of model convergence.
[0029] In some embodiments, the method further includes: obtaining a training data set based on multiple face images and the target 3D head model, training a neural network model based on the training data set, and performing 3D face reconstruction on a target face image based on the neural network model; wherein, the input of the neural network model is a face image, and the output is a 3D face model obtained by performing 3D face reconstruction on the target face image.
[0030] In a second aspect, the present invention provides a computer device, which includes a processor and a storage device. The storage device is adapted to store multiple program codes, and the program codes are adapted to be loaded and run by the processor to execute the digital human 3D face reconstruction method described in any one of the technical solutions of the above digital human 3D face reconstruction method.
[0031] In a third aspect, a computer-readable storage medium is provided, which stores multiple program codes therein, and the program codes are adapted to be loaded and run by a processor to execute the digital human 3D face reconstruction method described in any one of the technical solutions of the above digital human 3D face reconstruction method.
[0032] One or more of the above technical solutions of the present invention have at least one or more of the following beneficial effects: The target 3D head model created in this application has both shape parameters and expression parameters at the same time. Compared with the existing face reconstruction implemented by the expression model based on ARKit, it can express richer head shapes and facial expressions; on the other hand, using the target 3D head model created in this application, a large amount of data from face pictures to face 3D models can be constructed. According to the constructed dataset, a neural network model can be trained to achieve the effect of real-time generating 3D faces. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Referring to the accompanying drawings, the disclosure of the present invention will become more understandable. It is easy for those skilled in the art to understand that these drawings are only for illustrative purposes and are not intended to limit the protection scope of the present invention. In addition, similar numbers in the figures are used to represent similar components, where:
[0034] Figure 1 is a schematic diagram of the main step flow of a digital human 3D face reconstruction method according to an embodiment of the present invention;
[0035] Figure 2 is a schematic diagram of the effects of various expression models under the ARKit framework obtained;
[0036] Figure 3 is a schematic diagram of the effects of the expression model that can be fitted based on the target 3D head model created in this application.
[0037] Figure 4 is a schematic diagram of the main step flow of a digital human 3D face reconstruction method based on a neural network model according to an embodiment of the present invention;
[0038] Figure 5 is a schematic diagram of the structural block diagram of a digital human 3D face reconstruction system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The following describes some embodiments of the present invention with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.
[0040] In the description of the present invention, a "module" and a "processor" may include hardware, software, or a combination of both. A module may include a hardware circuit, various appropriate sensors, communication ports, a memory, and may also include a software part, such as program code, or may be a combination of software and hardware. A processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. The processor has data and / or signal processing functions. The processor may be implemented in software, in hardware, or in a combination of both. A non-transitory computer-readable storage medium includes any suitable medium for storing program code, such as a magnetic disk, a hard disk, an optical disk, a flash memory, a read-only memory, a random access memory, and so on. The term "A and / or B" represents all possible combinations of A and B, such as only A, only B, or A and B. The term "at least one of A or B" or "at least one of A and B" has a meaning similar to "A and / or B" and may include only A, only B, or A and B. The singular terms "a" and "the" may also include the plural form.
[0041] Here, some terms related to the present invention will be explained first.
[0042] ARKit: An augmented reality (AR) development framework developed by Apple for iOS devices, used to create and display augmented reality experiences on iOS devices.
[0043] Maya: A 3D modeling and animation software widely used in the creation of digital special effects for movies, television, advertisements, computer games, and console games.
[0044] MetaHuman: Refers to a virtual digital human created through 3D modeling technology, which is a highly realistic and customizable digital human model, usually used in scenarios such as the metaverse, virtual reality, and augmented reality. MetaHuman binding can be directly used in the UE5 engine or in Maya.
[0045] Refer to the appendix Figure 1 , Figure 1 is a schematic diagram of the main step flow of a 3D face reconstruction method for a digital human according to an embodiment of the present invention, mainly including the following steps S11 to step S14.
[0046] Step S11: Create multiple different expression models based on the ARKit framework;
[0047] It should be understood that ARKit describes facial expressions through a method called BlendShapes (also known as morphers in 3D Max). This concept was originally used to describe the displacement of a model mesh through parameter control, and in ARKit, it is used to represent the technology of driving a model through human face expression factors. Technically, BlendShapes is a dictionary that stores the motion factors of the user's facial expression features, and it contains a total of 52 sets of feature motion data. ARKit can set the corresponding motion factors in real time according to the collected user expression feature values, and these motion factors can be used to drive 2D or 3D human face models to make these models present expressions consistent with the user. Among these 52 sets of factors, there are 7 sets of left-eye motion factors, 7 sets of right-eye motion factors, 27 sets of mouth and chin motion factors, 10 sets of eyebrow, cheek, and nose motion factors, and a set of tongue motion factor data. Each set of motion factors represents a facial expression feature recognized by ARKit, which contains a locator representing a specific human face expression and a floating-point type value representing the degree of expression. The range of the expression degree value is [0, 1], where 0 means no expression and 1 means full expression.
[0048] In this embodiment, the implementation manner of this step specifically includes: for the 52 sets of motion factors under the ARKit framework, obtaining 52 expression models by setting the expression parameters (i.e., the above-mentioned expression degree values) corresponding to one set of motion factors to 1 each time and setting the expression parameters corresponding to other sets of motion factors to 0. The schematic diagram of the effect is specifically as Figure 2 shown.
[0049] Step S12: Create a basic 3D head model based on 3D modeling software;
[0050] In this embodiment, the implementation manner of this step may specifically include the following steps 121 and 122:
[0051] Step 121: Draw a basic head structure model based on 3D modeling software, where the shape parameters of the basic head structure model are configured as parameter values representing a neutral head shape;
[0052] In this embodiment, the key points (or marker points) of the basic head structure model drawn based on 3D modeling software at least include the eyeballs, the head, and the neck. For example, the 3D modeling software is Maya software.
[0053] Step 122: Obtain the human head bone data, and bind the head bones to the basic head structure model based on the human head bone data to obtain the basic 3D head model. Among them, the shape parameters of the basic 3D head model are the same as those of the basic head structure model, and the expression parameters of the basic 3D head model are configured as parameter values representing no expression.
[0054] In this embodiment, the human head bone data can be directly created in Maya or can be data pre-imported into Maya. For example, based on the digital human production tool (MetaHuman), the MetaHuman asset is imported into the Maya software to bind the head bones for the head basic structure model based on the MetaHuman head bone data. Here, the asset can be understood as a bound model file that can be directly used in the game engine or Maya.
[0055] Step S13: Use the basic 3D head model to fit each expression model respectively to determine multiple groups of expression parameters and shape parameters, and create a target 3D head model based on the multiple groups of expression parameters, shape parameters, and the basic 3D head model;
[0056] In this embodiment, the representation of the target 3D head model created based on the multiple groups of expression parameters, shape parameters, and the basic 3D head model is as follows:
[0057] T P (β, θ, ψ) = T + B S (β; S) + B P (θ; P) + B E (ψ; E);
[0058] Wherein, β, θ, and ψ are model parameters, β represents the shape parameter, ψ represents the expression parameter, θ represents the rotation parameter, T is the basic 3D head model, and B S (β; S) is the blendshape function, and B P (θ; P) is the human head pose correction function, and B E (ψ; E) is the human head expression correction function, and the correction function is used to compensate for the human head mesh defects generated by blendshape.
[0059] It can be understood that in this embodiment, all 52 expression models can be used, that is, the basic 3D head model is respectively fitted with the 52 expression models, or a part of the 52 expression models can be used. Correspondingly, the created target 3D head model can be used to represent 52 kinds of human face expressions. Exemplarily, the target 3D head model created based on this application can fit the effect schematic diagrams of the expression models corresponding one-to-one to the ARKit expression models as shown in Figure 2 shown, and the schematic diagram is as shown in Figure 3 shown.
[0060] In this embodiment, the process of the basic 3D head model fitting each expression model specifically includes the following steps 131 and 132:
[0061] Step 131: Determine a first residual function based on the distance between the vertices in the mesh corresponding to the base 3D head model and the faces in the mesh corresponding to the expression model;
[0062] Step 132: Update the model parameters of the base 3D head model based on the first residual function and the gradient descent method until the model converges, and determine the expression parameters and shape parameters based on the model parameters when the model converges.
[0063] Among them, the update of the model parameters includes the synchronous update of the expression parameters and the shape parameters. When the expression changes, the positions of the key points change, and the change in the positions of the key points corresponds to a change in the shape. For example, when the facial expression is that the mouth is wide open, the head shape will also change accordingly. For example, the surface at the junction of the chin and the neck will become smoother.
[0064] One implementation of the above step 131 is as follows: Obtain all the vertices of the mesh corresponding to the base 3D head model. For each vertex, determine the face closest to the vertex from all the constituent faces of the mesh corresponding to the expression model, and determine the first residual function according to the calculation formula of the distance between each vertex and the determined closest face.
[0065] The first residual function is expressed as follows:
[0066] L dis = ∑||v i - Ψ(M ij )||; where, v i represents a vertex, and Ψ(M ij ) represents the face closest to the vertex v i among all the constituent faces of the expression model.
[0067] Step S14: Perform 3D face reconstruction on the face image based on the target 3D head model.
[0068] In this embodiment, the implementation of this step specifically includes the following steps 141 to 143:
[0069] Step 141: Extract 2D face key points based on the face image;
[0070] Step 142: Obtain the target shape parameters and target expression parameters matched for the face image based on the 2D face key points and the target 3D head model;
[0071] Step 143: Perform 3D face reconstruction based on the target shape parameters and the target expression parameters.
[0072] It can be understood that in practical applications, when the acquired picture is a human image, 3D face reconstruction can also be performed based on the target 3D head model. In a specific implementation: First, the acquired picture is preprocessed. For example, through basic white balance algorithms and wide dynamic range algorithms in digital image processing, so that there will be no obvious color cast, unclear mixing of dark and bright areas in the overall picture after processing; then, based on the face detection algorithm, a face image is identified and acquired from the preprocessed picture or 2D face key points are directly identified and acquired.
[0073] An implementation manner of the above step 142 may specifically include the following steps 1421 to 1424:
[0074] Step 1421: Initialize the model parameters of the target 3D head model, where the model parameters include shape parameters and expression parameters;
[0075] Exemplarily, the initialization process may be to set both the shape parameters and expression parameters of the target 3D head model to 0, obtaining a 3D head model with a neutral head shape and no expression.
[0076] Step 1422: Based on the 2D face key points and the target 3D head model, obtain 3D face key points corresponding to the 2D face key points, and obtain the projected key points by projecting the 3D face key points onto the face image;
[0077] Exemplarily, it may be to screen and obtain key points with the same ID as the 2D face key points from the key points of the target 3D head model as the 3D face key points. The key points of the target 3D head model include key points marked as the face, and each key point marked as the face has a unique ID. The ID of the key point is usually used for facial feature annotation. For example, the key point ID of the left eye may be from 42 to 48, and the key point ID of the lips may be 48 and 68.
[0078] Step 1423: Construct a function for calculating the Euclidean distance between the projected key points and the 2D face key points as the second residual function;
[0079] Among them, the second residual function is expressed as follows:
[0080] Among them, (x i , y i ) represents the position of the 2D face key point, and (x j , y j ) represents the position of the projected key point.
[0081] Step 1424: Update the model parameters of the target 3D head model based on the second residual function and the gradient descent method until the model converges, and determine the target shape parameters and target expression parameters according to the model parameters at the time of convergence.
[0082] Specifically: Calculate the gradient of the second residual function under the current model parameters, use the learning rate and the gradient to update the model parameters, repeat the gradient calculation and parameter update until the model converges, and save the shape parameters and expression parameters at the time of model convergence as the finally determined target shape parameters and target expression parameters.
[0083] In a specific implementation manner, based on the target 3D head model created in the above steps S11 to S13, a large amount of data of face images to 3D faces can be constructed and obtained. According to the constructed data set, a neural network model can be trained, so as to achieve the effect of real-time generating 3D faces based on the neural network model. Correspondingly, after the above step S13, it may also include steps S21 to S23 as shown in Figure 4 The method based on the neural network shown has the advantage over Figure 4 the non-linear optimization digital human 3D face reconstruction method in that the speed is fast, and the effect of real-time reconstruction and generating 3D faces can be achieved; the advantage of non-linear optimization is high reconstruction accuracy, but the speed is slow, and one reconstruction may take several minutes. The main implementation steps of the digital human 3D face reconstruction method based on the neural network model as shown in Figure 1 include the following steps S21 to S23. Figure 4 Step S21: Obtain a training data set based on multiple face images and the target 3D head model;
[0084] Step S22: Train a neural network model based on the training data set;
[0085] Step S23: Perform 3D face reconstruction on the target face image based on the neural network model to generate a 3D face model.
[0086] An implementation manner of the above step S21 specifically includes: obtaining the shape parameters and expression parameters matched for each portrait image based on the target 3D head model, and performing 3D face reconstruction on the target face image based on the matched shape parameters and expression parameters to obtain a 3D face model; obtaining a training data set by taking each face image and its corresponding 3D face model as a set of training data.
[0087]
[0088] It is understandable that the training process of training a neural network model based on a training dataset can be directly carried out using existing neural network model training methods, and the input, output, and loss function of the model can be configured according to actual needs. Among them, the neural network model incorporates a 3D face driving algorithm. When the input of the neural network model is a face image, the output is a 3D face model generated by reconstructing the face of the face image.
[0089] It should be noted that all user-related data collected in this application is collected with the consent and authorization of the users, and the collection, use, and processing of relevant user data need to comply with relevant laws, regulations, and standards of relevant countries and regions. For example, the face images or portrait images used in this application are all permitted by the users and comply with legal regulations.
[0090] Refer to the attached Figure 5 , Figure 5 is a schematic diagram of the composition structure of a digital human 3D face reconstruction system provided by an embodiment of this application. As shown in the figure, the system mainly includes an acquisition module 31, a head model construction module 32, and a face reconstruction module 33, where:
[0091] The acquisition module 31 is configured to acquire multiple different expression models under the ARKit framework;
[0092] The head model construction module 32 is configured to create a basic 3D head model using 3D modeling software, and use the basic 3D head model to fit each expression model respectively to determine multiple groups of expression parameters and shape parameters, and create a target 3D head model based on the multiple groups of expression parameters, shape parameters, and the basic 3D head model;
[0093] The face reconstruction module 33 is configured to perform 3D face reconstruction on a face image based on the target 3D head model created by the head model construction module 32.
[0094] Furthermore, as shown in the figure, the system may further include a deep learning module 34 and a 3D face generation module 35.
[0095] The deep learning module 34 is configured to obtain a training dataset based on multiple face images and the target 3D head model created by the head model construction module 32, and train a neural network model based on the training dataset.
[0096] The 3D face generation module 35 is configured to perform 3D face reconstruction on a target face image based on the neural network model obtained by the deep learning module 34 to generate a 3D face model.
[0097] Furthermore, it should be understood that since the settings of the respective modules are only for illustrating the functional units of the system of the present invention, the physical devices corresponding to these modules can be the processor itself, or a part of the software in the processor, a part of the hardware, or a part of the combination of software and hardware. Therefore, the number of each module in the figure is only illustrative.
[0098] Those skilled in the art can understand that the respective modules in the system can be adaptively split or combined. Such splitting or combining of the specific modules will not cause the technical solution to deviate from the principle of the present invention. Therefore, the technical solutions after splitting or combining will all fall within the protection scope of the present invention.
[0099] It should be noted that although the above-mentioned embodiments describe the respective steps in a specific order, those skilled in the art can understand that in order to achieve the effects of the present invention, it is not necessary for different steps to be executed in such an order. They can be executed simultaneously (in parallel) or in other orders, and these variations are all within the protection scope of the present invention.
[0100] Those skilled in the art can understand that all or part of the processes in the method of the above-mentioned embodiment of the present invention can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electrical carrier signal, telecommunication signal, and software distribution medium that can carry the computer program code, etc.
[0101] Furthermore, the present invention also provides a computer device. In an embodiment of a computer device according to the present invention, the computer device includes a processor and a storage device. The storage device can be configured to store a program for executing the digital human 3D face reconstruction method of the above-mentioned method embodiment. The processor can be configured to execute the program in the storage device, and the program includes but is not limited to the program for executing the digital human 3D face reconstruction method of the above-mentioned method embodiment. For the sake of convenience of description, only the parts related to the embodiments of the present invention are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present invention.
[0102] Furthermore, the present invention also provides a computer-readable storage medium. In an embodiment of the computer-readable storage medium according to the present invention, the computer-readable storage medium may be configured to store a program for executing the digital human 3D face reconstruction method in the above method embodiment. This program can be loaded and run by a processor to implement the above digital human 3D face reconstruction method. For the sake of convenience of description, only the parts related to the embodiments of the present invention are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present invention. The computer-readable storage medium may be a storage device formed by various electronic devices. Optionally, the computer-readable storage medium in the embodiments of the present invention is a non-transitory computer-readable storage medium.
[0103] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. A method for 3D human face reconstruction of a digital human, characterized in that, Including: Obtaining multiple different expression models under the ARKit framework, and creating a basic 3D head model based on 3D modeling software; Using the basic 3D head model to fit each expression model respectively to determine multiple groups of expression parameters and shape parameters, and creating a target 3D head model based on the multiple groups of expression parameters, shape parameters and the basic 3D head model; Performing 3D face reconstruction on a face image based on the target 3D head model.
2. The method according to claim 1, wherein The obtaining of multiple different expression models under the ARKit framework includes: For 52 groups of motion factors under the ARKit framework, obtaining 52 expression models by setting the expression parameters corresponding to one group of motion factors to 1 each time and setting the expression parameters corresponding to other groups of motion factors to 0.
3. The method according to claim 1, characterized in that, The creating of a basic 3D head model based on 3D modeling software includes: Drawing a basic head structure model based on 3D modeling software, where the shape parameters of the basic head structure model are configured as parameter values representing a neutral head shape; Obtaining human head bone data, and binding a head bone to the basic head structure model based on the human head bone data to obtain the basic 3D head model. Among them, the shape parameters of the basic 3D head model are the same as those of the basic head structure model, and the expression parameters of the basic 3D head model are configured as parameter values representing no expression.
4. The method according to claim 1, wherein The process of the basic 3D head model fitting each expression model includes: Determining a first residual function according to the distance between the vertices in the mesh corresponding to the basic 3D head model and the faces in the mesh corresponding to the expression model; Updating the model parameters of the basic 3D head model based on the first residual function and the gradient descent method until the model converges, and determining the expression parameters and shape parameters based on the model parameters at the time of model convergence.
5. The method according to claim 4, wherein The determining of the first residual function according to the distance between the vertices in the mesh corresponding to the basic 3D head model and the faces in the mesh corresponding to the expression model includes: Obtaining all the vertices of the mesh corresponding to the basic 3D head model, determining the face closest to each vertex from all the constituent faces of the mesh corresponding to the expression model for each vertex, and determining the first residual function according to the distance calculation formula between each vertex and the determined closest face.
6. The method according to claim 1, wherein The performing of 3D face reconstruction on a face image based on the target 3D head model includes: Extracting 2D face key points based on the face image; Obtaining target shape parameters and target expression parameters matched for the face image based on the 2D face key points and the target 3D head model; Performing 3D face reconstruction based on the target shape parameters and the target expression parameters.
7. The method according to claim 6, wherein The obtaining of target shape parameters and target expression parameters matched for the face image based on the 2D face key points and the target 3D head model includes: Initializing the model parameters of the target 3D head model, where the model parameters include shape parameters and expression parameters; Obtaining 3D face key points corresponding to the 2D face key points based on the 2D face key points and the target 3D head model, and obtaining the projected key points by projecting the 3D face key points onto the face image; Construct a function for calculating the Euclidean distance between the projected key points and the 2D face key points as the second residual function; Update the model parameters of the target 3D head model based on the second residual function and the gradient descent method until the model converges, and determine the target shape parameters and target expression parameters according to the model parameters at the time of model convergence.
8. The method according to any one of claims 1-7, characterized in that, The method further includes: obtaining a training data set based on multiple face images and the target 3D head model, training a neural network model based on the training data set, and performing 3D face reconstruction on a target face image based on the neural network model to generate a 3D face model; Wherein, the input of the neural network model is a face image, and the output is a 3D face model generated by performing 3D face reconstruction on the target face image.
9. A computer device, comprising a processor and a storage device, the storage device being adapted to store multiple program codes, characterized in that, The program code is adapted to be loaded and run by the processor to execute the digital human 3D face reconstruction method according to any one of claims 1-7 or claim 8.
10. A computer-readable storage medium storing multiple program codes, characterized in that, The program code is adapted to be loaded and run by the processor to execute the digital human 3D face reconstruction method according to any one of claims 1-7 or claim 8.
Citation Information
Cited By
Facial expression real-time closed-loop compensation method for bionic human head
CN122110754A
A real-time closed-loop compensation method for facial expression of a bionic human head
CN122110754B