Method, electronic device and storage medium for migrating facial expressions from two-dimensional to three-dimensional

Through computer vision and deep learning technology, the expression parameters of the two-dimensional face image are converted into the fusion deformation coefficient in the three-dimensional face deformation model, solving the problem of low efficiency of two-dimensional to three-dimensional face expression migration, and achieving efficient and flexible facial expression migration.

CN114926581BActive Publication Date: 2025-05-06INST OF SOFTWARE - CHINESE ACAD OF SCI
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210430797.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-22
Publication Date
2025-05-06
Estimated Expiration
2042-04-22

AI Technical Summary

Technical Problem

The existing technology is difficult to achieve efficient two-dimensional to three-dimensional facial expression migration, and the existing methods are coupled with identity information and expression information during the face pinching process, so it is impossible to achieve flexible facial expression migration.

Method used

Using computer vision technology and deep learning technology in the field of artificial intelligence, the model is extracted and fusion deformation coefficient estimation model through pre-trained three-dimensional face parameters, and the expression parameters of the two-dimensional face image are extracted and converted into the fusion deformation coefficient in the three-dimensional face deformation model, thereby realizing the migration of two-dimensional to three-dimensional face expressions.

Benefits of technology

It greatly reduces the manual burden of designers in creating virtual facial expressions and animations, improves the work efficiency of animators, and realizes flexible facial expression migration, and only a single two-dimensional face image can be completed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114926581B_ABST
    Figure CN114926581B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, an electronic device and a storage medium for migrating facial expressions from two dimensions to three dimensions, and belongs to the field of computer vision. The method comprises: obtaining a two-dimensional facial image in a real interactive scene as a source expression representation, and a three-dimensional facial deformation model in a virtual interactive scene as a target three-dimensional facial expression representation; extracting three-dimensional facial parameters for the two-dimensional facial image; obtaining the fusion deformation coefficient of the target expression through a fusion deformation coefficient estimation model based on the expression parameters in the three-dimensional facial parameters; and driving the three-dimensional facial model to generate the target expression representation according to the obtained target expression fusion deformation coefficient. The present invention can effectively extract parameters related to facial expressions, alleviate the cross-dimensional problem of facial expressions from two-dimensional space to three-dimensional space; and transform the expression parameters into target fusion deformation coefficients, thereby ensuring the wide applicability of the method. The present method can realize accurate and rapid facial expression migration, and can be used to improve the work efficiency of animators in the process of three-dimensional facial modeling and creation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the relevant technical fields of computer vision and three-dimensional modeling, and in particular relates to a two-dimensional to three-dimensional facial expression migration method, an electronic device and a storage medium. Background Art

[0002] The human face plays a vital role in daily life, reflecting the basic identity characteristics of each person. At the same time, rich facial expressions can convey a person's emotions and mental state, and express very rich social and cultural significance. In a virtual environment, if virtual humans can accurately imitate human facial expressions, the distance between virtual humans and real humans will be greatly shortened. With the introduction of new concepts such as the metaverse, virtual reality and artificial intelligence have become the two cornerstones of future technological development. People hope to be able to communicate and interact with virtual humans like real humans in a computer-simulated virtual environment. In the process of achieving this goal, giving virtual humans richer expressions is a very important step. Rich facial expressions can greatly enhance the realism and expressiveness of virtual humans, give users a more immersive experience, and make interactive behaviors more meaningful.

[0003] At present, in most virtual reality applications, the expression modeling and animation creation of 3D virtual humans need to be done manually by 3D animation designers. Specifically, designers need to first create a set of fusion deformation expression base models based on model deformation for 3D virtual humans to express facial muscle movements. In the process of creating expression animation, designers need to manually set a set of fusion deformation coefficients, and then repeatedly observe, evaluate and adjust the expression effects corresponding to this set of coefficients. This trial-and-error method consumes a lot of time and energy, and the creation efficiency is extremely limited. Among them, blendshape is a technology that deforms a single mesh of a 3D model to achieve a large number of predefined shapes and any number of combinations.

[0004] Facial expression transfer technology can transfer existing expressions from an existing model to a new model and create new facial expression animations for the model. This technology reduces the time cost of making expression animations for new models, greatly improves production efficiency, and provides new ideas for the synthesis of highly realistic facial expression animations. Compared with three-dimensional facial expressions, two-dimensional facial expression images have the characteristics of wide distribution and rich resources. Migrating two-dimensional facial expressions to three-dimensional facial models is a very practical technology, but also a very challenging problem.

[0005] At present, there are relatively few inventions and research related to the transfer of facial expressions from two-dimensional to three-dimensional. A task closely related to this issue is the automatic generation of three-dimensional game characters based on two-dimensional real face images.

[0006] Chinese patent application CN201811556498.4 (CN109636886A) provides an image processing method and related device, which uses a real scene face as a first facial image, and a virtual scene face rendered based on face pinching parameters as a second facial image, and adjusts the face pinching parameters multiple times by measuring the similarity between the first facial image and the second facial image, thereby achieving the goal of face pinching parameter estimation. This method requires multiple rounds of iterative operations and is not very practical.

[0007] Chinese patent application CN201911014108.5 (CN110717977A) provides a game character face processing method and related device, which first obtains the identity information and content feature information of a real face image, predicts the target face pinching parameters based on this information, and renders a virtual image. The method optimizes by minimizing the difference in identity information and content feature information between the real face image and the rendered virtual image. The disadvantage of this method is that it often requires multiple face images to achieve a better effect.

[0008] These methods have certain inspiration and reference significance for the transfer of facial expressions from 2D to 3D, but their common disadvantage is that the identity information and expression information are coupled together during the face pinching process, making it impossible to achieve flexible facial expression transfer. However, in actual facial expression transfer applications, it is often required to fix the identity information of the 3D face model and only transfer the expression information.

[0009] Based on the above background, constructing a method and device for transferring facial expressions from two-dimensional to three-dimensional with strong operability and wide application range is an urgent problem to be solved in this field. Summary of the invention

[0010] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides a method and electronic device for migrating facial expressions from two-dimensional to three-dimensional, which uses computer vision technology and deep learning technology in the field of artificial intelligence as the technical basis, and migrates facial expressions from two-dimensional images in real scenes to virtual three-dimensional facial deformation models, greatly reducing the manual burden of designers in making virtual facial expressions and animations.

[0011] In order to achieve the above object, the technical method of the present invention includes:

[0012] A method for migrating facial expressions from two dimensions to three dimensions comprises the following steps:

[0013] Obtain a two-dimensional face image in a real interaction scene as a source expression; obtain a three-dimensional face deformation model in a virtual interaction scene as a target three-dimensional face expression;

[0014] Using a pre-trained facial three-dimensional parameter extraction model, extracting three-dimensional facial parameters from the two-dimensional facial image; the three-dimensional facial parameters include identity parameters, expression parameters, and camera parameters;

[0015] Inputting the expression parameter corresponding to the two-dimensional face image into a pre-trained fusion deformation coefficient estimation model to obtain a fusion deformation coefficient;

[0016] The fusion deformation coefficient is input into the three-dimensional modeling software to drive the three-dimensional face deformation model to obtain the target three-dimensional face expression representation, which has the same expression as the two-dimensional face image.

[0017] Optionally, the three-dimensional face morphing model includes a set of fused morphing expression base models capable of expressing basic facial muscle movements.

[0018] Optionally, the face 3D parameter extraction model is a neural network model, which is trained in the following way:

[0019] Collecting a real sample training set, where the real sample training set includes multiple face images in real interaction scenes;

[0020] Acquire a 3D reconstruction renderer, wherein the 3D reconstruction renderer performs 3D reconstruction, rendering, and projection based on 3D parameters in a differentiable manner;

[0021] The real sample training set and the three-dimensional reconstruction renderer are used to train the three-dimensional facial parameter extraction model.

[0022] Optionally, the real sample training set and the 3D reconstruction renderer are used to train the face 3D parameter extraction model, including:

[0023] Inputting a plurality of real face images in the real sample training set into the face 3D parameter extraction model to be trained to obtain the face 3D parameters corresponding to each face image;

[0024] Performing three-dimensional reconstruction according to the three-dimensional face parameters corresponding to each face image to obtain a three-dimensional face reconstruction result;

[0025] According to the three-dimensional face reconstruction result, a corresponding two-dimensional face projection image is obtained;

[0026] According to the real face image, the two-dimensional face projection image, and a preset loss function, the neural network parameters of the three-dimensional face parameter extraction model are iteratively optimized to obtain the three-dimensional face parameter extraction model.

[0027] Optionally, the fusion deformation coefficient estimation model is a neural network model, which is trained in the following way:

[0028] Collecting a virtual sample training set, wherein the virtual sample training set includes multiple groups of partially randomly generated fusion deformation coefficients, and corresponding virtual human face images and virtual human face expression parameters in a virtual interactive scene;

[0029] The virtual sample training set is used to train the fusion deformation coefficient estimation model.

[0030] Optionally, the method for obtaining the multiple groups of partially randomly generated fusion deformation coefficients in the virtual sample training set includes:

[0031] Generate multiple sets of fusion deformation coefficients completely randomly;

[0032] Formulate a deformation rule library based on prior knowledge of human physiology and muscle movement;

[0033] The deformation rule library is used to filter unreasonable fusion deformation coefficients in the multiple groups of fusion deformation coefficients generated completely randomly, so as to obtain the multiple groups of fusion deformation coefficients generated partially randomly.

[0034] Optionally, the method of acquiring the virtual face image and the virtual face expression parameters includes:

[0035] Inputting the partially randomly generated fusion deformation coefficients into the three-dimensional modeling software, driving the three-dimensional face deformation model, and using the rendering function of the three-dimensional modeling software to obtain the virtual face image;

[0036] The virtual human face image is input into the pre-trained human face three-dimensional parameter extraction model to obtain the virtual human face expression parameters.

[0037] Optionally, the virtual sample training set is used to train the fusion deformation coefficient estimation model, including:

[0038] Inputting the virtual human face expression parameters into the fusion deformation coefficient estimation model to be trained to obtain the estimated virtual human face fusion deformation coefficient;

[0039] According to the partially randomly generated fusion deformation coefficients, the estimated virtual face fusion deformation coefficients, and a preset loss function, the neural network parameters of the fusion deformation coefficient estimation model to be trained are iteratively optimized to obtain the fusion deformation coefficient estimation model.

[0040] A storage medium stores a computer program and data, wherein the computer program and data are configured to execute the above method when running.

[0041] An electronic device comprises a memory and a processor, wherein the memory stores a computer program and data, and the processor is configured to run the computer program and data to execute the method described above.

[0042] The beneficial effects of the present invention are:

[0043] 1. Based on computer vision technology and deep learning technology in the field of artificial intelligence, this invention transfers facial expressions from two-dimensional images in real scenes to three-dimensional facial deformation models in virtual scenes, greatly improving the work efficiency of animators in the process of three-dimensional face modeling and creation;

[0044] 2. The present invention has no special format or quantity requirements for the source expression, and only requires a single two-dimensional face image. Considering the universality of two-dimensional face images, this feature greatly improves the practicality of the method and broadens the scope of application;

[0045] 3. The present invention uses the three-dimensional facial parameter extraction to parameterize the two-dimensional facial image in three-dimensional space, which effectively alleviates the cross-dimensional problem of facial expression from two-dimensional space to three-dimensional space;

[0046] 4. The present invention can effectively decouple information such as identity, expression and posture by extracting three-dimensional facial parameters, so that the model can focus on facial expressions and achieve more accurate and effective expression migration. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a structural schematic diagram of a two-dimensional to three-dimensional facial expression migration method of the present invention.

[0048] Figure 2 Is the construction Figure 1 Detailed process diagram of the 3D facial parameter extraction model.

[0049] Figure 3 This is a detailed flowchart for obtaining a virtual sample training set.

[0050] Figure 4 Is the construction Figure 1 A detailed flowchart of the fusion deformation coefficient estimation model.

[0051] Figure 5 It is an effect diagram of facial expression migration from two-dimensional to three-dimensional achieved by the present invention. DETAILED DESCRIPTION

[0052] The present invention will be described in detail below in conjunction with the accompanying drawings. It should be noted that the described embodiments are only for the purpose of illustration and are not intended to limit the scope of the present invention.

[0053] The present invention discloses a method for migrating facial expressions from two dimensions to three dimensions. Figure 1 shown.

[0054] This method can predict and estimate a set of fusion deformation coefficients based on a 2D face image in a real interactive scene as the source expression representation, so that it can reproduce the source expression in a 3D face deformation model in a virtual scene. This method has no special requirements on the format and quantity of the source expression, and only requires a single 2D face image to complete the facial expression migration task.

[0055] In one embodiment of the present invention, the real scene is a scene of the real world, the two-dimensional face image is a facial image taken of any person in the real world, and the virtual scene is a virtual three-dimensional space in a computer, such as a space in a three-dimensional modeling software. The three-dimensional face deformation model includes 50 fusion deformation expression base models for expressing basic facial muscle movements. After obtaining the two-dimensional face image of the real scene, a pre-trained three-dimensional face parameter extraction model is used to extract the three-dimensional face parameters of the two-dimensional face image. Afterwards, the expression parameters are input into the fusion deformation coefficient estimation model to obtain the target expression fusion deformation coefficient. Finally, the target fusion deformation coefficient is input into the three-dimensional modeling software to drive the three-dimensional face deformation model to obtain the target three-dimensional face expression representation, which has an expression consistent with the two-dimensional face image.

[0056] 1. A model for extracting 3D face parameters

[0057] The model aims to parametrically represent a two-dimensional face image in three-dimensional space. The specific structure of the model is Figure 2 Displayed in.

[0058] In this embodiment, the face 3D parameter extraction model is implemented using a deep residual neural network, and the face 3D parameters include identity parameters, expression parameters, and camera parameters. The identity parameters can be further divided into shape and texture, the expression parameters are the focus of this embodiment, and the camera parameters can be further divided into angle and illumination. This series of parameters can be used as a parametric representation of a 2D face image in 3D space. Based on these parameters, the face can be reconstructed and rendered in 3D.

[0059] In this embodiment, the weights in the deep residual neural network of the face three-dimensional parameter extraction model are the items to be optimized. The optimization process of the model is that, based on a large-scale real scene face image data set, the model first predicts the face three-dimensional parameters of the face image, and then reconstructs the face in three dimensions based on the face three-dimensional parameters. After that, a two-dimensional face image is obtained from the result of the face three-dimensional reconstruction in a projection manner. By performing a contrast measurement on the similarity between the real scene face image and the reconstructed projected face image, a joint loss is calculated, and then the loss is converted into a gradient vector through a gradient back-propagation algorithm and back-propagated back to each layer of the model. The weights in the model adjust their own values ​​based on the gradient vectors that are back-propagated.

[0060] In this embodiment, the loss function includes three items:

[0061] 1) Image loss: Lphoto = || II ′ ||2

[0062] 2) Key point loss:

[0063] 3) Perceptual loss:

[0064] Among them, I and I ′ Denote the real scene face image and the reconstructed projection face image, qn and q ′ n They represent the facial key points of the real scene face image and the reconstructed projection face image respectively, N represents the number of key points, f represents the face image perception network model, <> represents the vector inner product, ||||2 represents the mean square error, and |||| represents the vector modulus.

[0065] In this embodiment, the value of N is 68, and f adopts the VGG deep neural network pre-trained based on large-scale face images.

[0066] In this embodiment, the overall loss function is calculated by adding each loss function with weights. The weights of the three items are 1.9, 1.6e-3, and 0.2 respectively.

[0067] 2. Construct a virtual sample training set

[0068] For 3D face morphing models with different fusion deformation topology structures, it is necessary to build a personalized virtual sample training set for it. The specific process of this part is described in Figure 3 Displayed in.

[0069] In this embodiment, the virtual sample training set includes multiple groups of partially randomly generated fusion deformation coefficients, and corresponding virtual human face images and virtual human face expression parameters in the virtual interaction scene.

[0070] In this embodiment, in the process of obtaining the partially randomly generated fusion deformation coefficient, for each sample, a value between 0 and 1 is first randomly generated for each fusion deformation in the sample, that is, a completely randomly generated fusion deformation coefficient is obtained. However, this generation method often leads to some illegal value combinations. For example, it is almost impossible for a person's mouth corners to move both up and down at the same time, which means that the coefficients of the two fusion deformations of the mouth corner up and the mouth corner down should not have relatively large values ​​at the same time. For a personalized three-dimensional face deformation model, based on human prior knowledge and the topological structure of the specific fusion deformation in the model, a set of deformation rules such as this can be formulated to construct a deformation rule library for filtering illegal value combinations. After screening the deformation rule library, a set of reasonable fusion deformation coefficients can be obtained, which are the partially randomly generated fusion deformation coefficients.

[0071] In this embodiment, with respect to obtaining a virtual facial image, a camera facing the three-dimensional facial deformation model is used in the three-dimensional modeling software (Maya). Based on the above-mentioned partially randomly generated fusion deformation coefficients, the three-dimensional facial deformation model is driven to generate corresponding expressions, and the built-in rendering function of Maya is used to obtain the virtual facial image.

[0072] In this embodiment, regarding obtaining virtual facial expression parameters, based on a trained facial three-dimensional parameter extraction model, the facial three-dimensional parameters are extracted from the virtual facial image, and the expression parameters therein are saved.

[0073] In this embodiment, the virtual sample training set includes the above three parts. In a specific application, each sample can be regarded as a triplet, that is, each sample includes a set of fusion deformation coefficients, a virtual face image, and a set of virtual face expression parameters.

[0074] 3. A Fusion Deformation Coefficient Estimation Model

[0075] The model transforms the expression parameters in the 3D facial parameters into the target expression fusion deformation coefficients of the 3D facial deformation model in the virtual scene. The details of the model are Figure 4 Displayed in.

[0076] In this embodiment, the fusion deformation coefficient estimation model is implemented using a self-designed neural network. Specifically, it contains two fully connected layers and an activation function layer to implement nonlinear transformation operations. In addition, considering that the value of the fusion deformation coefficient must be between 0 and 1, the last layer of the model additionally sets a truncation layer to truncate the values ​​outside the range of the previous layer output. For example, if the value output by the second fully connected layer is 1.1, then in the truncation layer, the value will be set to 1.0. Such an operation ensures that the fusion deformation coefficient output by the entire model is legal within the numerical range.

[0077] In this embodiment, the weight of the fully connected layer in the fusion deformation coefficient estimation model is the item to be optimized. The training optimization of the model is performed based on the virtual sample training set constructed above. The fusion deformation coefficient estimation model aims to establish a mapping relationship between the virtual face expression parameters and the fusion deformation coefficients in the data set. The optimization process is to obtain the loss amount by calculating the difference between the fusion deformation coefficients predicted and estimated by the calculation model and the original partially randomly generated fusion deformation coefficients in the data set, and then optimize the weights in the fusion deformation coefficient estimation model through the gradient back propagation algorithm.

[0078] In this embodiment, the loss function of the model is the mean square error between the two sets of fused deformation coefficients.

[0079] It can be seen from the above embodiments that this embodiment first maps the two-dimensional face image in the real scene into the three-dimensional space, uses the three-dimensional face parameters for parameterized representation, and then transforms the expression parameters therein to the target expression fusion deformation coefficients of the three-dimensional face deformation model, and then drives the three-dimensional face deformation model to generate expressions based on the fusion deformation coefficients estimated by the model, thereby realizing the migration of facial expressions from two-dimensional to three-dimensional.

[0080] The experimental data of this application includes two parts. Table 1 is the error results of the present invention in this embodiment, and the comparison between the results brought by different model structures. The numerical value represents the mean absolute error between the fusion deformation coefficient during prediction and the fusion deformation coefficient during generation, reflecting the accuracy of the method for predicting the fusion deformation coefficient of three-dimensional facial expressions. As shown in Table 1, the present invention explores a set of settings with lower errors by comparing different network structures and hyperparameter settings, that is, using both the LeakyReLU activation function and the Clamp truncation layer in the model.

[0081] Table 1

[0082] Network structure Number of hidden layer nodes Mean absolute error Linear->ReLU->Linear 256 0.09 Linear->ReLU->Linear 100 0.10 Linear->ReLU->Linear 384 0.09 Linear->LeakyReLU->Linear 256 0.09 Linear->ReLU->Linear->Clamp 256 0.08 Linear->LeakyReLU->Linear->Clamp 256 0.07

[0083] Figure 5The figure is a rendering of the quantitative experimental results of the present invention. The picture is divided into three columns. In each column, the left side is a face image collected in a real scene, and the right side is the result of facial expression migration from two dimensions to three dimensions based on the method proposed by the present invention. By comparison, it can be observed that the three-dimensional face deformation model on the right after expression migration can vividly reproduce the expression on the real face on the left, regardless of gender, age, or posture, which verifies the effectiveness of the method proposed by the present invention and its strong robustness.

[0084] Another embodiment of the present invention provides a storage medium (such as ROM / RAM, magnetic disk, optical disk), in which a computer program and data are stored, wherein the computer program and data are configured to execute the above-mentioned method when running.

[0085] Another embodiment of the present invention provides an electronic device (computer, server, smart phone, etc.), including a memory and a processor, wherein the memory stores computer programs and data, and the processor is configured to run the computer programs and data to execute the method described above.

[0086] The above implementation is only used to illustrate the technical solution of the present invention rather than to limit it. Ordinary technicians in this field can modify or replace the technical solution of the present invention with equivalents without departing from the spirit and scope of the present invention. The scope of protection of the present invention shall be based on the claims.

Claims

1. A method for migrating facial expressions from two-dimensional to three-dimensional, characterized in that: The following steps are involved: Obtain a two-dimensional face image in a real interaction scenario as the source expression representation; Obtaining a 3D face deformation model in a virtual interaction scene as a target 3D face expression representation; Using a pre-trained three-dimensional face parameter extraction model to extract three-dimensional face parameters from the two-dimensional face image; the three-dimensional face parameters include identity parameters, expression parameters, and camera parameters; Inputting the expression parameter corresponding to the two-dimensional face image into a pre-trained fusion deformation coefficient estimation model to obtain a fusion deformation coefficient; Inputting the fusion deformation coefficient into the three-dimensional modeling software, driving the three-dimensional face deformation model, and obtaining a target three-dimensional face expression representation having an expression consistent with the two-dimensional face image; The fusion deformation coefficient estimation model is a neural network model, which is trained in the following way: Collecting a virtual sample training set, wherein the virtual sample training set includes multiple groups of partially randomly generated fusion deformation coefficients, and corresponding virtual human face images and virtual human face expression parameters in a virtual interactive scene; Using the virtual sample training set, training to obtain the fusion deformation coefficient estimation model; The virtual sample training set is used to train the fusion deformation coefficient estimation model, including: Inputting the virtual human face expression parameters into the fusion deformation coefficient estimation model to be trained to obtain the estimated virtual human face fusion deformation coefficient; Iteratively optimizing the neural network parameters of the fusion deformation coefficient estimation model to be trained according to the partially randomly generated fusion deformation coefficients, the estimated virtual face fusion deformation coefficients, and a preset loss function to obtain the fusion deformation coefficient estimation model; The fusion deformation coefficient estimation model includes two fully connected layers and an activation function layer, which are used to implement nonlinear transformation operations. Considering that the value of the fusion deformation coefficient must be between 0 and 1, a truncation layer is additionally set in the last layer of the model to truncate the values ​​outside the range of the output of the previous layer; the loss amount is obtained by calculating the difference between the fusion deformation coefficient predicted and estimated by the model and the original partially randomly generated fusion deformation coefficient in the data set, and then the weights in the fusion deformation coefficient estimation model are optimized through the gradient back propagation algorithm.

2. The method according to claim 1, characterized in that The three-dimensional face deformation model includes a group of fused deformation expression base models that can express basic facial muscle movements.

3. The method according to claim 1, characterized in that The face 3D parameter extraction model is a neural network model, which is trained in the following way: Collecting a real sample training set, wherein the real sample training set includes multiple face images in real interaction scenes; Acquire a 3D reconstruction renderer, wherein the 3D reconstruction renderer can perform 3D reconstruction, rendering, and projection based on 3D parameters in a differentiable manner; The real sample training set and the three-dimensional reconstruction renderer are used to train and obtain the three-dimensional face parameter extraction model.

4. The method according to claim 3, characterized in that The real sample training set and the 3D reconstruction renderer are used to train the face 3D parameter extraction model, including: Inputting a plurality of real face images in the real sample training set into the face 3D parameter extraction model to be trained to obtain the face 3D parameters corresponding to each face image; Performing three-dimensional reconstruction according to the three-dimensional face parameters corresponding to each face image to obtain a three-dimensional face reconstruction result; According to the three-dimensional face reconstruction result, a corresponding two-dimensional face projection image is obtained; According to the real face image, the two-dimensional face projection image, and a preset loss function, the neural network parameters of the three-dimensional face parameter extraction model are iteratively optimized to obtain the three-dimensional face parameter extraction model.

5. The method according to claim 1, characterized in that The method for obtaining the multiple groups of partially randomly generated fusion deformation coefficients in the virtual sample training set includes: Generate multiple sets of fusion deformation coefficients completely randomly; Formulate a deformation rule library based on prior knowledge of human physiology and muscle movement; The deformation rule library is used to filter unreasonable fusion deformation coefficients in the multiple groups of fusion deformation coefficients generated completely randomly, so as to obtain the multiple groups of fusion deformation coefficients generated partially randomly.

6. The method according to claim 1, characterized in that The method of obtaining the virtual face image and the virtual face expression parameters includes: Inputting the partially randomly generated fusion deformation coefficients into the three-dimensional modeling software, driving the three-dimensional face deformation model, and using the rendering function of the three-dimensional modeling software to obtain the virtual face image; The virtual human face image is input into the pre-trained human face three-dimensional parameter extraction model to obtain the virtual human face expression parameters.

7. A storage medium storing a computer program and data, wherein: The computer program and data are arranged to execute the method of any one of claims 1-6 when run.

8. An electronic device, comprising a memory and a processor, wherein the memory stores a computer program and data, and the processor is configured to run the computer program and data to perform the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image processing method and device, storage medium and electronic device

    CN109636886A

  • Image processing methods, apparatus, storage media, and electronic devices

    CN109636886B

  • Game role face processing method and device, computer equipment and storage medium

    CN110717977A

  • Methods, devices, computer equipment, and storage media for processing game character faces

    CN110717977B