A method for driving the expression generation of a digital human

Through the feature fusion technology of the three-dimensional face reconstruction model and the expression generation model, the problem of rigid digital human expressions is solved, and natural and coherent realistic expressions are generated, which improves the human-computer interaction experience.

CN119379872BActive Publication Date: 2025-07-04NANJING SILICON INTELLIGENCE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411941998.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-07-04
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

The prior art is difficult to generate realistic, subtle, exaggerated, asymmetrical or continuous digital human expressions, resulting in insufficient emotional communication and reducing human-computer interaction experience.

Method used

By obtaining the driver video and the driven video, the target three-dimensional face reconstruction model and expression generation model are used for feature extraction and analysis, and feature fusion is combined with the adaptive attention normalization module and the residual network to generate realistic digital human expressions.

Benefits of technology

It realizes the natural, coherent and real expression of digital human expressions, enhances the immersion and emotional resonance of human-computer interaction, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119379872B_ABST
    Figure CN119379872B_ABST
Patent Text Reader

Abstract

The present application provides a method for driving digital human expression generation, which relates to the field of digital human generation technology. The generation method includes obtaining a driving video and a driven video; the driving video is the expression provider of the driven video; inputting the driving video into a target three-dimensional face reconstruction model to obtain target facial coefficient features; inputting the target facial coefficient features and the driven video into a target expression generation model to obtain a driven image with a target expression; performing video encoding based on the driven image to obtain a target video; the target video is a video that combines the face information of the driven video and the target expression of the driving video. The present application solves the problem that the driven digital human expression generated based on the prior art is too rigid through the above generation method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital human generation, and particularly relates to a method for driving digital human expression generation. Background Art

[0002] The research on digital human expressions is an important part of the field of natural human-computer interaction, aiming to accurately apply facial expressions to digital humans, enabling digital humans to present realistic emotional expressions, enhancing the user experience, promoting emotional resonance, and bringing higher immersion and authenticity to human-computer interaction. Facial coefficient-driven digital human expressions encode by obtaining the facial expression coefficients in the driving video to drive the digital human to generate expressions.

[0003] Currently, the research on driving digital human expressions through facial coefficients mainly focuses on the optimization of expression effects. Most facial coefficient-driven digital human technologies can have good transfer effects in simple expressions such as smiling and being happy or in face reconstruction, but they cannot achieve the transfer of subtle, exaggerated, asymmetric, or continuous expression changes. In many application scenarios, the unrealistic digital human expressions will lead to insufficient emotional communication, reduce the interaction experience, result in poor communication effects, and thus affect the user's willingness to participate and overall satisfaction. Currently, there is no method in the related technologies that can achieve the generation of digital human expressions based on facial coefficient driving, which is the bottleneck that needs to be solved for facial coefficient-driven digital human expressions. Summary of the Invention

[0004] This application provides a method for driving digital human expression generation to solve the problem that the digital human expressions generated based on the existing technology are too rigid.

[0005] The generation method includes:

[0006] Obtain a driving video and a driven video; the driving video is the expression provider of the driven video;

[0007] Input the driving video into a target three-dimensional face reconstruction model to obtain target facial coefficient features; the target three-dimensional face reconstruction model extracts features from the driving video based on the skip connection method;

[0008] Input the target facial coefficient features and the driven video into a target expression generation model, and the target expression generation model is configured to perform global feature analysis and local feature analysis based on the target facial coefficient features and the driven video to obtain a driven image with a target expression;

[0009] Perform video encoding based on the driven image to obtain a target video; the target video is a video that combines the face information of the driven video and the target expression of the driving video.

[0010] Preferably, the training process of the target 3D face reconstruction model includes:

[0011] Obtain an expression video, and perform a first preprocessing on the expression video to obtain a target expression video;

[0012] Perform model training and convergence according to the target expression video and the 3D face reconstruction model to obtain the target 3D face reconstruction model.

[0013] Preferably, the step of performing model training according to the target expression video and the 3D face reconstruction model includes:

[0014] Select several frames of training expression images from the target expression video;

[0015] Perform a first processing on the training expression images to obtain facial coefficient features;

[0016] Repeat the training of the 3D face reconstruction model, and use a first loss function to converge the facial coefficient features until the 3D face reconstruction model meets the first requirement to obtain the target 3D face reconstruction model.

[0017] Preferably, the step of performing a first processing on the training expression images includes:

[0018] Perform a reconstruction process and a masking process on the training expression images respectively to obtain a first reconstruction feature and a first masked image;

[0019] Perform a skip feature fusion process according to the first reconstruction feature and the first masked image to obtain a first fusion feature;

[0020] Perform a decoding process on the first fusion feature to obtain the facial coefficient features.

[0021] Preferably, the step of performing a reconstruction process on the training expression images includes:

[0022] Perform an encoding process on the training expression images to obtain a first encoded feature;

[0023] Perform 3D face mesh reconstruction on the first encoded feature to obtain a first reconstruction feature.

[0024] Preferably, the step of using a first loss function to converge the facial coefficient features includes:

[0025] Use the L1 loss function to calculate a first error between the facial coefficient features and the training expression images;

[0026] Calculate a second error between the facial coefficient feature and the training expression image by using a cycle loss function;

[0027] When both the first error and the second error meet a first requirement, stop training the 3D face reconstruction model to obtain the target 3D face reconstruction model.

[0028] Preferably, the target 3D face reconstruction model includes a first convolutional encoding unit, an image masking unit, a 3D face mesh reconstruction unit, a feature fusion unit, and a second convolutional encoding unit;

[0029] The feature fusion unit includes an encoding layer and a decoding layer. The encoding layer includes a first encoding layer, a second encoding layer, and a third encoding layer connected in sequence. The decoding layer includes a first decoding layer, a second decoding layer, and a third decoding layer connected in sequence. The third encoding layer is connected to the first decoding layer; there are skip connections between the encoding layer and the decoding layer;

[0030] The first encoding layer, the second encoding layer, and the third encoding layer are all configured to downsample the features input by the previous layer; the first decoding layer, the second decoding layer, and the third decoding layer are all configured to upsample the features input by the previous layer.

[0031] Preferably, the training process of the target expression generation model includes:

[0032] Obtain a training driving video, and perform a second preprocessing on the training driving video to obtain a target training driving video;

[0033] Perform model training and convergence according to the target training driving video, the facial coefficient feature, and the expression generation model to obtain the target expression generation model.

[0034] Preferably, the step of performing model training according to the target training driving video, the facial coefficient feature, and the expression generation model includes:

[0035] Perform convolutional encoding on the target training driving video and the facial coefficient feature respectively to obtain a second encoded feature and a third encoded feature;

[0036] Fuse the second encoded feature and the third encoded feature by using a second process to obtain a second fused feature; the second process includes an adaptive normalization process with an attention mechanism;

[0037] Perform residual calculation on the second fused feature to obtain a first residual feature;

[0038] Perform deconvolution processing on the first residual feature to obtain a first deconvolution feature;

[0039] Perform a third process on the first deconvolution feature to obtain a training driven image; the third process includes extracting deep features in the first deconvolution feature by using spatial normalization processing and uniformly and adaptively adjusting the normalization parameters of the spatial normalization processing.

[0040] Perform video encoding according to the training driven image to obtain a training target video.

[0041] Use the L1 loss function to converge the training target video until the expression generation model meets the second requirement to obtain the target expression generation model.

[0042] Preferably, the second process includes:

[0043] Divide the second encoded feature and the third encoded feature into multiple subsequences, perform feature encoding processing on each subsequence, calculate the attention weights corresponding to different subsequences through a spatial attention module, and fuse the results of multiple feature encoding processes based on the attention weight calculation results to obtain a local feature encoding; the spatial attention module is preset in the expression generation model.

[0044] Divide the local feature encoding into multiple subsequences, perform feature encoding processing on each subsequence, calculate the attention weights corresponding to different subsequences through the spatial attention module, and fuse the results of multiple feature encoding processes based on the attention weight calculation results to obtain a local feature encoding.

[0045] Iteratively update and independently output the local feature encoding obtained in each round to obtain multiple local feature encodings.

[0046] Calculate the attention weights for multiple local feature encodings according to a channel attention module, and perform feature processing based on the attention weight calculation results to obtain a global feature encoding; fuse multiple local feature encodings and the global feature encoding to obtain the second fusion feature; the channel attention module is preset in the expression generation model.

[0047] Preferably, the step of using the L1 loss function to converge the training generated video includes:

[0048] Use the L1 loss function to calculate the third error between the training target video and the target expression video.

[0049] When the third error meets the second requirement, stop training the expression generation model to obtain the target expression generation model.

[0050] As can be seen from the above, the present application provides a method for driving digital human expression generation. The generation method includes obtaining a driving video and a driven video; the driving video is the expression provider of the driven video; inputting the driving video into a target three-dimensional face reconstruction model to obtain target facial coefficient features; inputting the target facial coefficient features and the driven video into a target expression generation model to obtain a driven image with a target expression; performing video encoding based on the driven image to obtain a target video; the target video is a video combining the face information of the driven video and the target expression of the driving video. The present application solves the problem that the driven digital human expressions generated based on the prior art are too rigid through the above generation method. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0052] Figure 1 It is a flowchart of a method for driving digital human expression generation according to the present application;

[0053] Figure 2 It is a training flowchart of a three-dimensional face reconstruction model in a method for driving digital human expression generation according to the present application;

[0054] Figure 3 It is a specific training flowchart of a three-dimensional face reconstruction model in a method for driving digital human expression generation according to the present application;

[0055] Figure 4 It is a flowchart of obtaining facial coefficient features in a method for driving digital human expression generation according to the present application;

[0056] Figure 5 It is a flowchart of the convergence mode of a three-dimensional face reconstruction model in a method for driving digital human expression generation according to the present application;

[0057] Figure 6 It is a schematic diagram of a target three-dimensional face reconstruction model in a method for driving digital human expression generation according to the present application;

[0058] Figure 7 It is a schematic diagram of a feature fusion unit in a target three-dimensional face reconstruction model;

[0059] Figure 8 It is a training flowchart of an expression generation model in a method for driving digital human expression generation according to the present application;

[0060] Figure 9This is the specific flowchart for training the expression generation model in a method for driving digital human expression generation in this application;

[0061] Figure 10 This is the flowchart of the convergence method of the expression generation model in a method for driving digital human expression generation in this application;

[0062] Figure 11 This is the comparison diagram of the human face before and after occlusion processing in a method for driving digital human expression generation in this application;

[0063] Figure 12 This is the schematic diagram of the expression generation model in a method for driving digital human expression generation in this application;

[0064] Figure 13 This is the flowchart of the second processing in a method for driving digital human expression generation in this application. Detailed implementation mode

[0065] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0066] Figure 1 This is the flowchart of a method for driving digital human expression generation in this application.

[0067] See Figure 1 It can be seen that this embodiment provides a method for driving digital human expression generation, and the generation method includes:

[0068] S10. Obtain a driving video and a driven video. Specifically, in this embodiment, before generating a digital human, it is necessary to first obtain the driving object and the driven object of the digital human, where the driving video is the expression providing object of the driven video.

[0069] Among them, the driving video is a recorded video of an actual user, and the driven video can be a virtual anchor image video.

[0070] It should be noted that the driven video can only be obtained after permission. It can be understood that some virtual character image videos have usage permissions, so they must be used after obtaining permission.

[0071] The generation method further includes:

[0072] S20. Input the driving video into the target 3D face reconstruction model. Specifically, in this embodiment, before generating the expression video, it is necessary to first perform video reconstruction on the driving video to improve the accuracy and detail richness of the facial expression images in the driving video.

[0073] Among them, in this embodiment, by inputting the driving video into the target 3D face reconstruction model for face reconstruction processing, the features after face reconstruction of the corresponding face, that is, the target facial coefficient features, are obtained.

[0074] The generation method further includes:

[0075] S30. Input the target facial coefficient features and the driven video into the target expression generation model. Specifically, in this embodiment, in the conventional driving digital human expression generation technology, generally, the expression provider and the expression recipient are directly subjected to expression transfer, which often results in problems such as incoherence and disharmony in the generated expression; in step S20 above, the face of the expression provider, that is, the driving video, has been reconstructed. If you want the generated expression to be more natural, it is necessary to perform certain processing on the expression recipient as well, and perform corresponding facial processing on both the expression provider and the expression recipient, so as to further improve the authenticity of the generated expression.

[0076] Among them, in this embodiment, since it is necessary to transfer the expression from one video object to another video object, it is not possible to simply perform corresponding processing on the video object; for this, in this embodiment, the target facial coefficient features obtained in step S20 and the driven video for receiving the expression are input into the target expression generation model for expression transfer, so as to obtain the image after expression transfer, that is, the driven image with the target expression.

[0077] The generation method further includes:

[0078] S40. Perform video encoding based on the driven image. Specifically, in this embodiment, in step S30, the driven image with the transferred expression has been obtained, and a single image cannot form a complete video. Therefore, it is necessary to perform video encoding on the driven image to convert the image into a video and obtain the target video.

[0079] Among them, the target video is a video combining the face information of the driven video and the target expression of the driving video.

[0080] Figure 2 This is the training flowchart of the 3D face reconstruction model in a method for generating driving digital human expressions of this application.

[0081] SeeFigure 2 It can be seen that, further, in some embodiments, the training process of the target three-dimensional face reconstruction model includes:

[0082] S100, obtain an expression video and perform a first preprocessing on the expression video. Specifically, in this embodiment, since the target three-dimensional face reconstruction model is used to perform expression reconstruction and feature extraction on the expression provider, therefore, to train the target three-dimensional face reconstruction model, the corresponding expression video needs to be obtained. Although videos include various types, for the target three-dimensional face reconstruction model, the requirements of complete and clear expressions need to be met.

[0083] Among them, since the expression videos obtained for training are generally obtained from open-source video databases, and the quality of such videos does not perfectly meet the requirements of the training data required by the target three-dimensional face reconstruction model, therefore, the first preprocessing needs to be performed on the expression videos to obtain videos that meet the training requirements of the target three-dimensional face reconstruction model.

[0084] Among them, the first preprocessing may include data augmentation operations such as face detection, key point detection, random brightness, contrast, color jitter, Gaussian noise, and blur processing according to different situations. It should be noted that during the training process of the target three-dimensional face reconstruction model, preprocessing needs to be performed on the expression video objects of the expression provider. However, in the actual application of the target three-dimensional face reconstruction model, the expression video objects do not need to be preprocessed. Preprocessing the data during the training stage is to better train the model, and each step in the first preprocessing is not a necessary condition for model training.

[0085] After performing the first preprocessing on the expression video, a preprocessed target expression video is obtained.

[0086] S200, perform model training and convergence according to the target expression video and the three-dimensional face reconstruction model. Specifically, in this embodiment, during the process of using the expression video to perform model training on the three-dimensional face reconstruction model, loss convergence needs to be performed on the three-dimensional face reconstruction model to obtain the target three-dimensional face reconstruction model that meets the target requirements.

[0087] Figure 3 This is the specific flowchart of the training of the three-dimensional face reconstruction model in a method for driving digital human expression generation according to the present application.

[0088] See Figure 3 It can be seen that, further, in some embodiments, the step of performing model training according to the target expression video and the three-dimensional face reconstruction model includes:

[0089] S210, Select a number of training expression images from the target expression video. Specifically, in this embodiment, during the training process of the 3D face reconstruction model, a certain number of image frames need to be selected from the target expression video and used for training based on these image frames.

[0090] The steps of training the model according to the target expression video and the 3D face reconstruction model further include:

[0091] S220, Perform a first process on the training expression images. Specifically, in this embodiment, input the target expression video into the 3D face reconstruction model, and use the 3D face reconstruction model to perform a first process on a number of image frames of the target expression video to obtain the facial coefficient features corresponding to the face in the target expression video.

[0092] The steps of training the model according to the target expression video and the 3D face reconstruction model further include:

[0093] S230, Repeat the training of the 3D face reconstruction model and use the first loss function to converge the facial coefficient features until the 3D face reconstruction model meets the first requirement. Specifically, in this embodiment, the facial coefficient features obtained in step S220 above are the output of the model. However, if you want to train the model, you need to continuously converge the loss of the facial coefficient features to continuously improve the 3D face reconstruction model.

[0094] Among them, during the training process of the 3D face reconstruction model, use the first loss function to calculate the loss of the facial coefficient features output each time, so as to achieve the convergence of the 3D face reconstruction model and finally obtain the target 3D face reconstruction model.

[0095] Figure 4 This is a flowchart for obtaining facial coefficient features in a method for driving digital human expression generation in this application.

[0096] See Figure 4 It can be seen that, further, in some embodiments, the steps of performing a first process on the training expression images include:

[0097] S221. Reconstruct and mask the training facial expression images respectively. Specifically, in this embodiment, the first step of the first process includes reconstructing and masking the training facial expression images respectively. It can be understood that the training facial expression images are divided into two sequences, one sequence is reconstructed, and the other sequence is masked. For the training facial expression images undergoing reconstruction, the first reconstruction features are obtained. For the training facial expression images undergoing masking, the first masked images are obtained.

[0098] Among them, the reconstruction process includes encoding the training facial expression images to obtain the first encoded features, and performing 3D face mesh reconstruction on the first encoded features to obtain the first reconstruction features.

[0099] For the images after the masking process and the images before the masking process, reference can be made to Figure 11 , reference Figure 11 It can be seen that the masking process masks the relatively important regions of the face in the face image. The reason is that the acquisition of face features needs to meet the features of the entire face. Therefore, it is necessary to blacken the important regions of the features of the entire face, so as to provide an effective data basis for subsequent feature fusion.

[0100] The steps of the first process according to the training facial expression images further include:

[0101] S222. Perform skip feature fusion processing based on the first reconstruction features and the first masked images. Specifically, in this embodiment, the first reconstruction features and the first masked images obtained in step S221 are subjected to feature fusion processing. This feature fusion processing is skip fusion processing, and its principle is based on the neural network structure with shortcut connections.

[0102] Among them, by performing skip feature fusion processing on the first reconstruction features and the first masked images, the first fusion features are obtained.

[0103] The steps of the first process according to the training facial expression images further include:

[0104] S223. Encode the first fusion features. Specifically, in this embodiment, after completing step S222, the first fusion features obtained in step S222 are encoded to obtain the corresponding features, that is, the facial coefficient features.

[0105] It should be noted that although the above steps S221 to S223 are processes in model training, they are the same as steps S221 to S223 in the actual application of the model.

[0106] Figure 5 This is a flowchart of the convergence method of the 3D face reconstruction model in a method for driving digital human expression generation in this application.

[0107] Referring to Figure 5 It can be seen that, further, in some embodiments, the step of using the first loss function to converge the facial coefficient features includes:

[0108] S231, calculating the first error between the facial coefficient features and the training expression image by using the L1 loss function;

[0109] S232, calculating the second error between the facial coefficient features and the training expression image by using the cycle loss function;

[0110] Specifically, in this embodiment, by simultaneously using the L1 loss function and the cycle loss function to calculate the error between the facial coefficient features and the training expression image, the first error and the second error are obtained.

[0111] The step of using the first loss function to converge the facial coefficient features further includes:

[0112] S233, when both the first error and the second error meet the first requirement, stop training the 3D face reconstruction model. Specifically, in this embodiment, when both the first error and the second error meet the first requirement, it means that the 3D face reconstruction model has completed training. Therefore, the training can be stopped, and the target 3D face reconstruction model can be obtained.

[0113] Among them, we use the L1 loss function to calculate the L1 reconstruction error between the facial coefficient features and the training expression image. In order to make the predicted facial coefficients more accurate and stable, we add a cycle loss function, that is, the cycle loss between the facial coefficient features and the facial coefficients of the training expression image.

[0114] Figure 8 This is a flowchart of the training of the expression generation model in a method for driving digital human expression generation in this application.

[0115] Referring to Figure 8 It can be seen that, further, in some embodiments, the training process of the target expression generation model includes:

[0116] S300. Obtain the training driving video and perform a second preprocessing on the training driving video. Specifically, in this embodiment, since the target expression generation model is used to extract expression features from the expression recipient and the expression provider, the training of the target expression generation model also requires obtaining the corresponding training driving video. Since the training driving video is the same as the expression video and is diverse, the target expression generation model also needs to meet certain requirements.

[0117] Among them, the training driving video obtained for training can be obtained from an open-source video database or recorded by the training model personnel. The quality of these videos is also uneven. Therefore, it is necessary to perform a second preprocessing on the training driving video to obtain a video that meets the training requirements of the target expression generation model.

[0118] Among them, the second preprocessing is quite different from the first preprocessing. The second preprocessing may include video beautification, face detection, and cropping according to different situations. It should be noted that during the training process of the target expression generation model, preprocessing is required for the expression recipient video object of the expression. In the actual application of the target expression generation model, the expression recipient does not need to perform preprocessing. Preprocessing the data during the training stage is to better train the model, and each step in the second preprocessing is not a necessary condition for model training.

[0119] After performing the second preprocessing on the training driving video, a preprocessed target training driving video is obtained.

[0120] The training process of the target expression generation model further includes:

[0121] S400. Perform model training and convergence according to the target training driving video, the facial coefficient feature, and the expression generation model. Specifically, in this embodiment, during the process of using the target training driving video and the facial coefficient feature to perform model training on the expression generation model, it is also necessary to perform loss convergence on the expression generation model to obtain a loss target expression generation model that meets the target requirements.

[0122] Figure 9 This is a specific flowchart of the training of the expression generation model in a method for driving a digital human expression generation in this application.

[0123] See Figure 9 It can be seen that, further, in some embodiments, the step of performing model training and convergence according to the target training driving video, the facial coefficient feature, and the expression generation model includes:

[0124] S410. Convolutionally encode the target training driving video and the facial coefficient features respectively to obtain a second encoded feature and a third encoded feature. Specifically, in this embodiment, the target training driving video and the facial coefficient features are respectively convolutionally encoded to obtain the second encoded feature and the third encoded feature, where the second encoded feature corresponds to the target training driving video and the third encoded feature corresponds to the facial coefficient features.

[0125] The step of training and converging the model according to the target training driving video, the facial coefficient features and the expression generation model further includes:

[0126] S420. Feature-fuse the second encoded feature and the third encoded feature using a second process to obtain a second fused feature; the second process includes adaptive normalization processing with an attention mechanism.

[0127] S430. Calculate the residual of the second fused feature to obtain a first residual feature.

[0128] S440. Perform deconvolution processing on the first residual feature to obtain a first deconvolution feature.

[0129] S450. Perform a third process on the first deconvolution feature to obtain a trained driven image; the third process includes extracting deep features in the first deconvolution feature using spatial normalization processing and uniformly and adaptively adjusting the normalization parameters of the spatial normalization processing.

[0130] Specifically, in this embodiment, through steps S420 to S450, the encoded features of the expression provider and the encoded features of the expression receiver are processed using an improved neural network to fuse the two together and enhance the correlation between the face and the expression, thereby enhancing the authenticity of the subsequent generated expression video.

[0131] The step of training and converging the model according to the target training driving video, the facial coefficient features and the expression generation model further includes:

[0132] S460. Perform video encoding on the trained driven image to obtain a training target video.

[0133] S470. Use the L1 loss function to converge the training target video until the expression generation model meets the second requirement to obtain the target expression generation model.

[0134] Specifically, in this embodiment, after the processing of steps S420 to S450 is completed, the features are converted into a video, and the generated video is continuously subjected to loss convergence, so as to finally obtain the target expression generation model.

[0135] Figure 10 It is a flowchart of the convergence method of the expression generation model in a method for driving a digital human expression generation of the present application.

[0136] See Figure 10 It can be seen that, further, in some embodiments, the step of using the L1 loss function to converge the training generated video includes:

[0137] S471, calculating a third error between the training target video and the target expression video by using the L1 loss function;

[0138] S472, when the third error meets the second requirement, stop training the expression generation model to obtain the target expression generation model.

[0139] Specifically, in this embodiment, similar to the convergence method of the three-dimensional face reconstruction model, the loss convergence of the expression generation model calculates the third error between the training target video and the target expression video by using the L1 loss function, and determines whether the expression generation model is trained completed according to this error.

[0140] Figure 6 It is a schematic diagram of a target three-dimensional face reconstruction model in a method for driving a digital human expression generation of the present application;

[0141] Figure 7 It is a schematic diagram of a feature fusion unit in the target three-dimensional face reconstruction model.

[0142] See Figure 6 and Figure 7 It can be seen that, further, in some embodiments, the target three-dimensional face reconstruction model includes a first convolutional encoding unit, an image masking unit, a three-dimensional face mesh reconstruction unit, a feature fusion unit, and a second convolutional encoding unit;

[0143] The feature fusion unit includes an encoding layer and a decoding layer. The encoding layer includes a first encoding layer, a second encoding layer, and a third encoding layer connected in sequence. The decoding layer includes a first decoding layer, a second decoding layer, and a third decoding layer connected in sequence. The third encoding layer is connected to the first decoding layer; there are skip connections between the encoding layer and the decoding layer;

[0144] The first encoding layer, the second encoding layer, and the third encoding layer are all configured to perform downsampling on the features input from the previous layer; the first decoding layer, the second decoding layer, and the third decoding layer are all configured to perform upsampling on the features input from the previous layer.

[0145] Specifically, in this embodiment, the target 3D face reconstruction model takes the target facial coefficient feature and the driven video as data inputs, passes through the first convolutional encoding unit to obtain the face coefficients, and performs 3D face mesh reconstruction through the 3D face mesh reconstruction unit; through the feature fusion unit and the second convolutional encoding unit, the target facial coefficient feature is extracted.

[0146] The feature fusion unit consists of an improved neural network structure, including 3 encoding layers for downsampling and 3 decoding layers for upsampling, and each corresponding layer uses shortcut (skip connection) for splicing; the feature fusion unit takes the output of the image masking unit and the output of the 3D face mesh reconstruction unit as inputs, and can better capture global information such as the edge information, face texture, and expression of the human face through residuals and skip connections. Specifically, the above 3 encoding layers gradually reduce the spatial dimension of the image while increasing the number of feature channels; correspondingly, the 3 decoding layers gradually restore the spatial dimension of the image while reducing the number of feature channels; the 3 encoding layers and the 3 decoding layers show a corresponding relationship. The output of each encoding layer, in addition to being transmitted to the next encoding layer, will also be directly transmitted to the corresponding decoding layer (the data transmission relationship here can be understood as the above skip connection), thereby ensuring the fusion of information and alleviating the problem of gradient disappearance.

[0147] Figure 12 It is a schematic diagram of the expression generation model in a method for driving digital human expression generation according to this application.

[0148] Figure 13 It is a flowchart of the second processing in a method for driving digital human expression generation according to this application.

[0149] See Figure 12 and Figure 13 It can be seen that, further, in some embodiments, the second processing includes:

[0150] S421. Divide the second coding feature and the third coding feature into multiple subsequences, perform feature coding processing on each subsequence, calculate the attention weights corresponding to different subsequences through the spatial attention module, and fuse the results of multiple feature coding processes based on the attention weight calculation results to obtain local feature coding. The spatial attention module is preset in the expression generation model. Specifically, in this embodiment, through the sequence division, attention weight calculation, and local feature coding based on the attention weights performed on the second coding feature and the third coding feature in sequence, a spatial attention mechanism is introduced in the calculation process of local features. The spatial attention mechanism focuses on the important points at each spatial position in the feature map, so it has a better presentation result when calculating the attention of the local area.

[0151] Among them, the operation of dividing the second coding feature and the third coding feature into multiple subsequences further includes the following operations:

[0152] Perform sequence division on the second coding feature and the third coding feature respectively, and finally perform normalization processing on the results of both; or, first perform normalization operations on the second coding feature and the third coding feature in advance, and then perform sequence division on the fused feature.

[0153] It should be noted that the division operations of the second coding feature and the third coding feature may be different according to different application scenarios, but the preferred solution is to perform sequence division on the second coding feature and the third coding feature respectively, and finally perform normalization processing on the results of both.

[0154] The second processing further includes:

[0155] S422. Divide the local feature coding into multiple subsequences, perform feature coding processing on each subsequence, calculate the attention weights corresponding to different subsequences through the spatial attention module, and fuse the results of multiple feature coding processes based on the attention weight calculation results to obtain local feature coding. Specifically, in this embodiment, it is the same as the above step S421, both are performing local feature analysis on the feature coding, but the difference from step S421 is that after completing step S422, the obtained feature coding needs to be iteratively updated, perform cyclic local feature analysis on the updated feature coding, and independently output the obtained local feature coding after each round of local feature analysis.

[0156] The second processing further includes:

[0157] S423. Calculate the attention weights for multiple local feature encodings according to the channel attention module, and perform feature processing based on the calculation result of the attention weights to obtain the global feature encoding; fuse the multiple local feature encodings and the global feature encoding to obtain the second fused feature; the channel attention module is preset in the expression generation model. Specifically, in this embodiment, after several rounds of local feature analysis, several local feature encodings can be obtained. In step S423, further feature analysis needs to be performed on these local feature encodings. The core purpose of step S423 is to introduce the channel attention mechanism in the feature analysis process. The channel attention mechanism focuses on calculating the importance of different feature channels relative to the global, so it has a better performance in calculating the global features.

[0158] Through steps S421 to S423, the local analysis and global analysis of features are realized, so as to enhance the understanding of features by the expression generation model.

[0159] The above target 3D face reconstruction model helps to improve the accuracy and detail richness of the reconstructed face expression image; at the same time, the feature fusion unit also has other advantages:

[0160] Faster convergence speed: The introduction of the residual structure makes it easier to converge;

[0161] Stronger generalization ability: It has multi-scale feature fusion, and integrating local and global information helps to enhance the generalization ability of the model and generate more natural and reasonable images.

[0162] The training process adopts the pre-training technology and trains 20 rounds on a large data set of 1 million open-source expression video data collected and rich expression video data collected from the network, ensuring the accuracy of facial coefficient extraction and the generalization of the model;

[0163] Improve the 3D face reconstruction model through the feature fusion unit and the second convolutional encoding unit, and perform regression training on the face and facial coefficients of the face in the large data set of diverse expression videos collected, so that the model can accurately extract the facial coefficients of the expression face.

[0164] Correspondingly, the specific structure, training and functions of the target expression generation model are as follows:

[0165] During the training process of the target expression generation model, first, based on the aforementioned target three-dimensional face reconstruction model, the input driving person video is used to obtain the corresponding facial coefficients. Then, the obtained facial coefficients and the image of the driven person are used as inputs to output the corresponding expression images of the driven person, and the expression images are combined to form the final expression video of the driven person. The facial coefficients are input into a convolutional module for encoding to obtain expression features. Similarly, the image of the driven person is input into the convolutional module for encoding to obtain face features. The above expression features and face features are input into an adaptive attention normalization module for feature fusion.

[0166] The adaptive attention normalization module is an upgraded version of the attention normalization module. Conventional adaptive normalization does not consider local feature statistics and is prone to outputting unnatural expression faces and local distortions. The adaptive attention normalization module, however, dynamically adjusts the correlation between expression features and face features by introducing an attention mechanism, performs adaptive weighting and feature fusion, thereby reducing information loss.

[0167] Specifically, during the convolutional encoding process of the facial coefficients and the image of the driven person, both shallow features (i.e., the outputs of convolutional layers with a relatively forward position in the hierarchical relationship) and deep features (i.e., the outputs of convolutional layers with a relatively backward position in the hierarchical relationship) are simultaneously used as the corresponding expression features and face features and input into the adaptive attention normalization module, so that the adaptive attention normalization module can comprehensively consider the deep and shallow features of the corresponding images, and thus learn the spatial information of the expression and the face, especially being able to better capture local details.

[0168] On this basis, the adaptive attention normalization module first calculates an attention weight distribution map based on the above expression features and face features that contain spatial information. This weight distribution map is used to characterize which regions in the face features are the most critical for achieving expression transformation (such as eyes, mouth, etc. These parts are usually the most sensitive to expression changes, and different expressions will further result in different attention weights among some details in the above parts).

[0169] On this basis, the similarity between the expression features and the face features can be further measured (such as methods like cosine similarity), and based on the similarity between the expression features and the face features, the aforementioned attention weights are further adjusted.

[0170] After the above adjustments are completed, the calculated attention weights are used to weight the face features to enhance the features related to the expression features among them. On this basis, the weighted face features and the expression features are combined to generate a new feature representation, which contains both the identity information of the face features and the changes in the target expression.

[0171] The advantages of the above-mentioned adaptive attention normalization module also include:

[0172] Enhancing model capabilities: Due to the introduction of the attention mechanism, the adaptive attention normalization module can more finely control feature fusion and generation results, thereby improving the quality and details of the generated images, making the output more coherent and natural.

[0173] Improving robustness: The adaptive attention normalization module adaptively adjusts the weights of feature fusion through the attention mechanism, better focuses on important features, is more robust to input perturbations, and enhances the generalization and robustness of the model.

[0174] The above-mentioned fused features are input into the residual network module and the transposed convolution module. The residual network module learns residuals through skip connections, which can effectively avoid gradient vanishing, capture deep features, accelerate the model convergence speed, and improve the generalization ability of the model; while the transposed convolution can ensure that more effective information is retained during the upsampling process, can better restore high-resolution image details, and improve the image generation quality.

[0175] The output features of the above-mentioned transposed convolution are input into the spatial normalization module. The spatial normalization module mainly consists of 5 normalized residual layers, and the parameter adjustment of each normalized residual layer uses the spatial normalization module to learn; the normalized residual layer is similar to batch normalization and performs normalization in a channel-wise manner when activated; the main function of this module is to further extract deep features, and at the same time, adaptively adjust the normalization parameters through the spatial normalization module, apply different features to different positions of the facial expression face, can better retain image details, avoid blurring and distortion problems, make the generated facial expression face more accurate and natural, and can better reconstruct facial texture details.

[0176] Through the above self-developed facial expression generation model, it is possible to realize the image reconstruction of the face of the driven person and the corresponding expression of the driving person, and to achieve accurate facial expression transfer and facial texture reconstruction.

[0177] For the training of the above-mentioned facial expression generation model, we use the L1 loss function. We calculate the L1 error between the real facial expression image and the predicted facial expression image to make the predicted expression more accurate; at the same time, to ensure that the facial expression changes in the generated face are more natural and the facial texture is more accurate, we add dynamic face region mask enhancement and corresponding regional L1 error calculation, and set its weight to 10. The above respectively correspond to 2 face discriminators, and perform the L1 reconstruction loss and discriminant loss of the facial expression images respectively, making the face, corresponding expression, and facial texture more accurate and natural.

[0178] The above training also uses pre-training technology. Facial coefficients are extracted based on the large dataset of 3 million facial expression videos and pre-trained for 20 rounds. For each user, 5 minutes of diverse facial expression videos can be provided for fine-tuning, and only 5 rounds of training are required, with the rest remaining the same, which can reduce the training cost and achieve the effect of personal customization.

[0179] When actually applying the above method to generate a driven digital human, the video of the person to be driven is decoded into an image frame sequence, and the video stream of the driving person is obtained through a video capture device. During training, the driving person and the person to be driven are the same person mainly for the convenience of calculating the loss when regressing the facial expression image. During testing, the driving person and the person to be driven can be different people or the same person; an improved 3D face reconstruction model is used to extract the facial coefficients of the driving video, and the image of the person to be driven is input into the trained facial expression generation model, and the image of the person to be driven with the corresponding facial expression can be output. The continuous facial expression images are encoded into a video for display on the terminal device.

[0180] Exemplarily, in this exemplary embodiment, taking a virtual psychotherapist as an example, the beautified video of the person to be driven and the video of the driving person are uploaded for cloud training of the facial expression generation model. After the training is completed, by extracting the facial coefficients of the driving person in real time and combining them with the video frames of the person to be driven image, the generation of the corresponding facial expression of the person to be driven image is realized. At the same time, with the technology related to audio-driven mouth shapes, the effect is displayed through streaming, which can provide rich emotional feedback for patients and enhance interaction and communication.

[0181] It should be noted that the solution provided in this embodiment can be used as an independent facial expression migration function in live streaming with goods, virtual chat, virtual diagnosis and treatment, or short video creation, or can be integrated into hardware devices as an extended function.

[0182] This embodiment has the following advantages:

[0183] By improving the 3D face reconstruction model, the facial coefficients of diverse facial expression faces can be extracted more accurately.

[0184] By introducing an attention mechanism to dynamically adjust the correlation between facial expression features and face features, adaptive weighting and feature fusion are performed, and important features are focused on to avoid local distortion.

[0185] By extracting deep features and combining spatial adaptive normalization, different features are applied to different positions of the facial expression face, which can better retain the image details and make the generated facial expression face and facial texture details more accurate and natural.

Claims

1. A method for driving the expression generation of a digital human, characterized in that, The generation method includes: Obtain a driving video and a driven video; the driving video is the expression provider of the driven video. Input the driving video into a target 3D face reconstruction model to obtain target facial coefficient features; the target 3D face reconstruction model extracts features from the driving video based on the skip connection method. Input the target facial coefficient features and the driven video into a target expression generation model, and the target expression generation model is configured to perform global feature analysis and local feature analysis based on the target facial coefficient features and the driven video to obtain a driven image with a target expression. Perform video encoding based on the driven image to obtain a target video; the target video is a video that combines the facial information of the driven video and the target expression of the driving video. The target 3D face reconstruction model includes a first convolutional encoding unit, an image masking unit, a 3D face mesh reconstruction unit, a feature fusion unit, and a second convolutional encoding unit. The feature fusion unit includes an encoding layer and a decoding layer. The encoding layer includes a first encoding layer, a second encoding layer, and a third encoding layer connected in sequence. The decoding layer includes a first decoding layer, a second decoding layer, and a third decoding layer connected in sequence. The third encoding layer is connected to the first decoding layer; there is a skip connection between the encoding layer and the decoding layer. The first encoding layer, the second encoding layer, and the third encoding layer are all configured to perform downsampling on the features input from the previous layer; the first decoding layer, the second decoding layer, and the third decoding layer are all configured to perform upsampling on the features input from the previous layer. The first convolutional encoding unit and the 3D face mesh reconstruction unit are configured to perform reconstruction processing on the driving video; the image masking unit is configured to perform masking processing on the driving video; the feature fusion unit and the second convolutional encoding unit are configured to perform feature fusion processing based on residuals and skip connections on the features processed by the image masking unit and the 3D face mesh reconstruction unit.

2. The method for driving digital human expression generation according to claim 1, wherein The training process of the target 3D face reconstruction model includes: Obtain an expression video and perform a first preprocessing on the expression video to obtain a target expression video. Perform model training on the basis of the target expression video and the 3D face reconstruction model and converge to obtain the target 3D face reconstruction model.

3. The method for driving the generation of digital human expressions according to claim 2, wherein The step of performing model training based on the target expression video and the 3D face reconstruction model includes: Select several frames of training expression images from the target expression video. Perform a first processing on the training expression images to obtain facial coefficient features. Repeat the training of the 3D face reconstruction model and use a first loss function to converge the facial coefficient features until the 3D face reconstruction model meets the first requirement to obtain the target 3D face reconstruction model.

4. A method for driving digital human expression generation according to claim 3, wherein, The step of performing the first processing on the training expression images includes: Perform reconstruction processing and masking processing on the training expression images respectively to obtain first reconstruction features and a first masked image. Perform skip feature fusion processing based on the first reconstruction feature and the first occluded image to obtain a first fused feature; Perform encoding processing on the first fused feature to obtain the facial coefficient feature.

5. The method for driving digital human expression generation according to claim 4, wherein The steps of reconstructing the training expression image include: Perform encoding processing on the training expression image to obtain a first encoded feature; Perform 3D face mesh reconstruction on the first encoded feature to obtain a first reconstruction feature.

6. A method for driving digital human expression generation according to claim 4, characterized in that, The steps of converging the facial coefficient feature using the first loss function include: Calculate a first error between the facial coefficient feature and the training expression image using an L1 loss function; Calculate a second error between the facial coefficient feature and the training expression image using a cycle loss function; When both the first error and the second error meet the first requirement, stop training the 3D face reconstruction model to obtain the target 3D face reconstruction model.

7. A method for driving the generation of digital human expressions according to claim 3, characterized in that The training process of the target expression generation model includes: Obtain a training driving video, and perform second preprocessing on the training driving video to obtain a target training driving video; Perform model training and convergence according to the target training driving video, the facial coefficient feature, and the expression generation model to obtain the target expression generation model.

8. A method for driving the generation of digital human expressions according to claim 7, characterized in that, The steps of performing model training and convergence according to the target training driving video, the facial coefficient feature, and the expression generation model include: Perform convolutional encoding on the target training driving video and the facial coefficient feature respectively to obtain a second encoded feature and a third encoded feature; Perform feature fusion on the second encoded feature and the third encoded feature using a second process to obtain a second fused feature; the second process includes adaptive normalization processing with an attention mechanism; Perform residual calculation on the second fused feature to obtain a first residual feature; Perform deconvolution processing on the first residual feature to obtain a first deconvolution feature; Perform a third process on the first deconvolution feature to obtain a training driven image; the third process includes extracting deep features from the first deconvolution feature using spatial normalization processing and uniformly and adaptively adjusting the normalization parameters of the spatial normalization processing; Perform video encoding according to the training driven image to obtain a training target video; Use an L1 loss function to converge the training target video until the expression generation model meets the second requirement to obtain the target expression generation model.

9. The method for driving digital human expression generation according to claim 8, wherein The second process includes: Divide the second encoded feature and the third encoded feature into multiple subsequences, perform feature encoding processing on each subsequence, calculate the attention weights corresponding to different subsequences through a spatial attention module, and fuse the results of multiple feature encoding processes based on the attention weight calculation results to obtain a local feature encoding; the spatial attention module is preset in the expression generation model; Divide the local feature encoding into multiple subsequences, perform feature encoding processing on each subsequence, calculate the attention weights corresponding to different subsequences through the spatial attention module, and fuse the results of multiple feature encoding processes based on the calculation results of the attention weights to obtain the local feature encoding; Iteratively update, and independently output the local feature encoding obtained in each round to obtain multiple local feature encodings; Calculate the attention weights for multiple local feature encodings according to the channel attention module, and perform feature processing based on the calculation results of the attention weights to obtain the global feature encoding; fuse multiple local feature encodings and the global feature encoding to obtain the second fused feature; the channel attention module is preset in the expression generation model.

10. A method for driving digital human expression generation according to claim 8, characterized in that, The step of using the L1 loss function to converge the training target video includes: Calculate the third error between the training target video and the target expression video using the L1 loss function; When the third error meets the second requirement, stop training the expression generation model to obtain the target expression generation model.

Citation Information

Patent Citations

  • Expression migration method and device, equipment and medium

    CN115330914A