Digital Human Expression Driving Method, Device, Storage Medium and Computer Equipment

By combining multimodal feature extraction and cross-time and space attention mechanisms, high-quality digital human expression parameter sequences are generated and used to drive digital human expression videos, the problem of lack of space-time correlation of digital human expression drivers in the prior art is solved, and the video-level continuity and nature are achieved.

CN119538011BActive Publication Date: 2025-07-01GUANGZHOU QUWAN NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411784356.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-07-01
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

In the prior art, the digital human expression driving method is limited to feature processing at the image level and lacks in-depth correlation operations in the space-time dimension, so it is impossible to achieve high-quality video-level digital human expressions.

Method used

By obtaining the target audio and text, a multimodal feature extraction model is used to extract multimodal features and emotional categories, and an expression parameter prediction model trained across time and space attention mechanisms is used to generate an expression parameter sequence. Then, based on these emoticon parameter sequences, a face generation model is used to express the target character image to generate digital human expression-driven videos.

Benefits of technology

The digital human expression drive is realized in time and space, improving the quality and nature of the digital human expression drive in video scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119538011B_ABST
    Figure CN119538011B_ABST
Patent Text Reader

Abstract

The digital human expression driving method, device, storage medium and computer device provided by this application, when driving the expression of a digital human, first obtain the target audio and its target text, and use the target emotion feature extraction model to extract features from the target audio and target text to obtain multi-modal features and their corresponding emotion categories, so as to improve the accuracy of expression driving through multi-modal information; then determine the human expression data corresponding to the emotion category and the target expression parameter prediction model. Since this model introduces a cross-time-and-space attention mechanism, the expression parameter sequence generated after the model performs expression parameter prediction on the multi-modal features, emotion categories and human expression data contains feature correlation operations at the time and space levels. Therefore, after obtaining the target human image, the target face generation model can be used to drive the expression of the target human image based on the expression parameter sequence to generate a digital human expression driving video with time and space continuity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a digital human expression driving method, device, storage medium, and computer device. Background Art

[0002] With the rapid progress of digital human technology, its application scope is also continuously expanding, and it brings people and machines closer in an unprecedented interaction form. In daily communication, the human face, as an important social medium, the accurate transmission of its expressions and emotions is crucial for digital human technology. The process of transmitting expressions and emotions mainly relies on motion capture devices or facial reconstruction technologies to obtain expression parameters, and then realizes the dynamic changes of expressions through driving algorithms.

[0003] Currently, digital human expression driving methods mostly rely on deep neural networks, extract expression features by combining text or audio data, and then shape expressions through image generation technologies. However, these methods are limited to feature processing at the image level and lack in-depth correlation operations in the spatio-temporal dimension, so they cannot achieve high-quality video-level digital human expressions. Summary of the Invention

[0004] The purpose of this application aims to at least solve one of the above technical defects, especially the technical defect that the digital human expression driving method in the prior art is limited to feature processing at the image level and lacks in-depth correlation operations in the spatio-temporal dimension, so it cannot achieve high-quality video-level digital human expressions.

[0005] This application provides a digital human expression driving method, and the method includes:

[0006] Obtain the target audio and the target text corresponding to the target audio, and use the target emotion feature extraction model to extract features from the target audio and the target text to obtain multi-modal features and the emotion category corresponding to the multi-modal features;

[0007] Determine the human expression data corresponding to the emotion category, and determine the target expression parameter prediction model; the target expression parameter prediction model is trained using a cross-spatio-temporal attention mechanism;

[0008] Input the multi-modal features, the emotion category, and the human expression data into the target expression parameter prediction model to obtain an expression parameter sequence output by the target expression parameter prediction model;

[0009] Obtain the target human image, and use the target face generation model to drive the expression of the target human image based on the expression parameter sequence to generate a digital human expression driving video.

[0010] Optionally, the target emotion feature extraction model includes an audio feature extraction layer, a text feature extraction layer, a multi-modal feature fusion layer, and a classification layer;

[0011] The feature extraction of the target audio and the target text by using the target emotion feature extraction model to obtain multi-modal features and the emotion category corresponding to the multi-modal features includes:

[0012] Performing feature extraction on the target audio through the audio feature extraction layer to obtain audio features, and performing feature extraction on the target text through the text feature extraction layer to obtain text features;

[0013] Using the multi-modal feature fusion layer to splice and fuse the audio features and the text features, and outputting to obtain multi-modal features;

[0014] Inputting the multi-modal features into the classification layer, so that the classification layer performs emotion classification on the multi-modal features and outputs to obtain an emotion category.

[0015] Optionally, the determination of the facial expression data corresponding to the emotion category includes:

[0016] Retrieving the emotion category in the facial expression database to obtain the facial expression parameters corresponding to the emotion category;

[0017] Wherein, an index structure of multiple facial expression parameters is pre-established in the facial expression database.

[0018] Optionally, the construction process of the facial expression database includes:

[0019] Collecting a person dataset from multiple data sources, and annotating the emotion category corresponding to each person data in the person dataset;

[0020] For each emotion category, using face reconstruction technology to perform three-dimensional recognition on multiple face data under the emotion category to obtain the face key points of each face data, and determining the facial expression parameters corresponding to the emotion category based on each face key point;

[0021] Establishing an index structure between each emotion category and the corresponding facial expression parameters, and constructing a facial expression database according to each index structure.

[0022] Optionally, the determination of the target expression parameter prediction model includes:

[0023] Input the pre-acquired sample character emotion data into a preset initial expression parameter prediction model to obtain a predicted expression parameter sequence output by the initial expression parameter prediction model; wherein, the sample character emotion data includes multi-modal features, emotion categories, and character expression data;

[0024] Aim at making the predicted expression parameter sequence approach the real expression parameter sequence corresponding to the sample character emotion data, and use a cross-temporal and cross-spatial attention mechanism to train the initial expression parameter prediction model;

[0025] When the initial expression parameter prediction model meets the preset training end condition, use the trained initial expression parameter prediction model as the target expression parameter prediction model.

[0026] Optionally, the target expression parameter prediction model includes a text branch network, a video branch network, a CrossAttention network, and a fully connected layer;

[0027] The step of inputting the multi-modal features, the emotion categories, and the character expression data into the target expression parameter prediction model to obtain an expression parameter sequence output by the target expression parameter prediction model includes:

[0028] Perform high-dimensional mapping on the multi-modal features and the emotion categories through the text branch network to obtain text branch features, and perform high-dimensional mapping on the character expression data through the video branch network to obtain video branch features;

[0029] Use the Cross Attention network to calculate the cross-modal attention weights of the text branch features, the video branch features, the multi-modal features, and the character expression data, and output cross-modal attention features;

[0030] Input the cross-modal attention features and the character expression data into the fully connected layer to obtain an expression parameter sequence output by the fully connected layer.

[0031] Optionally, the step of using a target face generation model to perform expression driving on the target character image based on the expression parameter sequence to generate a digital human expression-driven video includes:

[0032] Input the expression parameter sequence and the target character image into the target face generation model, so that the target face generation model extracts the parameter sequence features of the expression parameter sequence and the image features of the target character image, and uses a cross-temporal and cross-spatial attention mechanism to adjust the weights of the parameter sequence features and the image features, and perform expression driving on the target character image according to the adjustment result to generate a digital human expression-driven video.

[0033] The present application also provides a digital human expression driving device, including:

[0034] A feature extraction module, configured to obtain a target audio and a target text corresponding to the target audio, and perform feature extraction on the target audio and the target text by using a target emotion feature extraction model to obtain multi-modal features and an emotion category corresponding to the multi-modal features;

[0035] A model determination module, configured to determine character expression data corresponding to the emotion category, and determine a target expression parameter prediction model; the target expression parameter prediction model is trained by using a cross-time and space attention mechanism;

[0036] A parameter prediction module, configured to input the multi-modal features, the emotion category, and the character expression data into the target expression parameter prediction model to obtain an expression parameter sequence output by the target expression parameter prediction model;

[0037] An expression driving module, configured to obtain a target human image, and perform expression driving on the target human image based on the expression parameter sequence by using a target face generation model to generate a digital human expression driving video.

[0038] The present application also provides a storage medium, in which computer-readable instructions are stored, and when the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the digital human expression driving method according to any one of the above embodiments.

[0039] The present application also provides a computer device, including: one or more processors, and a memory;

[0040] The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the digital human expression driving method according to any one of the above embodiments are executed.

[0041] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0042] The digital human expression driving method, device, storage medium, and computer device provided by this application can, when driving the expression of a digital human, first obtain the target audio and the target text corresponding to the target audio, and use the target emotion feature extraction model to extract features from the target audio and the target text to obtain multi-modal features and the emotion categories corresponding to the multi-modal features, so as to realize emotion feature extraction through multi-modal information, thereby improving the accuracy of digital human expression driving; then it can determine the character expression data corresponding to the emotion category and the target expression parameter prediction model. Since the target expression parameter prediction model introduces a cross-temporal and cross-spatial attention mechanism, the expression parameter sequence generated after the model predicts the expression parameters for the multi-modal features, emotion categories, and character expression data contains feature correlation operations at the temporal and spatial levels. Therefore, after obtaining the target human image, this application can use the target face generation model to drive the expression of the target human image based on the expression parameter sequence to generate a digital human expression driving video with temporal and spatial continuity. All in all, by combining multi-modal input information and the cross-temporal and cross-spatial attention mechanism, this application can achieve the ability to predict expression parameters at the video level, thereby enhancing the continuity of digital human expression driving in time and space in a video scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0044] Figure 1 It is a schematic flowchart of a digital human expression driving method provided by an embodiment of this application;

[0045] Figure 2 It is a schematic flowchart of the application process of a target emotion feature extraction model provided by an embodiment of this application;

[0046] Figure 3 It is a schematic flowchart of the application process of a target expression parameter prediction model provided by an embodiment of this application;

[0047] Figure 4 It is a schematic structural diagram of a digital human expression driving device provided by an embodiment of this application;

[0048] Figure 5 It is a schematic internal structure diagram of a computer device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0050] Currently, most digital human expression driving methods rely on deep neural networks to extract expression features by combining text or audio data, and then use image generation technology to shape the expressions. However, these methods are limited to feature processing at the image level and lack in-depth correlation operations in the spatio-temporal dimension, so high-quality video-level digital human expressions cannot be achieved.

[0051] Based on this, the present application proposes the following technical solutions. For details, please refer to the following:

[0052] In one embodiment, as Figure 1 shown, Figure 1 is a schematic flowchart of a digital human expression driving method provided by an embodiment of the present application; the present application provides a digital human expression driving method, which specifically includes the following:

[0053] S110: Obtain the target audio and the target text corresponding to the target audio, and use the target emotion feature extraction model to extract features from the target audio and the target text to obtain multi-modal features and the emotion category corresponding to the multi-modal features.

[0054] In this step, when it is necessary to configure the corresponding digital human expression driving for film and television-level audio, the computer device can first obtain the target audio and the target text corresponding to the target audio, and then use the pre-trained target emotion feature model to extract features from the target audio and the target text to obtain multi-modal features and the emotion category corresponding to the multi-modal features, so as to improve the accuracy of digital human expression driving by using multi-modal information.

[0055] Among them, the target emotion feature model refers to a model obtained by using machine learning or deep neural network learning as a pre-training model and then improving and training it. It is mainly used to analyze and extract the emotion features of relevant data such as audio and text, and then can accurately infer the emotion category in the audio and text according to the extracted emotion features, so as to drive the expression of the digital human according to these emotion information, making the driving effect more vivid and real.

[0056] Specifically, the target emotion feature model of the present application can extract multi-modal features representing emotions from the target audio and target text, such as features like pitch, volume, rhythm in the target audio, and features like emotional words, sentence patterns in the text. These features can all provide fine-grained information related to emotions. In addition, the model can also map the extracted multi-modal features to specific emotion categories, such as happiness, sadness, anger, etc., enabling the digital human to show expressions highly matching the emotions in the expression drive. Therefore, through the joint analysis of multi-modal features, the present application can capture more comprehensive emotion information, making the expression drive more accurate and natural. For example, the tone in the target audio may express an emotion, and the text content can provide additional context, thus helping the digital human to show more delicate expressions.

[0057] Furthermore, after training the target emotion feature model, the computer device can store it so that when performing feature extraction later, it can directly call the pre-stored target emotion feature model to extract features from the target audio and target emotion. In addition, the target emotion feature model of the present application can choose to improve and train the ImageBind model to use the improved ImageBind model to perform operations on deep features, realizing the extraction of multi-modal features and the recognition of emotion categories. Here, the ImageBind model refers to a multi-modal deep learning model that can map data of different modalities to a shared embedding space, enabling the data to find natural similarities between different modalities, and can also automatically align the features of different modalities, enabling the model to perform feature alignment while extracting multi-modal features, reducing information loss, and improving the quality of multi-modal information integration.

[0058] S120: Determine the human expression data corresponding to the emotion category, and determine the target expression parameter prediction model; the target expression parameter prediction model is trained using a cross-temporal and cross-spatial attention mechanism.

[0059] In this step, after obtaining the emotion category corresponding to the multi-modal through step S110, the computer device can determine the human expression data corresponding to this emotion category, and determine the target expression parameter prediction model for predicting the expression parameters of the human expression data. Since the target expression parameter prediction model is trained using a cross-temporal and cross-spatial attention mechanism, the results predicted by the model include feature correlation operations at the temporal and spatial levels, and thus can ensure the continuity of the subsequent digital human expression drive in time and space.

[0060] Among them, the human facial expression data refers to the expression information composed of multiple 3D facial key points, which are the core data for driving the expressions of digital humans. These key points usually include the positions and postures of parts such as eyes, mouths, and eyebrows, and can accurately describe the subtle changes in facial features, thereby reflecting specific emotional states. The target expression parameter prediction model is a model trained using a cross-temporal and spatial attention mechanism, which is used to predict and generate the expression parameters of digital humans in a time series, ensuring the continuity of digital human facial expressions in time and space.

[0061] It can be understood that the cross-temporal and spatial attention mechanism of this application is a mechanism for processing the spatio-temporal relationships in data, which extends the traditional attention mechanism to the time and space dimensions. Therefore, based on the cross-temporal and spatial attention mechanism, the target expression parameter prediction model can capture the associations of feature data in both the time and space dimensions, thereby understanding the dynamic changes of the object and the mutual influence of spatial features, and ensuring the coherence and naturalness of the model output results.

[0062] For example, in the time dimension, the target expression parameter prediction model can focus on the changes in data at different times to capture dynamic information; for example, when analyzing facial expressions, cross-temporal and spatial attention can focus on how a smile gradually unfolds from nothing to something, so that the model can understand the progress of events, the connection of actions, and the changes in facial expressions over time through temporal associations. In the space dimension, the target expression parameter prediction model can focus on the spatial positions and distribution relationships of data to capture the interactions of an object at different spatial points at the same time; for example, when analyzing a specific facial expression formed by the coordination relationship between a person's eyebrows, eyes, and mouth, cross-temporal and spatial attention can assist the model in capturing the spatial mutual influence of features at the same time to understand the spatial distribution characteristics of facial expressions or actions.

[0063] S130: Input the multi-modal features, emotion categories, and human facial expression data into the target expression parameter prediction model to obtain the expression parameter sequence output by the target expression parameter prediction model.

[0064] In this step, after obtaining the target expression parameter prediction model through step S120, the computer device can input the multi-modal features, emotion categories, and human facial expression data into the target expression parameter prediction model to obtain the expression parameter sequence output by the target expression parameter prediction model. Since the target expression parameter prediction model introduces a cross-temporal and spatial attention mechanism, the expression parameter sequence output by the model is coherent and natural in time and space.

[0065] It can be understood that the expression parameter sequence, as the output result of the target expression parameter prediction model, can drive the facial expression of the digital human to present a dynamic effect that conforms to the emotion category. The cross-temporal and cross-spatial attention mechanism introduced by the target expression parameter prediction model can focus on the change trend of expressions between different time points in the time dimension and the relationships and linkages between three-dimensional face key points in the spatial dimension. Therefore, when the emotion expression in the expression parameter sequence output by the model matches the multi-modal features and the facial expression data of the person, it can achieve a high degree of coherence in time and space, ensuring a smooth transition and natural performance of the expression change, and making the digital human expression drive present a more realistic and vivid effect in consecutive multi-frame images.

[0066] S140: Obtain the target person image, and use the target face generation model to perform expression drive on the target person image based on the expression parameter sequence to generate a digital human expression drive video.

[0067] In this step, after obtaining the expression parameter sequence through step S130, the computer device can obtain the target person image as the digital human for expression drive, and then use the target face generation model to perform expression drive on the target person image based on the expression parameter sequence to generate a digital human expression drive video with temporal and spatial continuity.

[0068] It can be understood that the target face generation model of the present application can use the U-Net model as a pre-training model for improvement and training. Among them, the U-Net model is a deep learning model used to generate a segmented output. It can perform accurate pixel-level prediction in terms of details and uses skip connections to directly transfer the information of the early layers in the encoder to the decoder, thereby retaining more spatial details.

[0069] Furthermore, the present application can further improve and train the structure of the U-Net model to better adapt to the spatio-temporal characteristics of the expression changes in the expression parameter sequence. Driven by the expression parameter sequence provided by the target expression parameter prediction model, the U-Net model can accurately update the expression of the target person image, enabling the digital human to show a reasonable expression evolution in the time dimension, and the relationships between facial features such as eyes, eyebrows, and mouth can be accurately mapped in space, ensuring that the finally generated digital human expression drive video has a sense of reality and smoothness in facial expressions.

[0070] In the above embodiments, when driving the expression of a digital human, the target audio and the target text corresponding to the target audio can be obtained first, and the target emotion feature extraction model can be used to extract features from the target audio and the target text to obtain multi-modal features and the emotion categories corresponding to the multi-modal features. Thus, emotion feature extraction can be achieved through multi-modal information, and further, the accuracy of digital human expression driving can be improved. Then, the human expression data corresponding to the emotion category and the target expression parameter prediction model can be determined. Since the target expression parameter prediction model introduces a cross-time-and-space attention mechanism, the expression parameter sequence generated after the model performs expression parameter prediction on the multi-modal features, emotion categories, and human expression data contains feature correlation operations at the time-and-space level. Therefore, after obtaining the target human image, the present application can use the target face generation model to drive the expression of the target human image based on the expression parameter sequence to generate a digital human expression driving video with time and space continuity. In summary, by combining multi-modal input information and a cross-time-and-space attention mechanism, the present application can achieve the ability to predict expression parameters at the video level, and further improve the continuity of digital human expression driving in terms of time and space.

[0071] In one embodiment, as Figure 2 shown, Figure 2 is a schematic flowchart of the application process of a target emotion feature extraction model provided by an embodiment of the present application; Figure 2 In [the figure], in step S110, the target emotion feature extraction model may include an audio feature extraction layer, a text feature extraction layer, a multi-modal feature fusion layer, and a classification layer. Among them, the process of using the target emotion feature extraction model to extract features from the target audio and the target text to obtain multi-modal features and the emotion categories corresponding to the multi-modal features may include:

[0072] S111: Extract audio features from the target audio through the audio feature extraction layer, and extract text features from the target text through the text feature extraction layer.

[0073] S112: Use the multi-modal feature fusion layer to splice and fuse the audio features and the text features, and output to obtain multi-modal features.

[0074] S113: Input the multi-modal features into the classification layer, so that the classification layer performs emotion classification on the multi-modal features and outputs to obtain the emotion category.

[0075] In this embodiment, the target emotion extraction model may be composed of an audio feature extraction layer, a text feature extraction layer, a multimodal feature fusion layer, and a classification layer. Therefore, when using the target emotion feature extraction model to extract multimodal features, the computer device may first input the target audio into the audio feature extraction layer and input the target text into the text feature extraction layer, so as to obtain the audio features output by the audio feature extraction layer and the text features output by the text feature extraction layer. Then, the computer device may use the multimodal feature fusion layer to splice and fuse the audio features and text features to obtain multimodal features. Furthermore, the multimodal features may be input into the classification layer, so that the classification layer performs emotion classification on the multimodal features and outputs the emotion category.

[0076] It should be noted that the target emotion feature extraction model of the present application may be obtained by improving the structure of the ImageBind model and then training it. Specifically, the computer device may set the feature extraction capabilities of two modalities, audio and text, in the multi-feature extraction layer of the ImageBind model, and then build a multimodal feature fusion layer based on the Transformer structure to perform feature splicing and feature extraction on the extracted audio and text features to obtain multimodal features with semantic information, and the feature dimension information thereof may be 1x768 and the numerical type is floating point number; finally, the computer device may create a classification layer according to the number of emotion categories, and then perform emotion recognition and classification based on the above-mentioned extracted multimodal features, calculate the loss value according to the classification result, and perform backpropagation based on the calculated loss value to update the weights of the model until the model meets the training end condition.

[0077] In one embodiment, determining the human facial expression data corresponding to the emotion category in step S120 may include:

[0078] S121: Retrieve the emotion category in the facial expression database to obtain the human facial expression parameters corresponding to the emotion category.

[0079] In this embodiment, the computer device is configured with a facial expression database, and an index structure of multiple human facial expression parameters is established in advance in the facial expression database. Therefore, the computer device may directly retrieve the emotion category in the facial expression database, so as to obtain the human facial expression parameters in the index structure corresponding to the emotion category.

[0080] It can be understood that the index structure in the expression database contains the corresponding human expression parameters under different emotion categories. These human expression parameters are accurately labeled and can accurately reflect the expression characteristics under different emotional states. When the system needs to generate an expression of a specific emotion category for the digital human, the computer device can directly perform a quick search in the expression database and extract the corresponding human expression parameters from the index structure based on the emotion category. Through this retrieval mechanism, the computer device can quickly obtain expression data that highly matches the emotion category and apply it to the expression driving of the digital human, thereby improving the efficiency of expression generation and reducing the calculation time for generating expression parameters from scratch.

[0081] In one embodiment, the construction process of the expression database in step S121 may include:

[0082] S1211: Collect a human dataset from multiple data sources and label the emotion category corresponding to each piece of human data in the human dataset.

[0083] S1212: For each emotion category, use face reconstruction technology to perform three-dimensional recognition on multiple face data under that emotion category, obtain the face key points of each face data, and determine the human expression parameters corresponding to that emotion category based on each face key point.

[0084] S1213: Establish an index structure between each emotion category and the corresponding human expression parameters, and construct an expression database according to each index structure.

[0085] In this embodiment, when constructing the expression database, the computer device can first collect multiple pieces of human data from multiple data sources to form a human dataset, and then label the corresponding emotion category for each piece of human data in the human dataset. For each emotion category, the computer device can use face reconstruction technology to perform three-dimensional recognition on multiple face data under that emotion category, obtain the face key points of each face data, and determine the human expression parameters corresponding to that emotion category; finally, the computer device can establish an index structure between each emotion category and the corresponding human expression parameters, and construct an expression database according to each index structure.

[0086] Among them, face reconstruction technology refers to a technology that extracts three-dimensional face shape and structure information from two-dimensional images or videos through computer vision and image processing methods, and restores and models facial details. It transforms ordinary two-dimensional facial images into accurate three-dimensional face models and provides precise positioning of facial features such as eyes, nose, mouth, etc.

[0087] It can be understood that the computer device can not only provide the facial expression features of the person in a single image in the emotion category, but also synthesize the facial expression features of different individuals, and then extract representative emotion features. After calculation and analysis, these features can form a set of high-dimensional human facial expression parameters, which cover the change range, movement trajectory and mutual coordination of each key point on the face, providing a data basis for subsequent expression driving.

[0088] Furthermore, when the computer device suggests an index structure, it can be constructed through the ElasticSearch retrieval tool. Elastic Search realizes faster filtering than relational databases through the inverted index technology of Lucene. Therefore, when using the Elastic Search retrieval tool, the human facial expression parameters corresponding to each emotion category can be obtained first, and then the Elastic Search retrieval tool can be used to establish an index between each emotion category and its human facial expression parameters, so as to obtain the index structure in the expression database.

[0089] In one embodiment, the process of determining the target expression parameter prediction model in step S120 may include:

[0090] S122: Input the pre-acquired sample human emotion data into the preset initial expression parameter prediction model to obtain the predicted expression parameter sequence output by the initial expression parameter prediction model; wherein, the sample human emotion data includes multi-modal features, emotion categories and human facial expression data.

[0091] S123: Aim at the predicted expression parameter sequence approaching the real expression parameter sequence corresponding to the sample human emotion data, and use the cross-time and space attention mechanism to train the initial expression parameter prediction model.

[0092] S124: When the initial expression parameter prediction model meets the preset training end condition, take the trained initial expression parameter prediction model as the target expression parameter prediction model.

[0093] In this embodiment, the computer device can first obtain sample human emotion data composed of multi-modal features, emotion categories, and human facial expression data, as well as the corresponding true expression parameter sequence of the sample human emotion data, as the basic data for model training. Then, the computer device can input the sample human emotion data into a preset initial expression parameter prediction model to obtain the predicted expression parameter sequence output by the initial expression parameter prediction model, and use a cross-temporal attention mechanism to train the initial expression parameter prediction model with the goal of making the predicted expression parameter sequence approach the true expression parameter sequence corresponding to the sample human emotion data. When the initial expression parameter prediction model meets the preset training end condition, the trained initial expression parameter prediction model is used as the target expression parameter prediction model.

[0094] Specifically, when training the target expression parameter prediction model, the computer device can use different types of sample human emotion data as training samples and label each training sample with a sample label, that is, the corresponding true expression parameter sequence. After all the training samples are labeled, the computer device can input the training samples with sample labels into the preset initial expression parameter prediction model for forward propagation to train the model, and use a preset target loss function to optimize the model parameters during the reverse propagation of the model. When the model meets certain training conditions or parameter convergence conditions, such as the number of iterations reaches a set value, it is considered that the training is completed. At this time, the trained model can be used as the final target expression parameter prediction model.

[0095] Furthermore, after the computer device trains the target expression parameter prediction model, it can store it so that when performing expression parameter prediction later, it can directly call the pre-stored target expression parameter prediction model to perform expression parameter prediction on multi-modal features and human facial expression data. In addition, the target expression parameter prediction model of this application can choose the VindLU model for improvement and training; the VindLU model here is a multi-modal model based on computer vision and natural language understanding, which can effectively process and fuse visual and language information, improve the understanding ability of multi-modal data, and then enhance the model's recognition and response to emotional changes, so as to improve the overall effect of expression driving and ensure that the emotional expression of the digital human is more delicate and real.

[0096] In one embodiment, as Figure 3 shown, Figure 3 is a schematic flowchart of the application process of a target expression parameter prediction model provided by an embodiment of this application; Figure 3In this case, in step S130, the target expression parameter prediction model may include a text branch network, a video branch network, a Cross Attention network, and a fully connected layer. Among them, the process of inputting multi-modal features, emotion categories, and human expression data into the target expression parameter prediction model to obtain the expression parameter sequence output by the target expression parameter prediction model may include:

[0097] S131: Perform high-dimensional mapping on the multi-modal features and emotion categories through the text branch network to obtain text branch features, and perform high-dimensional mapping on the human expression data through the video branch network to obtain video branch features.

[0098] S132: Use the Cross Attention network to calculate the cross-modal attention weights of the text branch features, video branch features, multi-modal features, and human expression data, and output the cross-modal attention features.

[0099] S133: Input the cross-modal attention features and the human expression data into the fully connected layer to obtain the expression parameter sequence output by the fully connected layer.

[0100] In this embodiment, the target expression parameter prediction model may be composed of a text branch network, a video branch network, a Cross Attention network, and a fully connected layer. Therefore, when the computer device predicts the expression parameter sequence, it can first input the multi-modal features and emotion categories into the text branch network, and input the human expression data into the video branch network, so as to obtain the text branch features output by the text branch network and the video branch features output by the video branch network. Then, the computer device can use the Cross Attention network to calculate the cross-modal attention weights of the text branch features, video branch features, multi-modal features, and human expression data to output the cross-modal attention features, and input the cross-modal attention features and the human expression data into the fully connected layer to obtain the expression parameter sequence output by the fully connected layer.

[0101] It should be noted that the target expression parameter prediction model of the present application can be obtained by improving the structure of the VindLU model and then training. Specifically, the computer device can set the input of the first-layer structure of the text branch network of the VindLU model to the multimodal features of 1x768 dimensions and the token values corresponding to the emotion labels, and set the input of the first-layer structure of the video branch network to the human face expression data; here, the human face expression data is the coordinate values of three-dimensional human face key points, and its dimension information is N×3, where N represents the number of key points, and 3 represents the coordinate values in the x, y, and z directions; then, the computer device can also set a Cross Attention network in the VindLU model with the structure modified, so that it can combine the multimodal features and the human face expression data to calculate the cross-modal attention weights, and use the calculated weights in subsequent feature operations, so that the model has the ability to extract attention features at the spatio-temporal level; finally, the computer device can create a new fully connected layer in the model, and its output dimension is set to B×N×3, which is used to represent the expression parameter sequence predicted and output by the model, where B represents the number of frames.

[0102] In one embodiment, the process of using the target human face generation model to perform expression driving on the target human image based on the expression parameter sequence in step S140 to generate a digital human expression-driven video may include:

[0103] S141: Input the expression parameter sequence and the target human image into the target human face generation model, so that the target human face generation model extracts the parameter sequence features of the expression parameter sequence and the image features of the target human image, and uses the cross-spatio-temporal attention mechanism to adjust the weights of the parameter sequence features and the image features, and, according to the adjustment result, perform expression driving on the target human image to generate a digital human expression-driven video.

[0104] In this embodiment, when the computer device performs expression driving on the digital human, it can first obtain the target human face generation model, so that it can input the expression parameter sequence and the target human image into the target human face generation model, so that the target human face generation model extracts the parameter sequence features of the expression parameter sequence and the image features of the target human image, and uses the cross-spatio-temporal attention mechanism to adjust the weights of the parameter sequence features and the image features. Finally, the target human face generation model can perform expression driving on the target human image according to the adjustment result to generate and output a digital human expression-driven video.

[0105] It should be noted that the target expression parameter prediction model of the present application can be obtained by improving the structure of the U-Net model and then training. Specifically, the computer device can modify the shallow structure of the U-Net model, and set its input as the target person image and the expression parameter sequence. Among them, the dimension of the target person image can be Bx3xHxW, and the dimension of the expression parameter sequence can be BxNx3, where B represents the number of image frames, N represents the number of three-dimensional face key points, H represents the image height, and W represents the image width. Then, the computer device can also introduce a CrossAttention module into the U-Net model after the structure modification, so that the model can calculate cross-modal attention weights by combining image features and parameter sequence features, and use the calculated weights for subsequent feature operations. Finally, the computer device can set the output dimension of the output layer in the network structure of the U-Net model to Bx3xHxW, which is used to represent the digital human expression-driven video composed of multiple frames of person images generated by the model.

[0106] The digital human expression-driven device provided by the embodiments of the present application will be described below. The digital human expression-driven device described below can be mutually corresponding and referred to the digital human expression-driven method described above.

[0107] In one embodiment, as Figure 4 shown, Figure 4 is a schematic structural diagram of a digital human expression-driven device provided by an embodiment of the present application; the present application also provides a digital human expression-driven device, including a feature extraction module 210, a model determination module 220, a parameter prediction module 230, and an expression driving module 240, specifically including the following:

[0108] The feature extraction module 210 is used to obtain the target audio and the target text corresponding to the target audio, and use the target emotion feature extraction model to extract features from the target audio and the target text to obtain multi-modal features and the emotion category corresponding to the multi-modal features.

[0109] The model determination module 220 is used to determine the person expression data corresponding to the emotion category, and determine the target expression parameter prediction model; the target expression parameter prediction model is trained by using a cross-space-time attention mechanism.

[0110] The parameter prediction module 230 is used to input the multi-modal features, the emotion category, and the person expression data into the target expression parameter prediction model to obtain the expression parameter sequence output by the target expression parameter prediction model.

[0111] The expression driving module 240 is used to obtain the target person image, and use the target face generation model to perform expression driving on the target person image based on the expression parameter sequence to generate a digital human expression-driven video.

[0112] In the above embodiments, when driving the expression of the digital human, the target audio and the target text corresponding to the target audio can be obtained first, and the target emotion feature extraction model can be used to extract features from the target audio and the target text to obtain multi-modal features and the emotion categories corresponding to the multi-modal features, so that emotion feature extraction can be realized through multi-modal information, and further improve the accuracy of digital human expression driving; then, the human expression data corresponding to the emotion category and the target expression parameter prediction model can be determined. Since the target expression parameter prediction model introduces a cross-time-and-space attention mechanism, the expression parameter sequence generated after the model performs expression parameter prediction on the multi-modal features, emotion categories, and human expression data includes feature correlation operations at the time and space levels. Therefore, after obtaining the target human image, the present application can use the target face generation model to perform expression driving on the target human image based on the expression parameter sequence to generate a digital human expression driving video with temporal and spatial continuity. In short, by combining multi-modal input information and a cross-time-and-space attention mechanism, the present application can achieve the ability to predict expression parameters at the video level, and further improve the temporal and spatial continuity of digital human expression driving in the video scenario.

[0113] In one embodiment, the target emotion feature extraction model in the feature extraction module 210 may include an audio feature extraction layer, a text feature extraction layer, a multi-modal feature fusion layer, and a classification layer; wherein, the feature extraction module 210 may further include:

[0114] An audio-text feature extraction sub-module, configured to extract features from the target audio through the audio feature extraction layer to obtain audio features, and extract features from the target text through the text feature extraction layer to obtain text features.

[0115] A feature splicing and fusion sub-module, configured to splice and fuse the audio features and the text features by using the multi-modal feature fusion layer, and output to obtain multi-modal features.

[0116] An emotion classification sub-module, configured to input the multi-modal features into the classification layer, so that the classification layer performs emotion classification on the multi-modal features, and outputs to obtain emotion categories.

[0117] In one embodiment, the model determination module 220 may include:

[0118] A category retrieval sub-module, configured to retrieve the emotion category in the expression database to obtain the human expression parameters corresponding to the emotion category.

[0119] Wherein, an index structure of multiple human expression parameters is pre-established in the expression database.

[0120] In one embodiment, the category retrieval sub-module may include:

[0121] A data acquisition unit for collecting a person dataset from multiple data sources and annotating the emotion category corresponding to each person data in the person dataset.

[0122] A parameter calculation unit for, for each emotion category, using face reconstruction technology to perform three-dimensional recognition on multiple face data under the emotion category to obtain the face key points of each face data, and determining the person expression parameters corresponding to the emotion category based on each face key point.

[0123] An index construction unit for establishing an index structure between each emotion category and the corresponding person expression parameters, and constructing an expression database according to each index structure.

[0124] In one embodiment, the model determination module 220 may further include:

[0125] A model prediction sub-module for inputting pre-acquired sample person emotion data into a preset initial expression parameter prediction model to obtain a predicted expression parameter sequence output by the initial expression parameter prediction model; wherein, the sample person emotion data includes multi-modal features, emotion categories, and person expression data.

[0126] A model training sub-module for aiming at the predicted expression parameter sequence approaching the true expression parameter sequence corresponding to the sample person emotion data, and training the initial expression parameter prediction model by using a cross-temporal and cross-spatial attention mechanism.

[0127] A model generation sub-module for, when the initial expression parameter prediction model meets the preset training end condition, taking the trained initial expression parameter prediction model as the target expression parameter prediction model.

[0128] In one embodiment, the target expression parameter prediction model in the parameter prediction module 230 may include a text branch network, a video branch network, a Cross Attention network, and a fully connected layer; wherein, the parameter prediction module 230 may further include:

[0129] A high-dimensional mapping sub-module for performing high-dimensional mapping on the multi-modal features and emotion categories through the text branch network to obtain text branch features, and performing high-dimensional mapping on the person expression data through the video branch network to obtain video branch features.

[0130] A weight adjustment sub-module for calculating the cross-modal attention weights of the text branch features, video branch features, multi-modal features, and person expression data by using the Cross Attention network, and outputting cross-modal attention features.

[0131] A sequence output sub-module, configured to input the cross-modal attention features and the human expression data into a fully-connected layer, and obtain an expression parameter sequence output by the fully-connected layer.

[0132] In one embodiment, the expression driving module 240 may include:

[0133] A model driving sub-module, configured to input the expression parameter sequence and the target human image into a target face generation model, so that the target face generation model extracts the parameter sequence features of the expression parameter sequence and the image features of the target human image, and uses a cross-temporal and spatial attention mechanism to adjust the weights of the parameter sequence features and the image features, and further, drive the expression of the target human image according to the adjustment result to generate a digital human expression driving video.

[0134] In one embodiment, the present application further provides a storage medium, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the digital human expression driving method according to any one of the above embodiments.

[0135] In one embodiment, the present application further provides a computer device, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the digital human expression driving method according to any one of the above embodiments.

[0136] Schematically, as Figure 5 shown, Figure 5 is an internal structural schematic diagram of a computer device provided by an embodiment of the present application. The computer device 300 may be provided as a server. Referring to Figure 5 , the computer device 300 includes a processing component 302, which further includes one or more processors, and memory resources represented by a memory 301 for storing instructions executable by the processing component 302, such as application programs. The application programs stored in the memory 301 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 302 is configured to execute instructions to perform the digital human expression driving method of any of the above embodiments.

[0137] The computer device 300 may further include a power supply component 303 configured to perform power management of the computer device 300, a wired or wireless network interface 304 configured to connect the computer device 300 to a network, and an input / output (I / O) interface 305. The computer device 300 may operate based on an operating system stored in the memory 301, such as Windows Server TM, Mac OS XTM, Unix TM, Linux TM, Free BSDTM, or the like.

[0138] Those skilled in the art can understand that Figure 5 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0139] Finally, it should also be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0140] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0141] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A digital human expression driving method, characterized in that: The method comprises: Acquire a target audio and a target text corresponding to the target audio, and use a target emotion feature extraction model to perform feature extraction on the target audio and the target text to obtain a multimodal feature and an emotion category corresponding to the multimodal feature; Determine the facial expression data corresponding to the emotion category, and determine the target facial expression parameter prediction model; the target facial expression parameter prediction model is trained by a cross-temporal and spatial attention mechanism; the cross-temporal and spatial attention mechanism focuses on the changing trend of facial expressions at different time points in the time dimension, and focuses on the relationship and linkage between key points of the three-dimensional face in the spatial dimension; Inputting the multimodal features, the emotion categories and the character expression data into the target expression parameter prediction model to obtain an expression parameter sequence output by the target expression parameter prediction model; A target person image is obtained, and a target person face generation model is used to perform expression driving on the target person image based on the expression parameter sequence to generate a digital human expression driven video.

2. The digital human expression driving method according to claim 1, characterized in that: The target emotion feature extraction model includes an audio feature extraction layer, a text feature extraction layer, a multimodal feature fusion layer and a classification layer; The target emotion feature extraction model is used to extract features from the target audio and the target text to obtain multimodal features and emotion categories corresponding to the multimodal features, including: Performing feature extraction on the target audio through the audio feature extraction layer to obtain audio features, and performing feature extraction on the target text through the text feature extraction layer to obtain text features; Using the multimodal feature fusion layer to concatenate and fuse the audio features and the text features, and output to obtain a multimodal feature; The multimodal features are input into the classification layer, so that the classification layer performs emotion classification on the multimodal features and outputs the emotion category.

3. The digital human expression driving method according to claim 1, characterized in that: The determining of the character expression data corresponding to the emotion category includes: Retrieving the emotion category in an expression database to obtain character expression parameters corresponding to the emotion category; Wherein, an index structure of multiple character expression parameters is pre-established in the expression database.

4. The digital human expression driving method according to claim 3, characterized in that: The process of constructing the expression database includes: Collecting a character data set from multiple data sources, and marking the emotion category corresponding to each character data in the character data set; For each emotion category, face reconstruction technology is used to perform three-dimensional recognition on multiple face data under the emotion category to obtain the face key points of each face data, and the character expression parameters corresponding to the emotion category are determined based on each face key point; An index structure is established between each emotion category and the corresponding character expression parameter, and an expression database is constructed based on each index structure.

5. The digital human expression driving method according to claim 1, characterized in that: The step of determining a target expression parameter prediction model comprises: Inputting the pre-acquired sample character emotion data into a preset initial expression parameter prediction model to obtain a predicted expression parameter sequence output by the initial expression parameter prediction model; wherein the sample character emotion data includes multimodal features, emotion categories and character expression data; The predicted expression parameter sequence is targeted to be close to the real expression parameter sequence corresponding to the sample character emotion data, and the initial expression parameter prediction model is trained using a cross-temporal and spatial attention mechanism; When the initial expression parameter prediction model meets the preset training end condition, the trained initial expression parameter prediction model is used as the target expression parameter prediction model.

6. The digital human expression driving method according to claim 1, characterized in that: The target expression parameter prediction model includes a text branch network, a video branch network, a Cross Attention network and a fully connected layer; The step of inputting the multimodal features, the emotion category and the character expression data into the target expression parameter prediction model to obtain an expression parameter sequence output by the target expression parameter prediction model comprises: Performing high-dimensional mapping on the multimodal features and the emotion categories through the text branch network to obtain text branch features, and performing high-dimensional mapping on the character expression data through the video branch network to obtain video branch features; The Cross Attention network is used to calculate the cross-modal attention weights of the text branch features, the video branch features, the multimodal features, and the character expression data, and outputs the cross-modal attention features; The cross-modal attention features and the character expression data are input into the fully connected layer to obtain an expression parameter sequence output by the fully connected layer.

7. The digital human expression driving method according to claim 1, characterized in that: The method of using the target human face generation model to drive the target human image by expression based on the expression parameter sequence to generate a digital human expression driven video includes: The expression parameter sequence and the target person image are input into a target face generation model, so that the target face generation model extracts parameter sequence features of the expression parameter sequence and image features of the target person image, and uses a cross-temporal and spatial attention mechanism to adjust the weights of the parameter sequence features and the image features, and, based on the adjustment results, the target person image is expression-driven to generate a digital human expression-driven video.

8. A digital human expression driving device, characterized in that: include: A feature extraction module is used to obtain a target audio and a target text corresponding to the target audio, and use a target emotion feature extraction model to perform feature extraction on the target audio and the target text to obtain a multimodal feature and an emotion category corresponding to the multimodal feature; A model determination module is used to determine the character expression data corresponding to the emotion category and determine a target expression parameter prediction model; the target expression parameter prediction model is trained using a cross-temporal and spatial attention mechanism; The cross-temporal and spatial attention mechanism focuses on the changing trend of expressions between different time points in the temporal dimension, and focuses on the relationship and linkage between key points of the three-dimensional face in the spatial dimension; A parameter prediction module, used for inputting the multimodal features, the emotion category and the character expression data into the target expression parameter prediction model to obtain an expression parameter sequence output by the target expression parameter prediction model; The expression driving module is used to obtain a target person image, and use a target face generation model to perform expression driving on the target person image based on the expression parameter sequence to generate a digital human expression driving video.

9. A storage medium, characterized in that: The storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the digital human expression driving method as claimed in any one of claims 1 to 7.

10. A computer device, characterized in that: include: one or more processors, and memory; The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the digital human expression driving method according to any one of claims 1 to 7 are executed.

Citation Information

Patent Citations

  • Emotional interaction apparatus

    US20170365277A1

  • Human-computer interaction method and apparatus, and terminal device

    US20240402989A1