Digital human animation generation method, model training method and computing equipment

By training an animation generation model, and using audio encoding networks and emotion prediction networks to generate digital human animations, the problem of time-consuming and costly generation of digital human facial animations in the existing technology is solved, and efficient and accurate personification is achieved.

CN120339471APending Publication Date: 2025-07-18湖北思极科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510329159.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the generation process of digital human facial animation is time-consuming, costly and poor reusability, making it difficult to achieve efficient and accurate personification improvement.

Method used

By training an animation generation model, audio coding network is used to extract audio timing features, combining emotion prediction network, feature connection network and implicit memory network, to generate expression control parameters, and achieve efficient generation of digital human animations.

Benefits of technology

It improves the efficiency and cost-effectiveness of digital human animation generation, the generated expression control parameters are more accurate, and it is adapted to a variety of scenes, which improves the anthropomorphism and fidelity of digital human facial animation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339471A_ABST
    Figure CN120339471A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a digital human animation generation method, a model training method and computing equipment. The method comprises the following steps: based on target audio data, extracting target audio time sequence characteristics of the target audio data by utilizing a frequency coding network in an animation generation model; based on the target audio time sequence features, predicting target emotion intensity features of the target audio data by using an emotion prediction network in an animation generation model; based on the target audio time sequence feature and the target emotion intensity feature, a feature connection network in an animation generation model is utilized to generate a target connection feature; generating target feature offset information of the target connection feature by using an implicit memory network in an animation generation model based on the target connection feature; generating a target expression control parameter by using an expression output network in an animation generation model based on the target connection feature and the target feature offset information; and generating a target digital human animation according to the target expression control parameter. According to the technical scheme of the embodiment of the invention, the anthropomorphic degree of the animation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of the metaverse, and in particular, to a method for generating digital human animations, a method for model training, and a computing device. Background Art

[0002] Under the upsurge of the metaverse, digital humans have begun to be involved in many fields such as culture and entertainment, services, and education. The types are also becoming more and more abundant, including functional digital humans, social digital humans, and companion digital humans, etc.

[0003] Since digital humans are digital human images created through digital technology and similar to human images, the anthropomorphism of digital human facial animations has always been an important indicator for measuring the quality of digital humans. At present, in order to ensure the anthropomorphism of digital human facial animations, it is usually manually made by experienced animators or realized based on facial motion capture technology, which is time-consuming, costly, and can only be applied to specific scenarios with poor reusability.

[0004] Therefore, how to efficiently and accurately improve the anthropomorphism of digital human facial animations is still an urgent problem to be solved at present. Summary of the Invention

[0005] The embodiments of the present application provide a method for generating digital human animations, a method for model training, and a computing device, which are used to solve the problems of long time consumption, high cost, and poor reusability in the process of digital human generation in the prior art.

[0006] In a first aspect, the embodiments of the present application provide a method for generating digital human animations, including:

[0007] Based on the target audio data, using the audio encoding network in the animation generation model, extract the target audio temporal features of the target audio data;

[0008] Based on the target audio temporal features, using the emotion prediction network in the animation generation model, predict the target emotion intensity features of the target audio data;

[0009] Based on the target audio temporal features and the target emotion intensity features, using the feature connection network in the animation generation model, generate target connection features;

[0010] Based on the target connection features, using the implicit memory network in the animation generation model, generate the target feature offset information of the target connection features;

[0011] Based on the target connection features and the target feature offset information, using the expression output network in the animation generation model, generate target expression control parameters;

[0012] Generate a target digital human animation according to the target expression control parameter.

[0013] In a second aspect, an embodiment of the present application provides a model training method, including:

[0014] Obtain sample audio data and the true expression control parameter corresponding to the sample audio data; wherein, the sample audio data is recorded when a sample user reads a preset text in a candidate emotion manner;

[0015] Based on the sample audio data, use the audio encoding network in the model to be trained to extract the sample audio time series features of the sample audio data;

[0016] Based on the sample audio time series features, use the emotion prediction network in the model to be trained to predict the sample emotion intensity features of the sample target audio data;

[0017] Based on the sample audio time series features and the sample emotion intensity features, use the feature connection network in the model to be trained to generate sample connection features;

[0018] Based on the sample connection features, use the implicit memory network in the model to be trained to generate the feature offset information of the sample connection features;

[0019] Based on the sample connection features and the sample feature offset information, use the expression output network in the model to be trained to generate sample expression control parameters;

[0020] Train the model to be trained according to the sample expression control parameter and the true expression control parameter to obtain an animation generation model.

[0021] In a third aspect, an embodiment of the present application provides a computing device, including a processing component and a storage component;

[0022] The storage component stores a computer program; the computer program is used to be called and executed by the processing component to implement the digital human animation generation method as described in the first aspect above or the model training method as described in the second aspect above.

[0023] In the embodiments of the present application, through the audio encoding network in the animation generation model, the target audio temporal features of the target audio data are extracted and input into the emotion prediction network of the animation generation model to predict the target emotion intensity features of the target audio data; then, through the feature connection network in the animation generation model, the target audio temporal features and the target emotion intensity features are spliced to generate target connection features, and the target connection features are input into the implicit memory network of the animation generation model to determine the target feature offset information; furthermore, through the expression output network in the animation generation model, based on the target connection features and the target feature offset information, target expression control parameters are generated; finally, according to the target expression control parameters, a target digital human animation is generated. In the embodiments of the present application, the animation generation model generates a digital human animation based on the temporal features and emotion intensity features of the audio data. Compared with the prior art of manual production or implementation based on facial motion capture technology, the generation efficiency of the digital human animation is greatly improved, the generation cost is reduced, and the trained animation generation model can be adapted to a variety of different scenarios with high reusability. In addition, when determining the expression control parameters, this solution not only combines the temporal features and emotion intensity features of the audio data, but also introduces the offset information of the features, making the generated expression control parameters more accurate, thereby achieving efficient and accurate improvement of the anthropomorphism of the digital human facial animation.

[0024] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0026] Figure 1 The system architecture diagram of the method for generating a digital human animation provided by the present application is shown;

[0027] Figure 2 The flowchart of an embodiment of a method for generating a digital human animation provided by the present application is shown;

[0028] Figure 3 The structural schematic diagram of the working principle of an animation generation model provided by the present application is shown;

[0029] Figure 4 The flowchart of an embodiment of a model training method provided by the present application is shown;

[0030] Figure 5 The structural schematic diagram of an embodiment of a digital animation generation device provided by the present application is shown;

[0031] Figure 6The structure diagram of an embodiment of a model training device provided by the present application is shown;

[0032] Figure 7 The structure diagram of an embodiment of a computing device provided by the present application is shown. Detailed implementation manners

[0033] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.

[0034] It should be noted that when the embodiments of the present application involve user information, the data involved in the embodiments of the present application (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select authorization or rejection. In addition, various models involved in the present application (including but not limited to animation generation models, models to be trained, AI large models, etc.) comply with relevant laws and standards.

[0035] The technical solution of the embodiment of the present application can be applied to the scenario of generating a digital human facial animation adapted to the target audio data. The inventor found in the process of implementing the present application that in the traditional method, whether it is manually made by a painter or implemented based on facial motion capture technology, there are problems of long time consumption, high cost and poor reusability. In order to efficiently and accurately improve the anthropomorphism of the digital human facial animation, the inventor found through research that with the rapid development of artificial intelligence technology, a neural network model can be used to realize the mapping from audio to animation control parameters, thereby improving the efficiency of digital human animation generation and reducing the labor cost. Therefore, the inventor proposed the technical solution of the embodiment of the present application. By training an animation generation model and using the model to analyze the temporal characteristics and emotional characteristics of audio data simultaneously, a digital human animation is generated, so that the generated animation contains emotional information, and thus the generated digital human animation is more vivid and diverse. In addition, on the basis of the above solution, the inventor carried out a series of considerations. Since there may be certain offset information in the features extracted by the model, which affects the fidelity of the digital human animation. Therefore, through the above innovative thinking, the inventor proposed to add an implicit memory network to analyze the offset information of the features to ensure the accuracy of the expression control parameters output by the model, and further improve the anthropomorphism and fidelity of the digital human animation on the premise of improving the efficiency.

[0036] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.

[0037] Figure 1 The system architecture diagram of the digital human animation generation method provided by the present application is shown. The system architecture may include a client 101 and a server 102. (Or: It may include a server and multiple clients; or it may include a server, a first client and a second client, etc.)

[0038] Among them, a connection can be established between the client 101 and the server 102 through a network. The network provides a medium for the communication link between the client 101 and the server 102. The network can include various connection types, such as wired, wireless or fiber optic cable, etc. The client 101 can interact with the server 102 through the network to receive or send messages, etc.

[0039] Among them, the client 101 can be a browser, an APP (Application), or a web application such as an H5 (HyperText Markup Language 5) application, or a light application (also known as a mini-program, a lightweight application) or a cloud application, etc. The client 101 can be deployed in a visualization electronic device and needs to rely on the device or certain apps in the device to run, etc. The electronic device can, for example, have a display screen and support information browsing, etc., such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, a desktop computer, a smart speaker, a smart watch, etc. For the sake of easy understanding, Figure 1 the client is mainly represented by the device image in Figure 1 . Various other types of applications can usually be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc. The electronic device can refer to a device used by a user and having functions such as computing, Internet access, and communication required by the user. The electronic device usually can include at least one processing component and at least one storage component. The electronic device may also include basic configurations such as a network card chip, an IO (input / output) bus, and audio-video components. This application does not limit this. Optionally, according to the implementation form of the electronic device, some peripheral devices may also be included, such as a keyboard, a mouse, an input pen, a printer, etc. This application does not limit this.

[0040] The server 102 can include servers that provide various services, such as a server that processes the interaction information sent by the client.

[0041] It should be noted that the server 102 can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms, or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology.

[0042] It should be noted that the digital human animation generation method provided in the embodiments of the present application is generally executed by the server 102. Correspondingly, the digital human animation generation device is generally set in the server 102. However, in other embodiments of the present application, the user terminal 101 can also have a similar function as the server 102, so as to execute the digital human animation generation method provided in the embodiments of the present application. In other embodiments, the digital human animation generation method provided in the embodiments of the present application can also be jointly executed by the user terminal 101 and the server 102.

[0043] It should be understood that Figure 1 the number of user terminals and servers in

[0044] The implementation details of the technical solutions in the embodiments of the present application will be elaborated in detail below.

[0045] Figure 2 The digital human animation generation method shown can include the following steps:

[0046] 201: Based on the target audio data, use the audio encoding network in the animation generation model to extract the target audio temporal features of the target audio data.

[0047] Among them, the target audio data can be the audio data that needs to be broadcast by the digital human. For example, if the digital human is a virtual teacher, the target audio data at this time can be the audio data of the teacher's teaching content. If the digital human is a virtual assistant providing question answering services for users, the target audio data at this time can be the audio data formed by the answers searched for the user's questions.

[0048] The target audio temporal features can refer to the characteristics shown by the target audio data in the time dimension, which reflect the variation law of the audio signal over time. The target audio temporal features can be used to reflect the expression change situation of the digital human in different time dimensions.

[0049] The animation generation model can be a model used to predict and control the expression control parameters of the digital human's expression change situation based on the target audio data. The animation generation model can be obtained through supervised training in advance based on the sample audio data and the corresponding real expression control parameters of the sample audio data. The specific training process will be introduced in detail in subsequent embodiments and will not be elaborated here. Optionally, Figure 3 shows a schematic structural diagram of the working principle of an animation generation model provided by the present application. As Figure 3As shown, the animation generation model 30 of this embodiment may be composed of an audio encoding network 301, an emotion prediction network 302, a feature connection network 303, an implicit memory network 304, and an expression output network 305. Among them, the audio encoding network 301 may be a neural network constructed based on a convolutional neural network for extracting audio temporal features from audio data. The other networks in the animation generation model 30 will be introduced in detail in subsequent embodiments.

[0050] Optionally, in this embodiment, the target audio data may be input into the audio encoding network 301 of the animation generation model 30, and the audio encoding network 301 will perform the operation of extracting the temporal features of the target audio data, so as to obtain the target audio temporal features of the target audio data. Specifically, the audio encoding network 301 may first sample the target audio data based on a preset sampling frequency, and then use the convolutional network therein to extract the temporal features of the sampled audio data based on the algorithms and parameters during training, and output the extraction result according to the preset driving output frequency, that is, output to the emotion prediction network 302 and the feature connection network 303. Optionally, the preset sampling frequency of this embodiment may be set to 16000 Hz to ensure the audio effect of the sampled audio data and prevent the loss of audio information. The driving output frequency of this embodiment may be set to 60 fps to ensure the accuracy of the output audio temporal features.

[0051] 202: Based on the target audio temporal features, use the emotion prediction network in the animation generation model to predict the target emotion intensity features of the target audio data.

[0052] Among them, the target emotion intensity features may be the feature values corresponding to the intensities of each candidate emotion when the digital human broadcasts the target audio data predicted by the emotion prediction network. The candidate emotions may be the basic emotions that the digital human may adopt preset in advance. For example, they may include, but are not limited to: anger, disgust, contempt, fear, happiness, sadness, surprise, etc. Optionally, the emotion prediction network 302 may analyze and predict the intensity features of each candidate emotion (i.e., the target emotion intensity features) that the digital human may show when broadcasting the target audio data based on the audio feature temporal features. In this embodiment, the target audio temporal features output by the audio encoding network 301 may be input into the emotion prediction network 302, and the emotion prediction network 302 will parse and predict the target emotion intensity features of the target audio data from the target audio temporal features based on the algorithms and parameters during training, and output the prediction result to the feature connection network 303.

[0053] In some embodiments, the above-mentioned emotion prediction network 302 includes a stacking layer and an embedding layer. The stacking layer is composed of at least two combined module residuals stacked; among them, the combined module can be composed of one-dimensional convolution, batch normalization, and rectified linear unit. For example, the combined module can be a CONV1D-BN-RELU module. Optionally, the way that at least two combined module residuals stack to form the stacking layer can be to stack multiple combined modules together to form a deep network, and add residual connections between every two combined modules, so that each combined module can directly obtain the input from the previous combined module and add the input to the output of the current combined module. The embedding layer of this embodiment is composed of a fully connected layer and a rectified linear unit. For example, the embedding layer can be an FC-RELU module.

[0054] Correspondingly, this step can be implemented in the following way at this time: use the stacking layer to perform emotion analysis on the target audio temporal features to obtain the emotion intensity of each candidate emotion; use the embedding layer to embed the emotion intensity of each candidate emotion into the target emotion intensity feature.

[0055] Among them, the emotion intensity of each candidate emotion can be the intensity value of each candidate emotion adopted when the digital human broadcasts the target audio data predicted by the emotion prediction network. The emotion intensity of each candidate emotion in this embodiment can be represented by an intensity vector, and the dimension of the intensity vector is the same as the type of preset candidate emotion. For example, if there are 7 preset candidate emotions, then the emotion intensity of each candidate emotion predicted at this time can be a 7-dimensional intensity vector, and each dimension represents the intensity value of a candidate emotion.

[0056] Specifically, it can be to first use the stacking layer (Net) constructed based on the residual stacking method to analyze the input target audio temporal feature audio_in, and parse the emotion intensity emotion_out = Net(audio_in) corresponding to each candidate emotion when the target audio data is broadcast, and then pass the emotion intensity emotion_out corresponding to each candidate emotion through the embedding layer and embed it into a feature vector, that is, the target emotion intensity feature.

[0057] In this embodiment, the method of introducing residual stacking when constructing the emotion prediction network effectively solves the problems of gradient disappearance and network degradation in the training of deep neural networks, improves the depth, training efficiency and stability of the network, ensures the accuracy of the emotion intensity of each candidate emotion predicted, and further ensures the accuracy of the generated target emotion intensity feature.

[0058] 203: Based on the target audio temporal feature and the target emotion intensity feature, use the feature connection network in the animation generation model to generate the target connection feature.

[0059] Among them, the feature connection network 303 can be a network for feature splicing. The target connection feature can be the feature after splicing the target audio time series feature and the target emotion intensity feature.

[0060] Optionally, one implementation can be to concatenate the target audio time series feature and the target emotion intensity feature through the feature connection network 303 to obtain the target connection feature. For example, the target emotion intensity feature can be added after the target audio time series feature, so as to merge the two features into a new feature as the target connection feature.

[0061] Another implementation can be to perform feature mapping and fusion on the target audio time series feature and the target emotion intensity feature through the feature connection network 303, and use the fused new feature as the target connection feature.

[0062] It should be noted that other splicing techniques can also be sampled in this embodiment, and this is not limited.

[0063] 204: Based on the target connection feature, use the implicit memory network in the animation generation model to generate the target feature offset information of the target connection feature.

[0064] Even though the animation generation model 30 in this embodiment has been trained, there will inevitably be some errors in its prediction results compared with the real results. The target feature offset information represents the offset amount of the target connection feature predicted by the model relative to the real target connection feature. The implicit memory network 304 is a neural network for evaluating the feature offset information of the target connection feature.

[0065] Optionally, in this embodiment, the target connection feature output by the feature connection network 303 can be input into the implicit memory network 304, and the implicit memory network 304 predicts the feature offset amount of the target connection feature based on the algorithm and parameters during training as the target feature offset information.

[0066] 205: Based on the target connection feature and the target feature offset information, use the expression output network in the animation generation model to generate the target expression control parameter.

[0067] Among them, the target expression control parameter can be control data for driving the expression of the digital human face model. Optionally, the target emotion control parameter can be a vector within a certain range. For example, it can be a vector between -1 and 1. The expression output network 305 can be a model for parsing the target link feature and the target feature offset information to determine the target expression control parameter. Optionally, the expression output network 305 can be similar to the stacked layer of the emotion prediction network 302, and is composed of at least two combined module residuals stacked by combining one-dimensional convolution, batch normalization, and rectified linear unit. The combined module can be a CONV1D-BN-RELU module.

[0068] Optionally, in this embodiment, the expression output network 305 can first perform a de-offset process on the target connection feature based on the target feature offset information. For example, the target feature offset information can be added to or subtracted from the target connection feature. Then, the control parameter analysis is performed on the de-offset target connection feature to generate the target expression control parameter.

[0069] 206: Generate a target digital human animation according to the target expression control parameter.

[0070] Among them, the target digital human animation can be a digital human animation adapted to the target audio data when playing the target audio data, that is, the facial expression of the digital human seems to be a real person broadcasting the target audio data.

[0071] Optionally, this embodiment can use digital human animation generation software to generate a target digital human animation based on the target expression control parameter. For example, in the parameter input interface of the digital human animation generation software, the corresponding target expression control parameter value can be input, and the pre-constructed expressionless digital human model can be imported, and then the generate button is clicked to obtain the target digital human animation generated by the digital human animation generation software.

[0072] It should be noted that the operation of generating the target digital human animation according to the target expression control parameter in this embodiment can be executed locally by the server. It can also be sent by the server to other terminals (such as other visualization terminals) for execution by other terminals, and this is not limited. The following embodiments will elaborate on these two scenarios in detail.

[0073] In this embodiment, the technical solution of the present application extracts the target audio temporal features of the target audio data through the audio encoding network in the animation generation model, and inputs them into the emotion prediction network of the animation generation model to predict the target emotion intensity features of the target audio data; then, through the feature connection network in the animation generation model, the target audio temporal features and the target emotion intensity features are spliced to generate target connection features, and the target connection features are input into the implicit memory network of the animation generation model to determine the target feature offset information; furthermore, through the expression output network in the animation generation model, based on the target connection features and the target feature offset information, target expression control parameters are generated; finally, based on the target expression control parameters, a target digital human animation is generated. The embodiment of the present application generates a digital human animation based on the temporal features and emotion intensity features of audio data through an animation generation model. Compared with the prior art of manual production or implementation based on facial motion capture technology, the generation efficiency of digital human animation is greatly improved, the generation cost is reduced, and the trained animation generation model can be adapted to a variety of different scenarios, with high reusability. In addition, when determining the expression control parameters, this solution not only combines the temporal features and emotion intensity features of audio data, but also introduces the offset information of the features, making the generated expression control parameters more accurate, thereby achieving an efficient and accurate improvement in the anthropomorphism of digital human facial animation.

[0074] In some embodiments, the implementation manner of the implicit memory network generating the target feature offset information of the target connection feature based on the target connection feature may be: using the implicit memory network in the animation generation model to perform weighted processing on the target connection feature and the feature vector parameter to obtain the target feature offset information of the target connection feature.

[0075] Among them, the feature vector parameter can be maintained in the implicit memory network and generated during the training process of the implicit memory network. That is to say, the feature vector parameter is a network parameter in the implicit memory network, and the specific value of the feature vector parameter is obtained through model training.

[0076] Optionally, an implementation manner may be to use the target connection feature as a weight to perform weighted processing on the feature vector parameter, and use the processing result as the target feature offset information of the target connection feature.

[0077] In order to improve the accuracy of determining the target feature offset information, another implementation manner of this embodiment may be to design the feature vector parameter to include a first matrix and a second matrix, and each row of the first matrix and the second matrix corresponds one by one. It should be noted that the first matrix and the second matrix are also network parameters of the implicit memory network. At this time, the first matrix and the second matrix may be matrices with the same dimension as the target connection feature, and the specific element values in the matrix can be determined during model training.

[0078] At this time, the implicit memory network in the animation generation model can be used to determine the similarity between the first matrix and the target connection feature, and the second matrix can be weighted according to the similarity to obtain the target feature offset information of the target connection feature. For example, the similarity between the first matrix and the target connection feature can be calculated by cosine similarity, Euclidean distance method or Pearson correlation coefficient method, and then the similarity is used as the weight to extract the features in the second matrix as the target feature offset information by weighted processing. Optionally, it can be implemented based on the following formula;

[0079] atten(Q,K,V)=[sim(Q,K)VW V ]W O ;

[0080] Among them, atten() is the weighted processing function; Q is the target connection feature; K is the first matrix; V is the second matrix; sim() is the similarity measurement function; Wv and W0 are weight parameters, which are also obtained based on the model training process.

[0081] This embodiment samples and calculates similarity, uses similarity as a weight, and extracts target offset feature information through weighted processing, thereby greatly improving the accuracy of the target offset feature information.

[0082] As can be seen from the foregoing description, the technical solution of the embodiment of the present application can be applied to the end-to-end generation of digital human animation scenarios, which can specifically involve multiple fields such as virtual customer service, virtual tour guides, and virtual family members. Figure 1 The user terminal 101 and the server terminal 102 shown in the figure are interactively implemented. The user terminal 101 in this scenario can be a visualization terminal that can provide query and answer services for users. The digital human animation generation method at this time may include: the visualization terminal receives user input information and sends it to the server terminal 102, the server terminal 102 responds to the user input information sent by the visualization terminal, generates target audio data based on the user input information, and then uses the animation generation model to generate target expression control parameters based on the target audio data, and then sends the target expression control parameters and the target audio data to the visualization terminal, so that the visualization terminal maps the target expression control parameters to the digital human face model, obtains the target digital human animation, and plays the target digital human animation and the target audio data at the same time.

[0083] Specifically, in this scenario, the user input information can be a question that the user needs to query or answer, and the target audio data is used to reply to the user input information. After receiving the user input information, the server can directly query the target audio data corresponding to the user input information by calling an AI large model or local search, or it can first find the text content for replying to the user input information by means of an AI large model or local search, and then select a corresponding voice model according to the attributes of the digital human to generate the audio corresponding to the text content as the target audio data. After obtaining the target audio data, based on the animation generation model, the target expression control parameters corresponding to the target audio data are generated in the manner described in the above embodiments, and the target expression control parameters and the target audio data are sent to the visualization terminal at the same time. The visualization terminal will first map the target expression control parameters onto the standard digital human model for rendering, that is, drive the expressions and emotions of the digital human model to obtain a digital human animation with changing expressions and emotions, and then play the target audio data and the target digital human animation to the user at the same time. Thus, it is realized that the user can see the digital human audio-visual animation that broadcasts the reply information on the visualization terminal. This method further improves the anthropomorphism and vividness of the played digital human animation on the premise of improving the efficiency of end-to-end generation of digital human animations.

[0084] It should be noted that the AI large model involved in this article refers to a large-parameter model trained using large-scale data and powerful computing capabilities, a machine learning model with a complex structure, which can process massive data and complete various complex tasks, such as natural language processing, computer vision, speech recognition, etc. It can include large language models (LLMs) or multimodal large models (MLMs), etc. The AI large model involved in this article can be a pre-trained model, and is retrained through model fine-tuning to adapt to different processing tasks. It not only utilizes the powerful capabilities of the pre-trained model, but also can adapt to the new data distribution. Therefore, it can ensure the generalization ability of the model and reduce the overfitting phenomenon.

[0085] As described above, the technical solution of the embodiment of the present application can also be applied to the scenario where a digital human animation scene is generated locally on a terminal. For example, in a scenario where the server needs to generate its own animation assets. At this time, the solution of the present application can be executed locally by the server. The digital human animation generation method at this time is similar to the method introduced in Embodiment 1. First, use the animation generation model to generate target expression control parameters based on the target audio data. Then, generate a parameter file in a preset format according to the target expression control parameters; based on this parameter file, use an animation production application to generate a target digital human animation. Among them, the preset format can be the file format supported by the animation production application for input. In this embodiment, the target expression control parameters can be first converted into a parameter text in the format supported by the animation production application (i.e., the preset format) for storage. Then, when generating the digital human animation, a 3D software for animation production (such as maya or 3dmax, etc.) is used to call the saved parameter file to produce a digital human animation in a general format as its own animation asset. In this scenario, compared with the prior art where animators manually draw animation assets or produce animation assets based on facial motion capture technology, the present solution can efficiently and accurately generate digital human animations using the animation generation model and can be reused in multiple scenarios.

[0086] As Figure 4 shown, it is a flowchart of an embodiment of a model training method provided by the present application. This method can be executed by the server before executing the digital human animation generation method and can include the following steps:

[0087] 401: Obtain sample audio data and the true expression control parameters corresponding to the sample audio data.

[0088] Among them, the sample audio data is recorded when the sample user reads the preset text in a candidate emotion manner. The candidate emotions can be the basic emotions that the digital human may adopt and are preset, for example, can include but are not limited to: angry, disgusted, contemptuous, fearful, happy, sad, surprised, etc. The preset text can be the text content that needs to be read and is selected in advance.

[0089] Specifically, in this embodiment, multiple different users can be found as sample users, and then each sample user is allowed to read the preset text in a different candidate emotion and record the audio data during the user's reading as the sample audio data. Then, for each sample audio data, the true expression control parameters of the sample audio data can be determined according to the actual expression and emotion during the reading and the parameter setting rules set in advance. Facial video data of the sample user reading the preset text can also be recorded and analyzed to determine the true expression control parameters. This embodiment does not limit this.

[0090] Optionally, when obtaining the sample audio data in this embodiment, in addition to asking the sample user to read the preset text with the above candidate emotions, it is also possible to read it in a non-emotional manner to improve the diversity and richness of the sample data.

[0091] 402: Based on the sample audio data, use the audio encoding network in the model to be trained to extract the sample audio time series features of the sample audio data.

[0092] 403: Based on the sample audio time series features, use the emotion prediction network in the model to be trained to predict the sample emotion intensity features of the sample target audio data.

[0093] 404: Based on the sample audio time series features and the sample emotion intensity features, use the feature connection network in the model to be trained to generate the sample connection features.

[0094] 405: Based on the sample connection features, use the implicit memory network in the model to be trained to generate the feature offset information of the sample connection features.

[0095] 406: Based on the sample connection features and the sample feature offset information, use the expression output network in the model to be trained to generate the sample expression control parameters.

[0096] It should be noted that the specific implementation methods of 402-406 in this embodiment are similar to the method of using the animation generation model in the above digital human animation generation method to generate the target sample control parameters based on the target audio data, and will not be elaborated here.

[0097] 407: According to the sample expression control parameters and the real expression control parameters, train the model to be trained to obtain the animation generation model.

[0098] Optionally, this embodiment can use the real expression control parameters as the supervision data, use the pre-set loss function to analyze the difference between the real expression control parameters and the sample expression control parameters as the loss value. For example, this embodiment can use the mean square error between the real expression control parameters and the sample expression control parameters, that is, the model reconstruction loss, as the loss function at this time. Then, based on the determined loss value, perform multiple rounds of supervised training on the model to be trained until the loss value is less than the preset value, that is, minimize the loss, and the training process is completed. At this time, the trained model to be trained is the animation generation model.

[0099] In this embodiment, the sampled audio data adopted is recorded by multiple users with various candidate emotions, which improves the diversity of training samples. When training the animation generation model, the temporal features and emotion features of the audio data are combined, and the feature offset information is also considered, which greatly ensures the accuracy of the expression control parameters output by the trained animation generation model. Subsequently, by using this animation generation model, realistic digital human animations can be generated efficiently and accurately.

[0100] In some embodiments, to further improve the model training accuracy, during the model training process, a multi-dimensional loss function can be introduced to train the model. At this time, when obtaining the true expression control parameters corresponding to the sampled frequency data and the sampled audio data in 401, the true emotion intensity and the true expression parameter speed corresponding to the sampled audio data can also be obtained. The true emotion intensity can be the intensity value of the emotion adopted by the sampled user when reading the sampled audio data in an actual situation. The actual expression parameter speed can be the speed of obtaining the user's true expression control parameters. In this embodiment, the method of obtaining the true emotion intensity and the true expression parameter speed can be similar to the method of obtaining the true expression control parameters described above. For example, it can be determined according to the actual expressions and emotions of the sampled user when reading the preset text according to the pre-set mapping rules. The facial video data of the sampled user reading the preset text can also be recorded, and the recorded facial video data can be analyzed to determine the true emotion intensity and the true expression parameter speed. This embodiment does not limit this.

[0101] It should be noted that the output of the animation generation model trained in this embodiment may further include data related to the emotion intensity and the expression parameter speed in addition to the expression control parameters. However, since these two types of data are not required during the process of generating digital human animations, the above embodiments only give the preferred solution of the animation generation model outputting the target expression control parameters. However, during the model training stage, to improve the model accuracy, when calculating the loss, the emotion intensity and the expression parameter speed output by the animation generation model can be further combined. At this time, when generating the sample expression control parameters based on the sample connection features and the sample feature offset information using the expression output network in the model to be trained in the above step 406, the sample emotion intensity and the sample expression parameter speed can also be generated simultaneously (that is, based on the sample connection features and the sample feature offset information, using the expression output network in the model to be trained, the sample emotion intensity and the sample expression parameter speed are generated). The sample emotion intensity can be the intensity value of the emotion predicted by the model when the sampled user reads the sampled audio data. The sample expression parameter speed can be the speed of the model predicting the sample expression control parameters.

[0102] Correspondingly, at this time, step 407 trains the model to be trained according to the sample expression control parameters and the real expression control parameters to obtain the specific implementation manner of the animation generation model, including the following sub-steps:

[0103] Sub-step 1: Determine the model reconstruction loss according to the sample expression control parameters and the real expression control parameters.

[0104] Optionally, the specific implementation manner of this sub-step can be similar to the implementation manner of step 407 in the above embodiment, and the loss determined in step 407 is used as the model reconstruction loss of this sub-step. Details are not described here.

[0105] Sub-step 2: Determine the emotion matching loss according to the sample emotion intensity and the real emotion intensity.

[0106] Among them, the emotion matching loss can be a loss function representing the degree of emotion matching.

[0107] Optionally, this sub-step can use a pre-set loss function to analyze the difference between the sample emotion intensity and the real emotion intensity as the emotion matching loss. For example, in this embodiment, the mean square error between the sample emotion intensity and the real emotion intensity can be used as the emotion matching loss.

[0108] Sub-step 3: Determine the reconstruction speed loss according to the sample expression parameter speed and the real expression parameter speed.

[0109] Among them, the reconstruction speed loss can be a loss function representing the model prediction speed.

[0110] Optionally, this sub-step can use a pre-set loss function to analyze the difference between the sample expression parameter speed and the real expression parameter speed as the reconstruction speed loss. For example, in this embodiment, the mean square error between the sample expression parameter speed and the real expression parameter speed can be used as the reconstruction speed loss.

[0111] Sub-step 4: Determine the implicit orthogonality loss according to the feature vector parameters maintained in the implicit memory network of the model to be trained.

[0112] Among them, the implicit orthogonality loss can be a loss used to measure whether the feature vector parameters in the implicit memory network are accurate.

[0113] Optionally, if the feature vector parameter maintained in the implicit memory network is a parameter matrix, then at this time, each row element of the feature vector parameter can be orthogonally operated with the other row elements, and the operation result is used as the implicit orthogonal loss. If the implicit memory network includes a first matrix and a second matrix, then at this time, the second matrix can be obtained from the feature vector parameter maintained by the implicit memory network of the model to be trained; then each row element of the second matrix is orthogonally operated with the other row elements except this row to obtain the implicit orthogonal loss.

[0114] Sub-step 5: Train the model to be trained according to the model reconstruction loss, emotion matching loss, reconstruction speed loss, and implicit orthogonal loss to obtain an animation generation model.

[0115] Optionally, when training the model to be trained according to the model reconstruction loss, emotion matching loss, reconstruction speed loss, and implicit orthogonal loss in this embodiment, weight values can be set for each type of loss in advance according to the importance of each type of loss, and then, according to the preset weight values, the model reconstruction loss, emotion matching loss, reconstruction speed loss, and implicit orthogonal loss are weighted and summed, and the sum result is used as the target loss function. Then, based on the target loss function, the model to be trained is trained to obtain an animation generation model.

[0116] It should be noted that when training the model, the model reconstruction loss, emotion matching loss, and reconstruction speed loss need to be minimized to ensure that the predicted value is closer to the true value. And the implicit orthogonal loss needs to be maximized to ensure that the data of each row of the feature vector parameter maintained in the implicit memory network are as different as possible, thereby improving the accuracy of determining the offset feature information.

[0117] In this embodiment, a multi-dimensional loss function of model reconstruction loss, emotion matching loss, reconstruction speed loss, and implicit orthogonal loss is used to train the model, which greatly improves the accuracy of model training. Thereby ensuring the accuracy of the target expression control parameters output by the trained animation generation model, and further ensuring the anthropomorphism and vividness of the digital human animation mapped based on the target expression control parameters.

[0118] It should be noted that in some of the processes described in the above embodiments and accompanying drawings, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations are only used to distinguish the different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., do not represent a sequence, and do not limit that "first" and "second" are of different types.

[0119] Figure 5 FIG. 4 is a schematic structural diagram of an embodiment of a digital human animation generation device provided by an embodiment of the present application. The device may include:

[0120] A model operation module 501, configured to extract target audio timing features of the target audio data by using an audio encoding network in an animation generation model based on the target audio data; predict target emotion intensity features of the target audio data by using an emotion prediction network in the animation generation model based on the target audio timing features; generate target connection features by using a feature connection network in the animation generation model based on the target audio timing features and the target emotion intensity features; generate target feature offset information of the target connection features by using an implicit memory network in the animation generation model based on the target connection features; and generate target expression control parameters by using an expression output network in the animation generation model based on the target connection features and the target feature offset information.

[0121] An animation generation module 502, configured to generate a target digital human animation according to the target expression control parameters.

[0122] In some embodiments, the model operation module 501 is specifically configured to perform weighted processing on the target connection features and feature vector parameters by using an implicit memory network in the animation generation model to obtain target feature offset information of the target connection features; wherein, the feature vector parameters are maintained in the implicit memory network and are generated during the training process of the implicit memory network.

[0123] In some embodiments, the feature vector parameters include a first matrix and a second matrix, and each row of the first matrix and the second matrix corresponds one by one; correspondingly, the model operation module 501 is specifically configured to determine the similarity between the first matrix and the target connection features by using an implicit memory network in the animation generation model, and perform weighted processing on the second matrix according to the similarity to obtain target feature offset information of the target connection features.

[0124] In some embodiments, the emotion prediction network includes a stacked layer and an embedding layer, and the stacked layer is composed of at least two combined module residuals stacked; correspondingly, the model operation module 501 is specifically configured to use the stacked layer to perform emotion analysis on the target audio time series features to obtain the emotion intensities of each candidate emotion; and use the embedding layer to embed the emotion intensities of each candidate emotion into target emotion intensity features.

[0125] In some embodiments, the animation generation module 502 is specifically configured to send the target expression control parameter and the target audio data to a visualization terminal, so that the visualization terminal maps the target expression control parameter to a digital human face model to obtain a target digital human animation, and simultaneously play the target digital human animation and the target audio data.

[0126] In some embodiments, the device may further include:

[0127] An audio data generation module, configured to generate target audio data in response to user input information sent by the visualization terminal, based on the user input information; wherein, the target audio data is used to reply to the user input information.

[0128] Figure 6 The figure is a schematic structural diagram of an embodiment of a model training device provided by an embodiment of the present application. The device may include:

[0129] A data acquisition module 601, configured to acquire sample audio data and true expression control parameters corresponding to the sample audio data; wherein, the sample audio data is recorded when a sample user reads a preset text in a candidate emotion manner.

[0130] A parameter determination module 602, configured to extract sample audio time series features of the sample audio data by using an audio encoding network in a to-be-trained model based on the sample audio data; predict sample emotion intensity features of the sample target audio data by using an emotion prediction network in the to-be-trained model based on the sample audio time series features; generate sample connection features by using a feature connection network in the to-be-trained model based on the sample audio time series features and the sample emotion intensity features; generate feature offset information of the sample connection features by using an implicit memory network in the to-be-trained model based on the sample connection features; and generate sample expression control parameters by using an expression output network in the to-be-trained model based on the sample connection features and the sample feature offset information.

[0131] A model training module 603, configured to train the to-be-trained model according to the sample expression control parameters and the true expression control parameters to obtain an animation generation model.

[0132] In some embodiments, the data acquisition module 601 is further configured to acquire the true emotion intensity and the true expression parameter speed corresponding to the sample audio data.

[0133] The parameter determination module 602 is further configured to generate a sample emotion intensity and a sample expression parameter speed by using the expression output network in the to-be-trained model based on the sample connection feature and the sample feature offset information.

[0134] The model training module 603 is specifically configured to determine a model reconstruction loss according to the sample expression control parameter and the true expression control parameter; determine an emotion matching loss according to the sample emotion intensity and the true emotion intensity; determine a reconstruction speed loss according to the sample expression parameter speed and the true expression parameter speed; determine an implicit orthogonality loss according to the feature vector parameters maintained in the implicit memory network of the to-be-trained model; and train the to-be-trained model according to the model reconstruction loss, the emotion matching loss, the reconstruction speed loss, and the implicit orthogonality loss to obtain an animation generation model.

[0135] In some embodiments, the model training module 603 is further specifically configured to obtain a second matrix from the feature vector parameters maintained in the implicit memory network of the to-be-trained model; perform an orthogonality operation on each row element of the second matrix with the other row elements except this row to obtain the implicit orthogonality loss.

[0136] Figure 5 The digital human animation generation device described above can execute Figure 2 the digital human animation generation method described in the illustrated embodiment, Figure 6 The model training device described above can execute Figure 4 the model training method described in the illustrated embodiment, and its implementation principle and technical effects will not be elaborated further. For the Figure 5 and Figure 6 devices, the specific manners in which each module and unit perform operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0137] Figure 7 FIG. is a schematic structural diagram of an embodiment of a computing device provided by the present application. As Figure 7 shown, in practice, the computing device may include: a storage component 1001 and a processing component 1002.

[0138] The storage component 1001 is configured to store computer programs and can be configured to store various other data to support operations on the computing device. Examples of these data include instructions for any application program or method for operating on the computing device, data structures, contact data, phone book data, messages, pictures, videos, etc.

[0139] The processing component 1002, coupled to the storage component 1001, is configured to execute the computer program in the storage component 1001 to implement the digital human animation generation method as shown in Figure 2 or the model training method as shown in Figure 4 .

[0140] Furthermore, as shown in Figure 7 , the computing device may further include other components such as a communication component 1003, a display component 1004, a power supply component 1005, and an audio component 1006. Figure 7 Only some components are schematically shown in Figure 7 , which does not mean that the computing device only includes the components shown in Figure 7 . In addition, the components within the dashed box in Figure 7 are optional components, not mandatory components, and are specifically determined according to the product form of the computing device. The computing device of this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, or an IOT (Internet of Things) device, or can also be a server device such as a conventional server, a cloud server, or a server array. If the computing device of this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, or a smart phone, it may include the components within the dashed box in Figure 7 ; if the computing device of this embodiment is implemented as a server device such as a conventional server, a cloud server, or a server array, it may not include the components within the dashed box in

[0141] The above-mentioned processing component includes one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component can also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for executing the above method.

[0142] The above storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0143] The above communication component is configured to facilitate communication between the device where the communication component is located and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on a communication standard, such as a mobile communication network, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel.

[0144] The above display component may include a screen, and the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation.

[0145] The above power supply component provides power to various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device where the power supply component is located.

[0146] The above audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), and when the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory or transmitted via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals.

[0147] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above method embodiment. Among them, the computer-readable storage medium can be implemented by volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium

[0148] Accordingly, an embodiment of the present application further provides a computer program product, which includes a computer program or instruction that, when executed by a processor, enables the processor to implement the steps in the above method embodiment. It should be understood that each process or a combination of multiple processes in the above method flow can be implemented by the computer program or instruction. In addition, these computer programs or instructions can be applied to the processors of general-purpose computers, special-purpose computers, embedded processors or other programmable data processing devices, so that the processors of general-purpose computers, special-purpose computers, embedded processors or other programmable data processing devices can be used as devices to implement the corresponding functions in the above method embodiment.

[0149] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0150] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0151] Finally, it should be noted that the above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for generating digital human animations, characterized in that, The method includes: Based on the target audio data, using the audio encoding network in the animation generation model, extract the target audio temporal features of the target audio data; Based on the target audio temporal features, using the emotion prediction network in the animation generation model, predict the target emotion intensity features of the target audio data; Based on the target audio temporal features and the target emotion intensity features, using the feature connection network in the animation generation model, generate target connection features; Based on the target connection features, using the implicit memory network in the animation generation model, generate the target feature offset information of the target connection features; Based on the target connection features and the target feature offset information, using the expression output network in the animation generation model, generate target expression control parameters; Generate a target digital human animation according to the target expression control parameters.

2. The method according to claim 1, wherein The generating the target feature offset information of the target connection features based on the target connection features and using the implicit memory network in the animation generation model includes: Using the implicit memory network in the animation generation model, perform weighted processing on the target connection features and the feature vector parameters to obtain the target feature offset information of the target connection features; Wherein, the feature vector parameters are maintained in the implicit memory network and are generated during the training process of the implicit memory network.

3. The method according to claim 2, wherein The feature vector parameters include a first matrix and a second matrix, and each row of the first matrix and the second matrix corresponds one by one; The performing weighted processing on the target connection features and the feature vector parameters using the implicit memory network in the animation generation model to obtain the target feature offset information of the target connection features includes: Using the implicit memory network in the animation generation model, determine the similarity between the first matrix and the target connection features, and according to the similarity, perform weighted processing on the second matrix to obtain the target feature offset information of the target connection features.

4. The method according to claim 1, wherein The emotion prediction network includes a stacking layer and an embedding layer, and the stacking layer is composed of at least two combined module residuals stacked; The predicting the target emotion intensity features of the target audio data based on the target audio temporal features and using the emotion prediction network in the animation generation model includes: Using the stacking layer, perform emotion analysis on the target audio temporal features to obtain the emotion intensity of each candidate emotion; Using the embedding layer, embed the emotion intensity of each candidate emotion into the target emotion intensity features.

5. The method according to any one of claims 1-4, characterized in that, The generating a target digital human animation according to the target expression control parameters includes: Send the target expression control parameters and the target audio data to a visualization terminal, so that the visualization terminal maps the target expression control parameters to the digital human face model to obtain a target digital human animation, and simultaneously play the target digital human animation and the target audio data.

6. The method according to claim 5, wherein It further includes: In response to the user input information sent by the visualization terminal, generate target audio data based on the user input information; wherein, the target audio data is used to reply to the user input information.

7. A model training method, characterized in that, The method includes: Obtain sample audio data and the true expression control parameters corresponding to the sample audio data; wherein, the sample audio data is recorded when a sample user reads a preset text in a candidate emotion manner; Based on the sample audio data, use the audio encoding network in the model to be trained to extract the sample audio time series features of the sample audio data; Based on the sample audio time series features, use the emotion prediction network in the model to be trained to predict the sample emotion intensity features of the sample target audio data; Based on the sample audio time series features and the sample emotion intensity features, use the feature connection network in the model to be trained to generate sample connection features; Based on the sample connection features, use the implicit memory network in the model to be trained to generate the feature offset information of the sample connection features; Based on the sample connection features and the sample feature offset information, use the expression output network in the model to be trained to generate sample expression control parameters; According to the sample expression control parameters and the true expression control parameters, train the model to be trained to obtain an animation generation model.

8. The method according to claim 7, wherein The method further includes: Obtain the true emotion intensity and the true expression parameter speed corresponding to the sample audio data; Based on the sample connection features and the sample feature offset information, use the expression output network in the model to be trained to generate the sample emotion intensity and the sample expression parameter speed; The step of training the model to be trained according to the sample expression control parameters and the true expression control parameters to obtain an animation generation model includes: Determine the model reconstruction loss according to the sample expression control parameters and the true expression control parameters; Determine the emotion matching loss according to the sample emotion intensity and the true emotion intensity; Determine the reconstruction speed loss according to the sample expression parameter speed and the true expression parameter speed; Determine the implicit orthogonality loss according to the feature vector parameters maintained in the implicit memory network of the model to be trained; Train the model to be trained according to the model reconstruction loss, the emotion matching loss, the reconstruction speed loss and the implicit orthogonality loss to obtain an animation generation model.

9. The method according to claim 8, wherein The step of determining the implicit orthogonality loss according to the feature vector parameters maintained in the implicit memory network of the model to be trained includes: Obtain a second matrix from the feature vector parameters maintained in the implicit memory network of the model to be trained; Perform an orthogonality operation on each row element of the second matrix with the other row elements except this row to obtain the implicit orthogonality loss.

10. A computing device, characterized in that, Includes a processing component and a storage component; The storage component stores a computer program; the computer program is used to be called and executed by the processing component to implement the digital human animation generation method according to any one of claims 1-6, or the model training method according to any one of claims 7-9.