3D virtual digital human action generation method and device, readable storage medium and equipment

By constructing multilingual data sets and using language IDs, combined with a deep neural network of diffusion models, the problem of poor performance in multilingual speech is solved, and high-quality and diverse 3D virtual digital human action generation is achieved.

CN120014127APending Publication Date: 2025-05-16INNER MONGOLIA UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510071483.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art does not perform well in pronunciation in other languages ​​except English, and it is difficult to effectively generate multilingual collaborative pronunciation actions.

Method used

Using a deep neural network based on diffusion model, the model is trained to generate 3D virtual digital human animations that match the corresponding speech by constructing multilingual data sets and using language IDs.

Benefits of technology

The performance and generalization capabilities of the model on multilingual data are improved, and the generated actions are both high-quality and diverse, solving the problem of poor performance of existing methods in multilingual pronunciation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014127A_ABST
    Figure CN120014127A_ABST
Patent Text Reader

Abstract

The invention provides a 3D virtual digital human action generation method and device, a readable storage medium and equipment. The method comprises the steps that 1, a multi-language data set is constructed based on a BEAT data set; 2, constructing a deep neural network based on a diffusion model; 3, training a deep neural network based on a diffusion model by using the constructed multilingual data set; and 4, performing model reasoning based on the trained deep neural network, and generating a 3D virtual digital human animation matched with the corresponding voice. According to the method, the diffusion model is utilized, so that the generated action has high quality and diversity; the constructed multilingual data set and the language ID are utilized to help the model to distinguish different languages, so that the model can understand the characteristics of differences and different languages, and the performance and generalization ability of the model on multilingual data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of 3D virtual digital human, and in particular to a method, apparatus, readable storage medium and device for generating actions of a 3D virtual digital human. Background Art

[0002] Virtual digital human refers to a virtual image with human or human-like appearance and behavior patterns produced through modeling, motion capture, AI and other technical means, and presented through display devices. In the early days, virtual digital humans were mainly presented through 2D images and simple animations. This form of virtual digital human production was relatively simple, but the interactivity and realism were limited. With the development of computer graphics (CG) technology, virtual digital humans began to shift to 3D, which made the virtual image more three-dimensional and realistic, bringing a more immersive experience to users. After entering the 3D era, virtual digital humans began to adopt three-dimensional modeling technology, which not only increased the dimension of information, but also increased the required computing power. 3D virtual digital humans can show more realistic appearance and behavior patterns through more sophisticated modeling and rendering technology. The realism of virtual digital human movements is the most important part of measuring whether the generated digital human is natural. From a technical perspective, 3D virtual digital human motion drive can be divided into algorithm-driven (AI real-time) and real-person driven (motion capture). The current mainstream solution for real-person-driven is to expand 3D scanning to 4D scanning and add a time dimension, so as to collect the most subtle facial changes and body movements of the actors, and then apply the collected action data files to the virtual image. However, the amount of data is huge, the post-processing time and manpower cost are too heavy, and it is not suitable for large-scale asset production.

[0003] In recent years, with the advancement of artificial intelligence technology, especially the development of deep learning models, the motion generation technology of 3D virtual digital humans has been significantly improved. The application of AI technology enables virtual digital humans to generate corresponding actions and expressions based on multiple input modalities such as voice, text, and sensors, greatly improving their interactivity and application scope. The correlation between voice and facial expressions and body movements during speaking is very high. Among them, the change of lip shape is directly related to voice. In addition, people's speech is often accompanied by changes in facial expressions and certain body movements. Therefore, it is very important to explore how to use voice as input to directly generate the lip shape expression and action sequence of the virtual image.

[0004] Most of the existing methods for virtual human motion generation focus on deterministic deep learning methods, which have the problem of mode collapse, resulting in low synthesis quality, especially when the data used is not included in the training data, or a trade-off needs to be made between generation quality and diversity.

[0005] 3D virtual human co-speech action generation is an important part of the human-computer interaction process. Co-speech action refers to the non-verbal behaviors such as gestures, facial expressions, and body postures made by the speaker when speaking. These non-verbal behaviors can help the audience understand the speaker's meaning and enhance the communication effect. However, most of the current research on co-speech action generation uses English data for training, without focusing on the co-speech action generation of multilingual speech, nor studying the model's generalization ability to multilingual data. In addition, these methods use pre-trained models to extract audio features, and cannot handle speech in languages ​​other than English well for co-speech action generation. Summary of the invention

[0006] In view of this, the present application provides a 3D virtual digital human motion generation method, apparatus, readable storage medium and device to overcome the problem that the prior art methods perform poorly in speech of languages ​​other than English.

[0007] To achieve the above objectives, the technical solutions adopted in this application are as follows: A 3D virtual human action generation method, comprising: Step 1: Build a multilingual dataset based on the BEAT dataset; Step 2: Construct a deep neural network based on the diffusion model based on the principle of the diffusion model; Step 3: Use the constructed multilingual dataset to train the deep neural network based on the diffusion model; Step 4: Perform model inference based on the trained deep neural network to generate a 3D virtual human animation that matches the corresponding voice.

[0008] Furthermore, the principle of the diffusion model in step 2 includes a diffusion process and a denoising process, and the diffusion process is specifically: According to the Markov chain rule, Gaussian noise is added to the motion sequence data in the multilingual dataset to approximate the posterior After the diffusion process is completed, the data The probability distribution is equivalent to an isotropic Gaussian distribution: in, represents the diffusion process at the time step A sample of Indicates that at time step The intensity of the noise added, represents the chain process of noise distribution transition, represents the total number of steps in the diffusion process, represents a normal distribution, represents the mean, represents the covariance matrix, is the identity matrix; The denoising process is specifically as follows: in, represents the reverse process of recovering samples from noise, represents a normal distribution, represents the mean of the conditional distribution, represents the covariance matrix of the conditional distribution; make , , then at time Noise Action It can be expressed as: in, represents a normal distribution, represents the mean, represents the covariance matrix; The model is based on the input sample , denoising steps And condition C to predict the original signal , the condition C includes seed action, voice and language ID.

[0009] Furthermore, the specific method of using the constructed multilingual dataset in step 3 to train the deep neural network based on the diffusion model is: Step 3.1: extracting low-level speech features from the original audio of the multilingual dataset, wherein the low-level speech features are universal features in the speech signal that are independent of the language; Step 3.2: Combine the extracted low-level speech features into a sequence, concatenate it with the facial parameter sequence after noise processing, and then send it to the facial decoder; add noise to the seed gesture and concatenate it with the low-level speech features, and then concatenate the language ID to enable the model to learn the differences between speech in different languages, and then send it to the cross-local attention layer and self-attention layer to capture the relationship between low-level speech features and gestures; Step 3.3: Calculate the loss function between the facial parameter sequence decoded and output by the facial decoder and the facial parameter sequence in the multilingual dataset for training; calculate the loss function between the action sequence generated by the self-attention layer and the action sequence in the multilingual dataset for training.

[0010] Furthermore, the loss function between the action sequence generated by the self-attention layer and the action sequence in the multilingual dataset calculated in step 3.3 is specifically: in, Represents the generated action sequence, i.e., the model prediction sample; Denotes the action sequence in the multilingual dataset, i.e., the original data sample; Denoise denotes the trained denoising model; represents the loss function; Represents two random variables and joint expectations; Represents data samples From the data distribution Medium sampling; Represents the time step From the time range Uniform sampling in Represents data samples under condition C The true probability distribution of ; HuberLoss represents the loss function.

[0011] A 3D virtual human motion generation device, comprising: The dataset construction module is used to build a multilingual dataset based on the BEAT dataset; A neural network building module is used to build a deep neural network based on the diffusion model based on the principle of the diffusion model; A training module, used to train a deep neural network based on a diffusion model using the constructed multilingual dataset; The inference module is used to perform model inference based on the trained deep neural network and generate 3D virtual digital human animation that matches the corresponding voice.

[0012] Furthermore, the principle of the diffusion model includes a diffusion process and a denoising process, and the diffusion process is specifically: According to the Markov chain rule, Gaussian noise is added to the motion sequence data in the multilingual dataset to approximate the posterior After the diffusion process is completed, the data The probability distribution is equivalent to an isotropic Gaussian distribution: in, represents the diffusion process at the time step A sample of Indicates that at time step The intensity of the noise added, represents the chain process of noise distribution transition, represents the total number of steps in the diffusion process, represents a normal distribution, represents the mean, represents the covariance matrix, is the identity matrix; The denoising process is specifically as follows: in, represents the reverse process of recovering samples from noise, represents a normal distribution, represents the mean of the conditional distribution, represents the covariance matrix of the conditional distribution; make , , then at time Noise Action It can be expressed as: in, represents a normal distribution, represents the mean, represents the covariance matrix; The model is based on the input sample , denoising steps and condition C to predict the original signal , the condition C includes seed action, voice and language ID.

[0013] Furthermore, the training module is specifically used to perform the following steps: Step 3.1: extracting low-level speech features from the original audio of the multilingual dataset, wherein the low-level speech features are universal features in the speech signal that are independent of the language; Step 3.2: Combine the extracted low-level speech features into a sequence, concatenate it with the facial parameter sequence after noise processing, and then send it to the facial decoder; add noise to the seed gesture and concatenate it with the low-level speech features, and then concatenate the language ID to enable the model to learn the differences between speech in different languages, and then send it to the cross-local attention layer and self-attention layer to capture the relationship between low-level speech features and gestures; Step 3.3: Calculate the loss function between the facial parameter sequence decoded and output by the facial decoder and the facial parameter sequence in the multilingual dataset for training; calculate the loss function between the action sequence generated by the self-attention layer and the action sequence in the multilingual dataset for training.

[0014] Furthermore, the loss function between the action sequence generated by the self-attention layer and the action sequence in the multilingual dataset calculated in step 3.3 is specifically: in, Represents the generated action sequence, i.e., the model prediction sample; Denotes the action sequence in the multilingual dataset, i.e., the original data sample; Denoise denotes the trained denoising model; represents the loss function; Represents two random variables and joint expectations; Represents data samples From the data distribution Medium sampling; Represents the time step From the time range Uniform sampling in Represents data samples under condition C The true probability distribution of ; HuberLoss represents the loss function.

[0015] According to another aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the steps in a 3D virtual digital human action generation method of the present application are implemented.

[0016] According to another aspect of the present application, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in a method for generating actions of a 3D virtual digital human of the present application are implemented.

[0017] Compared with the prior art, the beneficial effects of this application are: 1. This application combines the theoretical principles of the diffusion model with the framework design of the neural network, and maps the forward diffusion process and reverse generation process of the diffusion model to the training and generation of the neural network, so that the generated actions have both high quality and diversity; 2. Use the constructed multilingual datasets and language IDs to help the model distinguish between different languages, so that the model can understand the differences and characteristics between different languages, thereby improving its performance and generalization ability on multilingual data; 3. The extracted low-level speech features are used for training, which solves the problem that existing methods perform poorly on speech in languages ​​other than English. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0019] Figure 1A flow chart of a method for generating 3D virtual human actions for this application; Figure 2 Schematic diagram of the training process and reasoning process in the deep neural network of this application; Figure 3 This is a structural block diagram of a 3D virtual digital human motion generation device for this application. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments.

[0021] like Figure 1 and 2 As shown, a 3D virtual human action generation method includes: Step 1: Build a multilingual dataset based on the BEAT dataset; A multilingual dataset for training was built based on the public BEAT dataset, which contains speech data, corresponding facial Blendshape data, and body motion capture data. The BEAT dataset contains 76 hours of data from 30 speakers in four languages: English (60 hours), Chinese (12 hours), Spanish (2 hours), and Japanese (2 hours). Body movements were captured by professional motion capture equipment, and a total of 75 joint movements were recorded. A mobile phone with a 3D depth camera was used to capture facial expression and lip shape parameters, and the weights of 52 facial Blendshapes were extracted using the ARKit framework developed by Apple.

[0022] Based on the BEAT dataset, this application extracts, for example, Speakers 2 and 10 as English subsets, the Chinese part of speakers 12, 13, 14, 22 and 24 as a Chinese subset, the Spanish part of speakers 15 and 16 as a Spanish subset, and the Japanese part of speakers 17 and 25 as a Japanese subset.

[0023] Step 2: Construct a deep neural network based on the diffusion model based on the principle of the diffusion model; Specifically, the principle of the diffusion model includes a diffusion process and a denoising process, and the diffusion process is: According to the Markov chain rule, Gaussian noise is added to the motion sequence data in the multilingual dataset to approximate the posterior After the diffusion process is completed, the data The probability distribution is equivalent to an isotropic Gaussian distribution: in, represents the diffusion process at the time step A sample of Indicates that at time step The intensity of the noise added, usually is an increasing sequence between [0,1], indicating that As it increases, the noise added to the sample gradually increases. represents the chain process of noise distribution transition, represents the total number of steps in the diffusion process, represents a normal distribution, represents the mean, represents the covariance matrix, is the identity matrix; The diffusion process from Start by adding noise until you get , is an almost completely random noise sample. In this process, the sample gradually "diffuses" into the noise space, the noise gradually increases, and the final distribution can be described by a chain process of distribution transition, that is, the above formula (1).

[0024] The denoising process is specifically as follows: in, represents the reverse process of recovering samples from noise; represents normal distribution; Represents the mean of the conditional distribution, which determines the main trend of the generated samples and represents the result of denoising; The covariance matrix representing the conditional distribution is usually a diagonal matrix, which represents the residual part of the uncertainty or noise in generating the samples; make , used to control the retention ratio of noise; , represents the cumulative retention ratio; then at time Noise Action It can be expressed as: in, represents normal distribution; Represents the mean, which retains the components of the original data, but after multiple steps of diffusion, the components gradually decay; represents the covariance matrix, which describes the cumulative strength of the noise over time As the noise increases, it gradually becomes dominant; the model , denoising steps and condition C to predict the original signal , the condition C includes seed action, voice and language ID.

[0025] In order to recover samples from noise, it is necessary to train a reverse process model that learns how to recover the original data from noise. In the diffusion model, the reverse process is usually expressed as a conditional probability distribution, as shown in the above formula (3).

[0026] This application is based on a diffusion model, which can make the generated actions both high-quality and diverse.

[0027] Step 3: Use the constructed multilingual dataset to train the deep neural network based on the diffusion model; The specific method is: Step 3.1: extracting low-level speech features from the original audio of the multilingual dataset, wherein the low-level speech features are universal features in the speech signal that are independent of the language; MFCC (Mel Cepstral Coefficient), Mel Spectrum, Pitch, Energy, and Onsets (note or beat features) are extracted from the original audio and concatenated into low-level speech features. These low-level speech features are universal features extracted from speech signals and are independent of the language. In addition, human gestures and facial movements change with the intensity, rhythm, and tempo of speech. The extracted features are concatenated to obtain low-level speech features. Because different representations can complement each other, using more speech features can achieve better performance. After processing the speech in the entire data and training it, the model can learn the association between low-level features of speech in multiple languages ​​and expressions, lip shapes, and movements.

[0028] Step 3.2: Combine the extracted low-level speech features into a sequence, concatenate it with the facial parameter sequence after noise processing, and then send it to the facial decoder; add noise to the seed gesture and concatenate it with the low-level speech features, and then concatenate the language ID to enable the model to learn the differences between speech in different languages, and then send it to the cross-local attention layer and self-attention layer to capture the relationship between low-level speech features and gestures; The language ID may be different encoding data, and specifically may use one-hot encoding, for example, using 0001 to represent English, 0010 to represent Chinese, 0011 to represent Spanish, and 0100 to represent Japanese.

[0029] The constructed multilingual datasets and language IDs are used to help the model distinguish between different languages, so that the model can understand the differences and characteristics between different languages, thereby improving its performance and generalization ability on multilingual data.

[0030] Step 3.3: Calculate the loss function between the facial parameter sequence decoded and output by the facial decoder and the facial parameter sequence in the multilingual dataset for training; calculate the loss function between the action sequence generated by the self-attention layer and the action sequence in the multilingual dataset for training.

[0031] The extracted low-level speech features are used for training, which solves the problem that existing methods perform poorly on speech in languages ​​other than English. Finally, the trained model can generate high-quality and diverse parameters of facial expressions, lip shapes, and body movements through the input of speech in any language (which may not be in the dataset).

[0032] Furthermore, the loss function between the action sequence generated by the self-attention layer and the action sequence in the multilingual dataset calculated in step 3.3 is specifically: in, Represents the generated action sequence, i.e., the model prediction sample; Denotes the action sequence in the multilingual dataset, i.e., the original data sample; Denoise denotes the trained denoising model; represents the loss function; Represents two random variables and joint expectations; Represents data samples From the data distribution Medium sampling; Represents the time step From the time range Uniform sampling in Represents data samples under condition C The true probability distribution of ; HuberLoss represents the loss function.

[0033] In the diffusion process, at each time step, the data samples will be noisy, making them gradually close to Gaussian noise; in the denoising process, the model will perform denoising training on the noisy data at different time steps, so it is necessary to Take samples.

[0034] Step 4: Perform model inference based on the trained deep neural network to generate a 3D virtual human animation that matches the corresponding voice.

[0035] The trained model is used for model inference to generate the corresponding action file, and then a digital human animation matching the corresponding voice is generated through operations such as virtual character parameter binding and animation rendering.

[0036] The inference of the model is an iterative process with a time step decreasing from T to 1. The initial noise is represented by the actual noise in the normal distribution N(0,1). At each step, the network is provided with conditions such as audio and noisy action sequences. The predicted motion is then diffused again and fed to the next step of the iteration. Since the model uses language-independent low-level speech features during training, the trained model can handle some languages ​​and voices that were not used during training, and the generated lip shapes, expressions, and actions have a high degree of correlation and matching with the speech.

[0037] like Figure 3 As shown, a 3D virtual digital human motion generation device includes: A data set construction module 310, used to construct a multilingual data set based on the BEAT data set; A neural network construction module 320 is used to construct a deep neural network based on the diffusion model based on the principle of the diffusion model; A training module 330, for training a deep neural network based on a diffusion model using the constructed multilingual dataset; The reasoning module 340 is used to perform model reasoning based on the trained deep neural network to generate a 3D virtual digital human animation that matches the corresponding voice.

[0038] Furthermore, the neural network construction module 320 is specifically used for: The principle of the diffusion model includes a diffusion process and a denoising process. The diffusion process is specifically as follows: According to the Markov chain rule, Gaussian noise is added to the motion sequence data in the multilingual dataset to approximate the posterior After the diffusion process is completed, the data The probability distribution is equivalent to an isotropic Gaussian distribution: in, represents the diffusion process at the time step A sample of Indicates that at time step The intensity of the noise added, represents the chain process of noise distribution transition, represents the total number of steps in the diffusion process; The denoising process is specifically as follows: in, represents the reverse process of recovering samples from noise, Indicates that the model parameters and noise samples To predict the recovery sample; make , , then the noise action at time t It can be expressed as: The model is based on the input sample , denoising steps And condition C to predict the original signal , the condition C includes seed action, voice and language ID.

[0039] Furthermore, the training module 330 is specifically configured to perform the following steps: Step 3.1: extracting low-level speech features from the original audio of the multilingual dataset, wherein the low-level speech features are universal features in the speech signal that are independent of the language; Step 3.2: Combine the extracted low-level speech features into a sequence, concatenate it with the facial parameter sequence after noise processing, and then send it to the facial decoder; add noise to the seed gesture and concatenate it with the low-level speech features, and then concatenate the language ID to enable the model to learn the differences between speech in different languages, and then send it to the cross-local attention layer and self-attention layer to capture the relationship between low-level speech features and gestures; Step 3.3: Calculate the loss function between the facial parameter sequence decoded and output by the facial decoder and the facial parameter sequence in the multilingual dataset for training; calculate the loss function between the gesture action sequence generated by the self-attention layer and the gesture action sequence in the multilingual dataset for training.

[0040] Furthermore, the loss function between the action sequence generated by the self-attention layer and the action sequence in the multilingual dataset calculated in step 3.3 is specifically: in, Represents the generated action sequence, i.e., the model prediction sample; Denotes the action sequence in the multilingual dataset, i.e., the original data sample; Denoise denotes the trained denoising model; represents the loss function; Represents two random variables and joint expectations; Represents data samples From the data distribution Medium sampling; Represents the time step From the time range Uniform sampling in Represents data samples under condition C The true probability distribution of ; HuberLoss represents the loss function.

[0041] As for the device embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is only illustrative, and the units described as separate components may or may not be physically separated.

[0042] The components may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units, such as distributed on a server and a client. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art may understand and implement it without creative work.

[0043] According to another aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the steps in a 3D virtual digital human action generation method of the present application are implemented.

[0044] According to another aspect of the present application, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in a method for generating actions of a 3D virtual digital human of the present application are implemented.

[0045] This application uses a diffusion model to make the generated actions both high-quality and diverse; uses a constructed multilingual dataset and language ID to help the model distinguish different languages, so that the model can understand the differences and characteristics between different languages, thereby improving its performance and generalization ability on multilingual data; uses the extracted low-level speech features for training to solve the problem that previous methods perform poorly on speech in languages ​​other than English; finally, the trained model can generate high-quality and diverse facial expressions, lip shapes, and related parameters of body movements through the input of speech in any language (which may not be in the dataset).

[0046] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A method for generating 3D virtual human motion, characterized in that: include: Step 1: Build a multilingual dataset based on the BEAT dataset; Step 2: Construct a deep neural network based on the diffusion model based on the principle of the diffusion model; Step 3: Use the constructed multilingual dataset to train the deep neural network based on the diffusion model; Step 4: Perform model inference based on the trained deep neural network to generate a 3D virtual human animation that matches the corresponding voice.

2. A 3D virtual human motion generation method as claimed in claim 1, characterized in that: The principle of the diffusion model in step 2 includes a diffusion process and a denoising process, and the diffusion process is specifically: According to the Markov chain rule, Gaussian noise is added to the motion sequence data in the multilingual dataset to approximate the posterior After the diffusion process is completed, the data The probability distribution is equivalent to an isotropic Gaussian distribution: ; in, represents the diffusion process at the time step A sample of Indicates that at time step The intensity of the noise added, represents the chain process of noise distribution transition, represents the total number of steps in the diffusion process, represents a normal distribution, represents the mean, represents the covariance matrix, is the identity matrix; The denoising process is specifically as follows: in, represents the reverse process of recovering samples from noise, represents a normal distribution, represents the mean of the conditional distribution, represents the covariance matrix of the conditional distribution; make , , then at time Noise Action It can be expressed as: in, represents a normal distribution, represents the mean, represents the covariance matrix; The model is based on the input sample , denoising steps and condition C to predict the original signal , the condition C includes seed action, voice and language ID.

3. A 3D virtual human motion generation method as claimed in claim 2, characterized in that: The specific method of using the constructed multilingual dataset in step 3 to train the deep neural network based on the diffusion model is: Step 3.1: extracting low-level speech features from the original audio of the multilingual dataset, wherein the low-level speech features are universal features in the speech signal that are independent of the language; Step 3.2: Combine the extracted low-level speech features into a sequence, concatenate it with the facial parameter sequence after noise processing, and then send it to the facial decoder; add noise to the seed gesture and concatenate it with the low-level speech features, and then concatenate the language ID to enable the model to learn the differences between speech in different languages, and then send it to the cross-local attention layer and self-attention layer to capture the relationship between low-level speech features and gestures; Step 3.3: Calculate the loss function between the facial parameter sequence decoded and output by the facial decoder and the facial parameter sequence in the multilingual dataset for training; The loss function between the action sequence generated by the self-attention layer and the action sequence in the multilingual dataset is calculated for training.

4. A 3D virtual human motion generation method as claimed in claim 3, characterized in that: The loss function between the action sequence generated by the self-attention layer and the action sequence in the multilingual dataset is calculated in step 3.3 as follows: in, Represents the generated action sequence, i.e., the model prediction sample; Denotes the action sequence in the multilingual dataset, i.e., the original data sample; Denoise denotes the trained denoising model; represents the loss function; Represents two random variables and joint expectations; Represents data samples From the data distribution Medium sampling; Represents the time step From the time range Uniform sampling in Represents data samples under condition C The true probability distribution of ; HuberLoss represents the loss function.

5. A 3D virtual human motion generation device, characterized in that: include: The dataset construction module is used to build a multilingual dataset based on the BEAT dataset; A neural network building module for building a deep neural network based on the diffusion model based on the principle of the diffusion model; A training module, used to train a deep neural network based on a diffusion model using the constructed multilingual dataset; The inference module is used to perform model inference based on the trained deep neural network and generate 3D virtual digital human animation that matches the corresponding voice.

6. A 3D virtual human motion generation device as claimed in claim 5, characterized in that: The principle of the diffusion model includes a diffusion process and a denoising process. The diffusion process is specifically as follows: According to the Markov chain rule, Gaussian noise is added to the motion sequence data in the multilingual dataset to approximate the posterior After the diffusion process is completed, the data The probability distribution is equivalent to an isotropic Gaussian distribution: in, represents the diffusion process at the time step A sample of Indicates that at time step The intensity of the noise added, represents the chain process of noise distribution transition, represents the total number of steps in the diffusion process, represents a normal distribution, represents the mean, represents the covariance matrix, is the identity matrix; The denoising process is specifically as follows: in, represents the reverse process of recovering samples from noise, represents a normal distribution, represents the mean of the conditional distribution, represents the covariance matrix of the conditional distribution; make , , then at time Noise gesture It can be expressed as: in, represents a normal distribution, represents the mean, represents the covariance matrix; The model is based on the input sample , denoising steps and condition C to predict the original signal , the condition C includes seed action, voice and language ID.

7. The 3D virtual human motion generation device according to claim 6, characterized in that: The training module is specifically used to perform the following steps: Step 3.1: extracting low-level speech features from the original audio of the multilingual dataset, wherein the low-level speech features are universal features in the speech signal that are independent of the language; Step 3.2: Combine the extracted low-level speech features into a sequence, concatenate it with the facial parameter sequence after noise processing, and then send it to the facial decoder; add noise to the seed gesture and concatenate it with the low-level speech features, and then concatenate the language ID to enable the model to learn the differences between speech in different languages, and then send it to the cross-local attention layer and self-attention layer to capture the relationship between low-level speech features and gestures; Step 3.3: Calculate the loss function between the facial parameter sequence decoded and output by the facial decoder and the facial parameter sequence in the multilingual dataset for training; The loss function between the gesture action sequence generated by the self-attention layer and the gesture action sequence in the multilingual dataset is calculated for training.

8. The 3D virtual human motion generation device according to claim 7, characterized in that: The loss function between the gesture sequence generated by the self-attention layer and the gesture sequence in the multilingual dataset is calculated in step 3.3 as follows: in, Represents the generated action sequence, i.e., the model prediction sample; Denotes the action sequence in the multilingual dataset, i.e., the original data sample; Denoise denotes the trained denoising model; represents the loss function; Represents two random variables and joint expectations; Represents data samples From the data distribution Medium sampling; Represents the time step From the time range Uniform sampling in Represents data samples under condition C The true probability distribution of ; HuberLoss represents the loss function.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the method for generating actions of a 3D virtual digital human described in any one of claims 1 to 4 are implemented.

10. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the method for generating actions of a 3D virtual digital human as described in any one of claims 1 to 4 are implemented.