Method and device for generating 3D digital human motion sequence, electronic equipment and storage medium
By inputting multimodal data into a trained limb motion generation model, and combining it with a diffusion model and a body and hand motion generation model, high-quality 3D digital human motion sequences are generated. This solves the problem of poor motion sequence quality in existing technologies and achieves better coordination and fluency between 3D digital human motion and real human motion.
Patent Information
- Application Number
- CN202510276216.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-03-10
AI Technical Summary
The motion sequences of 3D digital humans in the current technology are of poor quality and there is a large gap between them and the motion sequences of real humans, resulting in a large difference between the movements of 3D digital humans and those of real humans.
A multimodal data input model for training body motion generation is used, including speech data, key motion data, style type data, and lower body walking motion sequences. A diffusion model is used to generate motion sequences of a 3D digital human. Noise is generated by generating Gaussian noise data. The body motion generation model and hand motion generation model are combined to generate body motion sequences and hand motion sequences.
The generated motion sequences are of higher quality, with better performance in terms of overall coordination, smoothness, rationality, and aesthetics. The 3D digital human's movements are closer to those of a real human.
Smart Images

Figure CN120355823B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of 3D motion generation technology, and in particular to a method, apparatus, electronic device, and storage medium for generating 3D digital human motion sequences. Background Technology
[0002] 3D virtual digital humans, or "3D digital humans" for short, are digitally created avatars that closely resemble real humans. They possess human characteristics such as appearance, gender, and personality, and have expressive abilities including voice, facial expressions, and body language. Understandably, the more closely a 3D digital human's appearance and expressive abilities resemble a real human, the better. Considering the natural consistency in voice, facial expressions, and body language among real humans, 3D digital humans should also exhibit consistency in these aspects to make them more closely resemble real humans overall.
[0003] For motion sequences used to drive 3D digital humans to perform actions, existing technologies can generate motion sequences by inputting speech data into the Speech To Animation (STA) model. However, the quality of motion sequences generated by the STA model in existing technologies is poor, and there is a large gap between them and real human motion sequences, resulting in a large difference between the actions performed by 3D digital humans and those of real humans.
[0004] Therefore, in order to make the movements of 3D digital humans more similar to those of real humans, it is necessary to consider improving the quality of the motion sequence. Summary of the Invention
[0005] This invention provides a method, apparatus, electronic device, and storage medium for generating 3D digital human motion sequences, in order to solve the technical problem of poor quality of 3D digital human motion sequences in the prior art.
[0006] This invention provides a method for generating 3D digital human motion sequences, comprising:
[0007] Receive first multimodal data of a 3D digital human, the first multimodal data including voice data;
[0008] The first multimodal data is input into the trained limb motion generation model to generate the motion sequence of the 3D digital human;
[0009] The training set of the limb movement generation model includes a first multimodal data sample and corresponding movement sequence samples.
[0010] According to a method for generating a 3D digital human motion sequence provided by the present invention, the first multimodal data further includes at least one of key motion data, style type data, and lower body walking motion sequence.
[0011] According to the present invention, a method for generating a 3D digital human motion sequence is provided, wherein the limb motion generation model is a diffusion model;
[0012] The step of inputting the first multimodal data into the trained limb movement generation model to generate the movement sequence of the 3D digital human includes:
[0013] Generate Gaussian noise data;
[0014] Using the first multimodal data as a condition of the diffusion model, the Gaussian noise data is denoised to generate the motion sequence of the 3D digital human.
[0015] According to a method for generating a 3D digital human motion sequence provided by the present invention, the first multimodal data includes a lower body walking motion sequence, the limb motion generation model includes a body motion generation model and a hand motion generation model, and the motion sequence includes a body motion sequence and a hand motion sequence.
[0016] The step of inputting the first multimodal data into the trained limb movement generation model to generate the movement sequence of the 3D digital human includes:
[0017] The first multimodal data is input into the body motion generation model to generate the body motion sequence;
[0018] The lower body walking action sequence in the first multimodal data is replaced with the body action sequence to obtain the second multimodal data;
[0019] The second multimodal data is input into the hand motion generation model to generate the hand motion sequence;
[0020] The training set of the body motion generation model includes the first multimodal data samples and the corresponding body motion sequence samples, and the training set of the hand motion generation model includes the second multimodal data samples and the corresponding hand motion sequence samples.
[0021] According to the present invention, a method for generating 3D digital human motion sequences is provided, wherein the diffusion model is a UNet network architecture, the UNet network consists of several encoding blocks and decoding blocks, and each encoding block or decoding block contains a convolutional neural network and a Transformer neural network.
[0022] According to a method for generating a 3D digital human motion sequence provided by the present invention, the step of inputting the first multimodal data into a trained limb motion generation model includes:
[0023] Use a speech recognition model to extract the speech features from the speech data;
[0024] The speech features are input into the body movement generation model.
[0025] According to a method for generating a 3D digital human motion sequence provided by the present invention, before receiving the first multimodal data of the 3D digital human, the method further includes:
[0026] Receive dialogue text;
[0027] The dialogue text is converted into speech data using speech synthesis technology.
[0028] The present invention also provides a 3D digital human driving method, comprising:
[0029] Generate a 3D digital human motion sequence using any of the methods described above.
[0030] The 3D digital human motion sequence is redirected.
[0031] The 3D digital human is driven using the redirected motion sequence of the 3D digital human.
[0032] This invention also provides a method for generating 3D digital human videos, comprising:
[0033] Drive the 3D digital human using the 3D digital human driving method described above;
[0034] The 3D digital human is rendered to generate a 3D digital human video.
[0035] The present invention also provides a device for generating 3D digital human motion sequences, comprising:
[0036] A receiving module is used to receive the first multimodal data of the 3D digital human, the first multimodal data including voice data;
[0037] The generation module is used to input the first multimodal data into the trained limb movement generation model to generate the movement sequence of the 3D digital human;
[0038] The training set of the limb movement generation model includes a first multimodal data sample and corresponding movement sequence samples.
[0039] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements any of the above-described methods for generating 3D digital human motion sequences, driving 3D digital humans, or generating 3D digital human videos.
[0040] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for generating a 3D digital human motion sequence, a 3D digital human driving method, or a 3D digital human video generation method as described above.
[0041] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above-described methods for generating 3D digital human motion sequences, driving 3D digital humans, or generating 3D digital human videos.
[0042] The present invention provides a method, apparatus, electronic device, and storage medium for generating 3D digital human motion sequences. It inputs multimodal data, including speech data, into a limb motion generation model to generate motion sequences for a 3D digital human. Compared to existing technologies that use Speech-to-Animation (STA) technology to generate 3D digital human motion sequences, the generated motion sequences are of higher quality, and the resulting 3D digital human exhibits superior performance in terms of overall coordination, motion smoothness, motion rationality, and aesthetic appeal. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0044] Figure 1 This is a schematic diagram of a frame of motion data in a 3D digital human motion sequence.
[0045] Figure 2 This is a flowchart illustrating the method for generating 3D digital human motion sequences provided by the present invention.
[0046] Figure 3 This is a schematic diagram illustrating the principle of forward diffusion in the diffusion model provided by this invention.
[0047] Figure 4 This is a schematic diagram illustrating the principle of reverse diffusion in the diffusion model provided by this invention.
[0048] Figure 5 This is a flowchart illustrating step S2 provided by the present invention.
[0049] Figure 6 This is a schematic diagram of generating body movement sequences provided by the present invention.
[0050] Figure 7This is a schematic diagram of the generation of hand movement sequences provided by the present invention.
[0051] Figure 8 This is a schematic diagram of the UNet network architecture provided by the present invention.
[0052] Figure 9 This is a flowchart illustrating the 3D digital human driving method provided by the present invention.
[0053] Figure 10 This is a flowchart illustrating the 3D digital human video generation method provided by the present invention.
[0054] Figure 11 This is a schematic diagram of the structure of the 3D digital human motion sequence generation device provided by the present invention.
[0055] Figure 12 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0057] It should be noted that in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. The terms "upper," "lower," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Unless otherwise expressly specified and limited, the terms "installed," "connected," and "linked" should be interpreted broadly, for example, as a fixed connection, a detachable connection, or an integral connection; a mechanical connection or an electrical connection; a direct connection or an indirect connection through an intermediate medium; or a connection within two elements. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0058] The terms "first," "second," etc., used in this invention are used to distinguish similar objects, not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0059] 3D virtual digital humans (or simply "3D digital humans") refer to digitally created human figures that closely resemble human appearances. They possess specific physical characteristics such as appearance, gender, and personality, and can express themselves through voice, facial expressions, and body language. Furthermore, 3D digital humans can be categorized into interactive and non-interactive types. Interactive 3D digital humans also possess the ability to recognize their environment and interact with people.
[0060] In simple terms, a 3D digital human needs at least a specific physical appearance as its external image, as well as expressive abilities such as voice, facial expressions, and body movements. The closer its appearance and expressive abilities are to those of a real human, the better. This is understandable, as real humans naturally possess consistency in their speech, facial expressions, and body movements. This consistency includes not only the fluency and coherence of speech, the naturalness of facial expressions, and the coordination of body movements, but also the overall consistency among speech, facial expressions, and body movements.
[0061] In related industries, full-body motion capture systems are typically used to collect motion sequences from real humans and record them as three-dimensional spatial coordinates of key points in the body. These motion sequences are then input into software, where redirection algorithms can drive any 3D digital human model to perform the same movements as a real human. In other words, by continuously collecting motion sequences from real humans, 3D digital humans can be made to perform various movements just like real humans.
[0062] A frame of motion data in a 3D digital human motion sequence, such as Figure 1As shown, the motion sequence refers to a set of motion data frames arranged chronologically, used to drive 3D digital humans to simulate the movement and behavior of real humans. It's important to note that in the video world, each second of change consists of several frames, called "frames," and the number of frames per second is called the "frame rate." Each frame is a still image; displaying frames rapidly and continuously creates the illusion of motion. Therefore, animations with higher frame rates are smoother and more realistic. In scenarios requiring higher smoothness, such as live sports broadcasts and e-sports video production, higher frame rates of 50 or 60 frames per second are used. Thus, the continuous images in a video are actually composed of several still images. Similarly, the motion sequence used to drive 3D digital humans to simulate human movement and behavior is also composed of frame-by-frame motion data arranged chronologically.
[0063] Figure 1 The single frame of motion data shown contains the positions of all skeletal points in space. Hundreds of consecutive frames of motion data represent the positional changes of these skeletal points within a corresponding time period. By using these hundreds of consecutive frames of motion data to drive the 3D digital human, the 3D digital human can continuously perform corresponding actions within a given time period.
[0064] However, the above solutions are costly, requiring not only expensive motion capture equipment but also professional motion capture actors to perform various movements in order to capture motion sequences.
[0065] With the continuous development of computer software technology, it is now possible to automatically generate motion sequences using computer software to drive 3D digital human models to perform corresponding body movements. Although some movements are still somewhat stiff, they can already meet the basic needs of users.
[0066] As explained above, 3D digital humans express themselves through various means, including speech, facial expressions, and body movements. To make a 3D digital human's expressive abilities closely resemble those of a real human, not only is coordinated body movement necessary, but the different forms of expression must also maintain overall consistency. For example, when a 3D digital human expresses happiness or excitement through speech, its facial expressions and body movements should also convey happiness or excitement, which aligns with the expressive habits of real humans.
[0067] To ensure consistency between the speech and body movements of 3D digital humans, related products employ Speech-to-Animation (STA) technology to generate motion sequences for the 3D digital humans, achieving uniformity between their speech and body movements. Specifically, by inputting a speech into the STA model, it automatically generates corresponding motion sequences based on the speech's content and tone, driving the 3D digital human model to speak while simultaneously performing the corresponding body movements. However, there are still discrepancies between the motion sequences generated by the STA model and those directly captured from real humans, resulting in differences in overall coordination, smoothness, and plausibility of the driven 3D digital human's body movements compared to real humans.
[0068] The following is combined Figures 2-12 This invention describes the method, apparatus, electronic device, and storage medium for generating 3D digital human motion sequences.
[0069] Figure 2 This is a flowchart illustrating the method for generating 3D digital human motion sequences provided by the present invention, as shown below. Figure 2 As shown, including but not limited to the following steps:
[0070] Step S1: Receive the first multimodal data of the 3D digital human, which includes voice data.
[0071] Multimodal data refers to datasets containing multiple types of data, such as images, text, audio, and video. Speech data exists in the form of audio. Therefore, multimodal data must include at least two types of data.
[0072] In some implementations, voice data can be acquired using an audio acquisition device. Audio is understood to be data that changes continuously over time, possessing temporal coherence and continuity. Audio corresponds to a timeline, the length of which determines the duration of the audio, and each point in time on the timeline can be marked with a timestamp.
[0073] Step S2: Input the first multimodal data into the trained limb motion generation model to generate a motion sequence of a 3D digital human.
[0074] The training set for the limb motion generation model includes first multimodal data samples and corresponding motion sequence samples.
[0075] In some implementations, the limb motion generation model can be one of the following: Generative Adversarial Networks (GAN), Transformer model, diffusion model, pre-trained model, and Contrastive Language-Image Pre-Training (CLIP).
[0076] Compared to existing technologies that use Speech To Animation (STA) to generate motion sequences for 3D digital humans, this invention inputs multimodal data, including speech data, into a training model to generate motion sequences for 3D digital humans. The generated motion sequences are of higher quality, and the 3D digital humans driven by this invention perform better in terms of overall coordination, motion smoothness, motion rationality, and aesthetic appeal.
[0077] In some embodiments, the first multimodal data of the present invention may further include at least one of key motion data, style type data, and lower body walking motion sequences.
[0078] In this dataset, speech data, key motion data, and lower body walking sequence are all time-related. Speech data is continuous, while key motion data and lower body walking sequence are discrete. Style type data can be encoded to map different style types to different discrete numbers. For example, key motion data can be in matrix form, where the number of columns represents the number of frames corresponding to the key motion (i.e., the duration of the key motion), and the set of numbers in each column represents one frame of key motion data. Similarly, lower body walking sequence can also be in matrix form, where the number of columns represents the number of frames in the lower body walking sequence, and the set of numbers in each column represents one frame of lower body walking data. Style type data can be in vector form, where the discrete number combinations within the vector correspond to different style types.
[0079] The key motion data refers to the data corresponding to the key actions that the user needs the 3D digital human to perform. For example, if the user needs the 3D digital human to wave at a specified time point (determined by a timestamp on the timeline), then the key motion data is the data corresponding to the specified time point and the wave action. If the first multimodal data includes key motion data, then step S2 can generate a motion sequence that matches the key motion data, further improving the quality of the motion sequence. After driving the 3D digital human based on the motion sequence, the 3D digital human can perform actions that are coordinated with the key actions.
[0080] In this invention, style type data represents the style of the generated action sequence, such as happy, sad, serious, lively, etc. The user can specify the style type for the action sequence to be generated and input it into the trained limb movement generation model in the form of style type data to generate action sequences of the specified style type. For example, if the user pre-specifies the style type of the generated action sequence as happy, then the action sequence generated by the limb movement generation model will drive the 3D digital human to perform a large number of actions representing happiness. That is, if the first multimodal data includes style type data, then step S2 can generate action sequences of the corresponding style type, allowing the user to specify the style type of the action sequence. After driving the 3D digital human based on the action sequence, the 3D digital human can perform actions consistent with the style type specified by the user.
[0081] In some scenarios, the 3D digital human needs to walk along a user-specified route. The lower body walking motion sequence is a collection of frames of data corresponding to the walking motion. The lower body can be understood as the area below the waist. If the first multimodal data includes the lower body walking motion sequence, step S2 can generate a motion sequence that matches the lower body walking motion sequence, allowing the user to specify the walking route based on the motion sequence. After driving the 3D digital human based on the motion sequence, the 3D digital human can perform actions that are coordinated with the walking motion.
[0082] It is understood that the first multimodal data of the present invention contains at least two types of data. In order to facilitate input into the limb movement generation model, the present invention can pre-fuse multiple types of data to obtain fused data, and input the fused data as a whole into the trained limb movement generation model.
[0083] In some embodiments, the limb movement generation model of the present invention can be a diffusion model;
[0084] Step S2 may include:
[0085] Generate Gaussian noise data;
[0086] Using the first multimodal data as a condition for the diffusion model, the Gaussian noise data is denoised to generate the motion sequence of the 3D digital human.
[0087] The working mechanism of the diffusion model consists of two stages: forward diffusion and reverse diffusion. Figure 3 This is the positive diffusion stage. Figure 4 This is the back diffusion stage. T represents the total number of noise addition or denoising steps, t represents the numerical marker corresponding to the current step (the numerical markers for noise addition steps are an integer sequence increasing from 0 to T, and the numerical markers for denoising steps are an integer sequence decreasing from T to 0), x0 represents the regular data, and x... T Given Gaussian noise data, x tThis is the data after adding noise at step t.
[0088] like Figure 3 As shown, forward diffusion is a process of gradually increasing data complexity. A simple sample is randomly drawn from a sample set that follows a target distribution. Through a series of reversible, stepwise modifications, a certain amount of noise (structured and controllable) is introduced into the sample at each step. As noise is continuously introduced, the complexity of the sample increases, eventually resulting in a sample very similar to the desired complex data distribution (such as a Gaussian distribution), thus transforming a data distribution into a noise distribution. In other words, it involves continuously introducing noise into regular data x0 to obtain Gaussian noise x. T .
[0089] like Figure 4 As shown, backdiffusion is the process of generating regular data x0, which can be obtained from Gaussian noise x. T Recover the regular data x0. First, randomly generate Gaussian noise x. T Then, the operation opposite to the forward diffusion is performed step by step to remove noise and thus recover the regular data x0.
[0090] It's important to note that the step-by-step noise removal process in backdiffusion can be defined as backsampling. Backsampling is implemented using a backsampler, so the key is to construct a backsampler that meets the requirements. Backsamplers include denoising diffusion probabilistic models (DDPM) and denoising diffusion implicit models (DDIM). The DDPM sampler defines a stochastic mapping, while the DDIM sampler defines a deterministic mapping. Specifically, the DDPM sampler is used to sample x... T After noise reduction processing, the output x0 and the input x T Almost independent, x t Independent of x T It only depends on x from the previous step. t-1 This forms a reverse Markov chain. Using the DDIM sampler on x... T After denoising, the output x0 strongly depends on the input x. T For a given x T The output x0 has a high degree of consistency, and the entire denoising process is a non-Markovian process.
[0091] Understandably, in a Markov process, the data from consecutive time steps are tightly bound together, thus requiring step-by-step denoising. In a non-Markov process, skip-step denoising can be implemented, meaning denoising is performed at intervals of several steps, thus accelerating data generation by reducing the number of steps. For example, if T is 1000, the DDPM sampler needs to perform denoising 1000 times to denoise based on the input x. T x0 is generated, and the DDIM sampler can be spaced 5 steps apart each time, so only 200 denoising processes are needed to achieve the same processing effect.
[0092] The training process of the diffusion model includes: when adding noise at step t, x t-1 By introducing random noise, we obtain the noisy data x. t The random noise at step t is recorded and used as training data for the corresponding denoising step; during denoising at step Tt, the data is denoised based on the conditional samples corresponding to the regular data sample x0 and the noisy data x. t Given the values of (Tt), predict the random noise introduced in step t, calculate the loss between the predicted noise and the actual random noise introduced in step t, adjust the parameters of the diffusion model based on the loss calculated in each denoising step, and perform denoising and calculate the loss again. Repeat the above steps multiple times until the diffusion model denoising converges.
[0093] Flow matching is a novel training paradigm for diffusion models. It first assumes that x... T There is a certain mapping relationship between x and x0, and from x T The transformation to x0 is continuous and differentiable, so from x... T The transformation process to x0 involves at least one link (i.e., a flow). This can be understood as the process from x... T The transformation to x0 is a dynamic motion process, therefore, at step t, there exists a velocity vector v. t , used to characterize the dynamic change at step t. That is, from x T The transformation to x0 can be viewed as the corresponding v at each step. t Superposition effect on x T This generates x0. Therefore, it is only necessary to determine v for each step. t This will complete the process from x T The transformation to x0. Specifically, this can be predicted using a neural network model, based on the data x at step t. t Predict the velocity vector v at step t t .
[0094] During the training phase of the diffusion model, x0 serves as the training sample, and the data x at step t can be easily determined.t Then, the corresponding velocity vector v is calculated. t The aforementioned neural network model is optimized using a loss function, continuously improving its predictive performance. By repeatedly training the model with different values of x0, the neural network model can be optimized towards a balanced approach across all speed angles.
[0095] In the diffusion model usage phase, the step size is set to Δt, so the time corresponding to each processing step is Δt, 2Δt, ..., 1. Let x... T go through In one step, x0 can be obtained.
[0096] It should be noted that from x T There are actually multiple links between x0 and x0. Although the conversion can eventually be completed along different links, there exists an optimal transport path (OT). Along this path, x... T It can move along a straight line at a constant speed to x0, and the path is simple and direct.
[0097] Compared with traditional diffusion model training methods, using Flow Matching to train diffusion models can find the optimal transmission path from multiple paths, so that Gaussian noise data can be transformed into regular data in the most direct and effective way during the model usage phase.
[0098] It should be noted that the limb motion generation model of this invention is a diffusion model. The input to the model is Gaussian noise data, the condition is first multimodal data, and the output is a sequence of motion of a 3D digital human. Therefore, when training the model, the training dataset consists of first multimodal data samples and corresponding motion sequence samples. t-1 and x t The samples are obtained by adding noise to the action sequence samples at steps t-1 and t, respectively. The noise is added at step t. t-1 By introducing random noise, we obtain the noisy data x. t The random noise at step t is recorded and used as training data for the corresponding denoising step; during denoising at step Tt, based on the first multimodal data sample and the noisy data x, t Given the values of (Tt), predict the random noise introduced in step t, calculate the loss between the predicted noise and the actual random noise introduced in step t, adjust the parameters of the diffusion model based on the loss calculated in each denoising step, and perform denoising and calculate the loss again. Repeat the above steps multiple times until the diffusion model denoising converges, thus completing the training of the model.
[0099] This invention generates a 3D digital human motion sequence using a diffusion model. Based on a trained diffusion model, it uses first multimodal data as a condition to perform multi-step denoising on randomly generated Gaussian noise data, thereby generating the motion sequence. Specifically, in the Tt-th denoising step, based on the input first multimodal data, the data obtained after the previous denoising step, and the value of (Tt), the noise to be removed in the Tt-th step is obtained. Removing this noise from the data obtained after the previous denoising step completes the Tt-th denoising step. After completing all denoising steps, the 3D digital human motion sequence is generated.
[0100] To improve the accuracy of the generated 3D digital human's motion sequence, the motion sequence of this invention can be divided into two parts: a body motion sequence and a hand motion sequence. In some embodiments, the first multimodal data can be input into the limb motion generation model to simultaneously generate the body motion sequence and the hand motion sequence. However, the body motion of a 3D digital human usually affects the hand motion. If the body motion sequence and the hand motion sequence are generated simultaneously, the body motion sequence and the hand motion sequence are relatively independent, which will result in poor quality of the generated hand motion sequence and a significant difference between the hand motion of the 3D digital human and that of a real human.
[0101] To improve the quality of hand motion sequences, in some embodiments, the first multimodal data of the present invention may include a lower body walking motion sequence, the limb motion generation model may include a body motion generation model and a hand motion generation model, and the motion sequence may include a body motion sequence and a hand motion sequence.
[0102] like Figure 5 As shown, step S2 may include:
[0103] Step S21: Input the first multimodal data into the body motion generation model to generate a body motion sequence;
[0104] Step S22: Replace the lower body walking action sequence in the first multimodal data with a body action sequence to obtain the second multimodal data;
[0105] Step S23: Input the second multimodal data into the hand motion generation model to generate a hand motion sequence;
[0106] The training set for the body motion generation model includes first multimodal data samples and corresponding body motion sequence samples, while the training set for the hand motion generation model includes second multimodal data samples and corresponding hand motion sequence samples.
[0107] When training the body motion generation model, random noise is introduced into the body motion sequence samples at step t to obtain the noisy data x. tThe random noise at step t is recorded and used as training data for the corresponding denoising step; during denoising at step Tt, the first multimodal data sample corresponding to the body action sequence sample and the noisy data x are used. t Given the values of (Tt), predict the random noise introduced in step t, calculate the loss between the predicted noise and the actual random noise introduced in step t, adjust the parameters of the body motion generation model based on the loss calculated in each denoising step, and perform denoising and calculate the loss again. Repeat the above steps multiple times until the denoising of the body motion generation model converges.
[0108] When training the hand motion generation model, random noise is introduced into the hand motion sequence samples at step t to obtain the noisy data x. t The random noise at step t is recorded and used as training data for the corresponding denoising step; during denoising at step Tt, the second multimodal data sample corresponding to the hand action sequence sample and the noisy data x are used. t Given the values of (Tt), predict the random noise introduced in step t, calculate the loss between the predicted noise and the actual random noise introduced in step t, adjust the parameters of the hand motion generation model based on the loss calculated in each denoising step, and perform denoising and calculate the loss again. Repeat the above steps multiple times until the denoising of the body motion generation model converges.
[0109] For example, the first multimodal data may include voice data, key motion data, style type data, and lower body walking motion sequences, while the second multimodal data may include voice data, key motion data, style type data, and body motion sequences.
[0110] The generated hand motion sequences are matched with the body motion sequences, improving the quality of the hand motion sequences. Based on these motion sequences, the 3D digital human can perform coordinated body movements, walking movements, and hand movements.
[0111] Of course, both the body motion generation model and the hand motion generation model can be diffusion models, in which case:
[0112] like Figure 6 As shown, step S21 may include:
[0113] Generate Gaussian noise data;
[0114] The first multimodal data is used as a condition for the body motion generation model. The Gaussian noise data is denoised to generate a 3D digital human body motion sequence.
[0115] like Figure 7 As shown, step S23 may include:
[0116] The second multimodal data was used as a condition for the hand motion generation model. The Gaussian noise data was denoised to generate a sequence of hand motions for a 3D digital human.
[0117] In some embodiments, the diffusion model of the present invention can be a UNet network architecture, which consists of several encoding blocks and decoding blocks, each of which contains a convolutional neural network (CNN) and a Transformer neural network. That is, both the body motion generation model and the hand motion generation model can be based on a UNet network architecture using convolutional neural networks and Transformer neural networks. Figure 8 As shown, the UNet network architecture is a U-shaped structure with left and right symmetry. It can include a 4-layer structure and 8 sub-networks. Each sub-network consists of a CNN module and a Transformer module.
[0118] In this invention, considering that speech data exists in audio format, it is impossible to directly input audio data into convolutional neural networks and Transformer neural networks; therefore, speech data needs to be converted. To this end, step S2 of this invention, inputting the first multimodal data into the trained limb movement generation model, may include:
[0119] Use a speech recognition model to extract speech features from speech data;
[0120] Voice features are input into the body movement generation model.
[0121] In some implementations, the speech recognition model can be a Hubert (Hidden-unit Bidirectional Encoder Representations from Transformers) model or a Wav2Vec model. The Hubert model's process for extracting speech features includes: first, the input speech data is converted into speech features by a feature extractor (such as MFCC or a convolutional neural network); then, the speech features are clustered using a clustering algorithm (such as k-means) to generate pseudo-labels, which represent hidden units (such as phonemes or sub-word units) in the speech; the input features are randomly masked, and the model is then trained to predict the pseudo-labels for the masked portions. Hubert uses a Transformer encoder to model the contextual information of the speech data and can capture long-range dependencies in the speech through a multi-layer self-attention mechanism. By predicting the pseudo-labels for the masked portions, Hubert can learn useful representations of speech without relying on manually labeled data.
[0122] In some embodiments, before step S2, the method for generating 3D digital human motion sequences of the present invention further includes: establishing an initial limb motion generation model, a training set and a test set for the limb motion generation model; training the initial limb motion generation model using the training set; and testing the trained limb motion generation model using the test set.
[0123] Prior to step S1, the method for generating 3D digital human motion sequences provided by the present invention further includes the following steps: receiving dialogue text; and using speech synthesis technology to convert the dialogue text into speech data.
[0124] As described above, the first multimodal data of the 3D digital human received by the 3D digital human motion sequence generation method provided by the present invention includes speech data. In some embodiments, the speech data can be obtained by users directly uploading audio files, or by audio acquisition and recording through audio acquisition devices, or by speech synthesis technology.
[0125] Common speech synthesis technologies include TTS (Text To Speech), which uses artificial neural networks to convert dialogue text (text) into speech data (natural speech stream), achieving real-time conversion of dialogue text.
[0126] This invention allows users to input dialogue text, and after automatically converting the dialogue text into speech data using speech synthesis technology, it is used as one type of data in the first multimodal data, and input together with other data in the first multimodal data into a trained limb movement generation model, thereby generating a 3D digital human movement sequence.
[0127] like Figure 9 As shown, the 3D digital human driving method provided by the present invention includes:
[0128] Step S3: Apply the aforementioned method for generating 3D digital human motion sequences to generate 3D digital human motion sequences;
[0129] Step S4: Redirect the 3D digital human motion sequence;
[0130] Step S5: Drive the 3D digital human using the redirected 3D digital human motion sequence.
[0131] This invention uses multimodal data, including voice data, to train a limb movement generation model. The resulting limb movement generation model has higher accuracy, generates better quality movement sequences, and has a smaller gap with real human movement sequences. By using movement sequences to drive 3D digital humans, the movements of 3D digital humans can be made closer to those of real humans.
[0132] like Figure 10As shown, the 3D digital human video generation method provided by the present invention includes:
[0133] Step S6: Apply the aforementioned 3D digital human driving method to drive the 3D digital human;
[0134] Step S7: Render the 3D digital human to generate a 3D digital human video.
[0135] This invention uses multimodal data, including voice data, to train a limb motion generation model. The resulting limb motion generation model has higher accuracy, generates better quality motion sequences, and has a smaller gap with real human motion sequences. Using motion sequences to drive 3D digital humans and generate 3D digital human videos can make 3D digital human videos more realistic.
[0136] like Figure 11 As shown, the 3D digital human motion sequence generation device of the present invention includes:
[0137] The receiving module is used to receive the first multimodal data of the 3D digital human, which includes voice data;
[0138] The generation module is used to input the first multimodal data into the trained limb motion generation model to generate a motion sequence of a 3D digital human.
[0139] The training set for the limb motion generation model includes first multimodal data samples and corresponding motion sequence samples.
[0140] It should be noted that the 3D digital human motion sequence generation device provided by the present invention can execute the 3D digital human motion sequence generation method described in any of the above embodiments during specific operation, and will not be elaborated on in this embodiment.
[0141] Figure 12 This is a schematic diagram of the structure of an electronic device provided by the present invention. The electronic device may include a processor, a communication interface, a memory, and a communication bus. The processor, communication interface, and memory communicate with each other via the communication bus. The processor can call logical instructions in the memory to execute methods for generating 3D digital human motion sequences, driving 3D digital humans, or generating 3D digital human videos.
[0142] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0143] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, and when the program instructions are executed by a computer, the computer is able to execute the 3D digital human motion sequence generation method, 3D digital human driving method or 3D digital human video generation method provided in the above embodiments.
[0144] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for generating 3D digital human motion sequences, the 3D digital human driving method, or the 3D digital human video generation method provided in the above embodiments.
[0145] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0146] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating a 3D digital human motion sequence, characterized in that, include: Receive first multimodal data of a 3D digital human, the first multimodal data including voice data and lower body walking motion sequence; The first multimodal data is input into the trained limb motion generation model to generate the motion sequence of the 3D digital human; The training set of the limb movement generation model includes a first multimodal data sample and a corresponding movement sequence sample. The limb movement generation model includes a body movement generation model and a hand movement generation model, and the movement sequence includes a body movement sequence and a hand movement sequence; The step of inputting the first multimodal data into the trained limb movement generation model to generate the movement sequence of the 3D digital human includes: The first multimodal data is input into the body motion generation model to generate the body motion sequence; the body motion sequence is used to replace the lower body walking motion sequence in the first multimodal data to obtain the second multimodal data; the second multimodal data is input into the hand motion generation model to generate the hand motion sequence. The training set of the body motion generation model includes the first multimodal data samples and the corresponding body motion sequence samples, and the training set of the hand motion generation model includes the second multimodal data samples and the corresponding hand motion sequence samples.
2. The method for generating 3D digital human motion sequences according to claim 1, characterized in that, The first multimodal data also includes at least one of key action data and style type data.
3. The method for generating a 3D digital human motion sequence according to claim 1, characterized in that, The limb movement generation model is a diffusion model; The step of inputting the first multimodal data into the trained limb movement generation model to generate the movement sequence of the 3D digital human includes: Generate Gaussian noise data; Using the first multimodal data as a condition of the diffusion model, the Gaussian noise data is denoised to generate the motion sequence of the 3D digital human.
4. The method for generating a 3D digital human motion sequence according to claim 3, characterized in that, The diffusion model is a UNet network architecture, which consists of several encoding blocks and decoding blocks. Each encoding block or decoding block contains a convolutional neural network and a Transformer neural network.
5. The method for generating a 3D digital human motion sequence according to claim 1, characterized in that, The step of inputting the first multimodal data into the trained limb movement generation model includes: Use a speech recognition model to extract the speech features from the speech data; The speech features are input into the body movement generation model.
6. The method for generating a 3D digital human motion sequence according to claim 1, characterized in that, Before receiving the first multimodal data of the 3D digital human, the method further includes: Receive dialogue text; The dialogue text is converted into speech data using speech synthesis technology.
7. A 3D digital human driving method, characterized in that, include: A 3D digital human motion sequence is generated using the method for generating a 3D digital human motion sequence as described in any one of claims 1-6; The 3D digital human motion sequence is redirected. The 3D digital human is driven using the redirected motion sequence of the 3D digital human.
8. A method for generating 3D digital human videos, characterized in that, include: The 3D digital human is driven using the 3D digital human driving method as described in claim 7; The 3D digital human is rendered to generate a 3D digital human video.
9. A device for generating 3D digital human motion sequences, characterized in that, include: The receiving module is used to receive the first multimodal data of the 3D digital human, the first multimodal data including voice data and a sequence of lower body walking movements; The generation module is used to input the first multimodal data into the trained limb movement generation model to generate the movement sequence of the 3D digital human; The training set of the limb movement generation model includes a first multimodal data sample and a corresponding movement sequence sample. The limb movement generation model includes a body movement generation model and a hand movement generation model, and the movement sequence includes a body movement sequence and a hand movement sequence; The step of inputting the first multimodal data into the trained limb movement generation model to generate the movement sequence of the 3D digital human includes: The first multimodal data is input into the body motion generation model to generate the body motion sequence; the body motion sequence is used to replace the lower body walking motion sequence in the first multimodal data to obtain the second multimodal data; the second multimodal data is input into the hand motion generation model to generate the hand motion sequence. The training set of the body motion generation model includes the first multimodal data samples and the corresponding body motion sequence samples, and the training set of the hand motion generation model includes the second multimodal data samples and the corresponding hand motion sequence samples.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for generating a 3D digital human motion sequence as described in any one of claims 1 to 6, the 3D digital human driving method as described in claim 7, or the 3D digital human video generation method as described in claim 8.
11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for generating a 3D digital human motion sequence as described in any one of claims 1 to 6, the 3D digital human driving method as described in claim 7, or the 3D digital human video generation method as described in claim 8.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it is a method for generating a 3D digital human motion sequence as described in any one of claims 1 to 6, a 3D digital human driving method as described in claim 7, or a 3D digital human video generation method as described in claim 8.
Citation Information
Patent Citations
Joint generation method of face and body motion parameters and related equipment
CN116977499A
Voice-driven 3D digital human motion generation method, system and device and medium
CN117831126A
Emotion-enhanced digital human driving and presenting system and method
CN119516063A