Training method and device of action diffusion model, equipment and storage medium
By adding a similarity training objective between the sample noisy sequence and the predicted noisy sequence to the action diffusion model, the action diffusion model can learn the noise intensity distribution, solving the problems of slow generation rate and smooth action, and achieving rapid generation of high-quality actions.
Patent Information
- Application Number
- CN202410545578.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-30
- Publication Date
- 2025-10-31
AI Technical Summary
Existing motion diffusion models adjust model parameters based solely on the difference between the motion sequence before noise addition and the predicted motion sequence output by the motion diffusion model during training. This makes it difficult to quickly generate high-quality motions, and existing methods suffer from slow generation rates and motion smoothness issues when generating complex motions.
By adding low-intensity second and third noise to the sample action sequence and the predicted action sequence respectively, and increasing the similarity between the sample noise sequence and the predicted noise sequence as a training target, the action diffusion model can learn the distribution of noise intensity between the sample noise sequence and the sample action sequence, thereby reducing noise with a large step when generating actions.
It enables the rapid generation of expressive actions, improves the action generation rate and quality, and solves the problems of slow generation rate and smooth action in existing technologies.
Smart Images

Figure CN120877014A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a training method, apparatus, device, and storage medium for an action diffusion model. Background Technology
[0002] The Motion Diffusion Model (MDM) is a diffusion model that generates a sequence of motions by denoising a noisy sequence, which can be used to control the actions of an object.
[0003] The application of the action diffusion model is based on the forward diffusion process and the backward denoising process. The forward diffusion process involves adding multiple steps of noise to the action sequence to obtain a noisy sequence, while the backward denoising process involves removing multiple steps of noise from the noisy sequence to reconstruct the action sequence. The purpose of training the action diffusion model is to enable it to fit the aforementioned backward denoising process.
[0004] However, during the training of the action diffusion model, the model parameters are usually adjusted only based on the difference between the action sequence before noise addition and the predicted action sequence output by the action diffusion model. The action diffusion model can only learn the overall process of reconstructing the action sequence before noise addition based on the noise sequence, and it is difficult to achieve a fast action generation rate. Summary of the Invention
[0005] This application provides a training method, apparatus, device, and storage medium for an action diffusion model. The technical solution provided by this application is as follows:
[0006] According to one aspect of the embodiments of this application, a training method for an action diffusion model is provided, the method comprising:
[0007] Obtain a sample action sequence and a sample condition representation, wherein the sample action sequence is used to characterize the action of the first object;
[0008] Add first noise to the sample action sequence to obtain a sample noise sequence;
[0009] The sample noise sequence and the sample conditional representation are input into the action diffusion model, and the action diffusion model outputs a predicted action sequence, which is used to characterize the action of the first object predicted by the action diffusion model under the guidance of the sample conditional representation.
[0010] A second noise is added to the sample action sequence to obtain a sample noisy sequence, and a third noise is added to the predicted action sequence to obtain a predicted noisy sequence, wherein the noise intensity of the second noise and the noise intensity of the third noise are both lower than the noise intensity of the first noise.
[0011] The action diffusion model is trained with the goal of increasing the similarity between the sample noisy sequence and the predicted noisy sequence.
[0012] According to one aspect of the embodiments of this application, a training apparatus for an action diffusion model is provided, the apparatus comprising:
[0013] The acquisition module is used to acquire sample action sequences and sample condition representations, wherein the sample action sequences are used to characterize the actions of the first object;
[0014] A noise-adding module is used to add first noise to the sample action sequence to obtain a sample noise sequence;
[0015] A generation module is used to input the sample noise sequence and sample conditional representation into the action diffusion model, and the action diffusion model outputs a predicted action sequence, which is used to characterize the action of the first object predicted by the action diffusion model under the guidance of the sample conditional representation.
[0016] The noise-adding module is further configured to add a second noise to the sample action sequence to obtain a sample noise-adding sequence, and to add a third noise to the predicted action sequence to obtain a predicted noise-adding sequence, wherein the noise intensity of the second noise and the noise intensity of the third noise are both lower than the noise intensity of the first noise.
[0017] The training module is used to train the action diffusion model with the training objective of increasing the similarity between the sample noisy sequence and the predicted noisy sequence.
[0018] According to one aspect of the present application, a computer device is provided, the computer device including a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the above-described training method for the action diffusion model.
[0019] According to one aspect of the present application, a computer-readable storage medium is provided, wherein a computer program is stored in the storage medium, the computer program being loaded and executed by a processor to implement the training method of the above-described action diffusion model.
[0020] According to one aspect of the present application, a computer program product is provided, the computer program product including a computer program stored in a computer-readable storage medium, wherein a processor reads from the computer-readable storage medium and executes the computer program to implement the above-described training method for the action diffusion model.
[0021] The technical solutions provided in this application have at least the following beneficial effects:
[0022] By inputting the sample noise sequence and sample conditional representation into the action diffusion model, the action diffusion model outputs a predicted action sequence. Then, the action diffusion model is trained with the goal of increasing the similarity between the sample noise sequence and the predicted noise sequence. In this method, since the sample noise sequence is obtained by adding a first noise to the sample action sequence, the sample noise sequence is obtained by adding a second noise, and the predicted noise sequence is obtained by adding a third noise, and the noise intensities of the second and third noises are both lower than the noise intensity of the first noise, the action diffusion model can learn the distribution of noise intensities in sample noise sequences between the sample noise sequence and the sample action sequence, i.e., the action diffusion model can learn the process of denoising the sample noise sequence into a sample noise sequence. Therefore, by employing the above technical solution, the action diffusion model can learn the process of denoising the sample noise sequence into a sample noise sequence under the guidance of the sample conditional representation, thus learning the intermediate process of reconstructing the sample action sequence from the sample noise sequence. Furthermore, when using the action diffusion model to generate actions, the action diffusion model can denoise the input noise sequence with the denoising step size of the aforementioned intermediate process, thereby achieving a faster action generation rate. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of a computer system provided in one embodiment of this application;
[0024] Figure 2 This is a flowchart of a training method for an action diffusion model provided in one embodiment of this application;
[0025] Figure 3 This is a schematic diagram illustrating the use of action sequences to drive a virtual character, as provided in one embodiment of this application.
[0026] Figure 4 This is a schematic diagram of the training objective of an action diffusion model provided in one embodiment of this application;
[0027] Figure 5 This is a flowchart of a training method for an action diffusion model provided in another embodiment of this application;
[0028] Figure 6 This is a schematic diagram of the training process of an action diffusion model provided in one embodiment of this application;
[0029] Figure 7 This is a flowchart of an action generation method based on an action diffusion model provided in one embodiment of this application;
[0030] Figure 8This is a flowchart of a training method for an action diffusion model in a video production scenario provided in one embodiment of this application;
[0031] Figure 9 This is a flowchart of an action generation method based on an action diffusion model in a video production scenario provided by one embodiment of this application;
[0032] Figure 10 This is a flowchart of a training method for an action diffusion model in a game character control scenario provided in one embodiment of this application;
[0033] Figure 11 This is a flowchart of an action generation method based on an action diffusion model in a game character control scenario provided in one embodiment of this application;
[0034] Figure 12 This is a schematic diagram illustrating the result of using action sequences generated by different models to drive a virtual character, provided in one embodiment of this application.
[0035] Figure 13 This is a block diagram of a training device for an action diffusion model provided in one embodiment of this application;
[0036] Figure 14 This is a structural block diagram of a computer device provided in one embodiment of this application. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0038] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0039] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0040] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning. Pre-trained models are the latest development in deep learning, integrating all of these techniques.
[0041] Speech technology encompasses key technologies such as Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and voiceprint recognition. Aiming to enable computers to hear, see, speak, and feel, it represents the future direction of human-computer interaction, with speech emerging as one of the most promising methods. Large-scale modeling has revolutionized speech technology. Pre-trained models using the Transformer architecture, such as WavLM (Waveform Language Model) and UniSpeech, possess strong generalization and versatility, enabling them to excel in various speech processing tasks.
[0042] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP deals with natural language, the language people use in daily life, and is closely related to linguistics research; furthermore, it involves computer science and mathematics. Pre-trained models, a crucial technique for model training in artificial intelligence, evolved from large language models in NLP. After fine-tuning, large language models can be widely applied to downstream tasks. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.
[0043] The technical solutions provided in this application involve artificial intelligence technologies such as machine learning, speech technology, and natural language processing, which are specifically described and illustrated through the following embodiments.
[0044] Before introducing the technical solution of this application, some terms involved in this application will be explained. The following related explanations are optional and can be combined with the technical solutions of the embodiments of this application in any way, all of which fall within the protection scope of this application. The embodiments of this application include at least some of the following contents.
[0045] Generative Adversarial Network (GAN): A deep learning framework consisting of a generator and a discriminator, used to generate realistic data.
[0046] Denoising Diffusion Probabilistic Models (DDPM): A probabilistic model used for modeling and processing data, generating data by removing noise from the input data.
[0047] Transformer: A deep learning architecture based on attention mechanisms, widely used in natural language processing tasks such as machine translation and text generation.
[0048] Denoising Diffusion Implicit Models (DDIM): An implicit model for modeling and processing data, which accelerates data generation by speeding up the backward denoising process.
[0049] Fréchet Gesture Distance: A metric used to measure the similarity between gestures, often used to evaluate the performance of motion diffusion models.
[0050] Non-Playable Character (NPC): In video games, a character is controlled by the computer instead of a player.
[0051] Multi-Layer Perceptron (MLP): A neural network architecture consisting of multiple fully connected layers.
[0052] Adaptive Moment Estimation (Adam): An optimization algorithm for updating neural network parameters. Adam designs independent adaptive learning rates for different parameters.
[0053] AdamW: An improved optimization algorithm for training neural networks that combines Adam with weight decay.
[0054] Exponential Moving Average (EMA): An algorithm that weights sequential data over time. In deep learning, EMA is used to smooth the updates of model parameters.
[0055] In the field of motion generation, especially gesture generation, there is a challenge in accurately and quickly generating complex movements. Taking human motion as an example, the motion generation method using Variational Autoencoders (VAEs) often results in stiff arm postures and unsmooth movements, making the generated actions less realistic. GAN-based motion generation methods show promise in this field, but when generating complex and refined movements, the design of the GAN training process still presents many challenges.
[0056] Action generation methods based on action diffusion models can capture object actions comprehensively and generate high-quality actions. The application of action diffusion models is based on a forward diffusion process and a backward denoising process. The forward diffusion process adds multiple noise steps to the action sequence to obtain a noisy sequence, while the backward denoising process removes multiple noise steps from the noisy sequence to reconstruct the action sequence. The forward diffusion process includes multiple noise-adding steps, with each step of adding noise completing one noise-adding step. The backward denoising process includes multiple denoising steps corresponding to the multiple noise-adding steps, with each step of removing noise completing one denoising step. The purpose of training the diffusion model is to enable it to fit the aforementioned backward denoising process. Taking DDPM as an example, during the training process of DDPM, the model parameters are adjusted only based on the difference between the action sequence before noise addition and the action sequence output by the diffusion model. In this process, DDPM can only learn to predict the conditional distribution of the action sequence before noise addition based on the distribution of the input noisy action sequence, but cannot learn the edge distribution of the denoised sequence corresponding to each denoising step in the process of reconstructing the action sequence. Therefore, when using DDPM to generate actions, to ensure the quality and accuracy of the generated actions, a large number of denoising steps needs to be set for DDPM, so that each denoising step fitted by DDPM denoises with small steps. For example, the number of denoising steps for DDPM is set to 1000. It can be seen that due to the limitation of small-step denoising, the action generation rate of DDPM is relatively slow. To address the above problem, DDIM accelerates action generation by skipping some denoising steps in the backward denoising process; however, this method results in generated actions that are too smooth and lack the subtle complexity of natural human movements.
[0057] The technical solutions provided in this application can quickly generate expressive actions, which are described in detail through the following embodiments.
[0058] Please refer to Figure 1 The diagram illustrates a computer system provided in one embodiment of this application. The computer system includes a model training device 10 and a model usage device 20.
[0059] The model training device 10 can be an electronic device such as a mobile phone, desktop computer, tablet computer, laptop computer, vehicle terminal, server, intelligent robot, smart TV, multimedia playback device, or other electronic devices with strong computing power; this application does not limit this. The model training device 10 is used to train the action diffusion model 30.
[0060] In this embodiment, the action diffusion model 30 is a machine learning model. Optionally, the model training device 10 trains the action diffusion model 30 using machine learning to achieve better performance.
[0061] Optionally, the training process of the action diffusion model 30 is as follows: A sample action sequence and a sample conditional representation are obtained. The sample action sequence is used to represent the action of the first object, and the sample conditional representation is used to guide the action diffusion model to generate the action of the first object. First noise is added to the sample action sequence to obtain a sample noise sequence. The sample noise sequence and the sample conditional representation are input into the action diffusion model 30, and the action diffusion model 30 outputs a predicted action sequence. The predicted action sequence is used to represent the action of the first object predicted by the action diffusion model under the guidance of the sample conditional representation. Second noise is added to the sample action sequence to obtain a sample noisy sequence, and third noise is added to the predicted action sequence to obtain a predicted noisy sequence. The noise intensity of the second noise and the third noise are both lower than the noise intensity of the first noise. The action diffusion model 30 is trained with the training objective of increasing the similarity between the sample noisy sequence and the predicted noisy sequence.
[0062] The model-using device 20 can be an electronic device such as a mobile phone, desktop computer, tablet computer, laptop computer, in-vehicle terminal, server, intelligent robot, smart TV, multimedia playback device, or other electronic devices with strong computing power; this application does not limit this. Optionally, the model-using device 20 performs the action generation task through a trained action diffusion model. For example, the model-using device 20 uses the trained action diffusion model to generate an action sequence for controlling the virtual character based on the audio features of the virtual character.
[0063] The model training device 10 and the model usage device 20 can be two separate devices or the same device.
[0064] The method provided in this application uses a computer device as the execution subject for each step. This computer device refers to an electronic device capable of data computation, processing, and storage. When the computer device is a server, the server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The computer device can be... Figure 1 The model training device 10 can also be the model usage device 20.
[0065] Please refer to Figure 2 The diagram illustrates a flowchart of a training method for an action diffusion model provided in one embodiment of this application. The execution entity for each step of this method is a computer device, for example, the computer device is... Figure 1 The model training device 10 in the computer system shown. The method includes at least one of the following steps 210 to 250.
[0066] Step 210: Obtain the sample action sequence and sample condition representation.
[0067] The sample action sequence is used to characterize the action of the first object, and the sample condition representation is used to guide the action diffusion model to generate the aforementioned action of the first object.
[0068] In this application embodiment, the first object can be a person, animal, or any other object capable of movement; this application does not limit this. Optionally, the first object can be a real object, such as a real person or animal. Optionally, the first object can be a virtual object, which refers to a computer-generated virtual object, such as a virtual person or animal.
[0069] In some embodiments, the sample action sequence includes n action frames, where n is a positive integer.
[0070] An action frame represents an action, and each action frame in the sample action sequence represents an action of the first object.
[0071] For example, please refer to Figure 3 The illustration shows a schematic diagram of a virtual character driven by an action sequence according to an embodiment of the present application, in which each action of the virtual character 41 corresponds to an action frame.
[0072] In some embodiments, an action frame includes at least one of the following data: root node data and body node data. The root node data and body node data are divided into root node and body node, where the root node is used to control the object's position in space, and the body node is used to control the object's body movements.
[0073] In some embodiments, the root node data includes overall position data and overall rotation data. For a given action frame, the overall position data refers to the object's position coordinates in space, and the overall rotation data represents the rotation angle of the object from the previous action frame to this action frame.
[0074] In some embodiments, body node data is used to describe at least one of the following data for at least one joint of an object: rotation angle, position, linear velocity, angular velocity, and contact state with the ground. Each joint represents a part of the object. For a given action frame, the rotation angle represents the change in rotation angle of the joint from the previous action frame to the current action frame; the position refers to the spatial coordinates of the joint; the linear velocity describes the linear velocity of the joint as it changes from the previous action frame to the current action frame; the angular velocity describes the rotational angular velocity of the joint as it changes from the previous action frame to the current action frame; and the contact state with the ground refers to the contact state between the object's foot joint and the ground, which can be either in contact or not in contact.
[0075] In some embodiments, the actions generated by various parts of an object are referred to as body movements. Optionally, body movements are represented as a combination of data from multiple joints.
[0076] It should be noted that the root node data mentioned above is used to represent the overall movement of an object in space, such as moving from one position to another along a certain direction, and does not correspond to a specific part of the object. For example, when the root node data is not used and only the body node data is used to drive the movement of the virtual character, the virtual character can exhibit body movements, but there will be no change in position or orientation in the virtual space.
[0077] Additionally, it should be noted that the "object" in the above embodiments is any object capable of generating actions. For example, the "object" is the first object, or the "object" is the second object in the embodiments below.
[0078] In some embodiments, the sample action sequence is represented as a vector, that is, the data in the n action frames included in the sample action sequence is represented as a long vector.
[0079] In some embodiments, the sample condition representation includes at least one of the following: sample audio features, sample style encoding, sample seed sequence, and sample time steps.
[0080] Sample audio features are used to characterize the features of the audio corresponding to the sample action sequence. Optionally, the audio corresponding to the sample action sequence is manually annotated and adapted to the action represented by the sample action sequence. For example, if the sample action sequence is used to drive the action of a character in a film, then the audio corresponding to the sample action sequence is the audio emitted by the character while performing the action represented by the sample action sequence. Optionally, the audio corresponding to the sample action sequence is audio describing the action represented by the sample action sequence, such as the text content of the audio being "a person is dancing". Using sample audio features as a sample conditional representation enables the action diffusion model to learn the ability to generate actions based on audio, thereby making the actions generated by the action diffusion model and the audio mutually coordinated.
[0081] Sample style encoding is used to characterize the emotional style of a sample action sequence. For example, the emotional style of a sample action sequence may be "sad," "happy," "joyful," or "serious." In some embodiments, the sample action sequence is labeled with textual style tags such as "sad," "happy," "joyful," and "serious," and the sample style encoding is obtained by textually encoding these style tags. Using sample style encoding as a sample conditional representation enables the action diffusion model to learn the emotional style of the sample action sequence, thereby enriching the emotional content that the actions generated by the action diffusion model can express.
[0082] The sample seed sequence includes action sequences that precede and are sequential to the sample action sequences. For example, the sample seed sequence includes 8 action frames, the sample action sequence includes 1800 action frames, and the first action frame of the sample action sequence can be sequentially connected to the eighth action frame of the sample seed sequence. Using the sample seed sequence as a sample conditional representation enables the action diffusion model to learn the ability to generate action sequences by sequentially connecting existing action sequences, thus enhancing the smoothness of the actions generated by the action diffusion model.
[0083] The sample time step count indicates the number of steps in which noise is added to the sample action sequence during the process of obtaining the sample noise sequence. For a detailed explanation of the sample time step count, please refer to the examples below; it will not be repeated here.
[0084] In some embodiments, the sample condition representation includes the above-mentioned sample audio features, sample seed sequence, and sample time step.
[0085] By introducing the above sample condition representation, the action diffusion model can generate more natural, smooth and expressive actions.
[0086] It should be noted that this application does not limit the specific content of the sample condition representation. Any data that can play a guiding role in the process of generating actions by the action diffusion model can be used as the sample condition representation in the embodiments of this application. In some embodiments, the sample condition representation further includes at least one of the following: sample semantic encoding, sample text features, sample facial expression data, and sample pose data.
[0087] Sample semantic encoding is used to describe the actions represented by the sample action data. For example, if the actions represented by the sample action data include the "OK" gesture, then the semantic encoding can be obtained from the text "OK".
[0088] Sample text features are used to characterize the text content of the audio corresponding to the sample action sequence. For example, sample text features are extracted from the dialogue of a virtual character in a film.
[0089] Sample facial expression data is used to represent the facial expression of the first object. For example, if sample motion data is used to represent the motion of a character in a film, then facial expression data is used to represent the facial expression that character has while performing the motion represented by the sample motion sequence.
[0090] Sample pose data is used to characterize the overall pose of the first object, such as standing, sitting, lying down, etc.
[0091] Step 220: Add first noise to the sample action sequence to obtain the sample noise sequence.
[0092] In this embodiment, adding first noise to the sample action sequence means fusing the first noise into the n action frames included in the sample action sequence. Specifically, after representing the sample action sequence as a vector, adding first noise to the sample action sequence means fusing the first noise into each element of the vector form of the sample action sequence, thereby obtaining a new vector, which is the sample noise sequence.
[0093] In some embodiments, the first noise includes t-step first sub-noise.
[0094] In some embodiments, t-step first sub-noise is added to the sample action sequence to obtain a sample noise sequence, where t is an integer greater than 1.
[0095] Adding t-step first sub-noise to a sample action sequence can also be called adding the first sub-noise to the sample action sequence in t-step steps. Specifically, in the first step of the noise addition step, the first sub-noise of the first step is added to the sample action sequence x0 to obtain x1. In the second step of the noise addition step, the first sub-noise of the second step is added to the sequence x1 obtained in the first step of the noise addition step to obtain x2, and so on. In the t-th step of the noise addition step, the first sub-noise of the second step is added to the sequence x0 obtained in the (t-1)-th step of the noise addition step. t-1 Adding the first sub-noise at step t yields the sample noise sequence x. t .
[0096] In some embodiments, the first sub-noise includes first Gaussian noise. t steps of first Gaussian noise are added to the sample action sequence to obtain a sample noise sequence. Specifically, for the sample action sequence x0, each t-step addition of first Gaussian noise involves adding first Gaussian noise to the sequence obtained in the previous denoising step according to the following formula 1:
[0097]
[0098] Where q(x) v |x v-1 () represents the sequence x obtained from the known (v-1)th noisy addition step. v-1 The probability distribution q(x) v-1 Predict the sequence x obtained from the v-th noisy step. v The conditional distribution of . v is a positive integer less than or equal to t. It follows a Gaussian distribution, where β v is the variance of the first Gaussian noise at step v, which is between 0 and 1, and I is a vector in which all elements are 1.
[0099] The sequence x obtained in step v (the noise addition step) v Gaussian distribution obtained from prediction Obtained by sampling from x. v-1Adding the first Gaussian noise in step v yields x v The process can also be intuitively represented by the following formula 2:
[0100]
[0101] Wherein, ∈ is sampled from the standard Gaussian distribution N(0,1).
[0102] After the above t noise-adding steps, the sample noise sequence can be obtained. It should be noted that in some embodiments, t is a positive integer less than or equal to T, where T is the number of steps to remove noise from the input noise sequence during the action generation task after the action diffusion model has been trained. T is set by the technician as needed. In this embodiment, since the action diffusion model can learn the ability to reduce noise with large strides, T can be set relatively small. For example, T = 10 steps, and t is any integer between 1 and 10.
[0103] It should be noted that the process of adding noise to the sample action sequence only changes the numerical value of the data in the sample action sequence. Therefore, when the sample action sequence includes the root node data and body node data mentioned above, the data in the sample noise sequence can also correspond to the root node and body node respectively.
[0104] In the above embodiment, the number of steps t in which the first sub-noise is added to the sample action sequence is the sample time step. The larger t is, the higher the noise intensity in the sample noise sequence. Using the sample time step t as a sample condition representation can instruct the action diffusion model to remove the noise in the sample noise sequence, thereby enabling the action diffusion model to learn the ability to remove noise of different intensities to generate actions.
[0105] Step 230: Input the sample noise sequence and sample conditional representation into the action diffusion model, and output the predicted action sequence from the action diffusion model.
[0106] The predicted action sequence is used to characterize the action of a first object predicted by the action diffusion model under the guidance of the sample conditional representation. In some embodiments, the predicted action sequence includes n action frames predicted by the action diffusion model.
[0107] In some embodiments, the sample noise sequence, sample condition representation, and predicted action sequence are all represented in vector form.
[0108] In some embodiments, the action diffusion model is built upon the Transformer encoder. Specifically, the action diffusion model has an m-layer encoder architecture, where m is a positive integer, and each layer includes multiple attention heads and a feedforward fully connected network. In some embodiments, the input of each layer in the action diffusion model (such as the input of the first layer, which is the aforementioned sample noise sequence and sample conditional representation) is assigned to multiple attention heads. Based on the assigned sequence data, each attention head in that layer obtains a Q(query), K(key), and V(value) matrix. Then, based on Q and K, it obtains an attention weight matrix to characterize the correlation between Q and K. Finally, based on the attention weight matrix and V, it obtains the output of that attention head. The outputs of each attention head in that layer are concatenated and input into the feedforward fully connected network of that layer. This feedforward fully connected network performs non-linear processing on the input information to obtain the output of that layer. The output of the last layer in the m-layer encoder architecture is the predicted action sequence output by the action diffusion model.
[0109] Step 240: Add second noise to the sample action sequence to obtain a sample noisy sequence, and add third noise to the predicted action sequence to obtain a predicted noisy sequence.
[0110] The noise intensity of the second and third noises is lower than that of the first noise. Therefore, the noise intensity in the sample noisy sequence is lower than that in the sample noisy sequence.
[0111] Noise intensity is used to reflect the degree to which noise disturbs the data. In the embodiments of this application, the noise intensity of the second noise and the noise intensity of the third noise are both lower than the noise intensity of the first noise. That is, the degree to which the second noise disturbs the sample action sequence and the degree to which the third noise disturbs the predicted action sequence are lower than the degree to which the first noise disturbs the sample action sequence.
[0112] In some embodiments, the second noise includes k steps of second sub-noise, and the third noise includes k steps of third sub-noise, where k is a positive integer less than t.
[0113] In some embodiments, k steps of second sub-noise are added to the sample action sequence to obtain a noisy sample sequence, and k steps of third sub-noise are added to the predicted action sequence to obtain a noisy predicted sequence. Wherein, the variance of the j-th step of the second sub-noise in the k-step second sub-noise is the same as the variance of the j-th step of the first sub-noise in the t-step first sub-noise, and the variance of the j-th step of the third sub-noise in the k-step third sub-noise is the same as the variance of the j-th step of the first sub-noise in the t-step first sub-noise, where j is a positive integer less than or equal to k.
[0114] It should be noted that noise variance can be used to reflect noise intensity. Therefore, the first sub-noise, second sub-noise, and third sub-noise in the same step (e.g., step j) can be considered to have the same noise intensity. Furthermore, when k is less than t, the noise intensity of the second noise including the second sub-noise of step k and the noise intensity of the third noise including the third sub-noise of step k are both lower than the noise intensity of the first noise including the first sub-noise of step t. Accordingly, the noise intensity in the sample noisy sequence obtained by the above method is lower than the noise intensity in the sample noise sequence. Therefore, in some embodiments, the second noise and the third noise can also be considered to have the same intensity.
[0115] In some embodiments, k is t-1, the second noise includes the second sub-noise at step t-1, and the third noise includes the third sub-noise at step t-1.
[0116] In some embodiments, a second sub-noise step t-1 is added to the sample action sequence to obtain a sample noisy sequence, wherein the variance of the i-th step second sub-noise in the t-1 step second sub-noise is the same as the variance of the i-th step first sub-noise in the t-th step first sub-noise, and i is a positive integer less than or equal to t-1.
[0117] In some embodiments, the first sub-noise and the second sub-noise are the same. The first sub-noise of step t-1 is added to the sample action sequence to obtain the sample noisy sequence. The first sub-noise of step t is added to the sample noisy sequence to obtain the sample noise sequence.
[0118] In some embodiments, the second sub-noise includes a second Gaussian noise, and t-1 steps of second Gaussian noise are added to the sample action sequence to obtain a sample noisy sequence.
[0119] In some embodiments, a third sub-noise of step t-1 is added to the predicted action sequence to obtain a predicted noisy sequence, wherein the variance of the third sub-noise of step t-1 is the same as the variance of the first sub-noise of step t-1.
[0120] In some embodiments, the third sub-noise includes third Gaussian noise, and t-1 steps of third Gaussian noise are added to the predicted action sequence to obtain a predicted noisy sequence.
[0121] The first, second, and third Gaussian noises mentioned above are all Gaussian noise, which refers to noise whose probability distribution follows a Gaussian distribution. As can be seen from the above embodiments, adding the first Gaussian noise to a sample action sequence includes sampling from a Gaussian distribution. Therefore, adding the second and third Gaussian noises also includes sampling from a Gaussian distribution. Optionally, for the i-th step first Gaussian noise, i-th step second Gaussian noise, and i-th step third Gaussian noise with the same variance, they correspond to three different sampling processes (the Gaussian distributions sampled by these three different sampling processes have the same variance). Optionally, the process of adding the i-th step first Gaussian noise and the i-th step second Gaussian noise corresponds to the same sampling process, that is, the i-th step first Gaussian noise and the i-th step second Gaussian noise are the same.
[0122] The specific steps for adding t-1 steps of second Gaussian noise to the sample action sequence and adding t-1 steps of third Gaussian noise to the predicted action sequence can be referred to in the above embodiment, specifically the first t-1 noise addition steps for adding t steps of first Gaussian noise to the sample action sequence. This application will not elaborate on these steps further.
[0123] In the above embodiment, the sample noise-added sequence and the predicted noise-added sequence are obtained by t-1 noise-added steps. With multiple rounds of training on the action diffusion model, the above method enables the model to learn the intermediate denoising steps that restore the input sample noise sequence to the sample action sequence, thereby giving the model a large-step denoising capability and improving its action generation rate. For example, if T = 10 steps, meaning the number of denoising steps in the action diffusion model is set to 10, then during training, when t is 10, the model can learn the sequence distribution after the 9th noise-added step, i.e., the sequence distribution after the 1st denoising step. When t is 9, the model can learn the sequence distribution after the 8th noise-added step, i.e., the sequence distribution after the 2nd denoising step, and so on. The model can learn the sequence distribution after all 10 denoising steps. Therefore, when the number of denoising steps is set to 10, the action diffusion model can ensure the quality of the generated actions while generating action sequences at high speed.
[0124] Step 250: Train the action diffusion model with the training objective of increasing the similarity between the sample noisy sequence and the predicted noisy sequence.
[0125] In some embodiments, both the sample noisy sequence and the predicted noisy sequence are represented in vector form.
[0126] In some embodiments, the action diffusion model is trained with the objective of minimizing the vector distance between the sample noisy sequence and the predicted noisy sequence. The distance can be, for example, Euclidean distance, Manhattan distance, etc., and this application does not limit it to any particular type.
[0127] The technical solution provided in this application involves inputting a sample noise sequence and a sample conditional representation into an action diffusion model, which then outputs a predicted action sequence. The model is trained with the goal of increasing the similarity between the sample noise sequence and the predicted noise sequence. In this method, since the sample noise sequence is obtained by adding a first noise to the sample action sequence, the sample noise sequence is obtained by adding a second noise, and the predicted noise sequence is obtained by adding a third noise, and the noise intensities of the second and third noises are both lower than the noise intensity of the first noise, the action diffusion model can learn the distribution of noise intensities in sample noise sequences between the sample noise sequence and the sample action sequence, i.e., it can learn the process of denoising a sample noise sequence into a sample noise sequence. Therefore, by employing this technical solution, the action diffusion model can learn the process of denoising a sample noise sequence into a sample noise sequence under the guidance of the sample conditional representation, thereby learning the intermediate process of reconstructing a sample action sequence from a sample noise sequence. Furthermore, when using the action diffusion model to generate actions, the action diffusion model can denoise the input noise sequence with the denoising step size of the aforementioned intermediate process, thereby achieving a faster action generation rate.
[0128] The following describes, in conjunction with the embodiments described above, a noise-adding method that adds t-step first sub-noise to a sample action sequence to obtain a sample noise sequence, adds t-1-step second sub-noise to a sample action sequence to obtain a sample noise-added sequence, and adds t-1-step third sub-noise to a predicted action sequence to obtain a predicted noise sequence, wherein the variance of the i-th step second sub-noise in the t-1-step second sub-noise is the same as the variance of the i-th step first sub-noise in the t-step first sub-noise, and the variance of the i-th step third sub-noise in the t-1-step third sub-noise is the same as the variance of the i-th step first sub-noise in the t-step first sub-noise, is used to illustrate the principle of accelerating action sequence generation in this application.
[0129] For example, please refer to Figure 4If the action diffusion model sets the number of steps T for removing noise from the noise sequence when performing the action generation task to 5, then during the training of the action diffusion model, the multiple training rounds will traverse all cases where the sample time step t takes the values t=1, t=2, t=3, t=4, and t=5. Specifically, adding the first sub-noise of step 1 to the sample action sequence x0 yields x1, adding the first sub-noise of step 2 to x1 yields x3, and so on, resulting in five noisy sequences x1, x2, x3, x4, and x5 (x5 being x...). T These five sequences correspond to five different noise intensities and have different distributions; that is, x1, x2, x3, x4, and x5 correspond to five different prior distributions (i.e., distributions predetermined before the model infers the predicted action sequence). Taking the training process at t=4 as an example, in this round of training, four steps of the first sub-noise are added to the sample action sequence to obtain the sample noise sequence x4. In other words, the action diffusion model needs to fit four denoising steps from x4 to x3, from x3 to x2, from x2 to x1, and from x1 to x0 to reconstruct the sample action sequence. At this time, adding three steps of the second sub-noise to the sample action sequence to obtain the sample noise-added sequence is a prior sampling process, that is, sampling the sample noise-added sequence A from the prior distribution corresponding to x3. The predicted action sequence... Adding three steps of third sub-noise to obtain the predicted noisy sequence B is a posterior sampling process. That is, it samples the predicted noisy sequence from a posterior distribution. The posterior distribution refers to the distribution of the re-predicted action sequence based on model inference (in this example, the distribution of the re-predicted x3). Therefore, the predicted noise sequence can be used to reflect the denoising result of the first denoising step fitted by the action diffusion model in this round of training. The goal is to increase the similarity between the predicted noisy sequence A and the sample noisy sequence B. The action diffusion model is trained to improve the denoising result of the first denoising step, which is then fitted by the action diffusion model. The distribution of x is close to the prior distribution of x3, thus learning the ability to infer x3 from x4 during the process of reconstructing x0 from x4. Similarly, after multiple rounds of training, the action diffusion model can learn each intermediate denoising step to infer x0 from x5. Furthermore, because the action diffusion model learns from x5 during training... T For each intermediate denoising step up to x0, even in order to accelerate action generation, T is set to a small value during the action generation task. This allows the action diffusion model to significantly reduce noise while ensuring the quality of the actions generated by the action diffusion model.
[0130] It should be noted that the noise addition method in the above example is only a preferred case provided by the embodiments of this application. When the noise intensity of the second noise and the noise intensity of the third noise are both lower than the noise intensity of the first noise, the action diffusion model is trained with the aim of minimizing the similarity between the predicted noise-added sequence and the sample noise-added sequence. This makes the posterior distribution obtained from the predicted action sequence output by the action diffusion model approach the prior distribution, which can improve the noise reduction stride of the action diffusion model to a certain extent, thereby accelerating action generation.
[0131] In summary, the technical solution provided in this application enables the action generation model to learn the ability to reduce noise over large strides, thereby accelerating action generation.
[0132] In some embodiments, by introducing a discriminator to measure the similarity between the sample noisy sequence and the predicted noisy sequence, step 250 includes at least one of the following sub-steps 252 to 258.
[0133] Sub-step 252: Input the sample noisy sequence, the predicted noisy sequence, and the sample conditional representation into the discriminator, and the discriminator outputs the discrimination value.
[0134] The discriminant value is used to reflect the similarity between the discriminator's predictions of the sample noisy sequence and the predicted noisy sequence, guided by the sample conditional representation.
[0135] In some embodiments, the discriminator includes an MLP (such as a 7-layer MLP). The discriminator categorizes the similarity between the sample noisy sequence and the predicted noisy sequence into two classes, 0 and 1, based on the sample condition representation. 0 indicates dissimilarity, and 1 indicates that the two are indistinguishable or similar. The discriminator can output its predicted probability distribution for each of the 0 and 1 classes. This probability distribution, or any probability from the probability distribution, can be used as the discriminant value.
[0136] Sub-step 254: Calculate the reconstruction loss based on the sample action sequence and the predicted action sequence. The reconstruction loss is used to reflect the difference between the sample action sequence and the predicted action sequence.
[0137] It should be noted that, in the embodiments of this application, one of the purposes of training the action diffusion model is to enable the action diffusion model to learn the ability to reconstruct sample action sequences based on sample noise sequences derived from sample action sequences with added noise, and the predicted action sequence is the result of the action diffusion model reconstructing the sample action sequence. Therefore, the loss used to reflect the difference between the sample action sequence and the predicted action sequence is called the reconstruction loss.
[0138] In some embodiments, the reconstruction loss is calculated using the following formula 3:
[0139]
[0140] Where E represents the search The expectation is given by c, where c represents the above sample conditions, x0~q(x0|c) represents the sample action sequence x0 sampled from the distribution q(x0|c), and t~[1,T] represents the sample time step t sampled from the interval [1,T]. This indicates a predicted sequence of actions. Please refer to Formula 4 below for the calculation:
[0141]
[0142] Where δ is the hyperparameter configured by the technicians.
[0143] Sub-step 256: Calculate the adversarial loss based on the discriminant value. The adversarial loss is negatively correlated with the similarity predicted by the discriminator for the noisy sequence and the predicted noisy sequence under the guidance of the sample conditional representation.
[0144] In some embodiments, the discriminant value is positively correlated with the aforementioned similarity. For example, the discriminant value is the probability that the sample noisy sequence output by the discriminator is similar to the predicted noisy sequence. The discriminant value is inverted or negative to obtain the adversarial loss. Alternatively, the adversarial loss is obtained based on the result of inverting or negatively taking the discriminant value, thereby making the adversarial loss negatively correlated with the aforementioned similarity.
[0145] In some embodiments, the discriminant value is negatively correlated with the aforementioned similarity. For example, the discriminant value is the probability that the sample noisy sequence output by the discriminator and the predicted noisy sequence are not similar. The discriminant value is directly determined as the adversarial loss, or the adversarial loss is obtained based on the discriminant value, thereby making the adversarial loss negatively correlated with the aforementioned similarity.
[0146] In some embodiments, the discriminant value is the probability distribution predicted by the discriminator regarding whether the sample noisy sequence and the predicted noisy sequence are similar. Based on this probability distribution and the pseudo-distribution, the adversarial loss is obtained. The adversarial loss is used to reflect the difference between the probability distribution and the pseudo-distribution. In the pseudo-distribution, the probability that the sample noisy sequence and the predicted noisy sequence are similar is 1, and the probability that they are not similar is 0.
[0147] In other words, the action diffusion model adjusts its parameters to maximize the similarity between the discriminator's prediction of the noisy sequence and the predicted noisy sequence.
[0148] Sub-step 258: Train the action diffusion model based on reconstruction loss and adversarial loss.
[0149] In some embodiments, the reconstruction loss and adversarial loss described above are weighted and summed to obtain the objective function value, which reflects the performance of the action diffusion model in performing the action generation task. Based on the objective function value, the parameters of the action diffusion model are adjusted to obtain the trained action diffusion model.
[0150] In some embodiments, the parameters of the action diffusion model are adjusted to minimize the objective function value, resulting in a trained action diffusion model.
[0151] In some embodiments, the objective function value can also be obtained through the following steps:
[0152] 1. Calculate the physical constraint value based on the predicted action sequence. The physical constraint value is used to reflect the physical naturalness of the action of the first object represented by the predicted action sequence.
[0153] The physical naturalness of an action refers to the degree to which an action conforms to the laws of natural physics.
[0154] In some embodiments, the physical constraint values include foot physical constraints, which reflect the physical naturalness (foot sliding) of the foot movements of the first object represented by the predicted action sequence. The calculation of the foot physical constraints is shown in Formula 5 below:
[0155]
[0156] Among them, L foot For the physical constraints of the feet, N is the number of action frames included in the predicted action sequence. and To predict two adjacent action frames in an action sequence.
[0157] In some embodiments, the physical constraint values include a tilt physical constraint, which reflects the degree of tilt of the first object represented by the predicted motion sequence. In some embodiments, the tilt physical constraint is calculated based on the coordinates of the center of gravity of the first object and the coordinates of the feet of the first object in contact with the ground in each motion frame included in the predicted motion sequence.
[0158] In some embodiments, the physical constraint value includes a clipping physical constraint, which reflects the degree of clipping of the action of a first object represented by the predicted action sequence. Exemplarily, the degree of clipping may refer to clipping of the first object's body passing through clothing worn by the first object, or clipping of the first object passing through a virtual boundary (such as a virtual wall).
[0159] The above physical constraint values are merely illustrative examples. Those skilled in the art can set them arbitrarily according to actual needs and the simulation environment in which the first object performs the action. This application does not limit them in this regard.
[0160] Introducing physical constraint values into the objective function can improve the physical naturalness and rationality of the actions generated by the action diffusion model.
[0161] 2. Calculate the regularization term based on the sample noisy sequence and the predicted noisy sequence. The regularization term is used to reflect the difference between the sample noisy sequence and the predicted noisy sequence.
[0162] In some embodiments, the regularization term is calculated using the following formula 6:
[0163]
[0164] Among them, L AFD Let E be the regularization term, representing the expression for finding... The expectation, q(x0) represents the probability distribution of the sample noise sequence x0, q(x t-1 |x0) represents the noisy sequence x obtained by predicting the sample obtained in the (t-1)th noisy step, given q(x0). t-1 The conditional distribution, q(x) t |x t-1 ) represents the known x t-1 The probability distribution q(x) t-1 Predict the sample noise sequence x t The conditional distribution, q(x0)q(x t-1 |x0)q(x t |x t-1 That is, the sample noise sequence x t The probability distribution q(x) t ), β t Let Variance be the variance of the first Gaussian noise at step t. To predict noisy sequences.
[0165] Introducing a regularization term into the objective function ensures the reliability of the adversarial loss when the action diffusion model adjusts its parameters based on the adversarial loss in the objective function. Specifically, a small regularization term and adversarial loss can only be achieved when the sample noisy sequence and the predicted noisy sequence are similar, and the discriminator also determines that they are similar.
[0166] 3. The objective function value is obtained by weighted summation of the reconstruction loss, physical constraint value, regularization term and adversarial loss.
[0167] In some embodiments, the reconstruction loss, physical constraint value, regularization term and adversarial loss are added together to obtain the objective function value.
[0168] In some embodiments, the weight parameters in the above weighted summation process are related to the number of steps T that the action diffusion model takes to remove noise from the noise sequence when performing the action generation task. In some embodiments, the larger T is, the larger the weight parameters corresponding to the reconstruction loss and physical constraint values, and the smaller the weight parameters corresponding to the regularization term and adversarial loss, so that the action diffusion model focuses more on the inference result (i.e., the predicted action sequence) obtained based on the sample noise sequence when large-step denoising is not required. The smaller T is, the smaller the weight parameters corresponding to the reconstruction loss and physical constraint values, and the larger the weight parameters corresponding to the regularization term and adversarial loss, so that the action diffusion model focuses more on the intermediate process of obtaining the inference result based on the sample noise sequence when large-step denoising is required.
[0169] This application does not impose any restrictions on the weighting parameters of the weighted summation process.
[0170] In some embodiments, the parameters of the discriminator are adjusted according to the discriminant value to obtain the trained discriminator.
[0171] In some embodiments, the discriminant value is positively correlated with the similarity mentioned above. For example, the discriminant value is the probability that the sample noisy sequence output by the discriminator is similar to the predicted noisy sequence. The discriminant parameters are adjusted with the goal of minimizing the discriminant value.
[0172] In some embodiments, the discriminant value is negatively correlated with the similarity mentioned above. For example, the discriminant value is the probability that the sample noisy sequence output by the discriminator and the predicted noisy sequence are not similar. The discriminant parameters are adjusted to maximize the discriminant value.
[0173] In some embodiments, the discriminant value is the probability distribution predicted by the discriminator regarding whether the sample noisy sequence and the predicted noisy sequence are similar. The discriminator's parameters are adjusted with the goal of minimizing the difference between the probability distribution and the true distribution. In the true distribution, the probability that the sample noisy sequence and the predicted noisy sequence are similar is 0, and the probability that they are dissimilar is 1.
[0174] In other words, the discriminator adjusts its parameters to minimize the similarity between its prediction of the noisy sequence and the predicted noisy sequence, thus countering the action diffusion model.
[0175] This application does not limit the specific methods for adjusting the parameters of the action spread model and the discriminator. For example, the parameters of the action spread model and the discriminator are adjusted by the Adam optimizer or the AdamW optimizer and the EMA algorithm to obtain the trained action spread model and the trained action spread model.
[0176] In some embodiments, the action diffusion model and discriminator are trained in multiple rounds according to the above embodiments, and the sample time step t corresponding to each round of training is randomly sampled from the interval [1,T].
[0177] In the above embodiment, the training of the action diffusion model and the discriminator forms an adversarial relationship, which enables the action diffusion model to both explicitly learn the ability to reduce the noise of the sample noise sequence to the sample action sequence from the reconstruction loss, and implicitly learn the distribution of the noise intensity of the sample noise sequence and the sample action sequence from the adversarial loss. This allows the action diffusion model to achieve a faster generation rate when generating actions.
[0178] Please refer to Figure 5 This illustrates a flowchart of a training method for an action diffusion model provided in another embodiment of this application. In this embodiment, the action diffusion model includes, for example... Figure 6 The first noise reduction unit 61 and the second noise reduction unit 62 shown, the discriminator includes, as follows Figure 6 The first discriminator 63 and the second discriminator 64 are shown. The execution entity for each step of this method is a computer device, for example, the computer device is... Figure 1 The model training device 10 in the computer system shown. The method includes at least one of the following steps 510 to 590.
[0179] Step 510: Obtain the sample action sequence and sample condition representation.
[0180] For a detailed explanation of the sample action sequence and sample condition representation, please refer to the above embodiment, which will not be repeated here.
[0181] Step 520: Input the sample local noise sequence and sample condition representation into the first denoiser, and the first denoiser outputs the sample local denoised sequence.
[0182] The sample local noise sequence corresponds to the body node of the first object, which is used to control the body movements of the first object. The sample local denoising sequence is used to characterize the body movements of the first object predicted by the first denoiser under the guidance of the sample conditional representation.
[0183] Step 530: Input the sample root noise sequence, sample local denoising sequence and sample conditional representation into the second denoiser, and output the sample root denoising sequence from the second denoiser.
[0184] The sample root noise sequence corresponds to the root node of the first object, which controls the position of the first object in space. The sample root denoising sequence is used to characterize the positional movement of the first object in space predicted by the second denoiser under the guidance of the sample conditional representation and the sample local denoising sequence.
[0185] The sample noise sequence includes the sample root noise sequence and the sample local noise sequence mentioned above. In some embodiments, the sample noise sequence is separated into the sample root noise sequence and the sample local noise sequence according to the attribution of the root node and the body node.
[0186] In some embodiments, the first and second noise reducers described above are both constructed based on the encoder of Transformer, which will not be elaborated further in this application.
[0187] Step 540: Merge the local denoised sequence of the sample and the root denoised sequence of the sample to obtain the predicted action sequence.
[0188] In the above steps, please refer to Figure 6 , the sample noise sequence x t Decoupling for local noise sequences of samples and sample root noise sequence First, let the first noise reduction unit 61 pair the sample local noise sequence. Noise reduction, to generate local denoised sequences Then, the local denoising sequence As new guiding data, the second noise reduction unit is guided by 62 pairs of sample root noise sequences. Denoising to generate the root denoised sequence In this process, the positional movement of the first object as a whole is guided by the local body movements of the first object, which can effectively avoid the problems of unnatural shaking and foot slippage in the generated movements, and effectively improve the stability, naturalness and complexity of the generated movements.
[0189] Step 550: Input the first sample noisy sequence, the first predicted noisy sequence, and the sample conditional representation into the first discriminator, and output the first discriminant value.
[0190] The first sample noisy sequence and the first predicted noisy sequence correspond to the first part of the first object, respectively. The first discriminant value is used to reflect the first similarity. The first similarity is the similarity between the first sample noisy sequence and the first predicted noisy sequence predicted by the first discriminator under the guidance of the sample condition representation.
[0191] Step 560: Input the second sample noisy sequence, the second predicted noisy sequence, and the sample conditional representation into the second discriminator, and output the second discriminant value from the second discriminator.
[0192] The second sample noisy sequence and the second predicted noisy sequence correspond to the remaining parts of the first object except for the first part. The second discriminant value is used to reflect the second similarity. The second similarity is the similarity between the second sample noisy sequence and the second predicted noisy sequence predicted by the second discriminator under the guidance of the sample condition representation.
[0193] The sample noise-added sequence includes the first sample noise-added sequence and the second sample noise-added sequence mentioned above, and the predicted noise-added sequence includes the first predicted noise-added sequence and the second predicted noise-added sequence.
[0194] In some embodiments, based on the attribution of different parts of the first object, the first sample noise sequence and the second sample noise sequence are separated from the sample noise sequence, and the first prediction noise sequence and the second prediction noise sequence are separated from the prediction noise sequence.
[0195] In some embodiments, based on the attribution of different parts of the first object, a first sample action sequence and a second sample action sequence are separated from the sample action sequence. The first sample action sequence corresponds to the first part of the first object, and the second sample action sequence corresponds to the remaining parts of the first object other than the first part. Then, t-1 steps of second sub-noise are added to the first sample action sequence and the second sample action sequence respectively to obtain the first sample noisy sequence and the second sample noisy sequence.
[0196] In some embodiments, based on the attribution of different parts of the first object, a first predicted action sequence and a second predicted action sequence are separated from the predicted action sequence. The first predicted action sequence corresponds to the first part of the first object, and the second predicted action sequence corresponds to the remaining parts of the first object other than the first part. Then, a third sub-noise step t-1 is added to the first predicted action sequence and the second predicted action sequence respectively to obtain the first predicted noisy sequence and the second predicted noisy sequence.
[0197] In the above steps, please refer to Figure 6 Add noise to the sample sequence x t-1 Decoupling for the first sample noisy sequence Second sample noisy sequence Predict the noisy sequence Decoupling for the first predicted noisy sequence Second predicted noisy sequence Let the first discriminator 63 discriminate the first sample noisy sequence. and the first predicted noisy sequence The degree of similarity allows the second discriminator 64 to distinguish the second sample noisy sequence. Second predicted noisy sequence The degree of similarity. In this process, different discriminators are used to perform adversarial learning on the data of different parts of the first object, thereby decoupling the body parts during the training of the action diffusion model, which helps to generate more detailed and natural actions for some parts of the first object.
[0198] In some embodiments, the first part includes the hand of the first object. The first sample action sequence and the first predicted action sequence correspond to the hand of the first object, and the second sample action sequence and the second predicted action sequence correspond to the remaining parts of the first object excluding the hand. Optionally, the hand of the first object includes the wrist, palm, and fingers. When the first part is the hand of the first object, the action diffusion model can generate more complex finger movements that are highly similar to real hand gestures, solving the problem of insufficient finger movements in generated actions in related technologies.
[0199] It should be noted that the first part is not limited to the hand. Technicians can choose any first part according to the needs of motion generation. For example, in the case of generating dance movements, the legs of the first object can be determined as the first part.
[0200] Step 570: Calculate the reconstruction loss based on the sample action sequence and the predicted action sequence.
[0201] Reconstruction loss is used to reflect the difference between the sample action sequence and the predicted action sequence.
[0202] Step 580: Calculate the first adversarial loss based on the first discriminant value, the first adversarial loss being negatively correlated with the first similarity; and calculate the second adversarial loss based on the second discriminant value, the second adversarial loss being negatively correlated with the second similarity.
[0203] In some embodiments, the first discriminant value is positively correlated with the first similarity. For example, the first discriminant value is the probability that the first sample noisy sequence output by the first discriminator and the first predicted noisy sequence are similar. The first discriminant value is inverted or negative to obtain the first adversarial loss. Alternatively, the first adversarial loss is obtained based on the result of inverting or negativening the first discriminant value, so that the first adversarial loss is negatively correlated with the first similarity.
[0204] In some embodiments, the first discriminant value is negatively correlated with the first similarity. For example, the first discriminant value is the probability that the first sample noisy sequence output by the first discriminator and the first predicted noisy sequence are not similar. The first discriminant value is directly determined as the first adversarial loss, or the first adversarial loss is obtained based on the first discriminant value, so that the first adversarial loss is negatively correlated with the first similarity.
[0205] In some embodiments, the first discriminant value is the probability distribution predicted by the first discriminator for whether the first sample noisy sequence and the first predicted noisy sequence are similar. Based on the probability distribution and the first pseudo distribution, the first adversarial loss is obtained. The first adversarial loss is used to reflect the difference between the probability distribution and the pseudo distribution. In the first pseudo distribution, the probability that the first sample noisy sequence and the first predicted noisy sequence are similar is 1, and the probability that they are not similar is 0.
[0206] In other words, the action diffusion model adjusts its parameters with the aim of maximizing the first similarity predicted by the first discriminator.
[0207] In some embodiments, the second discriminant value is positively correlated with the second similarity. For example, the second discriminant value is the probability that the second sample noisy sequence output by the second discriminator is similar to the second predicted noisy sequence. The second discriminant value is inverted or negative to obtain the second adversarial loss. Alternatively, the second adversarial loss is obtained based on the result of inverting or negativening the second discriminant value, so that the second adversarial loss is negatively correlated with the second similarity.
[0208] In some embodiments, the second discriminant value is negatively correlated with the second similarity. For example, the second discriminant value is the probability that the second sample noisy sequence output by the second discriminator is not similar to the second predicted noisy sequence. The second discriminant value is directly determined as the second adversarial loss, or the second adversarial loss is obtained based on the second discriminant value, so that the second adversarial loss is negatively correlated with the second similarity.
[0209] In some embodiments, the second discriminant value is the probability distribution predicted by the second discriminator for whether the second sample noisy sequence and the second predicted noisy sequence are similar. Based on the probability distribution and the second pseudo distribution, a second adversarial loss is obtained. The second adversarial loss is used to reflect the difference between the probability distribution and the pseudo distribution. In the second pseudo distribution, the probability that the second sample noisy sequence and the second predicted noisy sequence are similar is 1, and the probability that they are not similar is 0.
[0210] In other words, the action diffusion model adjusts its parameters to maximize the second similarity predicted by the second discriminator.
[0211] Step 590: Train the action diffusion model based on the reconstruction loss, the first adversarial loss, and the second adversarial loss.
[0212] In some embodiments, step 590 includes at least one of the following sub-steps 592 to 602.
[0213] Sub-step 592: Calculate the physical constraint value based on the predicted action sequence. The physical constraint value is used to reflect the physical naturalness of the action of the first object represented by the predicted action sequence.
[0214] Sub-step 594: Calculate the first regularization term based on the first sample noisy sequence and the first predicted noisy sequence. The first regularization term is used to reflect the difference between the first sample noisy sequence and the first predicted noisy sequence.
[0215] In some embodiments, the first regularization term is calculated using the following formula 7.
[0216]
[0217] Where E represents the search The expectation.
[0218] Introducing a first regularization term into the objective function ensures the reliability of the first adversarial loss when the first denoiser adjusts its parameters based on the first adversarial loss in the objective function. That is, a smaller first regularization term and a smaller first adversarial loss can only be achieved when the first sample denoised sequence and the first predicted denoised sequence are similar, and the first discriminator also determines that they are similar.
[0219] Sub-step 596: Calculate a second regularization term based on the second sample noisy sequence and the second predicted noisy sequence. The second regularization term is used to reflect the difference between the second sample noisy sequence and the second predicted noisy sequence.
[0220] In some embodiments, the second regularization term is calculated using the following formula 8.
[0221]
[0222] Where E represents the search The expectation.
[0223] Introducing a second regularization term into the objective function ensures the reliability of the second adversarial loss when the second denoiser adjusts its parameters based on the second adversarial loss in the objective function. That is, a smaller second regularization term and a smaller second adversarial loss can only be achieved when the second sample noisy sequence and the second predicted noisy sequence are similar to each other, and the second discriminator also determines that they are similar.
[0224] Sub-step 598: Calculate the objective function value based on the reconstruction loss, physical constraint value, first regularization term, second regularization term, first adversarial loss, and second adversarial loss.
[0225] In some embodiments, the reconstruction loss, physical constraint value, first regularization term, second regularization term, first adversarial loss and second adversarial loss are weighted and summed to obtain the objective function value.
[0226] In some embodiments, the weight parameter of the first regularization term is greater than the weight parameter of the second regularization term, and the weight parameter of the first adversarial loss is greater than the weight parameter of the second adversarial loss, thereby making the action diffusion model focus more on the construction of actions for the first part.
[0227] In some embodiments, the objective function value can be obtained by simply summing the above-mentioned reconstruction loss, first adversarial loss, and second adversarial loss.
[0228] Sub-step 602: Adjust the parameters of the action diffusion model according to the objective function value to obtain the trained action diffusion model.
[0229] In some embodiments, the parameters of the action diffusion model are adjusted to minimize the objective function value, resulting in a trained action diffusion model.
[0230] In the above steps, the objective function value of the action diffusion model is obtained by combining the reconstruction loss, physical constraint value, first regularization term, second regularization term, first adversarial loss, and second adversarial loss. During the training process of the action diffusion model, the reconstruction loss guides the explicit objective of adjusting the action diffusion model parameters, namely, learning the ability to reconstruct sample action sequences from sample noise sequences. The first and second adversarial losses guide the implicit objective of adjusting the action diffusion model parameters, namely, learning the intermediate process of reconstructing sample action sequences from sample noise sequences. Furthermore, the first and second adversarial losses are obtained by decoupling the parts of the aforementioned first object. Therefore, the action diffusion model can learn both the ability to reduce noise over large strides and the ability to finely process certain parts. In addition, the introduction of physical constraint value and regularization term in the objective function further optimizes the performance of the action diffusion model, ensuring the naturalness and stability of the actions generated by the action diffusion model.
[0231] In some embodiments, the parameters of the first discriminator are adjusted according to the first discriminant value to obtain the trained first discriminator.
[0232] In some embodiments, the first discriminant value is positively correlated with the first similarity. For example, the first discriminant value is the probability that the first sample noisy sequence output by the first discriminator is similar to the first predicted noisy sequence. The parameters of the first discriminator are adjusted with the goal of minimizing the first discriminant value.
[0233] In some embodiments, the first discriminant value is negatively correlated with the first similarity. For example, the first discriminant value is the probability that the first sample noisy sequence output by the first discriminator and the first predicted noisy sequence are not similar. The parameters of the first discriminator are adjusted to maximize the first discriminant value.
[0234] In some embodiments, the first discriminant value is the probability distribution predicted by the first discriminator regarding whether the first sample noisy sequence and the first predicted noisy sequence are similar. With the goal of minimizing the difference between the probability distribution and the first true distribution, the parameters of the first discriminator are adjusted. In the first true distribution, the probability that the first sample noisy sequence and the first predicted noisy sequence are similar is 0, and the probability that they are not similar is 1.
[0235] In other words, the first discriminator adjusts its parameters to minimize the first similarity, thereby creating an adversarial effect against the action diffusion model.
[0236] In some embodiments, the parameters of the second discriminator are adjusted according to the second discriminant value to obtain the trained second discriminator.
[0237] In some embodiments, the second discriminant value is positively correlated with the second similarity mentioned above. For example, the second discriminant value is the probability that the second sample noisy sequence output by the second discriminator is similar to the second predicted noisy sequence. The parameters of the second discriminator are adjusted with the goal of minimizing the second discriminant value.
[0238] In some embodiments, the second discriminant value is negatively correlated with the second similarity. For example, the second discriminant value is the probability that the second sample noisy sequence output by the second discriminator is not similar to the second predicted noisy sequence. The parameters of the second discriminator are adjusted to maximize the second discriminant value.
[0239] In some embodiments, the second discriminant value is the probability distribution predicted by the second discriminator for whether the second sample noisy sequence and the second predicted noisy sequence are similar. With the goal of minimizing the difference between the probability distribution and the true distribution, the parameters of the second discriminator are adjusted. In the second true distribution, the probability that the second sample noisy sequence and the second predicted noisy sequence are similar is 0, and the probability that they are not similar is 1.
[0240] In other words, the second discriminator adjusts its parameters to minimize the second similarity, thereby creating an adversarial effect against the action diffusion model.
[0241] In the above steps, the training of the action diffusion model and the discriminator (first discriminator and second discriminator) forms a mutual adversarial process. During this adversarial learning process, the action diffusion model can learn the generation process of actions of different parts of the first object, thereby more accurately capturing the characteristics of actions of certain parts and quickly generating complex and detailed actions.
[0242] The following is an introduction to the usage process of the action diffusion model through examples. The content involved in the training process and the content involved in the use of the model are corresponding to each other and are interconnected. If there is no detailed explanation on one side, you can refer to the description on the other side.
[0243] Please refer to Figure 7 The diagram illustrates a flowchart of an action generation method based on an action diffusion model according to an embodiment of this application. The action diffusion model includes a first noise reduction unit and a second noise reduction unit. The execution entity for each step of the method is a computer device, for example, the computer device is... Figure 1 The model in the computer system shown uses device 20. The method includes at least one of the following steps 710 to 740.
[0244] Step 710: Obtain the noise sequence and conditional representation.
[0245] Conditions represent actions used to guide the action diffusion model in generating a second object.
[0246] The second object can be a person, animal, or any other object capable of performing actions; this application does not limit this. The second object and the first object can be the same object or different objects.
[0247] Optionally, the first object and the second object are different virtual entities.
[0248] Optionally, the first object is a real object and the second object is a virtual object, for example, the first object is a real person and the second object is a virtual person.
[0249] In some embodiments, the second object and the first object have the same parts, such as both the second object and the first object having a hand.
[0250] A noisy sequence is a data sequence with an irregular data distribution. Specifically, a noisy sequence consists of n noisy frames, in which the data is distributed irregularly.
[0251] It should be noted that, in the embodiments of this application, the data in the noise sequence has no definite meaning, but those skilled in the art can still specify the correspondence between the data in the noise sequence and the root node and body node of the second object.
[0252] Step 720: Input the local noise sequence and conditional representation in the noise sequence into the first denoiser, and output the local denoised sequence from the first denoiser.
[0253] The local noise sequence corresponds to the body node of the second object, which is used to control the body action of the second object. The local denoising sequence is used to characterize the body action of the second object predicted by the first denoiser under the guidance of the conditional representation.
[0254] Step 730: Input the root noise sequence, local denoising sequence and conditional representation in the noise sequence into the second denoiser, and output the root denoising sequence from the second denoiser.
[0255] The root noise sequence corresponds to the root node of the second object, which controls the position of the second object in space. The local denoising sequence is used to characterize the positional movement of the second object in space predicted by the second denoiser under the guidance of the conditional representation and the local denoising sequence.
[0256] Step 740: Merge the local denoised sequence and the root denoised sequence to obtain the generated action sequence.
[0257] The generated action sequence is used to characterize the actions of the second object predicted by the action diffusion model under the guidance of the conditional representation.
[0258] In some embodiments, the generated action sequence includes n action frames predicted by the action diffusion model. Each action frame in the generated action sequence represents an action of a second object predicted by the action diffusion model.
[0259] The technical solution provided in this application uses a local denoised sequence to guide the generation of a root denoised sequence. This enables the second object to move as a whole by using the local body movements of the second object during the action generation process. This effectively avoids the problems of unnatural shaking and foot slippage in the generated action, and effectively improves the stability, naturalness and complexity of the generated action.
[0260] In some embodiments, the above conditions may include at least one of the following: audio features, style coding, seed sequence, and time steps.
[0261] Audio features are used to characterize the audio features corresponding to the action sequence to be generated. The audio corresponding to the action sequence to be generated can be arbitrarily selected by the technician according to the action generation requirements, and this application does not impose any restrictions on this.
[0262] Style coding is used to characterize the emotional style of the action sequence to be generated, such as "sad", "happy", "joyful", "serious", etc.
[0263] A seed sequence, which is a type of action sequence, is used to guide the action diffusion model in generating action sequences that are connected to the seed sequence. For example, the seed sequence includes eight action frames, which are used to guide the action diffusion model in generating action frames that can be connected to these eight action frames.
[0264] The time step count indicates the number of steps the action diffusion model takes to remove noise from the noise sequence during the process of generating the action sequence. It should be noted that in this embodiment, the model training process enables the action diffusion model to have a larger step size for noise reduction; therefore, the time step count can be set to a smaller number, such as T = 10 steps, thereby accelerating action generation.
[0265] By introducing the above conditions, the action diffusion model can generate more natural, fluid, and expressive actions.
[0266] It should be noted that the conditional representation in the model usage process corresponds to the sample conditional representation in the model training process, and this application does not limit the specific content of the conditional representation. Any data that can play a guiding role in the process of the action diffusion model generating actions can be used as the conditional representation in the embodiments of this application.
[0267] In some embodiments, the conditional representation may further include at least one of the following: semantic encoding, text features, facial expression data, and pose data.
[0268] Semantic encoding is used to describe the actions that need to be generated. For example, encoding the text describing the actions to be generated yields semantic encoding.
[0269] Text features are used to characterize the text content of the audio corresponding to the action sequence to be generated.
[0270] Facial expression data, used to characterize the facial expressions of a second object that match the action to be generated.
[0271] Pose data is used to characterize the overall pose of a second object that matches the action to be generated, such as standing, sitting, or lying down.
[0272] The action generation method based on the action diffusion model provided in this application can be applied to various scenarios. The specific implementation details are illustrated in the following examples.
[0273] In some embodiments, the training method for the action diffusion model provided in this application is applied to video production scenarios. Please refer to [reference needed]. Figure 8 The method includes at least one of the following steps 810 to 850:
[0274] Step 810: Obtain the sample action sequence and sample condition representation. The sample action sequence is used to characterize the actions of the character in the sample video, and the sample condition representation is used to guide the action diffusion model to generate the actions of the character in the sample video.
[0275] A sample video character is a character that shares the same body parts as the character to be driven in the first video to be produced. The sample video character can be a real object; for example, if the virtual object to be driven is a virtual person, the sample video character is a real person; if the virtual object to be driven is a virtual animal, the sample video character is a real animal. In some embodiments, the sample video character is a character from a pre-produced sample video.
[0276] In some embodiments, the sample condition representation includes at least one of the following: sample audio features, sample style encoding, sample seed sequence, sample seed step number, sample semantic encoding, sample text features, sample facial expression data, and sample pose data.
[0277] For a detailed explanation of the sample action sequence and sample condition representation, please refer to the above embodiment, which will not be repeated here.
[0278] Step 820: Add first noise to the sample action sequence to obtain the sample noise sequence.
[0279] In some embodiments, the first noise includes t-step first sub-noise, and t-step first sub-noise is added to the sample action sequence to obtain the sample noise sequence.
[0280] Step 830: Input the sample noise sequence and sample conditional representation into the action diffusion model, and output the predicted action sequence from the action diffusion model. The predicted action sequence is used to characterize the action of the sample video character predicted by the action diffusion model under the guidance of the sample conditional representation.
[0281] In some embodiments, the motion diffusion model includes a first denoiser and a second denoiser. The sample noise sequence includes a sample root noise sequence and a sample local noise sequence. The sample root noise sequence corresponds to the root node of the sample video character, and the sample local noise sequence corresponds to the body node of the sample video character. The root node of the sample video character is used to control the position of the sample video character in space, and the body node of the sample video character is used to control the body movements of the sample video character. Step 830 includes at least one of the following sub-steps 832 to 836.
[0282] Sub-step 832: Input the sample local noise sequence and sample conditional representation into the first denoiser, and output the sample local denoising sequence from the first denoiser. The sample local denoising sequence is used to characterize the body movements of the sample video characters predicted by the first denoiser under the guidance of the sample conditional representation.
[0283] Sub-step 834: Input the sample root noise sequence, sample local denoising sequence and sample conditional representation into the second denoiser, and output the sample root denoising sequence from the second denoiser. The sample root denoising sequence is used to characterize the spatial positional movement of the sample video character predicted by the second denoiser under the guidance of the sample conditional representation and sample local denoising sequence.
[0284] Sub-step 836: merge the local denoised sequence of the sample and the root denoised sequence of the sample to obtain the predicted action sequence.
[0285] Step 840: Add second noise to the sample action sequence to obtain a noisy sample sequence, and add third noise to the predicted action sequence to obtain a noisy predicted sequence.
[0286] The noise intensity of the second noise and the third noise are both lower than the noise intensity of the first noise.
[0287] In some embodiments, the second noise includes a second sub-noise at step t-1, and the third noise includes a third sub-noise at step t-1.
[0288] In some embodiments, a second sub-noise step t-1 is added to the sample action sequence to obtain a sample noisy sequence, wherein the variance of the i-th step second sub-noise in the t-1 step second sub-noise is the same as the variance of the i-th step first sub-noise in the t-th step first sub-noise, and i is a positive integer less than or equal to t-1.
[0289] In some embodiments, a third sub-noise of step t-1 is added to the predicted action sequence to obtain a predicted noisy sequence, wherein the variance of the third sub-noise of step t-1 is the same as the variance of the first sub-noise of step t-1.
[0290] Step 850: Train the action diffusion model with the training objective of increasing the similarity between the sample noisy sequence and the predicted noisy sequence.
[0291] In some embodiments, step 850 includes at least one of the following sub-steps 852 to 858.
[0292] Sub-step 852: Input the sample noisy sequence, the predicted noisy sequence, and the sample conditional representation into the discriminator, and have the discriminator output a discrimination value. The discrimination value is used to reflect the similarity predicted by the discriminator for the sample noisy sequence and the predicted noisy sequence under the guidance of the sample conditional representation.
[0293] In some embodiments, the discriminator includes a first discriminator and a second discriminator, the sample noise-adding sequence includes a first sample noise-adding sequence and a second sample noise-adding sequence, the predicted noise-adding sequence includes a first predicted noise-adding sequence and a second predicted noise-adding sequence, the first sample noise-adding sequence and the first predicted noise-adding sequence correspond to a first part of a character in the sample video, the second sample noise-adding sequence and the second predicted noise-adding sequence correspond to the remaining parts of the character in the sample video excluding the first part, and the discriminant value includes a first discriminant value and a second discriminant value.
[0294] Sub-step 852 includes: inputting a first sample noisy sequence, a first predicted noisy sequence, and a sample conditional representation into a first discriminator, and having the first discriminator output a first discriminant value, the first discriminant value being used to reflect a first similarity, the first similarity being the similarity predicted by the first discriminator for the first sample noisy sequence and the first predicted noisy sequence under the guidance of the sample conditional representation; and inputting a second sample noisy sequence, a second predicted noisy sequence, and a sample conditional representation into a second discriminator, and having the second discriminator output a second discriminant value, the second discriminant value being used to reflect a second similarity, the second similarity being the similarity predicted by the second discriminator for the second sample noisy sequence and the second predicted noisy sequence under the guidance of the sample conditional representation.
[0295] In some embodiments, the first part includes the hand of the character in the sample video.
[0296] Sub-step 854: Calculate the reconstruction loss based on the sample action sequence and the predicted action sequence. The reconstruction loss is used to reflect the difference between the sample action sequence and the predicted action sequence.
[0297] Sub-step 856: Calculate the adversarial loss based on the discriminant value. The adversarial loss is negatively correlated with the similarity predicted by the discriminator for the noisy sequence and the predicted noisy sequence under the guidance of the sample conditional representation.
[0298] Sub-step 858: Train the action diffusion model based on reconstruction loss and adversarial loss.
[0299] In some embodiments, sub-step 858 includes at least one of the following steps:
[0300] 1. Calculate the physical constraint value based on the predicted action sequence. The physical constraint value is used to reflect the physical naturalness of the actions of the sample video characters represented by the predicted action sequence.
[0301] 2. Calculate a first regularization term based on the first sample noisy sequence and the first predicted noisy sequence. The first regularization term is used to reflect the difference between the first sample noisy sequence and the first predicted noisy sequence. Calculate a second regularization term based on the second sample noisy sequence and the second predicted noisy sequence. The second regularization term is used to reflect the difference between the second sample noisy sequence and the second predicted noisy sequence.
[0302] 3. Based on the reconstruction loss, physical constraint value, first regularization term, second regularization term, first adversarial loss and second adversarial loss, the objective function value is calculated. The objective function value is used to reflect the performance of the action diffusion model in performing the action generation task.
[0303] 4. Adjust the parameters of the action diffusion model according to the objective function value to obtain the trained action diffusion model.
[0304] In some embodiments, the motion generation method based on the motion diffusion model provided in this application is applied to video production scenarios. Please refer to [link / reference]. Figure 9 The method includes at least one of the following steps 910 to 940.
[0305] Step 910: Obtain the noise sequence and conditional representation. The conditional representation is used to guide the trained motion diffusion model to generate the actions of the video characters.
[0306] In some embodiments, a video character is a computer-generated character in a first video to be produced.
[0307] In some embodiments, the conditional representation includes at least one of the following: audio features, style encoding, seed sequence, seed step, semantic encoding, text features, facial expression data, and pose data.
[0308] Audio features are used to characterize the audio features configured for the aforementioned video characters in the first video.
[0309] Style coding is used to characterize the emotional style of the actions of the aforementioned video characters in the first video.
[0310] Semantic encoding is used to describe the actions of the aforementioned video characters that need to be generated.
[0311] Text features are used to characterize the dialogue of the aforementioned video characters in the first video.
[0312] Facial expression data is used to represent the facial expressions of the aforementioned video characters in the first video.
[0313] Posture data is used to characterize the overall posture of the aforementioned video characters in the first video, such as standing, sitting, lying down, etc.
[0314] Step 920: Input the local noise sequence and conditional representation in the noise sequence into the first denoiser, and output the local denoised sequence from the first denoiser.
[0315] The local noise sequence corresponds to the body node of the video character, which is used to control the position of the video character in space. The local denoising sequence is used to characterize the body action of the video character predicted by the first denoiser under the guidance of the conditional representation.
[0316] Step 930: Input the root noise sequence, local denoising sequence and conditional representation in the noise sequence into the second denoiser, and the second denoiser outputs the root denoising sequence.
[0317] The root noise sequence corresponds to the root node of the video character, which controls the position of the video character in space. The local denoising sequence is used to characterize the positional movement of the video character in space predicted by the second denoiser under the guidance of the conditional representation and the local denoising sequence.
[0318] Step 940: Merge the local denoised sequence and the root denoised sequence to obtain the generated action sequence.
[0319] The generated action sequences are used to characterize the actions of video characters predicted by the trained action diffusion model under the guidance of conditional representations.
[0320] In some embodiments, a first video is created by using generated motion sequences to control the movement of video characters within a scene configured for the first video.
[0321] In some embodiments, the training method for the motion diffusion model provided in this application is applied to game character control scenarios. Please refer to [reference needed]. Figure 10 The method includes at least one of the following steps 1010 to 1050:
[0322] Step 1010: Obtain sample action sequences and sample condition representations. The sample action sequences are used to characterize the actions of the sample game characters, and the sample condition representations are used to guide the action diffusion model to generate the actions of the sample game characters.
[0323] The sample game character is a character that has the same body parts as the game character that needs to be driven in the game application.
[0324] The sample game characters can be real objects. For example, if the game character to be driven is a virtual character, the sample game character is a real person; if the game character to be driven is a virtual animal, the sample game character is a real animal.
[0325] In some embodiments, the sample condition representation includes at least one of the following: sample audio features, sample style encoding, sample seed sequence, sample time steps, sample semantic encoding, sample text features, sample facial expression data, and sample pose data.
[0326] For a detailed explanation of the sample action sequence and sample condition representation, please refer to the above embodiment, which will not be repeated here.
[0327] Step 1020: Add first noise to the sample action sequence to obtain the sample noise sequence.
[0328] In some embodiments, the first noise includes t-step first sub-noise, and t-step first sub-noise is added to the sample action sequence to obtain the sample noise sequence.
[0329] Step 1030: Input the sample noise sequence and sample conditional representation into the action diffusion model, and output the predicted action sequence from the action diffusion model. The predicted action sequence is used to characterize the actions of the sample game characters predicted by the action diffusion model under the guidance of the sample conditional representation.
[0330] In some embodiments, the motion diffusion model includes a first denoiser and a second denoiser. The sample noise sequence includes a sample root noise sequence and a sample local noise sequence. The sample root noise sequence corresponds to the root node of the sample game character, and the sample local noise sequence corresponds to the body node of the sample game character. The root node of the sample game character is used to control the position of the sample game character in space, and the body node of the sample game character is used to control the body movements of the sample game character. Step 1030 includes at least one of the following sub-steps 1032 to 1036.
[0331] Sub-step 1032: Input the sample local noise sequence and sample conditional representation into the first denoiser, and output the sample local denoising sequence from the first denoiser. The sample local denoising sequence is used to characterize the body movements of the sample game character predicted by the first denoiser under the guidance of the sample conditional representation.
[0332] Sub-step 1034: Input the sample root noise sequence, sample local denoising sequence and sample conditional representation into the second denoiser, and output the sample root denoising sequence from the second denoiser. The sample root denoising sequence is used to characterize the positional movement of the sample game character in space predicted by the second denoiser under the guidance of the sample conditional representation and sample local denoising sequence.
[0333] Sub-step 1036: Merge the local denoised sequence of the sample and the root denoised sequence of the sample to obtain the predicted action sequence.
[0334] Step 1040: Add second noise to the sample action sequence to obtain a noisy sample sequence, and add third noise to the predicted action sequence to obtain a noisy predicted sequence.
[0335] The noise intensity of the second noise and the third noise are both lower than the noise intensity of the first noise.
[0336] In some embodiments, the second noise includes a second sub-noise at step t-1, and the third noise includes a third sub-noise at step t-1.
[0337] In some embodiments, a second sub-noise step t-1 is added to the sample action sequence to obtain a sample noisy sequence, wherein the variance of the i-th step second sub-noise in the t-1 step second sub-noise is the same as the variance of the i-th step first sub-noise in the t-th step first sub-noise, and i is a positive integer less than or equal to t-1.
[0338] In some embodiments, a third sub-noise of step t-1 is added to the predicted action sequence to obtain a predicted noisy sequence, wherein the variance of the third sub-noise of step t-1 is the same as the variance of the first sub-noise of step t-1.
[0339] Step 1050: Train the action diffusion model with the goal of increasing the similarity between the sample noisy sequence and the predicted noisy sequence.
[0340] In some embodiments, step 1050 includes at least one of the following sub-steps 1052 to 1058.
[0341] Sub-step 1052: Input the sample noisy sequence, the predicted noisy sequence, and the sample conditional representation into the discriminator, and output the discriminator value. The discriminator value is used to reflect the similarity predicted by the discriminator for the sample noisy sequence and the predicted noisy sequence under the guidance of the sample conditional representation.
[0342] In some embodiments, the discriminator includes a first discriminator and a second discriminator, the sample noise sequence includes a first sample noise sequence and a second sample noise sequence, the prediction noise sequence includes a first prediction noise sequence and a second prediction noise sequence, the first sample noise sequence and the first prediction noise sequence correspond to a first part of the sample game character, the second sample noise sequence and the second prediction noise sequence correspond to the remaining parts of the sample game character other than the first part, and the discrimination value includes a first discrimination value and a second discrimination value.
[0343] Sub-step 1052 includes: inputting a first sample noisy sequence, a first predicted noisy sequence, and a sample conditional representation into a first discriminator, and having the first discriminator output a first discriminant value, the first discriminant value being used to reflect a first similarity, the first similarity being the similarity predicted by the first discriminator for the first sample noisy sequence and the first predicted noisy sequence under the guidance of the sample conditional representation; and inputting a second sample noisy sequence, a second predicted noisy sequence, and a sample conditional representation into a second discriminator, and having the second discriminator output a second discriminant value, the second discriminant value being used to reflect a second similarity, the second similarity being the similarity predicted by the second discriminator for the second sample noisy sequence and the second predicted noisy sequence under the guidance of the sample conditional representation.
[0344] In some embodiments, the first part includes the hand of the sample game character.
[0345] Sub-step 1054: Calculate the reconstruction loss based on the sample action sequence and the predicted action sequence. The reconstruction loss is used to reflect the difference between the sample action sequence and the predicted action sequence.
[0346] Sub-step 1056: Calculate the adversarial loss based on the discriminant value. The adversarial loss is negatively correlated with the similarity predicted by the discriminator for the noisy sequence and the predicted noisy sequence under the guidance of the sample conditional representation.
[0347] Sub-step 1058: Train the action diffusion model based on reconstruction loss and adversarial loss.
[0348] In some embodiments, sub-step 1058 includes at least one of the following steps:
[0349] 1. Calculate the physical constraint value based on the predicted action sequence. The physical constraint value is used to reflect the physical naturalness of the actions of the sample game characters represented by the predicted action sequence.
[0350] 2. Calculate a first regularization term based on the first sample noisy sequence and the first predicted noisy sequence. The first regularization term is used to reflect the difference between the first sample noisy sequence and the first predicted noisy sequence. Calculate a second regularization term based on the second sample noisy sequence and the second predicted noisy sequence. The second regularization term is used to reflect the difference between the second sample noisy sequence and the second predicted noisy sequence.
[0351] 3. Based on the reconstruction loss, physical constraint value, first regularization term, second regularization term, first adversarial loss and second adversarial loss, the objective function value is calculated. The objective function value is used to reflect the performance of the action diffusion model in performing the action generation task.
[0352] 4. Adjust the parameters of the action diffusion model according to the objective function value to obtain the trained action diffusion model.
[0353] In some embodiments, the motion training method based on the motion diffusion model provided in this application is applied to game character control scenarios. Please refer to... Figure 11 The method includes at least one of the following steps 1110 to 1140.
[0354] Step 1110: Obtain the noise sequence and conditional representation. The conditional representation is used to guide the trained action diffusion model to generate the actions of the game character.
[0355] In some embodiments, game characters are NPCs in the game.
[0356] In some embodiments, the conditional representation includes at least one of the following: audio features, style encoding, seed sequence, seed step, semantic encoding, text features, facial expression data, and pose data.
[0357] Audio features are used to characterize the audio features of the aforementioned game characters within a game application.
[0358] Optionally, the game character's audio includes the audio of the game character's response to the user's interactive actions.
[0359] Style coding is used to characterize the emotional style of the actions of the aforementioned game characters in a game application.
[0360] Semantic encoding is used to describe the actions of the game characters that need to be generated.
[0361] Text features are used to characterize the dialogue of the aforementioned game characters within a game application.
[0362] Optionally, the game character's dialogue includes the text content in response to the user's interactive actions.
[0363] Facial expression data, used to characterize the facial expressions of the aforementioned game characters within the game application.
[0364] Optionally, the expressions of the game characters include the expressions of the game characters when responding to the user's interactive operations.
[0365] Pose data is used to characterize the overall posture of the game character in the game application, such as standing, sitting, lying down, etc.
[0366] Optionally, the overall posture of the game character includes the overall posture of the game character when responding to the user's interactive operations.
[0367] Step 1120: Input the local noise sequence and conditional representation in the noise sequence into the first denoiser, and output the local denoised sequence from the first denoiser.
[0368] The local noise sequence corresponds to the body nodes of the game character, which are used to control the position of the game character in space. The local denoising sequence is used to characterize the body movements of the game character predicted by the first denoiser under the guidance of the conditional representation.
[0369] Step 1130: Input the root noise sequence, local denoising sequence and conditional representation in the noise sequence into the second denoiser, and the second denoiser outputs the root denoising sequence.
[0370] The root noise sequence corresponds to the root node of the game character, which controls the position of the game character in space. The local denoising sequence is used to characterize the positional movement of the game character in space predicted by the second denoiser under the guidance of the conditional representation and the local denoising sequence.
[0371] Step 1140: Merge the local denoised sequence and the root denoised sequence to obtain the generated action sequence.
[0372] The generated action sequences are used to represent the actions of game characters predicted by the trained action diffusion model under the guidance of conditional representations.
[0373] In some embodiments, generated motion sequences are used to control the movement of game characters in the game scene.
[0374] In the aforementioned game character control scenario, the game character can respond to user input and generate corresponding actions; therefore, the game character control scenario belongs to the human-computer interaction scenario. Due to the significant improvement in action generation speed achieved in this application—up to 272 frames per second—it is sufficient to support real-time applications in human-computer interaction scenarios, enabling action generation synchronized with audio, and thus possesses strong application value.
[0375] It should be noted that the above application scenarios are merely illustrative examples. The technical solutions provided in this application can be applied to various scenarios that require the generation of actions, and this application does not limit them.
[0376] The following section describes the specific performance of the action diffusion model provided in this application embodiment in performing action generation tasks, using test data.
[0377] Please refer to Figure 12 The illustration shows a schematic diagram of the results of driving a virtual character using action sequences generated by different models, according to an embodiment of this application. During the test, the condition representation includes audio features, the text of which reads, "I saw him waiting for another person bleeding on the ground."
[0378] As can be seen, the motion diffusion model provided in this application, regardless of whether the time step is set to 10 or 50, can generate motions that are better aligned with the audio melody, more distinctive in style, and more expressive than DDPM and DDIM.
[0379] Table 1 below compares the FGD and generation speed of various models in related technologies with the action diffusion model provided in the embodiments of this application on the same test set when performing action generation tasks.
[0380] Table 1
[0381] Model FGD Reasoning time Time Steps DSG 15.67 6.01s 1000 W / O SEMI-IMPLICIT 11.95 11.9s 1000 DDPM 84.7 0.28s 20 DDIM 21.4 0.93s 100 Action diffusion model 12.57 0.29s 10
[0382] As can be seen, the action diffusion model provided in this application can achieve a low FGD error while setting a small number of time steps and having a very fast inference speed, while ensuring the speed and quality of action generation.
[0383] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0384] Please refer to Figure 13 This diagram illustrates a block diagram of a training apparatus for an action diffusion model according to an embodiment of this application. The apparatus has the function of implementing the training method for the aforementioned action diffusion model; this function can be implemented in hardware or by hardware executing corresponding software. The apparatus can be a computer device or can be installed within a computer device. The apparatus 1300 may include: an acquisition module 1310, a first noise-adding module 1320, a generation module 1330, a second noise-adding module 1340, and a training module 1350.
[0385] The acquisition module 1310 is used to acquire a sample action sequence and a sample condition representation, wherein the sample action sequence is used to characterize the action of the first object.
[0386] The first noise-adding module 1320 is used to add first noise to the sample action sequence to obtain a sample noise sequence.
[0387] The generation module 1330 is used to input the sample noise sequence and sample condition representation into the action diffusion model, and the action diffusion model outputs a predicted action sequence, which is used to characterize the action of the first object predicted by the action diffusion model under the guidance of the sample condition representation.
[0388] The second noise-adding module 1340 is further configured to add a second noise to the sample action sequence to obtain a sample noise-adding sequence, and to add a third noise to the predicted action sequence to obtain a predicted noise-adding sequence, wherein the noise intensity of the second noise and the noise intensity of the third noise are both lower than the noise intensity of the first noise.
[0389] Training module 1350 is used to train the action diffusion model with the training objective of increasing the similarity between the sample noisy sequence and the predicted noisy sequence. In some embodiments, the discriminator includes a first discriminator and a second discriminator, the sample noisy sequence includes a first sample noisy sequence and a second sample noisy sequence, the predicted noisy sequence includes a first predicted noisy sequence and a second predicted noisy sequence, the first sample noisy sequence and the first predicted noisy sequence respectively correspond to a first part of the first object, the second sample noisy sequence and the second predicted noisy sequence respectively correspond to the remaining parts of the first object excluding the first part, and the discriminant value includes a first discriminant value and a second discriminant value.
[0390] In some embodiments, the training module 1350 includes: a discrimination submodule, a calculation submodule, and a training submodule.
[0391] The discriminant submodule is used to input the sample noisy sequence, the predicted noisy sequence, and the sample condition representation into the discriminator, and the discriminator outputs a discrimination value, which reflects the similarity predicted by the discriminator for the sample noisy sequence and the predicted noisy sequence under the guidance of the sample condition representation.
[0392] The calculation submodule is used to calculate the reconstruction loss based on the sample action sequence and the predicted action sequence, wherein the reconstruction loss is used to reflect the difference between the sample action sequence and the predicted action sequence;
[0393] The calculation submodule is also used to calculate the adversarial loss based on the discriminant value, wherein the adversarial loss is negatively correlated with the similarity.
[0394] The training submodule is used to train the action diffusion model based on the reconstruction loss and the adversarial loss.
[0395] In some embodiments, the discriminator includes a first discriminator and a second discriminator, the sample noise sequence includes a first sample noise sequence and a second sample noise sequence, the prediction noise sequence includes a first prediction noise sequence and a second prediction noise sequence, the first sample noise sequence and the first prediction noise sequence correspond to a first part of the first object, the second sample noise sequence and the second prediction noise sequence correspond to the remaining parts of the first object other than the first part, and the discrimination value includes a first discrimination value and a second discrimination value.
[0396] The discriminant submodule is used to input the first sample noisy sequence, the first predicted noisy sequence, and the sample condition representation into the first discriminator, and the first discriminator outputs the first discriminant value, which reflects the first similarity. The first similarity is the similarity predicted by the first discriminator for the first sample noisy sequence and the first predicted noisy sequence under the guidance of the sample condition representation. The second discriminator inputs the second sample noisy sequence, the second predicted noisy sequence, and the sample condition representation into the second discriminator, and the second discriminator outputs the second discriminant value, which reflects the second similarity. The second similarity is the similarity predicted by the second discriminator for the second sample noisy sequence and the second predicted noisy sequence under the guidance of the sample condition representation.
[0397] In some embodiments, the calculation submodule is configured to: calculate a first adversarial loss based on the first discriminant value, wherein the first adversarial loss is negatively correlated with the first similarity; and calculate a second adversarial loss based on the second discriminant value, wherein the second adversarial loss is negatively correlated with the second similarity.
[0398] In some embodiments, the training submodule is configured to: calculate physical constraint values based on the predicted action sequence, the physical constraint values reflecting the physical naturalness of the action of the first object represented by the predicted action sequence; calculate a first regularization term based on the first sample noisy sequence and the first predicted noisy sequence, the first regularization term reflecting the difference between the first sample noisy sequence and the first predicted noisy sequence; and calculate a second regularization term based on the second sample noisy sequence and the second predicted noisy sequence, the second regularization term reflecting the difference between the second sample noisy sequence and the second predicted noisy sequence; calculate an objective function value based on the reconstruction loss, the physical constraint values, the first regularization term, the second regularization term, the first adversarial loss, and the second adversarial loss, the objective function value reflecting the performance of the action diffusion model in performing the action generation task; and adjust the parameters of the action diffusion model based on the objective function value to obtain the trained action diffusion model.
[0399] In some embodiments, the first part includes the hand of the first object.
[0400] In some embodiments, the motion diffusion model includes a first noise denoiser and a second noise denoiser, the sample noise sequence includes a sample root noise sequence and a sample local noise sequence, the sample root noise sequence corresponds to the root node of the first object, the sample local noise sequence corresponds to the body node of the first object, the root node is used to control the position of the first object in space, and the body node is used to control the body motion of the first object.
[0401] The generation module 1330 is configured to input the sample local noise sequence and the sample condition representation into the first denoiser, and output a sample local denoising sequence from the first denoiser. The sample local denoising sequence is used to characterize the body movements of the first object predicted by the first denoiser under the guidance of the sample condition representation. The generation module 1330 is configured to input the sample root noise sequence, the sample local denoising sequence, and the sample condition representation into the second denoiser, and output a sample root denoising sequence from the second denoiser. The sample root denoising sequence is used to characterize the positional movement of the first object in the space predicted by the second denoiser under the guidance of the sample condition representation and the sample local denoising sequence. The generation module 1330 is configured to merge the sample local denoising sequence and the sample root denoising sequence to obtain the predicted action sequence.
[0402] In some embodiments, the generation module 1330 is further configured to acquire a noise sequence and a conditional representation, the conditional representation being used to guide the trained action diffusion model to generate the action of the second object; inputting a local noise sequence from the noise sequence and the conditional representation into the first denoiser, and outputting a local denoised sequence from the first denoiser, wherein the local noise sequence corresponds to the body nodes of the second object, the body nodes of the second object being used to control the position of the second object in the space, and the local denoised sequence being used to characterize the body action of the second object predicted by the first denoiser under the guidance of the conditional representation; and inputting a root noise sequence from the noise sequence and the conditional representation into the first denoiser. The local denoising sequence and the conditional representation are input to the second denoiser, which outputs a root denoising sequence. The root denoising sequence corresponds to the root node of the second object, which controls the position of the second object in the space. The local denoising sequence characterizes the positional movement of the second object in the space predicted by the second denoiser under the guidance of the conditional representation and the local denoising sequence. The local denoising sequence and the root denoising sequence are merged to obtain a generated action sequence, which characterizes the action of the second object predicted by the trained action diffusion model under the guidance of the conditional representation.
[0403] In some embodiments, the first noise includes t-step first sub-noise, the second noise includes t-1-step second sub-noise, and the third noise includes t-1-step third sub-noise; the first noise-adding module 1320 is used to add t-step first sub-noise to the sample action sequence to obtain the sample noise sequence, where t is an integer greater than 1;
[0404] The second noise-adding module 1340 is used to add the second sub-noise of step t-1 to the sample action sequence to obtain the sample noise-adding sequence, wherein the variance of the second sub-noise of step i in the second sub-noise of step t-1 is the same as the variance of the first sub-noise of step i in the first sub-noise of step t, and i is a positive integer less than or equal to t-1; and to add the third sub-noise of step t-1 to the predicted action sequence to obtain the predicted noise-adding sequence, wherein the variance of the third sub-noise of step i in the third sub-noise of step t-1 is the same as the variance of the first sub-noise of step i in the first sub-noise of step t.
[0405] In some embodiments, the sample condition representation includes at least one of the following: sample audio features, used to characterize the features of the audio corresponding to the sample action sequence; sample style encoding, used to characterize the emotional style of the sample action sequence; sample seed sequence, including an action sequence that precedes the sample action sequence and is sequential with the sample action sequence; and sample time steps, used to indicate the number of steps in which noise is added to the sample action sequence during the process of obtaining the sample noise sequence.
[0406] Please refer to Figure 14 The diagram illustrates a structural block diagram of a computer device provided in one embodiment of this application.
[0407] Typically, computer device 1400 includes a processor 1401 and a memory 1402.
[0408] Processor 1401 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1401 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1401 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1401 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 1401 may also include an AI processor for handling computational operations related to machine learning.
[0409] The memory 1402 may include one or more computer-readable storage media, which may be tangible and non-transitory. The memory 1402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1402 stores a computer program that is loaded and executed by the processor 1401 to implement the above-described training method for the action diffusion model, or to implement the above-described action generation method based on the action diffusion model.
[0410] Those skilled in the art will understand that Figure 14 The structure shown does not constitute a limitation on the computer device 1400, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0411] In some embodiments, a computer-readable storage medium is also provided, wherein a computer program is stored in the storage medium, the computer program being loaded and executed by a processor to implement the above-described training method for the action diffusion model, or to implement the above-described action generation method based on the action diffusion model.
[0412] Optionally, the computer-readable storage medium may include: ROM (Read-Only Memory), RAM (Random-Access Memory), SSD (Solid State Drives), or optical disc, etc. The random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0413] In some embodiments, a computer program product is also provided, the computer program product including a computer program stored in a computer-readable storage medium, the processor reading from the computer-readable storage medium and executing the computer program to implement the above-described training method for the action diffusion model, or to implement the above-described action generation method based on the action diffusion model.
[0414] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.
[0415] The above are merely exemplary embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application shall be included within the protection scope of this application.
Claims
1. A training method for an action diffusion model, characterized in that, The method includes: Obtain a sample action sequence and a sample condition representation, wherein the sample action sequence is used to characterize the action of the first object; Add first noise to the sample action sequence to obtain the first sample noisy sequence; The first sample noisy sequence and the sample conditional representation are input into the action diffusion model, and the action diffusion model outputs a predicted action sequence, which is used to characterize the action of the first object predicted by the action diffusion model under the guidance of the sample conditional representation. A second noise is added to the sample action sequence to obtain a second sample noisy sequence, and a third noise is added to the predicted action sequence to obtain a predicted noisy sequence, wherein the noise intensity of the second noise and the noise intensity of the third noise are both lower than the noise intensity of the first noise; The action diffusion model is trained with the goal of increasing the similarity between the second sample noisy sequence and the predicted noisy sequence.
2. The method according to claim 1, characterized in that, The training of the action diffusion model with the goal of increasing the similarity between the sample noisy sequence and the predicted noisy sequence includes: The sample noisy sequence, the predicted noisy sequence, and the sample conditional representation are input into the discriminator, and the discriminator outputs a discrimination value. The discrimination value is used to reflect the similarity between the sample noisy sequence and the predicted noisy sequence predicted by the discriminator under the guidance of the sample conditional representation. Based on the sample action sequence and the predicted action sequence, a reconstruction loss is calculated, which reflects the difference between the sample action sequence and the predicted action sequence. Based on the discriminant value, the adversarial loss is calculated, and the adversarial loss is negatively correlated with the similarity. The action diffusion model is trained based on the reconstruction loss and the adversarial loss.
3. The method according to claim 2, characterized in that, The discriminator includes a first discriminator and a second discriminator; the sample noise-adding sequence includes a first sample noise-adding sequence and a second sample noise-adding sequence; the predicted noise-adding sequence includes a first predicted noise-adding sequence and a second predicted noise-adding sequence; the first sample noise-adding sequence and the first predicted noise-adding sequence correspond to a first part of the first object, and the second sample noise-adding sequence and the second predicted noise-adding sequence correspond to the remaining parts of the first object other than the first part; the discrimination value includes a first discrimination value and a second discrimination value. The step of inputting the sample noisy sequence, the predicted noisy sequence, and the sample conditional representation into the discriminator, and having the discriminator output a discrimination value, includes: The first sample noisy sequence, the first predicted noisy sequence, and the sample condition representation are input into the first discriminator, which outputs a first discrimination value. The first discrimination value reflects a first similarity, which is the similarity predicted by the first discriminator for the first sample noisy sequence and the first predicted noisy sequence under the guidance of the sample condition representation. The second sample noisy sequence, the second predicted noisy sequence, and the sample condition representation are input into the second discriminator, which outputs a second discrimination value. The second discrimination value reflects a second similarity, which is the similarity predicted by the second discriminator for the second sample noisy sequence and the second predicted noisy sequence under the guidance of the sample condition representation.
4. The method according to claim 3, characterized in that, The step of calculating the adversarial loss based on the discriminant value includes: Based on the first discriminant value, a first adversarial loss is calculated, which is negatively correlated with the first similarity; and based on the second discriminant value, a second adversarial loss is calculated, which is negatively correlated with the second similarity.
5. The method according to claim 4, characterized in that, The step of training the action diffusion model based on the reconstruction loss and the adversarial loss includes: Based on the predicted action sequence, a physical constraint value is calculated, which reflects the physical naturalness of the action of the first object as represented by the predicted action sequence. Based on the first sample denoised sequence and the first predicted denoised sequence, a first regularization term is calculated, which is used to reflect the difference between the first sample denoised sequence and the first predicted denoised sequence. Based on the second sample denoised sequence and the second predicted denoised sequence, a second regularization term is calculated, which is used to reflect the difference between the second sample denoised sequence and the second predicted denoised sequence. The objective function value is calculated based on the reconstruction loss, the physical constraint value, the first regularization term, the second regularization term, the first adversarial loss, and the second adversarial loss. The objective function value is used to reflect the performance of the action diffusion model in performing the action generation task. Based on the objective function value, the parameters of the action diffusion model are adjusted to obtain the trained action diffusion model.
6. The method according to any one of claims 3 to 5, characterized in that, The first part includes the hand of the first object.
7. The method according to any one of claims 1 to 6, characterized in that, The motion diffusion model includes a first noise denoiser and a second noise denoiser. The sample noise sequence includes a sample root noise sequence and a sample local noise sequence. The sample root noise sequence corresponds to the root node of the first object. The sample local noise sequence corresponds to the body node of the first object. The root node of the first object is used to control the position of the first object in space. The body node of the first object is used to control the body motion of the first object. The step of inputting the sample noise sequence and the sample conditional representation into the action diffusion model, and having the action diffusion model output a predicted action sequence, includes: The sample local noise sequence and the sample condition representation are input into the first denoiser, and the first denoiser outputs a sample local denoising sequence. The sample local denoising sequence is used to characterize the body movements of the first object predicted by the first denoiser under the guidance of the sample condition representation. The sample root noise sequence, the sample local denoising sequence, and the sample conditional representation are input into the second denoiser, and the second denoiser outputs the sample root denoising sequence. The sample root denoising sequence is used to characterize the positional movement of the first object in the space predicted by the second denoiser under the guidance of the sample conditional representation and the sample local denoising sequence. The local denoised sequence of the sample and the root denoised sequence of the sample are merged to obtain the predicted action sequence.
8. The method according to claim 7, characterized in that, The method further includes: Obtain a noise sequence and a conditional representation, the conditional representation being used to guide the trained action diffusion model to generate the action of the second object; The local noise sequence and the conditional representation in the noise sequence are input into the first denoiser, and the first denoiser outputs a local denoised sequence. The local noise sequence corresponds to the body node of the second object, the body node of the second object is used to control the position of the second object in the space, and the local denoised sequence is used to characterize the body action of the second object predicted by the first denoiser under the guidance of the conditional representation. The root noise sequence, the local denoising sequence, and the conditional representation in the noise sequence are input into the second denoiser, and the second denoiser outputs the root denoising sequence. The root noise sequence corresponds to the root node of the second object, and the root node of the second object is used to control the position of the second object in the space. The local denoising sequence is used to characterize the positional movement of the second object in the space predicted by the second denoiser under the guidance of the conditional representation and the local denoising sequence. The local denoised sequence and the root denoised sequence are merged to obtain a generated action sequence, which is used to characterize the action of the second object predicted by the trained action diffusion model under the guidance of the conditional representation.
9. The method according to any one of claims 1 to 8, characterized in that, The first noise includes a first sub-noise at step t, the second noise includes a second sub-noise at step t-1, and the third noise includes a third sub-noise at step t-1. Adding first noise to the sample action sequence to obtain a sample noise sequence includes: Add t steps of the first sub-noise to the sample action sequence to obtain the sample noise sequence, where t is an integer greater than 1; Adding second noise to the sample action sequence to obtain a noisy sample sequence includes: Add the second sub-noise of step t-1 to the sample action sequence to obtain the sample noisy sequence, wherein the variance of the second sub-noise of step i in the second sub-noise of step t-1 is the same as the variance of the first sub-noise of step i in the first sub-noise of step t, and i is a positive integer less than or equal to t-1; Adding a third noise to the predicted action sequence to obtain a predicted noisy sequence includes: Add the third sub-noise of step t-1 to the predicted action sequence to obtain the predicted noisy sequence, wherein the variance of the third sub-noise of step i in step t-1 is the same as the variance of the first sub-noise of step i in step t.
10. The method according to any one of claims 1 to 9, characterized in that, The sample conditions include at least one of the following: Sample audio features are used to characterize the features of the audio corresponding to the sample action sequence; Sample style encoding is used to characterize the emotional style of the sample action sequence; The sample seed sequence includes an action sequence that precedes the sample action sequence and is sequential with the sample action sequence; The sample time step count indicates the number of steps in which noise is added to the sample action sequence during the process of obtaining the sample noise sequence.
11. A training device for an action diffusion model, characterized in that, The device includes: The acquisition module is used to acquire sample action sequences and sample condition representations, wherein the sample action sequences are used to characterize the actions of the first object; The first noise-adding module is used to add first noise to the sample action sequence to obtain a sample noise sequence; A generation module is used to input the sample noise sequence and sample conditional representation into the action diffusion model, and the action diffusion model outputs a predicted action sequence, which is used to characterize the action of the first object predicted by the action diffusion model under the guidance of the sample conditional representation. The second noise-adding module is further configured to add second noise to the sample action sequence to obtain a sample noise-adding sequence, and to add third noise to the predicted action sequence to obtain a predicted noise-adding sequence, wherein the noise intensity of the second noise and the noise intensity of the third noise are both lower than the noise intensity of the first noise. The training module is used to train the action diffusion model with the training objective of increasing the similarity between the sample noisy sequence and the predicted noisy sequence.
12. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program that is loaded and executed by the processor to implement the method as claimed in any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the method as described in any one of claims 1 to 10.
14. A computer program product, characterized in that, The computer program product includes a computer program stored in a computer-readable storage medium, which a processor reads from and executes to implement the method as described in any one of claims 1 to 10.