Human Action Generation Method Based on Spatio-Temporal Attention

By introducing space-time attention mechanisms and discrete cosine transformations into text-driven human motion generation, the problem of insufficient coherence of generated motion time is solved, and more natural and high-quality motion generation is achieved.

CN120014129BActive Publication Date: 2025-06-17JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510496895.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-06-17
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The prior art is difficult to ensure the time consistency of generated movements in text-driven human movement generation, which easily leads to motion jitter and unnatural transitions, affecting the quality of the movement.

Method used

A human body movement generation method based on space-time attention is proposed. By designing a multi-level attention mechanism in the time attention module, efficient self-attention calculation is carried out, and discrete cosine transformation is introduced in the space-time attention module to denoise and improve the quality of motion.

Benefits of technology

It effectively ensures the smoothness and coherence of the generated motion, significantly improves the quality of the motion generation, enhances the robustness of the model, and makes the generated motion more natural and realistic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014129B_ABST
    Figure CN120014129B_ABST
Patent Text Reader

Abstract

The present invention proposes a human motion generation method based on spatio-temporal attention. This method first obtains an original noisy motion sequence and a text description, and fuses the original noisy motion sequence and the text description based on cross-attention to obtain a first fused motion sequence; the first fused motion sequence is sequentially input into a temporal attention module and a spatial attention module for processing to obtain the output features of the spatial attention module; the output features of the spatial attention module are sequentially passed through a multi-layer perceptron and a stylization block and then subjected to a residual connection with the output features of the spatial attention module to obtain a predicted noise; the original noisy motion sequence is denoised in the reverse direction according to the predicted noise to reconstruct the motion sequence. In order to capture local-to-global long-term temporal dependencies, the present invention designs a multi-level attention mechanism in the temporal attention module to model the correlation between consecutive frames in the motion sequence, effectively ensuring the smoothness and coherence of the generated motion in terms of time series.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer text-driven motion generation, and particularly to a human motion generation method based on spatio-temporal attention. Background Art

[0002] Motion generation plays a crucial role in various applications such as games, movie production, and robot control. In recent years, deep learning-based methods have become increasingly common in this field, making it possible to create more diverse and delicate human motions. An area that has received increasing attention is the generation of realistic and dynamic human motion sequences through natural language descriptions, thus bridging the gap between text input and expressive motions.

[0003] Previous studies have explored the use of deep learning techniques for text-driven human motion generation. However, generating realistic motion sequences from text remains a challenging task. The main difficulties lie in how to ensure that the generated motions not only conform to natural human motion patterns but also maintain temporal continuity while accurately reflecting the attributes of the motions. Recently, some studies have utilized diffusion models to generate motions consistent with text inputs. Others have adopted the cross-attention mechanism of Transformer to fuse text features and thus generate high-quality motions consistent with the text. Although these methods have shown potential to some extent, they still face the problem of insufficient temporal coherence, which easily leads to artifacts such as motion jitter and unnatural transitions, thus affecting the overall motion quality. Summary of the Invention

[0004] In view of the above situation, the main object of the present invention is to propose a human motion generation method based on spatio-temporal attention to solve the above technical problems.

[0005] The present invention proposes a human motion generation method based on spatio-temporal attention, and the method includes the following steps:

[0006] Step 1, obtain an original noisy motion sequence and a text description, and fuse the original noisy motion sequence and the text description based on cross-attention to obtain a first fused motion sequence;

[0007] Step 2, input the first fused motion sequence into a temporal attention module, divide the first fused motion sequence into several segments, perform efficient self-attention calculations on each segment respectively, and fuse the efficient self-attention calculation results of each segment to obtain a first efficient self-attention output;

[0008] Step 3: Perform efficient self-attention calculation on the first fused motion sequence to obtain a second efficient self-attention output, and fuse the first efficient self-attention output and the second efficient self-attention output to obtain the output features of the temporal attention module;

[0009] Step 4: Repeat Steps 2 to 3 iteratively several times, and decrease the number of divided segments according to the iteration number during each iteration to obtain the temporal attention output for each iteration. Then fuse the temporal attention outputs obtained in each iteration to obtain the final output features of the temporal attention module;

[0010] Step 5: Pass the final output features of the temporal attention module through a stylization block and then perform a residual connection with the first fused motion sequence to obtain a second fused motion sequence;

[0011] Step 6: Input the second fused motion sequence into the spatial attention module to extract frequency domain information and obtain the output features of the spatial attention module;

[0012] Step 7: Pass the output features of the spatial attention module through a multi-layer perceptron and a stylization block in sequence and then perform a residual connection with the output features of the spatial attention module to obtain the predicted noise;

[0013] Step 8: Perform reverse denoising on the original noisy motion sequence according to the predicted noise to reconstruct the action sequence.

[0014] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0015] 1. The present invention designs a multi-level attention mechanism in the temporal attention module. By inputting the first fused motion sequence into the temporal attention module, the first fused motion sequence is divided into several segments, and efficient self-attention calculation is performed on each segment respectively. Then, the results of the efficient self-attention calculation for each segment are fused as the input for the next stage for iterative calculation to model the temporal dependence with a long time span, effectively ensuring the smoothness and coherence of the generated motion in time series.

[0016] 2. The present invention designs a smoothing step. By performing efficient self-attention calculation on the first fused motion sequence to obtain a second efficient self-attention output, and fusing the first efficient self-attention output and the second efficient self-attention output, seamless integration of features is achieved, thereby avoiding the problem that segmentation and re-splicing of segments may introduce jitter at the connection points.

[0017] 3. The present invention introduces the discrete cosine transform in the spatial attention module, which can eliminate the noise interference introduced in the time modeling process, not only significantly improving the quality of motion generation, but also enhancing the robustness performance of the model. The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the embodiments of the present invention. Description of the Drawings

[0018] Figure 1 It is a flowchart of the human action generation method based on spatio-temporal attention proposed by the present invention;

[0019] Figure 2 It is an architecture diagram of the human action generation method based on spatio-temporal attention proposed by the present invention;

[0020] Figure 3 It is an architecture diagram of the efficient attention calculation of the present invention;

[0021] Figure 4 It is an architecture diagram of the time attention module of the present invention;

[0022] Figure 5 It is an architecture diagram of the spatial attention module of the present invention;

[0023] Figure 6 It is a text-based action result diagram generated by the present invention;

[0024] Figure 7 It is a comparison diagram of the results of the present invention and MDM on the HumanML3D dataset. Detailed Embodiments

[0025] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0026] Referring to the following description and drawings, these and other aspects of the embodiments of the present invention will be clear. In these descriptions and drawings, some specific embodiments of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0027] Please refer to Figure 1 and Figure 2 , this embodiment provides a human action generation method based on spatio-temporal attention, and the method includes the following steps:

[0028] Step 1: Obtain the original noisy motion sequence and the text description, and fuse the original noisy motion sequence and the text description based on cross-attention to obtain the first fused motion sequence;

[0029] As a preferred embodiment of the present invention, fusing the original noisy motion sequence and the text description based on cross-attention to obtain the first fused motion sequence specifically includes the following steps:

[0030] Use a pre-trained CLIP text encoder to extract the token embeddings of the text description;

[0031] Apply positional embeddings to the token embeddings according to the position information of the token embeddings in the text description to obtain tokens containing position information;

[0032] Pass the tokens containing position information through a Transformer to obtain text features;

[0033] Perform cross-attention calculation on the text features and the original noisy motion sequence, and perform residual connection on the cross-attention calculation result with the original noisy motion sequence through a stylized block to obtain the first fused motion sequence.

[0034] Please refer to Figure 3 and Figure 4 , in the figure, represents matrix dot product, represents the Softmax activation function.

[0035] Step 2: Input the first fused motion sequence into the temporal attention module, divide the first fused motion sequence into several segments, and perform efficient self-attention calculation on each segment respectively, and fuse the efficient self-attention calculation results of each segment to obtain the first efficient self-attention output;

[0036] As a preferred embodiment of the present invention, the process of dividing the first fused motion sequence into several segments has the following relational formula:

[0037] ;

[0038] Among them, represents the first efficient self-attention output, represents the continuous features obtained by passing the divided motion sequence segments through the efficient attention module, represents the splicing operation of the motion sequence segments.

[0039] As a preferred embodiment of the present invention, performing efficient self-attention calculation on the segments specifically includes the following steps:

[0040] Use the segment to generate query vectors, key vectors, and value vectors;

[0041] Generate a global feature map based on the key vector and the value vector. The corresponding process has the following relational expression:

[0042] ;

[0043] Among them, represents the global feature map, , represents a real number, represents the dimension size of the feature, represents the key vector, represents the value vector, represents the normalization exponential function, represents the transpose operation, represents the matrix dot product;

[0044] Calculate the efficient self-attention output feature of the segment according to the global feature map and the query vector. The corresponding process has the following relational expression:

[0045] ;

[0046] Among them, represents the query vector, represents the efficient self-attention output feature of the segment.

[0047] Step 3: Perform efficient self-attention calculation on the first fused motion sequence to obtain the second efficient self-attention output, and fuse the first efficient self-attention output and the second efficient self-attention output to obtain the output feature of the temporal attention module;

[0048] Step 4: Repeat Step 2 to Step 3 several times in an iterative manner, and decrease the number of divided segments according to the number of iterations in each iteration process to obtain the temporal attention output of each iteration. Fuse the temporal attention outputs obtained in each iteration to obtain the final output feature of the temporal attention module;

[0049] As a preferred embodiment of the present invention, fusing the temporal attention outputs obtained in each iteration to obtain the final output feature of the temporal attention module. The corresponding process has the following relational expression:

[0050] ;

[0051] Among them, represents the final output feature of the temporal attention module, represents the second efficient self-attention output, represents the learnable weight, represents the first efficient self-attention output of the Mth iteration.

[0052] Step 5: Residually connect the output features of the final temporal attention module through the stylization block with the first fused motion sequence to obtain the second fused motion sequence;

[0053] Please refer to Figure 5 , Step 6: Input the second fused motion sequence into the spatial attention module for frequency domain information extraction to obtain the output features of the spatial attention module;

[0054] As a preferred embodiment of the present invention, input the second fused motion sequence into the spatial attention module for frequency domain information extraction to obtain Specifically, it includes the following steps:

[0055] Generate Z from the second fused motion sequence through the self-attention mechanism. The corresponding process has the following relationship:

[0056] ;

[0057] Among them, represents the feature representation that fuses global context information, represents the second fused motion sequence;

[0058] Apply the discrete cosine transform, multi-layer perceptron, and inverse discrete cosine transform to the feature representation that fuses global context information in sequence to obtain the low-frequency features. The corresponding process has the following relationship:

[0059] ;

[0060] Among them, represents the low-frequency features obtained by applying the discrete cosine transform to the feature representation that fuses global context information, represents the inverse discrete cosine transform, represents the multi-layer perceptron, represents the discrete cosine transform, represents layer normalization;

[0061] Apply the multi-layer perceptron to Z and fuse it with to obtain , and the corresponding process has the following relationship:

[0062] ;

[0063] Among them, represents the feature representation that fuses low-frequency features and original features;

[0064] Pass through the convolutional layer to obtain , and the corresponding process has the following relationship:

[0065] ;

[0066] Among them, represents the output feature of the spatial attention module.

[0067] Step 7: Pass the output feature of the spatial attention module through a multi-layer perceptron and a stylization block in sequence, and then perform a residual connection with the output feature of the spatial attention module to obtain the predicted noise;

[0068] Step 8: Perform reverse denoising on the original noisy motion sequence according to the predicted noise to reconstruct the action sequence.

[0069] As a preferred embodiment of the present invention, when performing reverse denoising on the original noisy motion sequence according to the predicted noise to reconstruct the action sequence, the corresponding process has the following relational formula:

[0070] ;

[0071] Among them, represents one step in the reverse diffusion process, represents the data distribution at time t, represents the timestamp, represents the Gaussian distribution, represents the predicted mean, represents the weight of the model.

[0072] As a preferred embodiment of the present invention, the calculation process of the stylization module has the following relational formula:

[0073] ;

[0074] Among them, represents the output of the stylization module, represents the input of the stylization module, represents the multiplicative offset of the original feature, represents the additive offset of the original feature.

[0075] To verify the effectiveness of the present invention, this embodiment evaluates the present invention on two widely used benchmark datasets: the HumanML3D and KIT-ML (KIT-Motion Language) datasets. HumanML3D contains the HumanAct12 and AMASS datasets, with a total of 14,616 actions and 44,970 text descriptions (including 5,371 unique words). The KIT-ML dataset provides 3,911 action sequences, accompanied by 6,353 sequence-level natural language descriptions.

[0076] The following metrics are used to evaluate the performance of the present invention:

[0077] R-precision (R-accuracy): For each generated text-action pair, 31 mismatched descriptions are randomly selected from the test set. Then, the Euclidean distances between the action features and all 32 candidate descriptions are calculated and they are ranked to obtain the top-k precision.

[0078] Frechet Inception Distance (FID): This metric measures the distribution similarity between the features extracted from the generated actions and the real actions.

[0079] Multi-Modal Distance: This is the distance between the text features and the corresponding generated action features, calculated for each given description.

[0080] Diversity: The overall variability of the generated actions is quantified by calculating the average pairwise Euclidean distance between randomly partitioned groups of actions among all descriptions.

[0081] Multimodality: For a single text description, 32 action sequences are generated and the similarity between them is calculated.

[0082] In this experiment, R-percision (R-accuracy) and FID (Frechet Inception Distance) are emphasized as the main performance metrics because they are key metrics for evaluating the overall quality of the generated actions. A diffusion model is adopted, using 1000 diffusion steps, and the noise variance increases linearly, ranging from 0.0001 to 0.002. In terms of optimization, the Adam optimizer is used and the learning rate is fixed at , with a batch size of 32. The training is conducted for 80K iterations on the HumanML3D dataset and 40K iterations on the KIT-ML dataset.

[0083] The results of the simulation show that: The present invention can generate text-driven human actions with diversity, including real movements of the arms, legs and body (such as walking, running, kicking, jumping, dancing). Example results are as Figure 6 shown, indicating that the present invention can generate high-quality and realistic actions. These simulation results show that the present invention has significant advantages and great potential in generating real, smooth and text-consistent human actions.

[0084] The present invention is also compared with MDM as a baseline. As Figure 7As shown on the left side, compared with the MDM method, the results of the present invention have significant advantages in multiple aspects: the generated dance movements show richer arm coordination, more rhythmic steps, and smoother long-distance movements, demonstrating higher accuracy. On the right side of Figure 6, the present invention can effectively depict the movement of a person standing up from a sitting position, while the results of the MDM method show that the person either fails to meet the movement requirements of standing up or only stands up partially.

[0085] In addition, this embodiment also evaluated the two methods (i.e., the present invention and the MDM method) in the treadmill running scenario. The present invention enables the runner to stay in the center of the treadmill, while the results of the MDM method show that the runner almost falls off the treadmill.

[0086] Based on the above results, through qualitative visual demonstrations and effect comparisons, it shows that the present invention has significant advantages in terms of diversity, realism, and visual expressiveness, and can generate diverse motion sequences that are highly consistent with the text description. These advantages make it highly competitive in text-driven motion generation.

[0087] In this embodiment, the present invention was also compared with several basic methods, including Seq2Seq, Language2Pose, TM2T, Text2Gesture, Guo24, MDM, MotionDiffuse, and TEMOS, as shown in Tables 1, 2, and 3. The present invention is superior to these basic methods in terms of FID (Frechet Inception Distance), and shows competitive or better performance in indicators such as R-Precision, MultiModal Distance, Diversity, and Multimodality, indicating that MotionSTA can generate high-quality motion sequences.

[0088] Among them, Table 1 shows the comparison between the present invention and some basic methods on the HumanML3D test set. All methods use the true length of the benchmark ground truth. Table 2 shows the comparison between the present invention and some basic methods in terms of FID and MultiModality on the HumanML3D test set. All methods use the true length of the benchmark ground truth. Table 3 shows the comparison between the present invention and some basic methods on the KIT-ML test set. All methods use the true length of the benchmark ground truth.

[0089] Table 1

[0090]

[0091] Table 2

[0092]

[0093] Table 3

[0094]

[0095] This embodiment further studied the impact of the latent dimension size and the number of layers on performance, namely the ablation study on the latent dim and the number of Transformer layers. All results are reported based on the HumanML3D test set. As can be seen from Table 4, when the latent dimension is 256, the performance is still not ideal regardless of the number of layers. When the number of layers is 4, increasing the latent dimension from 512 to 1024 brings only a negligible improvement in performance. However, in the case of 8 layers, the performance is optimal when the latent dimension is 512.

[0096] Table 4

[0097]

[0098] To demonstrate the importance of the multi-level self-attention design of the present invention and the spatial attention module combined with the discrete cosine transform, this embodiment conducted an ablation experiment, namely the ablation study on the spatial attention module and the temporal attention module. All results are reported based on the HumanML3D test set. As shown in Table 5, combining the temporal attention module slightly improves the motion quality compared to omitting the temporal attention module, demonstrating the advantages of the discrete cosine transform and the convolutional network in capturing motion features. It is worth noting that combining the spatial attention module and omitting the spatial attention module brings more significant improvements. A reasonable explanation is that the discrete cosine transform process in the SA module retains low-frequency features and filters out noise, thus improving the motion generation performance.

[0099] Table 5

[0100]

[0101] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown sequentially according to the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages does not necessarily have to be sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0102] It should be understood that each part of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0103] In the description of this specification, the description referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0104] The above-described embodiments merely represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed. However, it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.

Claims

1. A human motion generation method based on spatiotemporal attention, characterized in that: The method comprises the following steps: Step 1: obtaining an original noisy motion sequence and a text description, and fusing the original noisy motion sequence and the text description based on cross attention to obtain a first fused motion sequence; Step 2: Input the first fused motion sequence into the temporal attention module, divide the first fused motion sequence into several segments, perform efficient self-attention calculation on each segment, fuse the efficient self-attention calculation results of each segment, and obtain the first efficient self-attention output; Step 3: Perform efficient self-attention calculation on the first fused motion sequence to obtain a second efficient self-attention output, and fuse the first efficient self-attention output with the second efficient self-attention output to obtain the output feature of the temporal attention module; Step 4: Repeat steps 2 to 3 several times in an iterative manner, and reduce the number of divided segments according to the number of iterations in each iteration to obtain the temporal attention output of each iteration, and fuse the temporal attention outputs obtained in each iteration to obtain the final temporal attention module output features; Step 5: The final output feature of the temporal attention module is residually connected with the first fused motion sequence through the stylized block to obtain the second fused motion sequence; Step 6: Input the second fused motion sequence into the spatial attention module to extract frequency domain information, and obtain the output features of the spatial attention module; Step 7: The output features of the spatial attention module are sequentially passed through the multi-layer perceptron and the stylization block, and then residually connected with the output features of the spatial attention module to obtain the predicted noise; Step 8: Perform reverse denoising on the original noisy motion sequence according to the predicted noise to reconstruct the motion sequence.

2. The method for generating human motion based on spatiotemporal attention according to claim 1, characterized in that: In step 1, the original noisy motion sequence and the text description are fused based on cross attention to obtain a first fused motion sequence, which specifically includes the following steps: Use the pre-trained CLIP text encoder to extract word embeddings of text descriptions; Apply position embedding to the word-meta embedding according to the position information of the word-meta embedding in the text description to obtain a word-meta containing the position information; The word containing the position information is passed through Transformer to obtain text features; The text features and the original noisy motion sequence are cross-attention calculated, and the cross-attention calculation results are residually connected with the original noisy motion sequence through the stylized block to obtain the first fused motion sequence.

3. The method for generating human motion based on spatiotemporal attention according to claim 2, characterized in that: In step 2, the efficient self-attention calculation results of each segment are fused to obtain the process corresponding to the first efficient self-attention output, which has the following relationship: ; in, represents the first efficient self-attention output, It represents the continuity features of the divided motion sequence fragments obtained by the efficient attention module. Represents the splicing operation on motion sequence fragments.

4. The method for generating human motion based on spatiotemporal attention according to claim 3, characterized in that: In step 2, performing efficient self-attention calculation on the fragment specifically includes the following steps: Generate query vector, key vector and value vector using fragments; The global feature map is generated based on the key vector and the value vector. The corresponding process has the following relationship: ; in, represents the global feature map, , represents a real number, Indicates the dimension size of the feature, represents the key vector, represents a value vector, represents the normalized exponential function, represents the transpose operation, Represents matrix dot product; The efficient self-attention output feature of the fragment is calculated based on the global feature map and the query vector. The corresponding process has the following relationship: ; in, represents the query vector, Efficient self-attention output features for representing fragments.

5. The method for generating human motion based on spatiotemporal attention according to claim 4, characterized in that: In step 4, the temporal attention output obtained in each iteration is fused to obtain the final temporal attention module output feature. The corresponding process has the following relationship: ; in, represents the final temporal attention module output feature, represents the second most efficient self-attention output, represents the learnable weights, represents the first efficient self-attention output of the Mth iteration.

6. The method for generating human motion based on spatiotemporal attention according to claim 5, characterized in that: In step 6, the second fused motion sequence is input into the spatial attention module to extract frequency domain information, and the The specific steps include: The second fused motion sequence is generated through the self-attention mechanism to generate Z. The corresponding process has the following relationship: ; in, represents the feature representation that integrates global context information, represents the second fused motion sequence; The feature representation of the fusion of global context information is sequentially applied with discrete cosine transform, multi-layer perceptron and inverse discrete cosine transform to obtain low-frequency features. The corresponding process has the following relationship: ; in, represents the low-frequency features obtained by performing discrete cosine transform on the feature representation that integrates the global context information. represents the inverse discrete cosine transform, represents a multi-layer perceptron, represents discrete cosine transform, Representation layer normalization; Apply Z to a multilayer perceptron and combine Fusion, get , the corresponding process has the following relationship: ; in, It represents the feature representation that combines low-frequency features and original features; Will Through the convolution layer, we get , the corresponding process has the following relationship: ; in, Represents the output features of the spatial attention module.

7. The method for generating human motion based on spatiotemporal attention according to claim 6, characterized in that: In step 8, the original noisy motion sequence is reversely denoised according to the predicted noise to reconstruct the motion sequence. The corresponding process has the following relationship: ; in, represents the reverse diffusion process, represents the data distribution at time t, Indicates the timestamp, represents a Gaussian distribution, represents the predicted mean, Represents the parameters of the model.

8. The method for generating human motion based on spatiotemporal attention according to claim 7, characterized in that: The calculation process of the stylized module has the following relationship: ; in, represents the output of the stylization module, represents the input of the stylization module, represents the multiplicative offset of the original feature, Represents an additive offset to the original feature.

Citation Information

Patent Citations

  • Time sequence prediction method and system based on diffusion autoregression transformer

    CN119622248A

  • Method and server for training a neural network to generate a textual output sequence

    US20220108685A1