Human body action generation method based on space-time attention

By introducing multi-level attention mechanism based on space-time attention and discrete cosine transformation in text-driven human motion generation, the problems of insufficient motion jitter and time coherence in the prior art are solved, and high-quality and robust motion generation are achieved.

CN120014129AActive Publication Date: 2025-05-16JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS

Patent Information

Application Number
CN202510496895.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-16
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

In the text-driven human movement generation, it is difficult to ensure that the generated movement is not only in line with the natural human movement mode, but also maintains time continuity, resulting in artifacts such as motion jitter and unnatural transitions, affecting the overall movement quality.

Method used

A human body movement generation method based on space-time attention is proposed. By designing a multi-level attention mechanism in the time attention module, combining efficient self-attention calculation and iterative calculation, time dependence is ensured; discrete cosine transformation is introduced in the space-attention module to eliminate noise interference and improve movement quality.

Benefits of technology

It effectively ensures the smoothness and coherence of the generated motion, avoids motion jitter and unnatural transitions, and significantly improves the quality of motion generation and the robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014129A_ABST
    Figure CN120014129A_ABST
Patent Text Reader

Abstract

The invention provides a human motion generation method based on space-time attention, and the method comprises the steps: firstly obtaining an original noisy motion sequence and text description, and carrying out the fusion of the original noisy motion sequence and the text description based on cross attention, and obtaining a first fusion motion sequence; sequentially inputting the first fusion motion sequence into a time attention module and a space attention module for processing to obtain output features of the space attention module; enabling the output features of the space attention module to pass through a multi-layer perceptron and a stylized block in sequence, and then carrying out the residual connection with the output features of the space attention module, and obtaining prediction noise; and performing reverse denoising on the original motion sequence with noise according to the predicted noise so as to reconstruct an action sequence. In order to capture local-to-global long-term time dependence, a multi-level attention mechanism is designed in a time attention module to model correlation between continuous frames in a motion sequence, and smoothness and coherence of generated motion in time sequence are effectively guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer text-driven motion generation, and in particular to a human motion generation method based on spatiotemporal attention. Background Art

[0002] Motion generation plays a crucial role in a variety of applications such as gaming, filmmaking, and robotic control. In recent years, deep learning-based methods have become increasingly prevalent in this field, making it possible to create more diverse and detailed human motion. One area of ​​growing interest is the generation of realistic and dynamic human motion sequences from natural language descriptions, thus bridging the gap between textual input and expressive motion.

[0003] Previous studies have explored text-driven human motion generation using deep learning techniques. However, generating realistic motion sequences from text remains a challenging task. The main difficulty lies in how to ensure that the generated motion is consistent with natural human motion patterns, maintains temporal continuity, and accurately reflects the properties of the motion. Recently, some studies have used diffusion models to generate motion consistent with text input. Other studies have adopted the cross-attention mechanism of Transformer to fuse text features to generate high-quality and text-consistent motion. Although these methods have shown potential to a certain extent, they still face the problem of insufficient temporal coherence, which can easily lead to artifacts such as motion jitter and unnatural transitions, thus affecting the overall motion quality. Summary of the invention

[0004] In view of the above situation, the main purpose of the present invention is to propose a human motion generation method based on spatiotemporal attention to solve the above technical problems.

[0005] The present invention proposes a method for generating human motion based on spatiotemporal attention, the method comprising the following steps: Step 1: obtaining an original noisy motion sequence and a text description, and fusing the original noisy motion sequence and the text description based on cross attention to obtain a first fused motion sequence; Step 2: Input the first fused motion sequence into the temporal attention module, divide the first fused motion sequence into several segments, perform efficient self-attention calculation on each segment, fuse the efficient self-attention calculation results of each segment, and obtain the first efficient self-attention output; Step 3: Perform efficient self-attention calculation on the first fused motion sequence to obtain a second efficient self-attention output, and fuse the first efficient self-attention output with the second efficient self-attention output to obtain the output feature of the temporal attention module; Step 4: Repeat steps 2 to 3 several times in an iterative manner, and reduce the number of divided segments according to the number of iterations in each iteration to obtain the temporal attention output of each iteration, and fuse the temporal attention outputs obtained in each iteration to obtain the final temporal attention module output features; Step 5: The final output feature of the temporal attention module is residually connected with the first fused motion sequence through the stylized block to obtain the second fused motion sequence; Step 6: Input the second fused motion sequence into the spatial attention module to extract frequency domain information, and obtain the output features of the spatial attention module; Step 7: The output features of the spatial attention module are sequentially passed through the multi-layer perceptron and the stylization block, and then residually connected with the output features of the spatial attention module to obtain the predicted noise; Step 8: Perform reverse denoising on the original noisy motion sequence according to the predicted noise to reconstruct the motion sequence.

[0006] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention designs a multi-level attention mechanism in the temporal attention module. By inputting the first fused motion sequence into the temporal attention module, the first fused motion sequence is divided into several segments, and efficient self-attention calculation is performed on each segment respectively. The efficient self-attention calculation results of each segment are then fused as the input of the next stage for iterative calculation to model the time dependency of a long time span, effectively ensuring the smoothness and coherence of the generated motion in timing.

[0007] 2. The present invention designs a smoothing step, which performs efficient self-attention calculation on the first fused motion sequence to obtain a second efficient self-attention output, and fuses the first efficient self-attention output with the second efficient self-attention output to achieve seamless integration of features, thereby avoiding the problem that segmentation and re-splicing of fragments may introduce jitter at the connection points.

[0008] 3. The present invention introduces discrete cosine transform in the spatial attention module, which can eliminate the noise interference introduced in the temporal modeling process, not only significantly improves the quality of motion generation, but also enhances the robustness of the model. Additional aspects and advantages of the present invention will be given in part in the following description, and in part will become apparent from the following description, or will be understood through the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 A flow chart of a human motion generation method based on spatiotemporal attention proposed by the present invention; Figure 2 This is an architecture diagram of the human motion generation method based on spatiotemporal attention proposed by the present invention; Figure 3 The architecture diagram of the efficient attention calculation of the present invention; Figure 4 This is an architecture diagram of the temporal attention module of the present invention; Figure 5 Schematic diagram of the spatial attention module of the present invention; Figure 6 A text-based action-result graph generated for the present invention; Figure 7 A comparison chart of the results of the present invention and MDM on the HumanML3D dataset. DETAILED DESCRIPTION

[0010] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0011] These and other aspects of the embodiments of the present invention will be apparent with reference to the following description and accompanying drawings. In these descriptions and accompanying drawings, some specific implementations of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0012] See also Figure 1 and Figure 2 This embodiment provides a method for generating human motion based on spatiotemporal attention, the method comprising the following steps: Step 1: obtaining an original noisy motion sequence and a text description, and fusing the original noisy motion sequence and the text description based on cross attention to obtain a first fused motion sequence; As a preferred embodiment of the present invention, the original noisy motion sequence and the text description are fused based on cross attention to obtain a first fused motion sequence, which specifically includes the following steps: Use the pre-trained CLIP text encoder to extract word embeddings of text descriptions; Apply position embedding to the word-meta embedding according to the position information of the word-meta embedding in the text description to obtain a word-meta containing the position information; The word containing the position information is passed through Transformer to obtain text features; The text features and the original noisy motion sequence are cross-attention calculated, and the cross-attention calculation results are residually connected with the original noisy motion sequence through the stylized block to obtain the first fused motion sequence.

[0013] See also Figure 3 and Figure 4 , in the figure, represents matrix dot product, Represents the Softmax activation function.

[0014] Step 2: Input the first fused motion sequence into the temporal attention module, divide the first fused motion sequence into several segments, perform efficient self-attention calculation on each segment, fuse the efficient self-attention calculation results of each segment, and obtain the first efficient self-attention output; As a preferred embodiment of the present invention, the process of dividing the first fused motion sequence into several segments corresponds to the following relationship: ; in, represents the first efficient self-attention output, It represents the continuity features of the divided motion sequence fragments obtained by the efficient attention module. Represents the splicing operation on motion sequence fragments.

[0015] As a preferred embodiment of the present invention, performing efficient self-attention calculation on a fragment specifically includes the following steps: Generate query vector, key vector and value vector using fragments; The global feature map is generated based on the key vector and the value vector. The corresponding process has the following relationship: ; in, represents the global feature map, , represents a real number, Indicates the dimension size of the feature, represents the key vector, represents a value vector, represents the normalized exponential function, represents the transpose operation, Represents matrix dot product; The efficient self-attention output feature of the fragment is calculated based on the global feature map and the query vector. The corresponding process has the following relationship: ; in, represents the query vector, Efficient self-attention output features to represent fragments.

[0016] Step 3: Perform efficient self-attention calculation on the first fused motion sequence to obtain a second efficient self-attention output, and fuse the first efficient self-attention output with the second efficient self-attention output to obtain the output feature of the temporal attention module; Step 4: Repeat steps 2 to 3 several times in an iterative manner, and reduce the number of divided segments according to the number of iterations in each iteration to obtain the temporal attention output of each iteration, and fuse the temporal attention outputs obtained in each iteration to obtain the final temporal attention module output features; As a preferred embodiment of the present invention, the temporal attention output obtained in each iteration is fused to obtain the final temporal attention module output feature. The corresponding process has the following relationship: ; in, represents the final temporal attention module output feature, represents the second most efficient self-attention output, represents the learnable weights, represents the first efficient self-attention output of the Mth iteration.

[0017] Step 5: The final output feature of the temporal attention module is residually connected with the first fused motion sequence through the stylized block to obtain the second fused motion sequence; See also Figure 5 , step 6, input the second fused motion sequence into the spatial attention module to extract frequency domain information, and obtain the output features of the spatial attention module; As a preferred embodiment of the present invention, the second fused motion sequence is input into the spatial attention module to extract frequency domain information, and the obtained The specific steps include: The second fused motion sequence is generated through the self-attention mechanism to generate Z. The corresponding process has the following relationship: ; in, represents the feature representation that integrates global context information, represents the second fused motion sequence; The feature representation of the fusion of global context information is sequentially applied with discrete cosine transform, multi-layer perceptron and inverse discrete cosine transform to obtain low-frequency features. The corresponding process has the following relationship: ; in, represents the low-frequency features obtained by performing discrete cosine transform on the feature representation that integrates the global context information. represents the inverse discrete cosine transform, represents a multi-layer perceptron, represents discrete cosine transform, Representation layer normalization; Apply Z to a multilayer perceptron and combine Fusion, get , the corresponding process has the following relationship: ; in, It represents the feature representation that combines low-frequency features and original features; Will Through the convolution layer, we get , the corresponding process has the following relationship: ; in, Represents the output features of the spatial attention module.

[0018] Step 7: The output features of the spatial attention module are sequentially passed through the multi-layer perceptron and the stylization block, and then residually connected with the output features of the spatial attention module to obtain the predicted noise; Step 8: Perform reverse denoising on the original noisy motion sequence according to the predicted noise to reconstruct the motion sequence.

[0019] As a preferred embodiment of the present invention, the original noisy motion sequence is reversely denoised according to the predicted noise to reconstruct the motion sequence. The corresponding process has the following relationship: ; in, represents a step in the reverse diffusion process, represents the data distribution at time t, Indicates the timestamp, represents a Gaussian distribution, represents the predicted mean, Represents the weight of the model.

[0020] As a preferred embodiment of the present invention, the calculation process of the stylization module has the following relationship: ; in, represents the output of the stylization module, represents the input of the stylization module, represents the multiplicative offset of the original feature, Represents an additive offset to the original feature.

[0021] In order to verify the effectiveness of the present invention, this embodiment evaluates the present invention on two widely used benchmark datasets: HumanML3D and KIT-ML (KIT-Motion Language) datasets. HumanML3D contains HumanAct12 and AMASS datasets, with a total of 14,616 actions and 44,970 text descriptions (including 5,371 unique words). The KIT-ML dataset provides 3,911 action sequences and is equipped with 6,353 sequence-level natural language descriptions.

[0022] The following indicators are used to evaluate the performance of the present invention: R-precision: For each generated text-action pair, 31 unmatched descriptions are randomly selected from the test set. Then, the Euclidean distance between the action feature and all 32 candidate descriptions is calculated and ranked to obtain the top-k precision.

[0023] Frechet Inception Distance (FID): This metric measures the distribution similarity between features extracted from generated and real actions.

[0024] Multi-Modal Distance: This is the distance between text features and the corresponding generated action features, calculated for each given description.

[0025] Diversity: Quantifies the overall variability of generated actions by computing the average pairwise Euclidean distance between randomly partitioned groups of actions across all descriptions.

[0026] Multimodality: For a single text description, generate 32 action sequences and calculate the similarity between them.

[0027] In this experiment, we emphasize R-percision and FID (Frechet Inception Distance) as the main performance indicators, because they are key indicators for evaluating the overall quality of generated actions. A diffusion model is adopted, using 1000 diffusion steps and the noise variance It increases linearly from 0.0001 to 0.002. In terms of optimization, the Adam optimizer is used, and the learning rate is fixed at , the batch size is 32. The training is performed for 80K iterations on the HumanML3D dataset and 40K iterations on the KIT-ML dataset.

[0028] The simulation results show that the present invention can generate diverse text-driven human actions, including real movements of arms, legs and body (such as walking, running, kicking, jumping, dancing). Example results are as follows Figure 6 As shown, it is shown that the present invention can generate high-quality and realistic movements. These simulation results show that the present invention has significant advantages and great potential in generating realistic, smooth and text-consistent human body movements.

[0029] The present invention was also compared with MDM as a baseline. Figure 7 As shown on the left, the results of the present invention have significant advantages over the MDM method in many aspects: the generated dance movements show richer arm coordination, more rhythmic steps, and smoother long-distance movement, showing higher accuracy. On the right side of Figure 6, the present invention can effectively depict the action of a person standing up from a sitting position, while the results of the MDM method show that the character either fails to complete the action requirements of standing up or only stands up partially.

[0030] In addition, this embodiment also evaluates two methods (i.e., the present invention and the MDM method) in a treadmill running scenario. The present invention is able to keep the runner in the center of the treadmill, while the results of the MDM method show that the runner almost falls off the treadmill.

[0031] In summary, the qualitative visual display and effect comparison show that the present invention has significant advantages in diversity, realism and visual expression, and can generate diversified motion sequences that are highly consistent with the text description. These advantages make it highly competitive in text-driven motion generation.

[0032] In this embodiment, the present invention is also compared with several basic methods, including Seq2Seq, Language2Pose, TM2T, Text2Gesture, Guo24, MDM, MotionDiffuse and TEMOS, as shown in Tables 1, 2 and 3. The present invention outperforms these basic methods in terms of FID (Frechet Inception Distance) and shows competitive or better performance in indicators such as R-Precision, MultiModal Distance, Diversity and Multimodality, indicating that MotionSTA can generate high-quality motion sequences.

[0033] Among them, Table 1 is a comparison between the present invention and some basic methods on the HumanML3D test set. All methods use the true length of the benchmark true value. Table 2 is a comparison between the present invention and some basic methods on the HumanML3D test set in FID and MultiModality. All methods use the true length of the benchmark true value. Table 3 is a comparison between the present invention and some basic methods on the KIT-ML test set. All methods use the true length of the benchmark true value.

[0034] Table 1

[0035] Table 2

[0036] Table 3

[0037] This example further studies the impact of latent dimension size and number of layers on performance, and performs an ablation study on latent dim and number of Transformer layers. All results are reported based on the HumanML3D test set. As can be seen from Table 4, when the latent dimension is 256, the performance is still not ideal regardless of the number of layers. When the number of layers is 4, increasing the latent dimension from 512 to 1024 brings little performance improvement. However, in the case of 8 layers, the performance is best when the latent dimension is 512.

[0038] Table 4

[0039] In order to reflect the importance of the multi-level self-attention design of the present invention and the spatial attention module combined with discrete cosine transform, this embodiment conducted an ablation experiment to study the ablation of the spatial attention module and the temporal attention module. All results are reported based on the HumanML3D test set. As shown in Table 5, combining the temporal attention module slightly improves the motion quality compared to omitting the temporal attention module, proving the advantages of discrete cosine transform and convolutional networks in capturing motion features. It is worth noting that combining the spatial attention module and omitting the spatial attention module bring more significant improvements. A reasonable explanation is that the discrete cosine transform process in the SA module retains low-frequency features and filters noise, thereby improving the motion generation performance.

[0040] Table 5

[0041] It should be understood that, although each step in the flow chart of each embodiment of the present invention is shown in sequence according to the indication of the arrow, these steps are not necessarily performed in sequence according to the order indicated by the arrow. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.

[0042] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0043] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0044] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A human motion generation method based on spatiotemporal attention, characterized in that: The method comprises the following steps: Step 1: obtaining an original noisy motion sequence and a text description, and fusing the original noisy motion sequence and the text description based on cross attention to obtain a first fused motion sequence; Step 2: Input the first fused motion sequence into the temporal attention module, divide the first fused motion sequence into several segments, perform efficient self-attention calculation on each segment, fuse the efficient self-attention calculation results of each segment, and obtain the first efficient self-attention output; Step 3: Perform efficient self-attention calculation on the first fused motion sequence to obtain a second efficient self-attention output, and fuse the first efficient self-attention output with the second efficient self-attention output to obtain the output feature of the temporal attention module; Step 4: Repeat steps 2 to 3 several times in an iterative manner, and reduce the number of divided segments according to the number of iterations in each iteration to obtain the temporal attention output of each iteration, and fuse the temporal attention output obtained in each iteration to obtain the final temporal attention module output feature; Step 5: The final output feature of the temporal attention module is residually connected with the first fused motion sequence through the stylized block to obtain the second fused motion sequence; Step 6: Input the second fused motion sequence into the spatial attention module to extract frequency domain information, and obtain the output features of the spatial attention module; Step 7: The output features of the spatial attention module are sequentially passed through the multi-layer perceptron and the stylization block, and then residually connected with the output features of the spatial attention module to obtain the predicted noise; Step 8: Perform reverse denoising on the original noisy motion sequence according to the predicted noise to reconstruct the motion sequence.

2. The method for generating human motion based on spatiotemporal attention according to claim 1, characterized in that: In step 1, the original noisy motion sequence and the text description are fused based on cross attention to obtain a first fused motion sequence, which specifically includes the following steps: Use the pre-trained CLIP text encoder to extract word embeddings of text descriptions; Apply position embedding to the word-meta embedding according to the position information of the word-meta embedding in the text description to obtain a word-meta containing the position information; The word containing the position information is passed through Transformer to obtain text features; The text features and the original noisy motion sequence are cross-attention calculated, and the cross-attention calculation results are residually connected with the original noisy motion sequence through the stylized block to obtain the first fused motion sequence.

3. The method for generating human motion based on spatiotemporal attention according to claim 2, characterized in that: In step 2, the efficient self-attention calculation results of each segment are fused to obtain the process corresponding to the first efficient self-attention output, which has the following relationship: ; in, represents the first efficient self-attention output, It represents the continuity features of the divided motion sequence fragments obtained by the efficient attention module. Represents the splicing operation on motion sequence fragments.

4. The method for generating human motion based on spatiotemporal attention according to claim 3, characterized in that: In step 2, performing efficient self-attention calculation on the fragment specifically includes the following steps: Generate query vector, key vector and value vector using fragments; The global feature map is generated based on the key vector and the value vector. The corresponding process has the following relationship: ; in, represents the global feature map, , represents a real number, Indicates the dimension size of the feature, represents the key vector, represents a value vector, represents the normalized exponential function, represents the transpose operation, Represents matrix dot product; The efficient self-attention output feature of the fragment is calculated based on the global feature map and the query vector. The corresponding process has the following relationship: ; in, represents the query vector, Efficient self-attention output features to represent fragments.

5. The method for generating human motion based on spatiotemporal attention according to claim 4, characterized in that: In step 4, the temporal attention output obtained in each iteration is fused to obtain the final temporal attention module output feature. The corresponding process has the following relationship: ; in, represents the final temporal attention module output feature, represents the second most efficient self-attention output, represents the learnable weights, represents the first efficient self-attention output of the Mth iteration.

6. The method for generating human motion based on spatiotemporal attention according to claim 5, characterized in that: In step 6, the second fused motion sequence is input into the spatial attention module to extract frequency domain information, and the The specific steps include: The second fused motion sequence is generated through the self-attention mechanism to generate Z. The corresponding process has the following relationship: ; in, represents the feature representation that integrates global context information, represents the second fused motion sequence; The feature representation of the fusion of global context information is sequentially applied with discrete cosine transform, multi-layer perceptron and inverse discrete cosine transform to obtain low-frequency features. The corresponding process has the following relationship: ; in, represents the low-frequency features obtained by performing discrete cosine transform on the feature representation that integrates the global context information. represents the inverse discrete cosine transform, represents a multi-layer perceptron, represents discrete cosine transform, Representation layer normalization; Apply Z to a multilayer perceptron and combine Fusion, get , the corresponding process has the following relationship: ; in, It represents the feature representation that combines low-frequency features and original features; Will Through the convolution layer, we get , the corresponding process has the following relationship: ; in, Represents the output features of the spatial attention module.

7. The method for generating human motion based on spatiotemporal attention according to claim 6, characterized in that: In step 8, the original noisy motion sequence is reversely denoised according to the predicted noise to reconstruct the motion sequence. The corresponding process has the following relationship: ; in, represents the reverse diffusion process, represents the data distribution at time t, Indicates the timestamp, represents a Gaussian distribution, represents the predicted mean, Represents the parameters of the model.

8. The method for generating human motion based on spatiotemporal attention according to claim 7, characterized in that: The calculation process of the stylized module has the following relationship: ; in, represents the output of the stylization module, represents the input of the stylization module, represents the multiplicative offset of the original feature, Represents an additive offset to the original feature.

Citation Information

Patent Citations

  • Time sequence action detection method and device based on potential action interval feature integration

    CN118053107A

  • Weather data prediction method based on sequence segmentation and frequency domain attention

    CN119167067A

  • 3D human body posture estimation method based on diffusion model and continuous frame time-frequency information fusion

    CN119515983A

  • Time sequence prediction method and system based on diffusion autoregression transformer

    CN119622248A

  • Temporal bottleneck attention architecture for video action recognition

    US11270124B1

Cited By

  • Motion generation method and device based on multi-modal signal, equipment and storage medium

    CN120931775A

  • Multi-modal time sequence fusion voice drive gesture generation method

    CN121214502A