Deep Learning Method for Generating a Single Action Based on a Diffusion-GAN Framework

By combining diffusion model and GAN, a single action generation deep learning method based on the diffusion-generating adversarial framework is constructed, which solves the problem that action generation is difficult to take into account the rationality of kinematics and richness of details, and achieves high-quality action generation.

CN119888866BActive Publication Date: 2025-05-27JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510369711.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-05-27
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

The prior art is difficult to take into account the kinematic rationality and richness of detail in action generation, especially in the case of sparse data, and traditional single models are difficult to ensure both global coherence and local refinement.

Method used

A single action generation deep learning method based on the diffusion-generating adversarial framework is adopted, and the generator and discriminator are constructed by combining diffusion models and GANs, using matching and hybrid modules for action synthesis, and the generation quality is improved through multi-scale Patch-GAN discriminator and shallow U-Net architecture.

Benefits of technology

It achieves a dual breakthrough in kinematic rationality and detail richness in action generation, improves the generation quality, solves the problem that traditional models are difficult to take into account global coherence and local fineness, and supports the generation of sequences of arbitrary lengths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888866B_ABST
    Figure CN119888866B_ABST
Patent Text Reader

Abstract

The present invention proposes a deep learning method for single action generation based on a diffusion - generative adversarial framework. The method includes: dividing a motion sequence to obtain dynamic features and static features, and performing foot contact calculation through forward kinematics to obtain the calculated dynamic features; encoding the calculated dynamic features into Gaussian random noise and inputting them into a generator for denoising through a diffusion model; using a measurement of action similarity and an average voting strategy to synthesize actions for similar real actions; given an input action sequence to the calculated dynamic features, and injecting noise in an iterative form through a Markov denoising process in a specialized local attention layer of the diffusion model for denoising; independently evaluating the more accurate body actions and the reconstructed actions through a discriminator to generate classification results. The present invention effectively solves the problem of mode collapse under complex bone topologies through the joint optimization of gradient penalty and residual structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer graphics, and particularly to a deep learning method for single action generation based on a diffusion-generative adversarial framework. Background Art

[0002] Generating various realistic actions has always been a long-term goal in computer graphics and plays a crucial role in enhancing the authenticity and visual appeal of animated content. In recent years, generative adversarial networks (GANs) have emerged as a powerful solution, demonstrating impressive capabilities in the synthesis of high-quality motion sequences. In particular, even in the case of sparse data, GANs have shown remarkable generative abilities, producing convincing results with limited training samples.

[0003] Single-instance object generation methods in the motion field mainly rely on GANs. Among them, Ganimator is a notable method that uses a multi-level GAN to generate novel motions from a single motion sequence. However, its hierarchical cascaded generator and discriminator structures consume more computational resources. Similarly, LS-GAN utilizes the power of GANs and the compact representation of motion sequences in the latent space. However, it is mainly tailored for human motion and is limited by simple text-to-motion cues.

[0004] Although GANs focus on capturing the global features of motion synthesis - emphasizing the difference between real and generated samples, diffusion-based techniques are good at refining local details. For example, SinMDM adopts a shallow diffusion-based network and local attention layers, effectively reducing the receptive field to enhance motion diversity. However, this method encounters difficulties in synthesizing out-of-distribution motions. Summary of the Invention

[0005] In view of the above situation, the main objective of the present invention is to propose a deep learning method for single action generation based on a diffusion-generative adversarial framework to solve the above technical problems.

[0006] The present invention proposes a deep learning method for single action generation based on a diffusion-generative adversarial framework, and the method includes the following steps:

[0007] Step 1: Construct a generator and a discriminator based on the Diffusion-GAN architecture, and construct a matching and mixing module based on a matching and mixing mechanism. The generator, the discriminator, and the matching and mixing module constitute a diffusion-generative adversarial network model;

[0008] Among them, the generator includes a diffusion model, the diffusion model is combined with a shallow U-Net architecture, the discriminator operates in a multi-scale manner, and the matching and mixing module includes a matching module and a mixing module;

[0009] Step 2: Divide the motion sequence to obtain dynamic features and static features;

[0010] Convert the dynamic features to obtain a general motion representation, add foot contact labels to the general motion representation, and perform foot contact calculation through forward kinematics to obtain the calculated dynamic features;

[0011] Step 3: Add Gaussian random noise to the calculated dynamic features and input them into the generator to denoise through the diffusion model to obtain the denoised real action features;

[0012] Based on the matching and mixing module, use the measurement of action similarity and the average voting strategy to synthesize the denoised real action features to obtain the synthesized body action features;

[0013] Based on the calculated dynamic features, combined with the input action sequence, use the Markov denoising mechanism in the specialized local attention layer of the diffusion model to perform denoising processing in an iterative form, and output the reconstructed action features when the maximum number of iterations is reached;

[0014] Step 4: Independently evaluate the synthesized body action features and the reconstructed action features through the discriminator to generate classification results;

[0015] Step 5: Construct an adversarial loss based on the denoised real action features, construct a reconstruction loss based on the classification results, construct a diffusion loss based on the reconstructed action features, use the reconstruction loss and the diffusion loss to obtain the total reconstruction loss, and construct a matching loss based on the synthesized body action features;

[0016] Optimize the diffusion generative adversarial network model using the adversarial loss, reconstruction loss, diffusion loss, and matching loss to obtain the optimized diffusion generative adversarial network model;

[0017] Obtain the final classification result based on the optimized diffusion generative adversarial network model.

[0018] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0019] 1. By integrating the global feature capture ability of GAN and the local detail optimization mechanism of the diffusion model, the present invention achieves a double breakthrough in the kinematic rationality and detail richness of the generated actions, and solves the problem that it is difficult for traditional single models to balance global coherence and local fineness;

[0020] 2. The present invention constructs a joint evaluation system in the spatio-temporal dimension through a multi-scale Patch-GAN discriminator architecture, enabling the discriminator to simultaneously perceive the microscopic dynamic features and macroscopic motion patterns of the action sequence, and showing a higher generation quality compared to traditional single-scale discriminators;

[0021] 3. The present invention controls the receptive field of the model within the range of 8 - 12 frames through a shallow U-Net and attention collaborative architecture, supporting the generation of sequences of any length while ensuring action continuity, and successfully achieving the synthesis of 60-second long sequences with a single training;

[0022] 4. The present invention realizes the local feature adaptive fusion of the generated samples and the reference samples in the adversarial training through a matching-mixing dynamic adjustment mechanism, improving the inter-diversity index of the synthesized actions by 15.8% (please refer to Table 4 in the specification), while maintaining an action coverage rate of more than 97% (please refer to Table 2 in the specification);

[0023] 5. The present invention effectively solves the mode collapse problem under complex bone topologies through the joint optimization of gradient penalty and residual structure.

[0024] Additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is a flowchart of the steps of the deep learning method for single action generation based on a diffusion-generative adversarial framework proposed by the present invention.

[0026] Figure 2 It is a schematic diagram of the SinMDGan framework of the deep learning method for single action generation based on a diffusion-generative adversarial framework proposed by the present invention.

[0027] Figure 3 It is a motion representation diagram of the deep learning method for single action generation based on a diffusion-generative adversarial framework proposed by the present invention.

[0028] Figure 4 It is a spatial action combination diagram of the deep learning method for single action generation based on a diffusion-generative adversarial framework proposed by the present invention.

[0029] Figure 5 It is a motion expansion diagram of the deep learning method for single action generation based on a diffusion-generative adversarial framework proposed by the present invention.

[0030] Figure 6 It is a spatial composition diagram of the deep learning method for single action generation based on a diffusion-generative adversarial framework proposed by the present invention.

[0031] Figure 7 Motion sequence generation diagram of the deep learning method for single action generation based on the diffusion-generative adversarial framework proposed by the present invention. Detailed implementation manners

[0032] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals are the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.

[0033] Referring to the following description and drawings, these and other aspects of the embodiments of the present invention will be clear. In these descriptions and drawings, some specific implementation manners in the embodiments of the present invention are specifically disclosed as some ways to implement the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0034] Please refer to Figure 1 , an embodiment of the present invention proposes a deep learning method for single action generation based on a diffusion-generative adversarial framework. The method includes the following steps:

[0035] Step 1: Construct a generator and a discriminator based on the Diffusion-GAN architecture, and construct a matching and mixing module based on the matching and mixing mechanism. The generator, the discriminator, and the matching and mixing module constitute the Diffusion-GAN model;

[0036] Among them, the generator includes a diffusion model, the diffusion model is combined with a shallow U-Net architecture, the discriminator operates in a multi-scale manner, and the matching and mixing module includes a matching module and a mixing module.

[0037] Step 2: Divide the motion sequence to obtain dynamic features and static features;

[0038] Convert the dynamic features to obtain a general motion representation, add a foot contact label to the general motion representation, and perform foot contact calculation through forward kinematics to obtain the calculated dynamic features.

[0039] Please refer to Figure 3 , in Step 2, when converting the dynamic features to obtain a general motion representation, the relational expressions existing in the corresponding process are as follows:

[0040] ;

[0041] Among them, represents the general motion representation, represents a real number matrix with a dimension of , represents the total number of frames, Represents the feature dimension of each frame;

[0042] In the step of adding foot contact labels to the general motion representation and calculating foot contact through forward kinematics to obtain the calculated dynamic features, the relationships in the corresponding process are as follows:

[0043] ;

[0044] Among them, Represents the foot contact label, Represents the number of joints that frequently contact the ground during the movement process, Represents the number of skeletal joints, Represents the foot The Velocity of joint in the Represents forward kinematics, Represents the skeletal parameters in the static features, Represents the ground height of the frame, Represents the dynamic feature vector of joint in the foot kinematic model at the frame.

[0045] It should be noted that adding the foot contact label is to reduce common foot sliding artifacts. The general motion representation is the motion data in the skeletal motion (dynamic features), such as joint velocity and position, Represents the ground height or contact threshold of the nth frame, and this value changes dynamically.

[0046] Furthermore, in Figure 3 the 3D skeleton structure of the 6D rotation representation is presented as a cube, vividly representing the three-dimensional skeleton representation at the intersection of time and space.

[0047] It should be noted that Represents joint rotation, Represents local root joint displacement, Represents the number of rotation features of each joint, Represents rotation features + local contact features (used to describe the motion state of the skeleton in space-time), Represents rotation features + local contact features + displacement features (the displacement feature of the general root node is set to 3), Represents the total number of frames The subset frames in, which is a smaller unit, Represents the skeleton.

[0048] Step 3: Add Gaussian random noise to the calculated dynamic features and input them into the generator for denoising through the diffusion model to obtain the denoised real action features;

[0049] Based on the matching and mixing module, use the measurement of action similarity and the average voting strategy to perform action synthesis on the denoised real action features to obtain the synthesized body action features;

[0050] Based on the calculated dynamic features, combine the input action sequence and input it into the Markov denoising mechanism in the dedicated local attention layer of the diffusion model for iterative denoising processing. When the maximum number of iterations is reached, output the reconstructed action features.

[0051] Please refer to Figure 2 , in Step 3, based on the matching and mixing module, use the measurement of action similarity and the average voting strategy to perform action synthesis on the denoised real action features to obtain the synthesized body action features. The corresponding relationship is as follows:

[0052] ;

[0053] Among them, represents the synthesized body action features, represents the matching and mixing module, represents the generator network, represents time represents the sequence of

[0054] It should be noted that the matching module in measures the similarity between the example action patches and the synthesized action patches, while the mixing module uses the average voting strategy to combine the collected action patches to construct the synthesized partial body action features.

[0055] Based on the calculated dynamic features, combine the input action sequence, and use the Markov denoising mechanism in the dedicated local attention layer of the diffusion model to perform iterative denoising processing. When the maximum number of iterations is reached, output the reconstructed action features. The corresponding relationship is as follows:

[0056] ;

[0057] Among them, represents the probability distribution function in the Markov denoising process, represents the Gaussian distribution, represents the parameter that controls the noise addition intensity in the diffusion process, represents the identity matrix.

[0058] Furthermore, the diffusion model uses a probabilistic generation mechanism to effectively capture complex details in the data distribution. By gradually introducing and removing noise, the diversity of the generated samples is enhanced. In contrast, traditional Generative Adversarial Networks (GANs) rely on a generator that directly synthesizes samples, which sometimes leads to mode collapse and fails to capture complex motion variations. In the generator design of the present invention, the diffusion model is combined with a shallow U-Net architecture and a specialized local attention layer is incorporated. Experiments show that in the U-Net structure, using a narrower receptive field can significantly improve the motion realism and credibility in the field of motion.

[0059] It should be noted that in Figure 2 , removing noise: The model starts from the noise and gradually recovers the data through learning by a parameterized network, experiencing intermediate states → →...→ , and finally obtains the denoising result . Taking the → step as an example: When the model completes denoising of the state, the network calculates the deviation between the current state and the target distribution, and by adjusting the noise prediction gradient, makes updated along the descending direction of the loss function to a purer state.

[0060] Step 4: Independently evaluate the synthesized body motion features and the reconstructed motion features through a discriminator to generate classification results.

[0061] Furthermore, the discriminator of traditional GANs usually outputs a single scalar value to classify the input as real or fake. However, such an overly simple method may lead to mode collapse, that is, the generator overfits a limited training sample set, thereby reducing the diversity of the generated output. To solve this problem, the present invention adopts MultiScaleGAN, which is an enhanced version of PatchGAN. In this framework, multiple discriminators run at different scales, independently evaluate the input and produce real or fake classifications, and the total loss is calculated as a weighted sum of the multi-scale evaluations, thereby improving the robustness and diversity of motion generation.

[0062] Step 5: Construct an adversarial loss based on the denoised real motion features, a reconstruction loss based on the classification results, a diffusion loss based on the reconstructed motion features, obtain the total reconstruction loss using the reconstruction loss and the diffusion loss, and construct a matching loss based on the synthesized body motion features;

[0063] Optimize the diffusion generative adversarial network model using adversarial loss, reconstruction loss, diffusion loss, and matching loss to obtain the optimized diffusion generative adversarial network model;

[0064] Obtain the final classification result based on the optimized diffusion generative adversarial network model.

[0065] In step 5, construct the adversarial loss based on the denoised real action features, and the corresponding relationship is as follows:

[0066] ;

[0067] Among them, represents the adversarial loss, represents the expected value of sampling the first data from the real data distribution , represents the expected value of sampling the noise vector from the noise distribution , represents the gradient penalty term, represents the case of penalizing the gradient norm deviation from 1, represents the expected value of sampling the second data from the linearly interpolated distribution , represents the gradient operator, represents the discriminator.

[0068] It should be noted that represents the expected value of a randomly drawn data from the data distribution , represents the expected value of a randomly drawn noise vector from the noise distribution , represents the expected value of the second data randomly drawn from the data distribution .

[0069] Construct the reconstruction loss based on the classification result, and the corresponding relationship is as follows:

[0070] ;

[0071] Among them, represents the reconstruction loss, represents the number of different states of the discriminator, represents sampling and calculating the expected value of the sample generated from the distribution of the generator , represents the The discrimination score of a discriminator for the generated samples .

[0072] Based on the reconstructed actions, a diffusion loss is constructed, and the corresponding relationship is as follows:

[0073] ;

[0074] Among them, represents the diffusion loss, represents the sampling time step in the discrete uniform distribution , which is randomly selected with equal probability between 1 and the total number of time steps to generate an expected value, represents the action sequence data before processing.

[0075] It should be noted that represents the total number of steps in the diffusion and denoising processes. Specifically, in diffusion, it represents the number of steps for continuously adding noise to the input sequence until it reaches the Gaussian distribution; in denoising, it represents the number of steps for continuously denoising until the noise is removed.

[0076] Based on the synthesized body action features, a matching loss is constructed, and there is also a calculation of the patch distance matrix for the bone part. The corresponding relationship is as follows:

[0077] ;

[0078] Among them, represents the normalized distance between the th patch of the generated motion and the th patch of the reference motion in the patch distance matrix of the th bone part, represents the th bone part of the th motion patch extracted from the generated motion sequence, represents the th bone part of the th motion patch extracted from the reference motion sequence, represents a hyperparameter, represents finding the patch in all patches of the generated motion that is closest to the reference patch , that is, the minimum squared L2 distance.

[0079] Furthermore, the total reconstruction loss is obtained using the reconstruction loss and the diffusion loss. The corresponding relationship is as follows:

[0080] ;

[0081] Among them, Represents the total reconstruction loss.

[0082] The complete training objectives are summarized as follows:

[0083] ;

[0084] Represents the total loss, Represents the hyperparameter for balancing the adversarial objective, Represents the hyperparameter for the reconstruction objective.

[0085] For benchmarking, the model of the present invention was compared with SinMDM and Ganimator, as the SinMDM model and the Ganimator model represent the state-of-the-art baselines. To ensure the evaluation is valid and fair, the same metrics used in the established benchmarks were adopted.

[0086] The evaluation framework of the present invention prioritizes models that exhibit a balanced score distribution across multiple metrics over models that excel in a single aspect. To maintain rigor, this embodiment follows the methods established in previous studies and is consistent with the metric evaluation method introduced in SinMDM. Specifically, the well-recognized standard harmonic mean metric in machine learning is adopted to provide a comprehensive performance evaluation.

[0087] To ensure comparability between metrics, the following steps are taken to apply the standardization procedure:

[0088] Standardization range: Scores are scaled between zero and the maximum value of each corresponding metric;

[0089] Maximum value estimation: If there is no true maximum value, the 90th percentile of the observed scores is used for the maximization approximation;

[0090] Directional adjustment: For metrics where lower scores are more desirable, the scores are transformed by subtracting the obtained value from the determined maximum value; negative values are allowed while preserving the interpretability of the metrics;

[0091] The formula for the harmonic mean is:

[0092] ;

[0093] where, Represents the harmonic mean, Represents the total number of metrics, Represents the th numerical value in the harmonic mean formula;

[0094] Evaluation metrics: The present invention uses the following five metrics to evaluate performance:

[0095] Metric a. Coverage - Measures the proportion of the time window in the input motion sequence that is successfully reproduced in the synthesized motion;

[0096] Metric b. Global diversity - Calculates the distance between and where represents the tessellation that minimizes the distance from the input sequence. The function is equivalent to mapping into an optimization space, representing the tessellation that minimizes the distance from the input sequence;

[0097] Metric c. Local diversity - Measures the average distance between the time windows in the synthesized motion and the nearest corresponding windows in the input sequence;

[0098] Metric d. Mutual diversity - Quantifies the differences between different synthesized motion samples;

[0099] Metric e. Intra - diversity difference - Calculates the difference between the intra - sequence diversity of the synthesized motion and the intra - sequence diversity of the original input motion;

[0100] For Metrics a to d, a higher score indicates better performance; while for Metric e, a lower score indicates better performance.

[0101] Table 1: Quantitative comparison of the model of the present invention with Ganimator and SinMDM

[0102]

[0103] As can be seen from Table 1: Each metric is calculated separately on the benchmark actions and then averaged. The results show that: Except for diversity, the model of the present invention outperforms Ganimator in all metrics. It is worth noting that SinMDM always outperforms Ganimator in almost all metrics, especially having a strong advantage in the harmonic mean metric, highlighting its excellent overall performance.

[0104] Table 2: Comparison results of the model of the present invention with MotionTexture and acRNN on Salar - Dancing motion data

[0105]

[0106] As can be seen from Table 2, quantitative analysis shows that: the model of the present invention achieves state-of-the-art performance in terms of global diversity and harmonic mean, outperforming existing methods. In addition, the coverage rate and local diversity metrics show performance levels comparable to those of Ganimator and SinMDM, further verifying the effectiveness of the method of the present invention. The model of the present invention leads in terms of global diversity and harmonic mean score, and achieves comparable scores in terms of coverage rate and local diversity.

[0107] Given a reference motion sequence and a region of interest (ROI) mask, the goal is to synthesize a new motion. The region of interest is generated by random noise, while the remaining regions are as close as possible to the given motion. The relational expressions existing in the corresponding process are as follows:

[0108] ;

[0109] where, represents the synthesized new motion, represents element-wise multiplication, represents the mask of the region of interest (a binary tensor used to mark the region that needs to be synthesized), represents the reference motion sequence.

[0110] The model of the present invention aims to achieve a seamless transition between the original segment and the synthesized segment. The reference motion can be any sequence and is independent of the training data. When defining the region of interest using a binary mask, due to discontinuity, a sudden transition may occur between the fixed and generated segments. To solve this problem, the present invention introduces a linear interpolation boundary; specifically, the fixed motion segment remains unchanged, while the sampled regions that need to be synthesized are gradually blended; during the inference process, the present invention iteratively refines the motion to ensure that the relevant part of is merged into ; expressed as:

[0111] ;

[0112] If , the ROI is not modified, and the output remains the same as the reference example. On the contrary, if , the specified region is synthesized.

[0113] Temporal composition involves filling or extending the frames within a motion sequence; this concept is embodied in two key cases: intermediate frames and motion extension.

[0114] Intermediate frames create a smooth transition between key frames and play a crucial role in animation. The model of the present invention is good at generating realistic intermediate motions using only the initial and final poses. As Figure 4As shown, the synthesized intermediate frames create smooth transitions, preserving the realism of the motion and extending the sequence duration. In Figure 4 , the initial and final actions remain unchanged, while the intermediate actions are synthesized to achieve realistic and diverse variations.

[0115] See Figure 5 , the motion expansion extends a given motion sequence, similar to inpainting. Motion expansion is illustrated in Figure 5 , where the reference motion is extended forward and backward. In this case, the ROI mask is set to zero for the center frame, while the outer regions are specified as zero. Motion expansion not only extends the sequence but also reorganizes the motion dynamics of each frame, resulting in diverse and high-quality continuous actions; the actions on both sides are newly generated, while the actions in the middle represent the reference motion; the extended sequence demonstrates diverse synthesis from a single input.

[0116] See Figure 6 , spatial motion synthesis can also be applied by specifying specific joint indices in the ROI mask. In Figure 6 , the use of the ROI mask to edit upper body motion is illustrated; the upper body motion is guided by the reference motion, while the lower body motion is synthesized to maintain consistency.

[0117] See Figure 7 , the model of the present invention is capable of generating motion sequences of variable lengths, including extended animations, without additional training. Due to the model having a narrow receptive field, the model of the present invention can generate sequences of arbitrary length. In Figure 7 , an example of long action generation is shown; in it, a one-minute animation is synthesized based on the learned salar-dance sequence; the learned sequence is the salar dance, and the synthesized action is a 60-second sequence.

[0118] To evaluate the contributions of two key components, the multi-scale Patch-GAN structure and The Match and Blend Mechanism, the present invention conducted an ablation study, separated each component, and discussed the impact of the ablation module on the results.

[0119] Table 3: Results of internal diversity differences

[0120]

[0121] As can be seen from Table 3, the lower the mean and variance values, the smaller the difference between the internal diversity of the synthetic motion and the internal diversity of the input motion, thus achieving better global consistency. In addition, adding the multi-scale Patch-GAN can significantly reduce the difference in internal diversity, thereby improving the global motion consistency, highlighting its advantages in enhancing motion realism and structural coherence.

[0122] Table 4: Internal Diversity Results Model

[0123]

[0124] As can be seen from Table 4, the matching and mixing mechanism enhances the diversity between sub-windows in the motion sequence, thus improving the diversity index. Introducing the matching and mixing mechanism can improve the diversity metric, which measures the difference between sub-windows in the generated actions. This matching and mixing mechanism enhances the diversity of the synthetic actions while preserving the coherence of each sequence.

[0125] In this experiment, the present invention replaces the traditional Patch-GAN used in Ganimator with the multi-scale Patch-GAN. The original Ganimator framework uses a Patch-GAN classifier to prevent overfitting by restricting the receptive field of the discriminator. This classifier assigns confidence values at the patch level, and the final discriminator output is calculated as the average confidence of all patches; although this method is effective for evaluating local details, it may ignore the global consistency in the generated motion.

[0126] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0127] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0128] The embodiments described above merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation to the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.

Claims

1. A single action generation deep learning method based on a diffusion-generation adversarial framework, characterized in that: The method comprises the following steps: Step 1: Construct a generator and a discriminator based on the diffusion generative adversarial network architecture, and construct a matching and mixing module based on the matching and mixing mechanism. The generator, discriminator, matching and mixing modules constitute a diffusion generative adversarial network model. The generator includes a diffusion model, the diffusion model is combined with a shallow U-Net architecture, the discriminator operates in a multi-scale manner, and the matching and mixing module includes a matching module and a mixing module; Step 2: Divide the motion sequence to obtain dynamic features and static features; The dynamic features are converted to obtain a general motion representation, a foot contact label is added to the general motion representation, and foot contact is calculated by forward kinematics to obtain the calculated dynamic features; Step 3: Add Gaussian random noise to the calculated dynamic features, and input them into the generator for denoising through the diffusion model to obtain the denoised real action features; Based on the matching and mixing modules, the denoised real action features are synthesized using the measurement of action similarity and the average voting strategy to obtain the synthesized body action features. Based on the calculated dynamic features and the input action sequence, the Markov denoising mechanism in the special local attention layer of the diffusion model is used to perform denoising in an iterative manner, and the reconstructed action features are output when the maximum number of iterations is reached; Step 4: The synthesized body motion features and the reconstructed motion features are independently evaluated by the discriminator to generate classification results; Step 5: Construct adversarial loss based on the real motion features after denoising, construct reconstruction loss based on the classification results, construct diffusion loss based on the reconstructed motion features, use reconstruction loss and diffusion loss to get the total reconstruction loss, and construct matching loss based on the synthesized body motion features; The diffusion generative adversarial network model is optimized using adversarial loss, reconstruction loss, diffusion loss and matching loss to obtain the optimized diffusion generative adversarial network model; The final classification result is obtained based on the optimized diffusion generative adversarial network model.

2. The single action generation deep learning method based on the diffusion-generation adversarial framework according to claim 1 is characterized in that: In step 2, the dynamic features are transformed to obtain a general motion representation, and the relationship between the corresponding process is as follows: ; in, Indicates general movement expression, The dimension is A real matrix of Indicates the total number of frames. Represents the feature dimension of each frame; In the steps of adding foot contact labels to the general motion representation and calculating foot contact through forward kinematics to obtain the calculated dynamic features, the corresponding process has the following relationship: ; in, Indicates that the foot is touching the label, Indicates the number of joints that contact the ground at high frequency during movement. Indicates the number of bone joints, Indicates foot No. Joints in frame speed, represents forward kinematics, Represents the bone parameters in the static feature, Indicates The ground height of the frame, represents the first Frame time joint The dynamic feature vector of .

3. The single action generation deep learning method based on the diffusion-generation adversarial framework according to claim 2 is characterized in that: In step 3, based on the matching and mixing module, the denoised real action features are synthesized using the measured action similarity and average voting strategy to obtain the synthesized body action features. The corresponding process has the following relationship: ; in, represents the synthesized body movement characteristics, represents the matching and mixing modules, represents the generator network, Indicates time Represents a sequence.

4. The single action generation deep learning method based on the diffusion-generation adversarial framework according to claim 3 is characterized in that: Based on the calculated dynamic features, combined with the input action sequence, the Markov denoising mechanism in the special local attention layer of the diffusion model is used to perform denoising in an iterative form. When the maximum number of iterations is reached, the reconstructed action features are output. The corresponding process has the following relationship: ; in, represents the probability distribution function in the Markov denoising process, represents a Gaussian distribution, represents the parameter that controls the intensity of noise addition during the diffusion process, Represents the identity matrix.

5. The single action generation deep learning method based on the diffusion-generation adversarial framework according to claim 4 is characterized in that: In step 5, the adversarial loss is constructed based on the real action features after denoising, and the relationship between the corresponding process is as follows: ; in, Represents resistance to loss, Represents the distribution from real data Sample the first data The expected value of Represents the noise distribution The noise vector is sampled from The expected value of represents the gradient penalty term, Indicates the situation where the penalty gradient norm deviates from 1, represents the distribution from linear interpolation Sample the second data The expected value of represents the gradient operator, Represents the discriminator.

6. The single action generation deep learning method based on the diffusion-generation adversarial framework according to claim 5 is characterized in that: The reconstruction loss is constructed based on the classification results, and the relationship between the corresponding process is as follows: ; in, represents the reconstruction loss, represents the number of different states of the discriminator, Represents a generator Distribution Samples generated by sampling And calculate the expected value, Indicates The discriminator generates samples The discrimination score of .

7. The single action generation deep learning method based on the diffusion-generation adversarial framework according to claim 6 is characterized in that: Based on the reconstructed motion features, the diffusion loss is constructed, and the relationship between the corresponding process is as follows: ; in, represents the diffusion loss, represents the sampling time step in the discrete uniform distribution From 1 to the total time step Randomly select with equal probability and generate expected value, Represents action sequence data before processing.

8. The single action generation deep learning method based on the diffusion-generation adversarial framework according to claim 7 is characterized in that: Based on the synthesized body motion features, the matching loss is constructed, and there is also a patch distance matrix calculation for the skeleton part. The corresponding relationship in the process is as follows: ; in, Indicates The motion is generated from the patch distance matrix of the skeleton parts. Patch and reference motion The normalized distance of patches, represents the first The first part of the skeleton Sports patches, represents the first The first part of the skeleton Sports patches, represents the hyperparameter, Represents all patches in the generated motion Find the reference patch in The closest patch.

Citation Information

Patent Citations

  • Voice-driven posture action generation method and device based on diffusion model

    CN117292704A

  • Human body video generation method based on time sequence consistent hidden space guide diffusion model

    CN117994708A