Refined three-dimensional human body action sequence generation method and system based on SMPL model

Through the refined three-dimensional human body movement sequence generation method based on the SMPL model, the shortcomings of traditional methods in capturing subtle movement changes and individual differences are solved, and high-quality and natural three-dimensional human body movement sequence generation is achieved to adapt to the motion needs in complex scenarios.

CN120047584APending Publication Date: 2025-05-27ANOTHER ME (BEIJING) VIRTUAL TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510121047.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing three-dimensional human movement generation method based on traditional models has shortcomings in capturing subtle movement changes and individual differences, resulting in the generated human movements lacking a sense of nature and reality, and it is difficult to deal with the continuity and fluency of movement in dynamic and complex scenes.

Method used

The refined three-dimensional human body movement sequence generation method based on the SMPL model is adopted. By collecting and preprocessing the three-dimensional human body movement data, the SMPL model and adversarial network model are constructed, the adversarial network model is trained using the data set, SMPL parameters are generated, and the refined three-dimensional human body movement sequence is reconstructed through the body posture conversion process of the SMPL model.

Benefits of technology

It improves the accuracy and nature of three-dimensional human body modeling and animation generation, and can generate high-quality and refined three-dimensional human body movement sequences from the three-dimensional human body movement data obtained in real time, adapting to the continuity and fluency of movement in complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047584A_ABST
    Figure CN120047584A_ABST
Patent Text Reader

Abstract

The invention discloses a refined three-dimensional human body action sequence generation method and system based on an SMPL model, and belongs to the technical field of computer vision. Constructing an SMPL model and an adversarial network model, and training the adversarial network model by using the data set to obtain a parameter optimization model; preprocessing the three-dimensional human body motion data acquired in real time, inputting the parameter optimization model, and generating SMPL parameters; and according to the SMPL parameters, reconstructing a refined three-dimensional human body action sequence by utilizing a body posture conversion process of the SMPL model. According to the method, the powerful human body modeling ability of the SMPL model and the learning ability of the adversarial network are combined, so that a high-quality and refined three-dimensional human body action sequence is generated from the three-dimensional human body action data acquired in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more specifically, to a method and system for generating refined three-dimensional human motion sequences based on the SMPL model. Background Art

[0002] The SMPL model, with its elegant parametric structure, lays the foundation for modeling human postures and shapes. By combining deep learning techniques, it can generate highly realistic and complex human motion sequences. The development of this technology is of great significance in many fields such as entertainment, game design, virtual reality, and motion analysis. With the progress of technology, the demand for interactive experiences from users is increasing day by day, which requires that the generated motion sequences not only be consistent and credible but also adapt to the user's behavior in real time.

[0003] However, the current motion generation methods based on traditional models have deficiencies in capturing subtle motion changes and individual differences, often resulting in the lack of naturalness and realism in the generated human motions. In addition, the existing technologies often have difficulty in dealing with the motion continuity and smoothness in dynamic and complex scenarios, and are prone to presenting uncoordinated motion flows.

[0004] Therefore, how to propose a method and system for generating refined three-dimensional human motion sequences based on the SMPL model to improve the accuracy and naturalness of three-dimensional human modeling and animation generation is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides a method and system for generating refined three-dimensional human motion sequences based on the SMPL model to improve the accuracy and naturalness of three-dimensional human modeling and animation generation.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] On the one hand, the present invention discloses a method for generating refined three-dimensional human motion sequences based on the SMPL model, including the following steps:

[0008] Collect three-dimensional human motion data and perform preprocessing to obtain preprocessed data, and construct a data set according to the preprocessed data;

[0009] Construct an SMPL model and an adversarial network model, and train the adversarial network model using the data set to obtain a parameter-optimized model;

[0010] Perform preprocessing on the real-time obtained three-dimensional human motion data and input it into the parameter-optimized model to generate SMPL parameters;

[0011] Reconstruct a refined three-dimensional human motion sequence using the body pose conversion process of the SMPL model according to the SMPL parameters.

[0012] Preferably, the adversarial network model includes a generator and a discriminator;

[0013] The generator includes a first input layer, a first convolutional layer, an LSTM layer, a first fully connected layer, and a first output layer connected in sequence; the discriminator includes an input layer, a convolutional layer, a fully connected layer, and an output layer connected in sequence; the first output layer is connected to the input layer.

[0014] Preferably, the generator loss function L G consists of a three-dimensional key point loss function L 1 , a two-dimensional key point loss function L 2 , an identity pose parameter loss function L 3 , and an adversarial loss function L adv , and the formula is as follows:

[0015] L G =λ 1 L 1 +λ 2 L 2 +λ 3 L 3 +λ adv L adv ;

[0016] In the formula, λ 1 , λ 2 , λ 3 , λ adv are the weight parameters of the three-dimensional key point loss function, the two-dimensional key point loss function, the identity pose parameter loss function, and the adversarial loss function, respectively.

[0017] Preferably, the formula of the three-dimensional key point loss function L 1 is as follows:

[0018]

[0019] In the formula, N is the number of three-dimensional key points, is the i-th three-dimensional key point generated by the generator, and P i is the i-th real three-dimensional key point.

[0020] Preferably, the formula of the two-dimensional key point loss function L 2 is as follows:

[0021]

[0022] In the formula, M is the number of two-dimensional key points, The j-th two-dimensional key point generated by the generator, Q j is the real j-th two-dimensional key point.

[0023] Preferably, the identity pose parameter loss function L 3 has the following formula:

[0024]

[0025] In the formula, is the pose parameter generated by the generator, θ true is the real pose parameter.

[0026] Preferably, the identity pose parameter loss function L 4 has the following formula:

[0027] L 4 = -logD(G(Z));

[0028] In the formula, D(G(Z)) is the probability of the discriminator for the generated sample.

[0029] Preferably, the discriminator loss function L D is constructed based on the cross-entropy loss.

[0030] A refined three-dimensional human motion sequence generation system based on the SMPL model, used to implement the above-mentioned refined three-dimensional human motion sequence generation method based on the SMPL model, includes:

[0031] A data processing module, used to collect three-dimensional human motion data and perform preprocessing to obtain preprocessed data, and construct a data set according to the preprocessed data;

[0032] A model construction and training module, used to construct an SMPL model and an adversarial network model, and use the data set to train the adversarial network model to obtain a parameter optimization model;

[0033] A parameter generation module, used to preprocess the three-dimensional human motion data obtained in real time and input it into the parameter optimization model to generate SMPL parameters;

[0034] An action output module, used to reconstruct a refined three-dimensional human motion sequence according to the SMPL parameters by using the body pose conversion process of the SMPL model.

[0035] As can be seen from the above technical solutions, compared with the prior art, the present invention discloses a method and system for generating a refined three-dimensional human motion sequence based on the SMPL model, constructs an SMPL model and an adversarial network model, trains the adversarial network model using a data set to obtain a parameter optimization model; preprocesses the three-dimensional human motion data obtained in real time and inputs it into the parameter optimization model to generate SMPL parameters; according to the SMPL parameters, uses the body pose conversion process of the SMPL model to reconstruct a refined three-dimensional human motion sequence. The present invention combines the powerful human body modeling ability of the SMPL model and the learning ability of the adversarial network to realize the generation of high-quality and refined three-dimensional human motion sequences from the three-dimensional human motion data obtained in real time. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings according to the provided drawings without creative efforts.

[0037] Figure 1 It is a flowchart of the method provided by the present invention;

[0038] Figure 2 It is an architecture diagram of the system provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0040] On the one hand, the present invention discloses a method for generating a refined three-dimensional human motion sequence based on the SMPL model, as Figure 1 shown, the method includes the following steps:

[0041] S1. Collect three-dimensional human motion data and perform preprocessing to obtain preprocessed data, and construct a data set according to the preprocessed data.

[0042] Select data sets containing rich three-dimensional human motion and shape information, such as MPI-INF-3DHP, Human3.6M, 3DPW, etc. These data sets provide high-quality three-dimensional joint point information and SMPL parameters, which are helpful for training more accurate models.

[0043] Preprocess the dataset, including data cleaning, normalization, and necessary augmentation operations. Ensure that the format and quality of the input data meet the requirements of model training.

[0044] S2. Build the SMPL model and the adversarial network model, and use the dataset to train the adversarial network model to obtain a parameter-optimized model.

[0045] Select the shape parameters and pose parameters of the SMPL model.

[0046] Shape parameters: Use 10-dimensional shape parameters to describe the shape characteristics of the human body, such as height, weight, etc.

[0047] Pose parameters: Use 24×3-dimensional pose parameters to describe the action postures of the human body. The rotation angle of each joint point relative to its parent node is expressed in axis-angle form.

[0048] The adversarial network model includes a generator and a discriminator;

[0049] The generator includes a first input layer, a first convolutional layer, an LSTM layer, a first fully connected layer, and a first output layer connected in sequence.

[0050] Among them, the first input layer receives the preprocessed input data (for example, the coordinates of 2D joint points, video frames, etc.). The first convolutional layer uses a series of convolutional layers to extract spatial features, and the stacking of convolutional layers and pooling layers (such as max pooling) can be used to gradually reduce the size of the feature map; the LSTM layer is used to retain time series information and process the dynamic changes between frames; the first fully connected layer is used to map the extracted features to shape parameters (10 dimensions) and pose parameters (24×3 dimensions); the first output layer is used to output shape parameters and pose parameters.

[0051] The discriminator includes an input layer, a convolutional layer, a fully connected layer, and an output layer connected in sequence. The input layer is used to receive the output of the generator (SMPL parameters) and the ground truth (annotated SMPL parameters); the convolutional layer extracts the input features, and multiple convolutional layers and activation functions (such as ReLU) can usually be designed; the fully connected layer integrates the extracted features and outputs a probability value; the output layer generates a probability value indicating whether the input (parameters) come from real data.

[0052] Generator loss function L G Consists of the 3D keypoint loss function L 1 The 2D keypoint loss function L 2 The identity pose parameter loss function L 3 And the adversarial loss function L adv Composed as follows:

[0053] L G =λ 1L 1 + λ 2 L 2 + λ 3 L 3 + λ adv L adv ;

[0054] where λ 1 , λ 2 , λ 3 , λ adv are the weight parameters of the 3D key point loss function, 2D key point loss function, identity pose parameter loss function, and adversarial loss function, respectively.

[0055] Preferably, the formula of the 3D key point loss function L 1 is as follows:

[0056]

[0057] where N is the number of 3D key points, is the i-th 3D key point generated by the generator, and P i is the real i-th 3D key point.

[0058] Preferably, the formula of the 2D key point loss function L 2 is as follows:

[0059]

[0060] where M is the number of 2D key points, is the j-th 2D key point generated by the generator, and Q j is the real j-th 2D key point.

[0061] Preferably, the formula of the identity pose parameter loss function L 3 is as follows:

[0062]

[0063] where is the pose parameter generated by the generator, and θ true is the real pose parameter.

[0064] Preferably, the formula of the identity pose parameter loss function L 4 is as follows:

[0065] L 4 = -logD(G(Z));

[0066] where D(G(Z)) is the probability of the discriminator for the generated samples.

[0067] Preferably, the discriminator loss function LD Constructed based on cross - entropy loss, and the specific formula is as follows:

[0068]

[0069] In the formula, N is the total number of samples; y i is the true label, representing the true category of the i - th sample. If the sample x i is the actual true data, then y i = 1; if the sample G(z i ) is the generated sample, then y i = 0. D(x i ) is the output of the discriminator, representing the recognition probability of the discriminator for the true sample x i (true parameter), which is the result obtained through the forward propagation of the discriminator network and is used to determine whether the sample is real (1) or generated (0). G(z i ) is the output of the generator, which is a sample generated by inputting the noise z i . This sample is for the discriminator to evaluate the quality of the generator. D(G(z i )) is the output of the discriminator, representing the recognition probability of the discriminator for the generated sample G(z i ). This value reflects the possibility that the sample generated by the generator is considered real by the discriminator.

[0070] During the training process, by optimizing the loss functions of the generator and the discriminator, the parameters of the model are gradually updated. In this embodiment, the Adam optimization algorithm can be used to optimize the loss function.

[0071] S3. Pre - process the three - dimensional human motion data obtained in real - time and input it into the parameter optimization model to generate SMPL parameters.

[0072] S4. According to the SMPL parameters output by the parameter optimization model, use the body pose transformation process of the SMPL model to reconstruct a refined three - dimensional human motion sequence.

[0073] The body pose transformation process of the SMPL model mainly includes three steps: Shape Blend Shapes, Pose Blend Shapes, and Skinning.

[0074] Shape Blend Shapes deforms the human model using shape parameters to match the body type characteristics of a specific individual. The shape parameters are represented by the parameters of the body shape deformation basis extracted by PCA (Principal Component Analysis), and these parameters can describe the main changes in human shape.

[0075] Pose-based hybrid shaping uses pose parameters to adjust the pose of a human model to match the action pose at a specific moment. The pose parameters describe the rotation angles of human joint points, and these angles are encoded in the axis-angle representation. By adjusting these rotation angles, the joint point positions of the human model can be changed, thus achieving pose adjustment.

[0076] The skinning step is to apply the deformed shape and pose to the mesh vertices of the human model to generate the final 3D human model.

[0077] The SMPL model uses the Linear Blend Skinning technique to achieve this step.

[0078] The Linear Blend Skinning technique calculates the final position of each vertex based on the weights of each vertex and the positions of adjacent joint points, thus achieving a smooth deformation effect.

[0079] After obtaining the SMPL parameters and understanding the body pose conversion process, the reconstruction of a refined 3D human action sequence can be started. The specific steps are as follows:

[0080] Initialize the SMPL model:

[0081] Use a preset SMPL model template, which contains information such as the mesh vertices and joint point positions of the human model.

[0082] Apply the shape parameters:

[0083] Deform the SMPL model according to the shape parameters to match the body shape characteristics of the target individual.

[0084] Apply the pose parameters:

[0085] Adjust the pose of the SMPL model according to the pose parameters in the time series to generate a continuous 3D human action sequence.

[0086] At each time step, the pose parameters of the current time step are used to update the model.

[0087] Skin and generate the final model:

[0088] Apply the skinning technique to the deformed shape and pose at each time step to generate the final 3D human model.

[0089] Connect these models in time series to obtain a refined 3D human action sequence.

[0090] A refined 3D human action sequence generation system based on the SMPL model, as Figure 2 shown, the system includes:

[0091] A data processing module, which is used to collect three-dimensional human motion data, preprocess it to obtain preprocessed data, and construct a data set based on the preprocessed data;

[0092] A model construction and training module, which is used to construct an SMPL model and an adversarial network model, and train the adversarial network model using the data set to obtain a parameter-optimized model;

[0093] A parameter generation module, which is used to preprocess the three-dimensional human motion data obtained in real time and input it into the parameter-optimized model to generate SMPL parameters;

[0094] An action output module, which is used to reconstruct a refined three-dimensional human motion sequence according to the SMPL parameters by using the body pose conversion process of the SMPL model.

[0095] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description in the method part.

[0096] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A refined 3D human motion sequence generation method based on SMPL model, characterized in that: The following steps are involved: Collecting and preprocessing three-dimensional human motion data to obtain preprocessed data, and constructing a data set based on the preprocessed data; Constructing an SMPL model and an adversarial network model, and using the data set to train the adversarial network model to obtain a parameter optimization model; Preprocessing the three-dimensional human motion data acquired in real time and inputting the data into the parameter optimization model to generate SMPL parameters; According to the SMPL parameters, a refined three-dimensional human motion sequence is reconstructed using the posture conversion process of the SMPL model.

2. According to claim 1, a method for generating a refined three-dimensional human motion sequence based on the SMPL model is characterized in that: The adversarial network model includes a generator and a discriminator; The generator includes a first input layer, a first convolutional layer, an LSTM layer, a first fully connected layer and a first output layer connected in sequence; the discriminator includes an input layer, a convolutional layer, a fully connected layer and an output layer connected in sequence; the first output layer is connected to the input layer.

3. According to claim 2, a method for generating a refined three-dimensional human motion sequence based on the SMPL model is characterized in that: The generator loss function L G The loss function L1 of three-dimensional key point, the loss function L2 of two-dimensional key point, the loss function L3 of identity posture parameter and the loss function L adv The composition is as follows: L G =λ1L1+λ2L2+λ3L3+λ adv L adv ; In the formula, λ1, λ2, λ3, λ adv They are the weight parameters of the three-dimensional key point loss function, the two-dimensional key point loss function, the identity posture parameter loss function, and the adversarial loss function.

4. The method for generating a refined three-dimensional human motion sequence based on the SMPL model according to claim 3, characterized in that: The formula of the three-dimensional key point loss function L1 is as follows: Where N is the number of three-dimensional key points, is the i-th 3D key point generated by the generator, P i is the true i-th 3D key point.

5. The method for generating a refined three-dimensional human motion sequence based on the SMPL model according to claim 3, characterized in that: The formula of the two-dimensional key point loss function L2 is as follows: Where M is the number of two-dimensional key points, is the j-th two-dimensional key point generated by the generator, Q j is the real j-th two-dimensional key point.

6. The method for generating a refined three-dimensional human motion sequence based on the SMPL model according to claim 3, characterized in that: The formula of the identity posture parameter loss function L3 is as follows: In the formula, is the pose parameter generated by the generator, θ true is the actual posture parameter.

7. The method for generating a refined three-dimensional human motion sequence based on the SMPL model according to claim 3, characterized in that: The formula of the identity posture parameter loss function L4 is as follows: L4 = -logD(G(Z)); Where D(G(Z)) is the probability of the discriminator generating samples.

8. The method for generating a refined three-dimensional human motion sequence based on the SMPL model according to claim 2, characterized in that: The discriminator loss function L D Built on top of the cross entropy loss.

9. A system for generating a refined three-dimensional human motion sequence based on an SMPL model, used to implement a method for generating a refined three-dimensional human motion sequence based on an SMPL model as claimed in any one of claims 1 to 8, characterized in that: include: A data processing module, used for collecting and preprocessing three-dimensional human motion data to obtain preprocessed data, and constructing a data set based on the preprocessed data; A model building and training module is used to build an SMPL model and an adversarial network model, and train the adversarial network model using the data set to obtain a parameter optimization model; A parameter generation module, used for preprocessing the three-dimensional human motion data acquired in real time and inputting the data into the parameter optimization model to generate SMPL parameters; The motion output module is used to reconstruct a refined three-dimensional human motion sequence based on the SMPL parameters and using the posture conversion process of the SMPL model.