Training Method for 3D Human Generation Model Based on Single-View Tai Chi Movement Images
Through the training method of a three-dimensional human body generation model based on single-view Tai Chi action images, using SMPL model and deep learning technology, a three-dimensional human body model with similar body shape to the target user and similar posture to the standard athlete is generated, which solves the problem of poor accuracy in depth and three-dimensional space of traditional two-dimensional image analysis methods, and achieves high-accuracy three-dimensional action analysis and portable equipment requirements.
Patent Information
- Application Number
- CN202411266411.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-10
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-09-10
AI Technical Summary
The traditional two-dimensional image analysis method has poor accuracy in the depth and three-dimensional spatial display of Tai Chi movements, and the hardware equipment required for multi-view three-dimensional spatial display is complex.
Through the training method of a three-dimensional human body generation model based on single-view Tai Chi action images, SMPL model and deep learning technology are used to generate a three-dimensional human body model similar to the target user's body shape and the posture of a standard athlete.
It realizes the generation of accurate three-dimensional mannequin models without multiple viewing cameras, improving the accuracy of Tai Chi action analysis and the portability of the equipment.
Smart Images

Figure CN119180912B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision, and in particular, to a training method, system, computer device, and computer-readable storage medium for a three-dimensional human body generation model based on a single-view Tai Chi movement image. Background Art
[0002] Tai Chi, as a martial art with rich connotations, requires extremely high skills and coordination for each movement. Due to significant individual differences in the body shape, habits, and abilities of practitioners, traditional and unchanging standard movements are not always applicable. Therefore, it is particularly urgent to develop a method that can analyze and guide movements according to individual characteristics.
[0003] Traditional two-dimensional image analysis methods have obvious deficiencies in depth and three-dimensional space display, and too much hardware equipment is required for multi-view three-dimensional space display. Therefore, a portable method that can display multi-dimensional data based on a single view is very important. Summary of the Invention
[0004] Embodiments of the present application provide a training method, system, computer device, and computer-readable storage medium for a three-dimensional human body generation model based on a single-view Tai Chi movement image, so as to at least solve the problem of poor accuracy in depth and three-dimensional space display in the method of analyzing Tai Chi movements based on two-dimensional images in related technologies.
[0005] In a first aspect, embodiments of the present application provide a training method for a three-dimensional human body generation model based on a single-view Tai Chi movement image, and the method includes:
[0006] Obtain a training data set of Tai Chi movements, and preprocess the training data set to obtain body shape description information, where the training data set includes a template data set of standard athletes and an example data set of target users;
[0007] Based on each two-dimensional image in the training data set and its corresponding body shape description information, train a body shape parameter extraction model, where the trained body shape parameter extraction model is used to obtain corresponding two-dimensional body shape parameters according to the two-dimensional image;
[0008] Based on each two-dimensional image in the training data set, train a joint point reconstruction model, where the trained joint point reconstruction model is used to obtain corresponding three-dimensional pose parameters according to the two-dimensional image;
[0009] Input the two-dimensional body shape parameters and the three-dimensional pose parameters into the SMPL model to generate a three-dimensional human body model that is similar to the body shape of the target user and has a pose similar to that of the standard athlete.
[0010] In some of these embodiments, preprocessing the training dataset to obtain body shape description information includes:
[0011] By analyzing and processing each two-dimensional image in the training dataset, body shape description information of the standard athlete and the target user is obtained respectively, where the body shape description information includes: initial pose parameters and initial mesh vertices.
[0012] Among them, the initial pose parameters include variables of n×m dimensions, which are used to describe the action pose of the human body.
[0013] In some of these embodiments, training a body shape parameter extraction model according to each two-dimensional image and its corresponding body shape description information in the training dataset includes:
[0014] Construct a neural network model with the input of two-dimensional images from a preset perspective and the output of human body shape parameters, where the neural network model is connected to the SMPL model;
[0015] Based on the two-dimensional images, perform iterative optimization training on the neural network model. When the iterative optimization training reaches a preset threshold, obtain the trained body shape parameter extraction model;
[0016] Among them, the neural network model uses the initial mesh vertices corresponding to the two-dimensional images as labels, and determines the loss function based on the difference between the initial mesh vertices and the mesh vertices in the human body three-dimensional model output by the SMPL model.
[0017] In some of these embodiments, training a joint point reconstruction model based on each two-dimensional image in the training dataset includes:
[0018] Construct a deep learning model with the input of the two-dimensional images and the output of the three-dimensional pose parameters;
[0019] Taking predicting the coordinates of each joint point by regression based on multiple two-dimensional images as the learning task, train the deep learning model. When the number of iterations reaches a preset threshold, obtain the trained joint point reconstruction model;
[0020] Among them, the joint point reconstruction model extracts visual features in the two-dimensional images through a convolutional neural network, and enhances the visual features through a feature pyramid network.
[0021] Map the enhanced visual features to three-dimensional joint point coordinates through a pose regression network. The pose regression network includes multiple fully connected layers, which are used to process the correlation characteristics between the joints through the interaction of multiple fully connected layers.
[0022] In some of these embodiments, inputting the two-dimensional body type parameters and the three-dimensional pose parameters into the SMPL model to generate a three-dimensional human body model that is similar to the target user's body type and similar to the standard athlete's pose includes:
[0023] Based on the body type description information of the standard athlete and the target user respectively, obtain the distances between each joint point between the target user and the standard athlete, and determine the scaling ratio between each joint point between the standard athlete and the target athlete based on the distances between each joint point;
[0024] According to each of the scaling ratios, scale the three-dimensional pose parameters of the standard athlete to obtain optimized three-dimensional pose parameters;
[0025] Through the SMPL model, generate the three-dimensional human body model of the target user according to the optimized three-dimensional pose parameters and the two-dimensional body type parameters of the target user.
[0026] In a second aspect, an embodiment of the present application provides a single-view Tai Chi movement evaluation method, and the method includes:
[0027] Based on the Tai Chi movement images of the target user, construct a dataset to be evaluated, where the dataset to be evaluated includes a plurality of two-dimensional Tai Chi images and their corresponding movement type labels;
[0028] Train an RNN model according to the dataset to be evaluated to obtain a Tai Chi movement classification model;
[0029] Based on the dataset to be evaluated and the template dataset of the standard athlete, perform model training through the model training method of claim 1 to obtain a three-dimensional human body model;
[0030] Through the Tai Chi movement classification model, based on the dataset to be evaluated, obtain the image classification result corresponding to the two-dimensional image;
[0031] According to the two-dimensional image, its image classification result, and the three-dimensional human body model, obtain the Tai Chi movement evaluation result of the user to be evaluated.
[0032] In some of these embodiments, obtaining the Tai Chi movement evaluation result of the user to be evaluated according to the two-dimensional image, its image classification result, and the three-dimensional human body model includes:
[0033] Based on the image classification result, within the target time period corresponding to any Tai Chi movement classification, obtain the three-dimensional pose parameters of the user to be evaluated;
[0034] According to the three-dimensional human body model of the user to be evaluated, obtain the reference pose data within the target time period;
[0035] Optimize the three-dimensional pose parameters of the user to be evaluated through frame sampling or frame interpolation;
[0036] Compare the joint point coordinates of each image frame in the three-dimensional pose parameters of the user to be evaluated after optimization with the joint point coordinates of each image frame in the reference pose data to obtain the Tai Chi movement evaluation result of the user to be evaluated.
[0037] In a third aspect, an embodiment of the present application provides a training system for a three-dimensional human body generation model based on single-view Tai Chi movement images, and the system includes: an acquisition module, a parameter extraction module, and a human body model generation module;
[0038] The acquisition module is configured to acquire a training data set of Tai Chi movements, preprocess the training data set to obtain body shape description information, where the training data set includes a template data set of standard athletes and an example data set of target users;
[0039] The parameter extraction module is configured to train a body shape parameter extraction model based on each two-dimensional image and its corresponding body shape description information in the training data set, where the trained body shape parameter extraction model is used to obtain corresponding two-dimensional body shape parameters according to the two-dimensional image,
[0040] and, based on each two-dimensional image in the training data set, train a joint point reconstruction model, where the trained joint point reconstruction model is used to obtain corresponding three-dimensional pose parameters according to the two-dimensional image;
[0041] The human body three-dimensional model generation module is configured to input the two-dimensional body shape parameters and the three-dimensional pose parameters into the SMPL model to generate a human body three-dimensional model that is similar to the body shape of the target user and similar to the pose of the standard athlete.
[0042] In a fourth aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the method described in the first aspect above is implemented.
[0043] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in the first aspect above is implemented.
[0044] Compared with the related art, a training method for a three-dimensional human body generation model based on single-view Tai Chi action images provided by an embodiment of the present application uses a single-view based Tai Chi action analysis method. By introducing cutting-edge mathematical models and algorithms, a highly realistic three-dimensional digital human model is generated synchronously. This method has achieved remarkable breakthroughs in both the accuracy of action analysis and the realism of digital human generation. It can generate a three-dimensional human body model without the need for multiple-view cameras, reducing the requirements for equipment and venues and improving the accuracy of Tai Chi action analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0046] Figure 1 is a flowchart of a training method for a three-dimensional human body generation model based on single-view Tai Chi action images according to an embodiment of the present application;
[0047] Figure 2 is a flowchart of a single-view based Tai Chi action evaluation method according to an embodiment of the present application;
[0048] Figure 3 is a structural block diagram of a training system for a three-dimensional human body generation model based on single-view Tai Chi action images according to an embodiment of the present application;
[0049] Figure 4 is an internal structural schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be described and explained below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided by the present application without creative efforts fall within the scope of protection of the present application.
[0051] Obviously, the accompanying drawings in the following description are only some examples or embodiments of the present application. For those of ordinary skill in the art, without creative efforts, the present application can also be applied to other similar scenarios based on these drawings. In addition, it can also be understood that although the efforts made in such a development process may be complex and time-consuming, for those of ordinary skill in the art related to the content disclosed in the present application, some design, manufacturing, or production changes based on the technical content disclosed in the present application are only conventional technical means and should not be understood as the content disclosed in the present application being insufficient.
[0052] Reference to "embodiments" in the present application means that the specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict.
[0053] Unless otherwise defined, the technical terms or scientific terms involved in the present application should have the ordinary meaning understood by those of ordinary skill in the technical field to which the present application belongs. The words "a", "an", "one", "the", and similar words involved in the present application do not indicate a quantity limitation and can represent a singular or plural number. The terms "including", "comprising", "having", and any variations thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include unlisted steps or units, or may further include other steps or units inherent to these processes, methods, products, or devices. The terms "connected", "coupled", and similar words involved in the present application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The term "plurality" involved in the present application refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The terms "first", "second", "third", etc. involved in the present application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0054] In the analysis of Tai Chi movements, the diversity of human body shapes is a factor that cannot be ignored. For individuals with special body types, simply applying a standard movement template (such as the demonstration movements of Tai Chi world champions) often leads to distorted analysis results. To solve this problem, this application constructs a three-dimensional human body generation model that can convert the body postures of standard athletes into those of users, in order to analyze Tai Chi movements under the same body postures.
[0055] In a first aspect, an embodiment of this application provides a training method for a three-dimensional human body generation model based on single-view Tai Chi movement images. Figure 1 It is a flowchart of a training method for a three-dimensional human body generation model based on single-view Tai Chi movement images according to an embodiment of this application, as Figure 1 shown. The method includes the following steps:
[0056] S101, Obtain a training data set of Tai Chi movements, preprocess the training data set to obtain body shape description information, where the training data set includes a template data set of standard athletes and an example data set of target users;
[0057] Among them, the template data set of standard athletes should contain standardized Tai Chi movements, and the key point data of each movement should be as accurate as possible; while the example data set of target users can be recorded by any Tai Chi enthusiast by self-demonstrating or simulating Tai Chi movements.
[0058] In addition, data preprocessing is a process of extracting fixed information from the original data by a scanner for subsequent analysis and modeling. In this embodiment, the preprocessing steps specifically include:
[0059] S1, Data cleaning
[0060] Missing data processing: Check whether there are missing values in the data set. If there are missing values, measures such as filling in or deleting need to be taken.
[0061] Noise processing: Remove the noise in the data, such as incorrect key point coordinates or irrelevant movement segments.
[0062] S2, Data alignment
[0063] Time alignment: The movements of different users may not be synchronized in time, so time alignment is required. The Dynamic Time Warping (DTW) algorithm can be used to align the time series data of the movements.
[0064] Spatial alignment: The movements of standard athletes and those of target users may not be consistent in terms of spatial position. Normalization can be performed to compare movement data on a unified scale. For example, joint coordinate data can be centered (relative to the body's center point) or normalized.
[0065] S3. Analyze and obtain pose and body shape information
[0066] By analyzing each two-dimensional image in the training dataset, body shape description information is obtained, where the body shape description information includes: initial pose parameters and initial mesh vertices;
[0067] Specifically, pose estimation algorithms (such as OpenPose, HRNet, etc.) can be used to extract the two-dimensional coordinates of human key points from video frames and output the three-dimensional joint point coordinates of the human body in each frame, that is, the pose parameters θ;
[0068] Among them, the initial pose parameters θ have 24 × 3 dimensions and are used to describe the movement pose of the human body; among them, the number of joint points is 24, including 23 joint points and 1 root node: each joint point is described by 3-dimensional variables, and these 3 variables are usually rotation angles (such as Euler angles, quaternions, etc.) and are used to describe the rotation pose of this joint point relative to the root node.
[0069] Furthermore, the root node is the base point of the human body model, and its coordinates are the real positions relative to the entire space. The root node is usually located at the waist or pelvis of the human body and is the starting point of the entire skeleton system; other joint points except the root node are described by the coordinate system relative to their parent nodes. The pose (rotation) parameters of each joint point determine the position of this joint point in three-dimensional space.
[0070] Parent-child node relationship: Joint points are connected through a parent-child relationship, and each joint point inherits the coordinate system of its parent node. For example, if you have the pose parameters of the shoulder joint, it is relative to the root node (usually the pelvis). The position of the forearm joint is relative to the position of the upper arm joint, and so on.
[0071] In addition, if there is three-dimensional scan data, the vertex coordinates of the mesh can be directly extracted through three-dimensional mesh processing methods to describe the body shape. The initial mesh vertex coordinates obtained by scanning can describe the geometric body shape of the human body. These vertices are three-dimensional coordinates and usually form the outer surface of the human body; further, the human body surface is connected by these vertices through polygons (usually triangles or quadrilaterals) to form a three-dimensional mesh model.
[0072] It can be understood that through vertex coordinates and the mesh, the surface shape and size of the human body can be accurately described. Further, combined with the pose parameters θ, dynamic body shape description and motion simulation can be realized.
[0073] S102. Based on each two-dimensional image in the training dataset and its corresponding body shape description information, train a body shape parameter extraction model, where the trained body shape parameter extraction model is used to obtain the corresponding two-dimensional body shape parameters according to the two-dimensional image;
[0074] In this embodiment, the function of this step is to train a model for extracting body shape parameters. Specifically, this step includes the following sub-steps:
[0075] S1. Construct a neural network model with the input of two-dimensional images from a preset perspective and the output of its corresponding human body shape parameters, where the neural network model is configured to be connected to the SMPL model;
[0076] Among them, the function of step S1 is to design a neural network model for mapping human two-dimensional images to three-dimensional body shape parameters;
[0077] In an exemplary embodiment, the neural network model includes the following structure:
[0078] a. Convolutional Neural Network (CNN): Usually use CNN as the backbone network to extract features in the two-dimensional image, and these features may include contours, joint points, muscle distributions, etc.;
[0079] b. Fully connected layer: After the convolutional layer, use the fully connected layer to map the extracted features to the predicted values of three-dimensional body shape parameters;
[0080] c. Regression layer: Used to finally output body shape parameters, and usually predict continuous body shape variables through regression.
[0081] S2. Based on the two-dimensional image, perform iterative optimization training on the neural network model. When the iterative optimization training reaches a preset threshold, obtain the trained body shape parameter extraction model;
[0082] Among them, the neural network model uses the initial mesh vertices corresponding to the two-dimensional image as labels, and the difference between the initial mesh vertices and the mesh vertices in the human three-dimensional model output by the SMPL model as the loss function.
[0083] It can be understood that the loss function is designed as the difference between the initial mesh vertices (estimated based on the two-dimensional image) and the mesh vertices of the three-dimensional model generated by SMPL, which means that the training process aims to minimize the error between the human shape estimated from the two-dimensional image and the three-dimensional human shape generated by the SMPL model;
[0084] Optionally, a suitable optimization algorithm such as Adam or SGD can be used to minimize the loss function and update the model parameters; by repeatedly iterating to adjust the weights of the neural network model, the output body shape parameters can generate a 3D model that is closer to the actual shape of the 2D image. When the iterative optimization training reaches a preset threshold (e.g., the loss function converges to a relatively small value), a trained body shape parameter extraction model is obtained.
[0085] In an exemplary embodiment, the training process can divide the dataset into a training set, a validation set, and a test set. Usually, 80% is used for training, 10% for validation, and 10% for testing. In addition, the model is trained in multiple epochs, and the model weights are updated by the optimizer in each iteration; after each epoch, the validation set is used to evaluate the model performance, and the model hyperparameters (such as the learning rate, network structure, etc.) are adjusted according to the validation results.
[0086] Through the above step S102, a body shape parameter extraction model can be trained based on the preprocessed training dataset. This model can obtain the body shape parameters (β) of any standard athlete and the target experience user. These body shape parameters describe the individual differences of the human body, such as body shape, height, fatness, etc., and are usually represented by a 10-dimensional vector. This parameter controls the basic shape of the human body mesh.
[0087] S103, based on each 2D image in the training dataset, train a joint point reconstruction model, where the trained joint point reconstruction model is used to obtain the corresponding 3D pose parameters according to the 2D image;
[0088] It can be understood that the human body is a rigid object, and the distances between the joint points are fixed. Therefore, 3D data can be fitted based on the 2D human body image. In this embodiment, the model is trained with 2D images from any perspective as the input and 3D pose parameters as the output, so that the model fits 3D pose parameters based on the 2D image;
[0089] Specifically, this step includes the following sub-steps:
[0090] S1, construct a deep learning model with 2D images as the input and 3D pose parameters as the output;
[0091] Among them, the input data comes from the 2D images collected in step S101. These images are usually human body images from a certain perspective and contain different poses;
[0092] Furthermore, the label data is the 3D joint point coordinates corresponding to the input images, and these coordinates are also obtained through the preprocessing steps in step S101.
[0093] S2. Use multiple two-dimensional images to predict the coordinates of each joint point through regression, train the joint point reconstruction model, and obtain the trained joint point reconstruction model when the number of iterations reaches a preset threshold.
[0094] First, the joint point reconstruction model extracts visual features from the two-dimensional image through a convolutional neural network. The convolutional layer can identify edges, contours, and complex shape features in the image.
[0095] Second, through the Feature Pyramid Network (FPN) or Residual Network (ResNet): further extract multi-scale features and enhance the feature expression ability of the model.
[0096] Furthermore, the enhanced visual features are mapped to three-dimensional joint point coordinates through a pose regression network. It should be noted that this pose regression network maps the extracted features to three-dimensional joint point coordinates through a series of fully connected layers or RNNs (such as LSTM); this part of the network needs to handle the correlation between joint points and ensure that the output coordinates are reasonable.
[0097] Finally, through the output layer, the three-dimensional joint point coordinates are output. Usually, the regression method is used to predict the x, y, and z coordinates of each joint point.
[0098] In addition, a suitable optimization algorithm, such as Adam or RMSprop, can be used to minimize the loss function and optimize the model parameters. Among them, the difference between the measured predicted joint point coordinates and the true coordinates can be used to determine the loss function during model training.
[0099] Through the above step S103, the joint point reconstruction model can be trained based on the preprocessed training dataset. This model can obtain the pose parameters (θ) of any standard athlete and the target experience user during the entire Tai Chi process. This pose parameter describes the coordinates of each joint point of the human body during movement and is used to represent the overall pose.
[0100] S104. Input the two-dimensional body shape parameters and three-dimensional pose parameters into the SMPL model to generate a three-dimensional human model that is similar to the target user's body shape and similar to the standard athlete's pose.
[0101] Among them, the SMPL (Skinned Multi-Person Linear) model is a three-dimensional human model widely used in computer graphics and computer vision. It can represent various human postures and shapes through a parametric model, so as to be able to generate three-dimensional human meshes with different postures and body shapes.
[0102] The SMPL model is based on the human skeleton and consists of 24 joint points (including the root node and 23 other joint points), corresponding to the pose parameters mentioned in step S201. Secondly, the model uses a three-dimensional mesh with a fixed topology to represent the human body, usually consisting of 6890 vertices.
[0103] Furthermore, the SMP model controls the rotation of the human skeleton through pose parameters (θ), thereby affecting the relative positions of each joint point, and controls the basic shape of the human mesh through shape parameters (β).
[0104] It should be noted that due to the differences in human body shapes, the body shape parameters and pose parameters obtained under the same action will inevitably be different. Therefore, it is necessary to fine-tune the data of standard athletes so that the human body postures obtained through the SMPL model are suitable for the users, in order to provide better and more professional motion analysis.
[0105] Based on the above viewpoints, in the solution of this application, the three-dimensional pose information of standard athletes in the Taijiquan process is utilized, combined with the body shape information of the target user to create a three-dimensional model, and a three-dimensional human body model similar to the body shape of the target user and similar to the pose of the standard athlete is obtained.
[0106] It should be noted that considering that the pose parameters obtained from different body postures are also different, the positions of the joint points are relative, that is, the coordinates of each joint point are defined relative to its parent joint point. This means that if the body shapes of the user and the athlete are different (such as different leg lengths), then using the pose parameters of the athlete may cause the user's movements to look unnatural or distorted. Therefore, if the pose parameters of the athlete are directly combined with the body shape parameters of the target user, it will lead to deformation of the movements.
[0107] In this embodiment, the relative relationship between the joint points of the athlete is replaced with the relative relationship of the joint points of the user, and each joint point is scaled, and the scaling ratio is the ratio of the lengths of the joint points of the two.
[0108] Specifically, based on the body shape description information of the standard athlete and the target user respectively, the distances between each joint point between the target user and the standard athlete are obtained, and multiple joint scaling ratios are determined based on the distances between each joint point.
[0109] Furthermore, according to multiple joint scaling ratios, the three-dimensional pose parameters of the standard athlete are scaled to obtain optimized three-dimensional pose parameters, and the optimized three-dimensional pose parameters and the two-dimensional body shape parameters of the target user are used to obtain the three-dimensional human body model of the target user.
[0110] It can be understood that through the above steps, by using this scaling adjustment mechanism, the actions of the target user will be more natural and conform to the actual body posture, rather than showing uncoordinated actions due to body size differences.
[0111] Since this three-dimensional human body model is similar to the specific user in terms of body size, however, each action posture in the Taichi movement process is collected from the standard template. Therefore, this three-dimensional human body model can effectively make up for the differences between individuals, create a three-dimensional human body model applicable to different body sizes based on a single-view image, and achieve the creation of a three-dimensional human body model without the need to use multi-view hardware devices.
[0112] In an exemplary embodiment, the specific implementation method of constructing a three-dimensional human body model through the SMPL model includes the following steps:
[0113] Each body shape form constructed by the SMPL model is based on a linear combination of the base template. The coordinates of each mesh vertex under different body shape parameters β are set as , and its dimension is (6890, 3, 10); by linearly superimposing the standard base template, a human body posture under the current body shape can be obtained. The mathematical formula is as follows:
[0114]
[0115] Among them, is the obtained linear superposition offset, which represents the influence of each body shape parameter β on the human body mesh vertex; each represents the base mesh deformation under the body shape parameter. By superimposing these deformations, the mesh vertices under specific body shape parameters can be obtained, and a new human body posture can be obtained.
[0116] It can be understood that the base template refers to a standard human body mesh shape, and the linear combination means generating different body shapes through the linear combination of several predefined deformations (or called "shape vectors").
[0117] Furthermore, the pose parameter θ describes the poses of different joints through the rotation matrix R(θ). The adjustment of the pose is achieved through a linear evaluation function:
[0118]
[0119] Among them, is the rotation matrix generated by the pose parameter, is the weight matrix, which defines the influence of pose changes on the mesh vertices. Through this process, the mesh deformation based on the user's pose parameters can be obtained;
[0120] Each human body pose parameter is represented by a rotation matrix R. k = 23 is the number of joint points, the axis angle is 3, the rotation dimension of all axis angles is 3K, and the dimension after conversion to a rotation matrix is 9K. p is the weight matrix with a dimension of (6890, 3, 207). Similar to the human body pose, what is obtained is also a linear superposition offset in the base state, and the actual pose also needs to be obtained by superimposing the base pose.
[0121] Specifically, the axis angle rotation formula is:
[0122]
[0123] Among them, is the axis angle.
[0124] In addition, when the human body is in motion, the network vertices will move with the movement of the joints, and the linear skinning algorithm needs to be used to linearly weight and combine the mesh vertices.
[0125]
[0126] Among them, is the average mesh vertex of the human body, is the human body mesh vertex of the base template.
[0127] Through the linear skinning algorithm, the mesh vertices can be further adjusted to adapt to the movement of the joints, and the regression matrix between the network vertices and 24 joint points can be obtained, as well as the mixed weight matrix of linear skinning, so that the final relational expression can be obtained:
[0128]
[0129] This relational expression combines the regression matrix F of the mesh vertices and joint points and the mixed weight matrix W to generate the final three-dimensional human body model.
[0130] Through the above step S104, a three-dimensional human body model consistent with the user's body posture and consistent with the standard athlete's action posture can be constructed to ensure that the athlete's actions are more natural and accurate on the user.
[0131] Through the above steps S101 to S104, compared with the methods in the traditional technology, there is a problem of insufficient depth perception in analyzing Tai Chi movements based on two-dimensional images, while three-dimensional motion capture systems often require complex hardware support. This application takes a single-view Tai Chi movement image as input, and uses an innovative algorithm to generate a three-dimensional digital human in a three-dimensional perspective while maintaining the low cost and portability of the device, providing accurate analysis of Tai Chi movements; this method extracts joint position information from two-dimensional images, and based on the SMPL (Skinned Multi-Person Linear) model, matches the standard movements of Tai Chi world champions with the user's body shape, enabling motion analysis to be carried out under the same body posture, thus avoiding motion errors caused by body shape differences; it can automatically adjust the analysis algorithm according to the user's body shape, habits and abilities to provide more personalized guidance.
[0132] In a second aspect, an embodiment of the present application further provides a method for evaluating Tai Chi movements based on a single view, Figure 2 which is a flowchart of the method for evaluating Tai Chi movements based on a single view according to an embodiment of the present application, as Figure 2 shown. The process includes the following steps:
[0133] S201, based on the Tai Chi movement images of the target user, construct a dataset to be evaluated, where the dataset to be evaluated includes a plurality of two-dimensional Tai Chi images and their corresponding action type labels;
[0134] The target user is the target user in the above steps S101 to S103. The body shape, gender, and familiarity with Tai Chi movements of the target user have no impact on the solution of the present application and are not specifically limited in this embodiment;
[0135] S202, train an RNN model according to the dataset to be evaluated to obtain a Tai Chi movement classification model; based on the dataset to be evaluated and the template dataset of standard athletes, perform model training through the model training method provided in the above embodiment to obtain a human body three-dimensional model;
[0136] Among them, the input of the model is the two-dimensional image data of consecutive frames, and the output is the confidence of 25 action categories corresponding to each frame. Specifically, 24 are standard Tai Chi movements and 1 is a non-Tai Chi movement.
[0137] In this embodiment, an RNN model is used to process time series data. Since image data usually requires high-dimensional feature extraction, a CNN (Convolutional Neural Network) is used to extract the features of each frame first, and then the feature sequence is input into an RNN (such as LSTM or GRU).
[0138] The network structure of the convolutional neural network may include: a CNN layer for extracting features of each frame of image; an RNN layer for processing the extracted feature sequence; and a fully connected layer for outputting the classification confidence of the action.
[0139] When the model recognizes an action, record the timestamp of the current frame as the action end time, and the action start time can be set as the timestamp of the first frame stored in X. Additionally, if a more accurate time point is needed, the start time can be considered to be adjusted through the sliding window technique or based on the confidence threshold.
[0140] Finally, the specific generation process of the three-dimensional human model of the target user is the same as the above steps S101 to S103, which will not be elaborated here.
[0141] S203, through the tai chi action classification model, based on the dataset to be evaluated, obtain the image classification result corresponding to the two-dimensional image;
[0142] Among them, this step is interrelated with the model construction in the above step S202. It can be understood that the primary task of the action classification model is to identify the specific action that the user is performing. Since tai chi contains multiple standard actions, different actions have different postures and action trajectories. By classifying the actions, the system can determine which action the user is currently performing.
[0143] Only by accurately identifying the type of action performed by the user can the system perform subsequent action evaluation and feedback. This is also an important prerequisite for evaluating the standardization of the user's actions, and it divides the time period corresponding to each action, and provides data support for selecting the image frames for comparison during this time period;
[0144] That is, the classification model can help the system understand the specific type of the user's action, and according to the classification result, extract the corresponding standard action data for comparison. In action evaluation, the system needs to compare a certain category of the user's actions with the standard actions. If the classification is inaccurate, the user's actions may be compared with the wrong standard actions, resulting in inaccurate evaluation results. Therefore, the accuracy of action classification is directly related to the accuracy of action evaluation.
[0145] In addition, the action classification model can also distinguish between tai chi actions and non-tai chi actions. For example, when the user performs other irrelevant movements, the system can identify and ignore these non-tai chi actions through the classification model. This helps to filter out irrelevant action data and ensure that the system only evaluates tai chi-related actions, improving the efficiency and accuracy of the system.
[0146] S204, according to the tai chi action image set, its corresponding image classification result, and the three-dimensional human model, obtain the tai chi action evaluation result of the user to be evaluated.
[0147] Specifically, step S205 includes the following sub-steps:
[0148] Step1, based on the image classification result, within the target time period corresponding to any one type of Tai Chi movement classification, obtain the three-dimensional pose parameters of the user to be evaluated;
[0149] Step2, according to the three-dimensional human model of the user to be evaluated, obtain the reference pose data within the target time period;
[0150] Step3, optimize the three-dimensional pose parameters of the user to be evaluated through frame sampling or frame interpolation;
[0151] Step4, by comparing the joint point coordinates of each image frame in the optimized three-dimensional pose parameters of the user to be evaluated with the joint point coordinates of each image frame in the reference three-dimensional pose parameters, obtain the evaluation result of the Tai Chi movement of the user to be evaluated.
[0152] Through the above steps Step1 - Step3, according to the start time and end time recorded for each action category, extract the data for the corresponding time period from the joint point data of the target user. However, since the action start time may not be accurate, the start time can be corrected by moving the previous and next frames.
[0153] In addition, sample or interpolate the joint point data within the time period to make its length consistent with the standard action duration; assume that the standard action of the athlete within this time period consists of 50 frames of joint point data, but the actual action of the user consists of 40 frames of joint point data. To compare the action standardization of the two, the 40-frame data of the user needs to be adjusted to 50 frames, so that the actions of the user and the athlete can be compared frame by frame.
[0154] If the number of frames of the user is more than that of the standard action, some frames need to be evenly selected from the user's frame data to reduce the data length; if the number of frames of the user is less than that of the standard action, some frames may need to be deleted from the standard action, or the user's data may need to be interpolated and extended to make its number of frames consistent with the standard action.
[0155] If the number of frames of the user's action is less than that of the standard action, interpolation methods can be used to generate additional frames to make the length of the user's action sequence consistent with the standard action. Interpolation can be achieved through linear interpolation or more complex time series interpolation methods.
[0156] Through this adjustment and optimization, the joint point data of the user and the standard action data of the athlete have the same time length, which is convenient for frame-by-frame comparison and difference analysis, and then to judge the standardization of the user's actions.
[0157] Through the above steps S201 to S203, by means of precise three-dimensional pose comparison, it is evaluated whether the user's Tai Chi movements meet the standards. This evaluation is based on the dual dimensions of time and space. By comparing the joint point coordinates frame by frame, the subtle deviations in the user's movements can be accurately captured, helping the user continuously optimize the movement standardization during practice. Finally, the system provides the user with a data-driven and precise movement evaluation and improvement tool, which helps to improve the standard and effect of Tai Chi movements.
[0158] In a second aspect, the present embodiment also provides a training system for a three-dimensional human body generation model based on a single-view Tai Chi movement image. Figure 3 It is a structural block diagram of a training system for a three-dimensional human body generation model based on a single-view Tai Chi movement image according to an embodiment of the present application, as Figure 3 shown. The system includes: an acquisition module 30, a parameter extraction module 31, and a human body model generation module 32;
[0159] The acquisition module 30 is configured to acquire a training data set of Tai Chi movements, preprocess the training data set to obtain body shape description information, where the training data set includes a template data set of standard athletes and an example data set of target users;
[0160] The parameter extraction module 31 is configured to train a body shape parameter extraction model based on each two-dimensional image in the training data set and its corresponding body shape description information, where the trained body shape parameter extraction model is used to obtain the corresponding two-dimensional body shape parameters according to the two-dimensional image, and, based on each two-dimensional image in the training data set, train a joint point reconstruction model, where the trained joint point reconstruction model is used to obtain the corresponding three-dimensional pose parameters according to the two-dimensional image;
[0161] The human body three-dimensional model generation module 21 is configured to input the two-dimensional body shape parameters and the three-dimensional pose parameters into the SMPL model to generate a human body three-dimensional model that is similar to the target user's body shape and similar to the standard athlete's pose.
[0162] Through the above system, compared with the traditional technology, there is a problem of insufficient depth perception in analyzing Tai Chi movements based on two-dimensional images, while three-dimensional motion capture systems often require complex hardware support. This system uses single-view Tai Chi movement images as input and, with innovative algorithms, generates a three-dimensional digital human from a three-dimensional perspective while maintaining the low cost and portability of the device, providing accurate analysis of Tai Chi movements. This method extracts joint position information from two-dimensional images and, based on the SMPL (Skinned Multi-Person Linear) model, through computer graphics, deep neural networks, and motion analysis techniques, realizes accurate analysis and digital reproduction of Tai Chi movements. The generated three-dimensional digital human model has a high degree of authenticity and fidelity and can be applied to the research, teaching, and demonstration of Tai Chi movements.
[0163] This embodiment also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0164] Optionally, the above electronic device may further include a transmission device and an input / output device. Among them, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0165] Optionally, in this embodiment, the above processor may be configured to execute the following steps through a computer program:
[0166] S1. Obtain a training data set of Tai Chi movements, preprocess the training data set, and obtain body shape description information;
[0167] S2. Based on each two-dimensional image in the training data set and its corresponding body shape description information, train a body shape parameter extraction model;
[0168] S3. Based on each two-dimensional image in the training data set, train a joint point reconstruction model;
[0169] S4. Input the two-dimensional body shape parameters and three-dimensional pose parameters into the SMPL model to generate a three-dimensional human body model that is similar to the target user's body shape and similar to the standard athlete's pose.
[0170] It should be noted that the specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation manners, and will not be repeated here.
[0171] In one embodiment, Figure 4 is a schematic internal structure diagram of an electronic device according to an embodiment of the present application. As Figure 4 shown, an electronic device is provided. The electronic device may be a server, and its internal structure diagram may be as Figure 4As shown. The electronic device includes a processor, a network interface, an internal memory, and a non-volatile memory connected by an internal bus. Among them, the non-volatile memory stores an operating system, a computer program, and a database. The processor is used to provide computing and control capabilities. The network interface is used to communicate with external terminals through a network connection. The internal memory is used to provide an environment for the operation of the operating system. The computer program, when executed by the processor, implements a training method for a three-dimensional human body generation model based on single-view Tai Chi movement images. The database is used to store data.
[0172] Those skilled in the art can understand that Figure 4 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the electronic device to which the solution of this application is applied. Specifically, the electronic device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0173] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0174] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
Claims
1. A training method for generating a three-dimensional human body model based on single-view Tai Chi action images, characterized in that: The method comprises: Acquire a training data set of Tai Chi movements, and pre-process the training data set to obtain body shape description information, wherein the training data set includes a template data set of standard athletes and an example data set of target users; Based on each two-dimensional image in the training data set and its corresponding body shape description information, a neural network model is trained to obtain a body shape parameter extraction model, wherein a loss function used in training the neural network model is designed to be a difference between initial mesh vertices estimated based on the two-dimensional image and mesh vertices of the three-dimensional model generated by SMPL, and the trained body shape parameter extraction model is used to obtain corresponding two-dimensional body shape parameters according to the two-dimensional image; Based on each two-dimensional image in the training data set, a joint point reconstruction model is trained, wherein the trained joint point reconstruction model is used to obtain corresponding three-dimensional posture parameters according to the two-dimensional image; Inputting the two-dimensional body shape parameters and the three-dimensional posture parameters into the SMPL model to generate a three-dimensional human body model similar to the body shape of the target user and similar to the posture of the standard athlete includes: Based on the body shape description information of the standard athlete and the target user, respectively, the distances between the joints of the target user and the standard athlete are obtained, and based on the distances between the joints, the scaling ratios between the joints of the standard athlete and the target user are determined; Scaling the three-dimensional posture parameters of the standard athlete according to each of the scaling ratios to obtain optimized three-dimensional posture parameters; A three-dimensional human body model of the target user is generated through the SMPL model according to the optimized three-dimensional posture parameters and the two-dimensional body parameters of the target user.
2. The method according to claim 1, characterized in that Preprocessing the training data set to obtain body shape description information includes: By analyzing and processing each two-dimensional image in the training data set, the body shape description information of the standard athlete and the target user is obtained respectively, wherein the body shape description information includes: initial posture parameters and initial mesh vertices, The initial posture parameters include variables of n×m dimensions, which are used to describe the movement posture of the human body.
3. The method according to claim 2, characterized in that According to each two-dimensional image in the training data set and its corresponding body shape description information, training a body shape parameter extraction model includes: Constructing a neural network model with a two-dimensional image of a preset viewing angle as input and human body shape parameters as output, wherein the neural network model is connected to the SMPL model; Based on the two-dimensional image, the neural network model is iteratively optimized and trained, and when the iterative optimization training reaches a preset threshold, a trained body shape parameter extraction model is obtained; The neural network model uses the initial mesh vertices corresponding to the two-dimensional image as labels, and determines the loss function based on the difference between the initial mesh vertices and the mesh vertices in the three-dimensional human body model output by the SMPL model.
4. The method according to claim 1, characterized in that Based on each two-dimensional image in the training data set, training the joint point reconstruction model includes: Constructing a deep learning model whose input is the two-dimensional image and whose output is the three-dimensional posture parameter; The deep learning model is trained by predicting the coordinates of each joint point by regression based on the plurality of the two-dimensional images as a learning task, and a trained joint point reconstruction model is obtained when the number of iterations reaches a preset threshold; The joint reconstruction model extracts visual features from the two-dimensional image through convolutional neural networks and enhances the visual features through a feature pyramid network. The enhanced visual features are mapped into three-dimensional joint point coordinates through a posture regression network, and the posture regression network includes multiple fully connected layers, which are used to process the correlation characteristics between the joint points through the interaction of the multiple fully connected layers.
5. A Tai Chi action evaluation method based on a single perspective, characterized in that: The method comprises: Based on the Tai Chi action images of the target user, construct a data set to be evaluated, wherein the data set to be evaluated includes a plurality of two-dimensional Tai Chi images and their corresponding action type labels; Train the RNN model according to the data set to be evaluated to obtain a Tai Chi action classification model; Based on the dataset to be evaluated and the template dataset of a standard athlete, model training is performed by the model training method of claim 1 to obtain a three-dimensional human body model; Obtaining an image classification result corresponding to the two-dimensional image based on the data set to be evaluated by using the Tai Chi action classification model; Based on the two-dimensional image and its image classification result, as well as the three-dimensional human body model, the Tai Chi movement evaluation result of the user to be evaluated is obtained.
6. The method according to claim 5, characterized in that Based on the two-dimensional image and its image classification result, and the three-dimensional human body model, the Tai Chi action evaluation result of the user to be evaluated is obtained, including: Based on the image classification result, obtaining the three-dimensional posture parameters of the user to be evaluated within a target time period corresponding to any Tai Chi movement classification; Acquiring reference posture data within the target time period according to the three-dimensional human body model of the user to be evaluated; Optimizing the three-dimensional posture parameters of the user to be evaluated through frame sampling or frame interpolation; The joint point coordinates of each image frame in the three-dimensional posture parameters of the user to be evaluated after the optimization are compared with the joint point coordinates of each image frame in the reference posture data to obtain the Tai Chi movement evaluation result of the user to be evaluated.
7. A training system for generating a three-dimensional human body model based on single-view Tai Chi action images, characterized in that: The system comprises: an acquisition module, a parameter extraction module and a human body model generation module; The acquisition module is used to acquire a training data set of Tai Chi movements, pre-process the training data set, and obtain body shape description information, wherein the training data set includes a template data set of a standard athlete and an example data set of a target user; The parameter extraction module is used to train a neural network model to obtain a body shape parameter extraction model based on each two-dimensional image in the training data set and its corresponding body shape description information, wherein the loss function used in training the neural network model is designed to be the difference between the initial mesh vertices estimated based on the two-dimensional image and the mesh vertices of the three-dimensional model generated by SMPL, and the trained body shape parameter extraction model is used to obtain the corresponding two-dimensional body shape parameters according to the two-dimensional image, And, based on each two-dimensional image in the training data set, training a joint reconstruction model, wherein the trained joint reconstruction model is used to obtain corresponding three-dimensional posture parameters according to the two-dimensional image; The human body three-dimensional model generation module is used to input the two-dimensional body shape parameters and the three-dimensional posture parameters into the SMPL model to generate a human body three-dimensional model similar to the body shape of the target user and similar to the posture of the standard athlete, including: Based on the body shape description information of the standard athlete and the target user, respectively, the distances between the joints of the target user and the standard athlete are obtained, and based on the distances between the joints, the scaling ratios between the joints of the standard athlete and the target user are determined; Scaling the three-dimensional posture parameters of the standard athlete according to each of the scaling ratios to obtain optimized three-dimensional posture parameters; A three-dimensional human body model of the target user is generated through the SMPL model according to the optimized three-dimensional posture parameters and the two-dimensional body parameters of the target user.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Single-view-angle Tai Chi action analysis and assessment system based on artificial intelligence
CN112016497A
Motion capture method and system based on parameterized model
CN117541646A
Deep learning-based shadowboxing action scoring method, storage medium and electronic equipment
CN118053201A