A Video 3D Human Pose Estimation Method and System Based on Multi-Level Supervised Graph Convolution
Through the video three-dimensional human posture estimation method based on multi-level supervised graph convolution, the problems of human joint depth blur and self-occlusion in the video sequence are solved, the estimation accuracy and model flexibility are improved, and efficient three-dimensional human posture estimation is achieved.
Patent Information
- Application Number
- CN202210387182.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-14
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-04-14
AI Technical Summary
The existing three-dimensional human posture estimation method of videos is not effective when facing the depth blur and self-occlusion problems of human joints in video sequences, and the prediction accuracy is not high in complex action environments.
A video three-dimensional human pose estimation method based on multi-level supervised graph convolution is proposed. By acquiring video data and inputting it into the trained multi-level supervised graph convolution model, the three-dimensional human pose estimation result is output. The method includes using a CPN detector to obtain two-dimensional joint coordinates, perform pose correction and dimension-raising processing, extract spatial and temporal features in combination with adaptive graph attention unit and expansion time convolution model, and constructing a multi-level supervised loss function for iterative training.
This method can correct the local noise-free two-dimensional attitude, improve the flexibility of the model; capture rich spatial information through the dynamic graph attention module, and improve the inference effect in self-occlusion or depth blur; multi-level supervision strategy realizes prediction from coarse to fine, improves estimation accuracy, and obtains smooth video three-dimensional human posture motion results.
Smart Images

Figure CN114694261B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of human pose estimation, and particularly relates to a method and system for video three-dimensional human pose estimation based on multi-level supervised graph convolution. Background Technique
[0002] Three-dimensional human pose estimation is the basis for many research works in the field of computer vision and is also a hot research topic, which has very important practical significance in the fields of intelligent monitoring, medical rehabilitation, autonomous driving, game animation, etc. However, standard three-dimensional human motion capture systems usually require subjects to wear marker suits to collect actions in an indoor controlled environment. Such devices are expensive and complex, and it is impractical to collect actions outdoors. Therefore, how to directly regress the three-dimensional joint positions of the human body from the video stream has become a hot topic in the field of computer vision. Due to the depth ambiguity naturally existing in the ill-posed problem of the three-dimensional pose estimation task and often accompanied by self-occlusion problems, the three-dimensional human pose estimation task in videos remains challenging.
[0003] Before the wide application of deep learning, researchers mainly estimated three-dimensional human poses through some methods applied in the fields of traditional computer vision or machine learning. In recent years, with the popularization of three-dimensional pose annotation datasets and GPUs with high computing power, deep learning methods have become the main methods for three-dimensional human pose estimation, and they have also been greatly improved in terms of estimation accuracy, execution efficiency, etc.
[0004] In recent years, deep learning-based three-dimensional human pose estimation methods can be roughly divided into two types. One is the end-to-end method, which directly predicts three-dimensional pose coordinates from the input RGB image using a neural network. Its advantage is that the entire network model can achieve end-to-end training effects, but this method has high requirements for network structure and data preprocessing. The other is the two-stage two-dimensional information-based method, that is, first obtain two-dimensional information, and then predict three-dimensional pose coordinates from the two-dimensional pose. Two-dimensional pose estimation is relatively mature. The advantage is that the network model is relatively easy to learn the mapping from two-dimensional to three-dimensional, and at the same time, it is also relatively easy to introduce reprojection for semi-supervision. Therefore, this method is relatively mainstream.
[0005] In the methods of three-dimensional human pose estimation in videos, current mainstream methods are all two-stage. Many researchers have conducted research on this mainstream method for many years and achieved good results. However, when faced with the problems of depth ambiguity and self-occlusion of human joints in video sequences, the existing methods do not perform well. It is difficult for the estimator to learn the context information of joint movements in the frame sequence, and the obtained limb positions are far from the true positions. In addition, since the commonly used large public dataset Human3.6M is collected in an indoor controlled environment and does not contain some complex and diverse human actions outdoors, it is difficult for the estimator to make judgments when faced with complex actions, and the prediction accuracy is not high. Therefore, how to solve the problems of depth ambiguity and self-occlusion brought by the current video three-dimensional pose estimation task and improve the estimation accuracy is an urgent problem to be solved currently. Summary of the Invention
[0006] Aiming at the deficiencies of the existing technology, the present invention proposes a method and system for three-dimensional human pose estimation in videos based on multi-level supervised graph convolution. The method includes: obtaining video data to be estimated, inputting the video data into a trained three-dimensional human pose estimation model based on multi-level supervised graph convolution, and outputting the three-dimensional human pose estimation result;
[0007] The training process of the three-dimensional human pose estimation model based on multi-level supervised graph convolution is as follows:
[0008] S1: Obtain a training dataset;
[0009] S2: Use a CPN detector to obtain the two-dimensional joint coordinates of the human body in each video frame of the training dataset, and obtain a two-dimensional pose sequence according to the two-dimensional joint coordinates;
[0010] S3: Perform pose correction on the two-dimensional pose sequence to obtain a corrected two-dimensional pose sequence;
[0011] S4: Perform dimensionality elevation processing on the corrected two-dimensional pose sequence to obtain a two-dimensional pose sequence after dimensionality elevation;
[0012] S5: Cross-use an adaptive graph attention unit and a dilated temporal convolution model to extract the spatial features of the two-dimensional pose sequence and the temporal features of the two-dimensional pose sequence;
[0013] S6: Construct a multi-level supervised loss function of the model;
[0014] S7: Fuse the temporal features and spatial features and input them into a fully connected layer to obtain the three-dimensional human pose estimation result;
[0015] S8: Continuously adjust the parameters of the model, jointly optimize and solve the loss function, and perform iterative training on the model until the multi-level supervised loss function converges.
[0016] Preferably, the process of obtaining the two-dimensional joint coordinates of the human body includes:
[0017] Pre-training the CPN detector with the two-dimensional dataset COCO and fine-tuning the CPN detector with the two-dimensional projection of the three-dimensional pose dataset Human3.6M to obtain a trained CPN detector;
[0018] Using the trained CPN detector to perform two-dimensional pose estimation on each video frame to obtain the two-dimensional joint coordinates of the human body for each frame.
[0019] Preferably, the pose correction of the two-dimensional pose sequence includes: the CPN detector assigns a confidence score to the pose in each frame sequence, constructs a loss function weighted according to the confidence score, and uses the loss function for supervision. When the loss function is minimized, the corrected two-dimensional pose sequence is obtained.
[0020] Furthermore, the loss function is:
[0021]
[0022] where F represents the number of frame sequences, represents the reliability of the human joints, and a f represents the ground truth two-dimensional joint abscissa, represents the ground truth two-dimensional joint ordinate, and b f represents the noisy two-dimensional abscissa in the two-dimensional pose sequence, represents the noisy two-dimensional ordinate in the two-dimensional pose sequence.
[0023] Preferably, the process of using the adaptive graph attention unit to extract the spatial features of the two-dimensional pose sequence includes: using the dynamic graph unit to process the poses in the two-dimensional pose sequence to obtain a "constructed graph"; obtaining the first-order neighbor points and second-order neighbor points according to the "constructed graph"; constructing the adjacency matrix of the "constructed graph" based on the first-order neighbor points and second-order neighbor points, and using the adjacency matrix as the convolution kernel of graph convolution; using the graph convolution algorithm of the "constructed graph" to extract the spatial features of the two-dimensional pose sequence according to the convolution kernel of graph convolution; where the first-order neighbor points are the nodes in the "constructed graph" with a distance of 1 from the target joint point, and the second-order neighbor points are the nodes in the "constructed graph" with a distance of 2 from the target joint point.
[0024] Furthermore, the output of each layer of the graph convolution algorithm can be expressed as:
[0025]
[0026] where J (l+1) represents the (l + 1)-th layer of the network, J (l) represents the l-th layer of the network, C represents the number of channels, Denote the convolution kernel of graph convolution as w c Denote the c-th row vector in the transformation matrix W as M c Denote the weight matrix of the c-th channel. ρ and σ respectively denote the Softmax and ReLU non-linear activation functions.
[0027] Preferably, the process of using the dilated temporal convolution model to extract the temporal features of the 2D pose sequence includes:
[0028] Perform a temporal convolution with a dilation factor of d = k in the convolution block on the 2D pose sequence T to obtain intermediate temporal features, where k represents an odd number, and T represents the T-th temporal convolution block;
[0029] Perform a 1×1 convolution on the intermediate temporal features to obtain the intermediate temporal features after dilating the dimension;
[0030] Process the intermediate temporal features after dilating the dimension using batch normalization, ReLU activation function, and dropout to obtain non-overfitting intermediate temporal features; input the non-overfitting intermediate temporal features into a fully connected layer to obtain the final temporal features.
[0031] Furthermore, the process of using the dilated temporal convolution model to extract the temporal features of the 2D pose sequence further includes: implementing residual connection using max pooling in each layer of the convolution block to obtain temporal features with matching front and back dimensions.
[0032] Preferably, the multi-level supervised loss function of the model is:
[0033]
[0034] where denotes the loss function of the intermediate level, L final denotes the loss function of the last layer of the network, L total denotes the loss function of the entire network, T represents the number of temporal convolution blocks, α and β both denote balance factors, L Ref denotes the loss function of the optimization module, F represents the number of video frames input, Q represents the number of joint angles in one frame, denotes the joint angle predicted at the intermediate level, denotes the joint angle predicted at the network level, denotes the true angle of the joint.
[0035] A video three-dimensional human pose estimation system based on multi-level supervised graph convolution includes: an input module, a 2D pose sequence acquisition module, a network model loading module, and an output module;
[0036] The input module is used to input the human motion video to be estimated;
[0037] The two-dimensional pose sequence acquisition module is used to acquire the two-dimensional pose sequence in the input video sequence and input the two-dimensional pose sequence into the network model loading module;
[0038] The network model loading module includes a two-dimensional pose correction module, a dynamic graph attention module, an extended temporal convolution module, and an estimation module;
[0039] The two-dimensional pose correction module is used to correct the two-dimensional pose sequence to obtain the corrected two-dimensional pose sequence;
[0040] The dynamic graph attention module is used to extract the spatial features of the corrected two-dimensional pose sequence;
[0041] The extended temporal convolution module is used to extract the temporal features of the corrected two-dimensional pose sequence;
[0042] The estimation module obtains the three-dimensional human pose estimation result according to the spatial features of the two-dimensional pose sequence and the temporal features of the two-dimensional pose sequence;
[0043] The output module is used to output the three-dimensional human pose estimation result of the human motion video.
[0044] The beneficial effects of the present invention are as follows: The present invention can correct the two-dimensional pose with local noise, compensating for the influence caused by two-dimensional detection errors to a certain extent, making the model not restricted by a specific two-dimensional detector, thereby improving the flexibility of the model; The dynamic graph attention module is adopted for spatial feature extraction, which can capture not only the spatial relationship between joints with physical connections in the human skeleton diagram but also the spatial relationship between joints without physical connections but with high logical correlation, enabling the model to carry rich spatial information during inference and achieving a good inference effect in the case of self-occlusion or depth blur; The attention module of the present invention can enable the model to focus on the joints at the end of the human motion chain specifically. At the same time, the temporal convolution module and the adaptive graph attention module are arranged alternately to realize the modeling of spatio-temporal information by the model, thereby improving the estimation accuracy; A multi-level supervision strategy is proposed. The intermediate prediction results are roughly obtained after each temporal convolution module, and the multi-level features are fused at the end of the network to achieve a coarse-to-fine prediction; The present invention has high estimation accuracy. For video frames with a given input length, the model can achieve high-precision inference and obtain a smooth three-dimensional human pose motion result of the video, having good economic benefits. Description of the Drawings
[0045] Figure 1 It is the training flow chart of the three-dimensional human pose estimation model in the present invention;
[0046] Figure 2 It is the schematic diagram of the adaptive graph convolution module in the present invention;
[0047] Figure 3 Schematic diagram of the temporal convolutional block in the present invention;
[0048] Figure 4 Network model diagram of the present invention;
[0049] Figure 5 Schematic diagram of the training process of a preferred embodiment in the present invention;
[0050] Figure 6 Flowchart of a prototype system of the present invention;
[0051] Figure 7 Effect diagram of the three-dimensional human pose estimation function of a prototype system of the present invention. Detailed implementation manners
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0053] The present invention proposes a method and system for video three-dimensional human pose estimation based on multi-level supervised graph convolution. The method includes: obtaining video data to be estimated, inputting the video data into a trained video three-dimensional human pose estimation model based on multi-level supervised graph convolution, and outputting a three-dimensional human pose estimation result;
[0054] As Figure 1 shown, the training process of the video three-dimensional human pose estimation model based on multi-level supervised graph convolution is as follows:
[0055] S1: Obtain a training data set;
[0056] S2: Use a CPN detector to obtain the human body two-dimensional joint coordinates of each video frame in the training data set, and obtain a two-dimensional pose sequence according to the two-dimensional joint coordinates;
[0057] S3: Perform pose correction on the two-dimensional pose sequence to obtain a corrected two-dimensional pose sequence;
[0058] S4: Perform dimension elevation processing on the corrected two-dimensional pose sequence to obtain a dimension-elevated two-dimensional pose sequence;
[0059] S5: Cross-use an adaptive graph attention unit and a dilated temporal convolution model to extract the spatial features of the two-dimensional pose sequence and the temporal features of the two-dimensional pose sequence;
[0060] S6: Construct a multi-level supervised loss function of the model;
[0061] S7: Fuse the temporal features and spatial features and input them into the fully connected layer to obtain the 3D human pose estimation result;
[0062] S8: Continuously adjust the parameters of the model, jointly optimize and solve the loss function, and iteratively train the model until the multi-level supervised loss function converges.
[0063] The specific training process of the 3D human pose estimation model based on multi-level supervised graph convolution in the present invention is as follows:
[0064] Obtain the training data set. The process is as follows: Obtain the original data set, divide the Human3.6M data set as the original data set to obtain the training set and the test set; The training set is used to train the network model and perform multiple iterations, and the test set is used to test the model.
[0065] The process of obtaining the 2D human joint coordinates includes:
[0066] CPN uses a ResNet-50 backbone network with a resolution of 384×288, pre-trains CPN (Cascaded Pyramid Network, a 2D pose detector) using the 2D data set COCO, and fine-tunes the CPN detector using the 2D projection of the 3D pose data set Human3.6M. During the fine-tuning process, the CPN detector always maintains batch normalization, and initializes the last layer of the Global-Net and Refine-Net (convolution weights and batch normalization statistics) of the CPN detector. The batch size is 64 images, and then it is trained on the GPU using the learning rate gradual decay strategy to obtain the trained CPN detector.
[0067] Use the trained CPN detector to perform 2D pose estimation on each video frame to obtain the 2D coordinates of the human joints in each frame.
[0068] The pose correction for the 2D pose sequence includes: The CPN detector obtains the pose sequence of F frames (where J is the number of human joint points in each frame), and also assigns confidence scores to the poses in the F-frame sequence Construct a loss function by weighting according to the confidence scores, and use the loss function for supervision. When the loss function is minimized, the corrected 2D pose sequence is obtained.
[0069] The loss function is:
[0070]
[0071] where F represents the number of frame sequences, represents the reliability of the human joints, a fDenote the ground truth two-dimensional joint abscissa, Denote the ground truth two-dimensional joint ordinate, b f Denote the noisy two-dimensional abscissa in the two-dimensional pose sequence, Denote the noisy two-dimensional ordinate in the two-dimensional pose sequence.
[0072] Perform dimensionality elevation processing on the corrected two-dimensional human pose sequence to obtain the dimensionality-elevated two-dimensional pose sequence; for example, taking an 81-frame receptive field as an example, since there are 17 annotated human joints, each frame has 17×2 two-dimensional joint information, and the input tensor dimension of the dimensionality elevation module is 81×17×2. First, it passes through a convolution with a convolution kernel size of 3×1 to expand the dimension from 2D to 64D, and uses batch normalization, ReLU, and Dropout for conventional processing. The output tensor dimension becomes 79×17×64.
[0073] As Figure 2 shown, the process of using the adaptive graph attention unit to extract the spatial features of the two-dimensional pose sequence includes: using the dynamic graph unit to form a "constructed graph" according to the poses in the two-dimensional pose sequence, taking the adjacency matrix of the first-order plus second-order neighbors of the "constructed graph" as the convolution kernel of graph convolution, and using the graph convolution algorithm of the "constructed graph" to extract the spatial features of the two-dimensional pose sequence; the specific process is as follows:
[0074] To enable the model to comprehensively learn the spatial information of human poses, a dynamic graph unit is introduced. It can not only capture the relationship between physically connected human joints in space but also adaptively capture the relationship between joints that do not have physical connections but are highly correlated logically. The dynamic graph unit can form a new "constructed graph" for different poses. Compared with only using the pre-defined two-dimensional "skeleton graph" of the human body, this "constructed graph" generated adaptively according to the motion pose is more persuasive because an "implicit edge" is implicitly formed between joints that do not have physical connections but are highly affinity logically. In this way, the update of joint point features during training not only depends on the neighbor nodes of the "skeleton graph" but also on the neighbor nodes of the "constructed graph" generated by the dynamic graph unit.
[0075] In addition, since the end joints of the human motion chain (such as the left wrist, right wrist, head, left ankle, and right ankle) have only one first-order neighbor in the skeleton diagram (the number of first-order neighbors of other joints is greater than or equal to two), and large errors often occur at the end joint points. Once occlusion occurs, it will be very difficult to estimate their spatial positions. Therefore, the present invention introduces a graph attention module to enable the model to pay more attention to the end joints during training; the commonly used graph convolution operation uses the adjacency matrix of the graph as the convolution kernel. The elements in the adjacency matrix are either 0 or 1. 0 represents that the distance between two points is not one, and 1 represents that there is a direct edge between two points and the distance is 1. The matrix represented in this way is the adjacency matrix of the first-order neighbor points. The present invention obtains the first-order neighbor points and second-order neighbor points according to the "constructed graph". The first-order neighbor points are the nodes in the "constructed graph" that are at a distance of 1 from the target joint point, and the second-order neighbor points are the nodes in the "constructed graph" that are at a distance of 2 from the target joint point; an adjacency matrix of the "constructed graph" is constructed based on the first-order neighbor points and second-order neighbor points, that is, if the distance between two points in the matrix is 1 or 2, it is recorded as 1, and if the distance is greater than 2, it is recorded as 0. The adjacency matrix is used as the convolution kernel of graph convolution; according to the convolution kernel of graph convolution, the graph convolution algorithm of the "constructed graph" is used to extract the spatial features of the two-dimensional pose sequence. The reason for operating on the "constructed graph" is that there are two types of edges between the nodes in the "constructed graph": one is the "physical edge" that originally exists between the joint points, and the other is the "constructed edge" that has a high affinity but no physical connection between the joint points. In this way, during training, the features of the end joints have more choices during the update period. They can not only explicitly update the features according to the first- and second-order neighbors connected by the "physical edge", but also implicitly update the features according to the first- and second-order neighbors connected by the "constructed edge"; the process of using the graph convolution algorithm of the "constructed graph" to extract the spatial features of the two-dimensional pose sequence is as follows:
[0076] Denote the "skeleton graph" in the current frame as G F =(V F , E F ), where V F represents the set of joint nodes of the skeleton graph, and E F represents the set of edges connected by the joint nodes, represents the features of N joint nodes in the current frame, represents the feature vector of the Nth joint node, and c represents the number of features of each joint node.
[0077] For conventional graph convolution, the structure of the graph can be initialized using the first-order adjacency matrix A∈R N×N representing the connections between joints and the identity matrix I∈R N×N representing self-connections, is the adjacency matrix of the graph, and the degree matrix is used for row regularization, then As the convolution kernel of graph convolution. Then the output of each layer of graph convolution can be defined as where W is the learnable weight matrix, l represents the l-th layer of the network, σ is the ReLU non-linear activation function, and the graph convolution algorithm of "constructing a graph" in the present invention is described as:
[0078] For a certain joint j in the F-th frame, use the nearest neighbor algorithm to find its set of K neighbor nodes in the joint feature matrix J F in
[0079] For joint j, generate "constructed edges" among its k neighbors, then the adjacency matrix of the "constructed graph" is Then add the identity matrix I to obtain And perform row normalization using the degree matrix to obtain And construct a new convolution kernel
[0080] Use a learnable weight matrix M ∈ R N×N to learn the joints with different importance inside the first- and second-order neighbors, and use a learnable transformation matrix W to transform the output channels. At this time, the output of each layer of graph convolution can be briefly expressed as:
[0081]
[0082] Furthermore, adopt different weight matrices for each channel c of the output node features, and then perform connection at the channel level. At this time, the output of each layer of graph convolution can be specifically expressed as:
[0083]
[0084] where, J (l+1) represents the (l + 1)-th layer of the network, J (l) represents the l-th layer of the network, C represents the number of channels, represents the convolution kernel of graph convolution, w c represents the c-th row vector in the transformation matrix W, M c represents the weight matrix of the c-th channel, ρ and σ respectively represent the Softmax and ReLU non-linear activation functions; in addition, there are batch normalization and non-linear activation units after each layer of graph convolution, and Dropout is used to prevent overfitting.
[0085] As Figure 3 shown, the process of using the dilated temporal convolution model to extract the temporal features of the two-dimensional pose sequence includes: the dilated temporal convolution model consists of T temporal convolution blocks with residual connections, the convolution kernel size is k×1, and each temporal convolution block is sequentially set with an exponentially increasing dilation factor d = k TTo achieve precise control of the receptive field by the temporal convolutional model; perform temporal convolution with a dilation factor of d = k on the two-dimensional pose sequence in the convolutional block T to obtain intermediate temporal features, where k represents an odd number and T represents the T-th temporal convolutional block; perform 1×1 convolution on the intermediate temporal features to obtain the intermediate temporal features after dimension expansion; use batch normalization, ReLU activation function, and dropout to process the intermediate temporal features after dimension expansion to obtain non-overfitting intermediate temporal features; input the non-overfitting intermediate temporal features into the fully connected layer to obtain the final temporal features; in addition, max pooling is used in each convolutional block during the entire extraction process to implement residual connection, obtaining temporal features with matching front and back dimensions.
[0086] The present invention interleaves the graph convolutional module and the temporal convolutional module to achieve complementarity between spatial features and temporal features. Before entering the temporal convolutional block, the joint spatial features of each frame are extracted using the dynamic graph attention module, and the attention to the end joints is strengthened specifically to improve the inference ability of the model in the case of occlusion.
[0087] Fuse the temporal features and spatial features of the two-dimensional pose sequence and input them into the fully connected layer to obtain the three-dimensional human pose estimation result. For example, taking an 81-frame receptive field as an example, the dilated temporal convolutional model consists of 3 temporal convolutional blocks, the convolutional kernel size is 3×1, and the dilation factors of each temporal convolutional block are 3, 9, and 27 respectively. After the last convolutional block, the output tensor is 1×17×1024, and then through a fully connected layer, the final output tensor 1×17×3 is obtained, which is the three-dimensional human pose of the central frame.
[0088] The specific process of constructing the multi-level supervised loss function of the model is as follows: The present invention uses the error between the estimated joint angle and the true joint angle as the basic supervised loss to strengthen the constraint. For the labeled dataset Human3.6M, there are 17 human joints in total, so there are 16 human bones (16 spatial vectors), and the angle between joints can be calculated by the spatial vector included angle formula; the network model of the present invention is stacked by T temporal convolutional blocks, and low-level spatio-temporal features are output after each temporal convolutional block to obtain a rough estimated pose result, and the error between it and the ground truth joint angle is calculated to achieve the purpose of multi-level supervision and realize the optimization from rough to fine. The multi-level low-level features are fused at the last layer to obtain the final accurate three-dimensional pose. Define the loss function of each intermediate block as:
[0089]
[0090] where represents the loss of each intermediate level, F represents the number of video frames, Q represents the number of joint angles in a frame, Represents the predicted joint angles of each intermediate block, represents the ground truth joint angles.
[0091] The supervised loss function of the last layer of the network is defined as:
[0092]
[0093] where L final represents the supervised loss of the last layer of the network, represents the predicted joint angles of the last layer of the network, represents the ground truth joint angles.
[0094] Forming multi-level supervision with the intermediate predictions at each level and the final prediction is:
[0095]
[0096] where L 3d represents the multi-level supervision loss, α represents the balance factor, and T represents the number of temporal convolutional blocks. Finally, the overall loss function of the model is expressed as:
[0097] L total = L 3d + βL Ref
[0098] That is:
[0099]
[0100] where L total represents the total loss of the model, β represents the balance factor, and L Ref represents the supervised loss of the optimization module.
[0101] Continuously adjust the parameters of the model, jointly optimize and solve the loss function value, and perform iterative training on the model until convergence. The trained model is as Figure 4 shown.
[0102] The training process of a preferred embodiment in the present invention is as Figure 5 shown. Input the two-dimensional pose sequence into the video three-dimensional human pose estimation model based on multi-level supervised graph convolution, and perform supervised training under the guidance of ground truth annotations. Specifically, the present invention uses the Amsgrad optimizer for training, and adopts the learning rate adjustment strategy of cosine annealing. The momentum of BatchNorm starts from 0.1 and adopts the exponential decay strategy, reaching 0.001 at the last epoch, and the Dropout rate is 0.25. After training for 60 epochs, the neural network tends to be stable and the iterative training ends.
[0103] The process of a prototype system of the present invention is asFigure 6 As shown in the figure, after importing a video from the local, first select the video three-dimensional pose estimation function. If the two-dimensional detector is successfully called, use it to perform two-dimensional human pose estimation on the input video, and then input the two-dimensional human pose sequence obtained by the detector into the video three-dimensional human pose estimation network based on multi-level supervised graph convolution to perform the two-dimensional to three-dimensional lifting task, so as to obtain accurate frame-by-frame three-dimensional human pose estimation results.
[0104] A video three-dimensional human pose estimation system based on multi-level supervised graph convolution includes: an input module, a two-dimensional pose sequence acquisition module, a network model loading module, and an output module;
[0105] The input module is used to input the human motion video to be estimated;
[0106] The two-dimensional pose sequence acquisition module is used to obtain the two-dimensional pose sequence in the input video sequence and input the two-dimensional pose sequence into the network model loading module;
[0107] The network model loading module includes a two-dimensional pose correction module, a dynamic graph attention module, a dilated temporal convolution module, and an estimation module;
[0108] The two-dimensional pose correction module is used to correct the two-dimensional pose sequence to obtain the corrected two-dimensional pose sequence;
[0109] The dynamic graph attention module is used to extract the spatial features of the corrected two-dimensional pose sequence;
[0110] The dilated temporal convolution module is used to extract the temporal features of the corrected two-dimensional pose sequence;
[0111] The estimation module obtains the three-dimensional human pose estimation result according to the spatial features of the two-dimensional pose sequence and the temporal features of the two-dimensional pose sequence;
[0112] The output module is used to output the three-dimensional human pose estimation result of the human motion video.
[0113] The three-dimensional human pose estimation function effect of a prototype system of the present invention is as Figure 7 shown. After importing a video from the local path, select the video three-dimensional pose estimation function and call the two-dimensional detector to obtain the two-dimensional pose sequence, and then click to generate the three-dimensional human pose. The system will load the trained weights of the video three-dimensional human pose estimation network based on multi-level supervised graph convolution to obtain the estimation result. The human motion skeleton animation estimated by the model can be seen on the right.
[0114] The video three-dimensional human pose estimation method of the present invention can be applied to human action recognition scenarios. For example, it can be used to track the changes in a person's pose over a period of time, for activity and gait recognition, and can detect whether a person has fallen or has abnormal behavior. At the same time, the present invention helps researchers develop corresponding applications and apply them in the field of autonomous driving. The central processing unit analyzes the actions of road pedestrians captured in real time by the car camera, fully understands the actions of pedestrians, and predicts the subsequent movement trajectories of pedestrians, enabling the car to make further decisions and avoid traffic accidents in advance, improving the response ability and safety of autonomous driving in the face of complex road environments. Optionally, the face detection method provided in this application can also be applied to the following scenarios:
[0115] I. Human-computer interaction scenario;
[0116] For example, the subject does not need to wear complex motion capture devices, and can obtain relatively accurate body postures only through camera sensors. By tracking the changes in human postures, the machine can keenly and carefully discover the intentions of the subject, enabling the robot to follow the trajectory of the human body posture skeleton of the person performing the action, rather than manually programming the robot to follow the trajectory.
[0117] II. Video surveillance scenario;
[0118] For example, in an environment with a large number of people, such as a railway station, an airport, a bank, or a government building, through intelligent surveillance alone, it is possible to learn the movement trajectory of a person in the video, discover and analyze their abnormal behavior, and record the person, improving the public security prevention and control level in public places.
[0119] III. Game modeling scenario;
[0120] For example, in the development of large-scale 3D action games, 3D character modeling is a complex task. If the three-dimensional human pose can be estimated, then graphics, styles, fancy enhancements, devices, and artworks can be superimposed on the person. By tracking the changes in the three-dimensional human pose, the actions of virtual characters can be rendered, and the model animations used for rendering can "naturally fit" them when the person moves.
[0121] The present invention can correct the two-dimensional pose with local noise, compensating to a certain extent for the influence brought by two-dimensional detection errors, making the model not restricted by a specific two-dimensional detector, thereby improving the flexibility of the model; in spatial feature extraction, a dynamic graph attention module is adopted, which can capture the spatial relationship between joints with physical connections in the human skeleton graph and can also capture the spatial relationship between joints without physical connections but with high logical relevance, enabling the model to carry rich spatial information during inference and achieving a good inference effect in the case of self-occlusion or depth blur; the attention module of the present invention enables the model to focus specifically on the joints at the end of the human motion chain. At the same time, the temporal convolution module and the adaptive graph attention module are arranged alternately to realize the modeling of spatio-temporal information by the model, thereby improving the estimation accuracy; a multi-level supervision strategy is proposed. After each temporal convolution module, intermediate prediction results are roughly obtained, and multi-level features are fused at the end of the network to achieve coarse-to-fine prediction; the present invention has high estimation accuracy. For video frames with a given input length, the model can achieve high-precision inference and obtain smooth three-dimensional human pose motion results of the video, having good economic benefits.
[0122] It should be noted that in each embodiment of the present disclosure, each functional module can be integrated in one processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module; if the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product.
[0123] The above examples further elaborate in detail the purpose, technical solution and advantages of the present invention. It should be understood that the above examples are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modification, equivalent replacement, improvement, etc. made to the present invention within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for video three-dimensional human pose estimation based on multi-level supervised graph convolution, characterized in that, it includes: Obtain the video data to be estimated, input the video data into the trained video three-dimensional human pose estimation model based on multi-level supervised graph convolution, and output the three-dimensional human pose estimation result; The training process of the video three-dimensional human pose estimation model based on multi-level supervised graph convolution is: S1: Obtain the training data set; S2: Use the CPN detector to obtain the human body two-dimensional joint coordinates of each video frame in the training data set, and obtain the two-dimensional pose sequence according to the two-dimensional joint coordinates; S3: Perform pose correction on the two-dimensional pose sequence to obtain the corrected two-dimensional pose sequence; S4: Perform dimensionality elevation processing on the corrected two-dimensional pose sequence to obtain the two-dimensional pose sequence after dimensionality elevation; S5: Cross-use the adaptive graph attention unit and the dilated temporal convolution model to extract the spatial features of the two-dimensional pose sequence and the temporal features of the two-dimensional pose sequence; S6: Construct the multi-level supervised loss function of the model; the multi-level supervised loss function of the model is: Among them, represents the loss function of the intermediate level, L final represents the loss function of the last layer of the network, L total represents the loss function of the entire network, T represents the number of temporal convolutional blocks, and both α and β represent balance factors, L Ref represents the loss function of the optimization module, F represents the number of video frames in the input, Q represents the number of joints in each frame, represents the joint positions predicted at the intermediate level, represents the joint positions predicted at the network level, represents the true positions of the joints; S7: Fuse the temporal features and spatial features and input them into the fully connected layer to obtain the three-dimensional human pose estimation result; S8: Continuously adjust the parameters of the model, jointly optimize and solve the loss function, and perform iterative training on the model until the multi-level supervised loss function converges.
2. The method for video three-dimensional human pose estimation based on multi-level supervised graph convolution according to claim 1, characterized in that, The process of obtaining the human body two-dimensional joint coordinates includes: Pre-train the CPN detector with the two-dimensional data set COCO, and fine-tune the CPN detector with the two-dimensional projection of the three-dimensional pose data set Human3.6M to obtain the trained CPN detector; Use the trained CPN detector to perform two-dimensional pose estimation on each video frame to obtain the two-dimensional coordinates of the human joints in each frame.
3. The method for video three-dimensional human pose estimation based on multi-level supervised graph convolution according to claim 1, characterized in that, Performing pose correction on the two-dimensional pose sequence includes: The CPN detector assigns a confidence score to the pose in each frame sequence, constructs a loss function weighted according to the confidence score, and uses the loss function for supervision. When the loss function is minimized, the corrected two-dimensional pose sequence is obtained.
4. The method for video three-dimensional human pose estimation based on multi-level supervised graph convolution according to claim 3, characterized in that, The loss function is: Among them, F represents the number of frame sequences, represents the reliability of human joints, a f represents the abscissa of the ground truth two-dimensional joint, represents the ordinate of the ground truth two-dimensional joint, b f represents the noisy two-dimensional abscissa in the two-dimensional pose sequence, represents the noisy two-dimensional ordinate in the two-dimensional pose sequence.
5. The method for video three-dimensional human pose estimation based on multi-level supervised graph convolution according to claim 1, characterized in that, The process of extracting the spatial features of the two-dimensional pose sequence using the adaptive graph attention unit includes: processing the poses in the two-dimensional pose sequence using the dynamic graph unit to obtain the "constructed graph"; obtaining the first-order neighbor points and second-order neighbor points according to the "constructed graph"; constructing the adjacency matrix of the "constructed graph" based on the first-order neighbor points and second-order neighbor points, and using the adjacency matrix as the convolution kernel of graph convolution; extracting the spatial features of the two-dimensional pose sequence using the graph convolution algorithm of the "constructed graph"; where the first-order neighbor points are the nodes in the "constructed graph" with a distance of 1 from the target joint point, and the second-order neighbor points are the nodes in the "constructed graph" with a distance of 2 from the target joint point.
6. A method for video three-dimensional human pose estimation based on multi-level supervised graph convolution according to claim 5, characterized in that, The output of each layer of the graph convolution algorithm can be expressed as: Among them, J (l+1) represents the (l + 1)-th layer of the network, and J (l) represents the l-th layer of the network. C represents the number of channels, represents the convolution kernel of the graph convolution, and w c represents the c-th row vector in the transformation matrix W. M c represents the weight matrix of the c-th channel. ρ and σ represent the Softmax and ReLU non-linear activation functions respectively.
7. A method for video three-dimensional human pose estimation based on multi-level supervised graph convolution according to claim 1, characterized in that, The process of extracting the temporal features of the two-dimensional pose sequence using the dilated temporal convolution model includes: Perform temporal convolution with dilation factor d = k on the two-dimensional pose sequence in the convolutional block T to obtain intermediate temporal features, where k represents an odd number and T represents the T-th temporal convolutional block; Performing 1×1 convolution processing on the intermediate temporal features to obtain the intermediate temporal features after dilating the dimension; Processing the intermediate temporal features after dilating the dimension using batch normalization, ReLU activation function, and dropout to obtain non-overfitting intermediate temporal features; inputting the non-overfitting intermediate temporal features into the fully connected layer to obtain the final temporal features.
8. A method for video three-dimensional human pose estimation based on multi-level supervised graph convolution according to claim 7, characterized in that, The process of extracting the temporal features of the two-dimensional pose sequence using the dilated temporal convolution model further includes: using max pooling in each convolutional block to implement residual connection to obtain temporal features with matching front and back dimensions.
9. A video three-dimensional human pose estimation system based on multi-level supervised graph convolution, which is used to execute any one of the methods for video three-dimensional human pose estimation based on multi-level supervised graph convolution described in claims 1 to 8, characterized in that, including: An input module, a two-dimensional pose sequence acquisition module, a network model loading module, and an output module; The input module is used to input the human motion video to be estimated; The two-dimensional pose sequence acquisition module is used to acquire the two-dimensional pose sequence in the input video sequence and input the two-dimensional pose sequence into the network model loading module; The network model loading module includes a two-dimensional pose correction module, a dynamic graph attention module, a dilated temporal convolution module, and an estimation module; The two-dimensional pose correction module is used to correct the two-dimensional pose sequence to obtain the corrected two-dimensional pose sequence; The dynamic graph attention module is used to extract the spatial features of the corrected two-dimensional pose sequence; The dilated temporal convolution module is used to extract the temporal features of the corrected two-dimensional pose sequence; The estimation module obtains the three-dimensional human pose estimation result according to the spatial features of the two-dimensional pose sequence and the temporal features of the two-dimensional pose sequence; The output module is used to output the three-dimensional human pose estimation result of the human motion video.