A 3D Human Pose Estimation Method Based on Deep Learning

By using parameterless average pooling structure and step convolution in the three-dimensional human pose estimation method, the existing methods have solved the problems of high computational complexity and a lot of redundant information, and achieved more efficient pose estimation.

CN115376209BActive Publication Date: 2025-06-27CHANGZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211031409.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-26
Publication Date
2025-06-27
Estimated Expiration
2042-08-26

AI Technical Summary

Technical Problem

The existing three-dimensional human posture estimation method has high computational complexity when processing long sequence inputs, making it difficult to effectively learn spatiotemporal relationships, and there is a large amount of redundant information, resulting in a decrease in inference speed and insufficient estimation accuracy.

Method used

The average pooling structure without parameters is used to replace the attention mechanism in the Transformer encoder to reduce the computational complexity; the multi-frame pose information is aggregated into a single-frame representation through step-by-step convolution to reduce redundant information.

Benefits of technology

It effectively reduces the computational complexity of the model, improves the accuracy of pose estimation, reduces redundant information, and improves the inference speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115376209B_ABST
    Figure CN115376209B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of image detection, and particularly to a three-dimensional human pose estimation method based on deep learning, including generating multi-level initial features for the input pose sequence; capturing the temporal dependencies of multiple initial features in the time domain; establishing a feature refinement module to enhance the initial representation; performing interactive modeling on the refined multiple pose features; and aggregating multi-frame pose information into a single-frame representation through strided convolution to complete pose estimation. The present invention can compress the computational amount of the original Transformer model and at the same time ensure the position accuracy of human pose estimation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image detection, and in particular, to a three-dimensional human pose estimation method based on deep learning. Background Art

[0002] As one of the classic computer vision tasks, three-dimensional human pose estimation aims to estimate the three-dimensional joint positions of a human body from images or videos. The traditional method using a two-dimensional to three-dimensional (2D-to-3D) ascent approach is to first use a pre-trained 2D detector to estimate the 2D key point coordinates and then lift them to the 3D space. However, due to the lack of depth information, a unique 3D skeleton corresponding to the 2D cannot be obtained.

[0003] The currently adopted method is to use an original Transformer model without convolutional layers to capture spatio-temporal information from 2D pose sequences, use a spatio-temporal Transformer to model the global dependencies between 2D joints, and output the three-dimensional human pose of the intermediate frame. However, it ignores the motion differences between human joints, resulting in insufficient learning of spatio-temporal relationships, and due to the increase in the size of the time Transformer module, it cannot be applied to scenarios with longer input sequences either.

[0004] There is also a method that uses the encoder of the Transformer to capture long-sequence pose information, then a strided Transformer encoder aggregates the long-sequence information, and finally an accurate estimation result is obtained. However, the computational complexity and running time cost of this method are huge, resulting in a decrease in the inference speed.

[0005] The huge model parameters are the main reason for the high computational cost of the Transformer model, and there is also a problem of a large amount of redundant information in continuous pose sequences. Summary of the Invention

[0006] Aiming at the deficiencies of the existing algorithms, the present invention solves the problem of the computational complexity of the original model by using a parameter-free average pooling structure to replace the attention mechanism in the Transformer encoder; and solves the problem of redundancy in continuous pose sequences in video frames by strided convolution to aggregate multiple adjacent poses into a single representation.

[0007] The technical solution adopted by the present invention is: a three-dimensional human pose estimation method based on deep learning includes the following steps:

[0008] S1. Generate multi-level initial features for the input pose sequence;

[0009] Further, step S1 includes:

[0010] S11. Obtain the pose sequence composed of T-frame video frames, divide the human body in each frame into J joint points, and represent each joint point with the coordinates in the 2D space;

[0011] S12. Map the pose information of J joints in T frames into vector representations, and add them to the learnable positional encoding, which is used to retain the spatial information of human joints;

[0012] S13. Select the Transformer encoder module as the backbone network, replace the attention mechanism therein with pooling operations, and cascade 3 layers of this encoder module to extract features from the input pose sequence, and output the initial feature representations encoded by each layer.

[0013] S2. Capture the temporal dependencies of multiple initial features in the time domain;

[0014] Furthermore, map the 3 initial feature representations after encoding into high dimensions respectively, and add them to the learnable temporal positional encoding to retain the position information of the frames;

[0015] S3. Establish a feature refinement module to enhance the initial representation;

[0016] Furthermore, step S3 includes:

[0017] S31. Perform feature refinement operations on the 3 initial representations after temporal positional encoding respectively to enhance the initial representation. The feature refinement module consists of a residual structure composed of normalization and average pooling;

[0018] S32. Merge the above 3 refined features along the channels, pass through a residual structure composed of normalization and a multi-layer perceptron, and finally the obtained features are evenly divided into 3 non-overlapping blocks along the channel dimension to form a refined feature representation;

[0019] S4. Perform interactive modeling on the refined multiple pose features;

[0020] Furthermore, step S4 includes:

[0021] Regard the 3 refined feature representations alternately as the query (Q), key-value (K), and value (V) in the attention mechanism, send them into a residual structure composed of LayerNorm normalization and cross-attention mechanism, merge them into a single feature representation along the channel direction, and finally realize the information interaction between cross-features through a residual structure composed of normalization and a multi-layer perceptron.

[0022] S5. Aggregate the multi-frame pose information into a single-frame representation through strided convolution to complete pose estimation;

[0023] Further, step S5 specifically includes: using strided convolution to aggregate the interacted temporal information into a single pose representation, compressing the channels through convolution until the number of frames becomes a single frame, and finally obtaining the 3D pose of the predicted central frame.

[0024] Advantages of the present invention:

[0025] 1. The non-learning parameter average pooling structure is adopted to replace the attention mechanism, realizing information interaction between sequences, and effectively reducing the computational complexity of the model;

[0026] 2. The cross-attention mechanism can effectively model the dependence relationship between multiple input features. At the same time, strided convolution aggregates information into a single vector representation of the pose sequence, which is beneficial to solving the problem of a large amount of redundant information in the pose sequence and improving the accuracy of estimation. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a three-dimensional human pose estimation method based on deep learning according to the present invention;

[0028] Figure 2 is a structural diagram of the Transformer encoder module of the present invention;

[0029] Figure 3 is a structural diagram of the multi-level feature generation model of the present invention;

[0030] Figure 4 is a structural diagram of the feature refinement module and the feature interaction module of the present invention;

[0031] Figure 5 is a visualization result diagram of the present invention in the Eating action;

[0032] Figure 6 is a visualization result diagram of the present invention in the Purchases action;

[0033] Figure 7 is a visualization result diagram of the present invention in the Walking action;

[0034] Figure 8 is a visualization result diagram of the present invention in the WalkTogether action. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] The present invention will be further described below with reference to the drawings and embodiments. This figure is a simplified schematic diagram, which only illustrates the basic structure of the present invention in a schematic manner, so it only shows the components related to the present invention.

[0036] As Figure 1As shown, a 3D human pose estimation method based on deep learning inputs a continuous pose sequence, processes it through multiple intermediate modules, and finally outputs the target pose of the central frame, including the following steps:

[0037] S1. Generate multi-level initial features for the input pose sequence;

[0038] Further, specifically including:

[0039] S11. Based on the original Transformer encoder module, replace the attention mechanism therein with a spatial average pooling module with a pooling kernel size of 3×3, a stride of 1, and a padding of 1. Without any learnable parameters, let each pose average extract features from the adjacent pose sequence to aggregate the nearby feature information. Its complexity is linearly related to the length of the sequence, reducing the computational complexity of the model, as Figure 2 shown;

[0040] S12. Obtain a pose sequence composed of 81 video frames. Divide the human body in each frame into 17 joint points, represent each joint point with the coordinates in the 2D space, map the pose information of 17 joints in each frame into a vector representation, and add it to the learnable position encoding, which is used to retain the spatial information of the human joints;

[0041] S13. To extract the pose sequence features, select the Transformer encoder module as the backbone network, replace the attention mechanism therein with a pooling operation to prevent gradient explosion, and use residual connection to add the input features of the previous encoder module to the output. Concatenate 3 encoder modules, and each module loops 4 times. Output the encoded initial feature representation at each layer. The multi-level feature generation network is as Figure 3 shown;

[0042] S2. Capture the temporal dependencies of multiple initial features in the time domain;

[0043] Further, in order to effectively learn the relationship between multi-level features, map the 3 encoded initial feature representations to 256 dimensions respectively, and then add them to the learnable temporal position encoding to maintain the position information of the frames;

[0044] S3. Establish a feature refinement module to enhance the initial representation;

[0045] Further, the three initially feature representations after temporal position encoding respectively pass through a residual structure composed of normalization and average pooling; then refined feature representations are obtained, and multiple initial features are processed independently; the multiple refined 256-dimensional features are merged along the channels into 768 dimensions, pass through a residual connection composed of normalization and a multi-layer perceptron, and finally the merged features are evenly divided into non-overlapping blocks along the channel dimension, and the dimension of each block is 256 dimensions, forming a refined feature representation. The feature refinement module is as shown in Figure 4 the left half of the dotted box shown in

[0046] S4. Interactive modeling is performed on the multiple refined pose features;

[0047] Further, it specifically includes:

[0048] S41. The cross-attention mechanism is applicable to three different inputs. The three refined feature representations are alternately regarded as the query (Q), key-value (K), and value (V) in the attention mechanism, and then input into a residual structure composed of normalization and the cross-attention mechanism, merged into 768 dimensions along the channel direction to form a single feature representation, and finally information interaction between cross-features is realized through a residual structure composed of normalization and a multi-layer perceptron. The feature interaction module is as shown in Figure 4 the right half of the dotted box shown in

[0049] S5. Aggregate multi-frame pose information into a single-frame representation through strided convolution to complete pose estimation;

[0050] S51. Aggregate the temporal information obtained by the feature interaction module into a single pose representation through a strided convolution module to reduce the sequence length, and make full use of the local information of the pose while not losing a large amount of useful information;

[0051] The input of the entire model is a sequence of 81 frames. First, it passes through a convolution with a kernel size of 1×1 and a stride of 1 for channel compression, then passes through the Relu activation function, and a Dropout layer is added to prevent overfitting of the model. Finally, a convolution with a kernel size of 3×3 and a stride of 3 is used to compress the sequence. At this time, the number of frames in the sequence becomes 27. After 4 such cycles, a frame of pose information is output. Finally, through the MLP module and a linear layer, the output frame of pose information is mapped to a 51-dimensional vector, and this vector corresponds to the three-dimensional coordinates of 17 joint points, and finally the output of the predicted central frame 3D pose is realized;

[0052] Table 1 Vector changes of each module in the overall model of the present invention

[0053]

[0054] Compared with the original Transformer method that fully uses the attention mechanism in steps S1, S3, and S4 modules, the present invention uses an average pooling structure in steps S1 and S3 modules, and an attention mechanism in step S4 module. The number of model parameters is 6.5M, which is reduced by 30%. The MPJPE is 36.1mm, and the position accuracy is improved by 8.6%.

[0055] The Human3.6M dataset is divided into 11 groups. Groups 9 and 11 are used as the test sets, and the rest are used as the training sets. Visualization operations are performed on the test results of the entire network on the Human3.6M test set. Randomly select actions such as Eating, Purchases, Walking, and WalkTogether for visualization, and extract any frame of the picture. The left side is the Input, the middle is the result reconstructed by the method of the present invention Ours (Pooling), and the right side is the Ground Truth of the real pose. It can be seen that the present invention can accurately detect human key points, such as Figures 5-8 shown.

[0056] The present invention uses a parameter-free average pooling structure to replace the attention mechanism to achieve information interaction between sequences, effectively reducing the computational complexity of the model; the cross-attention mechanism can effectively model the dependencies between multiple input features, and at the same time, the strided convolution aggregates information into a single vector representation of the pose sequence, which is beneficial to solving the problem of a large amount of redundant information in the pose sequence and improving the accuracy of estimation.

[0057] Inspired by the above ideal embodiments according to the present invention, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of the present invention. The technical scope of the present invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.

Claims

1. A three-dimensional human pose estimation method based on deep learning, characterized in that, It includes the following steps: S1. Generate multi-level initial features for the input pose sequence; Step S1 specifically includes: S11. Obtain a pose sequence composed of T video frames, divide the human body in each frame into J joint points, and represent each joint point with the coordinates in 2D space; S12. Map the pose information of J joints in T frames into vector representations, and add them to the learnable position encoding, where the position encoding is used to retain the spatial information of human joints; S13. Select the Transformer encoder module as the backbone network, replace the attention mechanism therein with pooling operations, and cascade 3 layers of Transformer encoder modules to extract features from the input pose sequence, and output the initial feature representations encoded by each layer; S2. Capture the temporal dependencies of multiple initial features in the time domain; S3. Establish a feature refinement module to enhance the initial representation; Step S3 specifically includes: S31. Perform feature refinement operations on the 3 initial representations encoded with temporal position encoding respectively to enhance the initial representation. The feature refinement module consists of a residual structure composed of normalization and average pooling; S32. Merge the 3 refined features along the channels, pass through a residual structure composed of normalization and a multi-layer perceptron, and the obtained features are evenly divided into 3 non-overlapping blocks along the channel dimension to form a refined feature representation; S4. Perform interactive modeling on the refined multiple pose features; S5. Aggregate the multi-frame pose information into a single-frame representation through strided convolution to complete pose estimation.

2. The 3D human pose estimation method based on deep learning according to claim 1, characterized in that Step S2 specifically includes: Map the 3 initial feature representations encoded respectively to high dimensions, and add them to the learnable temporal position encoding to maintain the position information of the frames.

3. The 3D human pose estimation method based on deep learning according to claim 1, characterized in that Step S4 specifically includes: Alternately regard the 3 refined feature representations as the query, key, and value in the attention mechanism, pass through a residual structure composed of normalization and cross-attention mechanism, merge them into a single feature representation along the channel direction, and finally realize the information interaction between cross-features through a residual structure composed of normalization and a multi-layer perceptron.

4. The three-dimensional human pose estimation method based on deep learning according to claim 1, characterized in that Step S5 specifically includes: Use strided convolution to aggregate the interactive temporal information into a single pose representation, perform channel compression through convolution until the number of frames becomes a single frame, and finally obtain the 3D pose of the predicted central frame.

Citation Information

Patent Citations

  • Table sequence identification method and system based on contextual relationship attention mechanism

    CN114241497A

  • Sitting posture recognition method based on monocular video image sequence

    WO2021237913A1