A 3D Human Pose Estimation Method Based on Spatiotemporal Context Feature Perception

The method enhances 3D human pose estimation by integrating spatial and temporal context features to address depth ambiguity and self-occlusion, improving accuracy and stability in 3D pose estimation.

CN114241515BActive Publication Date: 2025-07-15ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111373663.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-19
Publication Date
2025-07-15
Estimated Expiration
2041-11-19

AI Technical Summary

Technical Problem

When facing the problems of depth ambiguity and self-occlusion, the existing three-dimensional human posture estimation method has low prediction accuracy and jitter in continuous videos, so it is impossible to effectively utilize timing information.

Method used

Using a method based on space-time context feature perception, a two-dimensional human posture is detected through a cascading pyramid structure, and a space context perception module and a time multi-layer perception network are combined to extract geometric dependence and temporal information to perform three-dimensional human posture estimation.

Benefits of technology

It improves the accuracy of three-dimensional human posture estimation, reduces jitter, and achieves stable prediction results. It has a simple network structure, fast and efficient calculations, and is suitable for real-time applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114241515B_ABST
    Figure CN114241515B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional human pose estimation method based on spatio-temporal context feature perception. The corresponding two-dimensional human poses are extracted from each frame of the video and composed into a two-dimensional human pose skeleton data sequence. The spatial context perception module is used to process the two-dimensional skeleton sequence in turn to obtain the geometric constraint information features implicit in the human body structure. The temporal context perception module extracts the inherent temporal features from the entire two-dimensional human skeleton sequence data. Finally, the regression module regresses the corresponding three-dimensional human poses from the features generated by the foregoing modules. The present invention significantly improves the accuracy of three-dimensional human pose estimation, consumes less computing resources, and has strong robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of three-dimensional human posture estimation, and in particular, relates to a three-dimensional human posture estimation method based on spatiotemporal context feature perception. Background Art

[0002] 3D human pose estimation is a basic research in the field of computer vision and also a hot research direction. It has a wide range of applications in virtual reality, human-computer interaction, behavior analysis and other fields. In recent years, although deep learning-based methods have made great progress, 3D human pose estimation is still a very challenging task due to the inherent depth ambiguity and widespread self-occlusion in 2D representation data.

[0003] The existing 3D human pose estimation methods are mainly divided into two categories: (1) estimating 3D human pose directly from images; (2) estimating 2D human pose from images first, and then regressing 3D human pose. The former requires a lot of computing resources and is limited by limited 3D annotated data. The latter splits the entire task of 3D human pose estimation, making the prediction easier. In addition, 2D pose detection has a large amount of annotated data and has achieved good accuracy. However, one 2D pose can often correspond to multiple different 3D poses, especially in the presence of self-occlusion. This inherent depth ambiguity problem in 2D representation data greatly affects the prediction accuracy.

[0004] In order to solve the problem of depth ambiguity, it is an effective way to use the attention mechanism to efficiently learn implicit geometric constraint information from two-dimensional human posture. In addition, the existing three-dimensional human posture estimation methods often produce incoherent and jittery prediction results when continuously predicting on videos. This is because the human body is a highly free and nonlinear soft structure, and self-occlusion often occurs. The existing three-dimensional human body estimation methods based on single-frame images lack the association and constraints between temporal information and are not competent for prediction tasks under continuous videos. Therefore, building an effective time extraction model is more conducive to the robustness and versatility of the model. Summary of the invention

[0005] The purpose of this application is to provide a 3D human posture estimation method based on spatiotemporal context feature perception to improve prediction accuracy.

[0006] In order to achieve the above purpose, the technical solution of this application is as follows:

[0007] A three-dimensional human body posture estimation method based on spatiotemporal context feature perception, characterized in that the three-dimensional human body posture estimation method based on spatiotemporal context feature perception includes:

[0008] Input consecutive F frames in the monocular video, detect the human body bounding box, and then use a two-dimensional human pose detector with a cascaded pyramid structure to detect the two-dimensional coordinates of human body joints for each frame, and form a two-dimensional human skeleton sequence;

[0009] Normalize each two-dimensional human skeleton in the two-dimensional human skeleton sequence, and raise the dimension of the joint point coordinates in the normalized two-dimensional human skeleton to obtain the skeleton features after dimension raising;

[0010] Input the skeleton features after dimension raising into the spatial context perception module to extract the dependency relationship features containing the geometric dependency information between human body joints;

[0011] Input the dependency relationship features into the time multi-layer perceptron network module to further extract time information in the time dimension and obtain the time context features;

[0012] Average the time context features in the time dimension, then normalize them, and then pass through a fully connected layer to predict the corresponding three-dimensional human pose results.

[0013] Further, the normalization process for each two-dimensional human skeleton in the two-dimensional human skeleton sequence includes:

[0014] For each two-dimensional human skeleton in the two-dimensional human skeleton sequence, subtract the two-dimensional coordinates of the hip joint from the two-dimensional coordinates of each joint point to obtain the normalized two-dimensional human skeleton.

[0015] Further, the process of inputting the skeleton features after dimension raising into the spatial context perception module to extract the dependency relationship features containing the geometric dependency information between human body joints includes:

[0016] 3.1), First, according to the preset human body structure, construct the structure matrix through the following formula

[0017]

[0018] where S (i,p) represents the element in the i-th row and p-th column of the structure matrix S, MD(i, p) represents the flow distance between the i-th human body joint and the p-th human body joint, the flow distance between joints is determined by the preset human skeleton structure diagram, and K represents a predefined hyperparameter.

[0019] 3.2), Input the structure matrix S and the skeleton features x after dimension raising new into the spatial context perception module for skeleton feature learning. The spatial context perception module is composed of N identical structured pose encoders connected in series; the structure matrix S and the skeleton features x after dimension raising newAfter passing through the first pose encoder, a feature matrix is obtained. This feature matrix has the same dimension as the skeleton feature x new The input of the next pose encoder is the feature matrix output by the previous pose encoder and the structure matrix S. After passing through N pose encoders, the output feature The output feature is normalized through a LayerNorm layer to obtain a dependency feature containing the geometric dependency information between human joint points

[0020] Furthermore, the pose encoder performs the following operations:

[0021] First, the structure matrix S is flattened into a one-dimensional vector with a dimension of 1×J 2 and input into the skeleton attention module. The skeleton attention module consists of a fully connected layer with J 2 neurons and a sigmoid activation function, and outputs an attention vector

[0022] The input feature matrix is first passed through a LayerNorm layer, then the dimension is changed to C s ×J through a transpose operation. Then, it passes through a fully connected layer with J 2 neurons and a GELU activation function to obtain an intermediate feature with a dimension of C s ×J 2 , and then the intermediate feature is element-wise multiplied with the attention vector W Att to obtain an attention feature matrix. Finally, the attention feature matrix passes through a fully connected layer with J neurons to obtain a skeleton attention feature matrix W SA with a dimension of C s ×J. Finally, the skeleton attention feature matrix W SA is transposed to change the dimension to J×C s and added to the input feature x new to obtain a residual feature value W Ra ;

[0023] Then, the residual feature value W Ra passes through a LayerNorm layer, as well as a fully connected layer with C s neurons and a GELU activation function to further learn the skeleton feature. Finally, after passing through a fully connected layer with C s neurons, the output is added to the residual feature value W RA to obtain a new residual feature W New_RA with a dimension of J×C s ; W New_RAThat is the feature matrix output by the current pose encoder.

[0024] Furthermore, inputting the dependency relationship features into the time multi-layer perceptron network module to further extract time information in the time dimension to obtain time context features, including:

[0025] 4.1), Concatenating the dependency relationship features of each two-dimensional human skeleton to form a skeleton feature sequence, and then flattening the second and third dimensions of the skeleton feature sequence into one dimension to form a new skeleton feature sequence;

[0026] 4.2) Inputting the new skeleton feature sequence into the time multi-layer perceptron network module, and normalizing the output features to obtain time context features.

[0027] Furthermore, the time multi-layer perceptron network module is composed of multiple multi-layer perceptron mixers with the same structure connected in series. Each multi-layer perceptron mixer performs the following operations:

[0028] First, perform normalization through the LayerNorm layer, then use the transpose operation to change the input feature dimension to C t ×F, then pass through a fully connected layer containing D s neurons, a layer of GELU activation function and a fully connected layer containing F neurons to obtain intermediate features with a dimension size of C t ×F, then transpose the intermediate features to change the dimension to F×C t , and add it to the input features to obtain the residual feature value

[0029] Then normalize the residual feature value F T_Ra through the LayerNorm layer, as well as a fully connected layer containing D c neurons and a layer of GELU activation function to further learn time features. Finally, after passing through a fully connected layer containing C t neurons, add the output to the residual feature value F T_Ra to obtain a new residual feature F New_T_Ra with a dimension size of F×C t , F New_T_Ra That is the time feature matrix output by the current multi-layer perceptron mixer.

[0030] Furthermore, averaging the time context features in the time dimension, then normalizing, and then passing through a fully connected layer to predict the corresponding three-dimensional human pose result, including:

[0031] Averaging the time context features F TCFirst, it is normalized through the LayerNorm layer, and then the mean operation is performed in the time dimension to obtain the final time feature

[0032] The time feature F T_Final is normalized through the LayerNorm layer again, and then followed by a fully connected layer containing J×3 neurons to obtain the final prediction result

[0033] Furthermore, the three-dimensional human pose estimation method based on spatio-temporal context feature perception further includes:

[0034] Construct a loss function:

[0035] where γ represents the predicted result, represents the real data result, and k represents the k-th joint point in the human skeleton.

[0036] Compared with the prior art, a three-dimensional human pose estimation method based on spatio-temporal context feature perception proposed in this application has the following advantages and beneficial effects:

[0037] 1. The scheme based on spatial context features proposed in this application can effectively learn the inherent geometric constraint information of the human skeleton, thereby alleviating the problems of self-occlusion and depth ambiguity in three-dimensional human pose estimation, and further improving the accuracy of three-dimensional human pose estimation.

[0038] 2. Currently, the three-dimensional human pose detection method based on single-frame images has serious jitter problems when detecting in continuous video streams. The scheme based on time context features proposed in this invention can significantly reduce the jittery prediction results and obtain stable prediction results.

[0039] 3. The networks in this application all use simple fully connected layers, with a simple network structure, fast and efficient calculation, saving computing resources, and thus can achieve the effect of real-time prediction. Description of the Drawings

[0040] Figure 1 is a flowchart of a three-dimensional human pose estimation method based on spatio-temporal context feature perception in this application;

[0041] Figure 2 is a schematic diagram of 17 predefined human skeleton joint points;

[0042] Figure 3 is a network framework diagram adopted by the three-dimensional human pose estimation method based on spatio-temporal context feature perception in this application. Detailed Embodiments

[0043] To make the objectives, technical solutions, and advantages of this application more clear and understandable, the following further details this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0044] A 3D human pose estimation method based on spatio-temporal context feature perception provided by this application, as Figure 1 shown, includes:

[0045] Step S1: Input consecutive F frames in a monocular video, detect the human body bounding box, and then use a two-dimensional human pose detector with a cascaded pyramid structure to detect the two-dimensional human joint point coordinates for each frame, and form a two-dimensional human skeleton sequence.

[0046] For 243 consecutive frames in the input monocular video, first use Mask R-CNN for human body bounding box detection, where Mask R-CNN uses ResNet101 as the backbone network, and then use a two-dimensional human pose detector (CPN) with a cascaded pyramid structure for two-dimensional human pose estimation. For CPN, this application uses ResNet-50 with a resolution of 384×288 as the backbone network. And both Mask R-CNN and CPN start from the pre-trained model on COCO and fine-tune the detector on Human3.6M to learn a new set of human joint points and form a two-dimensional human skeleton sequence.

[0047] Step S2: Normalize each two-dimensional human skeleton in the two-dimensional human skeleton sequence, and raise the dimension of the joint point coordinates in the normalized two-dimensional human skeleton to obtain the skeleton feature after dimension raising.

[0048] The normalization process for each two-dimensional human skeleton in the two-dimensional human skeleton sequence includes:

[0049] For each two-dimensional human skeleton in the two-dimensional human skeleton sequence, subtract the two-dimensional coordinates of the hip joint point from the two-dimensional coordinates of each joint point to obtain the normalized two-dimensional human skeleton.

[0050] That is, for the two-dimensional human skeleton sequence generated in step 1, we first process each two-dimensional human skeleton in the two-dimensional human skeleton sequence (as Figure 2 shown), where i represents the i-th human skeleton in seq and perform a normalization operation. The purpose is that we do not focus on the global position of the three-dimensional human skeleton, but on the relative positions between the joint points of the three-dimensional human skeleton. The specific operation is to subtract the two-dimensional coordinates of the hip joint point from the two-dimensional coordinates of each joint point in i x to obtain the normalized two-dimensional human skeleton. (as shown in Figure 3 , the normalized two-dimensional human coordinates).

[0051] Then, perform a dimensionality increase operation on , and pass it through a fully connected layer with 32 neurons as shown in Figure 3 (FC) to increase the dimension of the joint point coordinates in . After the dimensionality increase, the output data dimension is . Among them, the 32 neurons are the dimensions after the dimensionality increase.

[0052] Step S3: Input the skeleton features after the dimensionality increase into the spatial context awareness module to extract the dependency relationship features containing the geometric dependency information between human joint points.

[0053] In this step, the two-dimensional human skeleton after the dimensionality increase is input into the spatial context awareness module to extract the geometric dependency information between human joint points, including the following steps:

[0054] 3.1), First, construct a structure matrix according to the preset human structure.

[0055] As shown in Figure 2 , construct the structure matrix through the following formula

[0056]

[0057] where S (i,p) represents the element in the i-th row and p-th column of the structure matrix S, MD(i, p) represents the flow distance between the i-th human joint point and the p-th human joint point, the flow distance between joint points is determined by the preset human skeleton structure diagram, and K represents a predefined hyperparameter.

[0058] For example, according to Figure 2 , define the flow distance between the left hip and the hip as 1 because they are directly connected, and the flow distance between the left hip and the right hip as 2 because there is a hip joint point between them. In this embodiment, K is preset to 3.

[0059] 3.2), Input the structure matrix S and the skeleton features x new after the dimensionality increase into the spatial context awareness module for skeleton feature learning. The spatial context awareness module is composed of N pose encoders with the same structure connected in series; the structure matrix S and the skeleton features x new obtain a feature matrix after passing through the first pose encoder. The dimension of this feature matrix is the same as that of the skeleton features x new . The input of the latter pose encoder is the feature matrix output by the previous pose encoder and the structure matrix S; after passing through N pose encoders, the output feature is The output features Are normalized through the LayerNorm layer to obtain dependency features containing the geometric dependency information between human joint points

[0060] The structure matrix S and the upsampled skeleton features x new Are input into the spatial context awareness module for skeleton feature learning. The spatial context awareness module is composed of three pose encoders with the same structure in series. The structure matrix S and the upsampled skeleton features x new After passing through the first pose encoder, a feature matrix is obtained. The feature matrix has the same dimension as the skeleton features x new The input of the next pose encoder is the feature matrix output by the previous pose encoder and the structure matrix S. After passing through three pose encoders, the output features And are normalized through the LayerNorm layer to obtain the final features

[0061] Among them, the pose encoder performs the following operations:

[0062] First, the structure matrix S is flattened into a one-dimensional vector with a dimension of 1×J 2 And is input into the skeleton attention module. The skeleton attention module consists of a fully connected layer with J 2 Neurons and a sigmoid activation function, and outputs an attention vector

[0063] The input feature matrix is first passed through the LayerNorm layer, then the dimension is changed to C s ×J through a transpose operation, and then passed through a fully connected layer with J 2 Neurons and a GELU activation function to obtain intermediate features with a dimension of C s ×J 2 , and then the intermediate features are element-wise multiplied by the attention vector W Att To obtain an attention feature matrix. Finally, the attention feature matrix is passed through a fully connected layer with J neurons to obtain a skeleton attention feature matrix W SA With a dimension of C s ×J. Finally, the skeleton attention feature matrix W SA Is transposed to change the dimension to J×C s And added to the input feature x new To obtain a residual feature value W Ra ;

[0064] Then, the residual feature value W Ra Is passed through the LayerNorm layer again, and a layer containing Cs A fully connected layer of neurons and a layer of GELU activation function are used to further learn the skeleton features. Finally, after passing through a fully connected layer containing C s neurons, the output is added to the residual feature value W RA to obtain a new residual feature W New_RA with a dimensionality of J × C s ; W New_RA is the feature matrix output by the current pose encoder.

[0065] Specifically, as Figure 3 shown, first, the structure matrix S is flattened into a one-dimensional vector with a dimension of (1 × 289) and input into the skeleton attention module. The skeleton attention module consists of a fully connected layer of 289 neurons and a layer of sigmoid activation function, and outputs an attention vector The input feature matrix (if it is the first pose encoder, the input is the skeleton feature x new ) first passes through a LayerNorm layer, then undergoes a transpose operation to change the dimension to (32 × 17). Next, it passes through a fully connected layer containing 289 neurons and a layer of GELU activation function to obtain an intermediate feature with a dimension of (32 × 289). Then, this intermediate feature is element-wise multiplied with the attention vector W Att to obtain an attention feature matrix. Finally, the attention feature matrix passes through a fully connected layer containing 17 neurons to obtain a skeleton attention feature matrix W SA with a dimension of (32 × 17). Finally, the skeleton attention feature matrix W SA undergoes a transpose operation to change the dimension to (17 × 32) and is added to the input feature x new to obtain a residual feature value W Ra . Then, the residual feature value W Ra passes through a LayerNorm layer, as well as a fully connected layer containing 32 neurons and a layer of GELU activation function to further learn the skeleton features. Finally, after passing through a fully connected layer containing 32 neurons, the output is added to the residual feature value W RA to obtain a new residual feature W New_RA with a dimension of (17 × 32). W New_RA is the feature matrix output by the current pose encoder.

[0066] Step S4: Input the dependency relationship features into the temporal multi-layer perceptron network module to further extract temporal information in the temporal dimension and obtain temporal context features.

[0067] Inputting the dependency relationship features into the temporal multi-layer perceptron network module to further extract temporal information includes:

[0068] 4.1), Concatenate the dependency relationship features of each two-dimensional human skeleton to form a skeleton feature sequence, and then flatten the second and third dimensions of the skeleton feature sequence into one dimension to form a new skeleton feature sequence.

[0069] Extract the skeleton features of each two-dimensional human skeleton in the two-dimensional human skeleton sequence using step 3 And concatenate each skeleton feature, and then form a skeleton feature sequence Finally, flatten the second and third dimensions of the skeleton feature sequence β into one dimension to form a new skeleton feature sequence

[0070] 4.2), Input the new skeleton feature sequence into the time multi-layer perceptron network module, and normalize the output features to obtain the time context features.

[0071] Input the feature sequence β0 obtained in the previous step into the time context feature perception module to learn the time consistency information between frames. The time multi-layer perceptron network module is composed of multiple multi-layer perceptron mixers with the same structure connected in series. In this embodiment, the time multi-layer perceptron network module is composed of multiple multi-layer perceptron mixers with the same structure connected in series.

[0072] Each multi-layer perceptron mixer performs the following operations::

[0073] First, normalize through the LayerNorm layer, then use the transpose operation to change the input feature dimension to C t ×F, then pass through a fully connected layer containing D s neurons, a GELU activation function layer and a fully connected layer containing F neurons to obtain the intermediate feature with a dimension size of C t ×F, then transpose the intermediate feature to change the dimension to F×C t , and add it to the input feature to obtain the residual feature value

[0074] Then normalize the residual feature value F T_Ra through the LayerNorm layer, and a fully connected layer containing D c neurons and a GELU activation function layer to further learn the time features, and finally pass through a fully connected layer containing C t neurons and add the output to the residual feature value F T_Ra to obtain a new residual feature F New_T_Ra with a dimension size of F×C t , F New_T_Ra is the time feature matrix output by the current multi-layer perceptron mixer.

[0075] The feature sequence β0 passes through the first multi-layer perceptron mixer to obtain a temporal feature matrix, which has the same dimension as β0. The input of the subsequent multi-layer perceptron mixer is the temporal feature matrix output by the previous multi-layer perceptron mixer. After passing through 8 multi-layer perceptron mixers, the output features

[0076] Normalize the output features through the LayerNorm layer to obtain the temporal context features

[0077] For the described temporal multi-layer perceptron network module, the temporal feature matrix (if it is the first multi-layer perceptron mixer, the input is the feature sequence β0) is first normalized through the LayerNorm layer, then the dimension of the feature sequence is changed to (544×243) using the transpose operation, then passed through a fully connected layer with 256 neurons, a GELU activation function layer, and a fully connected layer with 243 neurons to obtain intermediate features with a dimension of (544×243). Then, the intermediate features are transposed to change the dimension to (243×544) and added to the input features to obtain the residual feature value Then the residual feature value F T_Ra is normalized through the LayerNorm layer, passed through a fully connected layer with 512 neurons and a GELU activation function layer to further learn the temporal features. Finally, after passing through a fully connected layer with 544 neurons, the output is added to the residual feature value F T_Ra to obtain a new residual feature F New_T_Ra with a dimension of (243×544). F New_T_Ra is the temporal feature matrix output by the current multi-layer perceptron mixer

[0078] Step S5: Average the temporal context features in the temporal dimension, then normalize them, and then pass them through a fully connected layer to predict the corresponding 3D human pose result

[0079] In this step, the temporal context feature F TC is first normalized through the LayerNorm layer, and then the averaging operation is performed in the temporal dimension to obtain the final temporal feature Then the temporal feature F T_Final is normalized through the LayerNorm layer again, and then followed by a fully connected layer with J×3 neurons to obtain the final prediction result

[0080] Specifically, averaging the temporal context features in the temporal dimension to obtain the 3D human pose result corresponding to the middle frame of the input 2D human skeleton sequence includes the following steps

[0081] The temporal context feature F TC First, it is normalized through a LayerNorm layer, and then a mean operation is performed in the temporal dimension to obtain the final temporal feature The temporal feature F T_Final After being normalized through a LayerNorm layer, it is then followed by a fully connected layer with (51) neurons to obtain the final prediction result

[0082] In a specific embodiment, the three-dimensional human pose estimation method based on spatio-temporal context feature perception of the present application further includes:

[0083] Construct a loss function:

[0084] where γ represents the predicted result, represents the real data result, and k represents the k-th joint point in the human skeleton. Through this loss function, the error between the network prediction result and the real data result can be accurately calculated, and then backpropagated to the neural network to update the network parameters, prompting the neural network to learn useful information and improving the prediction accuracy.

[0085] The above embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A three-dimensional human pose estimation method based on spatio-temporal context feature perception, characterized in that, The three-dimensional human pose estimation method based on spatio-temporal context feature perception includes: Inputting consecutive F frames in a monocular video, detecting the human body bounding box, and then using a two-dimensional human pose detector with a cascaded pyramid structure to detect the two-dimensional human joint coordinates for each frame, and forming a two-dimensional human skeleton sequence; Normalizing each two-dimensional human skeleton in the two-dimensional human skeleton sequence, and raising the dimension of the joint coordinates in the normalized two-dimensional human skeleton to obtain the skeleton features after dimension raising; Inputting the skeleton features after dimension raising into the spatial context perception module to extract the dependency relationship features containing the geometric dependency information between human joints; Inputting the dependency relationship features into the time multi-layer perceptron network module to further extract time information in the time dimension to obtain the time context features; Averaging the time context features in the time dimension, then normalizing, and then passing through a fully connected layer to predict the corresponding three-dimensional human pose result; Among them, the step of inputting the skeleton features after dimension raising into the spatial context perception module to extract the dependency relationship features containing the geometric dependency information between human joints includes: 3.1), First, according to the preset human body structure, construct a structure matrix through the following formula where S (i,p) represents the element in the \(i\)-th row and \(p\)-th column of the structure matrix \(S\), \(MD(i, p)\) represents the flow distance between the \(i\)-th human joint point and the \(p\)-th human joint point, the flow distance between joint points is determined by a preset human skeleton structure diagram, and \(K\) represents a predefined hyperparameter; 3.2), input the structure matrix S and the upsampled skeleton feature x new into the spatial context awareness module for skeleton feature learning. This spatial context awareness module is composed of N pose encoders with the same structure connected in series; the structure matrix S and the upsampled skeleton feature x new pass through the first pose encoder to obtain a feature matrix, which has the same dimension size as the skeleton feature x new The input of the next pose encoder is the feature matrix output by the previous pose encoder and the structure matrix S; after passing through N pose encoders, the output feature The output feature is normalized through the LayerNorm layer to obtain the dependency feature containing the geometric dependency information between human joint points The pose encoder performs the following operations: First, flatten the structure matrix S into a one-dimensional vector with a dimension of 1×J 2 and input it into the skeleton attention module, where the skeleton attention module consists of a fully connected layer with J 2 neurons and a sigmoid activation function, and outputs an attention vector The input feature matrix is first passed through a LayerNorm layer, and then its dimensions are changed to C through a transpose operation. s ×J, and then it passes through a fully connected layer with J 2 neurons and a GELU activation function to obtain intermediate features with dimensions of C s ×J 2 , then this intermediate feature is element-wise multiplied with the attention vector W Att to obtain the attention feature matrix. Finally, the attention feature matrix passes through a fully connected layer with J neurons to obtain the skeleton attention feature matrix W SA with dimensions of C s ×J. Finally, the skeleton attention feature matrix W SA has its dimensions changed to J×C through a transpose operation s and is added to the input feature x new to obtain the residual feature value W Ra ; Then, the residual eigenvalue W Ra passes through the LayerNorm layer, and a fully connected layer containing C s neurons and a GELU activation function to further learn the skeleton features. Finally, after passing through a fully connected layer containing C s neurons, the output is added to the residual eigenvalue W RA to obtain a new residual feature W New_RA with a dimension size of J×C s ; W New_RA is the feature matrix output by the current pose encoder; The step of inputting the dependency relationship features into the time multi-layer perceptron network module to further extract time information in the time dimension to obtain the time context features includes: 4.1), Concatenating the dependency relationship features of each two-dimensional human skeleton to form a skeleton feature sequence, and then flattening the second and third dimensions of the skeleton feature sequence into one dimension to form a new skeleton feature sequence; 4.2) Inputting the new skeleton feature sequence into the time multi-layer perceptron network module, and normalizing the output features to obtain the time context features; The time multi-layer perceptron network module is composed of multiple multi-layer perceptron mixers with the same structure connected in series, and each multi-layer perceptron mixer performs the following operations: First, it is normalized through the LayerNorm layer, and then the transpose operation is used to change the input feature dimension to C t ×F. Then, it passes through a fully connected layer containing D s neurons, a GELU activation function layer, and a fully connected layer containing F neurons to obtain intermediate features with a dimension size of C t ×F. Then, the intermediate features are transposed to change the dimension to F×C t , and added to the input features to obtain the residual feature value Then, the residual eigenvalue F T_Ra is normalized through a LayerNorm layer, and then passes through a fully connected layer with D c neurons and a GELU activation function to further learn temporal features. Finally, after passing through a fully connected layer with C t neurons, the output is added to the residual eigenvalue F T_Ra to obtain a new residual feature F New_T_Ra with a dimension size of F×C t , where F New_T_Ra is the temporal feature matrix output by the current multi-layer perceptron mixer.

2. The three-dimensional human pose estimation method based on spatio-temporal context feature perception according to claim 1, wherein The step of normalizing each two-dimensional human skeleton in the two-dimensional human skeleton sequence includes: For each two-dimensional human skeleton in the two-dimensional human skeleton sequence, the two-dimensional coordinates of each joint are subtracted from the two-dimensional coordinates of the hip joint to obtain the normalized two-dimensional human skeleton.

3. The three-dimensional human pose estimation method based on spatio-temporal context feature perception according to claim 1, wherein The step of averaging the time context features in the time dimension, then normalizing, and then passing through a fully connected layer to predict the corresponding three-dimensional human pose result includes: The time context feature F TC First, it is normalized through a LayerNorm layer, and then a mean operation is performed in the time dimension to obtain the final time feature The time feature F T_Final is then normalized by a LayerNorm layer and then followed by a fully connected layer containing J×3 neurons to obtain the final prediction result 4. The three-dimensional human body pose estimation method based on spatio-temporal context feature perception according to claim 1, wherein The three-dimensional human pose estimation method based on spatio-temporal context feature perception further includes: Construct a loss function: where γ represents the predicted result, represents the true data result, and k represents the k-th joint point in the human skeleton.

Citation Information

Patent Citations

  • Hand posture estimation method based on space-time context learning

    CN111178142A

  • Three-dimensional human body posture estimation method for monocular video

    CN113313731A