Human body posture estimation method based on double-layer hidden space
The dual latent space structure and feature reconstruction unit in human pose estimation methods address the challenge of occluded joint detection by capturing multi-scale inter-joint dependencies, improving accuracy and robustness in human pose estimation.
Patent Information
- Application Number
- CN202510345910.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-15
AI Technical Summary
The existing human pose estimation method based on convolutional neural networks is difficult to accurately detect the position of human joint nodes in occlusion scenarios, and it is not able to fully capture spatial dependence information at different scales.
A double-layer hidden space structure is constructed, and the human body posture is encoded through vector quantization variational autoencoder, and multi-scale feature fusion and feature reconstruction units are used to capture the joint dependence relationship between human body joint nodes and improve detection accuracy.
It improves the accuracy and robustness of human posture estimation, especially in the case of occlusion, which can accurately detect the location of the joint nodes, and enhances the feature expression ability.
Smart Images

Figure CN120318900A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human pose estimation, and in particular to a human pose estimation method based on a double-layer hidden space. Background Art
[0002] Human pose estimation has been widely studied in the field of computer vision, aiming to estimate the positions of human joint points from input data such as images and videos, and has great application prospects in the fields of human-computer interaction, motion analysis, augmented reality, virtual reality, healthcare, etc.
[0003] In recent years, human pose estimation methods based on convolutional neural networks (CNNs) have gradually become the mainstream technology in this field due to their excellent performance advantages. Typical CNN models such as SimpleBaseline, Hourglass, and HRNet can accurately capture local features near human joint points and related context information by extracting rich visual features from input images.
[0004] Although current human pose estimation methods based on convolutional neural networks are highly effective in image feature extraction and can effectively learn discriminant information for each human joint point based on the extracted features, these methods either ignore the joint dependence relationship between human joint points or only model the joint dependence relationship between human joint points at a single scale, failing to fully capture spatial dependence information at different scales. Therefore, for scenes with occlusion, current human pose estimation methods based on convolutional neural networks usually have difficulty accurately detecting the positions of occluded human joint points. Summary of the Invention
[0005] In view of the above situation, the purpose of the present invention is to provide a human pose estimation method based on a double-layer hidden space. By constructing a new double-layer hidden space structure, it models human poses, captures feature information of different scales of human poses, and explores the joint dependence relationship between multi-scale human joint points, including local dependence between adjacent human joint points (such as joint connection constraints) and global dependence between distant joint points (such as limb coordination and symmetry). Based on multi-scale feature fusion, a feature reconstruction unit is proposed. Multi-scale feature fusion enriches the representation of human pose features by using feature information of different scales in the double-layer hidden space and the joint dependence relationship between human joint points. The feature reconstruction unit further enhances the expression ability of human pose features by refining the feature structure, improves the detection accuracy of human joint point positions, including the detection accuracy of occluded joint point positions. Compared with current human pose estimation methods based on convolutional neural networks, the present invention improves the accuracy and robustness of human pose estimation.
[0006] The method of the present invention includes a pose encoding stage and a pose estimation stage:
[0007] S1: The specific steps of the pose encoding stage are as follows:
[0008] S1-1: Train an encoder using a vector quantization variational autoencoder to encode the human pose into discrete token sequences of different scales. At the same time, the joint dependencies between human joint points of different scales are also encoded to obtain a trained codebook and decoder.
[0009] S2: Freeze the codebook and decoder obtained in the pose encoding stage for pose estimation. The specific steps of the pose estimation stage are as follows:
[0010] S2-1: Extract image features from the preprocessed image through a neural network.
[0011] S2-2: First, perform preprocessing on the image features, and then use two classification heads to quantize the image features respectively through the nearest neighbor search algorithm in combination with two codebooks.
[0012] S2-3: Concatenate the two quantized features, input them into the decoder to reconstruct the pose, and output the human pose estimation result.
[0013] In the step S1-1, create and maintain two codebooks: a bottom-layer codebook an upper-layer codebook to represent the double-layer latent space for quantifying pose features of different scales. Among them, the bottom-layer codebook stores local human pose feature information, and the upper-layer codebook stores global human pose feature information. Since different-scale feature information is stored in the two codebooks, the encoding result can retain the joint dependencies between multi-scale human joint points. Among them, V b represents the number of tokens in the upper-layer codebook, V t represents the number of tokens in the bottom-layer codebook, S represents the size of the tokens in the codebook, and the codebook is updated using exponential moving average in combination with the token sequence output by the encoder.
[0014] Construct two encoders, and input the human pose joint point position sequence into the first encoder, where K represents the number of human joint points, and D is the dimension of each joint point. Among them, D = 2 for two-dimensional poses and D = 3 for three-dimensional poses, to obtain the bottom-layer token sequence T b =(t b1 , t b2 , …, t bM ), input the bottom-layer token sequence into the second encoder to obtain the upper-layer token sequence T t =(t t1 , t t2 , …, ttN )。
[0015] Quantify the marker sequences T at different scales using the nearest neighbor search algorithm b =(t b1 , t b2 , …, t bM ), T t =(t t1 , t t2 , …, t tN ), as follows:
[0016]
[0017] where ‖·‖2 represents the L2 norm, and c j is an element in the codebook.
[0018] Concatenate the quantified marker sequences to obtain the final pose feature sequence and input it into the decoder to reconstruct the pose P.
[0019] The loss function in the pose encoding stage includes the reconstruction loss and the commitment loss:
[0020]
[0021] The first part on the right side of the equal sign in the above formula is the reconstruction loss, which is used to calculate the difference between the pose input to the model and the pose reconstructed by the model, where smooth L1 represents the smooth L1 loss, and the second part is two commitment losses, corresponding to different latent spaces respectively, where α and β represent weight coefficients, and sg represents the stop gradient operator, represents the mean square error.
[0022] Based on the marker sequences T output by the encoder b =(t b1 , t b2 , …, t bM ), T t =(t t1 , t t2 , …, t tN) Combine with the element content in the corresponding codebook, and use Exponential Moving Average (EMA) to stably update the codebook, thereby effectively aggregating and adjusting the representations in the codebook, where each token sequence only participates in the update of the corresponding codebook. A token sequence participates in the update of one corresponding codebook. Specifically: The features output by the encoder are matched with the elements in the existing codebook (i.e., part of the subsequent quantization process, just finding the positions in the corresponding codebook without replacement), obtaining the indices of the elements in the corresponding codebook for each token. Then, according to the relationship between the indices and the features, by performing a dot product calculation on the matching information and the tokens output by the encoder, the contribution of each element in the codebook to all tokens in the token sequence output by the encoder is obtained. Subsequently, the Exponential Moving Average (EMA) method is used to update the codebook by combining the contributions and the previously obtained matching relationships.
[0023] In the step S2-1, during the training process, the human detection box is provided by an existing human detection algorithm, which detects and determines the position of the target human from the original image, generates the corresponding human detection box, and according to the position of the human detection box, crops the area containing only the target human from the original image. The cropped single-person human image is directly input into the backbone network for feature extraction and pose estimation. In this way, the training focuses on the learning of human pose features without having to pay attention to other parts of the image; during the inference process, first use the existing human detection algorithm to detect the human detection box in the image, and then input the image within the human detection box into the trained human pose estimation model for pose estimation.
[0024] In the step S2-2, a feature reconstruction unit is used to enhance the expressive ability of the feature map, thereby improving the performance of the classification head, which includes two steps: splitting and reconstruction. The specific steps are as follows:
[0025] Use group normalization to normalize the input feature map, where B is the batch size, F is the channel size, H and W are the height and width of the feature map respectively, as shown in the following formula:
[0026]
[0027] where μ and σ are the mean and standard deviation of X, ε is a small positive constant added for division stability, γ and β are trainable affine transformations, and the trainable parameter γ ∈ R F reflects the importance of the feature channels, and a larger γ value corresponds to a more important feature channel, which contains richer feature information.
[0028] Obtain the normalized weight W through the following formula γ ∈R F for expressing the importance of different feature channels:
[0029]
[0030] Map the weight W γ to the range (0, 1) through the sigmoid function, and set a threshold. Set the weights higher than the threshold to 1 to obtain the information weight W1, and set the weights lower than the threshold to 0 to obtain the non-information weight W2.
[0031] Multiply the standardized feature X out by W1 and W2 respectively to obtain the feature map X1 with rich information and the feature map X2 with less information, successfully splitting the input feature map.
[0032] Next, reconstruct the split feature maps. For the feature map X1 with rich feature information, send it into a 1×1 convolution to compress the channels to reduce the computational amount. Perform grouped convolution and pointwise convolution operations on the compressed X1 respectively, and then sum the output results to obtain the feature map Y1; for the feature map X2 with less feature information, first send it into a 1×1 convolution for channel compression in the same way, then use pointwise convolution operations to generate a feature map with shallow hidden details, and then splice it with the channel-compressed feature map X2 to obtain the feature map Y2 as a supplement to the feature map Y1.
[0033] Add the feature map Y1 and the feature map Y2 to obtain the feature Y. Extract the global channel features through adaptive average pooling, and apply Softmax on the channel dimension to generate attention weights, and then multiply the weights element-wise with the feature Y to dynamically adjust the importance of the channels, thereby highlighting the contribution of the key information channels.
[0034] For the pose estimation stage, use two losses to train the classification head. Both classification heads use the cross-entropy loss function, and the pose reconstruction loss uses the smooth L1 loss to minimize the difference between the predicted pose and the true pose to the greatest extent.
[0035] Advantages of the present invention:
[0036] The present invention models the human pose by constructing a double-layer hidden space. While fusing multi-scale human pose features to improve the accuracy of human pose estimation, it utilizes the joint dependence relationship between multi-scale human joint points to improve the detection accuracy of the occluded joint positions, thereby enhancing the accuracy and robustness of human pose estimation. In addition, the present invention introduces a feature reconstruction unit in the pose estimation stage, enhancing the expression ability of the feature map and further improving the performance of the algorithm. Secondly, the information in the codebook is learned from the training data, so the final pose reconstruction will not make any unrealistic assumptions, solving the ambiguity problem caused by occlusion and further improving the detection accuracy of the occluded joint positions. Description of the Drawings
[0037] Figure 1 This is the overall flowchart of the present invention;
[0038] Figure 2 This is the flowchart of the classification head;
[0039] Figure 3 This is the flowchart of the feature recombination unit;
[0040] Figure 4 This is the result of performing human pose estimation on each person in the first human body image selected from the COCO val2017 dataset using the HRNet human pose estimation method based on the HRNet-W48 network;
[0041] Figure 5 This is the result of performing human pose estimation on each person in the first human body image selected from the COCO val2017 dataset using the method based on the present invention;
[0042] Figure 6 This is the result of performing human pose estimation on each person in the second human body image selected from the COCO val2017 dataset using the HRNet human pose estimation method based on the HRNet-W48 network.
[0043] Figure 7 This is the result of performing human pose estimation on each person in the second human body image selected from the COCO val2017 dataset using the method based on the present invention. Detailed implementation manners
[0044] In order to better clarify the implementation method and structure of the present invention, the present invention will be explained in more detail below in conjunction with the accompanying drawings and embodiments.
[0045] This embodiment implements a human pose estimation method based on a double-layer hidden space. As Figure 1 shown, this method is divided into a pose encoding stage and a pose estimation stage.
[0046] The specific steps of the pose encoding stage are as follows:
[0047] S1-1: Input the sequence of human pose joint position points into the first encoder to obtain the underlying token sequence T b =(t b1 , t b2 , …, t bM ), where M is set to 34. Input the underlying pose token sequence into the second encoder to obtain the upper-layer token sequence T t =(t t1 , t t2 , …, t tN ), where N is set to 17.
[0048] Create and maintain two codebooks: the underlying codebook the upper - layer codebook to represent the double - layer latent space for quantifying pose features at different scales, where V t is set to 512, V b is set to 1024, and S is set to 512.
[0049] Use the nearest - neighbor search algorithm to quantify the labeled sequences T b =(t b1 ,t b2 ,…,t bM ), T t =(t t1 ,t t2 ,…,t tN ), as follows:
[0050]
[0051] where ‖·‖2 represents the L2 norm, and c j is an element in the codebook.
[0052] Concatenate the quantified labeled sequences to obtain the final pose feature sequence and input it into the decoder to reconstruct the pose where both the encoder and the decoder are composed of linear transformation layers and multiple MLP - Mixer blocks.
[0053] The loss function in the pose encoding stage includes the reconstruction loss and the commitment loss:
[0054]
[0055] The first part on the right - hand side of the above equation is the reconstruction loss, which is used to calculate the difference between the pose input to the model and the pose reconstructed by the model. The second part is two commitment losses, corresponding to different latent spaces respectively, where both α and β are set to 15.
[0056] For the pose encoding stage, the model was trained 50 times using the Adam optimizer, with a batch size of 128, an initial learning rate of 1e - 2, a weight decay of 0.15. In the first 500 iterations, the learning rate increased linearly, and then decayed according to the cosine annealing schedule.
[0057] Freeze the codebook and decoder obtained in the pose encoding stage for pose estimation. The specific steps in the pose estimation stage are as follows:
[0058] S2-1: The present invention does not involve the re-design of human body detection algorithms. During the training process, the human body detection box is provided by the existing HRNet human body detection algorithm. This algorithm detects and determines the position of the target human body from the original image, generates the corresponding human body detection box, and crops out the area containing only the target human body from the original image according to the position of the human body detection box. The cropped single-person human body image is directly input into the backbone network for feature extraction and pose estimation. In this way, the training focuses on the learning of human body pose features without having to pay attention to other parts of the image; during the inference process, first use the existing Cascade R-CNN human body detection algorithm to detect the human body detection box in the image, and then input the image within the human body detection box into the trained human body pose estimation model for pose estimation.
[0059] Use Swin Transformer V2 as the backbone network, and this model is pre-trained by SimMIM on ImageNet-1k and further trained on the COCO dataset based on heatmap supervision. In order to reduce the computational load, during the training of the pose estimation stage, the pre-trained backbone network is frozen.
[0060] S2-2: As Figure 2 shown, the features output by the backbone network pass through the classification head and are quantized into the features in the codebook; as Figure 3 shown, a feature reconstruction unit is used in the classification head, and the specific process is as follows:
[0061] Use group normalization to normalize the input feature map where B is set to 32, F is set to 256, and both H and W are set to 8 as shown in the following formula:
[0062]
[0063] where μ and σ are the mean and standard deviation of X, ε is a small positive constant added for division stability, and γ and β are trainable affine transformations. Use the trainable parameter γ ∈ R F to reflect the importance of the feature channels. Larger γ values correspond to more important feature channels, which contain richer feature information.
[0064] Obtain the normalized weight W through the following formula γ ∈ R F to express the importance of different feature channels:
[0065]
[0066] Multiply the weight W γIt is mapped to the range (0, 1) through the sigmoid function. The threshold is set to 0.5. The weights higher than the threshold are set to 1 to obtain the information weight W1, and the weights lower than the threshold are set to 0 to obtain the non-information weight W2.
[0067] The normalized feature X out is multiplied by W1 and W2 respectively to obtain the feature map X1 with rich information and the feature map X2 with less information, successfully splitting the input feature map.
[0068] Next, the split feature maps are reconstructed. For the feature map X1 with rich feature information, it is sent into a 1×1 convolution to compress the channels to reduce the computational amount. Group convolution and pointwise convolution operations are performed on X1 after channel compression respectively, and then the output results are summed to obtain the feature map Y1; for the feature map X2 with less feature information, it is also first sent into a 1×1 convolution for channel compression, then a pointwise convolution operation is used to generate a feature map with shallow hidden details, and then it is concatenated with the feature map X2 after channel compression to obtain the feature map Y2 as a supplement to the feature map Y1.
[0069] The feature map Y1 and the feature map Y2 are added to obtain the feature Y. Global channel features are extracted through adaptive average pooling, and Softmax is applied in the channel dimension to generate attention weights, and then the weights are multiplied element-wise with the feature Y to dynamically adjust the importance of the channels, thus highlighting the contribution of the key information channels.
[0070] S2-3: Concatenate the features obtained after quantization and input them into the decoder to reconstruct the pose.
[0071] For the pose estimation stage, two losses are used to train the classification head:
[0072] Both classification heads use the cross-entropy loss function:
[0073]
[0074] where represents the result predicted by the classification head, and L represents the result obtained by inputting the true pose into the encoder;
[0075] The pose reconstruction loss is used to minimize the difference between the predicted pose and the true pose to the greatest extent:
[0076]
[0077] The total loss function is as follows:
[0078]
[0079] For the pose estimation stage, the model was trained 210 times using the Adam optimizer, with a batch size of 128, an initial learning rate of 8e-4, and a weight decay of 0.05. In the first 500 iterations, the learning rate increased linearly, and then decayed according to the cosine annealing schedule.
[0080] To further illustrate the feasibility and effectiveness of the present invention, a large number of experiments were conducted on the widely used and large-scale COCO dataset and MPII dataset. In the experimental results, the average precision AP, AP 50 and AP 75 were used as evaluation metrics for the COCO dataset, where AP 50 and AP 75 refer to the average precision when the intersection over union (IoU) threshold is 50% and 75% respectively, which are used to more finely evaluate the model performance. The PCKh (Percentage of Correct Keypoints with head normalization) score was used as the evaluation metric for the MPII dataset, that is, the percentage of correct keypoints calculated under head normalization. Among them, PCKh@0.5 and PCKh@0.1 respectively represent the proportion of correct predictions when the Euclidean distance between the keypoint positions predicted by the model and the true annotations is less than 0.5 times and 0.1 times the head size. PCKh@0.1 shows the model's ability to require higher precision, demanding smaller prediction errors.
[0081] Table 1 shows the results of the method of the present invention and other state-of-the-art human pose estimation methods on the COCO test-dev2017 and COCO val2017 datasets. Among them, the method PCT is the reproduced result, and the experimental environment of the method PCT, including hyperparameter settings such as learning rate, batch size, and number of training epochs, is the same as that of the method of the present invention. Among them, ResNet-152 represents that the depth of the Residual Network (ResNet) is 152, HRNet-W48 represents using the HRNet network as the backbone network with 48 channels, ViT-Base represents using the Base scale of the backbone network, and Swin-Base represents using the Base scale of the Swin Transformer V2 backbone network.
[0082] Table 1 Results of the method of the present invention and other state-of-the-art human pose estimation methods on the COCO test-dev2017 and COCO val2017 datasets
[0083]
[0084] As can be seen from the data in Table 1, the effect of the method of the present invention is better than other methods on both the COCO test-dev2017 and COCO val2017 datasets.
[0085] Table 2 shows the results of the method of the present invention and the existing human pose estimation methods on the MPII validation set, where the evaluation index used for each joint point is PCKh@0.5.
[0086] Table 2 Results of the method of the present invention and the existing human pose estimation methods on the MPII validation set
[0087]
[0088] As can be seen from the data in Table 2, the detection accuracy of the method of the present invention for shoulders, elbows and knees is higher than that of the HRNet method and the PCT method, and the average index is higher than that of the HRNet method. Although the index PCKh@0.5 is the same as the method, for the more stringent index PCKh@0.1, the method of the present invention is 0.8% higher than the PCT method.
[0089] Figure 4 、 Figure 5 Correspondingly, the results of the existing HRNet human pose estimation method based on the HRNet-W48 network and the method of the present invention for human pose estimation of each human body in the first selected human body image in the COCO val2017 dataset are given. Comparison Figure 4 and Figure 5 , it can be clearly seen that the lower bodies of the four people on the right side of the figure are blocked by the table. Among them, the method of the present invention can more accurately detect the positions of the blocked lower body joint points, while the lower body joint points detected by the HRNet human pose estimation method based on the HRNet-W48 network are stacked horizontally together, which is obviously unreasonable and cannot accurately detect the positions of the blocked lower body joint points.
[0090] Figure 6 、 Figure 7 Correspondingly, the results of the existing HRNet human pose estimation method based on the HRNet-W48 network and the method of the present invention for human pose estimation of each human body in the second selected human body image in the COCO val2017 dataset are given. Comparison Figure 6 and Figure 7, it can be clearly seen that the lower body of the woman carrying the suitcase in the middle is blocked by the suitcase. The method of the present invention can accurately detect the position of the occluded lower body joint points, while the HRNet human pose estimation method based on the HRNet-W48 network cannot accurately detect the position of the occluded lower body joint points, and the detection result is unreasonable. For the man on the left side of the figure whose large area is blocked by the backpack, the method of the present invention more accurately detects the joint point positions of the occluded part and successfully detects the correct posture of standing upright, while the HRNet human pose estimation method based on the HRNet-W48 network still cannot accurately detect the joint point positions of the occluded part.
[0091] Through the above experiments, it is fully verified that when the human joint points are severely occluded, the method of the present invention can also accurately detect the positions of the occluded human joint points, reflecting the advantages of the present invention.
[0092] The above content is only a detailed description of the method provided by the present invention and does not limit the protection scope of the present invention. For those skilled in the art, various improvements and changes can be made without departing from the spirit and scope of the present invention. Therefore, the technical solutions obtained by equivalent substitution or equivalent transformation of the present invention should all be included in the protection scope of the present invention.
Claims
1. A human pose estimation method based on a double-layer hidden space, characterized in that, It includes the following steps: S1: Train an encoder using a vector quantization variational autoencoder to encode the joint dependencies between human postures and human joint points into discrete token sequences of different scales, obtaining a trained codebook and a decoder; S2: Freeze the obtained codebook and decoder and perform pose estimation: S2-1: Extract image features from the preprocessed image through a neural network; S2-2: First perform preprocessing on the image features, and then use two classification heads to quantize the image features respectively using the nearest neighbor search algorithm in combination with two codebooks; S2-3: Concatenate the two features obtained after quantization, input them into the decoder to reconstruct the pose, and output the human pose estimation result.
2. The human body pose estimation method based on a double-layer hidden space according to claim 1, wherein, The specific implementation process of step S1 is as follows: Create and maintain two codebooks: the underlying codebook the upper-level codebook represent a two-layer latent space, quantifying pose features at different scales, where the underlying codebook stores local human pose feature information and the upper-level codebook stores global human pose feature information, where V b represents the number of labels in the upper-level codebook, V t represents the number of labels in the underlying codebook, S represents the size of the labels in the codebook, and the codebook is updated using exponential moving average in combination with the label sequence output by the encoder; Construct two encoders for the sequence of human body pose joint point positions Input it into the first encoder, where K represents the number of human body joints and D is the dimension of each joint, to obtain the underlying token sequence T b =(t b1 ,t b2 ,…,t bM ), and input the underlying token sequence into the second encoder to obtain the upper-layer token sequences T at different scales t =(t t1 ,t t2 ,…,t tN ); Quantify the token sequences T at different scales using the nearest neighbor search algorithm b , T t as follows: where ‖·‖2 represents the L2 norm, and c j is an element in the codebook; The quantized marker sequence is concatenated to obtain the final pose feature sequence and input into the decoder to reconstruct the pose 3. The human body pose estimation method based on a double-layer hidden space according to claim 2, wherein The loss function in the encoding process includes a reconstruction loss and a commitment loss: The first part on the right side of the equal sign in the above formula is the reconstruction loss, which is used to calculate the difference between the pose input to the model and the pose reconstructed by the model, where smooth L1 represents the smooth L1 loss. The second part is the two submission losses, which correspond to different latent spaces respectively. Here, α and β represent the weight coefficients, and sg represents the stop gradient operator. represents the mean squared error.
4. The human pose estimation method based on a double-layer latent space according to claim 3, characterized in that The preprocessing operation in step S2-2 is specifically: Normalize the input feature map using group normalization where B is the batch size, F is the channel size, and H and W are the height and width of the feature map respectively, as shown in the following formula: μ and σ are the mean and standard deviation of X, ε is a positive constant, and γ and β are trainable affine transformations; The normalized weight W is obtained by the following formula γ ∈R F , which is used to express the importance of different feature channels: Map the weight W γ to the range (0, 1) through the sigmoid function, set a threshold, set the weights above the threshold to 1 to obtain the information weight W1, and set the weights below the threshold to 0 to obtain the non-information weight W2; The standardized feature X out is multiplied by W1 and W2 respectively to obtain the feature map X1 with rich information and the feature map X2 with less information. For the feature map X1, it is fed into a 1×1 convolution to compress the channels to reduce the computational amount. Then, grouped convolution and pointwise convolution operations are performed on the compressed X1 respectively, and the output results are summed to obtain the feature map Y1. For the feature map X2, it is also first fed into a 1×1 convolution for channel compression, and then a pointwise convolution operation is used to generate a feature map with shallow hidden details, which is then concatenated with the channel-compressed feature map X2 to obtain the feature map Y2 as a supplement to the feature map Y1. Add the feature map Y1 and the feature map Y2 to obtain the feature Y, extract global channel features through adaptive average pooling, apply Softmax in the channel dimension to generate attention weights, and then multiply the weights element-wise with the feature Y.
5. The human body pose estimation method based on a double-layer hidden space according to claim 4, wherein, In step S2, both classification heads use the cross-entropy loss function, and the pose reconstruction loss uses the smooth L1 loss to minimize the difference between the predicted pose and the true pose to the greatest extent.