Single-stage high-resolution human body posture estimation method for solving shielding
By combining Swin Transformer V2 and HRNet structures with a single-stage high-resolution human posture estimation method, the accuracy and robustness of human posture estimation in complex scenarios are solved, and more efficient and real-time multi-person posture estimation is achieved.
Patent Information
- Application Number
- CN202510028910.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-06
AI Technical Summary
Existing human pose estimation methods have low accuracy and robustness when dealing with occlusion and background interference in complex scenarios, especially in multi-person pose estimation.
A single-stage high-resolution human pose estimation method was used to combine the feature extraction network of Swin Transformer V2 and HRNet structures, and pose estimation was performed using Contextual Instance Decoupling (CID) head. This method improves the pose estimation ability in complex scenarios through multi-scale feature fusion and self-attention mechanism.
It significantly improves the real-time and robustness of multiple human posture estimation in complex scenarios, can effectively deal with occlusion and background interference, and improves the estimation accuracy of key points.
Smart Images

Figure CN119942644A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of computer vision, and in particular to a single-stage high-resolution human body posture estimation method for solving occlusion. Background Art
[0002] Human posture estimation is an important research direction in the field of computer vision and is widely used in many fields such as robotics, motion analysis, ergonomics, virtual reality, and human-computer interaction.
[0003] Existing methods have limitations in dealing with partial occlusion, real-time processing requirements, and multi-person pose estimation. In complex scenes, occlusion and background interference can significantly reduce the accuracy and robustness of pose estimation. Therefore, it is urgent to propose more robust and efficient methods to solve these problems.
[0004] The top-down method (arXiv:2012.13392v5) usually detects the entire human body first, and then performs key point detection. Occlusion will seriously affect the accuracy of the entire body detection, thereby affecting the pose estimation results. Since this method relies on the complete human body outline, the top-down method is relatively weak in dealing with occlusion, and the two-stage detection method will consume more resources and time, which is not conducive to real-time monitoring scenarios. Compared with the two-stage pose estimation method, the one-stage method eliminates the top-down preprocessing and the bottom-up post-processing. Regardless of the number of instances in the image, the inference time is consistent, which is convenient for practical application. Summary of the invention
[0005] In order to overcome the above technical problems, the object of the present invention is to provide a single-stage high-resolution human posture estimation method for solving occlusion, based on a single-stage algorithm SHR (Swin High-Resolution Network) to improve the real-time and robustness of multiple human posture estimation in complex scenes.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is:
[0007] A single-stage high-resolution human pose estimation method for solving occlusion, comprising the following steps;
[0008] Step 1: Obtain the CrowdPose public dataset for training and evaluating human pose estimation methods in multi-person crowded occlusion scenes;
[0009] The public dataset is divided into trainval images for training and validating the model, and test set images for evaluating the effect of the trained model;
[0010] Step 2: Input the trainval image and corresponding labels into the SHR model to obtain the trained SHR human posture model;
[0011] Step 3: Use the test set images to evaluate the quality of the SHR human posture model. If the training result of the SHR human posture model is poor, try to adjust the network parameters and retrain to obtain the optimal model; use the average precision and the mean of the average precision as evaluation indicators to evaluate the model performance and obtain the performance of the improved algorithm;
[0012] Step 4: Use the SHR human pose model with good evaluation results to perform human pose estimation.
[0013] The network structure of the SHR model is divided into five modules: input layer (input), feature extraction network (backbone), neck layer (neck), head (head) and output layer (output);
[0014] The input layer first inputs the trainval image and the corresponding labels of the key points of the human body, and uses the encoder to encode the image and the corresponding labels into a heat map mask;
[0015] Backbone is used to extract the features of key points of the human body;
[0016] The neck uses a multi-scale feature layer splicing module (FeatureMapProcessor) to perform multi-scale feature fusion on the features extracted by the backbone, and passes the fused features to the head prediction layer;
[0017] The head is responsible for model prediction; the Contextual Instance Decoupling (CID) head is used for human posture estimation;
[0018] The backbone module combines Swin Transformer V2 with the HRNet structure. The self-attention mechanism learns the relationship between different regions of the image, and the HRNet branch captures valuable information at different scales.
[0019] The output layer is used to output images with human key points.
[0020] Preferably, in step 2, the specific operations include the following steps:
[0021] Step 2.1: The feature extraction network of the SHR model combines the Swin Transformer V2 with the HRNet structure;
[0022] In the initial stage of HRNet, the bottleneck block is used to extract the initial high-resolution feature map. In the following three stages, feature maps of different resolutions are gradually introduced, and finally four branches are formed. Each branch contains multiple convolution blocks for in-depth feature extraction. After each stage is completed, the network fuses feature maps of different resolutions through convolution, upsampling or downsampling operations. Specifically, the high-resolution branch receives the upsampling information of the low-resolution branch; the low-resolution branch receives the downsampling information of the high-resolution branch; this operation is repeated many times in each stage, so that the feature map of each branch is integrated with multi-scale information from other branches;
[0023] The specific operation of combining Swin Transformer V2 with HRNet structure is to use HRNet structure, keep the bottleneck block in the initial stage unchanged, replace the convolution block in the last three stages with Transformer block to build the pose estimation network, the self-attention mechanism in Transformer block learns the relationship between different areas of the image, and the HRNet structure captures valuable information at different scales. The combination of the two improves the ability to infer occluded key points in crowded scenes;
[0024] Step 2.2: Transformer block introduces the adapted Swin Transformer, namely Swin Transformer V2, and uses the shift window partitioning method to map the feature layer obtained by the previous bottleneck block operation Divide it into a group of non-overlapping small windows, where each window size is K×K, and perform multi-head self-attention independently in each small window. The multi-head self-attention formula of the i-th window is:
[0025]
[0026]
[0027] in, h∈{1,…,H}, H represents the number of self-attention heads, C represents the number of channels, and N represents the input resolution. Represents the output representation of multi-head self-attention (MHSA), and then aggregates the information of each window and merges the output. At the same time, the residual post-normalization technology is used in the self-attention operation to replace the prenorm configuration of the Swin Transformer self-attention operation and the scaled cosine attention method is used to replace the dot product attention in the Swin Transformer self-attention operation. The logarithmic interval continuous position deviation method is used instead of the parameterized method to solve the problem of performance degradation when the model is transferred across window resolutions; Step 2.3: In the feature layers of four resolution sizes obtained by the feature extraction network, they are upsampled to the same resolution size, and then fused and added, and input into the Contextual Instance Decoupling (CID) head, and the output is an n-channel heat map, which represents the probability distribution map of each key point (n represents the number of key points); Contextual instance decoupling (CID) includes an instance information abstraction module and a global feature decoupling module, which decouples multi-person feature maps into a set of instance-aware feature maps, each of which represents clues about a specific person, and retains contextual clues to infer the key points of the person.
[0028] Step 2.4: Send the trainval dataset in step 1 to the SHR model for training to obtain a human posture estimation model.
[0029] The step 2.2 is specifically as follows:
[0030] Specifically, in Swin Transformer, the pre-normalization is used, that is, the normalization operation is located before the residual connection, in the following form:
[0031] y=x+F(Norm(x))
[0032] Swin Transformer V2 is changed to post-normalization, that is, the normalization operation is moved to the residual connection, and the form is as follows:
[0033] y=Norm(x+F(x))
[0034] Where x is the input, F(x) is the output of the residual block, and Norm is the normalization layer;
[0035] The scaled cosine attention is specifically:
[0036] The attention score is calculated using cosine similarity, and a learnable scaling factor τ is introduced to make the attention range more flexible. Cosine attention:
[0037]
[0038] in represents the cosine similarity between query Q and key K. τ is a learnable parameter. The normalized vector is used to calculate the cosine similarity instead of the dot product. The cosine similarity is divided by the learnable parameter τ to adjust the range of attention distribution. The value V is weighted and summed according to the attention weight.
[0039] The logarithmically spaced continuous position deviations are:
[0040] In Swin Transformer V1, the position deviation is predefined by a parameterized method, and each pair of positions has a fixed deviation value. Swin Transformer V2 introduces a logarithmically spaced continuous deviation method to dynamically calculate the relative position deviation. The formula is:
[0041] Δ i,j =log(|ij|+1)
[0042] where Δ i,j Represents the relative deviation between the i-th and j-th positions. For each pair of positions (i, j), the relative distance |ij| is calculated. The distance is transformed using a logarithmic function. In the attention calculation, the deviation Δ i,j Added to the attention score, the formula is:
[0043]
[0044] The step 2.3 is specifically as follows:
[0045] Instance information abstraction extracts locations and features to represent each person, and global feature decoupling operates on the original feature maps extracted by the feature extraction network to generate instance-aware feature maps, each of which is used to estimate the heat map and key points of a person respectively.
[0046] In the step 3, the experimental results of the SHR model are compared with the classic top-down posture estimation method HRNet, the posture estimation method based on the attention mechanism, and the posture estimation method combining HRNet with the attention mechanism;
[0047] Verify the effectiveness of SHR under pre-training, including the performance of SHR loaded with pre-trained weights.
[0048] Beneficial effects of the present invention:
[0049] 1. The feature extraction network in the model proposed in the present invention combines Swin Transformer V2 with the HRNet structure. HRNet is a posture estimation method based on convolutional neural networks, which uses parallel branches to extract feature maps of different resolutions. It introduces a feature sharing strategy that uses upsampling and downsampling operations to propagate information between feature maps of different resolutions. This enables HRNet to preserve information from different image scales, maintain a global perception of large-scale structures, and improve the recognition accuracy of small-scale details. It is worth noting that when dealing with complex data such as mutual occlusion of human limbs in crowded scenes, feature sharing between maps of different scales ensures the preservation of multi-scale information. Compared with conventional deep convolutional neural network methods, HRNet effectively reduces parameters and network layers and improves training efficiency. In the method of this application, HRNet is used as a parallel branch, and the convolution blocks in the last three levels are replaced with Transformer blocks to construct a posture estimation network. The self-attention mechanism learns the relationship between different areas of the image, and the HRNet branch captures valuable information at different scales. The two are combined to improve the ability to infer occluded key points in crowded scenes.
[0050] 2. The SHR model of the present invention uses Contextual Instance Decoupling (CID) head for human posture estimation. The feature layers of four resolution sizes obtained by the feature extraction network are upsampled to the same resolution size, then fused and added, and input into the CID head, and the output is an n-channel heat map, which represents the probability distribution map of each key point (n represents the number of key points). Contextual instance decoupling (CID) decouples multi-person feature maps into a set of instance-aware feature maps, each of which represents clues of a specific person, and retains context clues to infer the key points of the person. Specifically, instance information abstraction (IIA) extracts location and features to represent each person. Global feature decoupling (GFD) operates on the original feature map to generate instance-aware feature maps, each of which is used to estimate the heat map and key points of a person. CID head uses instance decoupling to decouple the key point prediction of each target instance from the influence of other instances. This decoupling helps to avoid interference from other instances, especially occlusion behavior, on the key point estimation of the target instance. By generating feature representations for each instance in the feature map and then combining these representations with context features for key point prediction, the position of key points can be estimated more effectively even in complex occlusion scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 The present invention is a single-stage high-resolution human body posture estimation algorithm flow chart.
[0052] Figure 2 This is a feature extraction network structure diagram that combines the Swin Transformer V2 algorithm of the present invention with the HRNet structure.
[0053] Figure 3 This is a diagram of the CID detection network structure introduced by the algorithm of the present invention.
[0054] Figure 4 These are some pictures of the testing phase of the present invention. DETAILED DESCRIPTION
[0055] The present invention will be further described in detail below in conjunction with the accompanying drawings.
[0056] like Figure 1 As shown, a single-stage high-resolution human posture estimation method for solving occlusion includes the following steps:
[0057] Step 1: Get the CrowdPose public dataset, which is a dataset specifically used to train and evaluate human pose estimation methods in multi-person crowded occlusion scenes. It includes a training set and a validation set for training and validating models, as well as a test set for evaluating the effect of the trained model.
[0058] Step 2: Input the training set images and the labels of the corresponding human posture key point position coordinates into the SHR model to obtain the trained SHR human posture model.
[0059] The network structure of the SHR model is divided into five modules: input layer (input), feature extraction network (backbone), neck layer (neck), head (head) and output layer (output);
[0060] The input layer first inputs the trainval image and the corresponding labels of the key points of the human body, and uses Encoder to encode the image and the corresponding labels into a heat map mask;
[0061] Backbone is used to extract the features of key points of the human body;
[0062] The neck is responsible for multi-scale feature fusion of the features extracted by the backbone and passing these features to the head prediction layer;
[0063] The head is responsible for model prediction; the backbone module combines Swin Transformer V2 with the HRNet structure, the self-attention mechanism learns the relationship between different regions of the image, and the HRNet branch captures valuable information at different scales;
[0064] The neck layer uses a multi-scale feature layer splicing module (FeatureMapProcessor); the head prediction layer uses the Contextual Instance Decoupling (CID) head for human posture estimation;
[0065] The output layer is used to output images with human key points.
[0066] Furthermore, in step 2, the specific operation of the human body posture estimation algorithm includes the following steps:
[0067] Step 2.1: The backbone network of the SHR model combines the Swin Transformer V2 with the HRNet structure, which is a pose estimation method based on convolutional neural networks. It includes the Stem module for extracting the initial high-resolution feature map. Through some convolution operations, the network first generates a high-resolution feature map as a basis; the multi-scale module processes the feature map in parallel through multiple branches to extract information at different scales. Each branch processes feature maps of different resolutions and fuses them through certain strategies (such as upsampling, downsampling, convolution, etc.); the multi-scale feature maps are fused through cross-scale convolution operations to maintain information consistency and context richness. The feature maps of each branch exchange information with other branches to optimize the multi-scale feature representation; in the last part of the network, the feature maps after multi-scale processing and information fusion are integrated together to form the final output. This enables HRNet to preserve information from different image scales, maintain a global perception of large-scale structures, and improve the recognition accuracy of small-scale details. It is worth noting that when dealing with complex data such as human limbs occluding each other in crowded scenes, feature sharing between maps of different scales ensures the preservation of multi-scale information. Compared with conventional deep convolutional neural network methods, HRNet effectively reduces parameters and network layers and improves training efficiency. Using HRNet as a parallel branch, the convolutional blocks in the last three levels are replaced with Transformer blocks to build a posture estimation network. The self-attention mechanism learns the relationship between different areas of the image, and the HRNet branch captures valuable information at different scales. The combination of the two improves the ability to infer occluded key points in crowded scenes;
[0068] Step 2.2: The Transformer block in the SHR model feature extraction network introduces the adapted SwinTransformer, namely Swin Transformer V2. Swin Transformer V2 uses the shift window partitioning method to improve modeling efficiency. At the same time, the residual post-normalization technology is used in the self-attention operation to replace the prenorm configuration in Swin Transformer and the scaled cosine attention method is used to replace the dot product attention in Swin Transformer to improve the stability of the model. The logarithmic interval continuous position deviation method is used instead of the parameterization method to solve the problem of performance degradation when transferring the model across window resolutions, so as to improve the robustness of the model and make the model training more stable;
[0069] Step 2.3: The feature layers of four resolution sizes obtained by the feature extraction network are upsampled to the same resolution size, fused and added, and input into the Contextual Instance Decoupling (CID) head, and the output is an n-channel heat map, which represents the probability distribution map of each key point (n represents the number of key points). Contextual instance decoupling (CID) includes an instance information abstraction module and a global feature decoupling module, which can decouple multi-person feature maps into a set of instance-aware feature maps, each of which represents clues about a specific person and retains contextual clues to infer the key points of the person.
[0070] Specifically, Instance Information Abstraction (IIA) extracts locations and features to represent each person, and Global Feature Decoupling (GFD) operates on the original feature maps extracted by the feature extraction network to generate instance-aware feature maps, each of which is used to estimate the heatmap and key points of a person respectively;
[0071] Step 2.4: Send the trainval dataset in step 1 to the SHR model for training to obtain the optimal human posture estimation model.
[0072] Step 3: After the training is completed, the image or video to be predicted is input into the optimal model to predict the key points of the person in the image or video, and finally obtain the recognition prediction result;
[0073] Furthermore, in step 3, the performance of the algorithm in this paper is evaluated with average precision (AP), mean average precision (mAP) and the number of frames per second (FPS) of predicted pictures as evaluation indicators. The PR curve is formed with recall (R) as the horizontal axis and precision (P) as the vertical axis. AP is the area under the PR curve, which indicates the average accuracy of all prediction results in the same category; mAP indicates that this is a key indicator to measure the overall performance of the pose estimation model on the entire test set. The mAP value is calculated by averaging the AP values of all key points; FPS is an important indicator for evaluating model performance. The higher the FPS value, the faster the model runs.
[0074] In summary. The present invention proposes a new posture estimation method SHR, whose feature extraction network combines the advantages of Swin Transformer V2 and HRNet models, and can effectively estimate the posture of multiple people. By introducing the self-attention mechanism, the method can accurately weight and focus on key key points, thereby achieving comprehensive modeling of postures with occlusions. The method significantly improves the posture estimation task with occlusions. The method combines the attention branch with the HRNet branch to extract high-resolution features and high-quality global features. Compared with traditional CNN-based methods, the method proposed in the present invention improves the accuracy and robustness of posture estimation with occlusions. During the training process, various data enhancement techniques, appropriate loss functions and optimizers are used to improve the generalization and convergence speed of the model.
[0075] like Figure 2 Shown is a feature extraction network structure diagram combining the Swin Transformer V2 algorithm of the present invention with the HRNet structure.
[0076] Its structure adopts the HRNet structure. Downsampling is performed at the beginning of each stage, and upsampling and downsampling are performed at the end to fuse the feature layers. In each stage, the solid line represents the transmission of the convolution block, and the dotted line represents the transmission of the attention module. Starting from a high-resolution convolution system as the first stage, a high-to-low resolution stream is gradually added as a new stage. Multi-resolution streams are connected in parallel. The main body consists of a series of stages. In each stage, the feature representation of each resolution stream is updated by multiple attention modules, and information of different resolutions is repeatedly exchanged through the convolution multi-scale fusion module.
[0077] like Figure 3The figure shows the structure of the CID detection network introduced in the present invention. Context instance decoupling (CID) includes an instance information abstraction (IIA) module and a global feature decoupling (GFD) module. The CID head identifies and describes the external features and spatial positions of each person in the IIA module, and obtains the central position coordinates and external feature information of each person. The GFD module uses the central position coordinates and external feature information obtained by the IIA module to supervise the attention mechanism to decouple the original feature map into multiple instance-aware feature maps;
[0078] Specifically, different people are decoupled into different spatial positions and channels of the feature map. After obtaining the feature maps processed by the IIA module and the GFD module, this set of feature maps representing different people is input into the heat map module to obtain the key point heat maps of each person.
[0079] like Figure 4 The following are some test results of the SHR model with the best training result. It can be seen from the figure that the SHR model has a good effect in detecting occluded key points.
Claims
1. A single-stage high-resolution human pose estimation method for solving occlusion, characterized in that: The steps include: Step 1: Obtain the CrowdPose public dataset for training and evaluating human pose estimation methods in multi-person crowded occlusion scenes; The public dataset is divided into trainval images for training and validating the model, and test set images for evaluating the effect of the trained model; Step 2: Input the trainval image and corresponding labels into the SHR model to obtain the trained SHR human posture model; Step 3: Use the test set images to evaluate the quality of the SHR human posture model. If the training result of the SHR human posture model is poor, try to adjust the network parameters and retrain to obtain the optimal model; use the average precision and the mean of the average precision as evaluation indicators to evaluate the model performance and obtain the performance of the improved algorithm; Step 4: Use the SHR human pose model with good evaluation results to perform human pose estimation.
2. A single-stage high-resolution human posture estimation method for solving occlusion according to claim 1, characterized in that: In step 2, the network structure of the SHR model is divided into five modules: input layer (input), feature extraction network (backbone), neck layer (neck), head (head) and output layer (output); The input layer first inputs the trainval image and the corresponding labels of the key points of the human body, and uses Encoder to encode the image and the corresponding labels into a heat map mask; Backbone is used to extract the features of key points of the human body; The neck uses a multi-scale feature layer splicing module (FeatureMapProcessor) to perform multi-scale feature fusion on the features extracted by the backbone, and pass the fused features to the head prediction layer; The head is responsible for model prediction, and the ContextualInstance Decoupling (CID) head is used for human posture estimation; The backbone module combines Swin Transformer V2 with the HRNet structure. The self-attention mechanism learns the relationship between different regions of the image, and the HRNet branch captures valuable information at different scales. The output layer is used to output images with human key points.
3. A single-stage high-resolution human posture estimation method for solving occlusion according to claim 2, characterized in that: In step 2, the specific operations include the following steps: Step 2.1: The feature extraction network of the SHR model combines the Swin Transformer V2 with the HRNet structure; In the initial stage of HRNet, the bottleneck block is used to extract the initial high-resolution feature map. In the following three stages, feature maps of different resolutions are gradually introduced, and finally four branches are formed. Each branch contains multiple convolution blocks for in-depth feature extraction. After completing each stage, the network fuses feature maps of different resolutions through convolution, upsampling or downsampling operations. The high-resolution branch receives the upsampling information of the low-resolution branch; the low-resolution branch receives the downsampling information of the high-resolution branch. This operation is repeated multiple times in each stage, so that the feature map of each branch is integrated with multi-scale information from other branches. Step 2.2: Swin Transformer V2 introduces the adapted SwinTransformer for the Transformer block; Using the shift window partitioning method, the feature layer map obtained by the bottleneck block operation is Divide it into a group of non-overlapping small windows, where each window size is K×K, and perform multi-head self-attention independently in each small window. The multi-head self-attention formula of the i-th window is: in, h∈{1,…,H}, H represents the number of self-attention heads, C represents the number of channels, and N represents the input resolution. Represents the output representation of Multi-Head Self-Attention (MHSA), after which the information of each window is aggregated and output together; At the same time, residual post-normalization and scaled cosine attention methods are used in self-attention operations to improve the stability of the model. The logarithmic interval continuous position deviation method is used to solve the problem of performance degradation when transferring the model across window resolutions, so as to improve the robustness of the model and make the model training more stable. Step 2.3: The feature layers of four resolution sizes obtained by the four stages of the HRNet structure in the feature extraction network are upsampled to the same resolution size, fused and added, and input into the Contextual Instance Decoupling (CID) head. The output is an n-channel heat map, which represents the probability distribution map of each key point; Context-instance decoupling (CID) includes an instance information abstraction module and a global feature decoupling module. It decouples multi-person feature maps into a set of instance-aware feature maps, where each map represents clues about a specific person and retains context clues to infer the key points of the person. The output result of the human posture model is the key points of the person. Step 2.4: Send the trainval dataset in step 1 to the SHR model for training to obtain a human posture estimation model.
4. A single-stage high-resolution human posture estimation method for solving occlusion according to claim 3, characterized in that: In step 2.1: The specific operation of combining Swin Transformer V2 with the HRNet structure is to use the HRNet structure, keep the bottleneck block in the initial stage unchanged, and build a posture estimation network through the Transformer block in the last three stages. The self-attention mechanism in the Transformer block learns the relationship between different areas of the image, and the HRNet structure captures valuable information at different scales. The combination of the two improves the ability to infer occluded key points in crowded scenes.
5. The single-stage high-resolution human posture estimation method for solving occlusion according to claim 3 is characterized in that: The step 2.2 is specifically as follows: Specifically, in Swin Transformer, the pre-normalization is used, that is, the normalization operation is located before the residual connection, in the following form: y=x+F(Norm(x)) Swin Transformer V2 is changed to post-normalization, that is, the normalization operation is moved to the residual connection, and the form is as follows: y=Norm(x+F(x)) Where x is the input, F(x) is the output of the residual block, and Norm is the normalization layer; The scaled cosine attention is specifically: The attention score is calculated using cosine similarity, and a learnable scaling factor τ is introduced to make the attention range more flexible. Cosine attention: in represents the cosine similarity between query Q and key K. τ is a learnable parameter. The normalized vector is used to calculate the cosine similarity instead of the dot product. The cosine similarity is divided by the learnable parameter τ to adjust the range of attention distribution. The value V is weighted and summed according to the attention weight. The logarithmically spaced continuous position deviations are: In Swin Transformer V1, the position deviation is predefined by a parameterized method, and each pair of positions has a fixed deviation value. Swin Transformer V2 introduces a logarithmically spaced continuous deviation method to dynamically calculate the relative position deviation. The formula is: Δ i,j =log(|i-j|+1) Where Δ i,j Represents the relative deviation between the i-th and j-th positions. For each pair of positions (i, j), the relative distance |ij| is calculated. The distance is transformed using a logarithmic function. In the attention calculation, the deviation Δ i,j Added to the attention score, the formula is:
6. A single-stage high-resolution human posture estimation method for solving occlusion according to claim 3, characterized in that: The step 2.3 is specifically as follows: Instance information abstraction extracts locations and features to represent each person, and global feature decoupling operates on the original feature maps extracted by the feature extraction network to generate instance-aware feature maps, each of which is used to estimate the heat map and key points of a person respectively.
7. A single-stage high-resolution human posture estimation method for solving occlusion according to claim 1, characterized in that: In step 3, the experimental results of the SHR model are compared with the classic top-down posture estimation method HRNet, the posture estimation method based on the attention mechanism, and the posture estimation method combining HRNet with the attention mechanism; when training the model, the number of training rounds, the learning rate size, and the number of model channels are adjusted according to the training results to obtain the model with the best effect; Verify the effectiveness of SHR under pre-training, including the performance of SHR loaded with pre-trained weights.