A human pose estimation method, device, medium and product
Patent Information
- Application Number
- CN202410599998.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-15
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2044-05-15
AI Technical Summary
因此,尽管CNNs和Transformer模型具有强大的学习和自校正能力,但当面对语义模糊的标记点时,它们可能无法准确预测与基本事实注释完全一致的输出
[0032]This invention inputs the image to be estimated into a deep learning network framework to obtain a low-resolution feature map, namely the first feature map. Periodic shuffling of the first feature map increases its resolution. Modifying the number of channels in the second feature map using depthwise separable convolution ensures consistency with the channels of the input image to be estimated and the input dimensions required by subsequent network layers, while reducing computational costs and significantly decreasing the number of parameters and computational complexity. The third feature map is then input into a convolutional layer with a set number of channels and kernels to obtain a predicted heatmap of keypoints. The coordinates of the highest predicted point offset one-quarter from the second-highest predicted point in the predicted heatmap are mapped to the corresponding coordinates in the image to be estimated, resulting in the final keypoint coordinate output, thereby generating an accurate human pose estimation result.
Smart Images

Figure CN118447537B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human posture estimation, and in particular to a method, device, medium and product for human posture estimation. Background Technology
[0002] Human pose estimation aims to detect the presence of one or more people in an image and then estimate the key locations of their body joints, such as shoulders, elbows, wrists, and ankles (see M. Andriluka, L. Pishchulin, P. Gehler, and B. Schiele, “2d human pose estimation: New benchmark and state of the art analysis,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014, pp. 3686-3693.), S. Jin, L. Xu, J. Xu, C. Wang, W. Liu, C. Qian, W. Ouyang, and P. Luo, “Whole-body human pose estimation in the wild,” in European Conference on Computer Vision (ECCV), 2020, pp. 196-214.), and A. Toshevand C. Szegedy, “Deeppose: Human pose estimation via deep neural networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014, pp. 1653-1660. This task is complicated by many external factors, such as occlusion, different body shapes, complex joints, different clothing, and different lighting conditions (see C. Zheng, W. Wu, C. Chen, T. Yang, S. Zhu, J. Shen, N. Kehtarnavaz, and M. Shah, “Deep learning-based humanposeestimation: A survey,” ACM Computing Surveys, vol. 56, no. 1, pp. 1-37, 2023).Accurate estimation of human pose is crucial for a wide range of applications, including action recognition and pose analysis (see Y. Kong and Y. Fu, “Human action recognition and prediction: A survey,” International Journal of Computer Vision, vol. 130, no. 5, pp. 1366-1401, 2022.), autonomous robotic systems (see H. Araujo, M. Murousavi, and M. Varshosaz, “Testing, validation, and verification of robotic and autonomous systems: a systematic review,” ACM Transactions on Software Engineering and Methodology, vol. 32, no. 2, pp. 1-61, 2023.), and animation content creation (see N.S. Willett, H.V. Shin, Z. Jin, W. Li, and A. Finkelstein, “Pose2pose: Pose selection and transfer for 2d character animation,” in International Conference on Intelligent User Interfaces (ICIUI), 2020, pp. 88-99.).
[0003] Existing human pose estimation methods can be broadly classified into two categories: top-down methods and bottom-up methods.
[0004] For example, in the papers "L. Pishchulin, E. Insafutdinov, S. Tang, B. Andres, M. Andriluka, PV Gehler, and B. Schiele, 'Deepcut: Joint subset partition and labeling for multi person pose estimation,' in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 4929-4937" and "Z. Cao, G. Hidalgo, T. Simon, S.-E. Wei, and Y. Sheikh, 'Openpose: Realtime multi-person 2dpose estimation using part affinity fields,' IEEE Transactions on Pattern Analysis and Machine Intelligence, pp. 172-186, 2021," the bottom-up approach first detects all possible, instance-independent body parts in the image and then classifies them into each instance. However, because keypoint detection is performed independently, it is susceptible to noise, leading to inaccurate pose estimation.
[0005] The top-down approach uses off-the-shelf detectors to provide bounding boxes (see S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” Advances in Neural Information Processing Systems, vol. 28, 2015.), Z. Cai and N. Vasconcelos, “Cascade r-cnn: Delving into high-quality object detection,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 6154-6162.), and C. Lyu, W. Zhang, H. Huang, Y. Zhou, Y. Wang, Y. Liu, S. Zhang, and K. Chen, “Rtmdet: An empirical study of designing real-time object detectors,” arXiv, 2022.), and then clips them to a uniform scale for pose estimation. This method achieves superior accuracy and efficiency on both Convolutional Neural Networks (CNNs) and Transformer frameworks, and has established a dominant position in public benchmarks in the field of pose estimation. For example, see the literature “L. Xu, S. Jin, W. Liu, C. Qian, W. Ouyang, P. Luo, and X. Wang, “Zoomnas: searching for whole-body human pose estimation in the wild,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol.45, no.4, pp.5296-5313, 2022” and “A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al.”,“An image is worth 16x16words:Transformers for image recognition at scale,”arXiv,2020.》、《D.Zhang,J.Tang,and K.-T.Cheng,“Graph reasoning transformer for image parsing,”in ACMInternational Conference on Multimedia(ACM MM),2022,pp.2380-2389》、《J.Li andM.Wang,“Multi-person pose estimation with accurate heatmap regression andgreedy association,”IEEE Transactions on Circuits and Systems for VideoTechnology,vol.32,no.8,pp.5521-5535,2022.》、《L.Zhao,J.Xu,C.Gong,J.Yang,W.Zuo,and X.Gao,“Learning to acquire the quality of human pose estimation,”IEEETransactions on Circuits and Systems for Video Technology,vol.31,no.4,pp.1555-1568,2020.》、《C.Du,Z.Yan,H.Yu,L.Yu,and Z.Xiong,“Hierarchicalassociative encoding and decoding for bottom-up human pose estimation,”IEEETransactions on Circuits and Systems for Video Technology,vol.33,no.4,pp.1762-1775,2022.》、《A.Garcia-Garcia,S.Orts-Escolano,S.Oprea,V.Villena-Martinez,and J.Garcia-Rodriguez, "A review on deep learning techniques applied to semantic segmentation," arXiv, 2017., D.Zhang, Y.Lin, H.Chen, Z.Tian, X.Yang, J.Tang, and K.T.Cheng, "Deep learning for medical image segmentation: tricks, challenges and future directions," arXiv, 2022., D.Zhang, C.Zuo, Q.Wu, L.Fu, and X.Xiang, "Unabridged adjacent modulation for clothing parsing," Pattern Recognition, vol.127, p.108594, 2022., K.Sun, B.Xiao, D.Liu, and J.Wang, "Deep high-resolution representation learning for human pose estimation," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp.5693-5703. and D.Zhang, L.Zhang, and J.Tang, "Augmented fcn: rethinking context modeling for semantic segmentation," Science China Information Sciences, vol.66, no.4, p.142105, 2023..
[0006] However, existing top-down human pose estimation methods based on CNNs and Transformer frameworks require upsampling of feature maps, which can pose challenges in complex application scenarios. This is because traditional upsampling operations (e.g., bilinear interpolation (R. Szeliski, Computer vision: algorithms and applications. Springer Nature, 2022) and transposed convolution (H. Gao, H. Yuan, Z. Wang, and S. Ji, “Pixel transposed convolutional networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 42, no. 5, pp. 1218-1227, 2019)) lack the ability to adapt to the specific feature requirements of the task being performed and may lead to the loss of fine-grained semantic details (see A. Toshev and C. Szegedy, “Deeppose: Human pose estimation via deep neural networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern) Recognition(CVPR), 2014, pp.1653-1660.", "Z.Zhou, H.Li, H.Liu, N.Wang, G.Yu, and R.Ji, "Star loss: Reducing semantic ambiguity in facial landmark detection," in Proceedings of the IEEE Conference on Computer Vision andPattern Recognition(CVPR), 2023, pp.15475-15484." and "D.Dumitrescu and C.-A.Boiangiu, "A study of image upsampling and downsampling filters," Computers, vol.8, no.2, p.30, 2019.").This limitation has spurred the development of many image upsampling methods, such as nonlinear image upsampling (see YJLee and J.Yoon, “Nonlinear image upsampling method based on radial basis function interpolation,” IEEE Transactions on Image Processing, vol.19, no.10, pp.2682-2692, 2010.) and attention-based image upsampling (see S.Kundu, H. Mostafa, SNSridhar, and S. Sundaresan, “Attention-based image upsampling,” arXiv, 2020.). These approaches have achieved good results in preserving fine-grained details and improving the accuracy and robustness of super-resolution tasks, as seen in DCLepcha, B.Goyal, A.Dogra, and V.Goyal, “Image super-resolution: A comprehensive review, recent trends, challenges and applications,” Information Fusion, vol.91, pp.230-260, 2023.).
[0007] For example, the documents "D.Zhang, J.Tang, and K.-T.Cheng, "Graphreasoning transformer for image parsing," in ACM International Conference on Multimedia (ACM MM), 2022, pp.2380-2389.", "H.Gao, H.Yuan, Z.Wang, and S.Ji, "Pixel transposedconvolutional networks," IEEE Transactions on PatternAnalysis and MachineIntelligence, vol.42, no.5, pp.1218-1227, 2019.", "C.Hanouti and H.Le Borgne, "Learning semantic ambiguities for zero-shot learning," Multimedia Tools andApplications, pp.1-15, 2023.", "H.-J.Ye, Y.Shi, and D.-C.Zhan, “Identifying ambiguous similarity conditions via semantic matching,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Despite these achievements in general vision tasks, the complexity of individual body poses often leads to awkward semantic ambiguity in the annotation space (see M. Andriluka, L. Pishchulin, P. Gehler, and B.).Schiele, "2d human pose estimation: New benchmark and stateofthe art analysis," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014, pp.3686-3693. applications," Information Fusion, vol.91, pp.230-260, 2023." and "H.-J.Ye, Y.Shi, and D.-C.Zhan, "Identifying ambiguous similarityconditions via semantic matching," in Proceedings of the IEEE Conference on Computer Vision and Pattern This poses an implicit challenge to human pose estimation and leads to unstable and inaccurate predictions (see C. Zheng, W. Wu, C. Chen, T. Yang, S. Zhu, J. Shen, N. Kehtarnavaz, and M. Shah, “Deep learning-based human pose estimation: A survey,” ACM Computing Surveys, vol. 56, no. 1, pp. 1-37, 2023) and Z. Zhou, H. Li, H. Liu, N. Wang, G. Yu, and R. Ji, “Star loss: Reducing semantic ambiguity in facial landmark detection,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 15475-15484.For example, the surface proximity and occlusion of body parts (especially the "shoulder," "wrist," and "elbow") (the "left profile of the face") make accurately determining the spatial location of body joints challenging. Even if the same annotator annotates these parts at different times, the labeled keypoints may still shift (see references: S. Jin, L. Xu, J. Xu, C. Wang, W. Liu, C. Qian, W. Ouyang, and P. Luo, “Whole-body human pose estimation in the wild,” in European Conference on Computer Vision (ECCV), 2020, pp. 196-214; Z. Zhou, H. Li, H. Liu, N. Wang, G. Yu, and R. Ji, “Star loss: Reducing semantic ambiguity in facial landmark detection,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 15475-15484; and DCLepcha, B. Goyal, A. Dogra, and V. Goyal, “Image super-resolution: A comprehensive review, recent trends, challenges, and applications,” Information). Therefore, despite the powerful learning and self-correcting capabilities of CNNs and Transformer models, they may fail to accurately predict outputs that are perfectly consistent with the basic fact annotations when faced with semantically ambiguous labeled points. Inflexible upsampling operations exacerbate the inherent semantic ambiguity problem, leading to frustratingly low accuracy (see M. Andriluka, L. Pishchulin, P. Gehler, and B. Schiele, “2d human pose estimation: New benchmark and state of the art analysis,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014, pp. 3686-3693)." and "A. Toshevand C. Szegedy, "Deeppose: Human pose estimation via deep neural networks," in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014, pp.1653-1660. "). . Summary of the Invention
[0008] To address the aforementioned problems in the existing technology, this invention provides a human posture estimation method, device, medium, and product.
[0009] To achieve the above objectives, the present invention provides the following solution:
[0010] A human pose estimation method includes:
[0011] The image to be estimated is input into the deep learning network framework to obtain the first feature map;
[0012] Perform a periodic shuffling operation on the first feature map to obtain the second feature map;
[0013] The number of channels in the second feature map is modified by using depthwise separable convolution to obtain the third feature map;
[0014] The third feature map is input into a convolutional layer with a set number of channels and a set kernel to obtain a predicted heat map of the key points;
[0015] Determine the position coordinates of the key point corresponding to the highest predicted value in the predicted heat map, offset by one-quarter from the key point corresponding to the second highest predicted value, and map the position coordinates onto the image to be estimated. Output the mapped coordinates as the coordinates of the key point.
[0016] The human pose estimation result is obtained based on the coordinate output of the key points.
[0017] Optionally, the deep learning network framework includes at least one of the following: CNNs, Transformers, and ResNet.
[0018] Optionally, the second feature map is represented as:
[0019] F SR =PS(F LR );
[0020] In the formula, F SR Represents the second feature map, PS(F) LR) represents the first feature map F LR Perform a periodic mixed washing operation.
[0021] Optionally, for the first feature map F LR Perform periodic mixed washing operation PS(F) LR ) is represented as:
[0022] c·r 2 +r·mod(x,r)+mod(y,r);
[0023] In the formula, (x,y,c) are the output pixel coordinates in the high-resolution space, x∈(0,H-1), y∈(0,W-1), c∈(0,C-1), r represents the upsampling ratio, mod(·,·) is a modulo function, and H×W×C represents the number of sampling channels.
[0024] Optionally, the third feature map is represented as:
[0025] H = DSConv(F SR );
[0026] In the formula, H represents the third feature map, DSConv(F) SR ) indicates that depthwise separable convolution is used on the second feature map F SR Modify the number of channels.
[0027] Optionally, the number of channels in the convolutional layer is set to 133, and the size of the convolutional kernel is set to 1.
[0028] A computer device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the computer program to implement the human pose estimation method described in any of the preceding claims.
[0029] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the human pose estimation method described in any of the preceding claims.
[0030] A computer program product includes a computer program that, when executed by a processor, implements the human pose estimation method described in any of the preceding claims.
[0031] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0032] This invention inputs the image to be estimated into a deep learning network framework to obtain a low-resolution feature map, namely the first feature map. Periodic shuffling of the first feature map increases its resolution. Modifying the number of channels in the second feature map using depthwise separable convolution ensures consistency with the channels of the input image to be estimated and the input dimensions required by subsequent network layers, while reducing computational costs and significantly decreasing the number of parameters and computational complexity. The third feature map is then input into a convolutional layer with a set number of channels and kernels to obtain a predicted heatmap of keypoints. The coordinates of the highest predicted point offset one-quarter from the second-highest predicted point in the predicted heatmap are mapped to the corresponding coordinates in the image to be estimated, resulting in the final keypoint coordinate output, thereby generating an accurate human pose estimation result. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1 A flowchart of the human pose estimation method provided in an embodiment of the present invention;
[0035] Figure 2 This diagram illustrates the implementation process of the human pose estimation method provided in this embodiment of the invention on a ResNet network structure.
[0036] Figure 3 This is a comparison chart of the results obtained by the human pose estimation method and existing sampling techniques in different network structures according to the embodiments of the present invention;
[0037] Figure 4 This is an array diagram of human pose estimation results obtained using the human pose estimation method in an embodiment of the present invention. Detailed Implementation
[0038] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0039] The purpose of this invention is to provide a human pose estimation method, device, medium, and product, which aims to overcome the limitations of existing human pose estimation image upsampling methods and generate accurate human pose estimation results.
[0040] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0041] Example 1
[0042] Inappropriate image upsampling techniques can exacerbate semantic ambiguity, particularly in human pose estimation tasks, amplifying the impact of semantic ambiguity in the annotation space on experimental results and severely affecting accuracy. To overcome the limitations of existing image upsampling methods, this invention proposes a simple yet effective two-step SIU strategy, which includes a trainable and efficient upsampling process, allowing the entire network to more flexibly adapt to different feature requirements for pose estimation. Therefore, this SIU strategy improves spatial information preservation; SIU can adapt to the specific feature requirements of human pose estimation and effectively preserve spatial information, which is crucial for the accuracy of pose estimation. Compared to previous image upsampling methods, the human pose estimation method proposed in this invention not only effectively reduces computational costs but also employs a strategy of increasing resolution before adjusting the channel size of features, ensuring alignment of output features in the channel dimension. This effectively alleviates the semantic ambiguity problem in human pose estimation tasks.
[0043] The SIU strategy proposed in this invention is a general method that can be applied to all deep learning frameworks, such as representative CNN-based and Transformer-based human pose estimation frameworks.
[0044] Based on the above description, such as Figure 1 As shown, this embodiment provides a human pose estimation method, including:
[0045] Step 100: Input the image to be estimated into the deep learning network framework to obtain the first feature map.
[0046] Step 101: Perform a periodic shuffling operation on the first feature map to obtain the second feature map.
[0047] The second feature map is represented as follows:
[0048] F SR =PS(F LR (1)
[0049] In the formula, F SR This represents the second feature map, whose size is... r represents the upsampling ratio; for example, the upsampling ratio is set to 2, 4, and 8. PS(F LR ) represents the first feature map F LR Perform a periodic mixed washing operation. c·r 2 +r·mod(x,r)+mod(y,r). (x,y,c) are the output pixel coordinates in the high-resolution space, x∈(0,H-1), y∈(0,W-1), c∈(0,C-1), mod(·,·) is a modulo function, and H×W×C represents the number of sampling channels.
[0050] Step 102: Modify the number of channels in the second feature map using depthwise separable convolution to obtain the third feature map. The third feature map is represented as follows:
[0051] H = DSConv(F SR (2)
[0052] In the formula, H represents the third feature map. C out DSConv(F) represents the number of output channels that match the number of channels in the input feature map (i.e., the image to be estimated) and the input requirements of subsequent network layers. SR ) represents depthwise separable convolution.
[0053] The implementation process of depthwise separable convolution is as follows: given the convolution kernel tensor k is the kernel size set to 3. During depthwise convolution, for each input channel i∈[1, C / r... 2 Perform convolution operations of size k×k on each side to obtain C / r. 2 Output feature map X i ∈R H×W Pointwise convolution performs pointwise convolution, i.e., 1×1 convolution, on the depthwise convolution output feature map along the channel dimension, to obtain the final output feature map, i.e., the third feature map H. Depthwise separable convolution not only helps to achieve excellent aggregation, but also significantly reduces the number of parameters involved and computational complexity.
[0054] Step 103: Input the third feature map into a convolutional layer with a set number of channels and a set kernel to obtain the predicted heatmap of the key points.
[0055] Step 104: Determine the position coordinates of the key point corresponding to the highest predicted value in the predicted heat map, offset by one-quarter from the key point corresponding to the second highest predicted value, and map the position coordinates onto the image to be estimated. Output the mapped coordinates as the coordinates of the key point.
[0056] Step 105: Obtain the human pose estimation result based on the coordinate output of the key points.
[0057] The commonly used ResNet network structure is used as the backbone network framework to explain the implementation process of the human pose estimation method provided in this embodiment.
[0058] The flowchart of the human pose estimation method based on ResNet network structure is as follows: Figure 2 As shown. First, the input image (i.e., the image to be estimated) passes through a ResNet backbone network to reduce its resolution and learn to extract features, resulting in a low-resolution feature map (i.e., the first feature map F). LR The ResNet backbone network consists of a series of basic residual networks. Next, the proposed SIU strategy is used for upsampling to upsample the low-resolution feature map to a high-resolution feature map (i.e., the second feature map H). Specifically, the SIU strategy execution process mainly consists of two steps:
[0059] The first step is to increase the resolution of the model feature map by using periodic shuffling. The calculation method can be found in the above formula (1).
[0060] The second step involves using convolutional layers to modify the channel size of the features. In this embodiment, the SIU convolutional layer portion is preferably a depthwise separable convolution to modify the number of channels in the feature map. The reason for using depthwise separable convolution instead of traditional convolution is to ensure a consistent match between the channels of the input feature map and the input dimensions required by subsequent network layers, while reducing computational costs. Finally, the third feature map H is obtained by modifying the number of channels in the feature map using depthwise separable convolution.
[0061] The high-resolution feature map (i.e. the third feature map H) output by the SIU strategy is passed through a convolutional layer with 133 channels and a kernel size of 1 to obtain a heat map for key point prediction. The coordinates of the highest predicted value point in the heat map, which is offset by one-quarter from the second highest predicted value point, are mapped to the corresponding coordinates in the original image to obtain the final key point coordinate output, thereby generating accurate human pose estimation results.
[0062] Furthermore, using experimental datasets, the implementation process and advantages of the human pose estimation method provided by the present invention are illustrated through comparison. Specifically:
[0063] I. Dataset and Evaluation Metrics
[0064] The challenging and representative COCO WholeBody dataset was chosen as the experimental dataset. This dataset is the first and mainstream large-scale benchmark dataset for full-body pose estimation, an extension of the COCO 2017 dataset with the same training and validation splits. Each instance in this dataset is annotated with 133 keypoints, including 17 body keypoints, 68 facial keypoints, 42 hand keypoints, and 6 foot keypoints. The total training set contains 250,000 instances. All models in the experiments were trained and validated on the COCO WholeBody V1.0 dataset.
[0065] Evaluation Metrics. The primary accuracy metrics used in this invention are the commonly used mean precision (mAP) and mean recall (mAR). These metrics are derived by evaluating the similarity of object keypoints within a threshold range of 0.5 to 0.95. Furthermore, parameters and floating-point operations (FLOPs) are considered to evaluate the model's efficiency.
[0066] II. Detailed Implementation Information
[0067] (1) Selection of benchmark models: SimpleBaseline based on ResNet, HRNet based on HRNet, and EfficientViT-L-SAM based on the Transformer framework were selected as benchmark models. Unless otherwise stated, all settings of the benchmark models are consistent with their respective papers.
[0068] (2) Training and Testing Details: For the SIU strategy based on SimpleBaseline and HRNet, most of the default training and evaluation settings of the benchmark models can be followed. Adam was chosen as the optimizer, aligning with the selection of SimpleBaseline and HRNet. For the learning rate settings, the base learning rate was set to 1e-3, decreasing to 1e-4 and 1e-5 at 370 and 400 epochs, respectively. The total number of training epochs was 400. Experiments were run on four NVIDIA TITAN RTX GPUs. For EfficientViT-L-SAM, the initial learning rate was 5e-4. Furthermore, to ensure the model learns sufficiently from the dataset, the total number of training epochs was modified to 610, and the second descent point of the learning plan was modified to 600 epochs. To ensure all other parameters are consistent with the benchmark model, a two-stage top-down paradigm can also be followed, using person detection boxes with 56.4 AP on the COCO val 2017 dataset, adhering to MMPose. All experiments were conducted using the MMPose framework based on the PyTorch deep learning platform, and initialized using pre-trained baseline model weights. The test results are shown in Table 1.
[0069] Table 1 Test Results
[0070]
[0071]
[0072]
[0073] III. Ablation Research.
[0074] The ablation studies provided in this invention aim to answer the following questions:
[0075] Question 1: Is it possible for the method provided by this invention to improve the performance of various benchmark models?
[0076] Question 2: Is progressive upsampling necessary for implementing the method provided in this invention?
[0077] Question 3: Compared with other related methods, does the method provided by this invention exhibit better performance?
[0078] Question 4: How does the method provided by this invention perform under different frameworks?
[0079] The following answers were obtained after conducting ablation studies to address the above issues:
[0080] Answer 1: The effectiveness of the SIU strategy provided by this invention on different benchmark models.
[0081] This section primarily analyzes the impact of the SIU strategy on different baseline models with varying backbones. Experimental results are shown in Table 2. It can be observed that the SIU strategy achieves significant improvements on different baseline models. The speed of all methods was recorded on a 3090 GPU with a batch size of 32. Specifically, compared to SimpleBaseline using ResNet-50 as the backbone and an input image size of 256×192, the model adopted in this invention achieves 53.7% mAP and 64.7% mAR with 23.6M parameters. This significantly improves the baseline model by 1.6% mAP and 1.4% mAR, while reducing the number of model parameters by more than 30%. With an input image size of 256×192, the model based on SimpleBaseline with a ResNet-101 backbone achieves 55.4% mAP and 66.7% mAR, representing gains of 2.3% and 2.2% in mAP and mAR, respectively, compared to the baseline model results. Similar to the results for HRNet, at resolutions of 256×192 and 384×288, the method provided in this invention achieves 58.6% and 64.0% mAP, respectively, compared to HRNet with HRNet-W48 as its backbone, representing improvements of 0.7% and 0.8%, significantly reducing computational complexity and achieving faster inference speeds. Furthermore, it is noteworthy that the results show improvements across four different body parts: body, feet, face, and hands. This observation underscores the comprehensive effectiveness of the method provided in improving performance for individual body parts.
[0082] For EfficientViT-L-SAM, the most efficient backbone network, EfficientViT-L0, can be selected for comparison. The method provided in this invention slightly outperforms the benchmark model in performance. In other aspects, it exhibits significant advantages. The method provided in this invention achieves 61.8% mAP at 10.3 G FLOPs and 616.0 FPS. Compared to the benchmark model, mAP is increased by 0.6%, while complexity is reduced by more than 15%, and inference speed is improved. These results demonstrate that the SIU strategy provided in this invention outperforms other benchmarks in terms of model parameters, computational complexity, and inference speed. Furthermore, the method provided in this invention achieves performance improvements on benchmark models of different categories and resolutions, demonstrating strong generalization ability. Moreover, the method provided in this invention is a one-time upsampling method, which is more efficient and effective compared to the stepwise upsampling method in the CNN head of the benchmark model.
[0083] Table 2 Experimental Results
[0084] SimpleBaseline - 52.1 63.3 5.5G 34.0M 1212.6 SIUours t=1 53.7 64.7 4.2G 23.6M 1517.0 SIUours t=2 53.7 64.8 4.2G 23.6M 1511.7 SIUours t=3 53.6 64.7 4.2G 23.6M 1509.4
[0085] Answer 2: Effectiveness of SIU with a Progressive Upsampling Strategy. The baseline model based on SimpleBaseline implements upsampling through several deconvolutional layers. Therefore, in this section, the effectiveness of the SIU strategy is tested through progressive upsampling. Table 3 compares the performance of the proposed SIU strategy based on ResNet-50, considering single, double, and triple upsampling scenarios with an input size of 256×192. Experimental results show that the frequency of upsampling has little impact on the model's performance, efficiency, and inference speed. As shown in Table 2, single upsampling achieves 53.7% mAP, 64.7% mAR, and 1517.0 FPS. Compared to single upsampling, the model with double upsampling shows a slight decrease in mAR (0.1%) but an increase in inference speed (5.3 FPS), highlighting a slight but significant improvement. To meet different deployment needs and achieve a balance between inference speed and performance, single upsampling can be chosen in practice.
[0086] Table 3. Experimental Results for Different Performance Levels
[0087]
[0088] Answer 3: Advantages of the SIU strategy compared to other upsampling methods. In pose estimation tasks, the most commonly used upsampling techniques include nearest neighbor interpolation, deconvolution, and DUpsampling. The method provided in this invention is compared with these upsampling techniques on SimpleBaseline based on ResNet-50 and HRNet based on HRNet-W32, with an input size of 256×192. Table 2 and... Figure 3 This demonstrates that the upsampling method provided by this invention not only produces excellent results but also maintains high execution efficiency. When integrated with the ResNet-50 backbone network, the SIU strategy provided by this invention outperforms DUpsampling and deconvolution methods. Specifically, the SIU strategy provided by this invention achieves 53.7% mAP and 64.7% mAR with 4.2G FLOPs and 23.6M Params. In terms of mAP, the SIU strategy provided by this invention improves by 2.3% and 1.6% compared to the DUpsampling and deconvolution methods, respectively. Furthermore, it exhibits a significant reduction in computational complexity, reducing it by more than 70% and 20% respectively compared to the comparison methods, and requires fewer parameters, reducing them by 0.9 times and 0.3 times, respectively. Experiments using HRNet-W32 as the backbone network further confirm the effectiveness and efficiency of the method provided by this invention, as it achieves optimal performance in terms of the required FLOPs and number of parameters.
[0089] Answer 4: The effectiveness of the SIU strategy on different computer vision frameworks. Table 1 shows the performance of the proposed SIU method on different frameworks, including CNN and Transformer. Experimental results show that CNN-based networks perform better with SIU compared to Transformer-based networks. Based on this, it can be observed that the performance and effectiveness of CNN-based networks are significantly improved when using the SIU strategy provided by this invention. However, applying the method provided by this invention to Transformer-based networks yields a significant improvement in efficiency, but only a moderate improvement in performance. This observation is attributed to the stronger alignment and consistency of the method provided by this invention for CNN-based networks compared to Transformer-based networks.
[0090] IV. Complexity Analysis.
[0091] For an input size of 384×288, the SIU policy model complexity based on ResNet-101 is 17.6 GFLOPs, based on HRNet-W48 is 33.9 GFLOPs, and based on EfficientViT-L0 is 10.3 GFLOPs. In comparison, OpenPose has a model complexity of 451.1 GFLOPs, SN has 272.3 GFLOPs, and ViTPose+-H has a model complexity of 122.9 GFLOPs at an input size of 256×192. Average inference speed time for excluding the detector part on COCO-WholeBody is reported on a 3090 GPU. The average processing time of the ViTPose+-H model is approximately 40.5 frames per second. In comparison, our ResNet-101-based SIU has an average processing time of 428.6 FPS at an input size of 384×288.
[0092] V. Comparison with the results of the latest technologies.
[0093] The method provided by this invention is quantitatively compared with other existing methods, and the results are shown in Table 4. Furthermore, the computational complexity (i.e., FLOPs) of each model can be evaluated and compared. The data in Table 4 show that the method provided by this invention significantly outperforms existing whole-body pose estimation techniques, such as SN and OpenPose, in terms of accuracy and efficiency. Moreover, the method provided by this invention also performs well in multi-person pose estimation, outperforming other benchmark methods. Specifically, compared with bottom-up methods such as PAF* and AE, the method provided by this invention shows a significant performance advantage. The SIU strategy provided by this invention is based on the SimpleBaseline, HRNet, and EfficientViT-L-SAM benchmark models. Without sacrificing performance, the method provided by this invention achieves a significant efficiency improvement compared to previous state-of-the-art top-down pose estimation models. In particular, the two implementations based on SimpleBaseline and HRNet with different input resolutions both outperform their respective benchmark methods with lower computational complexity. For example, the HRNet-W48-based model using SIU achieves 64.0% mAP and 73.2% mAR with an input size of 384×288. Meanwhile, the ResNet-101-based SIU model provided by this invention achieves 61.2% mAP and 70.9% mAR with an input size of 384×288, requiring only 17.6M FLOPs. This represents a complexity reduction of nearly 85% while maintaining mAP comparable to ViTPose+-H. The method provided by this invention can adapt to the specific needs of pose estimation, alleviate semantic ambiguity, and achieve an optimal balance between effectiveness and performance compared to other state-of-the-art techniques. This balance is particularly suitable for mobile phone chips and edge devices.
[0094] For an input size of 384×288, the SIU model complexity based on ResNet-101 is 17.6 G FLOPs, based on HRNet-W48 is 33.9 G FLOPs, and based on EfficientViT-L0 is 10.3 G FLOPs. In comparison, OpenPose has a model complexity of 451.1 G FLOPs, SN has 272.3 G FLOPs, and ViTPose+-H has a model complexity of 122.9 G FLOPs at an input size of 256×192. Average inference speed time excluding the detector portion on COCO-WholeBody is reported on a 3090 GPU. The average processing time of the ViTPose+-H model is approximately 40.5 frames per second. In comparison, the SIU based on ResNet-101 has an average processing time of 428.6 FPS at an input size of 384×288.
[0095] Table 4 Comparison of Results
[0096]
[0097]
[0098]
[0099]
[0100] VI. Visualization Results.
[0101] Figure 4 This paper demonstrates the qualitative effectiveness of the method provided by this invention in enhancing model handling of fuzzy annotations, improving specific functional requirements of pose estimation, and refining necessary spatial information and relationships to achieve accurate coordinate localization. It can be observed that the SIU provided by this invention shows significant advantages in overcoming challenges such as proximity of body parts, occlusion of body parts, and localization of small-scale keypoints (e.g., "hands," "face," and "feet"), which are among the most common and fundamental challenges in pose estimation tasks. In particular, as... Figure 4 As shown in the first row and first column, even when the left foot is obscured by the right foot, the method provided by this invention can accurately locate the joints of the left foot. The method also helps the model pay more attention to the spatial and semantic relationships between annotations, thereby more accurately locating facial contour keypoints prone to semantic ambiguity. These results confirm that appropriate upsampling methods capable of learning task-specific features are crucial for pose estimation. Furthermore, as... Figure 4 As shown, the method provided by this invention demonstrates improved performance in detecting hand / face keypoints with small scale and different poses. This indicates that the method provided by this invention has advantages in learning features of small objects.
[0102] In summary, based on the above quantitative and qualitative results, the method provided by this invention is accurate and efficient, has strong transferability, and improves the performance of different benchmark methods.
[0103] Example 2
[0104] A computer device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the human pose estimation method of Embodiment 1.
[0105] Example 3
[0106] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the human pose estimation method of Embodiment 1.
[0107] Example 4
[0108] A computer program product includes a computer program that, when executed by a processor, implements the human pose estimation method of Embodiment 1.
[0109] Example 5
[0110] A computer device, which may be a database, includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The database stores pending transactions. The I / O interfaces facilitate information exchange between the processor and external devices. The communication interface enables communication with external terminals via a network connection. When executed by the processor, the computer program implements the human pose estimation method described in Embodiment 1.
[0111] It should be noted that the object information (including but not limited to object device information, object personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this invention are all information and data authorized by the object or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0112] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided by this invention may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided by this invention may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0113] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0114] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Similar or identical parts between the various embodiments can be referred to mutually. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for estimating human pose, characterized in that, The method includes: The image to be estimated is input into the deep learning network framework to obtain the first feature map; Perform a periodic shuffling operation on the first feature map to obtain the second feature map; The number of channels in the second feature map is modified by using depthwise separable convolution to obtain the third feature map; The third feature map is input into a convolutional layer with a set number of channels and a set kernel to obtain a predicted heat map of the key points; Determine the position coordinates of the key point corresponding to the highest predicted value in the predicted heat map, offset by one-quarter from the key point corresponding to the second highest predicted value, and map the position coordinates onto the image to be estimated. Output the mapped coordinates as the coordinates of the key point. The human pose estimation result is obtained based on the coordinate output of the key points.
2. The human posture estimation method according to claim 1, characterized in that, The deep learning network framework includes at least one of the following: CNNs, Transformers, and ResNet.
3. The human posture estimation method according to claim 1, characterized in that, The second feature map is represented as follows: F SR =PS(F LR ); In the formula, F SR Represents the second feature map, PS(F) LR ) represents the first feature map F LR Perform a periodic mixed washing operation.
4. The human posture estimation method according to claim 3, characterized in that, For the first feature map F LR Perform periodic mixed washing operation PS(F) LR ) is represented as: In the formula, (x,y,c) are the output pixel coordinates in the high-resolution space, x∈(0,H-1), y∈(0,W-1), c∈(0,C-1), r represents the upsampling ratio, mod(·,·) is a modulo function, and H×W×C represents the number of sampling channels.
5. The human posture estimation method according to claim 1, characterized in that, The third feature map is represented as follows: H=DSConv(F SR ); In the formula, H represents the third feature map, DSConv(F) SR ) indicates that depthwise separable convolution is used on the second feature map F SR Modify the number of channels.
6. The human pose estimation method according to claim 1, characterized in that, The convolutional layer has 133 channels and a kernel size of 1.
7. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the human pose estimation method according to any one of claims 1-6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the human pose estimation method according to any one of claims 1-6.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the human pose estimation method according to any one of claims 1-6.