Occluded human pose estimation corrector and correction method based on key point interconnection

CN116403277BActive Publication Date: 2026-09-18CHANGCHUN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310269431.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-20
Publication Date
2026-09-18
Estimated Expiration
2043-03-20

AI Technical Summary

Technical Problem

[0009]本发明所要解决的技术问题是:提供一种基于关键点互联的被遮挡人体姿态估计矫正器及矫正方法用于解决现有技术中人体姿态估计网络中通过热力图进行估计关键点所在位置概率却并没有考虑姿态估计网络形成的热力图是否准确的技术问题

Benefits of technology

[0028] This invention discloses a pluggable, keypoint interconnection-based occluded human pose estimation corrector. This corrector is directly connected to the end of a traditional human pose estimation network to correct the heatmap estimated by the traditional network, resulting in a more accurate and realistic heatmap, thus further ensuring the accuracy of the output human pose estimation results. The corrector utilizes the interrelationships between different keypoints to establish relevant mathematical models, comprehensively considering the inherent connections between the neural network and keypoints, and rationally constructing an implicit keypoint interconnection network. It uses keypoints with better performance to infer areas of weaker performance, correcting potential erroneous predictions in the heatmaps generated by traditional pose neural networks. This invention is particularly effective for occluded human body parts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116403277B_ABST
    Figure CN116403277B_ABST
Patent Text Reader

Abstract

The key point interconnection-based occluded human posture estimation corrector and correction method belong to the technical field of human posture estimation computer vision recognition. The application discloses a plug-in type key point interconnection-based occluded human posture estimation corrector, which is directly connected at the end of a traditional human posture estimation network to correct a heat map estimated by the traditional network, so that a more accurate real heat map is obtained, and the accuracy of the output human posture estimation result is further ensured. In the corrector, a related mathematical model is established by using the mutual relationship between different key points, the internal relationship between the neural network and the key points is comprehensively considered, and a reasonable implicit key point interconnection network is built, the key points with good effects are used to infer the key point regions with poor effects, the incorrect prediction of the heat map generated in the traditional posture neural network is corrected, and the application has good effects on some occluded human body parts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision recognition technology for human pose estimation, and in particular relates to an occluded human pose estimation corrector and correction method based on key point interconnection. Background Technology

[0002] Human pose estimation is an important branch of modern image engineering. By accurately identifying and capturing key points of the human body, the results of key point recognition can be applied to other visual domains, providing technical support for virtual reality and human-computer interaction. However, traditional human pose estimation networks face enormous challenges. The flexible body structure and high degree of freedom of limbs can result in a wide variety of poses. Different colors on the body can cause visual interference, and complex external environments can lead to human occlusion and perspective differences.

[0003] With the continuous research into deep learning in computer science, applying convolutional neural networks (CNNs) to human pose estimation and recognition has become a promising solution. The neural network extracts features from the human body in an image, locates key points based on these features, and connects these key points through computation and connections within the network. After extensive training, simulating the human brain's recognition process, it automatically identifies the locations of these key points. Compared to traditional pose recognition methods, convolutional neural network-based pose recognition methods offer better accuracy and efficiency.

[0004] In the process of pose recognition in neural networks, a large number of convolutional kernels are used to extract features from images. During the downsampling feature extraction process, although a large receptive field is obtained, some information will inevitably be lost due to dilated convolution or max pooling operations. In the necessary feature extraction process, some very important feature information is lost. After multiple convolution stages, the accumulated loss of feature information is countless, which will cause the heat map of human body key points to be inaccurate and affect the subsequent key point recognition.

[0005] In the CPN (Cascaded Pyramid Network) and its variants, a pyramidal network structure similar to FPN is constructed using two cascaded modules (GlobalNet and RefineNet). GlobalNet is responsible for detecting all key points in the network, focusing on predicting eight key points that are not easily occluded. RefineNet corrects the prediction results of GlobalNet, addressing the issues of occluded or invisible joints or locations with large prediction errors for human key points in complex backgrounds, and employs an online hard mining strategy for key points.

[0006] HRNet is a parallel multi-scale fusion network structure that employs a multi-path parallel network architecture. Through multi-layer parallelism, it maintains a high resolution, and reducing downsampling effectively preserves the information retained at high resolution. Multi-scale fusion enables the fusion of contextual information. The two methods mentioned above only improve the network structure and do not address the interrelationships between key points.

[0007] Traditional network structures extract multidimensional features from images and further estimate keypoint heatmaps. They then use these heatmaps to estimate the probability of keypoint locations. However, they do not consider the accuracy of the heatmaps generated by the pose estimation network, nor do they consider further research on the generated heatmaps.

[0008] Therefore, there is an urgent need for a new technical solution to address this problem. Summary of the Invention

[0009] The technical problem to be solved by the present invention is to provide an occluded human pose estimation corrector and correction method based on key point interconnection to solve the technical problem in the prior art where the probability of key point location is estimated by heat map in human pose estimation network, but the accuracy of the heat map formed by pose estimation network is not considered.

[0010] An occluded human pose estimation corrector based on keypoint interconnection includes an input port, an output plug, and a central processing module. The input port is connected to the heatmap output plug of a traditional human pose estimation network; the output plug is connected to the heatmap input port of a traditional human pose estimation network. The central processing module includes a storage module, a human pose estimation correction processing module, and a corrected heatmap output module. The storage module stores the training set and the data obtained from the human pose estimation correction processing module. The corrected heatmap output module is used to output the corrected heatmap.

[0011] The method for occluded human pose estimation and correction based on keypoint interconnection, utilizing the aforementioned keypoint interconnection-based occluded human pose estimation and correction device, includes the following steps, which are performed sequentially:

[0012] Step 1: Select 9000 images containing human figures from the existing dataset as Dataset I. Randomly select 1000 human figures from Dataset I and randomly occlude key points to create new images. Store these new images in Dataset I to form Dataset II. The number of key points is 16. Dataset II contains a total of 10000 images, including normal and occluded human figures. Unify the image resolution of Dataset II to the same resolution. Generate Gaussian-distributed heatmaps based on the human figure recognition key points in the normal human figures, and denote them as the actual heatmaps. All actual heatmaps and dataset II are merged to form a training set;

[0013] Step 2: Store the training set into the traditional human pose estimation network. Connect the input plug of the keypoint interconnection-based occluded human pose estimation corrector to the heatmap output port of the traditional human pose estimation network. Connect the output plug of the keypoint interconnection-based occluded human pose estimation corrector to the heatmap input port of the traditional human pose estimation network. Train the corrector using the training set.

[0014] ① Use a traditional human pose estimation network to perform human pose estimation on the images in dataset II, obtain human pose estimation heatmaps and reorder them. The sorting rule is: sort the estimation accuracy of each identification key point from high to low to form a pseudo heatmap, denoted as P = H × W × K, where H is the height of the pseudo heatmap, W is the width of the pseudo heatmap, K is the length of the heatmap, and also the number of network key points, K = 16;

[0015] ② Divide the pseudo-heatmap P into two parts, represented as P = {P1, P2}, where the first part of the pseudo-heatmap P1 = H × W × K1, K1 = 12, and the second part of the pseudo-heatmap P2 = H × W × K2, K2 = 4, and reconstruct and arrange them.

[0016] The reconstructed permutation is a linear transformation mapping, which unfolds into a two-dimensional image, denoted as . Divide the region containing the two-dimensional image into and Two areas, of which, To guide key areas, The region is the key point area to be indexed, and K1 = 12, K2 = 4;

[0017] ③ For matrix multiplication calculations, Divide the region into three equal parts from top to bottom, denoted as G. Ⅰ G Ⅱ G ⅢThese three regions represent different correlations due to their varying accuracy levels, with G... Ⅰ =HW×K1 Ⅰ , K1 Ⅰ ={1≤K≤4}、K1 Ⅱ ={5≤K≤8}、

[0018] ④ Utilize the partitioned Calculate and obtain the key point area. With the key point region to be indexed The three correlation coefficient matrices represent the correlation between the guiding key point region and the key point region to be indexed.

[0019] ⑤ Because the accuracy of the three regions decreases sequentially, and the three correlation matrix coefficients have different effects on the total correlation coefficient matrix, different scaling factors are assigned to the three correlation matrix coefficients before they are summed to obtain the total correlation coefficient matrix C. M :

[0020]

[0021] ⑥ Use the relevant system matrix C M Treating the index key point region Perform matrix multiplication to obtain the index region F. The index region F is reconstructed and rearranged through a linear transformation mapping to obtain... The reconfiguration here is the reverse process of reconfiguring the permutation structure in step ②. The correlation coefficient matrix C represents the relationship between the correlation coefficient matrix and the correlation coefficient matrix C. M Activated pose feature image, K2 = 4;

[0022] ⑦ Take the pose feature image By performing a pixel-by-pixel addition operation with the pseudo-heatmap from the second part, the indexed keypoint regions can be obtained.

[0023] ⑧ Combine the pseudo-heatmap P1 from the first part with the indexed key point region By merging the data, a predicted heatmap is finally obtained.

[0024] ⑨ Utilize the predicted heatmap and the corresponding stored actual heatmaps in the training set to establish a loss function, and obtain the corresponding loss function value, where, The resulting predicted heatmap The actual heatmap representing the image, l is the evaluation coefficient, and K is the number of key points;

[0025] ⑩ Repeat steps ① to ⑨ to train the occluded human pose estimation corrector based on keypoint interconnection and the traditional human pose estimation network end-to-end using the training set. After reaching the specified number of training times, take the loss function with the smallest loss function value as the loss function of the corrector in the traditional human pose estimation network. The training of the occluded human pose estimation corrector based on keypoint interconnection and the traditional human pose estimation network is completed.

[0026] Step 3: Use the trained keypoint interconnected occluded human pose estimation corrector and the traditional human pose estimation network to perform human pose recognition in real time.

[0027] Through the above design scheme, the present invention can bring the following beneficial effects:

[0028] This invention discloses a pluggable, keypoint interconnection-based occluded human pose estimation corrector. This corrector is directly connected to the end of a traditional human pose estimation network to correct the heatmap estimated by the traditional network, resulting in a more accurate and realistic heatmap, thus further ensuring the accuracy of the output human pose estimation results. The corrector utilizes the interrelationships between different keypoints to establish relevant mathematical models, comprehensively considering the inherent connections between the neural network and keypoints, and rationally constructing an implicit keypoint interconnection network. It uses keypoints with better performance to infer areas of weaker performance, correcting potential erroneous predictions in the heatmaps generated by traditional pose neural networks. This invention is particularly effective for occluded human body parts. Attached Figure Description

[0029] The present invention will be further described below with reference to the accompanying drawings and specific embodiments:

[0030] Figure 1 This is a flowchart of the occluded human posture estimation corrector and correction method based on keypoint interconnection of the present invention.

[0031] Figure 2 This is a linear transformation mapping diagram in the occluded human posture estimation corrector and correction method based on key point interconnection of the present invention. Detailed Implementation

[0032] An occluded human pose estimation corrector based on keypoint interconnection includes an input port, an output plug, and a central processing module. The input port is connected to the heatmap output plug of a traditional human pose estimation network; the output plug is connected to the heatmap input port of a traditional human pose estimation network. The central processing module includes a storage module, a human pose estimation correction processing module, and a corrected heatmap output module. The storage module stores the training set and the data obtained from the human pose estimation correction processing module. The corrected heatmap output module is used to output the corrected heatmap.

[0033] The method for occluded human pose estimation and correction based on keypoint interconnection, utilizing the aforementioned keypoint interconnection-based occluded human pose estimation and correction device, includes the following steps, which are performed sequentially:

[0034] Step 1: Select 9000 images containing human figures from the existing dataset as Dataset I. Randomly select 1000 human figures from Dataset I and randomly occlude key points to create new images. Store these new images in Dataset I to form Dataset II. The number of key points is 16. Dataset II contains a total of 10000 images, including normal and occluded human figures. Unify the image resolution of Dataset II to the same resolution. Generate Gaussian-distributed heatmaps based on the human figure recognition key points in the normal human figures, and denote them as the actual heatmaps. All actual heatmaps and dataset II are merged to form a training set;

[0035] Step 2: Store the training set into the traditional human pose estimation network. Connect the input plug of the keypoint interconnection-based occluded human pose estimation corrector to the heatmap output port of the traditional human pose estimation network. Connect the output plug of the keypoint interconnection-based occluded human pose estimation corrector to the heatmap input port of the traditional human pose estimation network. Train the corrector using the training set.

[0036] ① Use a traditional human pose estimation network to perform human pose estimation on the images in dataset II, obtain human pose estimation heatmaps and reorder them. The sorting rule is: sort the estimation accuracy of each identification key point from high to low to form a pseudo heatmap, denoted as P = H × W × K, where H is the height of the pseudo heatmap, W is the width of the pseudo heatmap, K is the length of the heatmap, and also the number of network key points, K = 16;

[0037] ② Divide the pseudo-heatmap P into two parts, represented as P = {P1, P2}, where the first part of the pseudo-heatmap P1 = H × W × K1, K1 = 12, and the second part of the pseudo-heatmap P2 = H × W × K2, K2 = 4, and reconstruct and arrange them.

[0038] Reconstructing the permutation is essentially a linear transformation mapping, such as Figure 2 As shown, there are n elements a1 to a2. n up to n elements X1~X n The mapping process, the mapping parameter W in the mapping process 1n ~W nn The image, generated from network training, is unfolded into a two-dimensional image, denoted as... Divide the region containing the two-dimensional image into and Two areas, of which, To guide key areas, The region is the key point area to be indexed, and K1 = 12, K2 = 4;

[0039] ③ For matrix multiplication calculations, Divide the region into three equal parts from top to bottom, denoted as G. Ⅰ G Ⅱ G Ⅲ These three regions represent different correlations due to their varying accuracy levels, with G... Ⅰ =HW×K1 Ⅰ G Ⅱ =HW×K1 Ⅱ , K1 Ⅰ ={1≤K≤4}、

[0040] ④ Utilize the partitioned Calculate and obtain the key point area. With the key point region to be indexed The three correlation coefficient matrices represent the correlation between the guiding key point region and the key point region to be indexed.

[0041] ⑤ Because the accuracy of the three regions decreases sequentially, and the three correlation matrix coefficients have different effects on the total correlation coefficient matrix, different scaling factors are assigned to the three correlation matrix coefficients before they are summed to obtain the total correlation coefficient matrix C. M :

[0042]

[0043] ⑥ Use the relevant system matrix C M Treating the index key point region Perform matrix multiplication to obtain the index region F. The index region F is reconstructed and rearranged through a linear transformation mapping to obtain... The reconfiguration here is the reverse process of reconfiguring the permutation structure in step ②. The correlation coefficient matrix C represents the relationship between the correlation coefficient matrix and the correlation coefficient matrix C. M Activated pose feature image, K2 = 4;

[0044] ⑦ Take the pose feature image By performing a pixel-by-pixel addition operation with the pseudo-heatmap from the second part, the indexed keypoint regions can be obtained.

[0045] ⑧ Combine the pseudo-heatmap P1 from the first part with the indexed key point region By merging the data, a predicted heatmap is finally obtained.

[0046] ⑨ Utilize the predicted heatmap and the corresponding stored actual heatmaps in the training set to establish a loss function, and obtain the corresponding loss function value, where, The resulting predicted heatmap The actual heatmap representing the image, l is the evaluation coefficient, and K is the number of key points;

[0047] ⑩ Repeat steps ① to ⑨ to train the occluded human pose estimation corrector based on keypoint interconnection and the traditional human pose estimation network end-to-end using the training set. After reaching the specified number of training times, take the loss function with the smallest loss function value as the loss function of the corrector in the traditional human pose estimation network. The training of the occluded human pose estimation corrector based on keypoint interconnection and the traditional human pose estimation network is completed.

[0048] Step 3: Use the trained keypoint interconnected occluded human pose estimation corrector and the traditional human pose estimation network to perform human pose recognition in real time.

[0049] Example:

[0050] The unmodified traditional human pose estimation network and the traditional human pose estimation network with a corrector were trained on the same COCO2017 human pose detection training set, with identical training conditions. After training, the results were compared on the COCO2017 human pose detection test set. The comparison results are shown in the table below, where AP represents the average accuracy. 0.5 The average precision (AP) represents a threshold of 0.5. 0.75The average accuracy is represented by a threshold of 0.75. In the table below, HRNET stands for High-Resolution Net, SBL for Simple Baselines, and UDP for Unbiased Data Processing. These three networks are all common traditional networks for human pose estimation and have performed outstandingly in the field of human pose estimation.

[0051] The corrector was applied to these three networks, and the results were compared.

[0052]

[0053] The table shows that adding a designed human pose corrector to a traditional human pose estimation network helps the traditional network identify key points for human pose estimation. Compared with the traditional pose estimation network, the overall average accuracy is significantly improved. Therefore, the following results can be drawn: The pose corrector based on interconnected human key points designed in this invention can greatly improve the performance of traditional human pose estimation networks. The interdependence of key points corrects potential feature extraction errors in the pose estimation network. Under the same conditions, adding the designed corrector to the original traditional human pose estimation network improves the accuracy of human pose estimation, proving that the occluded human pose estimation corrector based on interconnected key points can strengthen the interrelationship between human key points and optimize the performance of the human pose estimation network.

Claims

1. A method for occluded human pose estimation and correction based on keypoint interconnection, utilizing an occluded human pose estimation corrector based on keypoint interconnection, the corrector including an input port, an output plug, and a central processing module; the input port is connected to the output plug of a traditional human pose estimation network's heatmap; the output plug is connected to the input port of a traditional human pose estimation network's heatmap; the central processing module includes a storage module, a human pose estimation and correction processing module, and a corrected heatmap output module; the storage module stores the training set and the data obtained from the human pose estimation and correction processing module; the corrected heatmap output module is used to output the corrected heatmap; its characteristic is: The steps include the following steps, and the following steps are performed in sequence: Step 1: Select 9000 images containing human figures from the existing dataset as Dataset I. Randomly select 1000 images from Dataset I and randomly occlude the key points for identification, creating new images. Store these new images in Dataset I to form Dataset II. The number of key points for identification is 16. Dataset II contains a total of 10000 images, including normal human images and occluded human images. Unify the image resolution of Dataset II to the same resolution. Generate Gaussian-distributed heatmaps based on the human identification key points in the normal human images, and denote them as the actual heatmaps. All actual heatmaps and dataset II are merged to form a training set; Step 2: Store the training set into the traditional human pose estimation network. Connect the input terminal of the keypoint interconnection-based occluded human pose estimation corrector to the heatmap output terminal of the traditional human pose estimation network. Connect the output terminal of the keypoint interconnection-based occluded human pose estimation corrector to the heatmap input terminal of the traditional human pose estimation network. Then train the corrector using the training set. ① Using a traditional human pose estimation network, human pose estimation is performed on the images in dataset II. A human pose estimation heatmap is obtained and reordered. The ordering rule is as follows: the estimation accuracy of each keypoint is ranked from highest to lowest, forming a pseudo-heatmap, denoted as . , It's a high value in a pseudo-heatmap. It is the width of the pseudo-heatmap. It is the length of the heatmap, and also the number of key points in the network. ; ②The pseudo-heatmap Divided into two parts, represented as The pseudo-heatmap in the first part , The pseudo-heatmap in the second part , Rearrange and restructure the arrangement; The reconstructed permutation is a linear transformation mapping, which unfolds into a two-dimensional image, denoted as . Divide the region where the two-dimensional image is located into and Two areas, of which, To guide key areas, The region is the key point area to be indexed, and , ; ③ For matrix multiplication calculations, The area is evenly divided into three regions from top to bottom, denoted as... , , These three regions represent different levels of correlation due to their varying accuracy. , , ,in , , ; ④ Utilize the partitioned Calculate and obtain the key point area. With the key point region to be indexed The three correlation coefficient matrices represent the correlation between the guiding key point region and the key point region to be indexed. , , ; ⑤ Because the accuracy of the three regions decreases sequentially, and the three correlation matrix coefficients have different effects on the total correlation coefficient matrix, different scaling factors are assigned to the three correlation matrix coefficients before they are summed to obtain the total correlation coefficient matrix. : ; ⑥ Use the total correlation coefficient matrix Treating the index key point region Perform matrix multiplication to obtain the index region. , index region The arrangement is reconstructed through a linear transformation mapping. The reconfiguration here is the reverse process of reconfiguring the permutation structure in step ②. Represented by the correlation coefficient matrix Activated pose feature image, , ; ⑦ Take the pose feature image By performing a pixel-by-pixel addition operation with the pseudo-heatmap from the second part, the indexed keypoint regions can be obtained. , ; ⑧ The pseudo-heatmap in the first part and indexed key point regions By merging the data, a predicted heatmap is finally obtained. , ; ⑨ Establish a loss function using the predicted heatmap and the corresponding actual heatmaps stored in the training set. The corresponding loss function value is obtained, where, , The resulting predicted heatmap The actual heatmap representing the image, It is an evaluation coefficient. The number of key points; ⑩ Repeat steps ① to ⑨ to train the occluded human pose estimation corrector based on keypoint interconnection and the traditional human pose estimation network end-to-end using the training set. After reaching the specified number of training times, take the loss function with the smallest loss function value. The training of the occluded human pose estimation corrector based on keypoint interconnection and the traditional human pose estimation network is completed. Step 3: Use the trained keypoint interconnected occluded human pose estimation corrector and the traditional human pose estimation network to perform human pose recognition in real time.