Differential joint point coordinate decoding method for human body posture estimation
The differentiable coordinate decoder method improves human pose estimation accuracy by smoothing and refining heatmaps to align with true joint locations, addressing target deviation issues in existing models through end-to-end optimization.
Patent Information
- Application Number
- CN202510462745.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-15
AI Technical Summary
In the existing heat map-based human pose estimation method, the difference between the prior distribution of prediction error and the real distribution leads to a decrease in prediction accuracy and it is difficult to effectively train the network.
Differentiable correlation node coordinate decoding method is used to generate single peak prediction heat maps through smooth multi-peak prediction heat maps, and the heat map of the correlation node area is cropped, and prediction is performed using layer normalization and multi-layer perceptron. Combining the threshold-based argmax function and discrete integral to calculate the coordinates of the integer maximum response point, the coordinate-based loss function is designed for end-to-end training.
Improve the accuracy of joint coordinate decoding, reduce quantization error and background noise interference, realize end-to-end optimization and dynamic adjustment of the model, enhance prediction accuracy, and seamless integration with existing models.
Smart Images

Figure CN120318862A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a differentiable joint point coordinate decoding method for human pose estimation, belonging to the technical field of computer vision. Background Art
[0002] The goal of human pose estimation is to determine the positions or spatial locations of the body key points (parts / joints) of a person from a given image or video, and this technology obtains the pose of a jointed human body based on the observation of the image. A jointed human body consists of joints and rigid parts.
[0003] As an important basic research in the field of computer vision, human pose estimation is the technical basis for tasks such as action recognition, human intention prediction, and intelligent monitoring. Current human pose estimation models mostly use methods based on joint point heatmaps with relatively high accuracy. This method uses a Gaussian distribution map centered on the joint position as the training target. This helps to achieve robust model fitting and also provides prior noise for point estimation. In addition, the heatmap-based method provides more dense supervision information than the method of directly regressing joint coordinates, thus improving the training efficiency.
[0004] For heatmap-based methods, the process of generating the true label heatmap usually involves using a manual prior distribution (such as a Gaussian distribution) to model the prediction error. However, this strategy introduces an object bias, that is, the difference between the prior true label distribution and the true distribution of the prediction error. Existing research has proved that the prediction noise of human pose estimation (HPE) tends to present an uneven and multimodal distribution, rather than a typical Gaussian or Laplace distribution. Therefore, the predicted heatmap deviates from the prior distribution in shape, and its maximum point does not align with the true joint position due to the prediction error. These differences (defined as object bias) lead to a decrease in prediction accuracy for the following reasons: 1) When decoding the predicted heatmap based on the prior distribution assumption, incorrect joint coordinates may be obtained; 2) The predicted heatmap may not be well aligned with the true heatmap, thus hindering effective network training. Summary of the Invention
[0005] Aiming at the problem that the difference between the prior distribution of the prediction error and the true distribution of the prediction error leads to an object bias in the heatmap of human pose estimation, the present invention provides a differentiable joint point coordinate decoding method for human pose estimation.
[0006] A differentiable joint point coordinate decoding method for human pose estimation according to the present invention includes:
[0007] Obtaining a multimodal predicted heatmap by passing a human pose image through a backbone network and a network head; smoothing the multimodal predicted heatmap to obtain a unimodal predicted heatmap, and determining the integer maximum response point coordinates of the unimodal predicted heatmap based on the peak point index;
[0008] Generate a grid on the unimodal prediction heatmap to obtain a unimodal grid map, and then, centered on the integer maximum response point coordinates, crop the joint region heatmap from the unimodal grid map based on a set length;
[0009] Perform prediction on the joint region heatmap using layer normalization and a multi-layer perceptron to obtain the offset between the integer maximum response point coordinates and the true joint point coordinates; then calculate the predicted coordinates by combining the offset with the integer maximum response point coordinates.
[0010] According to the differentiable joint point coordinate decoding method for human pose estimation of the present invention, the multi-modal prediction heatmap is a heatmap frame H containing one joint point u , and regard the heatmap frame H u as a two-dimensional discrete distribution array, expressed as the distribution P(u) of the predicted coordinates u in the two-dimensional coordinate system.
[0011] According to the differentiable joint point coordinate decoding method for human pose estimation of the present invention, use a learnable low-pass filter kernel k s to smooth P(u) through a convolution operation to obtain a unimodal prediction heatmap P(u)*k s .
[0012] According to the differentiable joint point coordinate decoding method for human pose estimation of the present invention, perform a threshold-based argmax function calculation on the unimodal prediction heatmap P(u)*k s to obtain the integer maximum response point coordinates u im :
[0013] First, convert the unimodal prediction heatmap P(u)*k s into a Dirac-like distribution δ l (u im ):
[0014]
[0015] where ∈ is a threshold parameter, is the set of real numbers, Sum(·) represents the sum of all elements, represents the indicator function;
[0016] Then, use discrete integration to calculate the integer maximum response point coordinates u im :
[0017]
[0018] where H is the height of the unimodal prediction heatmap P(u)*k s , and W is the width of the unimodal prediction heatmap P(u)*k s .
[0019] According to the differentiable joint point coordinate decoding method for human posture estimation of the present invention, the joint point area heat map φ is obtained by clipping the unimodal grid map L (P(u)*k s ,u im ), the size is (2s+1)×(2s+1), where s is the set length;
[0020] The joint point area heat map φ L (P(u)*k s ,u im ) of the sample grid g im It is expressed as:
[0021]
[0022] According to the differentiable joint point coordinate decoding method for human posture estimation of the present invention, U is used to represent the unimodal prediction heat map, and V is used to represent the cropped joint point area heat map:
[0023] for and
[0024]
[0025] Where i represents the horizontal coordinate of the joint point area heat map, j represents the vertical coordinate of the joint point area heat map, represents the domain of integers, Indicates g im The first number of the two-dimensional coordinates of the i-th row and j-th column, Indicates g im The second number of the two-dimensional coordinate of the i-th row and j-th column.
[0026] According to the differentiable joint point coordinate decoding method for human body posture estimation of the present invention, the predicted coordinate u is:
[0027]
[0028] Where u of is the offset between the integer maximum response point coordinate and the true joint point coordinate, MLP represents multi-layer perceptron, LN represents layer normalization, argmax thres (·) represents the coordinate decoding function.
[0029] According to the differentiable joint point coordinate decoding method for human posture estimation of the present invention, the loss function corresponding to the layer normalization and the multi-layer perceptron is L c :
[0030]
[0031] Where gk is the true joint coordinate of the k-th joint point, u k is the predicted coordinate of the k-th joint point, K is the total number of joint points, and t is the offset threshold.
[0032] According to the differentiable joint point coordinate decoding method for human pose estimation of the present invention, the loss function of the network structure from the human pose image to obtaining the heatmap corresponding to the joint point area is expressed as
[0033]
[0034] In the formula is the true heatmap of the k-th joint point, is the multi-modal predicted heatmap of the k-th joint point.
[0035] According to the differentiable joint point coordinate decoding method for human pose estimation of the present invention, the overall objective loss function L is:
[0036]
[0037] End-to-end training is adopted for the network structure.
[0038] Beneficial effects of the present invention: The method of the present invention can alleviate the target deviation caused by the difference between the prior distribution (such as Gaussian distribution) and the true distribution of the prediction error. It has the following advantages:
[0039] 1. Improve decoding accuracy: By smoothing the heatmap and local sampling, the joint coordinates can be decoded more accurately, reducing the interference of quantization error and background noise.
[0040] 2. End-to-end optimization: The entire decoding process is differentiable and can be optimized through supervised learning, making full use of the supervision information of the true coordinates to improve the training effect of the model.
[0041] 3. Dynamically adjust the prediction distribution: The coordinate-based loss function enables the network to dynamically adjust the prediction distribution of the heatmap according to the true coordinates, thereby realizing the transformation from the prior distribution to the result-driven distribution and improving the prediction accuracy.
[0042] 4. Strong compatibility: The decoder adopted by the method of the present invention can be seamlessly integrated with the existing heatmap-based human pose estimation model without large-scale modification of the model structure, having good compatibility and scalability. Description of the Drawings
[0043] Figure 1It is a schematic diagram of the decoding process of the Differentiable Coordinate Decoder (DCD) on which the differentiable joint point coordinate decoding method for human pose estimation according to the present invention is based; in the figure, δ(u of )δ(u im )δ(u) represent the Dirac distributions of u of , u im and u respectively, and G(g) is a Gaussian distribution centered on g. Detailed implementation manners
[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0046] The present invention will be further described below with reference to the accompanying drawings, but it is not a limitation of the present invention.
[0047] Combined with Figure 1 shown, the present invention provides a differentiable joint point coordinate decoding method for human pose estimation, including,
[0048] Obtaining a multi-peak prediction heat map by passing a human pose image through a backbone network and a network head; performing smoothing processing on the multi-peak prediction heat map to obtain a single-peak prediction heat map, and determining the integer maximum response point coordinates of the single-peak prediction heat map based on the peak point index;
[0049] Generating a grid on the single-peak prediction heat map to obtain a single-peak grid map, and then, centered on the integer maximum response point coordinates, cropping the single-peak grid map based on a set length to obtain a joint point region heat map;
[0050] Performing prediction on the joint point region heat map by using layer normalization and a multi-layer perceptron to obtain the offset between the integer maximum response point coordinates and the true joint point coordinates; and then calculating the predicted coordinates by combining the offset with the integer maximum response point coordinates.
[0051] Finally, a coordinate-based loss function can be used to train the model to optimize the prediction distribution and decoding accuracy of the heat map.
[0052] The differentiable joint point coordinate decoding method described in this embodiment is implemented based on a Differentiable Coordinate Decoder (DCD) and consists of the corresponding network structure in the method.
[0053] Combined with Figure 1 as shown in (a) of , the feature map output by the backbone network is first fed into a classical head structure for predicting the heatmap H u in the space, where B is the batch size, K is the number of channels (i.e., the number of human joints), and H and W are the height and width of the heatmap respectively. Subsequently, H u is reshaped into space so that each sample contains only one heatmap frame H u , which facilitates the DCD to process the heatmap.
[0054] In this embodiment, the multi-peak prediction heatmap is a heatmap frame H containing one joint point u . The heatmap frame H u is regarded as a two-dimensional discrete distribution array, expressed as the distribution P(u) of the predicted coordinate u in the two-dimensional coordinate system. Usually, P(u) will present multiple peaks around the maximum activation value, which may have a negative impact on the performance of the DCD.
[0055] To solve this problem, a learnable low-pass filter kernel k s is used to smooth P(u) through a convolution operation to obtain a single-peak prediction heatmap P(u)*k s .
[0056] To obtain the final joint position u from P(u), a local region is cropped, which is a neighborhood centered on the integer maximum activation coordinate u u of H im for precise coordinate regression.
[0057] Perform a threshold-based argmax function calculation on the single-peak prediction heatmap P(u)*k s to obtain the integer maximum response point coordinate u im , as shown in (a) of Figure 1 :
[0058] In the process of the threshold-based argmax function, first convert the single-peak prediction heatmap P(u)*k s into a Dirac-like distribution δ l (u im ):
[0059]
[0060] where ∈ is the threshold parameter. should be as large as possible, is the set of real numbers, Sum(·) represents the summation of all elements, represents the indicator function; the unimodal prediction heatmap P(u)*k s is normalized by its maximum value, and the maximum activation element is selected by ∈ and set to 1 by the indicator function is set to 1.
[0061] Then, the integer maximum response point coordinate u is calculated using discrete integration im :
[0062]
[0063] where H is the height of the unimodal prediction heatmap P(u)*k s and W is the width of the unimodal prediction heatmap P(u)*k s respectively.
[0064] (w,h) are the coordinates in the heatmap P(u)*k s . After obtaining u im , all integer coordinates centered at u im within the (2s + 1)×(2s + 1) neighborhood are used to generate a sampling grid. Then, the generated grid is used to crop a (2s + 1)×(2s + 1) local region from the heatmap P(u)*k of size H×W through differentiable grid sampling s . The processes of grid generation and grid sampling are denoted as φ L (·), as shown in Figure 1 (b).
[0065] Furthermore, the joint region heatmap φ L (P(u)*k s ,u im ) cropped from the unimodal grid map has a size of (2s + 1)×(2s + 1), where s is the set length;
[0066] Mathematically, the sample grid g L of the joint region heatmap φ s (P(u)*k im ,u im is expressed as:
[0067]
[0068] Let U represent the unimodal prediction heatmap and V represent the cropped joint region heatmap:
[0069] For and the process of grid sampling can be expressed as:
[0070]
[0071] Where \(i\) represents the abscissa of the heatmap of the joint point area, and \(j\) represents the ordinate of the heatmap of the joint point area. represents the integer domain. represents the first number of the two-dimensional coordinates of the \(i\)-th row and \(j\)-th column of \(g\) im im represents the second number of the two-dimensional coordinates of the \(i\)-th row and \(j\)-th column of \(g\) im L
[0072] The final coordinate calculation process is as shown in (c) in Figure 1 In order to ensure that the amplitudes of all heatmaps are consistent, first apply the layer normalization (LN) function to the cropped local region \(\varphi\) L (\(P(u) * k\) s ) for normalization. Subsequently, predict the integer maximum activation coordinate \(u\) from the normalized local region through an MLP im The offset \(u\) between the maximum activation coordinate \(u\) of Note that before feeding into the MLP, the local region is flattened into a one-dimensional space. Finally, add the offset \(u\) of to \(u\) im to obtain the predicted joint position \(u\). In summary, the formulas of all nested functions in the DCD module are expressed by predicting the coordinate \(u\) as:
[0073]
[0074] Where \(u\) of is the offset between the integer maximum response point coordinate and the true joint point coordinate, MLP represents the multi-layer perceptron, LN represents layer normalization, and argmax thres (·) represents the coordinate decoding function proposed in this embodiment.
[0075] Furthermore, in the DCD module, in order to directly predict the position \(u\), a coordinate-based loss function \(L\) c is designed. The coordinate-based loss functions corresponding to layer normalization and the multi-layer perceptron are \(L\) c :
[0076]
[0077] Where \(g\) k is the true joint point coordinate of the \(k\)-th joint point, \(u\) k is the predicted coordinate of the \(k\)-th joint point, \(K\) is the total number of joint points, \(t\) is the offset threshold, and the predicted coordinates with large deviations are excluded through the indicator function When \(L\) cWhen supplemented to heatmap-based supervision, this threshold setting can effectively improve the stability of training.
[0078] With the coordinate-based loss function being L c , the corresponding DCD in the method of the present invention can transform the fixed coordinate decoding process into a learnable operation and be end-to-end guided by the true coordinates. The coordinate-based loss function enables the network to dynamically adjust the prediction distribution of the heatmap according to the true coordinates, thus realizing the transformation from the prior distribution to the result-driven distribution.
[0079] The loss function of the network structure corresponding to obtaining the heatmap of the joint point area from the human body pose image is expressed as
[0080] In the formula is the true heatmap of the k-th joint point, is the multi-modal prediction heatmap of the k-th joint point.
[0081] Generally speaking, the training objectives of using DCD in the human body pose estimation model include the true heatmap H g and the coordinates g, and the corresponding loss functions are L h and L c . The model can be end-to-end trained by optimizing the overall objective loss function L shown as follows:
[0082]
[0083] Perform end-to-end training on the network structure.
[0084] Example:
[0085] In the heatmap-based human body pose estimation model, the differentiable coordinate decoder in the method of the present invention can be used as the decoding module in the last step. The specific steps are as follows:
[0086] 1. Use the existing human body pose estimation model structure to generate the prediction heatmap H u , and then input it into DCD to obtain the coordinate decoding result u.
[0087] 2. DCD needs to be trained together with the connected human body pose estimation network, and the loss function for training is:
[0088]
[0089] 3. Input the image containing the human body into the trained complete model, and the output of predicting the positions of the human body joint points can be obtained.
[0090] Verification experiment:
[0091] The test was carried out on the COCO dataset. The unified input image size was 256 pixels × 192 pixels. The hyperparameters s was set to 6, t was set to 2, and ∈ was set to 1 - 1e-5. The basic models used were HRNet-W32 respectively. The results are shown in Table 1.
[0092] Table 1 Validation of the effectiveness of the DCD module
[0093] Method AP AR Base Model 74.4 79.8 Base Model + DCD 76.2 81.3
[0094] The experimental results prove that the DCD proposed by the method of the present invention can effectively improve the prediction accuracy of the human pose estimation model.
[0095] Although the present invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed, as long as they do not depart from the spirit and scope of the present invention as defined by the appended claims. It should be understood that the different dependent claims and the features described herein can be combined in a manner different from that described in the original claims. It should also be understood that the features described in connection with a single embodiment can be used in other described embodiments.
Claims
1. A differentiable joint point coordinate decoding method for human pose estimation, characterized in that including obtaining a multi-peak prediction heat map by passing a human body pose image through a backbone network and a network head; smoothing the multi-peak prediction heat map to obtain a single-peak prediction heat map, and determining the integer maximum response point coordinates of the single-peak prediction heat map based on peak point indices; generating a grid on the single-peak prediction heat map to obtain a single-peak grid map, and then centering on the integer maximum response point coordinates, cropping a joint point region heat map from the single-peak grid map based on a set length; predicting the joint point region heat map using layer normalization and a multi-layer perceptron to obtain the offset between the integer maximum response point coordinates and the true joint point coordinates; and then calculating the predicted coordinates by combining the offset with the integer maximum response point coordinates.
2. The differentiable joint coordinate decoding method for human pose estimation according to claim 1, wherein The multi-peak prediction heat map is a heat map frame H that contains a joint point u , and the heat map frame H u is regarded as a two-dimensional discrete distribution array, which is expressed as the distribution P(u) of the predicted coordinates u in the two-dimensional coordinate system.
3. The differentiable joint point coordinate decoding method for human pose estimation according to claim 2, characterized in that Adopt a learnable low-pass filter kernel k s Smooth P(u) through a convolution operation to obtain a unimodal prediction heat map P(u)*k s .
4. The differentiable joint point coordinate decoding method for human pose estimation according to claim 3, characterized in that, For the unimodal prediction heatmap P(u)*k s Perform threshold-based argmax function calculation to obtain the coordinates u of the integer maximum response point im : First, convert the unimodal prediction heatmap P(u)*k s to a Dirac-like distribution δ l (u im ): where ∈ is a threshold parameter, is the set of real numbers, Sum(·) represents the summation of all elements, represents the indicator function; Then, use discrete integration to calculate the coordinates u of the integer maximum response point im : where H is the height of the unimodal predicted heatmap P(u)*k s and W is the width of the unimodal predicted heatmap P(u)*k s .
5. The differentiable joint point coordinate decoding method for human pose estimation according to claim 4, wherein The joint point region heat map φ is obtained by cropping the unimodal grid map L (P(u)*k s ,u im ), with the size of (2s + 1) × (2s + 1), where s is the set length; The heatmap φ of the joint point region L (P(u)*k s ,u im ) of the sample grid g im is expressed as:
6. The differentiable keypoint coordinate decoding method for human pose estimation according to claim 5, wherein Let U denote the single-peak prediction heat map and V denote the cropped joint point region heat map: For and where \(i\) represents the abscissa of the heat map of the joint point area, and \(j\) represents the ordinate of the heat map of the joint point area, represents the integer domain, represents the first number of the two-dimensional coordinates of the \(i\)-th row and \(j\)-th column of \(g\) im ; and represents the second number of the two-dimensional coordinates of the \(i\)-th row and \(j\)-th column of \(g\). im It should be noted that the original text seems to have some incomplete or unclear expressions in terms of the overall logic and variable definitions. The above translation is based on the best understanding of the provided text. If there are any inaccuracies, it may be necessary to further clarify the original content.
7. The differentiable joint point coordinate decoding method for human pose estimation according to claim 6, characterized in that, The predicted coordinate u is: where u of is the offset between the integer maximum response point coordinates and the true joint point coordinates, MLP represents the multi-layer perceptron, LN represents layer normalization, and argmax thres (·) represents the coordinate decoding function.
8. The differentiable joint point coordinate decoding method for human pose estimation according to claim 7, wherein The loss function corresponding to layer normalization and the multi-layer perceptron is L c : where g k is the true joint coordinate of the k-th joint point, u k is the predicted coordinate of the k-th joint point, K is the total number of joint points, and t is the offset threshold.
9. The differentiable joint point coordinate decoding method for human pose estimation according to claim 8, wherein The loss function of the network structure from the human body pose image to obtaining the heat map of the joint point area is expressed as where is the true heatmap of the k-th joint point, is the multi-peak predicted heatmap of the k-th joint point.
10. The differentiable joint point coordinate decoding method for human pose estimation according to claim 9, wherein The overall objective loss function L is: Performing end-to-end training on the network structure.