Human body posture estimation method based on multi-scale data self-adaption
The multi-scale data adaptive method with SDB modules dynamically adjusts network parameters to normalize features across varying scales, enhancing human pose estimation accuracy by addressing scale-related distortions.
Patent Information
- Application Number
- CN202510462742.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-15
AI Technical Summary
When processing human images of different scales, the prior art destroys the morphological characteristics of the original image by scaling them to the same scale, making it difficult to improve the accuracy, and using multiple neural networks to process complex and resource requirements.
A multi-scale data adaptive human posture estimation method is designed, and the feature map is dynamically adjusted through the backbone network and the scale-related domain bridge module to achieve scale-aware feature adaptation, and the same neural network is used to process images of different scales.
It improves the accuracy of human posture estimation, breaks through the representation deviation caused by scale changes, and significantly improves the model's adaptability.
Smart Images

Figure CN120318861A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a human pose estimation method based on multi-scale data adaptation, belonging to the technical field of computer vision. Background Technique
[0002] The goal of human pose estimation is to determine the positions or spatial positions of the body key points (parts / joints) of a person from a given image or video. This technology obtains the pose of a jointed human body based on the observation of the image, and the jointed human body is composed of joints and rigid parts.
[0003] As an important basic research in the field of computer vision, human pose estimation is the technical foundation for tasks such as action recognition, human intention prediction, and intelligent monitoring. Current human pose estimation models mostly use the top-down processing paradigm with relatively high accuracy. This paradigm first scales human body images of different scales to the same scale, and then uses the same deep neural network model to perform human pose estimation operations. Although existing methods have achieved good performance, scaling human body images of different scales to the same scale will destroy the morphological features and pixel distributions of the original and real human body images, making it difficult to further improve the accuracy of human pose estimation.
[0004] To solve the above problems, a feasible method is to use different neural networks to process the normalized images containing different scale features respectively, but this method is complex in operation and has high requirements for computing and storage resources. Summary of the Invention
[0005] Aiming at the problem of low accuracy in pose estimation using the same neural network model for original pedestrian image data of different scales, the present invention provides a human pose estimation method based on multi-scale data adaptation.
[0006] A human pose estimation method based on multi-scale data adaptation of the present invention includes:
[0007] Design a backbone network, including a first convolutional module, a first scale-related domain bridging module, a second convolutional module, a second scale-related domain bridging module, and a third convolutional module;
[0008] Preprocess the original image to obtain an input image of a set scale and obtain a scale scaling factor;
[0009] The input image is passed through the first convolutional module to obtain a low-level feature map. The first scale-related domain bridging module obtains a scale-normalized low-level feature map based on the low-level feature map and a scale factor; the scale-normalized low-level feature map is passed through the second convolutional module to obtain a human body structure feature map. The second scale-related domain bridging module obtains a scale-normalized human body structure feature map based on the human body structure feature map and the scale factor; the scale-normalized human body structure feature map is then passed through the third convolutional module to obtain a human joint point feature map. The human joint point feature map is passed through the network head to obtain a heatmap prediction value of the human joint points; human pose estimation is achieved based on the heatmap prediction value.
[0010] The network structures of the first scale-related domain bridging module and the second scale-related domain bridging module are the same.
[0011] The first scale-related domain bridging module and the second scale-related domain bridging module respectively dynamically adjust the input feature map in combination with the scale factor, so that the output scale-normalized feature map adapts to the scale of the human body instance bounding box in the original image, realizing scale-aware feature adaptation.
[0012] According to the human pose estimation method based on multi-scale data adaptation of the present invention, the process by which the first scale-related domain bridging module obtains the scale-normalized low-level feature map includes:
[0013] The low-level feature map is passed through downsampling convolution and upsampling convolution to obtain a first global feature, and then passed through the Sigmoid activation function to generate a first mask map with a value range between 0 and 1.
[0014] The scale factor is passed through a multi-layer perceptron and the Softmax function to generate a set of coefficients a1, and combined with a set of learnable parameters W1 for parameter weighting to obtain the convolution kernel of the scale convolution. Then the low-level feature map is passed through the scale convolution to obtain a first feature map after scale adaptation.
[0015] The first feature map after scale adaptation and the first mask map select the features with scale changes through pixel-by-pixel multiplication, and then add them to the low-level feature map to obtain the scale-normalized low-level feature map.
[0016] According to the human pose estimation method based on multi-scale data adaptation of the present invention, the process by which the second scale-related domain bridging module obtains the scale-normalized human body structure feature map includes:
[0017] The human body structure feature map is passed through downsampling convolution and upsampling convolution to obtain a second global feature, and then passed through the Sigmoid activation function to generate a second mask map with a value range between 0 and 1.
[0018] The scale scaling factor generates a set of coefficients a2 through a multi-layer perceptron and a Softmax function, combines a set of learnable parameters W2 for parameter weighting to obtain the convolution kernel of the scale convolution, and then makes the human body structure feature map pass through the scale convolution to obtain the second feature map after scale adaptation;
[0019] The second feature map after scale adaptation and the second mask map select the features with scale changes through pixel-by-pixel multiplication, and then add them to the human body structure feature map to obtain the scale-normalized human body structure feature map.
[0020] According to the human pose estimation method based on multi-scale data adaptation of the present invention, a loss function is calculated based on the predicted value and the ground truth of the heat map, and the network parameters of the backbone network and the network head are adjusted, as well as the learnable parameters W1 and the learnable parameters W2 are adjusted.
[0021] According to the human pose estimation method based on multi-scale data adaptation of the present invention, the scale scaling factor is expressed as s:
[0022]
[0023] where h o is the height of the original image, h i is the height of the input image, w o is the width of the original image, w i is the width of the input image. According to the human pose estimation method based on multi-scale data adaptation of the present invention, the first mask map is expressed as M1:
[0024] M1 = Sigmoid(Conv up (Conv down (F in1 ))),
[0025] where F in1 represents the low-level feature map, Conv down represents the downsampling convolution, Conv up represents the upsampling convolution.
[0026] According to the human pose estimation method based on multi-scale data adaptation of the present invention, the scale-normalized low-level feature map is expressed as F out1 :
[0027] F out1 = F in1 + F adapt1 × M1,
[0028] where F adapt1 represents the first feature map after scale adaptation;
[0029] F adapt1 = (a11 W 11 +...+a 1n W 1n )*F in1 ,
[0030] a1 = {a 11 ,…,a 1n} = Softmax(MLP(S)),
[0031] W1 = {W 11 ,...,W 1n},
[0032] where a 1n is the nth element of coefficient a1, n is the set number of elements, and W 1n is the nth element of learnable parameter W1, and MLP represents a multi - layer perceptron.
[0033] According to the human pose estimation method based on multi - scale data adaptation of the present invention, the second mask image is represented as M2:
[0034] M2 = Sigmoid(Conv up (Conv down (F in2 ))),
[0035] where F in2 represents the human body structure feature map.
[0036] According to the human pose estimation method based on multi - scale data adaptation of the present invention, the scale - normalized human body structure feature map is represented as F out2 :
[0037] F out2 = F in2 + F adapt2 ×M2,
[0038] where F adapt2 represents the second feature map after scale adaptation;
[0039] F adapt2 =(a 21 W 21 +...+ a 2n W 2n )*F in2 ,
[0040] a2 = {a 21 ,…,a 2n} = Softmax(MLP(S)),
[0041] W2 = {W 21 ,...,W 2n},
[0042] where a 2n is the n-th element of coefficient a2, and W 2n is the n-th element of learnable parameter W2.
[0043] According to the human pose estimation method based on multi-scale data adaptation of the present invention, the loss function is expressed as L h :
[0044]
[0045] where K represents the number of human joint points, is the ground truth heatmap of the k-th human joint point, is the predicted heatmap of the k-th human joint point.
[0046] Advantageous effects of the present invention: The method of the present invention helps to break through the performance bottleneck of human pose estimation by dynamically adjusting the neural network parameters to handle images with different scale features.
[0047] The method of the present invention proposes a Scale-Aware Domain Bridging (SDB) module for converting multiple scale-dependent feature maps into a unified intermediate domain. The SDB module dynamically adjusts the feature maps to adapt to the scale of the human instance bounding box in the original image, thereby achieving scale-based feature normalization.
[0048] The method of the present invention effectively solves the representation deviation problem caused by scale changes in human pose estimation. It can improve the adaptability of the model to human instances of different scales and significantly enhance the accuracy of human pose estimation. Brief Description of the Drawings
[0049] Figure 1 is a schematic diagram of the human pose estimation model of the human pose estimation method based on multi-scale data adaptation of the present invention;
[0050] Figure 2 is a schematic diagram of the network structure of the first scale-aware domain bridging module. Detailed Embodiments
[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0053] The present invention will be further described below in conjunction with the accompanying drawings, but it is not limited to the present invention.
[0054] Combined with Figure 1 and Figure 2 As shown, the present invention provides a human pose estimation method based on multi-scale data adaption, including,
[0055] Design a backbone network, including a first convolutional module, a first scale-related domain bridging module, a second convolutional module, a second scale-related domain bridging module, and a third convolutional module;
[0056] Preprocess the original image to obtain an input image of a set scale, and obtain a scale scaling factor;
[0057] The input image passes through the first convolutional module to obtain a low-level feature map. The first scale-related domain bridging module obtains a scale-normalized low-level feature map based on the low-level feature map and the scale scaling factor; the scale-normalized low-level feature map passes through the second convolutional module to obtain a human structure feature map. The second scale-related domain bridging module obtains a scale-normalized human structure feature map based on the human structure feature map and the scale scaling factor; the scale-normalized human structure feature map then passes through the third convolutional module to obtain a human joint point feature map, and the human joint point feature map passes through the network head to obtain a heatmap prediction value of the human joint point; human pose estimation is realized based on the heatmap prediction value;
[0058] The network structures of the first scale-related domain bridging module and the second scale-related domain bridging module are the same;
[0059] The first scale-related domain bridging module and the second scale-related domain bridging module respectively dynamically adjust the input feature map in combination with the scale scaling factor, so that the output scale-normalized feature map adapts to the scale of the human instance bounding box in the original image, realizing scale-aware feature adaptation.
[0060] In order to normalize the data features in the multi-scale domain to the same feature domain, first, taking the scaling factor as a condition, and using conditional convolution to perform domain adaptation processing on the features with high scale correlation in the data feature map.
[0061] Furthermore, combined with Figure 2 As shown, the process of the first scale-related domain bridging module obtaining the scale-normalized low-level feature map includes:
[0062] The low-level feature map passes through a downsampling convolution and an upsampling convolution to obtain a first global feature, and then passes through a Sigmoid activation function to generate a first mask map with a value range between 0 and 1;
[0063] The scale factor is passed through a multi-layer perceptron and the Softmax function to generate a set of coefficients a1. These coefficients are combined with a set of learnable parameters W1 for parameter weighting to obtain the convolution kernel of the scale convolution. Then, the low-level feature map is passed through the scale convolution to obtain the first feature map after scale adaptation.
[0064] The first feature map after scale adaptation and the first mask map select the features with scale changes through pixel-wise multiplication, and then add them to the low-level feature map to obtain the scale-normalized low-level feature map.
[0065] The process by which the second scale-related domain bridging module obtains the scale-normalized human body structure feature map includes:
[0066] The human body structure feature map is passed through downsampling convolution and upsampling convolution to obtain the second global feature, and then passed through the Sigmoid activation function to generate the second mask map with a value range between 0 and 1.
[0067] The scale factor is passed through a multi-layer perceptron and the Softmax function to generate a set of coefficients a2. These coefficients are combined with a set of learnable parameters W2 for parameter weighting to obtain the convolution kernel of the scale convolution. Then, the human body structure feature map is passed through the scale convolution to obtain the second feature map after scale adaptation.
[0068] The second feature map after scale adaptation and the second mask map select the features with scale changes through pixel-wise multiplication, and then add them to the human body structure feature map to obtain the scale-normalized human body structure feature map.
[0069] In this embodiment, the loss function is calculated based on the predicted value and the ground truth of the heat map, and the network parameters of the backbone network and the network head are adjusted, as well as the learnable parameters W1 and W2 are adjusted.
[0070] In the preprocessing stage of the human pose estimation task, the human body image of any size is resampled to a specific unified input size. The differences between different scale-related domains will reduce the model performance because it is a challenging task for a convolutional neural network to process features in different domains using invariant model parameters. In this embodiment, the term "representation bias" is used to define the difference in feature representation between different scale-related domains. To solve this bias, the scale factor between the original size and the resampled size of each image is first calculated.
[0071] The scale factor is denoted as s:
[0072]
[0073] where h o is the height of the original image, h i is the height of the input image, and w ois the width of the original image, w i is the width of the input image.
[0074] Then, using the scale factor, the corresponding feature maps are modulated through the first scale-related domain bridging module and the second scale-related domain bridging module, as Figure 2 shown. The role of the SDB module is to adapt the features in different domains to a unified intermediate domain. The input of the SDB module includes the feature maps output from the convolutional blocks of any existing backbone network and the scale factor. Considering that the head network structure is too simple to effectively process the adapted feature maps, the SDB module is designed to be inserted after each stage of the backbone network, except for the last stage, as shown in the appendix Figure 1 shown. The output of the SDB module is a feature map adapted to the scale-aware domain, which will be used as the input for the subsequent convolutional layers.
[0075] Given the low-level feature map, it is first input into the convolutional modules for downsampling and upsampling to extract features in a large range of regions. Subsequently, a mask map with a value range between 0 and 1 is generated through a Sigmoid activation function.
[0076] Denote the first mask map as M1:
[0077] M1 = Sigmoid(Conv up (Conv down (F in1 ))),
[0078] where F in1 represents the low-level feature map, Conv down represents the downsampling convolution, and Conv up represents the upsampling convolution.
[0079] The design inspiration for this kind of convolution block with downsampling first and then upsampling comes from the autoencoder. This type of network first compresses the feature map into a small-sized latent representation through the encoder (corresponding to the high-resolution to low-resolution block), and then reconstructs the feature map of the original size from the latent representation through the decoder (corresponding to the low-resolution to high-resolution block), thereby effectively extracting global features.
[0080] Here, the low-level feature map is also input into a scale convolution for feature adaptation. As Figure 2As shown. The scaling factor is input into a multi-layer perceptron (MLP) and a Softmax function to generate a set of coefficients. These coefficients are combined with a set of learnable parameters to serve as the convolution kernel of the scale convolution. Here, the Softmax function is used to normalize the parameters of the convolution kernel, which can effectively prevent overfitting during the training of the entire network. Thanks to the generated convolution kernel, the scale-aware convolution adapts the low-level feature map with domain shift to a unified intermediate domain, thus obtaining the adapted feature map F adapt1 .
[0081] The scale-normalized low-level feature map is denoted as F out1 :
[0082] F out1 = F in1 + F adapt1 × M1,
[0083] In the formula, F adapt1 represents the first feature map after scale adaptation;
[0084] F adapt1 = (a 11 W 11 +... + a 1n W 1n ) * F in1 ,
[0085] a1 = {a 11 ,..., a 1n} = Softmax(MLP(S)),
[0086] W1 = {W 11 ,..., W 1n},
[0087] In the formula, a 1n is the nth element of the coefficient a1, n is the set number of elements, and W 1n is the nth element of the learnable parameter W1, and MLP represents the multi-layer perceptron.
[0088] The coefficient a1 and the learnable parameter W1 are combined into the parameters of the scale-aware convolution kernel through a linear function. In the low-level feature map, the domain-shift features represent the features with scale changes, but some features are scale-invariant. In the design of this embodiment, the mask map M1 can select the features with scale changes through pixel-by-pixel multiplication and cover the scale-invariant features.
[0089] Intuitively, for the regions with high scale invariance under different scaling factors, F in1 can be directly used as F out1 . In contrast, for the regions with high scale variability, F out1Add to F out1 to achieve scale-aware feature adaptation. Here, the scale-aware convolution is dynamically customized according to the scale information of the original image, so the SDB module can help the human pose estimation network adapt to human instance images of various sizes.
[0090] Similarly, the second mask map is denoted as M2:
[0091] M2 = Sigmoid(Conv up (Conv down (F in2 ))),
[0092] where F in2 represents the human body structure feature map.
[0093] The scale-normalized human body structure feature map is denoted as F out2 :
[0094] F out2 = F in2 + F adapt2 × M2,
[0095] where F adapt2 represents the second feature map after scale adaptation;
[0096] F adapt2 = (a 21 W 21 +... + a 2n W 2n ) * F in2 ,
[0097] a2 = {a 21 , …, a 2n} = Softmax(MLP(S)),
[0098] W2 = {W 21 ,..., W 2n},
[0099] where a 2n is the nth element of the coefficient a2, and W 2n is the nth element of the learnable parameter W2.
[0100] In this embodiment, human pose estimation is performed by inserting the SDB module into an existing human pose estimation model, as Figure 1 shown. Then, the scale normalization adjustment of the feature map is performed in combination with the scale factor. The SDB module needs to be trained together with the connected human pose estimation network.
[0101] The loss function is denoted as L h :
[0102]
[0103] where \(K\) represents the number of human joint points, is the ground truth heatmap of the \(k\)-th human joint point, is the predicted heatmap of the \(k\)-th human joint point.
[0104] The gradient of the calculation result of this loss function is backpropagated to the trainable parameters of the scale-related domain bridging module through backpropagation, realizing the training of the scale-related domain bridging module.
[0105] Input an image containing a human body into the model composed of a trained backbone network and a network head, and the output of predicting the positions of human joint points can be obtained.
[0106] Verification experiment:
[0107] Test on the COCO dataset. The unified input image size is 256 pixels × 192 pixels, and the basic model used is HRNet-W32. The results are shown in Table 1.
[0108] Table 1 Verification of the effectiveness of the SDB module
[0109] Method AP AR Base model 74.4 79.8 Base model + SDB module 74.9 80.1
[0110] The experimental results prove that the SDB module proposed by the present invention can effectively improve the prediction accuracy of the human pose estimation model.
[0111] Although the present invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed, as long as they do not depart from the spirit and scope of the present invention as defined by the appended claims. It should be understood that different dependent claims and the features described herein can be combined in a manner different from that described in the original claims. It should also be understood that the features described in connection with a single embodiment can be used in other described embodiments.
Claims
1. A human pose estimation method based on multi-scale data adaption, characterized in that including, designing a backbone network, including a first convolutional module, a first scale-related domain bridging module, a second convolutional module, a second scale-related domain bridging module, and a third convolutional module; preprocessing the original image to obtain an input image of a set scale and obtaining a scale scaling factor; obtaining a low-level feature map by passing the input image through the first convolutional module, and obtaining a scale-normalized low-level feature map by the first scale-related domain bridging module based on the low-level feature map and the scale scaling factor; obtaining a human body structure feature map by passing the scale-normalized low-level feature map through the second convolutional module, and obtaining a scale-normalized human body structure feature map by the second scale-related domain bridging module based on the human body structure feature map and the scale scaling factor; obtaining a human body joint point feature map by passing the scale-normalized human body structure feature map through the third convolutional module, and obtaining a heatmap prediction value of the human body joint points by passing the human body joint point feature map through the network head; realizing human pose estimation based on the heatmap prediction value; the network structures of the first scale-related domain bridging module and the second scale-related domain bridging module are the same; the first scale-related domain bridging module and the second scale-related domain bridging module respectively dynamically adjust the input feature map in combination with the scale scaling factor, so that the output scale-normalized feature map adapts to the scale of the human body instance bounding box in the original image, realizing scale-aware feature adaptation.
2. The method for human pose estimation based on multi-scale data adaptation according to claim 1, wherein the process of the first scale-related domain bridging module obtaining the scale-normalized low-level feature map includes: passing the low-level feature map through downsampling convolution and upsampling convolution to obtain a first global feature, and then generating a first mask map with a value range between 0 and 1 through a Sigmoid activation function; passing the scale scaling factor through a multi-layer perceptron and a Softmax function to generate a set of coefficients a1, combining a set of learnable parameters W1 for parameter weighting to obtain a convolution kernel for scale convolution, and then passing the low-level feature map through the scale convolution to obtain a first feature map after scale adaptation; the first feature map after scale adaptation and the first mask map select features with scale changes through pixel-by-pixel multiplication, and then add them to the low-level feature map to obtain a scale-normalized low-level feature map.
3. The method for human pose estimation based on multi-scale data adaptation according to claim 2, wherein the process of the second scale-related domain bridging module obtaining the scale-normalized human body structure feature map includes: passing the human body structure feature map through downsampling convolution and upsampling convolution to obtain a second global feature, and then generating a second mask map with a value range between 0 and 1 through a Sigmoid activation function; passing the scale scaling factor through a multi-layer perceptron and a Softmax function to generate a set of coefficients a2, combining a set of learnable parameters W2 for parameter weighting to obtain a convolution kernel for scale convolution, and then passing the human body structure feature map through the scale convolution to obtain a second feature map after scale adaptation; the second feature map after scale adaptation and the second mask map select features with scale changes through pixel-by-pixel multiplication, and then add them to the human body structure feature map to obtain a scale-normalized human body structure feature map.
4. The human pose estimation method based on multi-scale data adaptation according to claim 3, wherein a loss function is calculated based on the predicted heatmap value and the ground truth heatmap, and network parameter adjustment of the backbone network and the network head is performed, and adjustment of the learnable parameter W1 and the learnable parameter W2 is performed.
5. The human pose estimation method based on multi-scale data adaptation according to claim 4, wherein the scale scaling factor is expressed as s: where h o is the height of the original image, h i is the height of the input image, w o is the width of the original image, w i is the width of the input image.
6. The human pose estimation method based on multi-scale data adaptation according to claim 5, wherein the first mask map is expressed as M1: M1 = Sigmoid(Conv up (Conv down (F in1 ))), Where F in1 represents the low-level feature map, Conv down represents the downsampling convolution, Conv up represents the upsampling convolution.
7. The human pose estimation method based on multi-scale data adaptation according to claim 6, wherein Denote the scale-normalized low-level feature map as F out1 : F out1 = F in1 + F adapt1 × M1, where F adapt1 represents the first feature map after scale adaptation; F adapt1 = (a 11 W 11 +... + a 1n W 1n ) * F in1 , a1 = {a 11 ,..., a 1n} = Softmax(MLP(S)), W1 = {W 11 ,..., W 1n}, where a 1n is the n-th element of coefficient a1, n is the set number of elements, W 1n is the n-th element of learnable parameter W1, and MLP represents a multi-layer perceptron.
8. The human pose estimation method based on multi-scale data adaptation according to claim 7, wherein the second mask map is expressed as M2: M2 = Sigmoid(Conv up (Conv down (F in2 ))), where F in2 represents the human body structure feature diagram.
9. The human pose estimation method based on multi-scale data adaptation according to claim 8, wherein Denote the scale-normalized human body structure feature map as F out2 : F out2 = F in2 + F adapt2 × M2, where F adapt2 represents the second feature map after scale adaptation; F adapt2 = (a 21 W 21 +...+ a 2n W 2n ) * F in2 , a2 = {a 21 ,..., a 2n} = Softmax(MLP(S)), W2 = {W 21 ,..., W 2n}, where a 2n is the n-th element of coefficient a2, and W 2n is the n-th element of learnable parameter W2.
10. The human pose estimation method based on multi-scale data adaptation according to claim 9, wherein Express the loss function as L h : where K represents the number of human joint points, is the ground truth heatmap of the k-th human joint point, is the predicted heatmap of the k-th human joint point.