A 2D Human Pose Estimation Method Based on Horizontal and Vertical Coordinate Decoupling and Offset Correction
By using the method of horizontal and vertical coordinate decoupling and offset correction in the 2D human posture estimation model, the low-precision problem caused by the quantization effect in the prior art is solved, and more efficient and more accurate posture estimation is achieved.
Patent Information
- Application Number
- CN202210089212.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-01-25
AI Technical Summary
There is a quantization effect in the existing 2D human pose estimation model based on thermal map regression, resulting in low estimation accuracy. Existing solutions such as deconvolution operations, empirical offset correction and two-dimensional offset map prediction have problems such as large calculation overhead, inaccuracy and poor robustness.
The 2D human posture estimation method based on horizontal and vertical coordinate decoupling and offset correction is adopted. By constructing a two-branch network, the relevant features of horizontal and vertical coordinates are extracted respectively, and the position classifier and offset regressor are used to obtain the position and relative offset of the joint nodes, and coordinate correction and refinement are carried out to improve the estimation accuracy.
It effectively reduces the calculation overhead and improves the estimation accuracy, especially under low-resolution input conditions, which performs more stably.
Smart Images

Figure CN114581941B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human pose estimation, and in particular, to a 2D human pose estimation method based on decoupling of horizontal and vertical coordinates and offset correction. Background Art
[0002] In recent years, the field of human pose estimation has attracted the attention of many researchers. Human pose estimation refers to detecting the positions of the main joints of the human body from a given image or video, which is the basis for high-level tasks such as action recognition and human-computer interaction, and has broad application prospects. With the development of convolutional neural networks, many excellent models have emerged in the field of human pose estimation. From the perspective of technical schools, they can be divided into methods based on heatmap regression and methods based on numerical coordinate regression, and the dominant one is the method based on heatmap regression. Although the method based on heatmap regression has brought significant development to the field of human pose estimation, there are still problems of quantization effects in most existing deep neural network models based on heatmap regression. The quantization effects include two aspects: on the one hand, due to the mismatch between the size of the output heatmap of the model and the size of the input image, there is a quantization effect between the input and the output; on the other hand, the numerical coordinates decoded from the heatmap are discrete values, and the decimal part is lost compared with the continuous true values. These problems will inevitably affect the estimation accuracy.
[0003] The existing solutions to the first problem usually increase the size of the output heatmap by adding a deconvolution layer at the end of the model, so as to reduce the difference in size between the input and the output. However, the deconvolution operation will bring a large amount of computational overhead, which is not efficient on devices with limited computing resources. For the solution to the second problem, the existing methods generally shift the position of the maximum peak towards the direction of the second peak by a quarter of a pixel on the heatmap generated by the model, which is equivalent to ±0.25 for the coordinates decoded from the heatmap. However, this method is empirical and not accurate. In addition, some researchers have proposed to use the form of a two-dimensional offset map to predict the relative offset of the joint point coordinates, so as to correct the coordinates decoded from the heatmap. However, the two-dimensional offset map additionally output by the network will bring a large amount of computation, and this method has poor robustness under low-resolution input conditions.
[0004] Generally speaking, the main problems of the existing methods for solving the quantization effects in the pose estimation model are:
[0005] 1. The operation of using deconvolution to increase the heatmap will bring a large amount of computational overhead, which is not efficient on devices with limited computing resources.
[0006] 2. Relying on empirically adding or subtracting 0.25 from the coordinates decoded from the heatmap, the obtained results are not accurate.
[0007] 3. Predicting the relative offset of joint coordinates in the form of a two-dimensional offset map not only increases a large amount of computation but also has low estimation accuracy under low-resolution input conditions. SUMMARY OF THE INVENTION
[0008] In view of the deficiencies of the prior art, the present invention aims to provide a 2D human pose estimation method based on decoupling of horizontal and vertical coordinates and offset correction, which can eliminate the quantization effect in the 2D human pose estimation model based on heatmap regression while overcoming the shortcomings of the above methods, thereby improving the accuracy of pose estimation.
[0009] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0010] A 2D human pose estimation method based on decoupling of horizontal and vertical coordinates and offset correction, having an input image with a size of 3×H i ×W i after preprocessing. The method includes extracting features from the input image, and then constructing a two-branch network for the obtained feature map to separately extract the relevant features of the horizontal and vertical coordinates, thereby decoupling the horizontal and vertical coordinates; wherein, for each network branch, a position classifier and a parallel offset regressor are further constructed to respectively obtain the position of the joint point and the corresponding relative offset, and then the relative offset is used to correct and refine the corresponding position coordinates, so as to obtain the accurate abscissa of the human joint point and the ordinate of the human joint point.
[0011] It should be noted that the image is input into the ResNet-50 backbone network. Assuming that the number of main joint points of the human body to be estimated is K, the network finally outputs K feature maps through a 1x1 convolutional layer, that is, the size of the feature map is K×H f ×W f , and each feature map respectively represents the information of a human joint point.
[0012] It should be noted that a two-branch network is constructed to decouple the horizontal and vertical coordinate features of each obtained feature map, that is, the two branches respectively extract the features related to the horizontal and vertical coordinates in the feature map. Among them, each network branch uses K convolutional kernels with a size of 3x3, a stride of 1, and a padding of 1 to perform depthwise convolution on the K feature maps respectively, batch-normalize the feature maps and pass through the ReLU activation function; each network branch ensures that the output feature map is the same as the input feature Figure 1 in size, that is, the size of the output feature map is also K×H f ×W f .
[0013] It should be noted that the K×H f ×W fThe feature maps are each recombined into a one-dimensional vector, thus obtaining K feature vectors of H f ·W f dimensions. Then, through two fully connected layers respectively, the dimensions of the feature vectors are both mapped to be the same as the width of the input image, that is, K position encoding vectors of W i dimensions are output respectively (k = 1, 2, …, K) and K offset encoding vectors of W i dimensions (k = 1, 2, …, K); among them, each dimension value in the k-th position encoding vector represents the confidence that the position of this column in the input image is the abscissa of the k-th joint point, while each dimension value in the k-th offset encoding vector represents the relative offset of this position from the abscissa of the k-th joint point
[0014] What the ordinate network branch outputs are K ordinate position encoding vectors of the joint points (k = 1, 2, …, K) and offset encoding vectors (k = 1, 2, …, K)
[0015] It should be noted that the coordinates of the joint points can be decoded from the output position encoding vectors and offset encoding vectors. First, the position encoding vector of the abscissa of the k-th joint point is input into the softmax function for probability normalization, and then one-dimensional Gaussian filters with standard deviations of σ 1 and σ 2 are respectively used to smooth the position encoding vector and offset encoding vector output by the model. Among them, at the position where the maximum activation value is taken in the position encoding vector, the rough discrete abscissa value of the k-th joint point can be obtained, that is Then, through the offset encoding vector of the k-th joint point, the position corresponding offset is obtained, that is In this way, the rough joint abscissa is adjusted, and the decimal part is increased to obtain the accurate abscissa of the joint point ordinate Furthermore, by using the position encoding vectors and offset encoding vectors of the abscissa and ordinate of the k-th joint point, the coordinate value of this joint on the input image can be decoded as
[0016] The beneficial effects of the present invention are as follows:
[0017] 1. The present invention uses the supervision information in the form of one-dimensional vectors to replace the two-dimensional heat map, avoiding the deconvolution operation, thereby reducing a large amount of additional computational overhead.
[0018] 2. The model of the present invention predicts the offset while estimating the position of the joint points, which is more reliable than empirical estimation.
[0019] 3. The present invention decouples the horizontal and vertical coordinates and predicts the relative offset of the joint coordinates in the form of one-dimensional offset vectors for the horizontal and vertical coordinates respectively, which has higher estimation accuracy than the form of two-dimensional offset maps under low-resolution inputs. Brief Description of the Drawings
[0020] Figure 1 is a schematic flow diagram of 2D human pose estimation based on decoupling of horizontal and vertical coordinates and offset correction of the present invention;
[0021] Figure 2 is the network structure diagram of ResNet-50 in the present invention;
[0022] Figure 3 is the schematic diagram of depthwise convolution in the present invention;
[0023] Figure 4 is the schematic diagram of 17 main key points of the human body in the present invention. Detailed Embodiment
[0024] The present invention will be further described below with reference to the drawings. It should be noted that this embodiment is based on the technical solution of the present application, and gives detailed implementation manners and specific operation processes, but the protection scope of the present invention is not limited to this embodiment.
[0025] As Figure 1 shown, the present invention is a 2D human pose estimation method based on decoupling of horizontal and vertical coordinates and offset correction, and has an input image with a size of 3×H i ×W i after preprocessing. The method includes extracting features from the input image, and then constructing a two-branch network for the obtained feature map to extract the relevant features of the horizontal and vertical coordinates respectively, so as to decouple the horizontal and vertical coordinates; wherein, a position classifier and a parallel offset regressor are constructed for each network branch to obtain the position of the joint point and the corresponding relative offset respectively, and then the relative offset is used to correct and refine the corresponding position coordinates, so as to obtain the accurate abscissa of the human joint point and the ordinate of the human joint point.
[0026] It should be noted that the image is input into the ResNet-50 backbone network. Assuming that the number of main joint points of the human body to be estimated is K, the network finally outputs K feature maps through a 1x1 convolutional layer, that is, the size of the feature map is K×H f ×W f , and each feature map respectively represents the information of a human joint point.
[0027] It should be noted that a two-branch network is constructed to decouple the horizontal and vertical coordinate features of each obtained feature map, that is, the two branches respectively extract the features related to the horizontal and vertical coordinates in the feature map. Among them, each network branch uses K convolutional kernels with a size of 3x3, a stride of 1, and a padding of 1 to perform depthwise convolution on K feature maps respectively, batch-normalize the feature maps, and pass through the ReLU activation function; each network branch ensures that the output feature map has the same size as the input feature Figure 1 map, that is, the size of the output feature map is also K×H f ×W f .
[0028] It should be noted that for the K×H f ×W f feature map related to the abscissa obtained by decoupling, each feature map is respectively recombined into a one-dimensional vector, that is, K H f ·W f -dimensional feature vectors are obtained, and then respectively passed through two fully connected layers to map the dimensions of the feature vectors to the same as the width of the input image, that is, K W i -dimensional position encoding vectors (k = 1, 2, …, K) and K W i -dimensional offset encoding vectors (k = 1, 2, …, K); among them, the value of each dimension in the k-th position encoding vector represents the confidence that the position of this column in the input image is the abscissa of the k-th joint point, while the value of each dimension in the k-th offset encoding vector represents the relative offset of this position from the abscissa of the k-th joint point.
[0029] What the vertical coordinate network branch outputs are K vertical coordinate position encoding vectors (k = 1, 2, …, K) and offset encoding vectors (k = 1, 2, …, K).
[0030] It should be noted that the coordinates of the joint points can be decoded from the output position encoding vectors and offset encoding vectors. First, the position encoding vector of the abscissa of the k-th joint point is input into the softmax function for probability normalization, and then one-dimensional Gaussian filters with standard deviations of σ 1 and σ 2 are respectively used to smooth the position encoding vector and offset encoding vector output by the model. Among them, at the position where the maximum activation value is taken in the position encoding vector, the rough discrete abscissa value of the k-th joint point can be obtained That is and then the position is obtained through the offset encoding vector of the k-th joint point The corresponding offset That is To adjust the rough joint abscissa And increase its decimal part to obtain the accurate abscissa of the joint point Ordinate Then, using the position encoding vector and the offset encoding vector of the abscissa and ordinate of the k-th joint point, the coordinate value of the joint on the input image can be decoded as
[0031] Embodiment
[0032] Training stage:
[0033] Use the Faster-RCNN detector to obtain the human body instances in each picture of the dataset, and crop out the bounding boxes containing the human body instances from the data and perform data augmentation transformation to expand the input sample space. Specifically, perform horizontal flipping with a probability of 0.5, randomly rotate -30 degrees to 30 degrees with a probability of 0.6, randomly zoom in by 0.75 to 1.25 times, and then scale to the input size 256×256 of the network to obtain the input image I∈R 3×256×256 As Figure 4 Shown, perform corresponding spatial transformation on the Ground Truth of the 17 main joint points of the human body
[0034] Construct supervision information from the Ground Truth of the joint points. Let the coordinates of the Ground Truth of the k-th joint point of the human body after transformation be (x k , y k ), then the supervision information of the abscissa position encoding vector is The supervision information of the abscissa offset encoding vector is R represents the range of the effective area of the joint point. In the present invention, R is set to 4, and the same applies to the ordinate
[0035] Input the input image I into the ResNet-50 backbone network, and obtain the output feature map F∈R with high representation ability through the last 17 1x1 convolutional kernels 17×8×8 , and each channel of the feature map F contains the spatial information of a main joint point of the human body
[0036] Decouple the coordinates of the joint points in the horizontal and vertical directions. Specifically, construct a two-branch network to separate the features related to the joint coordinates in the horizontal and vertical directions in each channel of the feature map F. Each network branch uses a convolutional kernel with a size of 3x3, a stride of 1, and a padding of 1 to perform depthwise convolution on the feature map, while keeping the output feature map the same size as F to avoid loss of spatial information as much as possible, and obtain the feature map F related to the horizontal direction coordinates after decoupling x ∈R 17×8×8 , and the feature map F related to the vertical direction coordinates y ∈R 17×8×8 .
[0037] For the independent horizontal and vertical coordinate network branches, construct a joint point position classifier and a parallel offset regressor on each branch to output the position and relative offset of the joint point in the horizontal / vertical direction respectively. Specifically, taking the horizontal coordinate network branch as an example, first reorganize each channel of the feature map F x into 17 64-dimensional feature vectors, that is, F x ∈R 17×8×8 is reorganized into V x ∈R 17×64 . Then, through two fully connected layers respectively, map the dimension of the feature vector to the same as the width of the input image I, that is, output 17 256-dimensional position encoding vectors (k = 1, 2, …, 17) and 17 256-dimensional offset encoding vectors (k = 1, 2, …, 17). Similarly, the vertical coordinate network branch outputs 17 vertical coordinate position encoding vectors of the joint points (k = 1, 2, …, 17) and offset encoding vectors (k = 1, 2, …, 17).
[0038] Training settings for the overall network: For the position encoding vectors, use the binary cross-entropy loss function for each dimension of the vector, as follows:
[0039]
[0040]
[0041] For the offset encoding vectors, use the Smooth L1 loss function, and the formula is as follows:
[0042]
[0043]
[0044] Where:
[0045]
[0046] The final loss function of the network is
[0047]
[0048] In the present invention, α = 1 and β = 5 are set. The model is trained for 140 epochs using the Adam optimizer, the initial learning rate is set to 0.001, the learning rate decays to 0.0001 after 90 epochs, and the learning rate decays to 0.00001 after the 120th epoch.
[0049] Testing phase:
[0050] When testing the input image, for the position encoding vector and offset encoding vector output by the model, the exact numerical coordinates of the joint points can be jointly decoded. Taking the abscissa of the k-th joint point as an example, first, the position encoding vector of the abscissa of the k-th joint point is input into the softmax function for probability normalization, and then one-dimensional Gaussian filters with standard deviations of σ 1 = 2 and σ 2 = 1 are used to smooth the position encoding vector and offset encoding vector output by the model respectively. Further, at the position where the position encoding vector takes the maximum activation value, the rough discrete abscissa value of the k-th joint point can be obtained That is Then, the position corresponding offset That is is used to adjust the rough joint abscissa and increase its decimal part to obtain the exact abscissa of the joint point Similarly, the ordinate can be obtained Therefore, using the position encoding vector and offset encoding vector of the abscissa and ordinate of the k-th joint point, the coordinate value of this joint on the input image can be decoded as The decoding process for other joint points is the same.
[0051] For those skilled in the art, various corresponding changes can be made according to the above technical solutions and concepts, and all such changes should be included within the protection scope of the claims of the present invention.
Claims
1. A 2D human pose estimation method based on decoupling of horizontal and vertical coordinates and offset correction, characterized in that, the method includes feature extraction of the input image; then a two-branch network is constructed for the obtained feature map to separately extract the relevant features of the horizontal and vertical coordinates, so as to decouple the horizontal and vertical coordinates; wherein, a position classifier and a parallel offset regressor are further constructed for each network branch to obtain the position of the joint point and the corresponding relative offset respectively; The decoupled K×H f ×W f feature map. Each feature map is separately reshaped into a one-dimensional vector, obtaining K H f ·W f dimensional feature vectors. Then, through two fully connected layers respectively, the dimensions of the feature vectors are both mapped to the same width as the input image, and K W i dimensional position encoding vectors and K W i dimensional offset encoding vectors are output. Among them, the value of each dimension in the k-th position encoding vector represents the confidence that the position in the column of the input image is the abscissa of the k-th joint point, while the value of each dimension in the k-th offset encoding vector represents the relative offset of this position from the abscissa of the k-th joint point; Similarly, the vertical coordinate network branch outputs the vertical coordinate position encoding vectors of K keypoints and the offset encoding vectors then the relative offset is used to correct and refine the corresponding position coordinates, so as to obtain the accurate abscissa of the human joint point and the ordinate of the human joint point; The coordinates of the joint points can be decoded from the output position encoding vector and offset encoding vector. First, the position encoding vector of the abscissa of the k-th joint point is input into the softmax function for probability normalization, and then the one-dimensional Gaussian filters with standard deviations of σ 1 and σ 2 are used to smooth the output position encoding vector and offset encoding vector of the model respectively. Among them, at the position where the position encoding vector takes the maximum activation value, the rough discrete abscissa value of the k-th joint point can be obtained That is Then, the offset corresponding to the position is obtained through the offset encoding vector of the k-th joint point That is This is used to adjust the rough joint abscissa and increase its decimal part to obtain the accurate abscissa of the joint point The ordinate Then, using the position encoding vector and offset encoding vector of the abscissa and ordinate of the k-th joint point, the coordinate value of the joint on the input image can be decoded as 2. The 2D human pose estimation method based on decoupling of horizontal and vertical coordinates and offset correction according to claim 1, characterized in that, Input the image into the ResNet-50 backbone network. Assuming the number of main human joint points to be estimated is K, the network finally outputs K feature maps through a 1x1 convolutional layer, that is, the size of the feature maps is K×H f ×W f , and each feature map represents the information of a human joint point respectively.
3. The 2D human pose estimation method based on decoupling of horizontal and vertical coordinates and offset correction according to claim 1, characterized in that, Construct a two-branch network to decouple the horizontal and vertical coordinate features of each obtained feature map, that is, the two branches respectively extract the features related to the horizontal and vertical coordinates in the feature map. Among them, each network branch uses K convolutional kernels with a size of 3x3, a stride of 1, and a padding of 1 to perform depthwise convolution on K feature maps respectively, batch-normalize the feature maps, and pass through the ReLU activation function; each network branch ensures that the output feature map is the same size as the input feature map, that is, the output feature map size is also K×H f ×W f .
Citation Information
Patent Citations
A method for recognizing objects in an image based on depth learning
CN109522938A