A heat map prediction method based on feature fusion and hourglass network

By adopting a heat map prediction method based on feature fusion and hourglass network in pose estimation, combining spatial attention and multi-feature fusion, the problems of insufficient heat map prediction accuracy and large background interference in the prior art are solved, and the high-precision and robust heat map prediction effect is achieved.

CN114998697BActive Publication Date: 2025-05-13SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210618197.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-01
Publication Date
2025-05-13
Estimated Expiration
2042-06-01

AI Technical Summary

Technical Problem

When the prior art improves the prediction accuracy of heat maps in pose estimation, the designed feature extraction structure is difficult to make full use of the prediction heat map spatial information produced by the intermediate process, is susceptible to background interference, and is not able to fully explore the context semantic relationship between feature extraction units at different levels.

Method used

Using a heat map prediction method based on feature fusion and hourglass network, a cascading hourglass network and heat map fusion module are designed by introducing spatial attention and multi-feature fusion, and combining multi-feature aggregation network and heat map aggregation module to enhance the connection between feature context semantic information, improve feature expression ability and stability of prediction results.

Benefits of technology

Higher precision heat map prediction is achieved, background interference is reduced, the robustness of the method is improved, the prediction heat map spatial information of the intermediate process can be more effectively utilized, and the context semantic relationship between feature extraction units at different levels is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114998697B_ABST
    Figure CN114998697B_ABST
Patent Text Reader

Abstract

The present invention discloses a heat map prediction method based on feature fusion and hourglass network, which includes: a heat map fusion module based on spatial attention, and a multi-feature fusion hourglass network that strengthens semantic connections. The main steps are: (1) selecting a public image data set and preprocessing it, the preprocessing including resizing and feature expansion; (2) using the hourglass network to obtain the predicted heat map and feature map of the input image; (3) using the heat map fusion module to introduce the spatial attention feature to improve the spatial guidance ability of the predicted heat map; (4) proposing a multi-feature aggregation hourglass network to take the fused heat map and feature map as input, aggregate semantic information at different levels, and output a more accurate predicted heat map. The present invention can make more full use of the spatial information in the heat map, integrate the associations between features at different levels, and achieve the acquisition of high-precision predicted heat maps.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of deep learning and computer vision, and in particular to a heat map prediction method based on deep learning. Background Art

[0002] In the field of posture estimation, there are two mainstream methods for obtaining key points from abstracted high-dimensional features, one is a regression-based method, and the other is a detection-based method. Although the regression-based method is simple and direct, its accuracy is somewhat insufficient, so the industry often uses the detection-based method. Heatmap is the most commonly used detection-based method. Using heatmap as the estimation target will make the network loss change softer and make it easier to predict key points. Therefore, obtaining a high-precision predicted heatmap has become a key goal in the field of posture estimation.

[0003] However, many methods to improve the accuracy of prediction heatmaps aim to design better feature extraction structures, and then use cascade iterations to optimize more accurate prediction heatmaps. However, this method is currently too costly. Many excellent feature extraction structures have been designed, and it is increasingly difficult to improve the accuracy of prediction heatmaps on this basis. Many current methods are simple cascade-designed feature extraction structures, which cannot mine the spatial information of the prediction heatmaps produced in the intermediate process and are easily disturbed by the background. Secondly, for many designed feature extraction structures, many methods do not continue to explore more possibilities of the original feature extraction structures. Summary of the invention

[0004] Purpose of the invention: In order to better utilize the predicted heat map of the intermediate process, reduce background interference, and improve the robustness of the heat map estimation, the present invention proposes a posture key point heat map prediction method based on feature fusion and hourglass network. This method introduces spatial attention and multi-feature fusion in the cascaded hourglass network, which can obtain a higher precision predicted heat map, and can reduce background interference, thereby improving the robustness of the method.

[0005] To solve the above problems, the present invention adopts the following technical solutions:

[0006] A heat map prediction method based on feature fusion and hourglass network, the method comprising the following steps:

[0007] S1. Select a public image dataset as the input image, then preprocess the input image, use the maximum pooling and residual module to reduce the width and height of the input image and expand the channel features to obtain the input features;

[0008] S2. According to the input features obtained in step S1, the hourglass network and the convolution layer are used to extract the input features to obtain a predicted heat map and a feature map;

[0009] S3. Use the heat map fusion module to splice the predicted heat map obtained in step S2 and the input features obtained in step S1 to obtain a fused heat map, and add the fused heat map, the input features obtained in step S1, and the feature map obtained in step S2 together as the output of S3;

[0010] 3.1: Design heatmap fusion module

[0011]

[0012] Among them, f1 is the input feature obtained in step S1, f2 is the feature extracted by the hourglass network in S2, w represents the convolution kernel parameter, * represents the convolution operation, Represents feature concatenation.

[0013] S4. Use the hourglass network and convolutional layer to extract features from the output of S3 to obtain a new prediction heat map and a new feature map, then use the heat map fusion module to splice the new prediction heat map with the input features obtained in step S1 to obtain a new fused heat map, and add the new fused heat map, the input features obtained in step S1, and the new feature map together as the output of S4;

[0014] S5. Add the input features obtained in step S1 and the predicted heat map and feature map obtained in step S4 together to obtain a final predicted heat map;

[0015] 5.1: Designing a multi-feature aggregation module

[0016] The multi-feature aggregation module consists of multiple convolution operations and a splicing operation. The multi-feature fusion hourglass network contains multiple hourglass networks. The multi-feature aggregation module aggregates the three different levels of features (input features, feature maps, and heat maps) of the previous hourglass network as the input of the next hourglass network. The calculation formula is:

[0017]

[0018] Among them, f in 、f fm 、f hm are input features, feature maps and heat maps, represents feature concatenation, and R represents the residual module.

[0019] 5.2: Design heatmap aggregation module

[0020] Heatmap aggregation module: The heatmap aggregation module is used to balance the prediction tendencies of multiple hourglass networks in the multi-feature fusion hourglass network, and aggregate different prediction heatmaps to obtain an average prediction heatmap; the formula is as follows:

[0021]

[0022] Among them, ρ t Represents different intermediate prediction heat maps in the multi-feature aggregation network, ρ fusion is the final heat map of the output.

[0023] S6. Set the loss function and train the network to obtain the mapping model.

[0024] Beneficial effects:

[0025] The present invention aims at the problems that the existing heat map prediction network based on stacked hourglass network does not fully utilize the spatial scale information of the intermediate prediction heat map, is seriously disturbed by the background, and does not fully explore the contextual semantic relationship between feature extraction units at different levels. The present invention proposes a human hand heat map prediction method based on feature fusion and hourglass network. The present invention combines the ideas of spatial attention and multi-stage feature extraction, and uses the heat maps estimated by the feature extraction network at different stages to continuously correct the original input of the network, strengthen the association between the image space and the feature space, and realize the fusion of spatial attention features of features at different stages. The heat map fusion module of the progressive hourglass network enhances the connection between feature context semantic information and improves the feature expression ability. The present invention also designs a multi-feature aggregation network, which integrates the association between local and global information, strengthens the connection between high-level semantic information, and improves the stability of prediction results through input aggregation, feature aggregation, and prediction aggregation, and finally realizes high-precision heat map prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a schematic diagram of the framework of the entire system of the invention.

[0027] Figure 2 This is a schematic diagram of the heat map fusion module.

[0028] Figure 3 It is a schematic diagram of multi-feature aggregation network. DETAILED DESCRIPTION

[0029] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments. Pytorch is selected as the deep learning framework for network implementation under the Linux operating system. The hardware environment is Nvidia RTX3070 and Intel (R) i5-10600KF CPU. The deep learning framework and the hardware environment are only for implementing and testing the proposed method to confirm the effectiveness of the super-resolution reconstruction method proposed in this patent. It should be understood that these examples are only used to illustrate the present invention and are not intended to limit the scope of the present invention. After reading the present invention, modifications to various equivalent forms of the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0030] A heat map prediction method based on feature fusion and hourglass network. The system framework diagram is as follows: Figure 1 shown.

[0031] The specific steps include:

[0032] Step 1: Select the public datasets Human3.6M and MPII Human Pose as training datasets. Use the preprocessed images in the image dataset as input images, and use convolution and maximum pooling operations to downsample and expand the input images. At this time, the width and height of the input features are 64×64, and the feature channels are changed from 3 to 256;

[0033] Step 2: Input the image processed in step 1 Figure 2 In the dashed box H1 shown, H1 is composed of an hourglass network, a convolutional layer, a residual module, a channel splicing and other operations. The input feature first passes through the hourglass network to obtain an intermediate feature, and then two convolution operations are used to obtain a feature map and a heat map respectively. The obtained heat map is spliced ​​with the input feature obtained in step 1, and the spliced ​​feature is input into the convolutional layer to obtain a fused heat map. Finally, the fused heat map is added together with the input feature and the feature map, and it is used as the output of H1;

[0034] Step 3: Similar to step 2, use the output of step 2 as Figure 2 The input of the dashed box H2 shown in the figure has the same structure as H1, and the added features are obtained;

[0035] Step 4: Design a multi-feature aggregation network and use the connection features obtained in step 3 as the multi-feature aggregation network (i.e. Figure 1 The structure of the multi-feature aggregation network is as follows: Figure 3 As shown;

[0036] 4.1: Multi-feature aggregation module

[0037] The publicly available hourglass network is used as the feature extraction unit. The multi-feature aggregation network consists of multiple hourglass networks, multi-feature aggregation modules, and heat map aggregation modules. The multi-feature aggregation module consists of three splicing operations and a residual module. The multi-feature aggregation module takes the different levels of features of the previous hourglass network as input and outputs the aggregated new features. The aggregated new features are input into the next hourglass network;

[0038] 4.2: Heatmap Aggregation Module

[0039] The heat map aggregation module is used to aggregate the heat map outputs of each level in the multi-feature aggregation network to balance the prediction tendencies of multiple hourglass networks in the multi-feature fusion network;

[0040] Step 5: Train a heatmap prediction network based on feature fusion and hourglass network

[0041] 5.1 Setting the loss function

[0042] The loss function consists of the loss of all heatmaps produced by the network, summed over them, with the weights kept the same, and the individual parameters in the network are calculated by minimizing the loss between the predicted heatmap and the heatmap generated by the true labels.

[0043] 5.2 Setting training parameters

[0044] The number of training rounds is set to 40, and the initial learning rate of the network is set to 0.0001, and it will be reduced to 1 / 10 after 30 rounds. As the number of training rounds increases, the learning rate is appropriately reduced. The network is iteratively trained using the Adam optimizer.

[0045] It should be noted that the above implementation examples are only cases made for clear explanation, and are not intended to limit the implementation methods. It is not necessary and impossible to list all the implementation methods here. All components not specified in this embodiment can be implemented using existing technologies. For those of ordinary skill in the art, several improvements and modifications can be made without departing from the principles of the present invention, and these improvements and modifications should also be considered as the scope of protection of the present invention.

Claims

1. A heat map prediction method based on feature fusion and hourglass network, comprising the following steps: S1. Select a public image dataset as the input image, then preprocess the input image, use the maximum pooling and residual module to reduce the width and height of the input image and expand the channel features to obtain the input features; S2. According to the input features obtained in step S1, the hourglass network and the convolution layer are used to extract the input features to obtain a predicted heat map and a feature map; S3. Use the heat map fusion module to splice the predicted heat map obtained in step S2 and the input features obtained in step S1 to obtain a fused heat map, and add the fused heat map, the input features obtained in step S1, and the feature map obtained in step S2 together as the output of S3; S4. Use the hourglass network and convolutional layer to extract features from the output of S3 to obtain a new prediction heat map and a new feature map, then use the heat map fusion module to splice the new prediction heat map with the input features obtained in step S1 to obtain a new fused heat map, and connect the new fused heat map with the input features obtained in step S1 and the new feature map as the output of S4; S5. Add the input features obtained in step S1 and the predicted heat map and feature map obtained in step S4 together, input them into the multi-feature aggregation hourglass network, and obtain the final predicted heat map; S6. Set the loss function and train the network to obtain a mapping model; Step S5 specifically includes the following steps: 5.1: Designing a multi-feature aggregation module The multi-feature aggregation module consists of three splicing operations and a residual module; the multi-feature fusion hourglass network consists of multiple hourglass networks. The multi-feature aggregation module aggregates the three different levels of features of the previous hourglass network, including input features, feature maps and heat maps, as the input of the next hourglass network. The calculation formula is: Among them, f in 、f fm 、f hm are input features, feature maps and heat maps, represents feature concatenation, and R represents the residual module; 5.2: Design heatmap aggregation module Heatmap aggregation module: The heatmap aggregation module is used to balance the prediction tendencies of multiple hourglass networks in the multi-feature fusion hourglass network, and aggregate different prediction heatmaps to obtain an average prediction heatmap; the formula is as follows: Among them, ρ t Represents different intermediate prediction heat maps in the multi-feature aggregation hourglass network, ρ fusion is the final heat map of the output.

2. The heat map prediction method based on feature fusion and hourglass network according to claim 1, characterized in that: Step S3 specifically includes the following steps: 3.1: Design heatmap fusion module Among them, f1 is the input feature obtained in step S1, f2 is the feature extracted by the hourglass network in S2, w represents the convolution kernel parameter, * represents the convolution operation, Represents feature concatenation.

Citation Information

Patent Citations

  • Human pose estimation based on directional image fusion

    CN109033946A

  • Human body posture estimation method based on hourglass network in combination with attention mechanism

    CN112232134A