A Multi-scale Fusion Human Pose Estimation Method Based on Improved HRNet
By improving the multi-scale fusion method of HRNet network, the problems of spatial information loss and insufficient semantic information of high-resolution feature maps in human pose estimation are solved, and higher precision joint node positioning is achieved, and pose estimation performance is improved.
Patent Information
- Application Number
- CN202411104226.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-13
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-08-13
AI Technical Summary
The prior art has problems of loss of spatial information of high-resolution feature maps and insufficient semantic information in human posture estimation, which affects the accuracy of node positioning.
Using the improved HRNet network structure, the multi-scale fusion method of encoder and decoder is combined with the lightweight module MBConv and the multi-scale feature fusion module to realize spatial characterization of high-resolution feature maps and fine-grained feature retention.
It improves the positioning accuracy of human joint nodes and improves the performance of human posture estimation, especially the pose estimation effect in different scenarios and conditions.
Smart Images

Figure CN119579641B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human posture estimation, and in particular to a multi-scale fusion human posture estimation method based on an improved HRNet. Background Art
[0002] The goal of human posture estimation is to accurately locate the joints of the human body. It has many applications, such as human behavior recognition, human-computer interaction, and animation production. It is a very challenging problem in the field of computer vision.
[0003] Representative work addressing human pose estimation involves progressive downsampling during encoding using strided convolution or pooling layers to produce low-resolution representations for extracting high-level semantic information. During decoding, low-resolution feature maps are gradually restored to high-resolution feature maps using transposed convolution or bilinear interpolation upsampling to obtain high-resolution representations. This process has significant drawbacks: during downsampling and subsequent restoration to high resolution, more spatial information in both length and width directions is often lost. This is detrimental for pixel-dense, position-sensitive prediction tasks like pose estimation.
[0004] In terms of feature fusion, most approaches use a skip connection mechanism to connect feature maps of the same scale between the encoder and decoder, helping the model fuse shallow and deep information, effectively integrating coarse-grained and fine-grained features. However, this process has limited effectiveness in improving the representational capabilities of high-resolution feature maps, which is crucial for dense prediction tasks such as human pose estimation.
[0005] Therefore, it is necessary to further explore and study new human pose estimation methods to enhance the spatial representation capabilities of high-resolution feature maps and enrich the semantic information of high-resolution feature maps. Summary of the Invention
[0006] In response to the shortcomings of the existing technology, the present invention develops a multi-scale fusion human posture estimation method based on improved HRNet, which can significantly improve the positioning accuracy of human joints, thereby improving the performance of human posture estimation.
[0007] The technical solution of the present invention to solve the technical problem is:
[0008] A multi-scale fusion human posture estimation method based on improved HRNet includes the following steps:
[0009] S1. Obtaining human pose estimation dataset , Represents the human pose estimation dataset The Zhang pictures, , for the dataset Perform data preprocessing to obtain the preprocessed human pose estimation dataset ;
[0010] S2. Construct an improved HRNet network structure, and use the Zhang pictures in the preprocessed human pose estimation dataset I as the input to the initial stage of the improved HRNet network structure to obtain a high-resolution feature map ;
[0011] S3. In the first stage of the encoder of the improved HRNet network structure, the high-resolution feature map passes through the Transition1 module to obtain a high-resolution feature map and a medium-resolution feature map . The feature maps and pass through the StageModule1 module to obtain a high-resolution feature map and a medium-resolution feature map ;
[0012] S4. In the second stage of the encoder of the improved HRNet network structure, the high-resolution feature map and the medium-resolution feature map pass through the Transition2 module to obtain a high-resolution feature map , a medium-resolution feature map and a low-resolution feature map . The feature maps , and pass through 2 StageModule2 modules to obtain a high-resolution feature map , a medium-resolution feature map and a low-resolution feature map ;
[0013] S5. In the third stage of the encoder of the improved HRNet network structure, the high-resolution feature map , the medium-resolution feature map and the low-resolution feature map pass through the Transition3 module to obtain a high-resolution feature map , a medium-resolution feature map , a low-resolution feature map and a small-resolution feature map . The feature maps , , and obtain a high - resolution feature map through the StageModule3 module and a medium - resolution feature map ;
[0014] S6. In the first stage of the decoder of the improved HRNet network structure, the high - resolution feature map and the medium - resolution feature map obtain a high - resolution feature map, a medium - resolution feature map, and a low - resolution feature map through the multi - scale fusion module Fusion4 and a medium - resolution feature map and a low - resolution feature map , and the feature maps , and obtain a high - resolution feature map, a medium - resolution feature map, and a low - resolution feature map through 2 StageModule2 modules and a medium - resolution feature map and a low - resolution feature map . The high - resolution feature map and obtain a high - resolution feature map through skip connection addition . The medium - resolution feature map and obtain a medium - resolution feature map through skip connection addition . The low - resolution feature map and obtain a low - resolution feature map through skip connection addition ;
[0015] S7. In the second stage of the decoder of the improved HRNet network structure, the high - resolution feature map , the medium - resolution feature map and the low - resolution feature map obtain a high - resolution feature map and a medium - resolution feature map through the multi - scale fusion module Fusion5 and a medium - resolution feature map . The feature maps , obtain a high - resolution feature map and a medium - resolution feature map through the StageModule1 module and a medium - resolution feature map . The high - resolution feature map and obtain a high - resolution feature map through skip connection addition . The medium - resolution feature map and obtain a medium - resolution feature map through skip connection addition ;
[0016] S8. In the third stage of the decoder of the improved HRNet network structure, the high-resolution feature map and the medium-resolution feature map are fused through the multi-scale fusion module Fusion6 to obtain a high-resolution feature map ;
[0017] S9. The high-resolution feature map is input into the detection head to obtain an output feature map . Each channel of the output feature map represents the position of each joint point, and finally the human pose estimation result is obtained.
[0018] Further, the specific steps for obtaining the preprocessed training set and the preprocessed validation set in step S1 are as follows:
[0019] All sample images included in the human pose estimation dataset are uniformly scaled, and then data augmentation strategies are performed, including: random horizontal flipping, random scaling and rotation of the human bounding box, and affine transformation, to obtain the preprocessed human pose estimation dataset ;
[0020] Labels are generated for training on the human pose estimation dataset after using the data augmentation strategy . Specifically, the operation is as follows: Taking the position of each human joint point in the image as the center, a ground-truth heat map is generated using a 2D Gaussian kernel function with a standard deviation of 1 pixel as the label.
[0021] Further, step S3 specifically includes:
[0022] The Transition1 module in the first stage of the encoder contains two branches: and . Branch includes a convolutional layer with a kernel size of 3x3, a stride of 1, and a padding of 1, a batch normalization layer, and a ReLU activation function. Branch also includes a convolutional layer with a kernel size of 3x3, a stride of 1, and a padding of 1, a batch normalization layer, and a ReLU activation function. The high-resolution feature map is mapped through branch to obtain a high-resolution feature map , and is mapped through branch to obtain a medium-resolution feature map ;
[0023] The StageModule1 module in the first stage of the encoder includes a feature extraction module Extract1 and a multi-scale fusion module Fusion1. The feature extraction module Extract1 contains two branches: and , the feature map after being mapped through branch , obtains a high-resolution feature map , the feature map after being mapped through branch , obtains a medium-resolution feature map ;
[0024] In the multi-scale fusion module Fusion1, the mapping relationship formula for obtaining the high-resolution feature map and the medium-resolution feature map is:
[0025] ,
[0026] ,
[0027] where represents the identity mapping, represents a 2x upsampling module for obtaining the high-resolution feature map, represents a 2x downsampling module for obtaining the medium-resolution feature map.
[0028] Furthermore, the specific steps of step S4 include:
[0029] The Transition2 module in the second stage of the encoder contains three branches: , and , branches are all identity mappings, branch includes a convolutional layer with a convolutional kernel size of 3x3, a stride of 2, and a padding of 1, a batch normalization layer, and a ReLU activation function; the high-resolution feature map after being mapped through branch obtains a high-resolution feature map , the medium-resolution feature map after being mapped through branch obtains a medium-resolution feature map , the medium-resolution feature map after being mapped through branch obtains a low-resolution feature map ;
[0030] The StageModule2 module of the second stage of the encoder includes a feature extraction module Extract2 and a multi-scale feature fusion module Fusion2. The feature extraction module Extract2 contains three branches: 、 and , high-resolution feature map Through the branch After mapping, a high-resolution feature map is obtained , medium-resolution feature map Through the branch After mapping, we get the medium resolution feature map , low-resolution feature map Through the branch After mapping, a low-resolution feature map is obtained ;
[0031] In the multi-scale fusion module Fusion2, a high-resolution feature map is obtained. , medium-resolution feature map and low-resolution feature maps The mapping relationship formula is:
[0032] ,
[0033] ,
[0034] ,
[0035] in, represents the identity mapping, Indicates a 4x upsampling module for obtaining high-resolution feature maps. Indicates a 2x upsampling module for obtaining a medium-resolution feature map. Indicates a 4-fold downsampling module for obtaining a low-resolution feature map. Represents a 2x downsampling module to obtain a low-resolution feature map.
[0036] Furthermore, the step S5 specifically includes:
[0037] The Transition3 module of the third stage of the encoder contains four branches: 、 、 and , branch are all identity mappings, branches Includes convolutional layers with a kernel size of 3x3, a stride of 2, and a padding of 1, a batch normalization layer, and a ReLU activation function; high-resolution feature maps Through the branch After mapping, a high-resolution feature map is obtained , a medium-resolution feature map After passing through the branch After mapping, a medium-resolution feature map is obtained , a low-resolution feature map After passing through the branch After mapping, a low-resolution feature map is obtained , the low-resolution feature map passes through the branch After mapping, a small-resolution feature map is obtained ;
[0038] The StageModule3 module in the third stage of the encoder includes a feature extraction module Extract3 and a multi-scale feature fusion module Fusion3. The feature extraction module Extract3 contains four branches: , , and , the resolution feature map After passing through the branch After mapping, a high-resolution feature map is obtained , a medium-resolution feature map After passing through the branch After mapping, a medium-resolution feature map is obtained , a low-resolution feature map After passing through the branch After mapping, a low-resolution feature map is obtained , a small-resolution feature map After passing through the branch After mapping, a small-resolution feature map is obtained ;
[0039] In the multi-scale fusion module Fusion3, the mapping relationship formula for obtaining the high-resolution feature map and the medium-resolution feature map is:
[0040] [[ID=6,4]]
[0041]
[0042] where represents an 8-fold upsampling module for obtaining the high-resolution feature map, represents a 4-fold upsampling module for obtaining the medium-resolution feature map.
[0043] Furthermore, in the multi-scale fusion module Fusion4 in step S6, the high-resolution feature map , the medium-resolution feature map and the low-resolution feature map The mapping relation formula is: , , .
[0044] Furthermore, in the multi-scale fusion module Fusion5 in step S7, the mapping relation formula for obtaining the high-resolution feature map and the medium-resolution feature map is: , .
[0045] Furthermore, in the multi-scale fusion module Fusion6 in step S8, the mapping relation formula for obtaining the high-resolution feature map is: .
[0046] Furthermore, the branches and in the feature extraction module Extract1, the branches , and in the feature extraction module Extract2, and the branches , , and in the feature extraction module Extract3 all contain 4 improved MBConv modules. The improved MBConv module contains 5 consecutive parts:
[0047] The first part: a convolutional layer with a kernel size of 1x1, a stride of 1, and a padding of 0, a batch normalization layer, and a SiLU activation function;
[0048] The second part: a depthwise separable convolution with a kernel size of 3x3, a stride of 1, and a padding of 1, a batch normalization layer, and a SiLU activation function;
[0049] The third part: an SE module: first perform an average pooling operation, then perform two fully connected layer operations. The first fully connected layer uses a SiLU activation function, and the second fully connected layer uses a Sigmoid activation function;
[0050] The fourth part: a convolutional layer with a kernel size of 1x1, a stride of 1, and a padding of 0, and a batch normalization layer;
[0051] The fifth part: perform a DropPath operation, then perform a residual connection between the input feature map and the output feature map in the fourth part, and finally obtain the output feature map of the improved MBConv module.
[0052] The present invention also provides a computing device, including a memory configured to store computer-executable instructions, and a processor configured to execute a multi-scale fusion human pose estimation method based on an improved HRNet when the computer-executable instructions are executed by the processor.
[0053] Generally speaking, through the above technical solutions of the inventive concept, the following beneficial effects are achieved:
[0054] A multi-scale fusion human pose estimation method based on an improved HRNet of the present invention mainly consists of an encoder and a decoder. The encoder is divided into three stages, and each stage generates a low-resolution branch to realize parallel connection of different-resolution branches. An improved lightweight module MBConv is used in each branch to further improve the feature extraction ability of the encoder, and a multi-scale fusion module is adopted at the end of each stage to complete repeated multi-scale feature fusion, which can enrich the semantic information of feature maps with different resolutions. The decoder is divided into three stages, which is a symmetric structure with the encoder. Each stage eliminates a low-resolution branch, so that the generated high-resolution feature map can retain more detailed information. Between the encoder and the decoder, skip connections are introduced to connect the features of the corresponding layers of the encoder and the decoder, thereby increasing the detailed information of the feature map and improving the accuracy of joint point prediction. The method of the present invention can better capture the spatial relationship between human joints, and at the same time ensure that fine-grained features are not lost, thus improving the performance of human pose estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention.
[0056] Figure 1 It is a network flow chart of the present invention.
[0057] Figure 2 It is an improved MBConv module of the present invention.
[0058] Figure 3 It is a multi-scale fusion - upsampling module of the present invention.
[0059] Figure 4 It is a multi-scale fusion - downsampling module of the present invention.
[0060] Figure 5 It is the human pose estimation result of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0061] To clearly illustrate the technical features of this solution, the present invention will be described in detail below through specific embodiments and in conjunction with its accompanying drawings. The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, the components and settings of specific examples are described below.
[0062] Embodiment 1
[0063] In this embodiment, as Figure 1 shown, a multi-scale fusion human pose estimation method based on improved HRNet includes the following steps:
[0064] a) Obtain a human pose estimation dataset , representing the th picture, , and perform data preprocessing on the dataset to obtain the preprocessed dataset .
[0065] a-1) Uniformly scale the size of all sample pictures included in the human pose estimation dataset , and then execute data augmentation strategies, including: random horizontal flipping, randomly scaling and rotating the human bounding box, and affine transformation.
[0066] a-2) For the human pose estimation dataset after using the data augmentation strategy, generate labels for training. The specific operation is: taking the position of each human joint point in the picture as the center, using a 2D Gaussian kernel function with a standard deviation of 1 pixel to generate a ground truth heatmap as the label.
[0067] b) Construct the network structure of improved HRNet, and input the th picture in the preprocessed human pose estimation dataset into the initial stage of the network structure to obtain a high-resolution feature map .
[0068] b-1) The size of the input sample picture is , and after mapping , the feature map is obtained, with a size of . The feature map is then mapped through to obtain the feature map , with a size of . The mapping It includes a convolutional layer, a batch normalization layer, and a ReLU activation function. Among them, the convolutional kernel size of the convolutional layer is 3x3, the stride is 2, and the padding is 1.
[0069] b-2) Feature map After passing through the Bottleneck module with a downsampling layer, a feature map is obtained , whose size is . The Bottleneck module with downsampling, for the input feature map , after mapping , the size of the obtained feature map is ; after mapping , the size of the obtained feature map is ; after mapping , the feature map is obtained, and the size is .; One downsampling operation is performed to obtain the feature map , and the size is . Then pass through the function: to obtain the feature map , and the size is . Among them, the mapping includes a convolutional layer with a convolutional kernel size of 1x1, a stride of 1, and a padding of 0, a batch normalization layer, and a ReLU activation function. includes a convolutional layer with a convolutional kernel size of 3x3, a stride of 1, and a padding of 1, a batch normalization layer, and a ReLU activation function. includes a convolutional layer with a convolutional kernel size of 1x1, a stride of 1, and a padding of 0, and a batch normalization layer. The downsampling operation includes a convolutional layer with a convolutional kernel size of 1x1, a stride of 1, and a padding of 0, and a batch normalization layer.
[0070] b-3) Feature map After passing through 3 Bottleneck modules without a downsampling layer, a high-resolution feature map is obtained, and its size is . The Bottleneck module without downsampling, for the input feature map , after mapping , the size of the obtained feature map is ; after mapping , the size of the obtained feature map is ; after mapping , the feature map is obtained, and the size is . Then pass through the function: to obtain the high-resolution feature map , and the size is . Among them, the mapping It includes a convolutional layer with a kernel size of 1x1, a stride of 1, and a padding of 0, a batch normalization layer, and a ReLU activation function. It includes a convolutional layer with a kernel size of 3x3, a stride of 1, and a padding of 1, a batch normalization layer, and a ReLU activation function. It includes a convolutional layer with a kernel size of 1x1, a stride of 1, and a padding of 0, and a batch normalization layer.
[0071] c) In the first stage of the encoder, the high-resolution feature map passes through the Transition1 module to obtain the high-resolution feature map , and the medium-resolution feature map ; After the feature maps , pass through the StageModule1 module, the high-resolution feature map and the medium-resolution feature map are obtained.
[0072] c-1) The Transition1 module contains two branches: , The high-resolution feature map is mapped through branch to obtain the high-resolution feature map , and is mapped through branch to obtain the medium-resolution feature map . The size of the feature map is , and the size of the feature map is . Branch contains a convolutional layer with a kernel size of 3x3, a stride of 1, and a padding of 1, a batch normalization layer, and a ReLU activation function. Branch also contains a convolutional layer with a kernel size of 3x3, a stride of 1, and a padding of 1, a batch normalization layer, and a ReLU activation function.
[0073] c-2) The StageModule1 module contains two parts: the feature extraction module Extract1 and the multi-scale fusion module Fusion1. The feature extraction module Extract1 contains two branches: , The feature map is mapped through branch to obtain the high-resolution feature map , with a size of ; The feature map is mapped through branch to obtain the medium-resolution feature map , with a size of . Branch and Both contain 4 improved MBConv modules. As Figure 2 shown, the improved MBConv module consists of 5 consecutive parts: (1) a convolutional layer with a kernel size of 1x1, a stride of 1, and a padding of 0, a batch normalization layer, and a SiLU activation function; the width and height remain unchanged, and the number of channels is adjusted to 6 times the number of input channels. (2) A depthwise separable convolution with a kernel size of 3x3, a stride of 1, and a padding of 1, a batch normalization layer, and a SiLU activation function; the width and height remain unchanged, and the number of channels remains 6 times the number of input channels. (3) SE module: First, perform an average pooling operation to adjust the feature map to a one-dimensional vector, with the dimension still being 6 times the number of input channels; then perform two fully connected layer operations. The first fully connected layer uses the SiLU activation function, and the resulting output dimension is of the number of input channels, and the second fully connected layer uses the Sigmoid activation function, and the resulting output dimension is 6 times the number of input channels. Multiply the input feature map element-wise with it to obtain the final output feature map of the SE module, with the width and height remaining unchanged and the number of channels being 6 times the number of input channels. (4) A convolutional layer with a kernel size of 1x1, a stride of 1, and a padding of 0, a batch normalization layer, and the number of channels is reduced by 6 times to obtain an output feature map with the same size as the input feature Figure 1 . (5) Perform the DropPath operation with a probability of 0.2, and then perform a residual connection between the input feature map and the output feature map in (4) to obtain an output feature map of the improved MBConv module with the same size as the input feature Figure 1 .
[0074] c-3) In the multi-scale fusion module Fusion1, the mapping relationship for obtaining the high-resolution feature map is: ; where represents the identity mapping, obtain the feature map , with a size of , represents the upsampling module. Obtain the feature map , with a size of . Therefore, the size of the high-resolution feature map is . As Figure 3 shown, the upsampling module upsamples the input feature map by 2 times, specifically including two parts: (1) a convolutional layer with a kernel size of 1x1, a stride of 1, and a padding of 0, a batch normalization layer, with the width and height remaining unchanged, and the number of output channels being of the number of input channels. (2) Upsample by 2 times using the nearest neighbor interpolation method to obtain the upsampling module The output feature map has its width and height adjusted to twice that of the input feature map, and the number of channels is adjusted to of the input feature map.
[0075] c-4) The mapping relationship of the medium-resolution feature map obtained in the multi-scale fusion module Fusion1 is: . Among them, represents the identity mapping, represents the downsampling module. The medium-resolution feature map has a size of . As Figure 4 shown, the downsampling module downsamples the input feature map by a factor of 2, specifically including: a convolutional layer with a kernel size of 3x3, a stride of 2, and a padding of 0, and a batch normalization layer. The width and height of the output feature map are of the input feature map, and the number of channels is twice that of the input feature map.
[0076] d) In the second stage of the encoder, the high-resolution feature map , and the medium-resolution feature map after passing through the Transition2 module, the high-resolution feature map , the medium-resolution feature map , and the low-resolution feature map are obtained. After passing the feature maps , , through 2 StageModule2 modules, the high-resolution feature map , the medium-resolution feature map , and the low-resolution feature map are obtained.
[0077] d-1) The Transition2 module contains 3 branches: . The high-resolution feature map after being mapped through branch obtains the high-resolution feature map , with a size of ; the medium-resolution feature map after being mapped through branch obtains the medium-resolution feature map , with a size of . The medium-resolution feature map after being mapped through branch obtains the low-resolution feature map , with a size of . Branches are all identity mappings and do not contain any operations. Branch It includes a convolutional layer with a convolutional kernel size of 3x3, a stride of 2, and a padding of 1, a batch normalization layer, and a ReLU activation function.
[0078] d-2) The StageModule2 module contains two modules: the feature extraction module Extract2 and the multi-scale feature fusion module Fusion2. The feature extraction module Extract2 contains 3 branches: . The high-resolution feature map passes through branch and is mapped to obtain the high-resolution feature map , with a size of ; the medium-resolution feature map passes through branch and is mapped to obtain the medium-resolution feature map , with a size of ; the low-resolution feature map passes through branch and is mapped to obtain the low-resolution feature map , with a size of . Branches both contain 4 improved MBConv modules.
[0079] d-3) In the multi-scale fusion module Fusion2, the mapping relationship for obtaining the high-resolution feature map is: ; is the identity mapping, obtains the feature map , with a size of . is the upsampling module. Obtains the feature map , with a size of , obtains the feature map , with a size of . Therefore, the size of the high-resolution feature map is . As Figure 3 shown, the upsampling module upsamples the input feature map by 4 times, which specifically includes two parts: (1) a convolutional layer with a convolutional kernel size of 1x1, a stride of 1, and a padding of 0, a batch normalization layer, with the width and height unchanged, and the number of output channels is of the number of input channels. (2) Using the nearest neighbor interpolation method to upsample by 4 times to obtain the output feature map of the upsampling module , with the width and height being 4 times that of the input feature map and the number of channels being of the number of input channels.
[0080] d-4) The mapping relationship for obtaining the medium-resolution feature map is: ; where is the identity mapping, obtain the feature map , with a size of ; is the upsampling module, obtain the feature map , with a size of ; is the downsampling module, obtain the feature map , with a size of . Therefore, the size of the medium-resolution feature map is . As Figure 3 shown, the upsampling module upsamples the input feature map by a factor of 2, which specifically includes two parts: (1) a convolutional layer with a kernel size of 1x1, a stride of 1, and a padding of 0, a batch normalization layer, with the width and height remaining unchanged, and the number of channels being of the input number of channels. (2) Upsample by a factor of 2 using the nearest neighbor interpolation method to obtain the output feature map of the upsampling module , with the width and height being 2 times that of the input feature map, and the number of channels being of the input feature map.
[0081] d-5) The mapping relationship of obtaining the low-resolution feature map is: . Among them is the identity mapping, obtain the feature map , with a size of ; is the downsampling module, obtain the feature map , with a size of ; obtain the feature map , with a size of . Therefore, the size of the low-resolution feature map is . As Figure 4 shown, the downsampling module downsamples the input feature map by a factor of 4, which specifically includes 2 parts: (1) a convolutional layer with a kernel size of 3x3, a stride of 2, and a padding of 1, a batch normalization layer, a ReLU activation function, with the width and height adjusted to of the input feature map, and the number of channels remaining unchanged. (2) A convolutional layer with a kernel size of 3x3, a stride of 2, and a padding of 1, a batch normalization layer, to obtain the output feature map of the downsampling module , with the width and height being of the input feature map, and the number of channels being 4 times the input number of channels. As Figure 4 shown, the downsampling module Downsample the input feature map by a factor of 2, specifically including: a convolutional layer with a kernel size of 3x3, a stride of 2, and a padding of 1, followed by a batch normalization layer, to obtain the downsampling module The output feature map has a width and height that are half of those of the input feature map, and the number of channels is twice that of the input channels.
[0082] e) In the third stage of the encoder, the high-resolution feature map , the medium-resolution feature map , and the low-resolution feature map pass through the Transition3 module to obtain the high-resolution feature map , the medium-resolution feature map , the low-resolution feature map , and the small-resolution feature map . The feature maps , , , pass through the StageModule3 module to obtain the high-resolution feature map , and the medium-resolution feature map .
[0083] e-1) The Transition3 module contains 4 branches: , the high-resolution feature map passes through branch and is mapped to obtain the high-resolution feature map , with a size of ; the medium-resolution feature map passes through branch and is mapped to obtain the medium-resolution feature map , with a size of ; the low-resolution feature map passes through branch and is mapped to obtain the low-resolution feature map , with a size of ; the low-resolution feature map passes through branch and is mapped to obtain the small-resolution feature map , with a size of . Branches are all identity mappings and do not contain any operations. Branch contains a convolutional layer with a kernel size of 3x3, a stride of 2, and a padding of 1, a batch normalization layer, and a ReLU activation function.
[0084] e-2) The StageModule3 module contains 2 parts: the feature extraction module Extract3 and the multi-scale fusion module Fusion3. The feature extraction module Extract3 contains 4 branches: The high-resolution feature map passes through a branch and after mapping, the high-resolution feature map is obtained , with a size of ; The medium-resolution feature map passes through a branch and after mapping, the medium-resolution feature map is obtained , with a size of ; The low-resolution feature map passes through a branch and after mapping, the low-resolution feature map is obtained , with a size of ; The small-resolution feature map passes through a branch and after mapping, the small-resolution feature map is obtained , with a size of . The branch all contains 4 improved MBConv modules.
[0085] e-3) In the multi-scale fusion module Fusion3, the mapping relationship for obtaining the high-resolution feature map is: ; where is the identity mapping, the feature map is obtained, with a size of , is the upsampling module. The feature map is obtained, with a size of ; The feature map is obtained, with a size of ; The feature map is obtained, with a size of . So the size of the high-resolution feature map is . As Figure 3 shown, the upsampling module upsamples the input feature map by 8 times, which specifically includes two parts: (1) A convolutional layer with a kernel size of 1x1, a stride of 1, and a padding of 0, a batch normalization layer, with the width and height remaining unchanged, and the number of channels being of the input number of channels. (2) Using the nearest neighbor interpolation method to upsample by 8 times, the output feature map of the upsampling module is obtained, with the width and height being 8 times that of the input feature map, and the number of channels being of the input number of channels.
[0086] e-4) The mapping relationship for obtaining the medium-resolution feature map is: ; Among them, is the identity mapping, to obtain the feature map , with a size of . is the upsampling module, to obtain the feature map , with a size of ; to obtain the feature map , with a size of . is the downsampling module. to obtain the feature map , with a size of . Therefore, the size of the medium-resolution feature map is . As shown in Figure 3 , the upsampling module upsamples the input feature map by a factor of 4, which specifically includes two parts: (1) A convolutional layer with a kernel size of 1x1, a stride of 1, and a padding of 0, followed by a batch normalization layer. The width and height of the output feature map remain unchanged, and the number of channels is of the input number of channels. (2) Using the nearest neighbor interpolation method to upsample by a factor of 4 to obtain the output feature map of the upsampling module , with a width and height 4 times that of the input feature map, and the number of channels is of the input number of channels.
[0087] f) In the first stage of the decoder, called the fourth stage, the high-resolution feature map , the medium-resolution feature map , after passing through the multi-scale fusion module Fusion4, the high-resolution feature map , the medium-resolution feature map , and the low-resolution feature map are obtained. After passing the feature maps , , through 2 StageModule2 modules, the high-resolution feature map , the medium-resolution feature map , and the low-resolution feature map are obtained. Adding to gives the high-resolution feature map ; adding to gives the medium-resolution feature map ; adding to gives the low-resolution feature map .
[0088] f-1) In the multi-scale fusion module Fusion4, a high-resolution feature map is obtained. The mapping relationship is: ; is the identity mapping, obtain the feature map , with a size of ; is the upsampling module, obtain the feature map , with a size of . So the high-resolution feature map has a size of . Obtain the medium-resolution feature map The mapping relationship is: ; is the identity mapping, obtain the feature map , with a size of ; is the downsampling module, obtain the feature map , with a size of . So the medium-resolution feature map has a size of . Obtain the low-resolution feature map The mapping relationship is: . is the downsampling module. obtain the feature map , with a size of , obtain the feature map , with a size of . So the low-resolution feature map has a size of .
[0089] f-2) After the feature maps , , pass through 2 StageModule2 modules, a high-resolution feature map is obtained, with a size of ; a medium-resolution feature map , with a size of ; a low-resolution feature map , with a size of . , with a size of ; ; with a size of ; , with a size of .
[0090] g) In the second stage of the decoder, called the fifth stage, the high-resolution feature map , the medium-resolution feature map , and the low-resolution feature map are passed through the multi-scale fusion module Fusion5 to obtain the high-resolution feature map , the medium-resolution feature map . The feature maps and are passed through the StageModule1 module to obtain the high-resolution feature map , the medium-resolution feature map . Adding and gives the high-resolution feature map ; adding and gives the medium-resolution feature map .
[0091] g-1) In the multi-scale fusion module Fusion5, the mapping relationship for obtaining the high-resolution feature map is: ; is the identity mapping, obtains the feature map , with a size of . is the upsampling module, obtains the feature map , with a size of ; obtains the feature map , with a size of . So the size of the high-resolution feature map is . The mapping relationship for obtaining the medium-resolution feature map is: . Among them, is the identity mapping, is the upsampling module, obtains the feature map , with a size of . is the downsampling module, obtains the feature map , with a size of . So the size of the medium-resolution feature map is .
[0092] g-2) After the feature maps and pass through the StageModule1 module, the high-resolution feature map is obtained , with a size of ; medium-resolution feature map , with a size of . , with a size of ; , with a size of .
[0093] h) In the third stage of the decoder, called the sixth stage, the high-resolution feature map , the medium-resolution feature map After passing through the multi-scale fusion module Fusion6, the high-resolution feature map is obtained.
[0094] h-1) In the multi-scale fusion module Fusion6, the mapping relationship for obtaining the high-resolution feature map is: ; where is the identity mapping, obtains the feature map , with a size of . is the upsampling module, obtains the feature map . Therefore, the size of the obtained high-resolution feature map is .
[0095] i) Input the obtained high-resolution feature map into the detection head to obtain the output feature map , and each channel of the output feature map is used to represent the position of each joint point.
[0096] i-1) The detection head is a convolutional layer with a kernel size of 1x1, a stride of 1, and a padding of 0. The size of the output feature map is , where represents the number of joint points in the human body, and each channel predicts the position of a joint point.
[0097] Example 2
[0098] In this example, Figure 5The following is the human pose estimation result of the method of the present invention on the COCO validation set. We selected a total of 17 pictures, including different scenarios, different perspectives, different human postures, different crowd densities, different occlusion degrees, different light intensities, and different human sizes. From these pictures, it can be seen that the method of the present invention has achieved good human pose estimation effects under various conditions, and the method of the present invention can be applied to many aspects, such as human tracking, disease rehabilitation, human-computer interaction, video surveillance, motion analysis, etc.
[0099] Example 3
[0100] Table 1:
[0101]
[0102] In this embodiment, Table 1 shows the experimental comparison results of the method of the present invention and other mainstream algorithms on the COCO validation set. We used six mainstream algorithms for experimental comparison, and the evaluation indicators were AP (average precision) indicator, AP indicator with a threshold of 0.5, AP indicator with a threshold of 0.75, AP indicator for medium-sized objects, AP indicator for large objects, and AR (average recall rate) indicator. It can be seen that the AP indicator of the method of the present invention reached 76.7%, which is higher than the AP indicators of the mainstream algorithm frameworks SimpleBaseline, Transpose, AggPose, MIPNet, HRNet-32, and HRNet-48. Compared with AggPose, which has the highest AP among the mainstream algorithm frameworks, the AP indicator of the present invention increased by 0.3%. The method of the present invention is an improvement based on HRNet-32 (the number of channels in the high-resolution branch is 32). Compared with the original HRNet-32 network, the AP indicator increased by 0.9%. In addition, compared with the more complex and higher-precision HRNet-48 (the number of channels in the high-resolution branch is 48), the AP indicator increased by 0.4%. The above data further illustrate that the method of the present invention has better human pose estimation performance.
[0103] Although the specific implementation manners of the invention are described above in conjunction with the accompanying drawings, it is not a limitation to the protection scope of the present invention. Based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A multi-scale fusion human pose estimation method based on improved HRNet, characterized in that Including the following steps: S1. Obtain a human pose estimation dataset , denote the th image in the human pose estimation dataset . Perform data preprocessing on the dataset to obtain the preprocessed human pose estimation dataset ; S2. Construct an improved HRNet network structure, where the improved HRNet network structure includes: an initial stage, an encoder, a decoder, and a detection head. Input the preprocessed human pose estimation dataset The th image into the initial stage to obtain a high-resolution feature map S3. The encoder includes three stages: the first stage, the second stage, and the third stage. Each stage includes a Transition module and a StageModule module. Among them, the StageModule module includes a feature extraction module Extract module and a multi-scale fusion module Fusion module. The Extract module uses an improved MBConv module, and the improved MBConv module includes five consecutive parts: the first part, a convolutional layer, a batch normalization layer, and a SiLU activation function; the second part, a depthwise separable convolution, a batch normalization layer, and a SiLU activation function; the third part, an SE module; the fourth part, a convolutional layer and a batch normalization layer; the fifth part, a DropPath operation and a residual connection; Specifically, the first stage of the encoder includes the Transition1 module and the StageModule1 module; the StageModule1 module includes the feature extraction module Extract1 module and the multi-scale fusion module Fusion1 module; the second stage of the encoder includes the Transition2 module and the StageModule2 module; the feature extraction module Extract2 module and the multi-scale feature fusion module Fusion2 module; the third stage of the encoder includes the Transition3 module and the StageModule3 module; the StageModule3 module feature extraction module Extract3 module and the multi-scale feature fusion module Fusion3 module; the high-resolution feature map After passing through the encoder, a high-resolution feature map is obtained and a medium-resolution feature map ; S4. The decoder also includes three stages, which is a symmetric structure with the encoder. The first stage and the second stage both include a StageModule module and 1 independent Fusion module, and the third stage includes 1 independent Fusion module; Specifically, the first stage of the decoder includes a multi-scale fusion module, the Fusion4 module, and two StageModule2 modules; the second stage of the decoder includes a multi-scale fusion module, the Fusion5 module, and one StageModule1 module; the third stage of the decoder includes one multi-scale fusion module, the Fusion6 module; the high-resolution feature map and the medium-resolution feature map are passed through the decoder to obtain a high-resolution feature map ; S5. Obtain the output feature map from the high-resolution feature map through the detection head . Each channel of the output feature map is used to represent the position of each joint point, and finally the human pose estimation result is obtained.
2. The multi-scale fusion human pose estimation method based on improved HRNet according to claim 1, characterized in that The preprocessed human pose estimation dataset obtained in step S1 specifically includes: Unify the scaling of all sample images included in the human pose estimation dataset and then perform data augmentation strategies, including: random horizontal flipping, randomly scaling and rotating the bounding boxes of the human body, and affine transformation, to obtain the preprocessed human pose estimation dataset ; For the human pose estimation dataset after using the data augmentation strategy Generate labels for training. The specific operation is as follows: taking the position of each human joint point in the picture as the center, use a 2D Gaussian kernel function with a standard deviation of 1 pixel to generate a ground-truth heatmap as the label.
3. The multi-scale fusion human pose estimation method based on the improved HRNet according to claim 2, wherein Specifically, the first stage of the encoder described in step S3 is: The Transition1 module in the first stage of the encoder contains two branches: and , the high-resolution feature map after passing through branches and is mapped to obtain the high-resolution feature map and the medium-resolution feature map respectively; The feature extraction module, i.e., the Extract1 module, contains two branches: and , the high-resolution feature map after being mapped by branch results in the high-resolution feature map , the medium-resolution feature map after being mapped by branch results in the medium-resolution feature map , the high-resolution feature map and the low-resolution feature map after being mapped by the multi-scale fusion module, i.e., the Fusion1 module, result in the high-resolution feature map and the medium-resolution feature map , and the mapping relationship formula is: , , Among them, represents the identity mapping, represents the 2x upsampling module for obtaining the high-resolution feature map, represents the 2x downsampling module for obtaining the medium-resolution feature map.
4. The multi-scale fusion human pose estimation method based on the improved HRNet according to claim 3, characterized in that Specifically, the second stage of the encoder described in step S3 is: The Transition2 module in the second stage of the encoder contains three branches: , and . The high-resolution feature map is mapped through branch to obtain the high-resolution feature map . The medium-resolution feature map is mapped through branches and to obtain the medium-resolution feature map and the low-resolution feature map ; The feature extraction module, the Extract2 module, contains three branches: , and . The high-resolution feature map after being mapped by branch results in the high-resolution feature map . The medium-resolution feature map after being mapped by branch results in the medium-resolution feature map . The low-resolution feature map after being mapped by branch results in the low-resolution feature map . The feature maps , and after being mapped by the multi-scale fusion module, the Fusion2 module, result in the high-resolution feature map , the medium-resolution feature map and the low-resolution feature map . The mapping relationship formula is: , , , Among them, represents the identity mapping, represents a 4-fold upsampling module for obtaining a high-resolution feature map, represents a 2-fold upsampling module for obtaining a medium-resolution feature map, represents a 4-fold downsampling module for obtaining a low-resolution feature map, represents a 2-fold downsampling module for obtaining a low-resolution feature map.
5. A multi-scale fusion human pose estimation method based on improved HRNet according to claim 4, characterized in that, Specifically, the third stage of the encoder described in step S3 is: The Transition3 module in the third stage of the encoder contains four branches: , , and . After being mapped by branch , the high-resolution feature map obtains the high-resolution feature map . After being mapped by branch , the medium-resolution feature map obtains the medium-resolution feature map . After being mapped by branches and , the low-resolution feature map obtains the low-resolution feature map and the small-resolution feature map respectively; The feature extraction module Extract3 contains four branches: , , and . The high-resolution feature map is mapped through branch to obtain the high-resolution feature map . The medium-resolution feature map is mapped through branch to obtain the medium-resolution feature map . The low-resolution feature map is mapped through branch to obtain the low-resolution feature map . The small-resolution feature map is mapped through branch to obtain the small-resolution feature map . The feature maps , , and are mapped through the multi-scale fusion module Fusion3 to obtain the high-resolution feature map and the medium-resolution feature map . The mapping relationship formula is: , , Among them, represents an 8x upsampling module for obtaining a high-resolution feature map, represents a 4x upsampling module for obtaining a medium-resolution feature map.
6. The multi-scale fusion human pose estimation method based on improved HRNet according to claim 5, characterized in that Specifically, the first stage of the decoder described in step S4 is: The feature map and After being mapped by the multi-scale fusion module Fusion4, a high-resolution feature map , a medium-resolution feature map and a low-resolution feature map are obtained. The mapping relationship formula is: , , ; The feature map , and obtain high-resolution feature maps , medium-resolution feature maps and low-resolution feature maps after passing through 2 StageModule2 modules. The high-resolution feature maps and are added together through skip connections to obtain high-resolution feature maps . The medium-resolution feature maps and are added together through skip connections to obtain medium-resolution feature maps . The low-resolution feature maps and are added together through skip connections to obtain low-resolution feature maps .
7. A multi-scale fusion human pose estimation method based on improved HRNet according to claim 6, characterized in that, Specifically, the second stage of the decoder described in step S4 is: The high-resolution feature map , and After being mapped by the multi-scale fusion module Fusion5, the high-resolution feature map and the medium-resolution feature map are obtained. The mapping relationship formula is: , ; The feature map , obtains a high-resolution feature map through the StageModule1 module and a medium-resolution feature map . The high-resolution feature map and are added together through skip connections to obtain a high-resolution feature map . The medium-resolution feature map and are added together through skip connections to obtain a medium-resolution feature map .
8. A multi-scale fusion human pose estimation method based on improved HRNet according to claim 7, characterized in that, Specifically, the third stage of the decoder described in step S4 is: The feature map and are mapped by the multi-scale fusion module Fusion6 to obtain a high-resolution feature map , and the mapping relationship formula is: .
9. A multi-scale fusion human pose estimation method based on improved HRNet according to claim 8, characterized in that, The branches in the feature extraction module Extract1 and , the branches in the feature extraction module Extract2 , and and the branches in the feature extraction module Extract3 , , and , and each branch contains 4 improved MBConv modules.
10. A computing device, characterized in that, Including a memory configured to store computer-executable instructions; a processor configured to execute a multi-scale fusion human pose estimation method based on the improved HRNet as described in any one of claims 1-9 when the computer-executable instructions are executed by the processor.
Citation Information
Patent Citations
Lightweight high-resolution network human body posture estimation method
CN117671779A
Human body posture estimation method based on scale feature and hierarchical feature fusion
CN117711023A