A human body pose estimation method and system

By using the YOLOv5 model in human posture estimation for object detection, combined with the improved hourglass model, and adding a spatial attention mechanism, the problem of difficult human posture estimation in forward-view sonar images is solved, and accurate human posture prediction for low-resolution, low signal-to-noise ratio images are achieved.

CN114842506BActive Publication Date: 2025-06-27INST OF ACOUSTICS CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210412076.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-19
Publication Date
2025-06-27
Estimated Expiration
2042-04-19

AI Technical Summary

Technical Problem

The prior art is difficult to accurately estimate human posture from low resolution, low signal-to-noise ratio forward-looking sonar images.

Method used

The YOLOv5 model is used to detect human objects, obtain images of the area where the human object is located, and input them into the improved hourglass model. The improved hourglass model adds a spatial attention mechanism before the four upsampling stages and before the thermal map prediction stages to optimize the feature extraction and prediction process.

Benefits of technology

Through the improved hourglass model, the key points of the human body can be accurately predicted from the forward-looking sonar images with low resolution and low signal-to-noise ratio, improving the accuracy of human body posture estimation, and has important application value in underwater target identification and rescue tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114842506B_ABST
    Figure CN114842506B_ABST
Patent Text Reader

Abstract

The present invention relates to a human body pose estimation method and system, including: S1. Preprocessing and labeling a forward-looking sonar image set to obtain a human body pose estimation data set, and dividing the human body pose estimation data set into a training set and a test set; S2. Using the training set to train a YOLOv5 model to obtain an image of the region where the human target is located; S3. Using the training set to train an improved hourglass model, inputting the image of the region where the human target is located into the trained improved hourglass model for processing to obtain predicted human body pose information. The improved hourglass model of the present invention is provided with a spatial attention mechanism, which can automatically learn to highlight important features and suppress background noise, so as to accurately predict human body key points from a forward-looking sonar image set with relatively low resolution and signal-to-noise ratio, obtain the human body pose, and has great application value for underwater target discrimination and salvage tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and particularly relates to a human pose estimation method and system. Background Art

[0002] Human pose estimation refers to a method of automatically analyzing the main parts and pose information of the human body from a given image through image processing technology. Human pose estimation based on forward-looking sonar images is of great significance for improving the autonomous signal processing ability of sonar, and can be applied to fields such as target rapid confirmation and rescue.

[0003] Since AlexNet won the championship in the ImageNet image classification task in 2012, deep learning methods have achieved breakthroughs in more and more computer vision tasks, such as image segmentation and object detection. DeepPose proposed in 2014 is one of the earliest models to apply convolutional neural networks to human pose estimation tasks. Most common human pose estimation technologies are mostly for optical images. Due to the high signal-to-noise ratio, automatic analysis and discrimination of multiple fine parts and complex postures of the human body can be achieved.

[0004] Forward-looking sonar images cannot reach the quality of optical images in terms of resolution, signal-to-noise ratio, etc., and it is difficult to perform human pose estimation. Summary of the Invention

[0005] In order to overcome the above-mentioned defects existing in the prior art, the present invention provides a human pose estimation method and system for solving the above-mentioned problems existing in the prior art.

[0006] A human pose estimation method for predicting human poses from forward-looking sonar images, the method comprising the following steps:

[0007] S1. Preprocess and label a forward-looking sonar image set to obtain a human pose estimation data set, and divide the human pose estimation data set into a training set and a test set;

[0008] S2. Use the training set to train the YOLOv5 model to obtain an image of the region where the human target is located;

[0009] S3. Use the training set to train an improved hourglass model, and input the image of the region where the human target is located into the trained improved hourglass model for processing to obtain predicted human pose information.

[0010] For the aspects and any possible implementation manners described above, a further implementation manner is provided. The human body pose data set includes first label information and second label information obtained by manually labeling all forward-looking sonar image sets. Among them, the first label information is the bounding box position of the human body target in the forward-looking sonar image set; the second label information is the positions of four key points including the head, torso, left leg, and right leg of the human body target in the forward-looking sonar image set.

[0011] For the aspects and any possible implementation manners described above, a further implementation manner is provided. Specifically, S2 is as follows: Use the forward-looking sonar images in the training set and the first label information to train the YOLOv5 model until convergence. After convergence, the YOLOv5 model detects the human body target from the forward-looking sonar image and crops out the region where the human body target is located to obtain the image of the region where the human body target is located.

[0012] For the aspects and any possible implementation manners described above, a further implementation manner is provided. Specifically, S3 is as follows:

[0013] S31. Use the first label information to obtain the image of the region where the human body target is located in the training data, input the image of the region where the human body target is located into the improved hourglass model for processing, and calculate the positions of the four key points of the human body target.

[0014] S32. Compare the positions of the four key points of the human body target calculated each time with the second label information, calculate the mean square error between the two as the input of the loss function, optimize the improved hourglass model, and use the positions of the four key points of the human body target when the output reaches the set requirements as the predicted positions of the four key points of the human body target.

[0015] For the aspects and any possible implementation manners described above, a further implementation manner is provided. Specifically, S31 includes:

[0016] S311. The improved hourglass model extracts features from the image of the region where the human body target is located, extracts feature maps before the four upsampling stages and before the heat map prediction stage of the model, uses the extracted feature maps as variables, and calculates the spatial attention map through the attention module.

[0017] S312. Multiply the spatial attention map by the extracted feature map to obtain the corrected feature map.

[0018] S313. Use the attention module as a part of the improved hourglass model and participate in the training together until convergence.

[0019] For the aspects and any possible implementation manners described above, a further implementation manner is provided. The feature maps extracted before the four upsampling stages and in the heat map prediction stage are F n, If n ∈ {1, 2, 3, 4, 5}, then the spatial attention map is calculated by the following formula: Mask(F n ) = sigmoid(Conv(AvgP(F n ))), where sigmoid is the activation function, Conv is the convolution function, and AvgP is the average pooling operation.

[0020] For the aspects and any possible implementation manners described above, a further implementation manner is provided. The modified feature map is obtained by the following formula: where Mask(F n ) is the spatial attention map, F n is the feature map, is the modified feature map, is the multiplication operation.

[0021] For the aspects and any possible implementation manners described above, a further implementation manner is provided. The loss function model is: where x is the input forward-looking sonar image containing the human target, are the positions of the four key points of the human target calculated each time, is the second label information in the training set, and θ are all the trainable parameters in the improved hourglass model.

[0022] The present invention also provides a human pose estimation system, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method of the present invention.

[0023] The present invention also provides a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the method of the present invention.

[0024] Beneficial technical effects

[0025] Before the four upsampling stages and before the heatmap prediction stage of the existing hourglass model, the present invention adds a spatial attention mechanism to optimize and adjust the feature map output by the residual module of this layer, obtaining an improved hourglass model. The improved hourglass model in the present invention can automatically learn to highlight the important features in the image and suppress background noise, so as to accurately predict human key points from the forward-looking sonar images with relatively low resolution and signal-to-noise ratio, estimate the human pose, improve the detection accuracy, accurately predict the human pose from the forward-looking sonar image set, which is of great significance for underwater target discrimination and salvage tasks and has great application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a schematic flowchart of the method in the embodiment of the present invention;

[0027] Figure 2 It is a schematic diagram of spatial attention in the embodiment of the present invention;

[0028] Figure 3 It is a schematic diagram of the residual module used in the embodiment of the present invention;

[0029] Figure 4 It is a schematic diagram of the relationship between the hourglass model and spatial attention in the embodiment of the present invention;

[0030] Figure 5 It is a schematic diagram of the estimation effect in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] To better understand the technical solution of the present invention, the content of the present invention includes but is not limited to the following specific embodiments, and similar technologies and methods should be regarded as within the scope of protection of the present invention. To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0032] It should be clear that the embodiments described in the present invention are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts belong to the scope of protection of the present invention.

[0033] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0034] As Figure 1 shown is the human pose estimation method of the present invention, and the method includes the following steps:

[0035] S1. Preprocess and label the forward-looking sonar image set to obtain a human pose estimation data set, and divide the human pose estimation data set into a training set and a test set;

[0036] S2. Use the training set to train the YOLOv5 model to obtain a human target, and use it to obtain the image of the area where the human target is located;

[0037] S3. Use the training set to train the improved hourglass model, and input the image of the area where the human target is located into the trained improved hourglass model for processing to obtain the predicted human pose information.

[0038] Preferably, the human pose estimation method of the present invention may also include the following steps:

[0039] S1’. Take the forward-looking sonar image as input, input it into the trained YOLOv5 model, output the bounding box of the area where the human target is located, and crop out the image of the area where the human target is located according to the bounding box.

[0040] S2’. Take the bounding box of the area where the human target is located as input, input it into the trained improved hourglass model, output the four key points of the human target, and obtain the predicted human pose.

[0041] The specific process of the present invention is as follows:

[0042] (1). Preprocess the forward-looking sonar image set and manually label the tag information. The obtained tag information is divided into two parts: the first tag information is the position of the bounding box of the human target in the forward-looking sonar image, and the second tag information is the positions of the four key points of the human target in the forward-looking sonar image: head, torso, left leg, and right leg. Divide the labeled data set into a training set and a test set. Both the test set and the training set include the first and second tag information to obtain the forward-looking sonar human pose estimation data set;

[0043] (2). Use the forward-looking sonar image and the first tag information in the training set in step (1) to train the YOLOv5 model until convergence. The trained YOLOv5 model can detect the human target from the input forward-looking sonar image set, and crop out the image of the area where the human target is located for the next operation;

[0044] (3) Input the image of the region where the human target is located obtained by cropping into the improved hourglass model for processing. After normalizing the image of the region where the human target is located obtained by cropping, input it into the improved hourglass model. First, it is processed by the initialization layer. The initialization layer consists of two convolutional layers. When performing convolutional processing, first, a convolutional layer is used to extract a feature map of 192x192x64, and then another convolutional layer is used to obtain a feature map of 192x192x128. The size of the convolutional kernel is 1 in both cases. The purpose is to perform initialization processing on the image of the region where the human target is located obtained by cropping to obtain an initial feature map. Then, the initial feature map is sent to the feature extraction module. The feature extraction module in the present invention is a residual module, and its structure is as Figure 3 shown. The number of input and output channels of the residual module is 128, and the size of the feature map input into it remains unchanged. The initial feature map undergoes four processes: after the residual module extracts features - max pooling for downsampling, its size is reduced from 192x192x128 to 12x12x128, that is, downsampling processing is performed on the initial feature map. By continuously performing downsampling processing, high-level abstract features of the input image can be extracted. Correspondingly, before each use of max pooling for downsampling, the feature map output by the current layer of the residual module is added to the feature map of the corresponding resolution layer during the upsampling process through the skip connection branch, so as to fuse high-resolution features. During the upsampling process of the downsampled feature map, it undergoes four more processes: bilinear interpolation - fusing the skip connection feature map - the operation of the residual module extracting features, and its size is restored from 12x12x128 to 192x192x128. The feature map with the restored size is input into a convolutional layer including four convolutional kernels, and finally, each convolutional kernel outputs a heat map respectively, that is, the convolutional layer outputs a total of four heat maps. The input of the prediction convolutional layer is the upsampled feature map with a size of 192x192x128, and the output is four heat maps with a size of 192x192x4. Calculate the coordinates of the point with the maximum pixel value in each heat map, and thus obtain the four different key point positions of the corresponding human target, thereby obtaining the predicted human pose information.

[0045] As Figure 4 shown, for the improved hourglass model of the present invention, the entire step process is as follows:

[0046] (1) Spatial attention mechanism modules are set before each of the four upsampling stages and before the heat map prediction stage. As Figure 2 shown, before each of the four upsampling stages and before the heat map prediction stage of the improved hourglass model, features are extracted by the residual module. Taking the feature map output by the residual module of this layer as a variable, in the spatial attention mechanism module, the above variable undergoes average pooling, convolution, and sigmoid activation function calculation to obtain a spatial attention map;

[0047] Let the feature maps obtained before the four upsampling stages and before the heatmap prediction stage be F n, n ∈ {1, 2, 3, 4, 5}. N = 1, 2, 3, 4 are before the four upsamplings, and n = 5 is before predicting the heatmap

[0048] Mask(F n ) = sigmoid(Conv(AvgP(F n ))), n ∈ {1, 2, 3, 4, 5} (1)

[0049] In the above calculation formula (1), sigmoid is an activation function, and its expression is: Its role is to perform a non - linear transformation on the feature map. Conv is a convolution operation, and the output is a feature map of HxWx1. AvgP is an average pooling operation. F n is the input feature map, with a size of HxWx128. The purpose of this step is to suppress the background noise in the feature map F n and make the key features more prominent.

[0050] Step (2): Multiply the spatial attention map obtained in step (1) with the feature map F output by the residual module n element - by - element to obtain the corrected feature map. The calculation process is as Figure 2 shown, and the calculation formula is as follows:

[0051]

[0052] In the above calculation formula (2), Mask(F n ) is the spatial attention map calculated in step (1). It is a set of two - dimensional weights used to highlight important features and suppress background noise. F n is the feature map output by the residual module in the last stage of the downsampling of the hourglass network and the four upsampling stages. is the adjusted feature map. is the element - by - element multiplication operation.

[0053] (3) Obtain the finally adjusted feature map from the calculation in step (2). Finally, four heatmaps are output by the prediction convolutional layer. The input of the prediction convolutional layer is the adjusted feature map with a size of 192x192x128, and the output is a heatmap with a size of 192x192x4, corresponding to the four predicted key points. The coordinates of the pixel with the maximum value in the heatmap are the coordinates of the corresponding key point in the input forward - looking sonar image.

[0054] Further, in step (3), the mean square error between the four predicted heatmaps corresponding to the predictions output by the improved hourglass model and the second label information manually marked in the training set is calculated as the loss function, and all trainable parameters in the improved hourglass model are updated through the backpropagation algorithm until the error between the output of the improved hourglass model and the second label information converges to a certain error range.

[0055] The loss function of the improved hourglass model is as follows:

[0056]

[0057] Wherein, is the improved hourglass model, x is the forward-looking sonar image containing the human target as input, is the predicted heatmap output by the improved hourglass model, is the manually marked heatmap, from the training set of the human pose estimation dataset of the aforementioned forward-looking sonar image, and θ are all trainable parameters in the improved hourglass model.

[0058] Specifically, in the embodiment of the present invention, the forward-looking sonar data needs to be preprocessed first. This step mainly maps the signal numerical values to grayscale values for the forward-looking sonar data, and uses mean filtering and erosion for processing to obtain a clearer forward-looking sonar image. Then, the Labelme software is used to make a human pose estimation dataset based on the forward-looking sonar image, that is, the position of the target box and the position of the target key points in the forward-looking sonar image are manually marked. There are at least 20 forward-looking sonar images in the dataset, which are expanded to 100 by random translation. Each image is manually marked with label information. The labels are divided into two parts: the first is the position of the bounding box of the human target in the forward-looking sonar image, and the second is the positions of the four key points of the human target in the forward-looking sonar image: head, torso, left leg, and right leg. The marked dataset is divided into a training set and a test set. The training set will be used to train the YOLOv5 model and the improved hourglass model of the present invention, and the test set will be used to evaluate the effect of the present invention.

[0059] Preferably, in the embodiments of the present invention, human target detection is a prerequisite for pose estimation. Human target detection means detecting human targets from the forward-looking sonar image, cropping out the human target region image, and handing it over to the improved hourglass model proposed by the present invention for further processing to obtain the human pose. The present invention uses the YOLOv5 model as the target detector, the input of which is the forward-looking sonar image, and the expected output is the bounding boxes of all human targets in the image. Using the forward-looking sonar images in the aforementioned forward-looking sonar image human pose estimation dataset and the bounding box positions of human targets in the corresponding forward-looking sonar images as the input and the human target detection standard, the YOLOv5 model is trained. After 1000 iterations of training, the recall rate and accuracy rate of the YOLOv5 model for detecting human targets in the forward-looking sonar image reach more than 90%, meeting the detection requirements, and the YOLOv5 model can accurately detect human targets from the forward-looking sonar image.

[0060] Preferably, in the embodiments of the present invention, when training the improved hourglass model, the initial learning rate is set to 0.000001, the Adam algorithm is selected as the optimizer, the training batch size is set to 10 images. After 8000 iterations of the improved hourglass model, the loss function value drops from about 0.5 to about 1e-05, indicating that the YOLOv5 model can converge.

[0061] Preferably, the present invention further includes step S4: evaluating and verifying the predicted human pose information obtained in S3 with the second label information in the test set, and based on the following evaluation indicators: The evaluation indicator used is the percentage of correctly predicted key points, that is, if the predicted human pose information, the Euclidean distance between the key points of the predicted heat map and the corresponding key points in the second label information manually marked in the test set is less than a preset threshold, then it is considered that the key points in the predicted heat map are successfully predicted and marked as 1; otherwise, the prediction fails and is marked as 0, and finally the proportion of key points marked as 1 for successful prediction is counted. The evaluation indicator consists of five parts: the percentage of all key points predicted correctly, the percentage of correctly predicted key points in the head, torso, left leg, right leg, etc. The thresholds are 20% and 30% of the torso length of the human target manually marked in the human pose estimation dataset, denoted as PCK@0.2 and PCK@0.3 respectively. Table 1 shows the results of the original hourglass model and the improved hourglass model of the present invention on the same test set. It can be seen from Table 1 that the proportion of successful predictions of the five key points of the whole body, head, torso, left leg, and right leg is higher than that of the prior art, and there is an improvement compared with the prior art under the two indicators.

[0062] This shows that the method of the present invention has high reliability due to the addition of the spatial attention mechanism module in the improved hourglass module. Therefore, it can quickly and accurately predict the human pose from the forward-looking sonar image. Table 1

[0063]

[0064] As shown by this Figure 5 the predicted human body posture is very similar to the human body posture in the target area image, indicating that the method of the present invention is very effective.

[0065] Preferably, the present invention further provides a human body posture estimation system, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method of the present invention.

[0066] Preferably, the present invention further provides a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the method of the present invention.

[0067] The above description shows and describes several preferred embodiments of the present invention. However, as mentioned above, it should be understood that the present invention is not limited to the form disclosed herein, should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications, and environments, and can be changed within the scope of the application concept described herein through the above teachings or the technology or knowledge in the relevant field. And any changes and variations made by those skilled in the art without departing from the spirit and scope of the present invention shall fall within the protection scope of the appended claims of the present invention.

Claims

1. A human body pose estimation method for predicting the human body pose from a forward-looking sonar image, characterized in that, The method includes the following steps: S1. Preprocess and label the forward-looking sonar image set to obtain a human pose estimation data set, and divide the human pose estimation data set into a training set and a test set; The human pose data set includes first label information and second label information obtained by manually labeling all forward-looking sonar image sets. Among them, the first label information is the bounding box position of the human target in the forward-looking sonar image set; the second label information is the positions of four key points including the head, torso, left leg, and right leg of the human target in the forward-looking sonar image set; S2. Use the training set to train the YOLOv5 model to obtain the image of the area where the human target is located; S3. Use the training set to train the improved hourglass model, and input the image of the area where the human target is located into the trained improved hourglass model for processing to obtain the predicted human pose information. Specifically: S31. Use the first label information to obtain the image of the area where the human target is located in the training data, input the image of the area where the human target is located into the improved hourglass model for processing, and calculate the positions of the four key points of the human target. S31 specifically includes: S311. The improved hourglass model extracts features from the image of the region where the human target is located. Feature maps are extracted before the four upsampling stages and before the heatmap prediction stage of the improved hourglass model. The extracted feature maps are used as variables, and after being calculated by the attention module, a spatial attention map is obtained. The extracted feature map is F n, If n ∈ {1, 2, 3, 4, 5}, then the spatial attention map is calculated by the following formula: Mask(F n ) = sigmoid(Conv(AvgP(F n ))), where sigmoid is the activation function, Conv is the convolution function, and AvgP is the average pooling operation; S312. Multiply the spatial attention map with the extracted feature map to obtain a corrected feature map, and the modified feature map is obtained by the following formula: where Mask(F n ) is the spatial attention map, F n is the feature map, is the modified feature map, is the multiplication operation; S313. Use the attention module as a part of the improved hourglass model and participate in the training together until convergence; S32. Compare the positions of the four key points of the human target obtained each time with the second label information, calculate their mean square error, and use it as the input of the loss function to optimize the improved hourglass model. When the output meets the set requirements, use the positions of the four key points of the human target as the predicted positions of the four key points of the human target. The loss function is as follows: where x is the forward-looking sonar image containing the human target as the input, is the positions of the four key points of the human target obtained each time, is the second label information in the training set, and θ are all trainable parameters in the improved hourglass model.

2. The human body posture estimation method according to claim 1, wherein Specifically, S2 is: Use the forward-looking sonar images and the first label information in the training set to train the YOLOv5 model until convergence. The converged YOLOv5 model detects the human target from the forward-looking sonar image and crops out the area where the human target is located to obtain the image of the area where the human target is located.

3. A human body pose estimation system, characterized in that, It includes a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor. The processor executes the computer-executable instructions to implement the method according to any one of claims 1-2.

4. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions prompt the processor to implement the method according to any one of claims 1-2.

Citation Information

Patent Citations

  • Single / multi-target key point positioning method based on stacked hourglass network

    CN111913435A

  • Method and device for intelligent estimation of human body movement posture based on convolutional neural network

    WO2022036777A1