A lightweight real-time human pose estimation method
By lightweight reconstruction of the hourglass network model and adding regression modules, the problems of complex human pose estimation model and high CPU occupation in the prior art are solved, and real-time human pose estimation and low memory usage are realized on embedded devices.
Patent Information
- Application Number
- CN202111265907.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-28
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-10-28
AI Technical Summary
In the prior art, the model structure of the human body pose estimation method is complex and has a large amount of calculation, which is difficult to meet the real-time processing requirements of embedded devices. The Heatmap prediction method requires post-processing calculation, resulting in high CPU and memory usage.
The lightweight reconstruction method of the hourglass network model Hourglass is adopted, and the lightweight codec module is built using MobilenetV2 to replace the residual module, and a lightweight regression module is added to directly predict the key points of the human body and cut off the Heatmap prediction steps.
It realizes real-time human pose estimation on embedded devices, reduces CPU and memory usage, and improves inference speed.
Smart Images

Figure CN113947784B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology. Specifically, it is a lightweight real-time human pose estimation method. Background Art
[0002] Human pose estimation is one of the important problems in the field of computer vision, aiming to estimate several key points of the human body from images. The problems existing in the human pose estimation methods in the prior art are as follows: 1. The model structure is often relatively complex, the model inference time is long, the calculation amount is large, and it is difficult to meet the requirements of real-time processing on the embedded side. Typical examples are models such as CPM and Hourglass; 2. Traditional methods basically use Heatmap for key point prediction, and CPU post-processing calculation is required, and its time-consuming cost and memory consumption are relatively too large, making it difficult to meet the CPU occupancy requirements of embedded devices such as TV terminals. Summary of the Invention
[0003] The purpose of the present invention is to provide a lightweight real-time human pose estimation method, which is used to solve the problems that the human pose estimation method in the prior art does not meet the real-time processing requirements and occupies a large amount of CPU and memory.
[0004] The present invention solves the above problems through the following technical solutions:
[0005] A lightweight real-time human pose estimation method includes:
[0006] Step S1: Load a data set and generate training data through a data augmentation method;
[0007] Step S2: Construct an original encoding and decoding module according to the Hourglass network model Hourglass stacked network, replace the basic network unit residual module of Hourglass with the inverted residual module in the lightweight network MobilenetV2 to form a lightweight encoding and decoding module; the lightweight encoding and decoding module includes a lightweight encoding module and a lightweight decoding module. The lightweight encoding module is used to extract human key point features on the image, and the lightweight decoding module is used to restore the image spatial position information for the human key points extracted by the lightweight encoding module and output a Heatmap feature map; a lightweight regression module is connected behind the output layer of the lightweight encoding module. The input of the lightweight regression module is the output result of the lightweight encoding module, and the output is the predicted coordinates of human key points;
[0008] Step S3: Train the lightweight encoding and decoding module and the lightweight regression module respectively. Specifically:
[0009] Train the lightweight encoding and decoding module. The input is the image in the training data, and the output is the Heatmap feature map. Use the mean squared error (MSE) loss function loss for prediction. Stop training when the loss decreases and stabilizes.
[0010] Freeze the parameters of the lightweight encoding and decoding module, and only train the parameters of the lightweight regression module. The input is the image in the training data, and the output is the predicted coordinates of the human key points. Use the mean squared error (MSE) loss function loss for prediction. Stop training when the loss decreases and stabilizes.
[0011] Step S4: Prune the model, remove the lightweight decoding module, and retain the lightweight regression module.
[0012] Step S5: Use the lightweight regression module to directly predict the human key points.
[0013] The present invention reconstructs the original network using a lightweight network, effectively improving the inference speed of the model and meeting the real-time human pose estimation based on embedded devices such as TVs. Compared with the traditional Heatmap prediction method that requires post-processing calculations and has high CPU and memory occupancy, the present invention uses the pruned lightweight regression model to directly predict the human key points, achieving direct regression prediction of the human key points and effectively reducing CPU and memory occupancy.
[0014] The lightweight regression module adopts the same structure as the lightweight encoding module, and adds a convolutional layer at the last layer for regression prediction of the human key points. The connection method of the intermediate layer is the same as that of the original encoding and decoding module.
[0015] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0016] (1) Aiming at the problem that the traditional model structure is complex and difficult to meet real-time processing, the present invention reconstructs the Hourglass network model using the lightweight network MobilenetV2 as the backbone network for prediction. The model is lightweight enough to perform real-time inference on embedded devices. Secondly, aiming at the problem that the Heatmap prediction method requires post-processing calculations and has high CPU and memory occupancy, a method of adding a regression module for prediction is proposed to directly predict the human key points, effectively reducing CPU and memory occupancy and meeting the requirements of practical applications.
[0017] (2) The present invention realizes a real-time human pose estimation method based on the TV embedded terminal. By reconstructing the original network using a lightweight network, the inference speed of the model is effectively improved, and real-time inference for practical applications can be satisfied. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is the structural diagram of the Hourglass network model Hourglass.
[0019] Figure 2 Schematic diagram of replacing the ResNet residual module with the MobilenetV2 inverted residual module;
[0020] Figure 3 Structure diagram of the lightweight encoding and decoding module and the lightweight regression module of the present invention;
[0021] Figure 4 Flowchart of the present invention. Detailed implementation manners
[0022] The present invention will be further described in detail below in conjunction with embodiments, but the implementation manners of the present invention are not limited thereto.
[0023] Embodiment:
[0024] Combined with the attached Figure 4 As shown, a lightweight real-time human pose estimation method includes the following steps:
[0025] Step 1, load training data sets such as COCO or MPII, and input RGB images with a size of 256*192*3.
[0026] Step 2, generate training data through data augmentation methods: the size of the Heatmap feature map is 64*48*17, and the human key point coordinates are (x, y, cls)*17, where (px, py) are the key point coordinates and cls is the key point confidence.
[0027] Step 3, the backbone network of the model draws on the Hourglass stacking network as Figure 1 As shown, construct the same encoding and decoding module structure. Replace the basic network unit ResNet residual module of the original model with the inverted residual module in the lightweight network MobilenetV2, as Figure 2 As shown, so that the number of network parameters is greatly reduced compared with the original network, forming a lightweight encoding and decoding model structure.
[0028] Step 4, the lightweight encoding and decoding model block includes a lightweight encoding module and a lightweight decoding module. The lightweight encoding module mainly extracts human key point features on the image. Its input image size is 256*192*3, and the output feature map size is 8*6*256. The lightweight decoding module restores the image spatial position information of the human key points. Its input is the output results of each layer of the lightweight encoding module C1-C6, as Figure 3 As shown, the output is a Heatmap feature map with a size of 64*48*17.
[0029] Step 5, behind the output layers of the lightweight encoding modules C1-C4, a lightweight regression module is newly added to regress and predict human key points, asFigure 3 As shown, its input is the output results of each layer of the lightweight encoding modules C1 - C4, and the output is the predicted coordinates of human key points (px, py, pcls) * 17, where (px, py) are the predicted coordinates of the key points and cls is the predicted confidence of the key points. The lightweight regression modules C1d - C4d are the same as the layers of the lightweight encoding modules C1 - C4. A new convolutional layer C8 is connected to the C4d layer for regression prediction of human key points. C1d - C4d are respectively connected to each layer of C1 - C4, and the connection method is the same as that of the original encoding - decoding network.
[0030] Step 6: First, train the lightweight encoding - decoding module. During training, the size of the input image is 256 * 192 * 3, and the size of the output Heatmap feature map is 64 * 48 * 17. The MSE (mean squared error) loss function is used. When the training loss drops and remains stable, stop training. At this time, the lightweight regression module does not participate in the training.
[0031] Step 7: After the lightweight encoding - decoding module is trained, freeze the parameters of the lightweight encoding - decoding module and only train the parameters of the lightweight regression module. The gradient does not backpropagate. During training, the size of the input image is 256 * 192 * 3, and the output is the predicted coordinates of human key points (px, py, pcls) * 17. The MSE loss function is used. When the training loss drops and remains stable, stop training.
[0032] Step 8: Finally, when deploying, trim the model structure, remove the Heatmap prediction of the lightweight decoding module, and retain the prediction of the lightweight regression module.
[0033] Step 9: Use the lightweight regression module to directly predict human key points.
[0034] Although the present invention has been described here with reference to the explanatory embodiments of the present invention, the above - mentioned embodiments are only the preferred embodiments of the present invention. The embodiments of the present invention are not limited by the above - mentioned embodiments. It should be understood that those skilled in the art can design many other modifications and embodiments, and these modifications and embodiments will fall within the scope and spirit of the principles disclosed in this application.
Claims
1. A lightweight real-time human pose estimation method, characterized in that, Including: Step S1: Load the dataset and generate training data through data augmentation methods; Step S2: Construct the original encoder-decoder module according to the Hourglass network model Hourglass stacking network. Replace the basic network unit residual module of Hourglass with the inverted residual module in the lightweight network MobilenetV2 to form a lightweight encoder-decoder module. The lightweight encoder-decoder module includes a lightweight encoding module and a lightweight decoding module. The lightweight encoding module is used to extract human keypoint features on the image, and the lightweight decoding module is used to restore the image spatial position information for the human keypoints extracted by the lightweight encoding module and output a Heatmap feature map. A lightweight regression module is connected behind the output layer of the lightweight encoding module. The input of the lightweight regression module is the output result of the lightweight encoding module, and the output is the predicted coordinates of human keypoints; Step S3: Train the lightweight encoder-decoder module and the lightweight regression module respectively, specifically: Train the lightweight encoder-decoder module. The input is the image in the training data, and the output is the Heatmap feature map. Use the mean squared error MSE loss function loss for prediction. When the loss drops and remains stable, stop training; Freeze the parameters of the lightweight encoder-decoder module and only train the parameters of the lightweight regression module. The input is the image in the training data, and the output is the predicted coordinates of human keypoints. Use the mean squared error MSE loss function loss for prediction. When the loss drops and remains stable, stop training; Step S4: Prune the model, remove the lightweight decoding module, and retain the lightweight regression module; Step S5: Use the lightweight regression module to directly predict human keypoints; The lightweight regression module adopts the same structure as the lightweight encoding module, and a new convolutional layer is added at the last layer for regression prediction of human keypoints. The connection method of the intermediate layer is the same as that of the original encoder-decoder module.
Citation Information
Patent Citations
Real-time hand posture estimation method based on MobileNet-v2
CN110188598A
System and Method for Generating Image Landmarks
US20210158023A1