Lightweight Human Pose Estimation Method and Device Based on Knowledge Distillation
Through knowledge distillation technology and rank-guided filter pruning, a lightweight student model is built, which solves the real-time reasoning problem of deep learning models on resource-limited devices, and realizes efficient and accurate human posture estimation, reducing the risk of sports injury.
Patent Information
- Application Number
- CN202510437685.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-09
AI Technical Summary
The existing deep learning models are huge in human pose estimation and have high computational cost, making it difficult to realize real-time inference on devices with limited resources, and misjudgment or delay in complex motion scenarios, making it difficult to meet the needs of efficient and accurate.
Using knowledge distillation technology, the student model is guided to perform feature distillation through the pruned teacher model intermediate feature map, combined with rank-guided filter pruning and feature correlation reconstruction, a lightweight student model is constructed to ensure the efficiency and accuracy of the pose estimation model.
While reducing model parameters, maintaining high accuracy, real-time posture estimation on edge devices is achieved, providing fast and accurate motion feedback, and reducing the risk of motion damage.
Smart Images

Figure CN119942655B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human pose estimation, and particularly to a lightweight human pose estimation method and device based on knowledge distillation. Background Art
[0002] With the popularization of the concept of national fitness, more and more people participate in daily exercises such as running, fitness, and strength training. Standard exercise movements are crucial for training effects and physical health. Irregular exercise postures may not only lead to poor training effects but also cause sports injuries. Therefore, how to analyze and provide real-time feedback on movements during daily exercises has become a widely concerned technical requirement.
[0003] Traditional exercise guidance relies on on-site supervision by coaches or video courses, but these methods cannot meet the needs in personalized, real-time, and large-scale scenarios. For example, in a gym, it is difficult for a coach to pay attention to each trainee at the same time; during home fitness or remote training, the lack of real-time feedback will also affect the training effect.
[0004] Pose estimation in artificial intelligence technology provides an effective solution to this problem. By capturing the coordinates of human key points and analyzing the differences between the user's movements and standard movements, the pose estimation algorithm can provide real-time feedback for users. However, existing deep learning models are usually based on large-scale neural networks, with large model volumes and high computational costs, making it difficult to achieve real-time inference on devices with limited resources (such as edge devices or mobile terminals). In addition, in complex motion scenarios, traditional models often have misjudgments or delays, making it difficult to meet the requirements of high efficiency and accuracy.
[0005] Knowledge distillation technology is an effective model compression method. By guiding a lightweight student model to learn with the help of a teacher model, a balance between accuracy and efficiency is achieved. Based on high-precision predictions, the teacher model transfers knowledge to the student model, which can maintain a high prediction accuracy while simplifying the network structure by learning the output of the teacher model. Summary of the Invention
[0006] To address the above problems, the present invention proposes a lightweight human pose estimation method and device based on knowledge distillation. By introducing knowledge distillation technology and using the intermediate feature maps of the pruned teacher model to guide the student model for feature distillation, while significantly reducing the model parameters, the efficiency and accuracy of the pose estimation model are ensured.
[0007] On the one hand, the lightweight human pose estimation method based on knowledge distillation has the following specific steps:
[0008] S1, construct an initial human video dataset and label the dataset;
[0009] S2. Construct a teacher model and a student model for human pose estimation;
[0010] S3. Use the labeled dataset to train the teacher model, and use the trained teacher model to guide the training of the student model to obtain a trained student model; the guided training includes a feature distillation stage and a logic distillation stage; the feature distillation is based on the intermediate feature maps of the student model and the pruned intermediate feature maps of the teacher model; the logic distillation is based on the logical outputs of the teacher model and the student model;
[0011] S4. Obtain human video data, and input the human video data into the trained student model for human pose estimation.
[0012] Preferably, the process of obtaining the pruned intermediate feature maps of the teacher model is as follows:
[0013] Slice the intermediate feature maps of the same batch in the teacher model by channel, calculate the average value of the ranks of all intermediate feature maps within each channel, and use the average value to represent the rank value of the feature map of this channel;
[0014] Sort the rank values of each layer of feature maps, and divide the intermediate feature maps into a retention set and a pruning set according to the sorting result of the rank values. The retention set stores the high-rank feature maps, and the pruning set stores the low-rank feature maps; the retention set is the pruned intermediate feature maps of the teacher model.
[0015] Preferably, the loss function in the feature distillation stage is the channel correlation information loss, which is specifically as follows:
[0016] Unfold the pruned intermediate feature maps of the teacher model to obtain the teacher model correlation matrix; unfold the intermediate feature maps of the student model to generate the student model correlation matrix;
[0017] Construct an L2 constraint based on the teacher model correlation matrix and the student model correlation matrix to obtain the channel correlation information loss; the channel correlation information loss is expressed as:
[0018] ;
[0019] Where, represents the channel correlation information loss; represents the student model correlation matrix after adjusting the dimension; represents the correlation matrix of the teacher model; represents the intermediate feature maps of the student model; represents the intermediate feature maps after pruning of the teacher model; represents the number of channels of the teacher model; represents the square of the L2 distance; is used to adjust the feature dimension of the student model to match the number of channels of the teacher model's features; Represents the ICC matrix.
[0020] Preferably, the logical outputs of the teacher model and the student model use the simcc algorithm for pose keypoint prediction, and the human pose estimation task is decomposed into independent classification tasks for horizontal and vertical coordinates.
[0021] Preferably, the loss function in the logical distillation stage is the logical distillation loss; the logical distillation loss combines the original loss, the distribution loss, and the soft loss; the distribution loss enables the student model to learn the non-target knowledge of the teacher model; the soft loss enables the student model to learn the target knowledge of the teacher model; the original loss is the cross-entropy loss of the student model.
[0022] Preferably, the logical distillation loss is expressed as:
[0023] ;
[0024] where represents the logical distillation loss; represents the hyperparameter for balancing the loss; N represents the number of samples in a batch; K represents the total number of keypoints; L represents the length of the localization region in the horizontal coordinate x or vertical coordinate y direction; represents the non-target cell score of the teacher model; represents the non-target cell score of the student model; represents the predicted score of the teacher model for the target cell; represents the predicted score of the student model for the target cell; t represents the target cell; i represents the i-th cell.
[0025] Preferably, the distribution loss is expressed as:
[0026] ;
[0027] ;
[0028] ;
[0029] where represents the distribution loss; N represents the number of samples in a batch; K represents the total number of keypoints; L represents the length of the localization region in the horizontal coordinate x or vertical coordinate y direction; t represents the target cell; represents the non-target cell score of the teacher model; represents the non-target cell score of the student model; i represents the i-th cell; represents the predicted score of the teacher model for the i-th cell; represents the predicted score of the teacher model for the target cell; Denote the predicted score of the student model for the \(i\)-th cell; Denote the predicted score of the student model for the target cell.
[0030] Preferably, the soft loss is expressed as:
[0031] ;
[0032] where, Denote the soft loss; Denote the predicted score of the teacher model for the target cell; Denote the predicted score of the student model for the target cell; \(t\) represents the target cell.
[0033] Preferably, the total loss function of the student model is expressed as:
[0034] ;
[0035] where, Denote the total loss function; Denote the loss of channel correlation information; Denote the logical distillation loss.
[0036] On the other hand, a lightweight human pose estimation device based on knowledge distillation includes the following:
[0037] A dataset construction and labeling module, which is used to construct an initial human video dataset and label the dataset;
[0038] A model construction module, which is used to construct a teacher model and a student model for human pose estimation;
[0039] A distillation training module, which uses the labeled dataset to train the teacher model, and uses the trained teacher model to guide the training of the student model to obtain a trained student model; the guided training includes a feature distillation stage and a logical distillation stage; the feature distillation is based on the intermediate feature maps of the student model and the pruned intermediate feature maps of the teacher model; the logical distillation is based on the logical outputs of the teacher model and the student model;
[0040] The process of obtaining the pruned intermediate feature maps of the teacher model is specifically as follows:
[0041] Slice the intermediate feature maps of the same batch in the teacher model by channel, calculate the average value of the ranks of all intermediate feature maps within each channel, and use the average value to represent the rank value of the channel feature map;
[0042] Sort the rank values of each layer of feature maps, and divide the intermediate feature maps into a retention set and a pruning set according to the sorting result of the rank values. The retention set stores the high-rank feature maps, and the pruning set stores the low-rank feature maps. The retention set is the intermediate feature map of the pruned teacher model.
[0043] The human pose estimation module is used to obtain human video data and input the human video data into the trained student model for human pose estimation.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] (1) By applying rank-guided filter pruning and feature correlation reconstruction, the present invention prunes the intermediate feature maps of the teacher model. While significantly reducing the model parameters, it ensures the efficiency and accuracy of the pose estimation model in the lightweight human pose estimation model based on knowledge distillation, further improving the real-time inference ability of the model on edge devices, and providing rock-solid technical support for the general action analysis system relying on this framework in daily motion scenarios.
[0046] (2) The present invention introduces an efficient backbone network architecture, EfficientNet, for extracting feature representations of multiple key points. EfficientNet achieves a good balance between speed and accuracy, and at the same time has low computational resource requirements, making it suitable for edge device deployment.
[0047] (3) The lightweight student model of the present invention can be deployed on edge devices to process image data in real time and quickly, accurately identify the coordinates of human key points, and carefully compare them with a preset standard library to immediately output the action matching degree and feedback, ultimately achieving fast and low-power pose estimation on edge devices. During the exercise process, users can obtain accurate real-time action feedback and practical suggestions, thereby effectively optimizing the exercise effect and significantly reducing the risk of sports injuries. Brief Description of the Drawings
[0048] The following further describes the present invention in detail with reference to the drawings;
[0049] Figure 1 It is a flowchart of the lightweight human pose estimation method based on knowledge distillation according to an embodiment of the present invention;
[0050] Figure 2 It is a network structure diagram of the lightweight human pose estimation method based on knowledge distillation according to an embodiment of the present invention;
[0051] Figure 3 It is a schematic flow diagram of the lightweight human pose estimation method based on knowledge distillation according to an embodiment of the present invention;
[0052] Figure 4Block diagram of the lightweight human pose estimation device based on knowledge distillation according to an embodiment of the present invention. Detailed implementation manners
[0053] The present invention will be further described below through specific implementation manners.
[0054] As Figure 1 shown, the lightweight human pose estimation method based on knowledge distillation is as follows:
[0055] S1. Construct an initial human video dataset and label the dataset.
[0056] Use a high-definition camera to collect video data covering various daily motion actions (such as running, push-ups, squats, etc.), and extract frames from each video to obtain an initial dataset. These data can cover common motion scenarios, ensuring action diversity and generalization.
[0057] Perform key point annotation on the daily motion actions in the dataset, and record the joint point coordinates (such as shoulders, knees, ankles, etc.) of each pose. The generated annotation data provides a high-quality supervision signal for model training, ensuring the accuracy of action analysis.
[0058] S2. Construct a teacher model and a student model for human pose estimation.
[0059] Introduce the efficient backbone network architecture EfficientNet to extract feature representations of multiple key points. EfficientNet achieves a good balance between speed and accuracy, and at the same time has low computational resource requirements, which is suitable for edge device deployment. In this embodiment, both the teacher and the student are EfficientNet models. The teacher model is EfficientNet-L, and the student model is EfficientNet-M. The size of the student model is smaller than that of the teacher model.
[0060] S3. Use the labeled dataset to train the teacher model, and use the trained teacher model to guide the training of the student model to obtain a trained student model; the guided training includes a feature distillation stage and a logic distillation stage; the feature distillation is based on the intermediate feature maps of the student model and the pruned intermediate feature maps of the teacher model; the logic distillation is based on the logical outputs of the teacher model and the student model.
[0061] Use the dataset with complete annotations to train the teacher model, and accurately learn the complex action features through an efficient convolutional network structure. Then, using knowledge distillation technology, the teacher model guides the lightweight student model to train, enabling the student model to significantly reduce the number of parameters and computational requirements while maintaining high accuracy.
[0062] SeeFigure 2 As shown in Figure 2 , the distillation of this embodiment includes two stages.
[0063] In the first-stage feature distillation, this embodiment proposes a rank-guided filter pruning method to provide learning assistance for the student model by retaining and extracting the guiding information in the intermediate feature maps of the teacher model. On this basis, this embodiment also adopts a method of reconstructing feature correlation to make up for the common loss of feature correlation information in the filter pruning process.
[0064] The pruning method is as follows:
[0065] For the rank-guided filter pruning method, in a convolutional neural network, reasonably selecting the batch size not only helps to fully utilize the memory capacity but also improves the running efficiency of the network. The resulting feature map data for each layer is stored in the form of a four-dimensional matrix , where N represents the batch number, C is the number of channels of the feature map, and W and H represent the width and height of the feature map respectively. Analyzing the convolution process of the feature map, it can be seen that for the feature maps of the same batch, the convolution calculations of the filter at each position are the same. Since the average rank of the feature maps generated by the same filter remains relatively stable, the feature maps in the same batch can be sliced by channel, and the average value of the ranks of all feature maps within each channel is calculated, and this mean value is used to represent the rank value of the feature map of this channel. The information metric is defined as follows:
[0066] ;
[0067] where represents the feature map, represents the channel quantity of the th feature map, is the set of ranks of the th layer of feature maps. Through the information metric function we obtain a one-dimensional vector , and each value represents the mean value of the corresponding feature map of each channel. The calculation result of the rank of the th layer of feature maps is expressed as ( represents the value of the rank of the 0th channel at the index position, represents the value of the rank of the th channel at the index position).
[0068] Sorting the calculated rank values of the feature maps, an importance sorting index based on channels can be obtained. Through the high-rank index values, the feature maps with higher importance can be identified, and the original feature maps are divided into two subsets based on these high-rank indexes, expressed as:
[0069]
[0070]
[0071] Reserved set Store high-rank feature maps, cropping set Store low-rank feature maps, and the above two sets constitute the original feature maps of the layer, and there is no intersection between the two sets. In this way, the reserved feature information set can be screened out.
[0072] The method for reconstructing feature correlation is as follows:
[0073] To address the possible loss of associated information in the refined feature maps, in this embodiment, the feature maps of the teacher and student models are respectively unfolded, and their respective correlation matrices are generated to reconstruct and transmit the associated information between the refined feature maps, thereby ensuring that the guiding role of the intermediate feature maps of the teacher model can be fully exerted.
[0074] Specifically, let the features of the teacher model and the student model be represented by and respectively. In practical applications, and can be any differentiable functions, and they are parameterized into a convolutional neural network (CNN). The embedding of the teacher model is represented as , where represents the number of output channels, and represent the height and width of the feature map respectively. Similarly, the embedding of the student model is represented as .
[0075] For two given feature channels, the correlation metric should return a value to reflect the degree of association between the channels. A high value indicates that the channels are homologous, otherwise they are heterologous. Finally, all correlation metrics are collected in order to characterize the overall diversity of the feature channels. The correlation between features is expressed as:
[0076]
[0077] where represents the th channel of feature , vectorizes the 2D feature map into a vector of length , is a function that measures the correlation of the input, where the inner product is used. Rewritten in the form of matrix multiplication to form the ICC matrix:
[0078]
[0079] Among them, is a dimensionality reducer that flattens the spatial dimensions. denotes the transpose of, regardless of the spatial dimensions and how, the size of the obtained ICC (Information from the correlated Channels) matrix is . In this embodiment, a linear transformation layer is added to the features of the student model. This layer consists of a convolutional layer with a kernel and a batch normalization (BN) layer without an activation function. When the number of output channels of the student model is inconsistent with that of the teacher model, the dimensions of the student features can be adjusted to match the number of channels of the teacher features . This adjustment process does not change the spatial dimensions of the feature map. In addition, by constraining the L2 distance between the ICC matrices of the student and the teacher, the student model is encouraged to achieve similar feature diversity to the teacher, which is expressed as follows:
[0080]
[0081] This embodiment adopts a method for extracting the guiding information of the teacher's intermediate feature map, ensuring the scale matching of the teacher-student feature maps, while retaining the key information in the teacher's feature map, overcoming the problem of loss of feature correlation caused by filter pruning, so that the student model can effectively receive the guidance of the teacher model.
[0082] In the logical distillation of the second stage, this embodiment uses the simcc (Simple Coordinate Classification) algorithm to predict the pose key points, and this algorithm regards the key point localization as a classification task of horizontal and vertical coordinates. Following this design, this embodiment innovatively applies the logit-based knowledge distillation method to it. For logit-based distillation, this embodiment follows the form of the original classification loss of EfficientNet, but discards the weight mask for distillation. Different from the label values, reasonable values can also be assigned to the invisible key points by the teacher model. The loss function in the logical distillation stage is the new logical distillation loss . Specifically as follows:
[0083] This embodiment proposes a simple and effective coordinate classification method, which decomposes the human pose estimation task into two independent classification tasks for horizontal and vertical coordinates. Specifically, after the backbone network, classifiers in the horizontal and vertical directions (each classifier is a single-layer linear structure) are connected to classify the coordinates of each sample, so that each key point can be divided into multiple pixel-level categories in the horizontal and vertical directions. Based on this classification method, a classification task is established for each key point, and the student model learns the classification results of the teacher model. The distillation loss is measured by the difference between logits, where logits are the outputs of the penultimate layer of the model or the layer before softmax. The loss expression is as follows:
[0084]
[0085] where N represents the number of samples in a batch, K is the total number of key points, and L represents the length of the localization region in the x or y direction. T represents the prediction score of the teacher model, and S is the prediction score of the student model. L represents the lengths on the x and y axes, i is the i-th cell, represents the prediction score of the teacher model for the i-th cell, represents the prediction score of the student model for the i-th cell.
[0086] Construct a distribution loss based on the knowledge of the non-target distribution. To transfer the knowledge of the non-target distribution, this embodiment proposes a distributed loss, which is expressed as follows:
[0087]
[0088]
[0089]
[0090]
[0091] where, represents the prediction score of the teacher model for the target cell; is the non-target cell score of the i-th cell of the teacher model, that is, the proportion of the prediction score of the i-th cell among all non-target cells. represents the prediction score of the student model for the target cell; represents the non-target cell score of the i-th cell of the student model.
[0092] Construct a soft loss based on the target knowledge. In this case, it can be seen that , making it easier for the student to learn the non-target knowledge of the teacher. However, Lack of the teacher's target knowledge. Some early knowledge distillation (KD) methods have shown that the prediction results of the teacher model can be used as soft labels to help the student model converge faster and improve its performance. Inspired by this soft label method, in this embodiment, the prediction score of the teacher model for the target cell (i.e., the target output probability of the teacher model) is used as the soft target. Based on the soft target provided by the teacher model, this embodiment proposes a soft loss for distillation in combination with the teacher model:
[0093] ;
[0094] Original loss. The cross-entropy loss (or original loss) uses the properties of the logarithmic function to penalize the difference between the model prediction and the true label. When the probability predicted by the model is close to the true label, will be larger, resulting in a smaller loss value; conversely, when the difference between the prediction and the truth is large, will become smaller and the loss value will increase. It can guide the model to learn in the direction of making the prediction distribution closer to the true distribution. During the optimization process, the model will continuously adjust the parameters to minimize the cross-entropy loss, thereby improving the accuracy of the prediction.
[0095] Finally, combining the original loss , the distribution loss and the soft loss , this embodiment proposes a new logical distillation loss, which is expressed as follows:
[0096] ;
[0097] where, is the hyperparameter for balancing the loss.
[0098] This embodiment adopts a method to extract the guiding information of the intermediate feature map of the teacher model, ensuring the scale matching of the teacher-student feature maps, while retaining the key information in the teacher feature map, overcoming the problem of loss of feature correlation caused by filter pruning, so that the student model can effectively receive the guidance of the teacher model.
[0099] The total loss formula of the final student model is:
[0100] ;
[0101] S4. Obtain the human body video data, and input the human body video data into the trained student model for human pose estimation.
[0102] Deploy the trained lightweight student model into the inference engine of edge devices (such as smartphones, sports bracelets or camera devices).
[0103] The edge device captures a real-time video stream through a camera, and the video frames are processed by a student model deployed in the edge device to detect and analyze the coordinates of human key points.
[0104] Action analysis and feedback. Based on the detected key point coordinates and the action standard library, calculate the matching degree of each action and output real-time feedback. For example, the matching degree of a squat can be used to evaluate the correctness of the action. Actions with a matching degree higher than 90% are judged as standard actions, while those lower than 60% prompt the user to adjust the posture. At the same time, provide an action analysis report, including the overall score, suggestions for improving action details, etc.
[0105] The following is the specific implementation of the process of this embodiment. See Figure 3 as shown:
[0106] Construct the initial dataset:
[0107] With the help of a high-definition camera, collect videos covering various daily exercise actions such as running, push-ups, squats, and pull-ups. Each video is about 1 minute long and is finely frame-extracted at a frequency of 5 frames per second to construct a dataset of about 10,000 frames of images, providing sufficient data support for the subsequent lightweight human pose estimation process based on knowledge distillation. The collected video data can be formatted and frame-extracted using the FFmpeg tool to make it fit the subsequent data processing and annotation processes, ensuring the high quality and reliability of the data source for the entire lightweight human pose estimation framework based on knowledge distillation.
[0108] Label the dataset:
[0109] Use the Labelme annotation software to accurately label the human key points in each frame of the image, such as the head, shoulders, elbows, knees, and ankles, etc., and then generate approximately 200,000 key point labels. The labeled data is properly saved in JSON format and preprocessed using OpenCV, such as normalization, size adjustment, etc., to comprehensively ensure the consistency and high quality of the training data, laying a solid foundation for training accurate models in this lightweight human pose estimation framework.
[0110] Build a pose estimation framework:
[0111] Leveraging the TensorFlow deep learning framework, we meticulously constructed a lightweight human pose estimation model using EfficientNet as its backbone network, leveraging EfficientNet's efficient computing performance and excellent scalability within this framework. During implementation, we loaded the pre-trained EfficientNet model via TensorFlow Hub and cleverly utilized transfer learning techniques to precisely adapt it to the requirements of keypoint detection for motion poses, creating a solid model core for this lightweight human pose estimation framework based on knowledge distillation.
[0112] Training the model:
[0113] The teacher model is fully trained in the TensorFlow environment, deeply exploring the potential of labeled data to improve model accuracy. Subsequently, knowledge distillation techniques (for example, using the Distiller library) are closely integrated, allowing the teacher model to guide the training of a lightweight student model. This significantly reduces the number of model parameters by approximately 50%, enabling the model to meet the stringent requirements of subsequent edge device deployment and ensuring that the lightweight human pose estimation framework based on knowledge distillation can run smoothly even on resource-constrained edge devices.
[0114] Model deployment:
[0115] The lightweight student model was converted to the TFLite format for edge devices using TensorFlow Lite, enabling it to perform low-latency inference tasks. Furthermore, advanced tools such as NVIDIA TensorRT were cleverly leveraged to further optimize model performance, ensuring that edge devices could consistently process over 10 frames per second of real-time data streams. This enabled the lightweight human pose estimation framework based on knowledge distillation to run efficiently and in real time in edge scenarios, accurately capturing changes in human pose.
[0116] Perform pose estimation:
[0117] The camera captures the user's real-time video stream and efficiently streams it via GStreamer, ensuring smooth, high-quality video input to the pose estimation module within the knowledge distillation-based lightweight human pose estimation framework. This module accurately detects the coordinates of key points on the human body in real time and uses OpenCV to visualize and fine-tune these key points. This provides precise input data for the subsequent motion scoring within the knowledge distillation-based lightweight human pose estimation framework, seamlessly integrating the entire analysis process.
[0118] Sports Action Rating:
[0119] Integrate the powerful functions of the SciPy and NumPy libraries in Python, calculate the similarity between the key point coordinates for real-time detection and the ideal postures in the standard action library, and generate accurate action scores frame by frame. Relying on the scoring system of this lightweight human pose estimation framework, it can provide real-time feedback, clearly display the key point matching degree, and generate a visual scoring curve with Matplotlib to provide highly targeted action improvement suggestions for users, helping to improve the action standardization and training effect, and fully demonstrating the practical value of the lightweight human pose estimation framework based on knowledge distillation in the analysis of daily sports scenarios.
[0120] See Figure 4 As shown, the present invention also discloses a lightweight human pose estimation device based on knowledge distillation, including:
[0121] A dataset construction and marking module 401, which is used to construct an initial human video dataset and mark the dataset;
[0122] A model construction module 402, which is used to construct a teacher model and a student model for human pose estimation;
[0123] A distillation training module 403, which uses the marked dataset to train the teacher model, and uses the trained teacher model to guide the training of the student model to obtain a trained student model; the guiding training includes a feature distillation stage and a logic distillation stage; the feature distillation is based on the intermediate feature maps of the student model and the pruned intermediate feature maps of the teacher model; the logic distillation is based on the logical outputs of the teacher model and the student model;
[0124] A human pose estimation module 404, which is used to obtain human video data and input the human video data into the trained student model for human pose estimation.
[0125] The specific implementation of the lightweight human pose estimation device based on knowledge distillation is the same as that of the lightweight human pose estimation method based on knowledge distillation, and will not be repeated in this embodiment.
[0126] The above is only the specific implementation manner of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantive modification made to the present invention using this concept shall fall within the scope of infringement of the protection of the present invention.
Claims
1. A lightweight human pose estimation method based on knowledge distillation, characterized in that, It includes the following steps: S1. Construct an initial human body video dataset and label the dataset; S2. Construct a teacher model and a student model for human pose estimation; S3. Use the labeled dataset to train the teacher model, and use the trained teacher model to guide the training of the student model to obtain a trained student model; the guidance training includes a feature distillation stage and a logic distillation stage; the feature distillation is based on the intermediate feature maps of the student model and the pruned intermediate feature maps of the teacher model; the logic distillation is based on the logical outputs of the teacher model and the student model; The process of obtaining the pruned intermediate feature maps of the teacher model is specifically as follows: Slice the intermediate feature maps of the same batch in the teacher model by channel, calculate the average value of the ranks of all intermediate feature maps within each channel, and use the average value to represent the rank value of the feature map of this channel; Sort the rank values of each layer of feature maps, and divide the intermediate feature maps into a retention set and a pruning set according to the sorting result of the rank values. The retention set stores the high-rank feature maps, and the pruning set stores the low-rank feature maps; the retention set is the pruned intermediate feature maps of the teacher model; S4. Obtain human body video data, and input the human body video data into the trained student model for human pose estimation; The loss function in the feature distillation stage is the channel correlation information loss, which is specifically as follows: Unfold the pruned intermediate feature maps of the teacher model to obtain the teacher model correlation matrix; unfold the intermediate feature maps of the student model to generate the student model correlation matrix; Construct an L2 constraint based on the teacher model correlation matrix and the student model correlation matrix to obtain the channel correlation information loss; the channel correlation information loss is expressed as: Among them, L cc represents the loss of channel correlation information; represents the associated matrix of the student model after adjusting the dimension; represents the associated matrix of the teacher model; F S represents the intermediate feature map of the student model; F T represents the intermediate feature map of the teacher model after pruning; c represents the number of channels of the teacher model; represents the square of the L2 distance; C l is used to adjust the feature dimension of the student model to match the number of channels of the teacher model features; G represents the ICC matrix; The logical outputs of the teacher model and the student model use the simcc algorithm to predict the pose key points, and decompose the human pose estimation task into independent classification tasks of horizontal coordinates and vertical coordinates; The loss function in the logic distillation stage is the logic distillation loss; the logic distillation loss combines the original loss, the distribution loss, and the soft loss; the distribution loss enables the student model to learn the non-target knowledge of the teacher model; the soft loss enables the student model to learn the target knowledge of the teacher model; the original loss is the cross-entropy loss of the student model.
2. The lightweight human pose estimation method based on knowledge distillation according to claim 1, wherein The logic distillation loss is expressed as: Among them, L newlogit represents the logical distillation loss; α represents the hyperparameter for balancing the loss; N represents the number of samples in a batch; K represents the total number of key points; L represents the length of the localization region in the horizontal coordinate x or vertical coordinate y direction; represents the non-target cell score of the teacher model; represents the non-target cell score of the student model; T t represents the predicted score of the teacher model for the target cell; S t represents the predicted score of the student model for the target cell; t represents the target cell; i represents the i-th cell.
3. The lightweight human pose estimation method based on knowledge distillation according to claim 1, characterized in that The distribution loss is expressed as: Among them, L distributed represents the distribution loss; N represents the number of samples in a batch; K represents the total number of key points; L represents the length of the positioning area in the horizontal coordinate x or vertical coordinate y direction; t represents the target cell; represents the score of the non-target cell of the teacher model; represents the score of the non-target cell of the student model; i represents the i-th cell; T i represents the predicted score of the teacher model for the i-th cell; T t represents the predicted score of the teacher model for the target cell; S i represents the predicted score of the student model for the i-th cell; S t represents the predicted score of the student model for the target cell.
4. The lightweight human pose estimation method based on knowledge distillation according to claim 1, wherein The soft loss is expressed as: L soft = -T t log(S t ) Among them, L soft represents the soft loss; T t represents the prediction score of the teacher model for the target cell; S t represents the prediction score of the student model for the target cell; t represents the target cell.
5. The lightweight human pose estimation method based on knowledge distillation according to claim 1, characterized in that The total loss function of the student model is expressed as: L tol = L newlogit + L cc ; Among them, L tol represents the total loss function; L cc represents the loss of channel correlation information; L newlogit represents the logical distillation loss.
6. A lightweight human pose estimation device based on knowledge distillation using the lightweight human pose estimation method based on knowledge distillation according to any one of claims 1-5, including the following: A dataset construction and labeling module for constructing an initial human body video dataset and labeling the dataset; A model construction module for constructing a teacher model and a student model for human pose estimation; A distillation training module that uses the labeled dataset to train the teacher model and uses the trained teacher model to guide the training of the student model to obtain a trained student model; the guidance training includes a feature distillation stage and a logic distillation stage; the feature distillation is based on the intermediate feature maps of the student model and the pruned intermediate feature maps of the teacher model; the logic distillation is based on the logical outputs of the teacher model and the student model; The process of obtaining the intermediate feature maps of the pruned teacher model is as follows: Slice the intermediate feature maps of the same batch in the teacher model by channel, calculate the average value of the ranks of all intermediate feature maps within each channel, and use the average value to represent the rank value of the feature map of this channel; Sort the rank values of each layer of feature maps, divide the intermediate feature maps into a retention set and a pruning set according to the sorting result of the rank values, store the high-rank feature maps in the retention set, and store the low-rank feature maps in the pruning set; the retention set is the intermediate feature maps of the pruned teacher model; The human pose estimation module is used to obtain human video data and input the human video data into the trained student model for human pose estimation.
Citation Information
Patent Citations
Underground coal mine human body action recognition method suitable for edge terminal
CN116189299A
Personalized human body action recognition method based on knowledge distillation
CN116844225A