Lightweight human body posture estimation method and device based on knowledge distillation
Through knowledge distillation technology, the pruned teacher model intermediate feature map is used to guide student models for feature distillation, which solves the real-time reasoning problem of deep learning models on resource-limited equipment in the existing technology, and realizes efficient and accurate human posture estimation, which is suitable for real-time action feedback in complex motion scenarios.
Patent Information
- Application Number
- CN202510437685.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-09
AI Technical Summary
Existing deep learning models are difficult to achieve real-time inference on devices with limited resources, and there are misjudgment or delays in complex motion scenarios, making it difficult to meet the needs of efficiency and accuracy.
A lightweight human pose estimation method based on knowledge distillation is used to guide student models to perform feature distillation through the pruned teacher model intermediate feature map to reduce model parameters, while ensuring the efficiency and accuracy of the pose estimation model.
Real-time inference capability on edge devices is realized, the accuracy and efficiency of the model in complex motion scenarios is improved, accurate real-time motion feedback is provided, and the risk of motion damage is significantly reduced.
Smart Images

Figure CN119942655A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human body posture estimation, and in particular to a lightweight human body posture estimation method and device based on knowledge distillation. Background Art
[0002] With the popularization of the concept of national fitness, more and more people are participating in daily sports such as running, fitness, and strength training. Standardized sports movements are crucial to training results and physical health. Irregular sports postures may not only lead to poor training results, but may also cause sports injuries. Therefore, how to analyze and feedback movements in real time during daily sports has become a widely concerned technical demand.
[0003] Traditional sports guidance relies on on-site supervision or video courses by coaches, but these methods cannot meet the needs of personalized, real-time and large-scale scenarios. For example, in the gym, it is difficult for coaches to pay attention to every trainer at the same time; when exercising at home or remotely, the lack of real-time feedback will also affect the training effect.
[0004] Posture estimation in artificial intelligence technology provides an effective solution to this problem. By capturing the coordinates of key points of the human body and analyzing the differences between user movements and standard movements, posture estimation algorithms can provide real-time feedback to users. However, existing deep learning models are usually based on large-scale neural networks. The models are bulky and computationally expensive, making it difficult to achieve real-time reasoning on devices with limited resources (such as edge devices or mobile terminals). In addition, in complex motion scenarios, traditional models often have misjudgments or delays, making it difficult to meet the needs of efficiency and accuracy.
[0005] Knowledge distillation technology is an effective model compression method that achieves a balance between accuracy and efficiency by guiding the lightweight student model to learn through the teacher model. Based on high-precision prediction, the teacher model transfers knowledge to the student model, which can simplify the network structure while maintaining high prediction accuracy by learning the output of the teacher model. Summary of the invention
[0006] In response to the above problems, the present invention proposes a lightweight human posture estimation method and device based on knowledge distillation. By introducing knowledge distillation technology, the pruned intermediate feature graph of the teacher model is used to guide the student model to perform feature distillation, which greatly reduces the model parameters while ensuring the efficiency and accuracy of the posture estimation model.
[0007] On the one hand, the lightweight human posture estimation method based on knowledge distillation has the following specific steps:
[0008] S1, construct the initial human video dataset and label the dataset;
[0009] S2, constructing a teacher model and a student model for human pose estimation;
[0010] S3, using the labeled data set to train the teacher model, and using the trained teacher model to guide the student model to obtain a trained student model; the guided training includes a feature distillation stage and a logic distillation stage; the feature distillation is based on the intermediate feature map of the student model and the pruned intermediate feature map of the teacher model; the logic distillation is based on the logic output of the teacher model and the student model;
[0011] S4, obtaining human body video data, and inputting the human body video data into the trained student model for human body posture estimation.
[0012] Preferably, the process of obtaining the intermediate feature map of the pruned teacher model is as follows:
[0013] The intermediate feature maps of the same batch in the teacher model are sliced by channel, and the average of the ranks of all intermediate feature maps in each channel is calculated, and the rank value of the feature map of the channel is represented by the average value;
[0014] The rank values of the feature maps of each layer are sorted, and the intermediate feature maps are divided into a retention set and a cropping set according to the rank sorting results. The retention set stores high-rank feature maps, and the cropping set stores low-rank feature maps. The retention set is the intermediate feature map of the teacher model after pruning.
[0015] Preferably, the loss function of the feature distillation stage is channel correlation information loss, which is as follows:
[0016] Expand the intermediate feature graph of the pruned teacher model to obtain the teacher model association matrix; expand the intermediate feature graph of the student model to generate the student model association matrix;
[0017] Based on the teacher model association matrix and the student model association matrix, the L2 constraint is constructed to obtain the channel association information loss; the channel association information loss is expressed as:
[0018] ;
[0019] in, Indicates the loss of channel-related information; represents the student model correlation matrix after adjusting the dimension; represents the association matrix of the teacher model; Represents the intermediate feature map of the student model; Represents the intermediate feature map after the teacher model is pruned; Indicates the number of channels of the teacher model; represents the L2 distance squared; The number of channels used to adjust the feature dimension of the student model to match the features of the teacher model; Represents the ICC matrix.
[0020] Preferably, the teacher model and the student model logic output use the SIMCC algorithm to predict posture key points, and decompose the human posture estimation task into independent classification tasks of horizontal coordinates and vertical coordinates.
[0021] Preferably, the loss function of the logic distillation stage is the logic distillation loss; the logic distillation loss combines the original loss, the distribution loss and the soft loss; the distribution loss enables the student model to learn the non-target knowledge of the teacher model; the soft loss enables the student model to learn the target knowledge of the teacher model; the original loss is the cross entropy loss of the student model.
[0022] Preferably, the logical distillation loss is expressed as:
[0023] ;
[0024] in, represents the logistic distillation loss; represents the hyperparameter of the balanced loss; N represents the number of samples in a batch; K represents the total number of key points; L represents the length of the positioning area in the horizontal coordinate x or vertical coordinate y direction; represents the non-target cell score of the teacher model; represents the non-target cell score of the student model; represents the prediction score of the teacher model for the target cell; represents the prediction score of the student model for the target cell; t represents the target cell; i represents the i-th cell.
[0025] Preferably, the distribution loss is expressed as:
[0026] ;
[0027] ;
[0028] ;
[0029] in, represents the distribution loss; N represents the number of samples in a batch; K represents the total number of key points; L represents the length of the positioning area in the horizontal coordinate x or vertical coordinate y direction; t represents the target cell; represents the non-target cell score of the teacher model; represents the non-target cell score of the student model; i represents the i-th cell; represents the prediction score of the teacher model for the i-th cell; represents the prediction score of the teacher model for the target cell; represents the prediction score of the student model for the i-th cell; Represents the prediction score of the student model for the target cell.
[0030] Preferably, the soft loss is expressed as:
[0031] ;
[0032] in, Indicates soft loss; represents the prediction score of the teacher model for the target cell; represents the prediction score of the student model for the target cell; t represents the target cell.
[0033] Preferably, the total loss function of the student model is expressed as:
[0034] ;
[0035] in, represents the total loss function; Indicates the loss of channel-related information; represents the logistic distillation loss.
[0036] On the other hand, a lightweight human posture estimation device based on knowledge distillation includes the following:
[0037] The dataset construction and labeling module is used to construct the initial dataset of human videos and label the dataset;
[0038] Model building module, used to build teacher model and student model for human pose estimation;
[0039] A distillation training module uses a labeled data set to train a teacher model, and uses the trained teacher model to guide the student model to obtain a trained student model; the guided training includes a feature distillation stage and a logic distillation stage; the feature distillation is based on the intermediate feature map of the student model and the pruned intermediate feature map of the teacher model; the logic distillation is based on the logic output of the teacher model and the student model;
[0040] The process of obtaining the intermediate feature map of the pruned teacher model is as follows:
[0041] The intermediate feature maps of the same batch in the teacher model are sliced by channel, and the average of the ranks of all intermediate feature maps in each channel is calculated, and the rank value of the feature map of the channel is represented by the average value;
[0042] Sort the rank values of the feature maps of each layer, and divide the intermediate feature maps into a reserved set and a pruned set according to the rank sorting results. The reserved set stores high-rank feature maps, and the pruned set stores low-rank feature maps. The reserved set is the intermediate feature map of the teacher model after pruning.
[0043] The human body posture estimation module is used to obtain human body video data and input the human body video data into the trained student model for human body posture estimation.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] (1) The present invention prunes the intermediate feature graphs of the teacher model by using rank-guided filter pruning and feature correlation reconstruction, which greatly reduces the model parameters while ensuring the efficiency and accuracy of the posture estimation model in the lightweight human posture estimation model based on knowledge distillation, further improving the real-time reasoning capability of the model on edge devices, and providing a rock-solid technical guarantee for the general motion analysis system based on this framework in daily sports scenarios;
[0046] (2) The present invention introduces an efficient backbone network architecture, EfficientNet, to extract feature representations of multiple key points. EfficientNet achieves a good balance between speed and accuracy, while having low computing resource requirements and is suitable for edge device deployment.
[0047] (3) The lightweight student model of the present invention can be deployed on edge devices to process image data in real time and quickly, accurately identify the coordinates of key points of the human body, and compare them with the preset standard library in detail, and instantly output action matching and feedback, ultimately achieving fast and low-power posture estimation on edge devices; during exercise, users can receive accurate real-time action feedback and practical suggestions, thereby effectively optimizing exercise effects and significantly reducing the risk of sports injuries. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The present invention is further described in detail below in conjunction with the accompanying drawings;
[0049] Figure 1 Flow chart of a lightweight human posture estimation method based on knowledge distillation according to an embodiment of the present invention;
[0050] Figure 2 A network structure diagram of a lightweight human posture estimation method based on knowledge distillation according to an embodiment of the present invention;
[0051] Figure 3 Schematic diagram of a flow chart of a lightweight human posture estimation method based on knowledge distillation according to an embodiment of the present invention;
[0052] Figure 44 is a structural block diagram of a lightweight human posture estimation device based on knowledge distillation according to an embodiment of the present invention. DETAILED DESCRIPTION
[0053] The present invention is further described below through specific implementation modes.
[0054] like Figure 1 As shown in the figure, the lightweight human posture estimation method based on knowledge distillation is as follows:
[0055] S1, construct the initial human video dataset and label the dataset.
[0056] We use high-definition cameras to collect video data covering a variety of daily sports movements (such as running, push-ups, squats, etc.), and extract frames from each video to obtain the initial data set. These data can cover common sports scenes and ensure the diversity and generalization of movements.
[0057] The key points of daily sports actions in the dataset are annotated, and the coordinates of the joint points of each posture (such as shoulders, knees, ankles, etc.) are recorded. The generated annotated data provides high-quality supervision signals for model training and ensures the accuracy of motion analysis.
[0058] S2, constructs a teacher model and a student model for human pose estimation.
[0059] An efficient backbone network architecture EfficientNet is introduced to extract feature representations of multiple key points. EfficientNet achieves a good balance between speed and accuracy, and has low computing resource requirements, making it suitable for edge device deployment. In this embodiment, both the teacher and the student are EfficientNet models, the teacher model is EfficientNet-L, and the student model is EfficientNet-M. The size of the student model is smaller than that of the teacher model.
[0060] S3, using the labeled data set to train the teacher model, and using the trained teacher model to guide the student model to obtain a trained student model; the guided training includes a feature distillation stage and a logic distillation stage; the feature distillation is based on the intermediate feature map of the student model and the pruned intermediate feature map of the teacher model; the logic distillation is based on the logic output of the teacher model and the student model.
[0061] The teacher model is trained using a well-labeled dataset, and the efficient convolutional network structure is used to accurately learn complex motion features. Then, using knowledge distillation technology, the teacher model guides the lightweight student model to train, which significantly reduces the number of parameters and computing requirements while maintaining high accuracy.
[0062] See also Figure 2 As shown, the distillation of this embodiment includes two stages.
[0063] In the first stage of feature distillation, this embodiment proposes a rank-guided filter pruning method, which provides auxiliary support for the student model in learning by retaining and extracting the guidance information in the intermediate feature map of the teacher model. On this basis, this embodiment also adopts a method for reconstructing feature correlation to compensate for the common problem of feature correlation information loss in the filter pruning process.
[0064] The pruning method is as follows:
[0065] For rank-guided filter pruning methods, in convolutional neural networks, a reasonable choice of batch size not only helps to fully utilize memory capacity, but also improves the network's operating efficiency. The resulting feature map data of each layer is stored in a four-dimensional matrix form , where N represents the number of batches, C is the number of channels of the feature map, and W and H represent the width and height of the feature map, respectively. Analysis of the convolution process of the feature map shows that for the feature maps of the same batch, the convolution calculation of the filter at each position is the same. Since the average rank of the feature map generated by the same filter remains relatively stable, the feature maps in the same batch can be sliced by channel, and the average of the ranks of all feature maps in each channel is calculated, and the average is used to represent the rank value of the feature map of the channel. The information measure used here is The definition is as follows:
[0066] ;
[0067] in, Indicates Feature map, Indicates The number of channels of the feature map, It is The collection of layer feature map ranks. Through the information metric function We obtain a one-dimensional vector , each value represents the mean of the feature map corresponding to each channel. The calculation result of the layer feature map rank is expressed as ( Indicates the value of the index 0 channel rank. Indicates the index position channel rank value).
[0068] By sorting the calculated feature map rank values, we can obtain the channel-based importance ranking index. Through the high-rank index value, we can identify the feature map with higher importance, and divide the original feature map into two subsets based on these high-rank indexes, expressed as:
[0069]
[0070]
[0071] Keep Set Store high-rank feature maps and crop sets Store low-rank feature maps. The above two sets constitute the first The original feature map of the layer, and there is no intersection between the two sets. In this way, the retained feature information set can be screened out.
[0072] The method for reconstructing feature association is as follows:
[0073] In order to address the problem of possible loss of correlation information in refined feature maps, this embodiment expands the feature maps of the teacher and student models respectively and generates their own correlation matrices to reconstruct and transfer the correlation information between the refined feature maps, thereby ensuring that the guiding role of the intermediate feature maps of the teacher model is fully exerted.
[0074] Specifically, let the characteristics of the teacher model and the student model be and In practical applications, and can be any differentiable functions, which are parameterized as convolutional neural networks (CNNs). The embedding of the teacher model is represented as ,in, Indicates the number of output channels, and Represent the height and width of the feature map respectively. Similarly, the embedding of the student model is expressed as .
[0075] For a given two feature channels, the correlation measure should return a value to reflect the degree of association between the channels. A high value indicates that the channels are homologous, otherwise they are heterologous. Finally, all correlation indicators are collected in order to characterize the overall diversity of feature channels. The correlation between features is expressed as:
[0076]
[0077] in, Representation characteristics No. channels, Vectorize the 2D feature map to a length of The vector of is a function that measures the correlation of the input pair, using the inner product. Rewriting it in terms of matrix multiplication forms the ICC matrix:
[0078]
[0079] in, It is a dimensionality reducer that flattens the spatial dimensions. express The transpose of and The size of the ICC (Information from the correlated Channels) matrix is This embodiment adds a linear transformation layer to the features of the student model , which consists of a The convolutional layer with a kernel and a batch normalization (BN) layer without activation function. When the number of output channels is inconsistent with that of the teacher model, Can adjust student characteristics The dimension of to match the teacher characteristics Number of channels This adjustment process does not change the spatial dimension of the feature map. In addition, the L2 distance between the ICC matrices of the student and the teacher is constrained to force the student model to achieve similar feature diversity as the teacher, as shown below:
[0080]
[0081] This embodiment adopts a method for extracting guidance information from the teacher's intermediate feature map, ensuring the matching of the teacher and student feature maps in scale, while retaining the key information in the teacher's feature map, overcoming the problem of feature correlation loss caused by filter pruning, so that the student model can effectively accept the guidance of the teacher model.
[0082] In the second stage of logical distillation, this embodiment uses the simcc (Simple Coordinate Classification) algorithm to predict posture key points, which treats key point positioning as a classification task of horizontal and vertical coordinates. Following this design, this embodiment innovatively applies the logit-based knowledge distillation method to it. For logit-based distillation, this embodiment follows the form of EfficientNet's original classification loss, but discards the weight mask used for distillation. Unlike label values, invisible key points can also be assigned a reasonable value by the teacher model. The loss function of the logical distillation stage is the new logical distillation loss The details are as follows:
[0083] This embodiment proposes a simple and effective coordinate classification method, which decomposes the human posture estimation task into two independent classification tasks of horizontal and vertical coordinates. Specifically, the backbone network is connected to the horizontal and vertical classifiers (each classifier is a single-layer linear structure) to classify the coordinates of each sample, so that each key point can be divided into multiple pixel-level categories in the horizontal and vertical directions. Based on this classification method, a classification task is established for each key point, and the student model is allowed to learn the classification results of the teacher model. The distillation loss is measured by the difference between logits, which is the output of the penultimate layer of the model or the previous layer of softmax. The loss expression is:
[0084]
[0085] Where N is the number of samples in a batch, K is the total number of key points, and L is the length of the localization region in the x or y direction. T is the prediction score of the teacher model, and S is the prediction score of the student model. L represents the length on the x and y axes, i is the i-th cell, represents the prediction score of the teacher model for the i-th cell, represents the prediction score of the student model for the i-th cell.
[0086] A distribution loss is constructed based on the knowledge of the non-target distribution. In order to transfer the knowledge of the non-target distribution, this embodiment proposes a distribution loss, which is expressed as follows:
[0087]
[0088]
[0089]
[0090]
[0091] in, represents the prediction score of the teacher model for the target cell; is the non-target cell score of the i-th cell of the teacher model, that is, the proportion of the predicted scores of the i-th cell in all non-target cells. represents the prediction score of the student model for the target cell; Represents the non-target cell score of the i-th cell of the student model.
[0092] The soft loss is constructed based on the target knowledge. In this case, we can see , making it easier for students to learn the teacher's non-target knowledge. However, Lack of target knowledge of the teacher. Some early knowledge distillation (KD) methods have shown that the prediction results of the teacher model can be used as soft labels to help the student model converge faster and improve its performance. Inspired by this soft label method, this embodiment directly uses the prediction score of the teacher model for the target cell (that is, the target output probability of the teacher model) As a soft target. Based on the soft target provided by the teacher model, this embodiment proposes a soft loss combined with the teacher model for distillation:
[0093] ;
[0094] Original loss. Cross entropy loss (or original loss) The properties of the logarithmic function are used to penalize the difference between the model prediction and the true label. When the probability predicted by the model is close to the true label, The value of will be larger, resulting in a smaller loss value; on the contrary, when the difference between the prediction and the actual value is large, It can guide the model to learn in the direction of making the predicted distribution closer to the true distribution. During the optimization process, the model will continuously adjust the parameters to minimize the cross entropy loss, thereby improving the accuracy of the prediction.
[0095] Finally, combined with the original loss , distribution loss and soft loss , this embodiment proposes a new logical distillation loss, which is expressed as follows:
[0096] ;
[0097] in, is the hyperparameter of the balance loss.
[0098] This embodiment adopts a method for extracting guidance information from the intermediate feature graph of the teacher model, ensuring the matching of the teacher and student feature graphs in scale, while retaining the key information in the teacher feature graph, overcoming the problem of feature correlation loss caused by filter pruning, so that the student model can effectively accept the guidance of the teacher model.
[0099] The total loss formula of the final student model is:
[0100] ;
[0101] S4, obtaining human body video data, and inputting the human body video data into the trained student model for human body posture estimation.
[0102] Deploy the trained lightweight student model to the inference engine of edge devices such as smartphones, sports bracelets, or camera devices.
[0103] The edge device captures real-time video streams through the camera, and the video frames are processed by the student model deployed in the edge device to detect and analyze the coordinates of key points of the human body.
[0104] Action analysis and feedback. Based on the detected key point coordinates and the action standard library, the matching degree of each action is calculated and real-time feedback is output. For example, the matching degree of squat can be used to evaluate the correctness of the action. Actions above 90% will be judged as standard actions, while those below 60% will prompt the user to adjust the posture. At the same time, an action analysis report is provided, including overall scores, action details improvement suggestions, etc.
[0105] The following is a specific implementation of the process of this embodiment, see Figure 3 As shown:
[0106] Construct the initial dataset:
[0107] With the help of high-definition cameras, we collected videos of various daily sports movements such as running, push-ups, squats, and pull-ups. Each video is about 1 minute long, and fine frame extraction is performed at a frequency of 5 frames per second to construct a data set of about 10,000 frames of images, providing sufficient data support for the subsequent lightweight human posture estimation process based on knowledge distillation. The collected video data can be format converted and frame extracted using the FFmpeg tool to make it compatible with the subsequent data processing and annotation process, ensuring that the data source of the entire lightweight human posture estimation framework based on knowledge distillation is high-quality and reliable.
[0108] Labeled dataset:
[0109] Labelme annotation software is used to accurately mark the key points of the human body in each frame, such as the head, shoulders, elbows, knees, and ankles, and then generate about 200,000 key point labels. The annotated data is properly saved in JSON format, and pre-processed with OpenCV, such as normalization and resizing, to ensure the consistency and high quality of the training data in all aspects, laying a solid foundation for training accurate models under this lightweight human posture estimation framework.
[0110] Build the pose estimation framework:
[0111] Relying on the TensorFlow deep learning framework, we carefully built a lightweight human posture estimation model with EfficientNet as the backbone network, giving full play to the advantages of EfficientNet's efficient computing performance and good scalability under this framework. In the specific implementation process, the pre-trained model of EfficientNet can be loaded through TensorFlow Hub, and the transfer learning method can be cleverly used to make it accurately adapt to the key point detection requirements of motion posture, creating a solid model core for the lightweight human posture estimation framework based on knowledge distillation.
[0112] Train the model:
[0113] The teacher model is fully trained in the TensorFlow environment, and the potential of labeled data is deeply explored to improve the model accuracy. Then, the teacher model is closely combined with knowledge distillation technology (for example, with the help of the Distiller library) to guide the lightweight student model to train, which greatly reduces the model parameters by about 50%, enabling the model to meet the stringent requirements of subsequent deployment on edge devices, and ensuring that the lightweight human posture estimation framework based on knowledge distillation can run smoothly on the resource-constrained edge.
[0114] Model deployment:
[0115] The lightweight student model is converted into the TFLite format adapted to edge devices with the help of TensorFlow Lite, enabling it to perform low-latency reasoning tasks. At the same time, the model performance is further optimized with the help of advanced tools such as NVIDIA TensorRT, ensuring that edge devices can stably process more than 10 frames of real-time data streams per second, allowing the lightweight human posture estimation framework based on knowledge distillation to run in real time and efficiently in edge scenarios, accurately capturing changes in human posture.
[0116] Perform pose estimation:
[0117] The camera is used to keenly capture the user's real-time video stream, and GStreamer is used for efficient streaming media transmission to ensure that high-quality video is smoothly input into the posture estimation module in the lightweight human posture estimation framework based on knowledge distillation. This module will accurately detect the key point coordinates of the human body in real time, and use OpenCV to visualize and fine-tune the key points, providing accurate input data for the subsequent action scoring link under the lightweight human posture estimation framework based on knowledge distillation, seamlessly connecting the entire analysis process.
[0118] Sports Action Rating:
[0119] Combining the powerful functions of Python's SciPy and NumPy libraries, the similarity calculation is performed on the key point coordinates detected in real time and the ideal posture in the standard action library, and accurate action scores are generated frame by frame. The scoring system based on this lightweight human posture estimation framework can provide real-time feedback, clearly display the key point matching degree, and generate a visual scoring curve with the help of Matplotlib, providing users with highly targeted action improvement suggestions, helping to improve the standardization of actions and training effects, and fully demonstrating the practical value of the lightweight human posture estimation framework based on knowledge distillation in daily sports scene analysis.
[0120] See also Figure 4 As shown, the present invention also discloses a lightweight human posture estimation device based on knowledge distillation, comprising:
[0121] The data set construction and labeling module 401 is used to construct an initial data set of human body videos and label the data set;
[0122] A model building module 402, used to build a teacher model and a student model for human posture estimation;
[0123] The distillation training module 403 uses the labeled data set to train the teacher model, and uses the trained teacher model to guide the student model to obtain a trained student model; the guided training includes a feature distillation stage and a logic distillation stage; the feature distillation is based on the intermediate feature map of the student model and the intermediate feature map of the teacher model after pruning; the logic distillation is based on the logic output of the teacher model and the student model;
[0124] The human body posture estimation module 404 is used to obtain human body video data and input the human body video data into the trained student model to perform human body posture estimation.
[0125] The specific implementation of the lightweight human posture estimation device based on knowledge distillation is the same as the lightweight human posture estimation method based on knowledge distillation, and will not be repeated in this embodiment.
[0126] The above is only a specific implementation of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantial changes to the present invention using this concept shall be deemed as an infringement of the protection scope of the present invention.
Claims
1. A lightweight human posture estimation method based on knowledge distillation, characterized in that: The steps include: S1, construct the initial human video dataset and label the dataset; S2, constructing a teacher model and a student model for human pose estimation; S3, using the labeled data set to train the teacher model, and using the trained teacher model to guide the student model to obtain a trained student model; the guided training includes a feature distillation stage and a logic distillation stage; the feature distillation is based on the intermediate feature map of the student model and the pruned intermediate feature map of the teacher model; the logic distillation is based on the logic output of the teacher model and the student model; The process of obtaining the intermediate feature map of the pruned teacher model is as follows: The intermediate feature maps of the same batch in the teacher model are sliced by channel, and the average of the ranks of all intermediate feature maps in each channel is calculated, and the rank value of the feature map of the channel is represented by the average value; Sort the rank values of the feature maps of each layer, and divide the intermediate feature maps into a reserved set and a pruned set according to the rank sorting results. The reserved set stores high-rank feature maps, and the pruned set stores low-rank feature maps. The reserved set is the intermediate feature map of the teacher model after pruning. S4, obtaining human body video data, and inputting the human body video data into the trained student model for human body posture estimation.
2. The lightweight human posture estimation method based on knowledge distillation according to claim 1 is characterized in that: The loss function of the feature distillation stage is the channel correlation information loss, which is as follows: Expand the intermediate feature graph of the pruned teacher model to obtain the teacher model association matrix; expand the intermediate feature graph of the student model to generate the student model association matrix; Based on the teacher model association matrix and the student model association matrix, the L2 constraint is constructed to obtain the channel association information loss; the channel association information loss is expressed as: ; in, Indicates the loss of channel-related information; represents the student model correlation matrix after adjusting the dimension; represents the association matrix of the teacher model; Represents the intermediate feature map of the student model; Represents the intermediate feature map after the teacher model is pruned; Indicates the number of channels of the teacher model; represents the L2 distance squared; The number of channels used to adjust the feature dimension of the student model to match the features of the teacher model; Represents the ICC matrix.
3. The lightweight human posture estimation method based on knowledge distillation according to claim 1 is characterized in that: The teacher model and the student model logic output use the SIMCC algorithm to predict the posture key points, and the human posture estimation task is decomposed into independent classification tasks of horizontal coordinates and vertical coordinates.
4. The lightweight human posture estimation method based on knowledge distillation according to claim 3 is characterized in that: The loss function of the logic distillation stage is the logic distillation loss; the logic distillation loss combines the original loss, distribution loss and soft loss; the distribution loss enables the student model to learn the non-target knowledge of the teacher model; the soft loss enables the student model to learn the target knowledge of the teacher model; the original loss is the cross entropy loss of the student model.
5. The lightweight human posture estimation method based on knowledge distillation according to claim 4 is characterized in that: The logistic distillation loss is expressed as: ; in, represents the logistic distillation loss; represents the hyperparameter of the balanced loss; N represents the number of samples in a batch; K represents the total number of key points; L represents the length of the positioning area in the horizontal coordinate x or vertical coordinate y direction; represents the non-target cell score of the teacher model; represents the non-target cell score of the student model; represents the prediction score of the teacher model for the target cell; represents the prediction score of the student model for the target cell; t represents the target cell; i represents the i-th cell.
6. The lightweight human posture estimation method based on knowledge distillation according to claim 4 is characterized in that: The distribution loss is expressed as: ; ; ; in, represents the distribution loss; N represents the number of samples in a batch; K represents the total number of key points; L represents the length of the positioning area in the horizontal coordinate x or vertical coordinate y direction; t represents the target cell; represents the non-target cell score of the teacher model; represents the non-target cell score of the student model; i represents the i-th cell; represents the prediction score of the teacher model for the i-th cell; represents the prediction score of the teacher model for the target cell; represents the prediction score of the student model for the i-th cell; Represents the prediction score of the student model for the target cell.
7. The lightweight human posture estimation method based on knowledge distillation according to claim 4 is characterized in that: The soft loss is expressed as: ; in, Indicates soft loss; represents the prediction score of the teacher model for the target cell; represents the prediction score of the student model for the target cell; t represents the target cell.
8. The lightweight human posture estimation method based on knowledge distillation according to claim 1, characterized in that: The total loss function of the student model is expressed as: ; in, represents the total loss function; Indicates the loss of channel-related information; represents the logistic distillation loss.
9. A lightweight human posture estimation device based on knowledge distillation, comprising: The dataset construction and labeling module is used to construct the initial dataset of human videos and label the dataset; Model building module, used to build teacher model and student model for human pose estimation; A distillation training module uses a labeled data set to train a teacher model, and uses the trained teacher model to guide the student model to obtain a trained student model; the guided training includes a feature distillation stage and a logic distillation stage; the feature distillation is based on the intermediate feature map of the student model and the pruned intermediate feature map of the teacher model; the logic distillation is based on the logic output of the teacher model and the student model; The process of obtaining the intermediate feature map of the pruned teacher model is as follows: The intermediate feature maps of the same batch in the teacher model are sliced by channel, and the average of the ranks of all intermediate feature maps in each channel is calculated, and the rank value of the feature map of the channel is represented by the average value; Sort the rank values of the feature maps of each layer, and divide the intermediate feature maps into a reserved set and a pruned set according to the rank sorting results. The reserved set stores high-rank feature maps, and the pruned set stores low-rank feature maps. The reserved set is the intermediate feature map of the teacher model after pruning. The human body posture estimation module is used to obtain human body video data and input the human body video data into the trained student model for human body posture estimation.
Citation Information
Patent Citations
Underground coal mine human body action recognition method suitable for edge terminal
CN116189299A
Personalized human body action recognition method based on knowledge distillation
CN116844225A
Human body motion state detection method and system
CN118298499A
Method and platform for pre-trained language model automatic compression based on multilevel knowledge distillation
US20220198276A1
Rank Distillation for Training Supervised Machine Learning Models
US20230206134A1
Cited By
Rotating machinery intelligent diagnosis method based on structured pruning and knowledge fusion distillation
CN121502525A