A 3D Model Compression Method Based on Spatiotemporal Information Transfer Knowledge Distillation Technology

Through the multi-layer feature distillation module using spatiotemporal information transfer knowledge distillation technology in the three-dimensional behavior recognition model, the problems of high model complexity and low training efficiency are solved, and the compression and recognition accuracy of the model are improved.

CN115035595BActive Publication Date: 2025-06-20NORTHWEST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210624609.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-02
Publication Date
2025-06-20
Estimated Expiration
2042-06-02

AI Technical Summary

Technical Problem

The prior art is difficult to effectively compress complex three-dimensional behavior recognition models, resulting in low training efficiency and high spatial complexity, making it difficult to deploy on devices with limited computing resources.

Method used

Using knowledge distillation technology based on spatiotemporal information transfer, the 3D model is compressed through a multi-layer feature distillation module, and multi-layer context loss control feature transmission is used to achieve lightweighting of the model.

Benefits of technology

It realizes effective compression of 3D models, improves training efficiency and recognition accuracy, promotes the lightweight development of the model, and accelerates the application of behavior recognition algorithms in practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115035595B_ABST
    Figure CN115035595B_ABST
Patent Text Reader

Abstract

The present invention discloses a 3D model compression method based on spatio-temporal information transfer knowledge distillation technology. This method uses a multi-layer feature distillation module (MFDM) as the core component of the 3D model to compress the model of the action recognition network. A publicly available dataset containing multiple action classes and multiple video segments is used as the experimental dataset. From each action video segment in the dataset, a part of the frames are taken at equal intervals and used as the input of the teacher model and the student model in the spatio-temporal feature distillation method for feature extraction; the features of each layer are extracted, a multi-level feature distillation algorithm is performed on the features of each layer, and the multi-layer spatio-temporal feature transfer loss is calculated; the features of the last layer are put into the classifier for classification to obtain the logistic regression probability, and the loss functions are calculated respectively with the soft label and the true label generated by the teacher model; finally, according to the overall loss function, all parameters of the student model are updated through backpropagation, and at the same time, the parameters of the multi-layer feature distillation module are updated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of model compression, and specifically relates to a 3D model compression method based on spatio-temporal information transfer knowledge distillation technology. Background Art

[0002] As a highly popular direction in computer vision, most models for action recognition are targeted at video data. If a higher recognition rate is desired, the model complexity will be relatively high, and thus its training efficiency will be reduced. This is also a common problem in the field of deep learning. Since the algorithms developed in the laboratory mostly pursue higher accuracy, if these problems occur in the laboratory, the training efficiency and model complexity do not seem to be that important. However, in practical applications, when the model needs to be deployed to a device, the computing resources and storage resources of the device are limited. For example, in applications such as autonomous driving and intelligent robots, there are more stringent requirements for the accuracy and computing resources of the model. Therefore, action recognition models with high complexity have problems of low training efficiency and high space complexity, and these two problems are the main issues to be considered.

[0003] The algorithm model based on video is more complex compared with the algorithm model based on image, because video has temporality in addition to pictures, and under the premise of the same resolution, the data volume of video is much higher than that of pictures. The existing knowledge distillation algorithms, including KD and FITNETS, are all for two-dimensional convolutional feature extraction networks such as VGG and ResNet. However, the number of parameters and the complexity of modules in three-dimensional feature extraction network models are higher than those of two-dimensional feature extraction networks, and thus more compression is needed. At the same time, three-dimensional feature extraction networks are not only widely used in action recognition, but also, in the latest three-dimensional object detection and tracking, when the input data is point cloud, using three-dimensional convolution for feature extraction has become a very common method. The most widely used scenario for three-dimensional object detection is the field of autonomous driving. For practical applications, models with small size and high accuracy will be more easily applied to actual products. Summary of the Invention

[0004] Aiming at the deficiencies existing in the above prior art, the purpose of the present invention is to provide a 3D model compression method based on spatio-temporal information transfer knowledge distillation technology.

[0005] To achieve the above task, the present invention adopts the following technical solutions:

[0006] A 3D model compression method based on spatio-temporal information transfer knowledge distillation technology, characterized in that the method uses a multi-layer feature distillation module MFDM as the core component of the 3D model to compress the model of the action recognition network, and specifically includes the following steps:

[0007] S1: Use a publicly available dataset containing various action classes and multiple video segments as the experimental dataset. Extract partial frames at equal intervals from each behavioral video segment in the dataset and use them as the inputs for the teacher model and the student model in the spatio-temporal feature distillation method to extract features.

[0008] S2: Extract the features of each layer, perform a multi-level feature distillation algorithm on the features of each layer, and calculate the multi-level spatio-temporal feature transfer loss.

[0009] S3: Put the features of the last layer into the classifier for classification to obtain the logistic regression probability. Calculate the loss functions with the soft labels and the true labels generated by the teacher model respectively.

[0010] S4: Finally, according to the overall loss function, update all the parameters of the student model through backpropagation, and at the same time update the parameters of the multi-level feature distillation module.

[0011] Specifically, use 3D Resnet-18, 3D Resnet-34, and 3D Resnet-50 as the backbones of the student network to extract features; use 3D Resnet-34, 3D Resnet-50, and 3D Resnet-101 as the backbones of the teacher model to extract features.

[0012] Furthermore, the multi-level feature distillation module takes into account both the transfer of spatial features and the transfer of temporal features, and uses multi-level context loss to control the feature transfer, which is called multi-level spatio-temporal feature transfer loss.

[0013] The 3D model compression method based on spatio-temporal information transfer knowledge distillation technology of the present invention compresses the model of the behavior recognition network. A multi-level feature distillation module is used, which mainly divides the features in time and space and compares them respectively to ensure that the output features of each layer of the student model are close to the output features of the corresponding layer of the teacher model as much as possible to ensure better feature transfer. The spatio-temporal feature transfer loss function is used to update the parameters of the student model through backpropagation with the loss function of the soft label and the loss function of the classifier. It can not only achieve model compression, but also improve the training efficiency and recognition accuracy of the model. It promotes the trend of the model towards lightweight development and speeds up the application of the behavior recognition algorithm in real life. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 It is a schematic diagram of the 3D network structure adopted by the 3D model compression method of the spatio-temporal information transfer knowledge distillation technology of the present invention.

[0015] Figure 2 It is a schematic diagram of the multi-level feature distillation module structure.

[0016] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. Specific Embodiments

[0017] The design concept of this application is that behavior recognition, as a highly popular direction in computer vision, most of its models are targeted at video data. If a higher recognition rate is desired, the model complexity will be relatively high, and thus its training efficiency will also decrease. This is also a common problem in the field of deep learning. In actual applications, when the model needs to be deployed to a device, the computing resources and storage resources of the device are limited. For example, in applications such as autonomous driving and intelligent robots, there are more stringent requirements for the accuracy and computing resources of the model. The behavior recognition model with high complexity has problems of low training efficiency and high space complexity.

[0018] Secondly, since video has one more time dimension than an image, therefore, when compressing a model for processing video, not only the learning of spatial features needs to be considered, but also the retention of temporal features needs to be taken into account. For the behavior recognition model, this method uses a three-dimensional convolutional network as the feature extraction network. Different from previous knowledge distillation, when the data is an image, only the transfer of spatial features is considered, but for video, this module needs to consider both the transfer of spatial features and the transfer of temporal features. For the above reasons, this application adopts a network framework of a multi-layer feature distillation module.

[0019] However, in the knowledge distillation of behavior recognition, extracting the temporal feature extraction ability of the teacher model is required for information transfer, which has the same effect as obtaining the temporal features of the video in behavior recognition.

[0020] Therefore, drawing on the idea of the temporal excitation and aggregation mechanism of temporal excitation and aggregation (TEA), the behavior recognition network is compressed by model compression, specifically using the method of knowledge distillation, which promotes the trend of the model towards lightweight development and speeds up the application of the behavior recognition algorithm in real life.

[0021] This application uses the features output by this layer of the teacher model as a standard to compare with the features output by the corresponding layer of the student model for training. When the difference between the output features is smaller, it means that the student model is closer to the feature extraction ability of the teacher model at this layer at this layer, thereby realizing the transfer of spatio-temporal information of different scales.

[0022] This embodiment provides a 3D model compression method based on spatio-temporal information transfer knowledge distillation technology, which uses a multi-layer feature distillation module MFDM as the core component of the 3D model to compress the behavior recognition network. The specific steps are as follows:

[0023] S1: Use the publicly available UCF101 dataset that contains various action classes and multiple video segments as the experimental dataset. Take partial frames at equal intervals from each video segment in the dataset as the inputs for the teacher model and the student model in the spatio-temporal feature distillation method.

[0024] S2: Extract the features of each layer, perform a multi-level feature distillation algorithm on the features of each layer, and calculate the multi-layer spatio-temporal feature transfer loss.

[0025] S3: Put the features of the last layer into a classifier for classification to obtain the logistic regression probability. Calculate the loss functions respectively with the soft labels and the true labels generated by the teacher model.

[0026] S4: Finally, according to the overall loss function, update all the parameters of the student model through backpropagation, and at the same time update the parameters of the multi-layer feature distillation module.

[0027] In this embodiment, the multi-layer feature distillation module (MFDM, Multilayer fearture distillation module) should consider both the transfer of spatial features and the transfer of temporal features, and uses a multi-layer context loss to control the transfer of features. This loss is called the multi-layer spatio-temporal feature transfer loss (STTL, Spatiotemporal feature transfer loss). The schematic diagram of the structure of the multi-layer feature distillation module is as Figure 2 shown.

[0028] The following is the specific implementation process:

[0029] S1: Obtain the experimental dataset and preprocess the data.

[0030] Use the publicly available UCF101 dataset as the experimental dataset. This dataset UCF101 focuses on human actions in videos. It consists of 101 action classes, with a total video duration of about 27 hours, specifically including 13,320 action-class videos. All the data are real videos uploaded by various users. Therefore, some data will be affected by camera movement and cluttered backgrounds, mainly including: human-computer interaction, body movement, interaction between people, playing musical instruments, and sports, a total of 5 categories.

[0031] S1.1: First, perform frame sampling on the video segments. Since different videos have different lengths, the method uses specific frames as the default length of the entire video segment, and uses the specific video frames to represent the entire video as the input frames for the network.

[0032] S1.2: The number of behavior categories in the training set video is 101. The official training set consists of 9,537 videos. In this experiment, it is re-divided. 9,337 videos are divided into the training set and 200 videos are divided into the validation set. The test set uses the official split01 as the test set, and the number of videos is 3,783.

[0033] S2: Extract the features of each layer, perform a multi-level feature distillation algorithm on the features of each layer, and calculate the multi-level spatio-temporal feature transfer loss.

[0034] S2.1: Conv in the three-dimensional spatio-temporal feature learning module represents the use of convolutional operations, and each feature map applies convolutions with a specific stride. Here, a set of hyperparameters α, β, γ, μ are set, and the feature map is transformed into αT×βH×γW,μC through three-dimensional convolution. The convolution size of the i-th layer of the student model is set to W×H×T, C. The output features of the student model and the teacher model are respectively transformed into αT×βH×γW,μC through convolutions with a specific stride. The schematic of the 3D network model is as Figure 1 shown.

[0035] S2.2: Respectively use 3D Resnet-18, 3D Resnet-34, 3D Resnet-50 as the backbone of the student network to extract features. The subscript s represents the student model. Input the video data X into the student model S, and let Y s = S(X), where Y s represents the feature map of the student model output by the backbone network. Since the entire network S can be divided into different parts, where S i represents the i-th layer network. Therefore, the output feature Y s of the backbone network can also be represented as Y s = S n …Δ…S i Δ…S 2 ΔS 1 (X). Among them, Y s represents the output feature map of the student model, and S i ΔS i-1 represents S i (S i-1 (X)), which means taking the feature map output by the i-1 layer as the input of the i-th layer to obtain the output feature map of S i .

[0036] S2.3: Respectively use 3D Resnet-34, 3D Resnet-50, 3D Resnet-101 as the backbone of the teacher model to extract features. For the teacher model, it is the same as the student model. The subscript t represents the teacher model.

[0037] S2.4: Effectively transfer the feature extraction ability in the teacher model to the student model. Since knowledge transfer is required for the intermediate feature extraction layers, the output feature map of the i-th layer network S i can be taken out and set as F i . The output feature matrix of the intermediate layer of the entire backbone network can be set as F = (F 1 , F 2 , …, F i , …, F n ).

[0038] S2.5: Single-layer knowledge distillation can be expressed as: represents the feature map obtained after putting the feature map of the i-th layer of the student model into the same feature space, represents the feature map obtained after putting the feature map of the i-th layer of the teacher model into the same feature space. Since each layer of L consists of multiple L2s, D represents the distance function between two feature maps. Among them, M represents putting the features into the same feature space. And the multi-layer knowledge distillation adopted by this method can be expressed as: where I represents the set of layers that need to be distilled, corresponding to the loss function of the entire multi-layer feature distillation module.

[0039] S3: Put the features of the last layer into the classifier for classification to obtain the logistic regression probability, and calculate the loss functions with the soft label and the true label generated by the teacher model respectively;

[0040] S3.1: In a neural network, generally softmax is often used as the output layer to output the probability z i of each class, and then by comparing with the sum of other logistic regression probability values, the calculated logistic regression probability of each class can be converted into the probability q i , then:

[0041]

[0042] where T is a temperature usually set to 1. Using a higher value for T will result in a weaker probability distribution over classes.

[0043] S3.2: The output after the fully connected layer is Z s = FCL(Y s ), Z s represents the logistic regression probability value output by the backbone network after passing through the fully connected layer.

[0044] S3.3: The network model also combines the "soft label" loss and the "hard label" loss. The soft label loss function is the cross-entropy with the soft label, which can be abbreviated as the "soft target", and the other objective function is the cross-entropy with the correct label, which can be called the "hard target". To obtain better results when using the "soft target" while using the "hard target", it is necessary to reduce the weight of the second objective function, and it is also necessary to use the scores of the classes that the data does not belong to as targets, which requires increasing the weight of the "soft target".

[0045] S4: Finally, according to the overall loss function, all parameters of the student model are updated through backpropagation, and at the same time, the parameters of the multi-layer feature distillation module are updated.

[0046] S4.1: Update all parameters of the student model through backpropagation.

[0047] S4.2: Update the parameters of the multi-layer feature distillation module.

[0048] S4.3: Obtain the accuracy rate after the model converges.

[0049] To verify the effectiveness of the 3D model compression method based on the spatio-temporal information transfer knowledge distillation technology in this embodiment, the applicant also conducted the following experiments and comparative analyses. Multiple comparative experiments were mainly carried out on the three-dimensional convolutional network, and the differences in the training efficiency and accuracy rate of the model trained by separate training and the spatio-temporal feature distillation method were analyzed. The results show that the spatio-temporal feature distillation method can not only achieve model compression, but also improve the training efficiency and recognition accuracy rate of the network model.

[0050] 1. The experimental environment setup is described as follows:

[0051] The experiment uses the Python language to implement the algorithm. The version of Python is 3.6. The specific code is written using the deep learning framework PyTorch. The important Python packages used in the experiment are: PyTorch version 1.1.0, ffmpeg version 1.4, mmcv version 1.3.17, opencv-python version 4.5, and tensorboard version 1.14.0. The experimental device has three graphics cards, and there are many models to be trained. One model is trained on a single GPU. The experimental device can train 3 models simultaneously. The batchsize of each model is the same, which is 64 videos. There are 9337 videos as the training set in one epoch, and the remaining 200 videos are used as the validation set. Generally, it takes 80 minutes to train one epoch. The total number of epochs for each model is set to 150. The optimizer is SGD, the learning rate is 0.01, the moment is 0.09, and the weight decay is 0.0005.

[0052] 2. Network Model Training

[0053] The main process of network model training is as follows:

[0054] (1) First, conduct ablation experiments: mainly analyze the multi-layer feature distillation module. Since the knowledge distillation scheme used in this method consists of three important loss functions, corresponding to the loss of hard labels, soft label loss, and STTL (spatiotemporal feature transfer loss), they will be trained separately in this section. The backbone network adopted by the model is 3D Resnet, which is also written as Res3D for simplicity.

[0055] (2) Study the information transfer ability of different teacher models to the same student model: Use 3DResnet-18, 3D Resnet-34, and 3D Resnet-50 as student models respectively, and use 3D Resnet-34, 3D Resnet-50, and 3D Resnet-101 as teacher models for experiments.

[0056] (3) Train the model: Compare the proposed model compression scheme with existing methods.

[0057] 3. Network Performance Evaluation

[0058] The evaluation metrics adopted are loss value and accuracy. If the prediction result is consistent with the true label, that is, the classification is correct, then the prediction is correct; otherwise, the prediction is incorrect.

[0059] The comparison with common behavior models in recent years is as follows:

[0060] Table 1: Ablation Experiment of Spatiotemporal Feature Distillation Method on UCF101 Dataset

[0061]

[0062]

[0063] As shown in Table 1, in the experiment, 3D Resnet-18 was used as the student model and 3D Resnet-50 was used as the teacher model. When the student model was trained alone until convergence, the accuracy rate was 84.4%. If only the teacher soft labels or only the spatio-temporal feature transfer loss was used to achieve knowledge distillation, it could be seen that there was an improvement in the accuracy rate compared to training alone. Therefore, it could be shown that the 3D model compression method based on spatio-temporal information transfer knowledge distillation technology of this embodiment (hereinafter referred to as the present design) could effectively transfer the feature extraction ability in the teacher model to the student model. It could be seen from the table that the accuracy rate of training using the spatio-temporal feature transfer loss designed in this paper was higher than that of training only using soft labels, indicating that the spatio-temporal feature transfer method designed in this paper had a stronger information transfer ability than using soft labels. Moreover, the classification accuracy rate of the student model obtained by the knowledge distillation method using both was the highest. Therefore, the effectiveness of the present design was demonstrated.

[0064] Table 2: Test results of different student models using different teacher models on UCF101

[0065]

[0066] It could be seen from Table 2 that as the network depth deepened, the time consumed for detecting a single video would increase. This was because as the network depth increased, the internal parameters would also increase, and the time consumed for the test data to be calculated through the network would also increase. Within a certain range, the deeper the network, the stronger the feature extraction ability for video data. At this time, its role as a teacher model in guiding the student model was more obvious. Therefore, more information was transmitted to the student model through the multi-layer feature distillation module, and the accuracy rate improvement was also greater.

[0067] Table 3: Comparison table of knowledge distillation algorithms

[0068]

[0069] In this experiment, the proposed model compression scheme was compared with existing methods. Since there was less research on applying knowledge distillation to action recognition, as shown in Table 3, for the improvement of the accuracy rate of the student model, it could be seen that the accuracy rate of the multi-level spatio-temporal distillation module designed in this paper was higher than that of STDDCN. The multi-layer feature distillation module designed in this paper not only targeted the feature map output at the end of the model, but also compared the feature map of each layer of the student model with the corresponding layer of the teacher model to obtain a spatio-temporal feature transfer loss, so as to approximate the feature map of each layer of the student model to the corresponding layer of the teacher model, demonstrating the superiority of the multi-level spatio-temporal distillation module proposed in this design.

[0070] 4. Conclusion

[0071] In this experiment, the 3D model compression method based on the spatio-temporal information transfer knowledge distillation technology of the embodiment was mainly compared with the existing 3D convolutional network in multiple comparative experiments, and the differences in the training efficiency and accuracy of the model were analyzed when training the model separately and using the spatio-temporal feature distillation method. The results show that the spatio-temporal feature distillation method can not only achieve model compression, but also improve the training efficiency and recognition accuracy of the model.

[0072] In summary, the 3D model compression method based on the spatio-temporal information transfer knowledge distillation technology given in this embodiment is characterized in that it compresses the model of the action recognition network, specifically adopts the spatio-temporal feature distillation method, and conducts multiple comparative experiments with the existing 3D convolutional network, and analyzes the differences in the training efficiency and accuracy of the model when training the model separately and using the spatio-temporal feature distillation method. Finally, it is shown that the spatio-temporal feature distillation method can not only achieve model compression, but also improve the training efficiency and recognition accuracy of the model.

[0073] Finally, it should be pointed out that the embodiments described above are only used to illustrate and understand the technical solutions of the present application, and the present invention is not limited to the above embodiments. Those of ordinary skill in the art should understand that in the technical solutions of the present application, simple modifications, substitutions or additions can be made to the technical features, and these simple modifications, substitutions or additions should belong to the protection scope of the present application.

Claims

1. A 3D model compression method based on spatio-temporal information transfer knowledge distillation technology, characterized in that This method uses a multi-layer feature distillation module MFDM as the core component of the 3D model to compress the model of the action recognition network, which specifically includes the following steps: S1: Use a public dataset containing multiple action classes and multiple video segments as the experimental dataset. Extract partial frames at equal intervals from each action video segment in the dataset and use them as the inputs of the teacher model and the student model in the spatio-temporal feature distillation method for feature extraction respectively; S2: Extract the features of each layer, perform a multi-layer feature distillation algorithm on the features of each layer, and calculate the multi-layer spatio-temporal feature transfer loss; S2.1: Conv in the three-dimensional spatio-temporal feature learning module represents the use of convolution operations, and a convolution with a specific stride is applied to each feature map; here, a set of hyperparameters α, β, γ, μ are set, and the feature map is transformed into αT×βH×γW,μC through three-dimensional convolution; set the convolution size of the i-th layer of the student model to W×H×T, C, and respectively transform the output features of the student model and the teacher model through convolutions with specific strides into αT×βH×γW, μC; S2.2: Use 3D Resnet-18, 3D Resnet-34, and 3D Resnet-50 as the backbone of the student network to extract features respectively; the subscript s represents the student model. Input the video data X into the student model S, and let Y s = S(X), where Y s represents the feature map of the student model output by the backbone network. Since the entire network S can be divided into different parts, where S i represents the i-th layer of the network. Therefore, the feature Y s output by the backbone network can also be represented as Y s = S n …Δ…S i Δ…S 2 ΔS 1 (X), where Y s represents the output feature map of the student model, and S i ΔS i-1 represents S i (S i-1 (X)), which means taking the feature map output by the (i - 1)-th layer as the input of the i-th layer to obtain the output feature map of S i . S2.3: Use 3D Resnet-34, 3D Resnet-50, and 3D Resnet-101 as the backbone of the teacher model to extract features; for the teacher model, it is the same as the student model, and the subscript t is used to represent the teacher model; S2.4: Transfer the feature extraction ability in the teacher model to the student model. Since knowledge transfer is required for the intermediate feature extraction layers, the output feature map of the i-th layer network S i can be taken out and set as F i . The output feature matrix of the intermediate layer of the entire backbone network is set as F = (F 1 , F 2 , …, F i , …, F n ); S2.5: Single-layer knowledge distillation can be expressed as: represents the feature map obtained after putting the feature maps of the i-th layer of the student model into the same feature space, represents the feature map obtained after putting the feature maps of the i-th layer of the teacher model into the same feature space; since L for each layer consists of multiple L2s, D represents the distance function between two feature maps; where M represents putting the features into the same feature space; multi-layer knowledge distillation is expressed as: where I represents the set of layers for which distillation is to be performed, corresponding to the loss function of the entire multi-layer feature distillation module; S3: Put the features of the last layer into the classifier for classification to obtain the logistic regression probability, and calculate the loss function by comparing it with the soft label and the true label generated by the teacher model respectively; S3.1: In a neural network, softmax is often used as the output layer to output the probability z of each class i , and then by adding it to the sum of other logistic regression probability values , the calculated logistic regression probability of each class can be converted into probability q i , then: Among them, T is a temperature usually set to 1. Using a higher value for T will result in a weaker probability distribution over classes; S3.2: The output after the fully connected layer is Z s = FCL(Y s ), where Z s represents the logistic regression probability value output by the backbone network after passing through the fully connected layer; S3.3: The loss function of the network model includes the soft label loss and the hard label loss. The soft label loss function is the cross-entropy with the soft label, and the hard label loss is the cross-entropy with the correct label; S4: Finally, according to the overall loss function, update all the parameters of the student model through backpropagation, and at the same time update the parameters of the multi-layer feature distillation module.

2. The method according to claim 1, characterized in that: Use 3D Resnet-18, 3D Resnet-34, and 3D Resnet-50 as the backbone of the student network to extract features respectively; Use 3D Resnet-34, 3D Resnet-50, and 3D Resnet-101 as the backbone of the teacher model to extract features respectively.

3. The model compression method according to claim 1, characterized in that The multi-layer feature distillation module takes into account both the transfer of spatial features and the transfer of temporal features, and uses a multi-layer context loss to control the transfer of features, which is called the multi-layer spatio-temporal feature transfer loss.

Citation Information

Patent Citations

  • Training data management method and related system

    US20130323687A1

  • Knowledge distillation-based compression method for pre-trained language model, and platform

    WO2021248868A1