Video identification method based on inter-frame motion attention mechanism and frame-level self-distillation

Through the inter-frame motion attention mechanism and frame-level self-distillation method, the accuracy and robustness of video recognition are improved, the problem of insufficient frame-level feature extraction in the prior art is solved, and more efficient video recognition is achieved without increasing computing resources.

CN120375249APending Publication Date: 2025-07-25NORTHWEST ELECTROMECHANICAL ENG RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510340789.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing video recognition methods still have room for improvement in frame-level feature extraction and temporal feature extraction, especially the lack of attention to the dynamic changes between frames in the video, resulting in low recognition accuracy.

Method used

Using video recognition method based on inter-frame motion attention mechanism and frame-level self-distillation, a frame-level self-distillation framework is constructed, and the inter-frame motion attention mechanism module and residual module are connected, and training is combined with self-distillation error loss and classifier loss to improve the effectiveness of frame-level feature extraction.

Benefits of technology

Without increasing computing resources, the accuracy and robustness of video recognition are significantly improved, especially on large-scale video datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375249A_ABST
    Figure CN120375249A_ABST
Patent Text Reader

Abstract

The invention discloses a video recognition method based on an inter-frame motion attention mechanism and frame-level self-distillation, and the method comprises the steps: obtaining video data to be recognized, inputting the video data into a trained video recognition model, and obtaining a video recognition result, the video recognition model comprises a frame-level feature extractor based on an inter-frame motion attention mechanism module and a frame-level self-distillation method; wherein in the frame-level feature extractor, four inter-frame motion attention mechanism modules are sequentially arranged from front to back, and the inter-frame motion attention mechanism modules are connected through residual error modules; and the output of the last residual module is subjected to feature extraction through a time feature extractor, and finally, prediction is executed through a classifier to obtain a final recognition result. According to the method, the model reasoning capability can be improved, and the feature expression is improved under the condition that computing resources are not increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of deep learning and computer vision, and relates to a video recognition method based on an inter-frame motion attention mechanism and frame-level self-distillation. Background Art

[0002] Installing cameras on ships to collect video information can be used for fault diagnosis prediction and health management of equipment, realizing the intelligence of equipment; video action recognition based on the collected video information is a necessary means for fault diagnosis prediction and health management. For video recognition, researchers mainly perform joint extraction of frame-level features and temporal features of videos. With the rapid development of deep learning, multi-cue and attention mechanisms have shown excellent performance in applications such as image classification, object detection, and semantic segmentation, and have gradually been applied to video recognition. In these methods, whether it is 2DCNN or 3DCNN, in the frame-level feature extraction stage, they mainly focus on static images in the video sequence. It should be noted that motion change is a significant clue for video recognition. Some existing solutions also perform video recognition by studying the idea of dynamic changes between frames. These solutions are more about improving the feature distribution in the global region for temporal features to enhance beneficial features, and only perform ordinary network layer sequential calculation processing for frame-level features, which still need to be improved. Summary of the Invention

[0003] The purpose of the present invention is to provide a video recognition method based on an inter-frame motion attention mechanism and frame-level self-distillation. By studying the dynamic changes between frames, a dynamic expression of image changes is obtained, and the distorted changes in the local motion area when generating action expressions are captured, thereby improving the video recognition accuracy.

[0004] To achieve the above task, the present invention adopts the following technical solutions:

[0005] A video recognition method based on an inter-frame motion attention mechanism and frame-level self-distillation, comprising:

[0006] Obtain the video data to be recognized, input the video data into a trained video recognition model, and obtain a video recognition result. Among them, the video recognition model includes a frame-level feature extractor based on an inter-frame motion attention mechanism module and a frame-level self-distillation method; in the frame-level feature extractor, four inter-frame motion attention mechanism modules are sequentially arranged from front to back, and each inter-frame motion attention mechanism module is connected through a residual module; the output of the last residual module is subjected to feature extraction by a temporal feature extractor, and finally a classifier is used to perform prediction to obtain a final recognition result;

[0007] The frame-level self-distillation method is as follows: construct a frame-level self-distillation framework, which includes a downsampling unit respectively set between the outputs of the second, third, and fourth inter-frame motion attention mechanism modules and the subsequent residual modules; use the feature maps output by the latter three residual modules as teacher features, and use the features obtained by processing the feature maps output by the latter three inter-frame motion attention mechanism modules through the downsampling units as student features. Calculate the self-distillation error loss for each corresponding pair of student and teacher features, and the weighted sum of all self-distillation error losses, together with the classifier loss and the VAE loss, constitute the final training loss function for joint training.

[0008] Furthermore, when the video recognition model is being trained, the preprocessing process of the video data in the dataset is as follows:

[0009] Perform random cropping on the video data, cropping each frame of video image in the video data to a random size; perform random flipping on the video data, specifically flipping each video image according to a preset flipping probability; perform temporal augmentation on the video data, randomly increasing or shortening the length of the video sequence within a preset range.

[0010] Furthermore, the processing process of the first inter-frame motion attention mechanism module is as follows:

[0011] The temporal length of the video data is T, and each frame of video image is input into the inter-frame motion attention mechanism module as a feature map, denoted as: where F input represents the input feature map sequence, f t represents the t-th feature map, t = 1, 2,..., T; H×W is the height and width of f t , and C is the number of channels of the feature map, represents the real number field;

[0012] First, the feature map sequence F input passes through a 3D convolution with a convolution kernel of N×1×1, and its calculation process is as follows:

[0013]

[0014] where f motor is the dynamic feature information after convolution calculation, and each element in f motor represents the result of weighted summation of the corresponding elements in the neighborhood feature maps; i represents the index of height H, j represents the index of width W, n represents the index of the defined convolution kernel size, m is the number of convolution kernels, w c (n + m) represents the (n + m)-th weight value in the weight vector w c corresponding to the c-th channel; and It means traversing its N neighboring feature maps centered on the feature map f t and performing calculations pixel by pixel;

[0015] The dynamic feature information f motor after the calculation passes through the activation function Relu to obtain the feature f m ' otor ;

[0016] The feature f m ' otor is then processed through three 3D convolutions + activation functions. Among them, the first two 3D convolutions are activated using the activation function Relu after processing, while the third 3D convolution is activated using the activation function Sigmod after processing. The activated result is the final dynamic information intensity distribution feature map

[0017] Then the feature map F out output by the inter-frame motion attention mechanism module is the result of multiplying this feature map input by the input feature map sequence F

[0018] Furthermore, the downsampling unit is a 2D convolution. By setting the channel and stride parameters, the channels of the feature map output by the previous inter-frame motion attention mechanism module are upsampled while the size of the feature map is downsampled to match the size of the feature map in the higher dimension.

[0019] Furthermore, the feature map F out1 output by the first inter-frame motion attention mechanism module is input to the second inter-frame motion attention mechanism module after being processed by the first residual module. The feature map F out2 output by the second inter-frame motion attention mechanism module is used as the first frame-level feature. The first frame-level feature is processed by the downsampling unit to obtain the first student feature, and the first frame-level feature is processed by the second residual module to obtain the first teacher feature. The first student feature and the first teacher feature are used as the input of the frame-level self-distillation module to obtain the first self-distillation error loss;

[0020] The feature map F out3 output by passing the first teacher feature through the third inter-frame motion attention mechanism module is used as the second frame-level feature. The second frame-level feature passes through the third residual module to obtain the second teacher feature. The second frame-level feature is processed by the downsampling unit to obtain the second student feature. The second student feature and the second teacher feature are used as the input of the frame-level self-distillation module to obtain the second self-distillation error loss;

[0021] The feature map F output by passing the second teacher feature through the third inter-frame motion attention mechanism moduleout4 As the third frame-level feature, the third frame-level feature is passed through the fourth residual module to obtain the third teacher feature. After the third frame-level feature is processed by the downsampling unit, the third student feature is obtained. The third student feature and the third teacher feature are used as the inputs of the frame-level self-distillation module to obtain the third self-distillation error loss.

[0022] Furthermore, the third frame-level feature is subjected to feature extraction by a time feature extractor composed of a one-dimensional CNN + BiLSTM network, and finally a CTC classifier is used to perform prediction to obtain the final recognition result.

[0023] Furthermore, when performing self-distillation training, the MSE loss function is used as the self-distillation loss function for calculating the self-distillation error loss, and its expression is as follows:

[0024]

[0025] where n is the number of features, y i is the i-th teacher feature, and y i ' is the i-th student feature;

[0026] The three self-distillation error losses are weighted and summed to obtain the total self-distillation loss loss MSE , as follows:

[0027]

[0028] where loss MSE is the total self-distillation loss, is the n-th self-distillation error loss, and α, β, and λ are weights, which assign different weights to each self-distillation error loss.

[0029] Furthermore, when the video recognition model is trained, the total self-distillation loss loss MSE is combined with the CTC classification loss and the VAE loss of the CTC classifier to form the final training loss function for joint training.

[0030] A terminal device includes a processor, a memory, and a computer program stored in the memory; when the processor executes the computer program, the simulation analysis method of the high-speed parachuting aircraft based on the parachuting model is implemented.

[0031] A computer-readable storage medium stores a computer program; when the computer program is executed by a processor, the simulation analysis method of the high-speed parachuting aircraft based on the parachuting model is implemented.

[0032] Compared with the prior art, the present invention has the following technical features:

[0033] 1. A novel inter-frame motion attention mechanism is proposed, which improves the model inference ability by paying more attention to the dynamic changes in consecutive video frames.

[0034] 2. A frame-level self-distillation method is proposed, which increases the feature expression ability without increasing computational resources by performing teacher-guided student operations on the frame-level features of consecutive video frames.

[0035] 3. A video recognition network model based on the inter-frame motion attention mechanism and the frame-level self-distillation method is proposed, which achieves new state-of-the-art accuracy on three large-scale video datasets, RWTH, RWTH-T, and CSL-Daily, under the condition of only RGB input. Brief Description of the Drawings

[0036] Figure 1 It is the overall network structure diagram of the present invention;

[0037] Figure 2 It is the network architecture diagram of the inter-frame motion attention mechanism module;

[0038] Figure 3 It is the experimental result curve diagram of the present invention on the RWTH dataset;

[0039] Figure 4 It is the experimental result curve diagram of the present invention on the RWTH-T dataset;

[0040] Figure 5 It is the experimental result curve diagram of the present invention on the CSL-Daily dataset;

[0041] Figure 6 Visualization heat map of the inter-frame motion attention mechanism module. Detailed Description of the Invention

[0042] With the development of deep learning and the wide use of CNN, video recognition has also made great progress. Video recognition models composed of networks combining CNN with traditional machine learning algorithms, such as CNN+HMM, CNN+LSTM+HMM, and networks combining CNN with various neural networks, such as CNN+RNN, CNN+LSTM, and CNN+BiLSTM, have laid a solid and classic foundation in video recognition research while achieving success. Researchers have used various dimensions of convolution in their development. It must be mentioned here that the common problem of all convolutional neural networks is that they only focus on the local features of the entire image. This solution takes a different approach according to the characteristics of videos, pays more attention to the changing parts in each frame of the image to obtain the dynamic changes between consecutive frames, and then recognizes the video.

[0043] The method of the present invention enables the model to learn the motion change regions worthy of attention end-to-end and self-learn the connection between consecutive frames without relying on additional cues, such as pre-extracted key points or multiple streams, which requires more computation to utilize this information. This solution aims to obtain an efficient model for video recognition. Additionally, different from the ordinary processing of frame-level features in other methods, the present invention also proposes a frame-level feature self-distillation method, which applies the self-distillation method to the extraction of frame-level features of videos, improves the feature expression without increasing computational resources, and further improves the model performance and robustness.

[0044] The present invention provides a video recognition method based on an inter-frame motion attention mechanism and frame-level self-distillation. By capturing the distortion changes in the local motion regions when actions are expressed in the video, a dynamic expression of the image changes is obtained. And the self-distillation method is applied to frame-level feature extraction. By performing self-distillation on the features of adjacent stages and using high-order features as teachers to guide low-order features, the feature expression is improved without increasing computational resources. The combination of the two constitutes an overall video recognition model, improving the inference ability and robustness of the model.

[0045] See the appendix Figure 1 , a video recognition method based on an inter-frame motion attention mechanism and frame-level self-distillation provided by the present invention includes:

[0046] Obtain the video data to be recognized, input the video data into the trained video recognition model, and obtain the video recognition result. Among them, the video recognition model includes a frame-level feature extractor based on an inter-frame motion attention mechanism module and a frame-level self-distillation method; in the frame-level feature extractor, four inter-frame motion attention mechanism modules are sequentially arranged from front to back, and each inter-frame motion attention mechanism module is connected by a residual module; the output of the last residual module is subjected to feature extraction by a temporal feature extractor, and finally a classifier is used to perform prediction to obtain the final recognition result;

[0047] The frame-level self-distillation method is as follows: construct a frame-level self-distillation framework, and the frame-level self-distillation framework includes a downsampling unit respectively arranged between the outputs of the second, third, and fourth inter-frame motion attention mechanism modules and the subsequent residual modules; the feature maps output by the last three residual modules are used as teacher features, and the features obtained by processing the feature maps output by the last three inter-frame motion attention mechanism modules through the downsampling unit are used as student features. The self-distillation error loss is calculated for each corresponding student feature and teacher feature, and the weighted sum of all self-distillation error losses, the classifier loss, and the VAE loss together constitute the final training loss function for joint training.

[0048] Among them, the training process of the video recognition model is described as follows:

[0049] 1. Preprocessing of the dataset.

[0050] Preprocess each video data in the dataset. Specifically, use some data augmentation methods to improve the generalization of the network model, including:

[0051] Perform random cropping on the video data, and crop each frame of the video image in the video data to a random size. For example, the original size is 256×256, and the cropped size is 224×224.

[0052] Perform random flipping on the video data. Specifically, flip each video image according to a preset flipping probability; the preset flipping probability is 0.5.

[0053] Perform temporal augmentation on the video data, and randomly increase or decrease the length of the video sequence within ±20%.

[0054] 2. Video recognition model; the backbone of the video recognition model includes a frame-level feature extractor based on an inter-frame motion attention mechanism module and frame-level self-distillation method, a temporal feature extractor composed of 1D CNN + BiLSTM, and a CTC classifier to perform predictions, using the VAC model as the basic model.

[0055] 2.1 Inter-frame motion attention mechanism module.

[0056] As Figure 1 shown, four inter-frame motion attention mechanism modules are sequentially set from front to back, and each inter-frame motion attention mechanism module is connected by a residual module; this residual module uses the residual block (Residual Blocks) in the ResNet34 network, which will not be elaborated here; the feature map output by the first inter-frame motion attention mechanism module enters the next inter-frame motion attention mechanism module after being processed by the residual module, and so on.

[0057] Among them, the processing process of the first inter-frame motion attention mechanism module is as follows, and the processing process of each subsequent inter-frame motion attention mechanism module is the same as it:

[0058] The inter-frame motion attention mechanism module uses multi-layer three-dimensional convolution to perform weighted summation on the corresponding pixels of the feature maps of adjacent image frames of the video data, and multiplies the summation result after normalization with the original feature map to enhance the dynamic motion information; the details are as Figure 2 shown, specifically as follows:

[0059] In the training stage, the temporal length of the preprocessed video data is T, and each frame of the video image is input into the inter-frame motion attention mechanism module as a feature map, expressed as: where F input represents the input feature map sequence, ft represents the t-th feature map, where t = 1, 2, ..., T; H×W is the height and width of f t and C is the number of channels of the feature map, denotes the real number field.

[0060] First, the feature map sequence F input passes through a 3D convolution with a convolution kernel of N×1×1, and its calculation process is as follows:

[0061]

[0062] where f motor is the dynamic feature information after convolution calculation, and each element in f motor represents the result of weighted summation of the corresponding elements in the neighborhood feature maps; i represents the index of height H, j represents the index of width W, n is the index of the defined convolution kernel size, m is the number of convolution kernels, w c (n + m) represents the (n + m)-th weight value in the weight vector w c corresponding to the c-th channel; and denotes traversing its N neighborhood feature maps centered on the feature map f t and performing calculations pixel by pixel; the dynamic feature information f motor obtained after calculation passes through the activation function Relu to get the feature f m ', otor and its expression is as follows:

[0063]

[0064] where C' is the size of the number of channels after feature activation.

[0065] As Figure 2 shown, the feature f m ' otor is further processed through three 3D convolutions + activation functions. Among them, the first two 3D convolution processes are activated using the activation function Relu, and after the third 3D convolution process, in order to convert the dynamic information into an intensity distribution, the feature values are constrained to the range of 0 - 1, and then the activation function Sigmod is used for activation. The result after activation is the final dynamic information intensity distribution feature map At this time, the number of channels is restored to the input size C; then the feature map F out output by the inter-frame motion attention mechanism module is the result of multiplying this feature map input by the input feature map sequence F, as shown below:

[0066]

[0067] 2.2 Frame-level self-distillation method.

[0068] For frame-level features, this solution proposes a frame-level self-distillation technique during training and constructs a frame-level self-distillation framework:

[0069] The frame-level self-distillation framework includes a downsampling unit respectively set between the outputs of the second, third, and fourth inter-frame motion attention mechanism modules and the subsequent residual modules; the downsampling unit is a 2D convolution. By setting the channel and stride parameters, the channels of the feature map output by the previous inter-frame motion attention mechanism module are upsampled while the size of the feature map is downsampled to make the size of its feature map match that of the higher-dimensional feature map.

[0070] The feature map output by each inter-frame motion attention mechanism module is used as the teacher feature, and the feature after being processed by the downsampling unit is called the student feature; where:

[0071] The feature map F output by the first inter-frame motion attention mechanism module out1 The feature map after being processed by the first residual module is input to the second inter-frame motion attention mechanism module, and the feature map F output by the second inter-frame motion attention mechanism module out2 As the first frame-level feature, the first frame-level feature is processed by the downsampling unit to obtain the first student feature, and the first frame-level feature is processed by the second residual module to obtain the first teacher feature. The first student feature and the first teacher feature are used as the inputs of the frame-level self-distillation module to obtain the first self-distillation error loss;

[0072] The first teacher feature is passed through the third inter-frame motion attention mechanism module to output the feature map F out3 As the second frame-level feature, the second frame-level feature is processed by the third residual module to obtain the second teacher feature, and the second frame-level feature is processed by the downsampling unit to obtain the second student feature. The second student feature and the second teacher feature are used as the inputs of the frame-level self-distillation module to obtain the second self-distillation error loss;

[0073] The second teacher feature is passed through the third inter-frame motion attention mechanism module to output the feature map F out4 As the third frame-level feature, the third frame-level feature is processed by the fourth residual module to obtain the third teacher feature, and the third frame-level feature is processed by the downsampling unit to obtain the third student feature. The third student feature and the third teacher feature are used as the inputs of the frame-level self-distillation module to obtain the third self-distillation error loss.

[0074] The third frame-level feature is extracted by a time feature extractor composed of a one-dimensional CNN + BiLSTM network, and finally a CTC classifier is used to perform prediction to obtain the final recognition result.

[0075] When performing self-distillation training, the MSE loss function is used as the self-distillation loss function for calculating the three self-distillation error losses, and its expression is as follows:

[0076]

[0077] where n is the number of features, y i is the i-th teacher feature, and y i ' is the i-th student feature.

[0078] The three self-distillation error losses are weighted and summed to obtain the self-distillation total loss loss MSE , as follows:

[0079]

[0080] where loss MSE is the self-distillation total loss, is the n-th self-distillation error loss, and α, β, λ are weights, which assign different weights to each self-distillation error loss.

[0081] When training the video recognition model, the self-distillation total loss loss MSE and the sum of the CTC classification loss and VAE loss of the CTC classifier constitute the final training loss function for joint training:

[0082] The Adam optimizer is used to train the video recognition model on the dataset. The initial learning rate and weight factor are set to 10 -4 , the batch size used is 2; there are 50 epochs in the training phase, and the learning rate is reduced by 80% at the 30th and 40th epochs; the beam search algorithm is used for decoding in the final CTC classifier stage, and the beam width is 10.

[0083] The method proposed in the present invention achieves a new state-of-the-art accuracy for continuous sign language recognition when applied to three large-scale video datasets RWTH, RWTH-T, and CSL-Daily. The experimental result curve graphs are as shown in Figure 3 , 4 , 5; Figure 6 shows the heatmap after the visualization experiment of the inter-frame motion attention mechanism module proposed in the present invention, so as to more directly show the value of the method proposed in the present invention. It can be seen that the dynamic attention module can enable the model to focus on the motion area of the image during the recognition process to obtain a dynamic feature expression that is helpful for reasoning.

[0084] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A video recognition method based on an inter-frame motion attention mechanism and frame-level self-distillation, characterized in that Including: Obtain video data to be recognized, input the video data into a trained video recognition model to obtain a video recognition result, wherein the video recognition model includes a frame-level feature extractor based on an inter-frame motion attention mechanism module and a frame-level self-distillation method; Among them, in the frame-level feature extractor, four inter-frame motion attention mechanism modules are sequentially arranged from front to back, and each inter-frame motion attention mechanism module is connected through a residual module; the output of the last residual module is subjected to feature extraction by a temporal feature extractor, and finally a classifier is used to perform prediction to obtain a final recognition result; The frame-level self-distillation method is as follows: construct a frame-level self-distillation framework, and the frame-level self-distillation framework includes respectively setting a downsampling unit between the outputs of the second, third, and fourth inter-frame motion attention mechanism modules and the subsequent residual modules; the feature maps output by the latter three residual modules are used as teacher features, and the features obtained by processing the feature maps output by the latter three inter-frame motion attention mechanism modules through the downsampling unit are used as student features. The self-distillation error loss is calculated for each corresponding student feature and teacher feature, and the weighted sum of all self-distillation error losses, the classifier loss, and the VAE loss constitute the final training loss function for joint training.

2. The video recognition method based on an inter-frame motion attention mechanism and frame-level self-distillation according to claim 1, wherein When the video recognition model is being trained, the preprocessing process of the video data in the dataset is as follows: Perform a random cropping operation on the video data, and crop each frame of video image in the video data to a random size; perform a random flipping operation on the video data, specifically flipping each video image according to a preset flipping probability; perform a temporal enhancement operation on the video data, and randomly increase or decrease the length of the video sequence within a preset range.

3. The video recognition method based on the inter-frame motion attention mechanism and frame-level self-distillation according to claim 1, wherein The processing process of the first inter-frame motion attention mechanism module is as follows: The temporal length of the video data is T, and each frame of the video image is input into the inter-frame motion attention mechanism module as a feature map, which is expressed as: Among them, F input represents the input feature map sequence, and f t represents the t-th feature map, where t = 1, 2,..., T; H×W is the height and width of f t , and C is the number of channels of the feature map, represents the real number field; First, the feature map sequence F input Passes through a 3D convolution with a convolution kernel of N×1×1, and its calculation process is as follows: where f motor is the dynamic feature information after convolution calculation, and each element in f motor represents the result of weighted summation of the corresponding elements in the neighborhood feature map; i represents the index of height H, j represents the index of width W, n represents the index of the defined convolution kernel size, m is the number of convolution kernels, w c (n + m) represents the (n + m)-th weight in the weight vector w c corresponding to the c-th channel; and means traversing its N neighborhood feature maps centered on the feature map f t and performing calculations pixel by pixel; The dynamic feature information f after calculation motor The feature f is obtained after passing through the activation function Relu m ' otor ; Feature f m ' otor It is further processed by three 3D convolutions + activation functions. Among them, the first two 3D convolutions are activated by the activation function Relu, and the third 3D convolution is activated by the activation function Sigmod. The activated result is the final dynamic information intensity distribution feature map Then the feature map F output by the inter-frame motion attention mechanism module out is this feature map multiplied by the input feature map sequence F input as a result.

4. The video recognition method based on the inter-frame motion attention mechanism and frame-level self-distillation according to claim 1, wherein, The downsampling unit is a 2D convolution. By setting the channel and stride parameters, the channels of the feature map output by the previous inter-frame motion attention mechanism module are upsampled while the size of the feature map is downsampled, so that the size of the feature map matches the size of the feature map in the higher dimension.

5. The video recognition method based on an inter-frame motion attention mechanism and frame-level self-distillation according to claim 1, wherein The feature map F output by the first inter-frame motion attention mechanism module out1 The feature map processed by the first residual module is input into the second inter-frame motion attention mechanism module, and the feature map F output by the second inter-frame motion attention mechanism module out2 As the first frame-level feature, the first frame-level feature is processed by the downsampling unit to obtain the first student feature, and the first frame-level feature is processed by the second residual module to obtain the first teacher feature. The first student feature and the first teacher feature are used as the input of the frame-level self-distillation module to obtain the first self-distillation error loss; Take the feature map F output by passing the first teacher feature through the third inter-frame motion attention mechanism module out3 As the second frame-level feature, pass the second frame-level feature through the third residual module to obtain the second teacher feature, pass the second frame-level feature through the downsampling unit to obtain the second student feature, and use the second student feature and the second teacher feature as the input of the frame-level self-distillation module to obtain the second self-distillation error loss; The feature map F output by passing the second teacher feature through the third inter-frame motion attention mechanism module out4 As the third frame-level feature, pass the third frame-level feature through the fourth residual module to obtain the third teacher feature, pass the third frame-level feature through the downsampling unit to obtain the third student feature, and use the third student feature and the third teacher feature as the input of the frame-level self-distillation module to obtain the third self-distillation error loss.

6. The video recognition method based on the inter-frame motion attention mechanism and frame-level self-distillation according to claim 1, characterized in that, The third frame-level feature is subjected to feature extraction by a temporal feature extractor composed of a one-dimensional CNN+BiLSTM network, and finally a CTC classifier is used to perform prediction to obtain a final recognition result.

7. The video recognition method based on the inter-frame motion attention mechanism and frame-level self-distillation according to claim 1, wherein When performing self-distillation training, the MSE loss function is used as the self-distillation loss function for calculating the self-distillation error loss, and its expression is as follows: where n is the number of features, and y i is the i-th teacher feature, and y i ' is the i-th student feature; The three self-distillation error losses are weighted and summed to obtain the total self-distillation loss loss MSE , as follows: where loss MSE is the total self-distillation loss, and loss MSEn is the n-th self-distillation error loss. α, β, and λ are weights, which assign different weights to each self-distillation error loss.

8. The video recognition method based on the inter-frame motion attention mechanism and frame-level self-distillation according to claim 7, characterized in that, When the video recognition model is trained, the total self-distillation loss loss MSE Combined with the CTC classification loss of the CTC classifier and the sum of the VAE losses, it constitutes the final training loss function for joint training.

9. A terminal device, comprising a processor, a memory, and a computer program stored in the memory; characterized in that, When the processor executes the computer program, it implements the simulation analysis method of the high-speed parachuting aircraft based on the parachute descent model according to any one of claims 1-8.

10. A computer-readable storage medium, in which a computer program is stored; when the computer program is executed by a processor, it implements the simulation analysis method of the high-speed parachuting aircraft based on the parachute descent model according to any one of claims 1-8.