An Aerial Video Classification Method Based on Spatiotemporal Multi-Scale Transformer
By adopting a space-time multi-scale Transformer method in aerial video classification, the limitations of traditional convolutional neural networks in processing aerial videos are solved, and higher recognition accuracy and lower computational complexity are achieved.
Patent Information
- Application Number
- CN202210844866.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-19
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-07-19
AI Technical Summary
Traditional convolutional neural networks have limitations when processing drone aerial videos, including difficulty in making full use of high perspectives and rich background information, the presence of induction bias, high computational cost and slow inference speed.
The aerial video classification method based on space-time multi-scale Transformer is adopted to perform feature extraction through the 2D Transformer network, and multi-scale information is introduced using the pooled multi-head self-attention module. The hollow time feature extraction module processes video data of any length, and improves the utilization of time information with the help of the feature offset module.
It improves the model's performance in recognition accuracy and computing efficiency, can better adapt to the characteristics of aerial video, enhances the utilization of time and space information, and reduces the computational complexity.
Smart Images

Figure CN115223082B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent analysis of remote sensing images, and specifically relates to an aerial video classification method based on spatio-temporal multi-scale Transformer. Background Art
[0002] With the rapid development of the drone industry and the field of computer vision, the massive video data with high quality, high resolution, and high flexibility collected by drones has greatly promoted the research of computer vision in aerial video analysis. At the same time, drones equipped with intelligent image analysis systems can complete various specific tasks, such as in industries such as agricultural and forestry plant protection, power inspection, aerial mapping, police security, and logistics transportation, and have extremely high practical value. The drones in our country are entering a new era of innovative leapfrog development, and the intelligence of drones is an important research direction in the future. However, there are the following difficulties in fully and accurately utilizing the imaging resources of drones with traditional convolutional neural networks:
[0003] (1) The imaging perspective of drones is high and the field of view is wide, and the background information is rich. Although convolutional neural networks can capture spatio-temporal information well, there are limitations in capturing global relationships within a local area.
[0004] (2) Convolutional neural networks have strong inductive biases, which are beneficial for training on small datasets, but when the data is sufficient, they limit the expression ability of the model for all examples.
[0005] (3) In the case of high-resolution and long-sequence videos, the computational cost of convolutional neural networks is high and the inference speed is slow.
[0006] Traditional video analysis and processing methods include three categories: one is the neural network method based on two streams, the second is the method based on 2D convolutional neural networks, and the third is the method based on 3D convolutional networks. The neural network based on two streams means that the input is a temporal stream and a spatial stream. The spatial stream processes single-frame images, and the temporal stream processes multi-frame optical flow images. However, optical flow cannot capture long-term temporal information, and the computational amount for extracting optical flow is huge, which limits its wide application in the industry. The method based on 2D convolutional neural networks is to introduce temporal information into spatial features by means of difference, feature transformation, multi-scale fusion, etc. while extracting spatial features using 2D convolution. Although the complexity of the 2D network can approach the accuracy of the 3D network, the limitations of convolutional neural networks are still not solved. The method based on 3D convolutional neural networks is to capture temporal features and spatial features from the extended temporal dimension using 3D convolution, and at the same time capture long-term temporal information by stacking 3D convolutions. However, the computational cost of 3D convolution is also huge, and it is difficult to deploy on devices. Summary of the Invention
[0007] To solve the problems existing in the above prior art, the present invention proposes an aerial video classification method based on spatio-temporal multi-scale Transformer. The method includes:
[0008] Obtain aerial video data and preprocess the aerial video data;
[0009] Input the preprocessed aerial video data into a trained aerial video recognition model based on multi-scale Transformer to output a recognition result;
[0010] The aerial video recognition model based on multi-scale Transformer includes using a 2D transformer network as the backbone network. The 2D transformer network includes a pre-coding module, a multi-scale spatio-temporal feature extraction module composed of a multi-level coding block structure, a dilated temporal feature extraction module ETM, and a classifier of a fully connected layer. Among them, each level of coding block structure includes multiple layers of coding blocks, and each layer of coding block includes two feature shift modules FS, a multi-layer perceptron MLP, and a pooled multi-head self-attention module PMHA or a standard multi-head self-attention module MHA. One feature shift module FS is located at the head of the layer of coding block, and the other feature shift module FS is inserted between the multi-layer perceptron MLP and the pooled multi-head self-attention module PMHA or the standard multi-head self-attention module MHA, and the pooled multi-head self-attention module PMHA is less than the standard multi-head self-attention module MHA. The pre-coding module is located at the head of the 2D transformer network, the classifier is located at the tail of the 2D transformer network, the feature extraction module and the dilated temporal self-attention module are located in the middle of the 2D transformer network, and the dilated temporal feature extraction module ETM is inserted between the feature extraction module and the classifier of the fully connected layer.
[0011] Preferably, the aerial video recognition model includes a 2D Transformer network Vision Transformer-base, which includes a pre-coding module, twelve coding block structures, and a classifier of a fully connected layer. The feature shift module FS is inserted into each coding block; some of the multi-head self-attention modules MHA in the original network are replaced with pooled multi-head self-attention PMHA modules; the dilated temporal feature extraction module ETM is inserted between the network feature extraction and the classifier to form the aerial video recognition model as a whole.
[0012] The beneficial effects of the present invention are as follows:
[0013] 1. The present invention utilizes the Pooling Multi-Head Self-Attention module (PMHA) to introduce multi-scale information, enabling the model to focus on the low-level visual information of the image at high resolution in the early stage and the deep semantic information of the image at low resolution in the later stage. That is, it utilizes the rich background information of aerial images and also pays attention to the deep detail information of the images, well adapting to the characteristics of aerial videos. Using pooling multi-head self-attention circumvents the limitations of convolutional neural networks in locally modeling the global situation on one hand, and on the other hand, it well addresses the problem of inconsistent semantic sizes of foreground and background objects in aerial videos. Not only is the recognition accuracy improved, but the pooling operation also shortens the length of the sequence, exponentially reducing the computational complexity of the spatial self-attention operation and significantly enhancing the computational efficiency. Due to the use of the pooling operation, the channel depth remains unchanged, and with global self-attention, the sequence length is not restricted. Therefore, the CLS token can be used as the spatial feature. The spatial expression of the CLS token participating in the self-attention calculation is stronger than the average aggregation of tokens. Compared with the full sequence, the input feature of the CLS token is shifted by the Feature Shift module (FS). The former operation is simple and avoids generating excessive offset padding zeros. The CLS token sequence can also be used as the input sequence for the subsequent Dilated Temporal Module (ETM).
[0014] 2. The present invention utilizes the Dilated Temporal Module (ETM), which can flexibly process video data of any length and fully utilizes the high-quality long-duration aerial data provided by drones. In the NLP field, the correlation between adjacent word vectors is extremely strong. Using window self-attention not only reduces the computational complexity but also directly avoids the attention allocation to irrelevant tokens, improving the accuracy. However, in the field of CV video analysis, adjacent frames often have very small differences. Using window self-attention results in redundant self-attention calculations and a relatively small receptive field in the time dimension. Therefore, in the present invention, dilated temporal self-attention is calculated here. Compared with window self-attention of the same length, it can calculate information of farther video frames and can well utilize long-term temporal information to enhance the expressive power of the model. Compared with global self-attention, it achieves a linear complexity in calculating dilated temporal self-attention. Calculating dilated temporal self-attention can improve the accuracy, effectiveness, and feasibility of the aerial video recognition model.
[0015] 3. The present invention utilizes the Feature Shift module (FS), which only needs to introduce temporal information in the spatial calculation without introducing additional parameters, thus reducing the computational complexity. In the model with separate calculation of spatio-temporal self-attention, while calculating the spatial self-attention, it maintains the attention to temporal information, enhancing the utilization of temporal information and also enhancing the expression of spatial features, alleviating the drawbacks of separate spatio-temporal processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic flowchart of the aerial video classification method of the present invention;
[0017] Figure 2 Schematic flowchart of the aerial video classification method based on spatio-temporal multi-scale Transformer of the present invention;
[0018] Figure 3 Schematic diagram of the spatio-temporal multi-scale Transformer network structure of the present invention;
[0019] Figure 4 Schematic diagram inside the encoding block in the spatio-temporal multi-scale Transformer network of the present invention;
[0020] Figure 5 Schematic diagram of pooling self-attention of the present invention;
[0021] Figure 6 Schematic diagram of the dilated temporal self-attention encoding block of the present invention. Detailed implementation manners
[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0023] Figure 1 An aerial video classification method provided by an embodiment of the present invention, as Figure 1 shown, the method includes: obtaining aerial video data and preprocessing the aerial video data; inputting the preprocessed aerial video data into a trained aerial video recognition model based on multi-scale Transformer, and outputting a recognition result.
[0024] Figure 2 An aerial video classification method based on spatio-temporal multi-scale Transformer provided by an embodiment of the present invention, as Figure 2 shown, in the classification method, the recognition process of the aerial video recognition model based on multi-scale Transformer is further defined; first, an equi-length video sequence needs to be obtained, and video frames are extracted at a fixed frequency to form a video frame sequence with a size of T frames; each video frame is cut into blocks to reduce the resolution of the video frame image; the two-dimensional video frame image is converted into a one-dimensional sequence, and the classification label CLS token is spliced from the one-dimensional sequence. After adding position encoding information, the aerial video recognition result is obtained through a multi-scale spatio-temporal feature extraction module and a dilated temporal feature extraction module.
[0025] The cleaning and preprocessing operations of aerial video data include: First, clip the video data to remove the blurred, unstable, and invalid parts, and obtain high-quality and equal-length video segments. The extraction frequency can be adjusted according to the existing computing resources, and video frames are extracted from the video segments at a fixed frequency to generate a video frame sequence of length T while adjusting the resolution of the video frames to a fixed size of 224×224.
[0026] Among them, the training set and the test set can be divided in a 1:1 ratio, and both the training set and the test set will be used as sample videos for training and testing.
[0027] In an exemplary embodiment, the sample video mentioned in the embodiments of the present application refers to the sample video based on which the aerial video recognition model based on multi-scale Transformer is trained once. The number of sample videos can be one or multiple, and the embodiments of the present application do not limit this. Exemplarily, the number of sample videos is multiple to ensure the model training effect. Exemplarily, for the case where the number of sample videos is multiple, the number of video frames in different sample videos is the same to ensure the model training effect.
[0028] In the embodiments of the present invention, the aerial video data is an aerial sample video for training and testing a video classification model. The aerial video data is clipped into effective and equal-length video segments as the training set and the test set according to requirements and computing conditions, ensuring that the acquisition results at the same frame rate are the same. To avoid the influence of data imbalance, the number of videos of different classes should be as equal as possible to ensure the model training effect. To avoid the influence of homologous videos on the test results, multiple video segments clipped from homologous videos need to be put into the same training set and test set to ensure the model training effect.
[0029] The sample video corresponds to a classification label, and the classification label is used to indicate the actual category corresponding to the sample video. The embodiments of the present application do not limit the representation form of the classification label corresponding to the sample video. Exemplarily, the classification label corresponding to the sample video is represented by the identification information of the actual category corresponding to the sample video, such as the name of the category, the code of the category, etc. Exemplarily, the actual category corresponding to the sample video is used to describe the content in the sample video. Exemplarily, the actual category corresponding to the sample video is a category in the candidate categories, and the candidate categories are set according to experience or flexibly adjusted according to the actual application scenario. The embodiments of the present application do not limit this.
[0030] It should be noted that the actual corresponding category of a sample video may be one or more, and the embodiments of the present application do not limit this. For example, if the content in a sample video is a performer playing a musical instrument, the actual corresponding category of this sample video is playing a musical instrument; or, if the content in a sample video is someone singing while walking, the actual corresponding category of this sample video is walking and singing; or, if the content in a sample video is rowing a boat, the actual corresponding category of this sample video is rowing a boat.
[0031] In the embodiments of the present invention, the training set and test set generated by clipping aerial video data are single-label, and the labels can be set according to requirements. For example, if the recognition task only has a basketball court and only global scene information, the label can be set as basketball court. However, if the recognition task has a basketball court and playing basketball, and both global scene information and motion information are required, then the video without people playing needs to be set as basketball court, and the video with people playing needs to be set as playing basketball. In this way, by setting the labels, the model can be made to learn to focus on different regions and information.
[0032] The acquisition method of the aerial video data in the embodiments of the present invention is not limited. The method of obtaining the video by oneself is as follows: operate the drone to shoot the target scene or subject in a multi-angle, long-time, partially occluded, and different lighting manner, and then select multiple different target scenes or subjects for shooting. The method of obtaining the sample video by computer is as follows: search for video clips containing the target scene or subject on the network as the aerial video data, which can be directly obtained from existing drone aerial video datasets, such as ERA, MOD20, UAVhuman.
[0033] Figure 3 This is a schematic diagram of the spatio-temporal multi-scale Transformer network structure based on the present invention. The aerial video recognition model based on the spatio-temporal multi-scale Transformer is as Figure 3 shown, including: using a 2D transformer network as the backbone network, the 2D transformer network includes a pre-coding module, a multi-scale spatio-temporal feature extraction module composed of a multi-level coding block structure, a dilated temporal feature extraction module ETM, and a classifier of a fully connected layer; wherein, each level of the coding block structure includes multiple layers of coding blocks, and each layer of coding block includes two feature shift modules FS, a multi-layer perceptron MLP, and a pooling multi-head self-attention module PMHA or a standard multi-head self-attention module MHA.
[0034] Among them, one feature offset module FS is located at the head of the layer encoding block, and the other feature offset module FS is inserted between the multi-layer perceptron MLP and the pooling multi-head self-attention module PMHA or between the multi-head self-attention module PMHA and the standard multi-head self-attention module MHA. And the pooling multi-head self-attention module PMHA is less than the standard multi-head self-attention module MHA, that is, the pooling multi-head self-attention module PMHA replaces a small part of the standard multi-head self-attention module MHA in the original Transformer network, so that the model focuses on the low-level visual information of the image at high resolution in the early stage and focuses on the deep language information of the image at low resolution in the later stage. That is, it utilizes the rich background information of aerial images and also pays attention to the deep detail information of the image, well adapting to the characteristics of aerial videos; the pre-encoding module is located at the head of the 2D transformer network, the classifier is located at the tail of the 2D transformer network, the feature extraction module and the dilated temporal self-attention module are located in the middle of the 2D transformer network, and the dilated temporal feature extraction module ETM is inserted between the feature extraction module and the classifier of the fully connected layer.
[0035] In some preferred embodiments, the 2D transformer neural network of the present invention is a Vision Transformer network, the feature offset module FS is twice the number of encoding blocks in the Vision Transformer network, and there are three pooling multi-head self-attention PMHA modules, and the network only includes one dilated temporal feature extraction module ETM.
[0036] As Figure 4 shown, the aerial video recognition model based on multi-scale Transformer includes a convolutional layer, a fully connected layer, a multi-scale spatio-temporal feature extraction module composed of twelve encoding block structures, a dilated temporal feature extraction module ETM, and a classifier structure. As Figure 4 shown, two feature offset modules FS are inserted in each encoding block structure, located at the input of the block and the output of the pooling multi-head self-attention PMHA module or the self-attention MHA module respectively, and the pooling multi-head self-attention PMHA module replaces the multi-head self-attention MHA in the first encoding block of stage2, stage3, and stage4.
[0037] In the embodiment of the present invention, the process of training the aerial video recognition model includes:
[0038] S1: Obtain a training video frame sequence with a length of T;
[0039] S2: Input the training video frame sequence into the pre - encoding module to reduce the resolution, compensate the channels, splice the classification token CLS token, and add position encoding information to generate a video frame sequence with the classification token CLS token;
[0040] S3: Input the video frame sequence with the classification token CLS token into the multi - scale spatio - temporal feature extraction module to obtain multi - scale spatio - temporal features. The multi - scale spatio - temporal features are video classification features containing other frame features at different resolutions;
[0041] S4: The classification token CLS token of each frame image, that is, the frame classification information, constitutes a frame classification information sequence. Splice the video classification features to the frame classification information sequence to form a CLS token sequence of length T + 1, and input the CLS token sequence into the dilated temporal module ETM to obtain the video classification features in the CLS token sequence;
[0042] S5: Input the video classification features in the CLS token sequence into the classifier, and the classification result with the largest score is the video classification result;
[0043] S6: Calculate the loss function of the classification process, update the network parameters through the loss function, and continuously update and iterate. When the loss function drops to the lowest, the model training is completed.
[0044] In some exemplary embodiments, in step S1, the obtained training video frame sequence of length T is pre - processed aerial video data. During the model training stage, this training video frame sequence is used for training. During the model validation stage, this training video frame sequence can be used for validation. During the model testing stage, this training video frame sequence can be used for testing. It can be understood that the training, validation, and testing processes of the model have similar or the same processing flows. Those skilled in the art should know that during the validation and testing stages, the pre - processed aerial video data can be processed in the same way as the training process to obtain the corresponding aerial video classification results.
[0045] In some exemplary embodiments, in step S2, such as Figure 3As shown, the preprocessed input sequence forming the T-frame is fed into the aerial video recognition model. The input T-frame sequence first passes through a convolutional layer to reduce the dimension of the input sequence and compensate the information in the channel dimension. For example, the details can be as follows: Through 96 convolutional kernels of size 4×4 with a stride of 4, the dimension of the picture is converted from 3×224×224 to 96×56×56, which is equivalent to changing the unit of the image from 1×1 pixel to a patch composed of 4×4 pixels. Then, the two-dimensional picture is converted into a one-dimensional sequence, and a learnable classification token CLS token with the same channel dimension as that involved in the attention calculation is concatenated to the one-dimensional sequence of each frame, and then the position encoding is generated and added to the sequence. The expression of the pre-encoding module is:
[0046] X 1 =Cat N (flat(Conv2d(imgT)), cls)+Pe
[0047] Among them, img T is the video frame with a length of T, cls is the classification token CLS token, Cat N is the concatenation operation of tensors in dimension N, flat(.) is the function to convert a two-dimensional matrix into a one-dimensional sequence, Pe is the position encoding information, and X 1 represents the input of the first layer encoding block.
[0048] In step S3, the specific process of using the multi-scale spatio-temporal feature extraction module to process the input video frame sequence with the classification token CLS token includes:
[0049] S31: The video frame sequence with the classification token CLS token is used as the input sequence and input into the feature shift module FS to shift the channels of the classification token CLS token, so as to establish the spatio-temporal information interaction between different video frame images in the input sequence;
[0050] S32: The shifted input sequence is input into the pooling multi-head self-attention module PMHA or the standard multi-head self-attention module MHA, and the self-attention at different scales is obtained by calculating the pooling multi-head self-attention or the standard multi-head self-attention;
[0051] S33: The input sequence after calculating the pooling multi-head self-attention is input into the MLP for dimension transformation, and a non-linear mapping relationship is introduced for the dimension transformation relationship of the linear input sequence.
[0052] X m =X i +Attention(SAS(X i ))
[0053] X i+1 = Xm + MLP(SAS(X m ))
[0054] where the encoding block is divided into three parts: Attention, MLP, and feature offset module FS, and X i is the input sequence of the encoding block, X m is the input sequence of the MLP, and X i+1 is the output sequence of the encoding block.
[0055] Then, the sequence passing through the pre-encoding module is input into the first encoding block structure. First, the CLS token in the input sequence is extracted as the first classification token matrix cls ∈ R T×1×C , and a dimensionality reduction operation is performed, cls → cls' ∈ R T×C , to obtain the second classification token matrix cls'; a third classification token matrix S of all zeros with the same dimension as cls' is set. The channels of cls' are divided into three parts. The first part and the second part are offset in opposite directions along the time dimension by a distance of 1 frame, and after the offset, the dimension of the original sequence cls is restored. The parts where the CLS token becomes a null value due to the offset will be filled with zeros. The channels of the third part remain unchanged. Through assignment, the third classification token matrix CLSS corresponding to the offset classification token CLS is obtained. The dimension of the offset third classification token matrix S is increased to be the same as that of the first classification token matrix cls and replaces the non-offset first classification token matrix clsCLS in the input sequence; in this embodiment, first, a third classification token matrix S of the same size as cls' and initially all zeros is designed. By offsetting cls', the offset classification token values in cls' are filled into S, changing the corresponding element values in S; by increasing the dimension of the S matrix, the non-offset classification token values in cls are filled into S, changing the corresponding element values in S, and finally, the all-zero matrix is updated to obtain the updated third classification token matrix. Among them, the expression for the offset of the CLS token is:
[0056] S[:, -1, :fold] = cls'[1:, :fold]
[0057] S[1:, fold:2*fold] = cls'[:-1, fold:2*fold]
[0058] S[:, 2*fold:] = cls'[:, 2*fold:]
[0059] where R represents the spatio-temporal matrix, T represents the length of the input sequence, C represents the number of channels, and X ∈ R T×1×CThe CLS token in the input sequence, and fold is the channel length offset between the first part and the second part.
[0060] This model designs twelve encoding blocks, which are divided into four stages. Each stage contains three encoding blocks, and each stage has different resolutions and channel numbers. In the first encoding block of each stage, pooling multi-head self-attention (PMHA) is used to reduce the resolution (except for the first encoding block, where the multi-head self-attention (MHA) module is used. The resolution has already been reduced and the channels have been increased in the pre-encoding module). The MLP in the last encoding block increases the number of channels. The data stream input to the network structure is shown in the figure.
[0061] Stage Tensor shape Patch Embedding T×3×224×224 stagel T×96×(56×56+1) stage2 T×192×(28×28+1) stage3 T×384×(14×14+1) stage4 T×768×(7×7+1) ETM T×768 MLP 1×768
[0062] The offset sequence is first normalized by Layer Normalization according to the encoding block, and then the pooling multi-head self-attention (PMHA) module or the self-attention (MHA) module is calculated according to the level of the encoding block. Here, the processing of the pooling multi-head self-attention (PMHA) module is mainly described. Figure 5 It is a schematic diagram of pooling self-attention.
[0063] For the output sequence passed through the offset module FS Perform a linear mapping to obtain the query tensor Q, the key tensor K, and the value tensor V.
[0064] Q = X S W Q K = X S W K V = X S W V
[0065] W Q W K W V They are trainable and parameter-shared parameter matrices implemented by three fully-connected layers.
[0066] Before performing the pooling operation on the sequence (X S , Q′, K′, V′), the CLS token is stripped to obtain Perform a max-pooling operation on X S QKV, reducing the resolution by half, and then concatenating the CLS token to obtain Then calculate the multi-head self-attention, and finally input it into the Dropout layer to reduce the overfitting phenomenon. The overall structure is a residual connection, and multi-scale spatio-temporal interaction features are obtained in the spatial dimension. The formulas for pooling multi-head self-attention (PMHA) and standard multi-head self-attention (MHA) are as follows:
[0067]
[0068]
[0069]
[0070]
[0071] Among them is max pooling which performs row-wise normalization on the inner product matrix
[0072] After calculating self-attention, the sequence X m , introduces spatio-temporal interaction through the FS module, and then inputs it into the MLP block for dimensional transformation and introduces non-linearity to increase the expressive power of the model. If the encoding block is the last one in the stage, the MLP needs to double the channel dimension. Both calculating self-attention and the MLP inside the encoder adopt the residual connection method plus the original input. The expression form is as follows
[0073] X m = X i + Attention(SAS(X i ))
[0074] X i+1 = X m + MLP(SAS(X m ))
[0075] Among them, the encoding block is divided into three parts: Attention, MLP, and the feature offset module FS. X i is the input sequence of the encoding block, X m is the input sequence of the MLP, and X i+1 is the output sequence of the encoding block
[0076] The output of the last encoding block is a CLS token sequence of length t Input the sequence X into the dilated temporal feature extraction module ETM. At this time, the CLS token represents the classification information of each frame of the image. Concatenate the video classification feature token cls to the CLS token sequence X t . If the length of the video frame sequence is less than the preset length, for example, the sequence length is short and less than or equal to 16 frames (preset length), global self-attention can be calculated. Otherwise, dilated temporal self-attention is calculated. For the convenience of self-attention operation, pad sequences of length The all-zero mask has a hole interval of d, the length of the hole is n, and it slides 1 frame distance along the time dimension each time. When calculating the dilated multi-head self-attention, the dilated self-attention is calculated for the frame features, and the global self-attention is calculated for the cls t Calculate the global self-attention.
[0077] In the preferred embodiment of the present invention, the dilated temporal feature extraction module ETM is designed as a three-layer encoding block structure. The schematic diagram of the dilated temporal self-attention encoding block is as Figure 6 shown. For each layer of encoding block j, the calculation formula of the encoding block in ETM is as follows:
[0078] CLS = Cat(cls j , Mask(X))
[0079] cls j+1 = Ω T-MSA (cls j , CLS)
[0080] X j+1 = Ω TW-MSA (X j , CLS)
[0081] Among them, Mask(·) is the concatenated all-zero mask, indicating that all-zero masks are concatenated at both ends of the video frame cls sequence. The length of the video frame cls sequence is equal to T, which is the frame classification information sequence. The length of Mask(X) is greater than T. cls j represents the video classification feature input to the j-th layer encoding block in ETM, with a length of 1, representing the video classification feature obtained by the multi-scale spatio-temporal feature extraction module. CLS is the complete calculation sequence, with a length greater than T + 1; Ω T-MSA is the global self-attention in the time dimension, and Ω TW-MSA is the dilated self-attention in the time dimension. cls j+1 represents the video classification feature output by the j-th layer encoding block in ETM, that is, the video classification feature input to the (j + 1)-th layer encoding block. Therefore, cls j+1 can be the output feature of ETM or the video classification feature input to the next-level block. This video classification feature cls j+1 is the feature after global self-attention calculation, and the initially input cls j is the video classification feature containing other frame features at different resolutions output by the multi-scale spatio-temporal feature extraction module; X j+1 is input as the CLS token sequence of the (j + 1)-th layer encoding block. This X j+1is the feature after the calculation of the dilated temporal self-attention. This feature can make good use of the long temporal information to enhance the expression ability of the model. Therefore, using this feature as the output feature of the ETM for classification and recognition can improve the accuracy, effectiveness, and feasibility of the aerial video recognition model, and the initially input X j is the output sequence of the multi-scale spatio-temporal feature extraction module, that is, the CLS token sequence
[0082] The computational formulas for the dilated self-attention and the global self-attention in the time dimension are as follows:
[0083] Ω T-MSA = tf(t)
[0084]
[0085] where t is the sequence length of the video frames, n is the full length of the dilation, d represents the dilation interval, and f(x) is the computational amount of the self-attention of the sequence with length x.
[0086] The classification feature token representing the video is obtained through training with the dilated time feature extraction module ETM The cls output by the last encoding block j Inputting the video classification feature into a fully connected layer can obtain the final classification information, and the expression form is as follows:
[0087]
[0088] where fc(·) is a fully connected layer classifier with an input dimension of 768 and an output dimension of the total number of classes, and Max(·) represents taking the maximum result of all scores is the final classification category of the video.
[0089] The loss function of the model is the cross-entropy loss function, and its expression is:
[0090]
[0091] where θ is the model parameter, χ is the input data, N represents the training batch, nc represents the total number of classes is the result of the feature passing through the fully connected layer classifier is the event indicator function, which judges whether the i-th sample is the category c. If it is, it is 1, and if not, it is 0. The expression of the event indicator function is as follows:
[0092]
[0093] Continuously calculate the loss function, update the network parameters through backpropagation, continuously update and iterate to improve the recognition accuracy of the model. When the loss function drops to the lowest, the model training is completed. After the above process, the aerial video result can be output from the fully connected layer fc of the last layer, taking Figure 3 as an example. After obtaining the aerial video data and preprocessing the aerial video data, through the trained aerial video recognition model based on multi-scale Transformer, the classification recognition result of rowing can be obtained.
[0094] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for classifying aerial videos based on spatio-temporal multi-scale Transformer, the method comprises: obtaining aerial video data and preprocessing the aerial video data; inputting the preprocessed aerial video data into a trained aerial video recognition model based on multi-scale Transformer to output a recognition result; characterized in that the aerial video recognition model based on multi-scale Transformer includes using a 2D transformer network as the backbone network, and the 2D transformer network includes a pre-coding module, a multi-scale spatio-temporal feature extraction module composed of a multi-level coding block structure, a dilated temporal feature extraction module ETM, and a classifier of a fully-connected layer; wherein, each level of coding block structure includes multiple layers of coding blocks, and each layer of coding block sequentially includes a first feature offset module FS, a first normalization layer, a pooling multi-head self-attention module PMHA or a standard multi-head self-attention module MHA, a first Dropout layer, a second feature offset module FS, a second normalization layer, a multi-layer perceptron MLP, and a second Dropout layer, wherein the first normalization layer, the pooling multi-head self-attention module PMHA, and the first Dropout layer form a residual structure, and the second normalization layer, the multi-layer perceptron MLP, and the second Dropout layer form a residual structure; the pre-coding module is located at the head of the 2D transformer network, the classifier is located at the tail of the 2D transformer network, the feature extraction module and the dilated temporal feature extraction module ETM are located in the middle of the 2D transformer network, and the dilated temporal feature extraction module ETM is inserted between the feature extraction module and the classifier of the fully-connected layer; Among them, the operation process of the dilated temporal feature extraction module ETM in the temporal dimension includes: the output of the last encoding block is a sequence X of classification tokens CLS tokens of length t. The sequence X is input into the dilated temporal feature extraction module ETM. At this time, the classification token CLS token represents the frame classification information of each frame of the image, and X represents the CLS token sequence spliced with the video classification feature token cls t ; if the length of the CLS token sequence is lower than the preset length, global self-attention is calculated, otherwise dilated temporal self-attention is calculated, and the CLS token sequence is determined according to the self-attention The process of calculating the temporal self-attention of the holes includes concatenating all-zero masks of on the left and right of the CLS token sequence, with a hole interval of d, a hole length of n, and sliding 1 frame distance along the time dimension each time; The calculation formulas of dilated temporal self-attention and global self-attention are as follows: Ω T-MSA = tf(t) Among them, Ω T-MSA is the global self-attention in the time dimension, and Ω TW-MSA is the dilated self-attention in the time dimension, t is the sequence length of video frames, and f(x) is the self-attention calculation amount of a sequence of length x.
2. The method for classifying aerial videos based on spatio-temporal multi-scale Transformer according to claim 1, characterized in that, the preprocessing operation of the aerial video data includes clipping the aerial video data to remove blurred, unstable and invalid parts, and obtaining equi-length video segments; extracting video frames from each equi-length video segment at a fixed frequency, and generating a video frame sequence with a length of T as the input of the aerial video recognition model based on multi-scale Transformer by adjusting the video frame resolution.
3. The method for classifying aerial videos based on spatio-temporal multi-scale Transformer according to claim 1, characterized in that, the process of training the aerial video recognition model based on multi-scale Transformer includes: S1: obtaining a training video frame sequence with a length of T; S2: inputting the training video frame sequence into the pre-coding module, reducing the resolution, compensating the channels, splicing the classification token CLS token and adding position encoding information to generate a video frame sequence with the classification token CLS token; S3: Input the video frame sequence with the classification token CLS into the multi-scale spatio-temporal feature extraction module to obtain multi-scale spatio-temporal features, where the multi-scale spatio-temporal features are video classification features containing other frame features at different resolutions; S4: Construct a frame classification information sequence from the classification token CLS of each frame image, i.e., the frame classification information. Concatenate the frame classification information sequence with the video classification features to form a CLS token sequence of length T + 1, and input the CLS token sequence into the dilated temporal module ETM to obtain the video classification features in the CLS token sequence; S5: Input the video classification features in the CLS token sequence into the classifier to obtain the classification result with the largest score, which is the video classification result; S6: Calculate the loss function of the classification process, update the network parameters through the loss function, and continuously update and iterate. When the loss function drops to the lowest, the model training is completed.
4. A method for aerial video classification based on spatio-temporal multi-scale Transformer according to claim 3, characterized in that, In step S3, the specific process of using the multi-scale spatio-temporal feature extraction module to process the input video frame sequence with the classification token CLS includes: S31: Construct the input sequence from the video frame sequence with the classification token CLS and input it into the feature shift module FS to shift the channels of the classification token CLS, so as to establish spatio-temporal information interaction between different video frames in the input sequence; S32: Input the shifted input sequence into the pooling multi-head self-attention module PMHA or the standard multi-head self-attention module MHA, and calculate the self-attention at different scales by calculating the pooling multi-head self-attention or the standard multi-head self-attention; S33: Input the input sequence after calculating the pooling multi-head self-attention into the MLP for dimensional transformation, and introduce a non-linear mapping relationship into the dimensional transformation relationship of the linear input sequence.
5. A method for aerial video classification based on spatio-temporal multi-scale Transformer according to claim 4, characterized in that, In step S31, the process of establishing spatio-temporal interaction between different video frames in the input sequence by using the feature offset module FS includes extracting the classification token CLS token in the input sequence and using it as the first classification token matrix cls. A dimensionality reduction operation is performed on the first classification token matrix cls to obtain a second classification token matrix cls'. A third classification token matrix S of all zeros with the same dimension as the second classification token matrix cls' is set. The channels of cls' are divided into three parts. The channels of the first part and the second part are offset in opposite directions along the time dimension. The parts where the classification token CLS token becomes a null value due to the offset are filled with zeros. The channels of the third part remain unchanged. The third classification token matrix S corresponding to the offset classification token CLS token is obtained through assignment. The offset third classification token matrix S is subjected to a dimensionality increase operation to be consistent with the dimension of the first classification token matrix cls, and replaces the non-offset first classification token matrix cls in the input sequence, where cls ∈ R T×1×C , cls' ∈ R T×C , R represents the spatio-temporal matrix, T represents the length of the input sequence, and C represents the number of channels.
6. A method for aerial video classification based on spatio-temporal multi-scale Transformer according to claim 4, characterized in that, In step S32, calculating the self-attention at different scales through the pooling multi-head self-attention includes normalizing the shifted input sequence, then performing a linear mapping to obtain the query tensor Q, key tensor K, and value tensor V, and stripping the classification token CLS of the input sequence; then performing a max pooling operation on the sequence after stripping the classification information and each tensor QKV; concatenating the classification tokens CLS in the input sequence to obtain the pooled input sequence and the QKV sequence, calculating the multi-head self-attention according to the pooled input sequence and the QKV sequence, and inputting the calculated multi-head self-attention into the Dropout layer to obtain the features of multi-scale spatio-temporal interaction containing the classification token CLS in the spatial dimension.
7. Aerial video classification method based on spatio-temporal multi-scale Transformer according to claim 3, characterized in that, the classification result expression form is as follows: where, cls j is the video classification feature in the CLS token sequence, fc(·) is the fully connected layer classifier, and Max(·) represents taking the classification result with the highest score. is the final classification category of the aerial video data.