Aerial Video Classification Model and Method

By using local semantic enhancement encoder and window semantic enhancement Transformer blocks in the aerial video classification model, local video features of key window areas are located and stripped away, and the problem of low efficiency and accuracy of drone video recognition in complex scenarios is solved, and efficient aerial video recognition is achieved.

CN119229319BActive Publication Date: 2025-05-30CHONGQING GEOMATICS & REMOTE SENSING CENT
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411269784.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-11
Publication Date
2025-05-30
Estimated Expiration
2044-09-11

AI Technical Summary

Technical Problem

The drone videos collected in complex scenarios have a large amount of background information that is insensitive to human behavior information, which leads to excessive self-attention calculations of drone videos, reducing the recognition efficiency and accuracy of drone videos.

Method used

Aerial video classification model is proposed, using local semantic enhancement encoder and window semantic enhancement Transformer block, and the key window area is positioned through the window positioning module, local video features are stripped away, and window time multi-head self-attention module is used to calculate window time multi-head self-attention of local video features, and background information that is insensitive to motion information is excluded.

Benefits of technology

It improves the efficiency and accuracy of aerial video recognition, avoids the problem of excessive calculation, enhances the local motion information of aerial video, and improves the recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229319B_ABST
    Figure CN119229319B_ABST
Patent Text Reader

Abstract

The present invention discloses an aerial video classification model and method. The encoder includes a window positioning module and a window temporal multi-head self-attention module. The window positioning module calculates the feature response of the input video features using a non-padding convolutional kernel with the same size as the local window, and thereby determines the key window area with the maximum characteristic response in the video features, and then extracts the local video features within the key window area. The window temporal multi-head self-attention module calculates the window temporal multi-head self-attention of the local video features, and adds the window temporal multi-head self-attention to the video features through a residual block. In this way, not only the background information insensitive to motion information is excluded, the excessive computational load caused by calculating self-attention for too long video sequences is avoided, and the efficiency of aerial video recognition is improved. Also, the local motion information of the aerial video is enhanced, and the accuracy of subsequent aerial video recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video recognition, and particularly relates to an aerial video classification model and method. Background Art

[0002] With the continuous development of aviation automation technology and remote sensing imaging technology, drones can capture a large amount of remote sensing images from different perspectives due to their high mobility, low cost, and easy operation. At the same time, drones equipped with intelligent image analysis systems can capture and analyze videos and images, which have extremely high practical value in many application fields, such as target reconnaissance, disaster detection, logistics distribution, pest analysis, etc.

[0003] Processing drone videos in a manual interpretation manner is costly and slow, and it is difficult to adapt to the large amount of data obtained by drones. Therefore, a more effective and efficient way is needed to automatically interpret the content of drone videos. Deep learning is an important research branch of machine learning. It learns complex features and representations through the targeted design of deep neural networks and is widely used in fields such as computer vision and natural language processing due to its excellent robustness and generalization.

[0004] Compared with the manual interpretation method, the method based on deep learning can automatically interpret the content of drone videos in a more effective and efficient way. Among them, convolutional neural networks and vision transformers are the mainstream deep learning methods in the field of computer vision.

[0005] Transformer is a deep neural network based on self-attention. Compared with convolutional neural networks, Transformer naturally has excellent capabilities for global feature modeling and parallel computing. It can effectively utilize the rich information in drone videos and efficiently process the large amount of data in drone videos. Currently, the research on Transformer in the field of video recognition mainly focuses on human behaviors in conventional videos. Such video events have the characteristics of clear targets and obvious target behaviors.

[0006] However, there is a large amount of background information in drone videos collected in complex scenarios that is insensitive to human behavior information. This not only leads to an excessive self-attention calculation amount in drone videos, reducing the recognition efficiency of drone videos, but also affects the recognition results of drone videos, reducing the recognition accuracy of drone videos. Summary of the Invention

[0007] Aiming at the deficiencies of the existing technology, the present invention proposes an aerial video classification model and method, which can improve the efficiency and accuracy of aerial video recognition. The specific technical solutions are as follows:

[0008] In a first aspect, a local semantic enhancement encoder is provided. In a first implementable manner of the first aspect, it includes:

[0009] A window localization module configured to locate the key window region of the video features using a non-padding convolutional kernel with the same size as the window, and strip the local video features within the key window region;

[0010] A window temporal multi-head self-attention module configured to calculate the window temporal multi-head self-attention of the local video features and perform a residual connection between the window temporal multi-head self-attention and the video features.

[0011] Combined with the first implementable manner of the first aspect, in a second implementable manner of the first aspect, the window localization module includes:

[0012] A video pooling unit configured to perform global average pooling and global max pooling on the input video features respectively to obtain corresponding pooled features;

[0013] A channel splicing unit configured to splice the two pooled features obtained by the video pooling unit.

[0014] To obtain corresponding spliced features;

[0015] A feature response unit configured to calculate the feature responses of the spliced features using a non-padding convolutional kernel with the same size as the window region, and normalize the obtained feature responses to determine the weights of the feature responses;

[0016] A window position unit configured to perform global average pooling on the weights of all feature responses based on the time dimension and locate the key window position of the video features according to the feature response with the maximum average weight;

[0017] A window region unit configured to determine the key window region according to the key window position and the size of the non-padding convolutional kernel.

[0018] Combined with the first implementable manner of the first aspect, in a third implementable manner of the first aspect, it further includes:

[0019] A first layer normalization module configured to perform normalization processing on the video features and input the processed video features into the window localization module;

[0020] A second layer normalization module configured to perform normalization processing on the video features with the window temporal multi-head self-attention added;

[0021] A first multi-layer perceptron configured to perform feature extraction on the video features output by the second layer normalization module, and the extracted features are residually connected to the video features input to the second layer normalization module.

[0022] In a second aspect, a window semantic enhancement Transformer block is provided, including:

[0023] The local semantic enhancement encoder as described in any one of the first to third implementable manners of the first aspect;

[0024] A standard encoder that extracts features from the video features output by the local semantic enhancement encoder to obtain the global scene information and local scene information of the video features.

[0025] Combined with the first implementable manner of the second aspect, in the second implementable manner of the second aspect, the standard encoder includes a third layer normalization module, a multi-head attention module, a fourth layer normalization module, and a second multi-layer perceptron connected in sequence.

[0026] Combined with the first implementable manner of the second aspect, in the third implementable manner of the second aspect, the output of the multi-head attention module is connected to the input of the standard encoder through a residual block;

[0027] The output of the second multi-layer perceptron is connected to the input of the fourth layer normalization module through a residual block.

[0028] In a third aspect, an aerial video classification model is provided. In the first implementable manner of the third aspect, it includes:

[0029] Multiple window semantic enhancement Transformer blocks as described in the third implementable manner of the second aspect, configured to extract features from the aerial video to obtain the global scene information and local scene information of the aerial video;

[0030] A global-local self-fusion Transformer block configured to fuse the global scene information and local scene information of the aerial video using a self-fusion attention mechanism to obtain a fused video feature representation;

[0031] A classifier configured to determine the classification result of the aerial video based on the fused video feature representation.

[0032] Combined with the first implementable manner of the third aspect, in the second implementable manner of the third aspect, the global-local self-fusion Transformer block includes:

[0033] A Transformer encoder configured to preprocess the global scene information and local scene information of the aerial video;

[0034] A self-fusion attention module configured to fuse the preprocessed global scene information and local scene information to obtain the video feature representation.

[0035] Combined with the second implementation manner of the third aspect, in the third implementation manner of the third aspect, the self-fusion attention module includes:

[0036] A query vector calculation unit, configured to respectively determine a global key-value vector and a local key-value vector through a single linear mapping according to the global scene information and the local scene information, and obtain a globally and locally mixed query vector through channel concatenation and linear mapping according to the global key-value vector and the local key-value vector;

[0037] A probability distribution calculation unit, configured to calculate a probability distribution matrix according to the query vector, the global key-value vector, and the local key-value vector;

[0038] A video feature calculation unit, configured to calculate a fused video feature representation according to the probability distribution matrix, the global key-value vector, and the local key-value vector.

[0039] In a fourth aspect, an aerial video classification method is provided, including:

[0040] Obtain aerial data, and preprocess the aerial data to obtain an aerial video;

[0041] Adopt the trained aerial video classification model according to any one of the first to third implementation manners of the third aspect to classify the preprocessed aerial video.

[0042] Beneficial effects: By using the aerial video classification model and method of the present invention, the key window area with the maximum feature response in the aerial video can be located through the window positioning module, and then the local video features within the key window area can be stripped. The window-time multi-head self-attention of the local video features is calculated through the window-time multi-head self-attention module, thereby excluding background information that is insensitive to motion information and avoiding excessive computational complexity caused by calculating self-attention for too long video sequences, improving the efficiency of aerial video recognition. And by adding the window-time multi-head self-attention to the aerial video through residual connection, the local motion information of the aerial video is enhanced, and the recognition accuracy of the aerial video is improved. Description of the Drawings

[0043] In order to more clearly illustrate the specific implementation manners of the present invention, the drawings required for the specific implementation manners will be briefly introduced below. In all the drawings, the components or parts do not necessarily draw according to the actual proportion.

[0044] Figure 1 It is a schematic structural diagram of a local semantic enhancement encoder provided by an embodiment of the present invention;

[0045] Figure 2 It is a schematic diagram of the positioning principle of the window positioning module provided by an embodiment of the present invention;

[0046] Figure 3 Structural schematic diagram of a window semantic enhancement Transformer block provided by an embodiment of the present invention;

[0047] Figure 4 Structural schematic diagram of an aerial video classification model provided by an embodiment of the present invention;

[0048] Figure 5 Structural schematic diagram of a global-local self-fusion Transformer block provided by an embodiment of the present invention;

[0049] Figure 6 Structural schematic diagram of a self-fusion attention module provided by an embodiment of the present invention;

[0050] Figure 7 Flowchart of an aerial video classification method provided by an embodiment of the present invention. Detailed implementation manners

[0051] Hereinafter, embodiments of the technical solutions of the present invention will be described in detail with reference to the accompanying drawings. The following embodiments are only used to illustrate the technical solutions of the present invention more clearly, and thus are only examples and cannot be used to limit the protection scope of the present invention.

[0052] It should be understood that in this embodiment, the purpose of aerial video recognition is to analyze and understand the spatio-temporal visual patterns existing in the video, and its events include multiple concepts such as objects, scenes, behaviors, etc., such as disaster events such as fires and floods, human activity events such as running and swimming, and traffic events such as car accidents and traffic jams. Event recognition and behavior recognition are not in an antagonistic relationship. When the video content to be recognized is the behavior of a certain target, the event recognition task can be regarded as a behavior recognition task.

[0053] As Figure 1 shown in the structural schematic diagram of the local semantic enhancement encoder, the encoder includes:

[0054] A window positioning module configured to locate the key window area of the video feature by using a non-padding convolutional kernel of the same size as the window and strip the local video feature within the key window area;

[0055] A window temporal multi-head self-attention module configured to calculate the window temporal multi-head self-attention of the local video feature and perform a residual connection between the window temporal multi-head self-attention and the video feature.

[0056] Specifically, the local semantic enhancement encoder includes a window positioning module and a window temporal multi-head self-attention module. Among them, the window positioning module can calculate the feature response of the input video features using a non-padding convolutional kernel with the same size as the local window, and thereby determine the key window area with the maximum characteristic response in the video features, and then extract the local video features within the key window area. Then, the window temporal multi-head self-attention of the local video features is calculated through the window temporal multi-head self-attention module, and the window temporal multi-head self-attention is added to the video features through a residual block to learn the identity mapping, alleviate the vanishing gradient, promote the rapid convergence of the model, and improve the efficiency of aerial video recognition.

[0057] In this way, not only the background information insensitive to motion information is excluded, the excessive computational load caused by calculating self-attention for too long video sequences is avoided, but also the targets with complete space are included, the efficient modeling of temporal information is achieved, and the efficiency of aerial video recognition is improved. It also enhances the local motion information of the aerial video and improves the accuracy of subsequent aerial video recognition.

[0058] In this embodiment, optionally, the encoder further includes:

[0059] The first layer normalization module is configured to perform normalization processing on the video features and input the processed video features into the window positioning module;

[0060] The second layer normalization module is configured to perform normalization processing on the video features added with the window temporal multi-head self-attention;

[0061] The first multi-layer perceptron is configured to perform feature extraction on the video features output by the second layer normalization module, and the extracted features are residually connected to the video features input to the second layer normalization module.

[0062] Specifically, the local semantic enhancement encoder is a feed-forward neural network structure, including a first layer normalization module, a window positioning module, a window temporal multi-head self-attention module, a second layer normalization module, and a first multi-layer perceptron connected in sequence.

[0063] Among them, the first layer normalization module can perform layer normalization processing on the input video features, prevent internal covariate shift during the learning process of the encoder, accelerate the learning convergence speed of the encoder, and improve the efficiency of aerial video recognition.

[0064] The window positioning module can calculate the feature response of the video features processed by the first layer normalization module using a non-padding convolutional kernel with the same size as the local window, and thereby determine the key window area with the maximum characteristic response in the video features, and then extract the local video features within the key window area.

[0065] In this embodiment, optionally, the window positioning module includes:

[0066] A video pooling unit configured to perform global average pooling and global max pooling on the input video features respectively to obtain corresponding pooled features;

[0067] A channel splicing unit configured to splice the two pooled features obtained by the video pooling unit.

[0068] To obtain corresponding spliced features;

[0069] A feature response unit configured to calculate the feature responses of the spliced features using a non-padding convolution kernel with a consistent window area size, and normalize the obtained feature responses to determine the weights of the feature responses;

[0070] A window position unit configured to perform global average pooling on the weights of all feature responses based on the time dimension, and locate the key window position of the video features according to the feature response with the maximum average weight;

[0071] A window area unit configured to determine the key window area according to the key window position and the size of the non-padding convolution kernel.

[0072] Specifically, as Figure 2 shown, the window positioning module includes a video pooling unit, a channel splicing unit, a feature response unit, a window position unit, and a window area unit. Among them, the video pooling unit can perform global average pooling and global max pooling on the input video features respectively through a global average pooling element and a global max pooling element based on the channel dimension, so as to obtain the global average pooling feature and the global max pooling feature corresponding to the video features.

[0073] The channel splicing unit can splice the global average pooling feature and the global max pooling feature calculated by the video pooling unit along the channel dimension. The specific calculation formula is as follows:

[0074] X p = Cat c (AvgPool(X), MaxPool(X));

[0075] Where T and C respectively represent the number of sampled frames and the channel dimension; N = H×W / P 2 represents the feature dimension, H and W respectively represent the height and width of the input video frame, that is, the spatial dimensions, P represents the convolution kernel size in the block embedding operation; Cat c (·) represents channel splicing, AvgPool C (·) and MaxPool C(·) respectively represent channel-based global average pooling and global max pooling.

[0076] The feature response unit can transform the feature dimension of the concatenated feature X p from one-dimensional to two-dimensional to perform subsequent convolution operations, where To ensure that the feature response can be mapped to the corresponding window area, a padding-free convolution with a convolution kernel the same size as the window can be used to calculate the feature response, thereby reducing the number of channels of the feature to 1. The feature response unit normalizes the feature response using the Softmax function to determine the feature response weights corresponding to each window obtained through convolution calculation. The specific calculation is as follows:

[0077]

[0078] In the formula, represents the feature response weight of each frame of the image; u = n - s + 1 represents the size of the feature response weight; Conv s×s (·) represents a padding-free convolution kernel of size s. The size s of the convolution kernel is the size of the window, and the window size is a hyperparameter of network training and can be determined according to the actual training task. If the drone is flying at low altitude and the recognition target is obvious, the window can be enlarged; if the drone is flying at high altitude and the recognition target is relatively small, the window can be reduced. During training, the parameter of the window size can be adjusted multiple times to ensure that the optimal window response area can be learned.

[0079] It should be understood that in an aerial video, due to the small size of the moving target, the spatial displacement of the moving target along the time axis is often more important than the movement itself. For example, in the event of cycling, the movement information of the bicycle in space is more valuable than the movement information of the athlete's legs, and the latter is difficult to capture.

[0080] To obtain this information, one window corresponds to one video. The window position unit can average the feature response weights X of all windows in the time dimension map . Then obtain the average weight of the maximum position to obtain the window position with high response in the input video feature X as the key window position. The specific calculation is as follows:

[0081]

[0082] In the formula, represents the key window position; AvgPool T (·) represents global average pooling based on the time dimension; Argmax(·) represents calculating the position of the highest value of the feature response weight.

[0083] After locating the key window position, the window area unit can infer the key window area based on the key window position and the size of the padding-free convolutional kernel, so as to strip out the video features within the key window area. This local video feature mainly focuses on the targets in motion while excluding the scene information that is not sensitive to spatio-temporal information. The specific calculation is as follows:

[0084] X w = Loc(X, pos, s),

[0085] where, represents the local video feature; Loc represents the operation of locating the local video feature based on the window position.

[0086] The window temporal multi-head self-attention module can calculate the window temporal multi-head self-attention of the local video feature stripped out by the window localization module, and add the window temporal multi-head self-attention to the video feature through a residual block. The specific calculation formula is as follows:

[0087]

[0088] where, is the input of the layer encoder, and X w is the local video feature.

[0089] It should be understood that several existing spatio-temporal information modeling methods based on Transformer include:

[0090] Joint spatio-temporal multi-head self-attention, which performs self-attention on the spatio-temporal joint video feature. This approach tends to focus on richer scene information, while the overly long feature sequence will lead to excessive computational complexity.

[0091] Temporal multi-head self-attention based on pixel blocks, which performs self-attention on pixel blocks with the same space but different times. Compared with joint spatio-temporal multi-head self-attention, although the computational complexity is lower, it breaks the integrity of the target in the spatial dimension.

[0092] In this embodiment, the window temporal multi-head self-attention adopted by the window temporal multi-head self-attention module only calculates self-attention for the local subjects that are time-sensitive and spatially complete, which can achieve efficient modeling of temporal information. On a video frame sequence of length T, the computational complexity comparison between window temporal multi-head self-attention and other methods is as follows:

[0093] Ω(WT-MSA) = 4s 2 TC 2 + 2s 4 T 2 C,

[0094] Ω(JST-MSA) = 4hwTC2 +2(hw) 2 T 2 C,

[0095] Ω(PT-MSA) = 4hwTC 2 +2hwT 2 C,

[0096] where JST-MSA is spatio-temporal joint multi-head self-attention; PT-MSA is pixel-block based temporal multi-head self-attention; since s 2 << hw, compared with the other two self-attentions, window temporal multi-head self-attention exhibits lower computational complexity. Coupled with a window positioning module, it realizes accurate modeling of temporal information.

[0097] The second layer normalization module can perform layer normalization on the video features with window temporal multi-head self-attention added, prevent internal covariate shift during the learning process of the encoder, accelerate the learning convergence speed of the encoder, and improve the efficiency of aerial video recognition.

[0098] The first multi-layer perceptron can extract features from the video features processed by the second layer normalization module, obtain high-level features of the video features, and add the extracted high-level features to the video features input to the second layer normalization module through a residual block to obtain video features with local motion information, so as to improve the recognition accuracy of subsequent aerial videos. The specific calculation formula is as follows:

[0099]

[0100] where LN represents layer normalization and MLP represents multi-layer perceptron.

[0101] Such as Figure 3 the structural schematic diagram of the window semantic enhancement Transformer block shown, this Transformer block includes:

[0102] the above-mentioned local semantic enhancement encoder;

[0103] a standard encoder, which extracts features from the video features output by the local semantic enhancement encoder to obtain the global scene information and local scene information of the video features.

[0104] Specifically, the window semantic enhancement Transformer block includes two encoders in series, one of which is a local semantic enhancement encoder and the other is a standard encoder. The local semantic enhancement encoder can enhance the local motion information of the aerial video, and the standard encoder can extract features from the video features output by the local semantic enhancement encoder, so as to obtain the global scene information and local scene information of the video features for subsequent aerial video recognition.

[0105] In this embodiment, optionally, the standard encoder includes a third-layer normalization module, a multi-head attention module, a fourth-layer normalization module, and a second multi-layer perceptron connected in sequence.

[0106] Specifically, the standard encoder is a feed-forward neural network structure, including a third-layer normalization module, a multi-head attention module, a fourth-layer normalization module, and a second multi-layer perceptron connected in sequence. Among them, the third-layer normalization module can perform layer normalization on the video features output by the local semantic enhancement encoder. The multi-head attention module can calculate the multi-head attention of the video features processed by the third-layer normalization module and add the multi-head attention to the input video features through a residual block. The specific calculation formula is as follows:

[0107]

[0108] Among them, MSA represents the calculation operation of performing multi-head attention.

[0109] The fourth-layer normalization module can perform layer normalization on the video features with multi-head attention added, and extract the global scene information of the video features after layer normalization through the second multi-layer perceptron. The specific calculation formula is as follows:

[0110]

[0111] Such as Figure 4 The structural schematic diagram of the aerial video classification model shown, the classification model includes:

[0112] Multiple layers of the above window semantic enhancement Transformer blocks, configured to extract features from the aerial video to obtain the global scene information and local scene information of the aerial video;

[0113] The global-local self-fusion Transformer block, configured to fuse the global scene information and local scene information of the aerial video by using a self-fusion attention mechanism to obtain a fused video feature representation;

[0114] The classifier, configured to determine the classification result of the aerial video according to the fused video feature representation.

[0115] Specifically, the classification model includes a multi-layer window semantic enhancement Transformer block, a global-local self-fusion Transformer block, and a classifier. Among them, the window semantic enhancement Transformer block can enhance the local motion information of the aerial video and quickly and accurately extract the global scene information and local scene information of the enhanced aerial video. The local scene information includes the local class embedding of the aerial video, and the global scene information includes the global class embedding sequence of the aerial video. By averaging the features of the global class embedding, the global class embedding of the aerial video can be obtained. The global-local self-fusion Transformer block can fuse the local class embedding and the global class embedding extracted by the window semantic enhancement Transformer block using the self-fusion attention mechanism, so as to obtain the fused video feature representation. The classifier can accurately identify the category of the aerial video based on the fused video feature representation and obtain the classification result of the aerial video.

[0116] In this embodiment, optionally, the classification model further includes a pixel block embedding module, a position embedding module, and a global-local class embedding module. Among them, since the high resolution of the image is not conducive to self-attention modeling, the aerial video can be processed by the pixel block embedding module for pixel block embedding to reduce the image resolution. The specific method is to perform a convolution operation on the aerial video through the pixel block embedding module, so as to convert P×P pixels in the aerial video into a pixel block.

[0117] Self-attention is different from convolution and does not naturally have an inductive bias. Therefore, a learnable position encoding can be introduced into the aerial video through the position embedding module to better process the aerial video sequence. For the convenience of classification, global class embedding information and local class embedding information different from the video feature semantics can also be embedded into the aerial video through the global-local class embedding module.

[0118] In this embodiment, optionally, the global-local self-fusion Transformer block includes:

[0119] A Transformer encoder configured to preprocess the global scene information and local scene information of the aerial video;

[0120] A self-fusion attention module configured to fuse the preprocessed global scene information and local scene information to obtain the video feature representation.

[0121] Specifically, as Figure 5As shown in the figure, the global-local self-fusion Transformer block includes a Transformer encoder and a self-fusion attention module. Among them, the Transformer encoder is a standard encoder, and its specific composition is the same as the above-mentioned standard encoder. Through the Transformer encoder, a preliminary interaction can be established between the global scene information and the local scene information. The specific calculation formula is as follows:

[0122] CLS gl = Cat(CLS global , CLS local )

[0123] [CLS g , CLS l = MSA(LN(CLS gl )) + CLS gl ;

[0124] Among them, CLS local is the local class embedding, CLS global is the global class embedding, and the self-fusion attention module.

[0125] The self-fusion attention module can perform deep fusion on the preprocessed local class embedding and global class embedding, so as to achieve adaptive feature fusion based on the self-attention method and obtain the corresponding video feature representation.

[0126] It should be understood that the common fusion strategies include feature averaging and score averaging, which take the average of feature or score weights. This strategy is easy to implement but too simple and easy to ignore detailed features. There are also some adaptive fusion strategies, such as channel adaptation, adaptive weights, etc. These methods are often based on convolution and may not be suitable for the deep features of the Transformer network.

[0127] In this embodiment, optionally, the self-fusion attention module includes:

[0128] A query vector calculation unit, configured to respectively determine a global key-value vector and a local key-value vector through a single linear mapping according to the global scene information and the local scene information, and obtain a globally-local mixed query vector through channel concatenation and linear mapping according to the global key-value vector and the local key-value vector;

[0129] A probability distribution calculation unit, configured to calculate a probability distribution matrix according to the query vector, the global key-value vector, and the local key-value vector;

[0130] A video feature calculation unit, configured to calculate a fused video feature representation according to the probability distribution matrix, the global key-value vector, and the local key-value vector.

[0131] Specifically, asFigure 6 As shown, the self-fusion attention module includes a query vector calculation unit, a probability distribution calculation unit, and a video feature calculation unit. Among them, the query vector calculation unit can obtain the corresponding global key-value vectors V g , K g and local key-value vectors V l , K l respectively through single linear mappings based on local class embeddings and global class embeddings. The global-local mixed query vector Q m is obtained through channel concatenation and linear mapping. The specific calculation formula is as follows:

[0132]

[0133] In the formula, V g , K g , V l , K l , Q m ∈R 1×C ; denotes channel dimension concatenation; W represents the weight of the linear mapping, and R is the real number field.

[0134] The probability distribution calculation unit can determine the channel probability distribution matrix of the mixed vector pair for the double input vectors based on the global key-value vectors V g , K g and local key-value vectors V l , K l , as well as the global-local mixed query vector Q m . The specific calculation formula is as follows:

[0135]

[0136] In the formula, map ∈ R C×C represents the channel probability distribution matrix of the mixed vector pair for the double input vectors; denotes the dimension of vector K.

[0137] The video feature calculation unit can splice the channel probability distribution matrices map g , map l and value vectors V g , V l in pairs, and calculate the video feature representation after the fusion of global class embeddings and local class embeddings through matrix multiplication. The specific calculation formula is as follows:

[0138]

[0139] Among them, CLS fusion ∈R 1×2CIt represents the feature vector after multi-modal fusion, that is, the video feature representation after the fusion of global class embedding and local class embedding.

[0140] Since the self-fusion attention module mixes the query vectors of the input features through a mixed linear mapping and performs self-fusion with the key and value vectors of their respective features. Therefore, compared with the existing feature averaging, score averaging, channel adaptation, and adaptive weights, the self-fusion attention module can obtain a richer video feature representation, which helps to further improve the accuracy of aerial video recognition.

[0141] After obtaining the video feature representation after the fusion of global class embedding and local class embedding, the video feature representation can be input into a classifier, and the final classification result of the video can be obtained through the classifier. The specific calculation formula is as follows:

[0142] C = Max(FC(CLS fusion ));

[0143] where FC(·) is a fully connected layer classifier with an input dimension of 768 and an output dimension of the total number of classes, Max(·) represents taking the maximum score result for all, and C is the final classification category of the video.

[0144] As Figure 7 shown in the flowchart of the aerial video classification method, the classification method includes:

[0145] Obtain aerial data and preprocess the aerial data to obtain an aerial video;

[0146] Use the trained aerial video classification model as described above to classify the preprocessed aerial video.

[0147] Specifically, first, aerial data collected by a drone can be obtained, and the aerial data can be preprocessed to obtain the corresponding aerial video. In the embodiments of the present invention, the acquisition method of aerial video data is not limited. The method of self-collecting videos includes: controlling the drone to shoot a target scene or object from multiple angles, for a long time, with partial occlusion, and under different lighting conditions, and selecting multiple different target scenes or objects for shooting. The method of obtaining sample videos through a computer includes: searching the network for video segments containing the target scene or object as aerial video data, or directly obtaining them from existing drone aerial video datasets (such as ERA, MOD20, UAVhuman).

[0148] Specifically, first, the aerial photography data can be clipped to delete the blurred, unstable, and invalid video frames in the aerial photography data, so as to obtain high-quality and equal-length video clips. Then, the video sampling frequency can be adjusted according to the existing computing resources, and the corresponding video frames can be extracted from the video clips according to the adjusted video sampling frequency to generate a video frame sequence of the corresponding length as the aerial photography video, and the resolution of the video frames can be adjusted to a preset size. Then, the preprocessed aerial photography video is input into the trained aerial photography video classification model, and the trained aerial photography video classification model can classify the aerial photography video.

[0149] In this embodiment, the training process of the aerial photography video classification model includes:

[0150] First, multiple groups of aerial photography video data are collected, and the same method as the above preprocessing is used to preprocess each group of aerial photography video data, so as to obtain the corresponding aerial photography videos to construct a data set. When preprocessing the aerial photography video data, according to the requirements and computing conditions, these aerial photography video data can be clipped into effective and equal-length video clips, which are used as the training set and the test set respectively, to ensure that the data results collected at the same frame rate are consistent. To avoid the impact of data imbalance on model training, the number of videos of different categories should be as equal as possible. In addition, to avoid the influence of homologous videos on the test results, multiple segments clipped from the same video need to be put into the same training set or test set to ensure the training effect of the model.

[0151] It should be understood that in this embodiment, there is no limit to the number of samples in the data set, and the number of sample videos can be one or multiple. Exemplarily, the number of sample videos is multiple to ensure the model training effect. Exemplarily, for the case where the number of sample videos is multiple, the number of video frames in different sample videos is the same to ensure the model training effect.

[0152] After that, the above-mentioned aerial photography video classification model can be trained with the training set. During the training process, the parameters of the aerial photography video classification model can be continuously updated through the loss function until the loss function is reduced to the lowest, and then the trained aerial photography video classification model can be verified with the test set. It can be understood that the training, verification, and testing processes of the model have similar or the same processing flows. Those skilled in the art should know that in the verification and testing stages, the preprocessed aerial photography video data can be processed according to the same process as the training to obtain the corresponding aerial photography video classification results.

[0153] In this embodiment, the parameters of the model can be updated by backpropagation of gradients, the cross-entropy loss function is used as the loss function, and the stochastic gradient descent optimizer is used for updating. The stochastic gradient descent optimizer can select the cosine annealing strategy to keep the learning rate updated linearly. The cross-entropy loss function is specifically:

[0154]

[0155] Among them, N represents the total number of categories of training data; q i represents the probability that the model is trained for the i-th category, and q i ∈{0, 1}; p i represents the probability that the model predicts the i-th category, and p i ∈{0, 1},

[0156] By continuously calculating the loss function and using backpropagation to update the network parameters, iterative optimization is continuously carried out to improve the recognition accuracy of the model. When the loss function drops to the lowest point, the model training is completed.

[0157] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered by the scope of the claims and the description of the present invention.

Claims

1. An aerial video classification model, characterized in that: include: Multi-layer window semantic enhancement Transformer block, global local self-fusion and Transformer block and classifier; The window semantic enhancement Transformer block includes a local semantic enhancement encoder and a standard encoder; Local semantic enhancement encoder, including: A window positioning module is configured to use an unfilled convolution kernel with the same size as the window to locate a key window area of ​​the video feature and strip off the local video features within the key window area; A windowed time multi-head self-attention module is configured to calculate the windowed time multi-head self-attention of the local video feature and perform a residual connection between the windowed time multi-head self-attention and the video feature; A standard encoder is used to extract features of the video features output by the local semantic enhancement encoder to obtain global scene information and local scene information of the video features; The standard encoder includes a third-layer normalization module, a multi-head attention module, a fourth-layer normalization module and a second multi-layer perceptron connected in sequence; The output of the multi-head attention module is connected to the input of the standard encoder through a residual block; The output of the second multi-layer perceptron is connected to the input of the fourth layer normalization module through a residual block; The global local self-fusion and Transformer block includes a self-fusion attention module, which includes: A query vector calculation unit is configured to determine a global key value vector and a local key value vector respectively through a single linear mapping according to the global scene information and the local scene information, and obtain a global-local mixed query vector through channel splicing and linear mapping according to the global key value vector and the local key value vector; A probability distribution calculation unit, configured to calculate a probability distribution matrix according to the query vector, the global key value vector and the local key value vector; A video feature calculation unit, configured to calculate a fused video feature representation according to the probability distribution matrix, the global key value vector and the local key value vector; The classifier is configured to determine the classification result of the aerial video according to the fused video feature representation.

2. The aerial video classification model according to claim 1, characterized in that: The window positioning module comprises: A video pooling unit, configured to perform global average pooling and global maximum pooling on the input video features respectively to obtain corresponding pooling features; A channel splicing unit, configured to splice the two pooling features obtained by the video pooling unit to obtain a corresponding splicing feature; A feature response unit is configured to calculate feature responses of the spliced ​​features using an unfilled convolution kernel with a consistent window area size, and normalize each obtained feature response to determine a weight of each feature response; A window position unit, configured to perform global average pooling on the weights of all feature responses based on the time dimension, and locate the key window position of the video feature according to the feature response with the largest average weight; A window area unit is configured to determine the key window area according to the key window position and the size of the unfilled convolution kernel.

3. The aerial video classification model according to claim 1, characterized in that: Also includes: A first-layer normalization module configured to perform normalization processing on the video features and input the processed video features into the window positioning module; A second layer normalization module is configured to perform normalization processing on the video features added with the window time multi-head self-attention; The first multi-layer perceptron is configured to extract features of the video output by the second-layer normalization module, and the extracted features are residually connected with the video features input to the second-layer normalization module.

4. The aerial video classification model according to claim 1, characterized in that: The global local self-fusion and Transformer block includes: The Transformer encoder is configured to preprocess the global scene information and the local scene information of the aerial video.

5. A method for classifying aerial videos, characterized in that: include: Acquire aerial photography data, and pre-process the aerial photography data to obtain an aerial video; The preprocessed aerial video is classified using the trained aerial video classification model as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Column semantic recognition method and system based on context awareness of GCN and RoBERTa

    CN117312989A