Aerial image target detection method based on MAE
By constructing a MAE-based aerial image object detection model, combining labeled and unlabeled image training, the problem of low aerial image object detection accuracy is solved by using Linformer self-attention mechanism and joint loss function, and efficient and accurate rotation frame object detection is achieved.
Patent Information
- Application Number
- CN202510337641.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-22
AI Technical Summary
In the prior art, the aerial image object detection efficiency is low and the unlabeled image cannot be effectively utilized, resulting in a low target detection accuracy.
Aerial image object detection model based on MAE is constructed, and the annotated and unlabeled images are jointly trained, and the Linformer self-attention mechanism and joint loss function are used to optimize the model training process and output the rotation box object detection results.
It improves the accuracy and efficiency of aerial image target detection, reduces model training costs, and is suitable for situations where target angles change greatly in aerial images, reduces the computational complexity, and is easy to deploy on drone edge devices.
Smart Images

Figure CN120356117A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image target detection, and specifically relates to a method for detecting aerial image targets based on MAE. Background Art
[0002] Aerial photography by unmanned aerial vehicle (UAV) is a photography method that relies on the UAV to take pictures of the target object from the air. This method can obtain a large number of aerial images containing the target object. By detecting the target object in these aerial images, information such as the position, shape, and size of the target object can be obtained. However, the target objects contained in these photos often have small pixels, and directly performing target detection often results in missed detection of the target object, leading to low accuracy of aerial image target detection.
[0003] Currently, mainly by manually annotating the target objects in aerial images, the problem of missed detection of the target during the target detection process of these images can be reduced to a certain extent. However, the number of aerial images is huge. Using the method of manual annotation usually has low efficiency and cannot annotate all the target objects contained in these aerial images in a short time. During actual target detection, usually only the annotated aerial images are used, and the remaining large number of unannotated original aerial images cannot be used for aerial image target detection, but this easily leads to incomplete information such as the position, shape, and size of the target object.
[0004] Therefore, how to use unannotated aerial images for target detection without increasing labor costs has always been an urgent problem for those skilled in the art. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for detecting aerial image targets based on MAE in response to the deficiencies of the prior art. By constructing a model for detecting aerial image targets based on MAE, and using unannotated images and annotated images to jointly train the model for detecting aerial image targets, and then using the trained model for detecting aerial image targets to detect the targets in aerial images, the target detection of aerial images can be achieved.
[0006] The purpose of the present invention is achieved by the following scheme:
[0007] A method for detecting aerial image targets based on MAE includes the following steps:
[0008] 1) Establish a model for detecting aerial image targets based on MAE for identifying the targets contained in aerial images;
[0009] 2) Input unannotated aerial images and divide the aerial images into several image patches;
[0010] 3) Input the several image patches obtained in step 2) into the aerial image target detection model, and use the aerial image target detection model to perform target detection;
[0011] 4) Fuse the results detected in step 3), remove the redundant detection frames in the overlapping part, and perform visual output on the fused results.
[0012] Preferably, the aerial image target detection model is constructed in the following manner:
[0013] 1-1) Construct an MAE encoder for extracting the features of the aerial image;
[0014] 1-2) Construct an aerial image target detection model based on the MAE encoder;
[0015] 1-3) Train the aerial image target detection model.
[0016] Preferably, in step 1-2), the specific manner of constructing the aerial image target detection model based on the MAE encoder includes:
[0017] 1-2-1) Add a prediction head to the MAE encoder and adjust the size of the finally output rotated box;
[0018] 1-2-2) Use the generalized intersection over union loss as the regression loss and the cross-entropy loss function as the classification loss;
[0019] 1-2-3) Construct a joint training loss function according to the regression loss function and the classification loss function:
[0020] Loss = α·L MSE + β·L GIoU + λ·L cls
[0021] In the formula, Loss is the joint training loss function, α is the first hyperparameter for adjusting each loss, L MSE is the image reconstruction loss, β is the second hyperparameter for adjusting each loss, L GIoU is the regression loss, λ is the third hyperparameter for adjusting each loss, L cls is the classification loss;
[0022] 1-2-4) Combine the joint training loss function constructed in step 1-2-2), use the MAE editor as the feature extractor, and construct an aerial image target detection model.
[0023] Preferably, in step 1-3), the specific manner of training the aerial image target detection model includes:
[0024] 1-3-1) Input several labeled aerial images and several unlabeled aerial images;
[0025] 1-3-2) Use the labeled aerial images and unlabeled aerial images to alternately train the aerial image object detection model for several rounds;
[0026] 1-3-3) Use the labeled aerial images to train the last Linformer block and the prediction head part of the aerial image object detection model for several rounds.
[0027] Preferably, the training process of the aerial image object detection model specifically includes:
[0028] S1) Crop the aerial image and divide the cropped aerial image into several image patches to form a pre-training set;
[0029] S2) Randomly select several image patches in the pre-training set for masking to obtain several masked image patches and several unmasked image patches, and represent each masked image patch as a mask label;
[0030] S3) Map each image patch in the pre-training set to an embedding vector and calculate the position encoding of each image patch in the aerial image;
[0031] S4) Use the encoder to calculate the feature vectors of each unmasked image patch;
[0032] S5) Recombine the feature vectors of each unmasked image patch and the mask symbols of each masked image patch according to the position encoding calculated in step S3) to obtain the feature sequence of the aerial image;
[0033] S6) Adjust the output of the decoder to the picture pixel distribution, and use the decoder to reconstruct the aerial image based on the feature sequence of the aerial image to obtain a reconstructed image;
[0034] S7) Adopt the mean square error loss function to calculate the loss value of the masked image patches between the reconstructed image and the aerial image;
[0035] S8) According to the calculated loss value, combined with the backpropagation algorithm, use the optimizer to update each hyperparameter of the model.
[0036] Preferably, the aerial image object detection model includes an encoder, a decoder, and a prediction head.
[0037] Preferably, the encoder consists of 6 Linformer blocks, each Linformer block containing 8 self-attention heads and a feature representation of 512 dimensions; the decoder consists of 1 lightweight Linformer block, each Linformer block containing 8 self-attention heads and a feature representation of 512 dimensions; the prediction head adjusts the output of the Linformer to a prediction output with a size of 80×80×20.
[0038] Preferably, in the Linformer block, a linear projection matrix with a fixed dimension k = 128 is used in its self-attention mechanism to compress the input sequence from the original length n = 1600 to 128.
[0039] Preferably, a first training round threshold is set to limit the number of times of alternately training the aerial image target detection model using the unlabeled training set and the labeled training set, and a second training round threshold is set to limit the number of times of training the aerial image target detection model using only the labeled training set.
[0040] Preferably, the value range of the first training round threshold is 150±10, and the value range of the second training round threshold is 50±10.
[0041] The beneficial effects of the present invention are as follows:
[0042] ① In the present invention, by using a linear projection matrix with a fixed dimension k = 128 in the Linformer self-attention mechanism and compressing the input sequence from the original length n = 1600 to 128, the computational complexity can be reduced, thereby optimizing the computational efficiency and making it easier for the MAE model to learn and optimize during the training process;
[0043] ② In the present invention, the labeled training set clarifies information such as the position and category of the targets in the aerial images. Training the aerial image target detection model in combination with the unlabeled training set in the early stage can enable the model to have a certain feature learning ability, and then training with the labeled data can enable the aerial image target detection model to accurately learn the features of the targets according to the annotation information, thereby improving the accuracy of target detection and the precision of classification;
[0044] ③ By setting the first training times threshold and the second training times threshold, the present invention can prevent abnormal situations such as overfitting and even model degradation of the aerial image target detection model while ensuring relatively high training efficiency of the model.
[0045] The advantages of the present invention are as follows:
[0046] ① The MAE algorithm is carried out on general object detection tasks, and the detected bounding boxes it outputs are horizontal rectangular boxes. The MAE-based aerial image object detection model in the present invention is for aerial images. It can output rectangular boxes with angles, is suitable for the situation where the angles of objects in aerial images change greatly, is more accurate in object localization in aerial images, has a higher effective pixel of the object contained in the box, and can improve the accuracy of aerial image object detection without increasing labor.
[0047] ② In the present invention, the training of MAE is mainly divided into the pre-training and fine-tuning processes. By weighting the reconstruction Loss of unlabeled images and the target Loss of labeled ones, the model training can be completed only once, which can avoid the catastrophic forgetting caused by subsequent fine-tuning and can greatly reduce the training cost and later maintenance cost of the model.
[0048] ③ The implementation of MAE is based on ViT (Vision Transformer), and ViT has a large number of parameters and high computational complexity. The present invention introduces Linformer. As a lightweight variant of Transformer, Linformer reduces the complexity of the self-attention mechanism from O(n2) to O(nk) (where k is the dimension of the low-rank projection), significantly reducing the number of parameters of the model and helping to deploy this algorithm on the edge devices of drones.
[0049] Explanation of Terms
[0050] MAE: Masked Autoencoder, a neural network architecture based on self-supervised learning. The core is to perform a masking operation on the input image data, and then let the model learn to reconstruct the original image data from the partially visible image data, so as to learn the internal structure and feature representation of the data.
[0051] Generalized Intersection over Union Loss (GIoU Loss): A loss function commonly used in object detection tasks to measure the difference between the predicted bounding box and the ground truth bounding box.
[0052] Focal Loss: A loss function designed to solve the problems of sample class imbalance and uneven learning of easy and difficult samples in object detection. Its core idea is to make the model pay more attention to difficult-to-classify samples and positive samples through a modulation factor.
[0053] FLOPs: (Floating-point Operations Per Second) is an important metric for measuring the computing power and complexity of a computer or a deep learning model. In the field of deep learning, it usually refers to the total amount of floating-point operations required for a single forward pass of the model, which is used to evaluate the computational efficiency and hardware requirements of the model.
[0054] Params: (Parameters) is one of the core attributes of a deep learning model, referring to the total number of all learnable parameters in the model (usually in millions, such as M). It directly reflects the complexity, storage requirements, and potential expressive power of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is a flowchart of the present invention;
[0056] Figure 2 is a schematic diagram of image slicing in this embodiment;
[0057] Figure 3 is a schematic diagram of the structure of the mask autoencoder in this embodiment;
[0058] Figure 4 is a schematic diagram of target detection for unlabeled images in this embodiment;
[0059] Figure 5 is a schematic diagram of the visual output of target detection for aerial images using the aerial image target detection model in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] As Figures 1 to 5 shown, a method for aerial image target detection based on MAE includes the following steps:
[0061] 1) Establish an aerial image target detection model based on MAE for identifying targets contained in aerial images;
[0062] 2) Input unlabeled aerial images and slice the aerial images into a number of image patches;
[0063] 3) Input the number of image patches obtained in step 2) into the aerial image target detection model and perform target detection using the aerial image target detection model;
[0064] 4) Fuse the results detected in step 3), remove the redundant detection boxes in the overlapping parts, and perform visual output on the fused results.
[0065] According to the above method, the following is an embodiment:
[0066] 1) The specific ways to establish a MAE-based aviation image target detection model for identifying targets contained in aviation images include:
[0067] 1-1) Construct a MAE encoder for extracting features of aviation images. The MAE encoder mainly includes an encoder and a decoder. The encoder consists of 6 Linformer blocks, each block containing 8 self-attention heads and a 512-dimensional feature representation. The decoder uses 1 lightweight Linformer block, maintaining the same number of attention heads and feature dimensions as the encoder;
[0068] In this implementation, to further optimize the computational efficiency, a linear projection matrix with a fixed dimension k = 128 is used in the Linformer self-attention mechanism to compress the input sequence from the original length n = 1600 to 128.
[0069] 1-2) The specific ways to construct an aviation image target detection model based on the MAE encoder include:
[0070] The loss function of the MAE-based aviation image target detection model mainly consists of two parts. The first part is the reconstruction loss of the image by the encoder and the decoder. Among them, the image reconstruction loss uses the mean squared error (MSE) loss, and the loss is only calculated for the masked part. The second part is the target box regression and target category loss of the labeled targets, which are specifically as follows:
[0071] 1-2-1) Add a prediction head to the MAE encoder and adjust the size of the finally output rotated box;
[0072] In this embodiment, the prediction head adjusts the output of Linformer to the prediction output, with a size of 80×80×20, where 80×80 is the grid division and 20 is the regression of 15 categories and (x, y, w, h, r) of the DOTA dataset. In rotated box target detection, the loss function of the target box needs to adapt to the characteristics of the rotated box. The definition of the rotated box includes the center point coordinates (x, y), width w, height h, and rotation angle r. In this embodiment, the regression loss uses the GIoU loss. IoU reflects the ratio between the intersection and union of two detection boxes, so it is independent of the size of the detection box and helps to improve the detection effect of small targets in aviation images.
[0073] 1-2-2) Use the generalized intersection over union loss as the regression loss and the cross-entropy loss function as the classification loss;
[0074] The formula for the GIoU loss is as follows:
[0075]
[0076] Wherein, L GIoU is the regression loss, 1 is the intersection area of the predicted bounding box Bp and the true bounding box Bg, U is the union area of Bp and Bg, C is the area of the smallest closed rectangle containing Bp and Bg, and IoU is the intersection over union of the true box and the predicted box.
[0077] The classification loss adopts Focal Loss, which can effectively handle the class imbalance problem of the aerial image dataset. The formula for the classification loss is as follows:
[0078]
[0079] Wherein, L cls is the classification loss, N is the number of all target boxes in the image, c is the number of target classes, y i,c is the true label that the i-th box belongs to class c, the predicted probability that the i-th box belongs to class c, α c is the class weight that can control the imbalanced classes, and γ is the weight for adjusting the easy-to-classify samples and the hard-to-classify samples (the value set in this embodiment is 2).
[0080] 1-2-3) Construct a joint training loss function according to the regression loss function and the classification loss function;
[0081] Loss = α·L MSE + β·L GIoU + λ·L cls
[0082] Wherein, Loss is the joint training loss function, α is the first hyperparameter for adjusting each loss, L MSE is the image reconstruction loss, β is the second hyperparameter for adjusting each loss, L GIoU is the regression loss, λ is the third hyperparameter for adjusting each loss, L cls is the classification loss;
[0083] In this embodiment, α, β, and λ can be regarded as weight parameters, and α is set to 0.5, β is set to 0.8, and λ is set to 0.5 respectively.
[0084] 1-2-4) Combine the joint training loss function constructed in step 1-2-2), use the MAE editor as the feature extractor, and build an aerial image target detection model.
[0085] The aerial image target detection model constructed by the above method mainly includes an encoder, a decoder, and a prediction head. The encoder consists of 6 Linformer blocks, each Linformer block contains 8 self-attention heads and a feature representation of 512 dimensions; the decoder consists of 1 lightweight Linformer block, each Linformer block contains 8 self-attention heads and a feature representation of 512 dimensions; the prediction head adjusts the output of Linformer to a prediction output with a size of 80×80×20, where 80×80 is the grid division and 20 is the regression of 15 categories in the DOTA dataset and (x, y, w, h, r). In the Linformer block, a linear projection matrix with a fixed dimension k = 128 is used in its self-attention mechanism to compress the input sequence from the original length n = 1600 to 128.
[0086] Table 1 Network structure of the aerial image target detection model based on MAE
[0087] layer component number of parameters linear projection (16×16×3)×512 6291456 encoder 6 Linformer Blocks 1179480 Linformer Block multi-head attention 1966080 decoder 1 Linformer Block 1966080 prediction head Reshape layer (80×80×20) 10240 total number of parameters linear projection + encoder + decoder + prediction head 20065256
[0088] 1-3) The specific method for training the aerial image target detection model includes:
[0089] 1-3-1) Input 1411 labeled aerial images and 2000 unlabeled aerial images. The labeled aerial images are obtained by manually annotating aerial images, and the unlabeled aerial images are the original aerial images collected by drone aerial photography;
[0090] 1-3-2) Use the labeled aerial images and unlabeled aerial images to alternately train the aerial image target detection model for several rounds;
[0091] 1-3-3) Use the labeled aerial images to train only the last Linformer block and the prediction head part of the aerial image target detection model for several rounds by freezing other parts of the model except the last Linformer block and the prediction head.
[0092] In this embodiment, during the process of training the aerial image target detection model, a first training round threshold is also set to limit the number of times of alternately training the aerial image target detection model using the unlabeled training set and the labeled training set. The value range of the first training round threshold is 150±10; a second training round threshold is set to limit the number of times of training the aerial image target detection model only using the labeled training set. The value range of the second training round threshold is 50±10.
[0093] In this embodiment, the training process of the aerial image target detection model specifically includes:
[0094] S1) Let the input aerial image be \(x\in\mathbb{R}\) H×W×C , crop the aerial image into pictures with a resolution of 640×640, and cut the cropped aerial image into 1600 image patches (patches) of size 16x16 pixels. Each patch is of size 16×16×3 dimensions to form a pre-training set;
[0095] S2) Use a random Mask with a ratio of 75%, put the picture patches in the pre-training set into a list, and randomly select 25% of the image patches (i.e., 400 image patches) in the pre-training set for masking to obtain a number of masked image patches and a number of unmasked image patches. The set of masked image patches (patches) is of size \(|M|\). The model reconstructs the masked patch as The true masked patch is \(X\) m . Represent each masked image patch as a mask label respectively;
[0096] S3) Map each image patch in the pre-training set into a one-dimensional embedding vector respectively for the attention mechanism calculation inside the Transformer block, and calculate the position encoding of each image patch in the aerial image;
[0097] ① Embed the sliced picture patches. The dimension of each picture patch is 16×16×3 (i.e., \(P\times P\times C\)), and the embedding dimension is 512. Therefore, the dimension of the tensor after passing through the linear projection layer is 400×512.
[0098] ② Position encoding of the picture patches. The picture encoding uses the cosine and sine methods, as shown in the following formula. The dimension of the encoding is the same as the picture patch embedding dimension, and its output uses the direct addition method. Therefore, the dimension of the output tensor of this module remains unchanged, still maintaining 400×512 dimensions.
[0099]
[0100] In the formula, \(PE\) (pos,2i) is the value of the position encoding calculated using the sine function (\(\sin\)) at even dimensions, \(PE\) (pos,2i+1) is the value of the position encoding calculated using the cosine function (\(\cos\)) at odd dimensions, pos represents the position of an element in the sequence (Position), i represents the encoding dimension index (i.e., the \(i\)-th dimension), and \(d\) dmodel is the dimension of the hidden layer of the model;
[0101] S4) Use the encoder to calculate the feature vectors of each unmasked image patch. The input of the encoder is the unmasked image patch, and the number of the unmasked image patches is 25% of the total picture patches, a total of 400 patches;
[0102] S5) Recombine the feature vectors of each unmasked image patch and the mask symbols of each masked image patch according to the position encoding calculated in step S3) to obtain the feature sequence of the aerial image;
[0103] It should be noted that the encoder is composed of 6 stacked Linformer blocks, and the dimension inside each Linformer block is 512. Therefore, the output dimension after stacking multiple modules remains unchanged, still 400×512 dimensions; the decoder is composed of two stacked Linformer blocks and a reshape module. The size of the Linformer block is the same as that of the encoder. Reshape projects the image patches back to 400×768 dimensions through a fully connected layer.
[0104] S6) Adjust the output of the decoder to the pixel distribution of the image, that is, adjust the 400×768-dimensional tensor output by the decoder to 20×20 image patches of size 16×16×3. Use the decoder to reconstruct the aerial image based on the feature sequence of the aerial image to obtain the reconstructed image;
[0105] S7) Use the MSE mean squared error loss function to calculate the loss value of the masked image patches between the reconstructed image and the aerial image;
[0106] The formula for MSE Loss is:
[0107]
[0108] In the formula, |M| is the number of masked patches, X m is the real masked image patch, is the masked image patch predicted by the model, is the square of the Euclidean distance between the real masked image patch and the masked image patch predicted by the model.
[0109] S8) According to the calculated loss value, combined with the backpropagation algorithm, use the optimizer to update each hyperparameter of the model.
[0110] In this embodiment, the model training parameters are shown in Table 2 (in Table 2, e is a symbol used to represent the power of 10 in scientific notation, such as 1e-6 which is 0.000001). The number of training rounds is 200. In the first 150 rounds, unlabeled aerial images and labeled aerial images are alternately used for training, and only the reconstruction loss is calculated for the unlabeled data. After 150 rounds, only labeled data is used for training. And, all images will be trained in each round, that is, in the first round, 1411 labeled aerial images are used to train the model, in the second round, 2000 unlabeled aerial images are used to train the model, and so on alternately until after 150 rounds, 1411 labeled aerial images are used to train the model for 50 rounds.
[0111] Table 2 Model Training Parameters
[0112] parameter name value optimizer AdamW initial learning rate 1e-4 weight decay 1e-4 Beta parameter <![CDATA[β1 = 0.9, β2 = 0.999]]> minimum learning rate 1e-6 Warm-up steps 20 batch size (unlabeled) 64 batch size (labeled) 32
[0113] It should be noted that in this embodiment, an image segmentation and position encoding module, an asymmetric autoencoder training module and a network prediction head output module, and a reconstruction of unlabeled images and a joint loss calculation module for labeled targets are set. The asymmetric autoencoder is mainly composed of a Linformer model. The definitions and uses of these modules are as follows:
[0114] a. Image Segmentation and Position Encoding Module
[0115] Definition: This module is used to segment the input aerial image and add position encoding to the segmented image blocks to make them suitable for the input of the Transformer structure.
[0116] Use:
[0117] The input image is segmented into multiple small blocks in units of 16×16 pixels to adapt to the self-attention mechanism.
[0118] Through sine-cosine position encoding, position information is added to each image block to help the model better understand the spatial relationship.
[0119] b. Asymmetric Autoencoder Training Module and Network Prediction Head Output Module
[0120] Definition: This module consists of an encoder, a decoder and a prediction head. It uses an asymmetric autoencoder structure for feature learning and finally outputs the rotation box target detection result.
[0121] Use:
[0122] Encoder: Use 6 layers of Linformer blocks to extract features from the input image.
[0123] Decoder: Use 1 layer of Linformer block to reconstruct the original image based on the features of the encoder.
[0124] Prediction Head: Adjust the features output by the encoder and finally output a rotation box prediction result of 80×80×20. Among them, the 20 dimensions include the category and bounding box parameters (x, y, w, h, r). x is the central coordinate x of the minimum bounding rectangle of the detection target, indicating the position of the geometric center of the rectangle in the image; y is the central coordinate y of the minimum bounding rectangle of the detection target, indicating the position of the geometric center of the rectangle in the image; w is the length of the bounding rectangle along its short axis; h is the length of the bounding rectangle along its long axis; r is the angle between the height direction (i.e., the long axis of the rectangle) of the bounding rectangle and the x-axis, measured counterclockwise, and the unit is degree.
[0125] c. Reconstruction of Unlabeled Images and Joint Loss Calculation Module for Labeled Targets
[0126] Definition: This module is used to calculate the reconstruction loss of self-supervised learning and the object detection loss of supervised learning, and perform joint training through a weighted method.
[0127] Purpose:
[0128] The image reconstruction loss (MSE Loss) is used for unlabeled data, and the loss is only calculated for the masked part to train the feature extraction ability of the encoder.
[0129] Object detection loss: It includes the GIoU loss of the rotated bounding box (used for bounding box optimization) and the Focal Loss (used for handling class imbalance).
[0130] The joint training strategy uses unlabeled and labeled data alternately in the first 150 rounds and only uses labeled data in the last 50 rounds.
[0131] The above modules together constitute an aviation image object detection model based on MAE (Masked Autoencoder) (i.e., the aviation image rotated bounding box object detection algorithm) to improve the utilization rate of unlabeled data and enhance the object detection performance.
[0132] 2) Input the aviation image and divide the aviation image into several image patches;
[0133] 3) Input the several image patches obtained in step 2) into the aviation image object detection model, and use the aviation image object detection model to perform object detection;
[0134] In this embodiment, when using the aviation image object detection model to perform object detection on the image patches, if there is no target object, the content of the image patch will not be boxed. Only when the image patch contains a target object, the aviation image object detection model will box the target object in the image patch and output this target object as the detected result. That is to say, in step 3), this kind of processing will be performed on several image patches to obtain several detection results.
[0135] 4) Fuse the results detected in step 3), remove the redundant detection boxes in the overlapping parts, and visually output the fused results.
[0136] In this embodiment, the inference device used is NVIDIA GeForce RTX 2080Ti with a video memory of 12GB. By saving and uploading the inference results to the DOTA dataset server for result comparison, the result mAP50 is obtained. From the comparison experiment results, it can be seen that by adding 2000 unlabeled data to Model_2, compared with not adding unlabeled data, the evaluation index mAP50 has increased by 0.8, that is, the mean average precision at the intersection over union (IoU) threshold of 0.5 has increased by 0.8.
[0137] Table 2 Comparison of Experimental Results
[0138]
[0139] Model_1: The object detection model established using the present invention. The training set of this model is 1411 labeled aerial images. The speed of detecting one image (i.e., Speed) is 210ms, the number of parameters (i.e., params) is 20.1M, the number of floating-point operations per second, i.e., FLOPs(B) is 21.25, and the mean average precision (i.e., mAP50) of this model at an intersection over union of 50% is 79.6.
[0140] Model_2: The object detection model established using the present invention. The training set of this model is 1411 labeled aerial images and 2000 unlabeled aerial images. The speed of detecting one image (i.e., Speed) is 210ms, the number of parameters (i.e., params) is 20.1M, the number of floating-point operations per second, i.e., FLOPs(B) is 21.25, and the mean average precision (i.e., mAP50) of this model at an intersection over union of 50% is 80.4.
[0141] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications made by those skilled in the art without departing from the spirit of the present invention fall within the protection scope of the present invention.
Claims
1. A method for detecting aviation image targets based on MAE, characterized in that, It includes the following steps: 1) Establish an MAE-based aerial image target detection model for identifying the targets contained in aerial images; 2) Input unannotated aerial images and divide the aerial images into several image patches; 3) Input the several image patches obtained in step 2) into the aerial image target detection model, and use the aerial image target detection model for target detection; 4) Fuse the results detected in step 3), remove the redundant detection frames in the overlapping parts, and visually output the fused results.
2. The aviation image target detection method according to claim 1, characterized in that The aerial image target detection model is constructed in the following way: 1-1) Construct an MAE encoder for extracting the features of aerial images; 1-2) Construct an aerial image target detection model according to the MAE encoder; 1-3) Train the aerial image target detection model.
3. The aviation image target detection method according to claim 2, wherein In step 1-2), the specific way to construct the aerial image target detection model according to the MAE encoder includes: 1-2-1) Add a prediction head to the MAE encoder and adjust the size of the finally output rotated box; 1-2-2) Use the generalized intersection over union loss as the regression loss and the cross-entropy loss function as the classification loss; 1-2-3) Construct a joint training loss function according to the regression loss function and the classification loss function: Loss = α·LMSE + β·LGIoU + λ·Lcls where Loss is the joint training loss function, α is the first hyperparameter for adjusting each loss, L MSE is the image reconstruction loss, β is the second hyperparameter for adjusting each loss, L GIoU is the regression loss, λ is the third hyperparameter for adjusting each loss, L cls is the classification loss; 1-2-4) Combine the joint training loss function constructed in step 1-2-2), use the MAE editor as a feature extractor, and construct an aerial image target detection model.
4. The aviation image target detection method according to claim 2, wherein In step 1-3), the specific way to train the aerial image target detection model includes: 1-3-1) Input several annotated aerial images and several unannotated aerial images; 1-3-2) Use the annotated aerial images and unannotated aerial images to alternately train the aerial image target detection model for several rounds; 1-3-3) Use the annotated aerial images to train the last Linformer block and the prediction head part of the aerial image target detection model for several rounds.
5. The aviation image target detection method according to claim 4, characterized in that The training process of the aerial image target detection model specifically includes: S1) Crop the aerial images and divide the cropped aerial images into several image patches to form a pre-training set; S2) Randomly select several image patches in the pre-training set for masking to obtain several masked image patches and several unmasked image patches, and represent each masked image patch as a mask label; S3) Map each image patch in the pre-training set to an embedding vector respectively, and calculate the position encoding of each image patch in the aerial image; S4) Use the encoder to calculate the feature vectors of each unmasked image patch; S5) Recombine the feature vectors of each unmasked image patch and the mask symbols of each masked image patch according to the position encoding calculated in step S3) to obtain the feature sequence of the aerial image; S6) Adjust the output of the decoder to the pixel distribution of the picture, and use the decoder to reconstruct the aerial image based on the feature sequence of the aerial image to obtain a reconstructed image; S7) Use the mean square error loss function to calculate the loss value of the masked image patches between the reconstructed image and the aerial image; S8) Update each hyperparameter of the model using an optimizer in combination with the backpropagation algorithm based on the calculated loss value.
6. The aviation image target detection method according to claim 2, characterized in that, The aerial image target detection model includes an encoder, a decoder, and a prediction head.
7. The aviation image target detection method according to claim 6, wherein The encoder consists of 6 Linformer blocks, each Linformer block containing 8 self-attention heads and a feature representation of 512 dimensions; the decoder consists of 1 lightweight Linformer block, each Linformer block containing 8 self-attention heads and a feature representation of 512 dimensions; the prediction head adjusts the output of the Linformer to a prediction output with a size of 80×80×20.
8. The aviation image target detection method according to claim 7, characterized in that, In the Linformer block, a linear projection matrix with a fixed dimension k = 128 is used in its self-attention mechanism to compress the input sequence from the original length n = 1600 to 128.
9. The aviation image target detection method according to claim 4, characterized in that Set a first training round threshold for limiting the number of times of alternately training the aerial image target detection model using the unlabeled training set and the labeled training set, and set a second training round threshold for limiting the number of times of training the aerial image target detection model only using the labeled training set.
10. The aviation image target detection method according to claim 9, wherein, The value range of the first training round threshold is 150±10, and the value range of the second training round threshold is 50±10.