Improved Remote Sensing Image Aircraft Target Detection Model Based on YOLOv5 and Its Method
By improving the remote sensing image aircraft target detection model of YOLOv5, combining feature extraction, fusion and detection heads, the problems of slow detection speed and low accuracy in traditional methods are solved, and more efficient aircraft target detection is achieved.
Patent Information
- Application Number
- CN202310621328.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-30
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-05-30
AI Technical Summary
Traditional remote sensing image processing methods have problems such as low accuracy, slow processing speed, and difficulty in dealing with complex environments in aircraft target detection.
Improve the remote sensing image aircraft target detection model of YOLOv5, and improve the speed and accuracy of the model through the combination of feature extraction module, feature fusion module and multiple detection heads, including standard focus module, CSP optimized C3 module, ViT encoder, fast feature pyramid pooling layer and Bi-FPN architecture.
The detection speed and accuracy of the model are improved, and it can better cope with remote sensing image aircraft detection tasks in complex environments.
Smart Images

Figure CN116863145B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image target detection, and in particular to a remote sensing image aircraft target detection model and method for improving YOLOv5. Background Art
[0002] In recent years, with the booming development of high-resolution earth observation technology, the resolution of optical remote sensing images has been increasing year by year. The ground resolution of advanced remote sensing satellites can even reach the decimeter level. The rapid development of remote sensing technology has greatly enriched the acquisition channels of remote sensing images. Aircraft detection in remote sensing images, as an important issue in remote sensing technology, has extensive applications in the fields of military, civil aviation, security supervision, etc. However, due to factors such as complex backgrounds, large amounts of data, and noise interference in remote sensing images, traditional remote sensing image processing methods have problems such as low accuracy, slow processing speed, and difficulty in coping with complex environments in aircraft target detection. Therefore, aircraft detection in remote sensing images has always been a challenging issue. Summary of the Invention
[0003] The purpose of the present invention is to overcome the above technical deficiencies, and propose a remote sensing image aircraft target detection model and method for improving YOLOv5, so as to solve the technical problems such as low accuracy, slow processing speed, and difficulty in coping with complex environments existing in traditional remote sensing image processing methods in aircraft target detection.
[0004] To achieve the above technical objectives, in the first aspect, the technical solution of the present invention provides a remote sensing image aircraft target detection model for improving YOLOv5, including:
[0005] A feature extraction module, which is used to extract features from the input image to obtain a feature map, including:
[0006] A standard focus module, which is used to divide the input image into multiple sub-images, perform convolution operations on each of the sub-images, and connect the feature maps of all the sub-images to reduce the resolution of the feature map and improve the speed and accuracy of the model;
[0007] Multiple C3 modules optimized by CSP, which are used to integrate the change of gradients into the feature map to reduce the number of model parameters and FLOPs;
[0008] A ViT encoder, which uses a self-attention mechanism to perform global modeling on the input image to help the model capture global information and improve the model's understanding ability and representation ability for the input image;
[0009] A fast feature pyramid pooling layer, which is used to process the input image at different scales and perform feature fusion at different levels to improve the model's object recognition and localization ability;
[0010] The feature fusion module of the Bi-FPN architecture introduces a horizontal shortcut on the PAN and is communicatively connected to the feature extraction module. The feature fusion module is used to perform feature fusion processing on the feature maps of multiple levels output by the feature extraction module to obtain feature decodings of multiple scales;
[0011] Multiple detection heads, communicatively connected to the feature fusion module, are used to parse the feature decodings of different scales to obtain the edge coordinates of the parsed prediction boxes and output the detection results of multiple scales corresponding to the categories of the targets.
[0012] Compared with the prior art, the beneficial effects of the improved YOLOv5 remote sensing image aircraft target detection model provided by the present invention include:
[0013] The improved YOLOv5 remote sensing image aircraft target detection model provided by the present invention includes: a feature extraction module, a feature fusion module, and multiple detection heads. Among them, the feature extraction module is a neural network module that can aggregate and form image features at different image fine-grain levels, and can provide several combinations of receptive field sizes and center strides to meet the target detection at different categories and scales. The present invention introduces various strategies in the feature extraction network to improve the model accuracy. The improvement strategies of the feature extraction module include:
[0014] A standard focus module, multiple C3 modules optimized by CSP, a ViT encoder, and a fast feature pyramid pooling layer. The standard focus module divides the input image into multiple sub-images, performs convolution operations on each sub-image, and connects the feature maps of all sub-images. The standard focus module can reduce the resolution of the feature map by half without losing information, thereby reducing the computational amount. In actual detection tasks, it can greatly improve the speed and accuracy of the model;
[0015] Multiple C3 modules optimized by CSP. The C3 module used for feature extraction comes from the optimized CrossStagePartialNetworks, which integrates the change of gradients into the feature map, reducing the number of model parameters and FLOPs. Therefore, the model is smaller. The C3 module adopts the optimized CSPBottleneck, in which there are only 3 convolutions related to IO instead of the original 4 convolutions, which can make the model smaller without loss of accuracy, while ensuring the inference speed and accuracy of the model, and can let the model learn more features without changing the resolution size;
[0016] The ViT encoder uses the self-attention mechanism to globally model the input, which can help the model capture global information and improve the model's understanding and representation ability of the input. The ViT (Vision Transformer) encoder is a visual model that processes image data. Therefore, some special processing needs to be done on the image in the input part: VIT divides the input image into blocks and vectorizes it, so that the same encoding model as the word vector can be used.
[0017] The Spatial Pyramid Pooling Fast (SPPF) layer can improve the network's detection ability for objects of different scales. Compared with traditional network architectures, the SPPF module can process different scales of the input image and perform feature fusion at different levels, so as to obtain more context information and improve the network's object recognition and localization ability. It can process input images of different sizes, perform spatial pyramid pooling at different levels, and finally concatenate the feature maps of different levels together to form a feature representation with better context information. In the SPPF module, the size of the pooling operation can be adjusted according to the size of different input images, so that the network can process input images of different sizes and has better generalization performance.
[0018] Feature fusion module: The feature fusion network is a series of networks that mix and combine image features and is used to transfer image features to the detection head. The Bi-FPN structure is introduced to optimize the transfer of information flow and accelerate the convergence of the model. The feature maps from multiple levels of the feature extraction network are fused and processed to enhance their expression ability, and the processed feature maps of the same size are output for use by the detection head.
[0019] Multiple detection heads. For example, 4 detection heads can be set. In all current detection models, the detection head part shares parameters, that is, P3, P4, and P5 share one detection head. This solves the problem of uneven numbers of multi-scale samples in multi-level detection. By sharing the detection head, the model parameters can be reduced, the accuracy can be improved, and the P6 layer is added to form four scales of small, medium, large, and extra-large for object detection to improve the model performance. To balance accuracy and speed, the present invention adopts the method of setting 4 detection heads. Through multi-scale detection, larger objects can be detected, and the entire model can benefit from training at a higher resolution.
[0020] In a second aspect, the present invention provides a method for detecting aircraft targets in remote sensing images based on improved YOLOv5, including the following steps:
[0021] Perform data cleaning and hierarchical annotation processing on the aircraft image detection data set;
[0022] Input the aircraft image detection dataset into the trained remote sensing image aircraft target detection model based on the improved YOLOv5 as described above, and use the feature extraction module to extract features from the aircraft image detection dataset to obtain feature maps at multiple scales to meet the target detection requirements at multiple scales;
[0023] Use the feature fusion module with the Bi-FPN architecture to perform fusion of different-level semantic information transfer on the feature maps at multiple scales to obtain feature decodings at multiple scales;
[0024] Use the detection head to perform target detection at multiple scales, and perform parsing operations on the feature decodings at different scales to obtain the edge coordinates of the parsed prediction boxes, and output the detection results at multiple scales of the corresponding classes of the targets;
[0025] Calibrate the original image in the aircraft image detection dataset according to the detection results, select the target positions, and obtain the class confidence levels.
[0026] According to some embodiments of the present invention, perform data cleaning and hierarchical annotation processing on the aircraft image detection dataset, including the steps of:
[0027] Verify whether the pictures in the aircraft image detection dataset contain aircraft targets, and eliminate the chaotic pictures in the aircraft image detection dataset;
[0028] Train a binary classification aircraft data recognition model using a noise data processing algorithm, and use the trained binary classification aircraft data recognition model to recognize the picture data in the aircraft image detection dataset to clean the picture data that does not contain aircraft targets;
[0029] Control the proportion of noise picture data in the aircraft image detection dataset to be the first proportion;
[0030] Perform hierarchical annotation on each image in the aircraft image detection dataset, and the annotation information includes: class information, attribute information, bounding boxes, and object instances.
[0031] According to some embodiments of the present invention, training a binary classification aircraft data recognition model using a noise data processing algorithm includes the steps of:
[0032] Select the first quantity of noise picture data and the second quantity of aircraft picture data in the aircraft image detection dataset to train the binary classification aircraft data recognition model, and the first quantity is much smaller than the second quantity;
[0033] Apply the binary classification aircraft data recognition model to the remaining aircraft picture data after cleaning to predict the noise picture data with a prediction threshold greater than 0.85.
[0034] According to some embodiments of the present invention, the training process of the improved YOLOv5 remote sensing image aircraft target detection model includes the steps:
[0035] Obtain a training data set containing images and labels of target objects, and each of the labels should contain the category information and location information of the target;
[0036] Configure the parameters of the improved YOLOv5 remote sensing image aircraft target detection model according to the size and number of categories of the training data set;
[0037] Use the RIFMosaic data augmentation strategy to mosaically splice multiple images containing target objects in the spatial dimension in equal parts to complete the data augmentation process;
[0038] Set the training size and number of training rounds to train the improved YOLOv5 remote sensing image aircraft target detection model to adjust the resolution and anchor box position;
[0039] Use the validation data set to verify the performance of the trained model and optimize the model according to the verification results.
[0040] According to some embodiments of the present invention, the data augmentation process applies a data augmentation model, and the training process of the data augmentation model is as follows:
[0041] Data preparation: Randomly select multiple pictures from the data set for splicing, and use an adaptive method to scale each picture, fill the background with the average value of the data set pixels and retain the label box;
[0042] Model training: Perform data augmentation on each spliced picture, use the augmented pictures for model training, use Mosaic with a probability of 90%, RIFMosaic with a probability of 10%, and Mixup with a probability of 10% during training, and turn off RIFMosaic, Mosaic, and Mixup data augmentation in the last 10 rounds;
[0043] Model evaluation: Use the validation set or test set to evaluate the model and adjust the model according to the evaluation results.
[0044] According to some embodiments of the present invention, setting the training size and number of training rounds to train the improved YOLOv5 remote sensing image aircraft target detection model to adjust the resolution and anchor box position includes the steps:
[0045] The training size starts from [640, 960, 1280] and increases in increments of 100 rounds;
[0046] Adjust the resolution using the rectangular inference method, obtain the results of the training dataset using K-Means clustering, and then adjust the anchor box settings. Conduct experiments and analyze the receptive field of the model to determine the appropriate anchor box settings.
[0047] In a third aspect, the present invention provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the remote sensing image aircraft target detection method based on the improved YOLOv5 as described in any one of the second aspects.
[0048] The additional aspects and advantages of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The above and / or additional aspects and advantages of the present invention will become obvious and easy to understand from the description of the embodiments in conjunction with the following drawings, in which the abstract drawing should be exactly the same as one of the specification drawings:
[0050] Figure 1 Schematic diagram of the remote sensing image aircraft detection model based on the improved YOLOv5 provided by an embodiment of the present invention;
[0051] Figure 2 Flowchart of the remote sensing image aircraft detection method based on the improved YOLOv5 provided by an embodiment of the present invention;
[0052] Figure 3 Flowchart of the remote sensing image aircraft detection method based on the improved YOLOv5 provided by an embodiment of the present invention;
[0053] Figure 4 SPPF module diagram of the remote sensing image aircraft detection model based on the improved YOLOv5 provided by an embodiment of the present invention;
[0054] Figure 5 C3 module diagram after CSP optimization of the remote sensing image aircraft detection model based on the improved YOLOv5 provided by an embodiment of the present invention;
[0055] Figure 6 C3TR encoding module diagram of the remote sensing image aircraft detection model based on the improved YOLOv5 provided by an embodiment of the present invention;
[0056] Figure 7 RepVGG module diagram of the remote sensing image aircraft detection model based on the improved YOLOv5 provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] In order to make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0058] It should be noted that although the functional modules are divided in the system schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the system or the sequence in the flowchart. The terms "first", "second", etc. in the specification, claims, and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0059] The present invention provides a remote sensing image aircraft detection model that improves YOLOv5, which can improve the accuracy and processing speed and can handle the remote sensing image aircraft detection tasks in complex environments.
[0060] The embodiments of the present invention will be further described below with reference to the accompanying drawings.
[0061] Refer to Figure 1 、 Figures 4 - 7 , Figure 1 which is a schematic diagram of the remote sensing image aircraft detection model that improves YOLOv5 provided by an embodiment of the present invention; Figure 4 which is a SPPF module diagram of the remote sensing image aircraft detection model that improves YOLOv5 provided by an embodiment of the present invention; Figure 5 which is a C3 module diagram after CSP optimization of the remote sensing image aircraft detection model that improves YOLOv5 provided by an embodiment of the present invention; Figure 6 which is a C3TR encoding module diagram of the remote sensing image aircraft detection model that improves YOLOv5 provided by an embodiment of the present invention; Figure 7 which is a RepVGG module diagram of the remote sensing image aircraft detection model that improves YOLOv5 provided by an embodiment of the present invention.
[0062] In one embodiment, the improved remote sensing image aircraft target detection model based on YOLOv5 includes: a feature extraction module for extracting features from the input image to obtain a feature map, including: a standard focus module for dividing the input image into multiple sub-images, performing convolution operations on each sub-image, and connecting the feature maps of all sub-images to reduce the resolution of the feature map and improve the speed and accuracy of the model; multiple C3 modules optimized by CSP for integrating the change of gradients into the feature map to reduce the number of model parameters and FLOPs; a ViT encoder for globally modeling the input image using the self-attention mechanism to help the model capture global information and improve the model's understanding and representation capabilities of the input image; a fast feature pyramid pooling layer for processing the input image at different scales and performing feature fusion at different levels to improve the model's object recognition and localization capabilities; a feature fusion module with a Bi-FPN architecture that introduces a horizontal shortcut on the PAN and is communicatively connected to the feature extraction module. The feature fusion module is used to perform feature fusion processing on the feature maps of multiple levels output by the feature extraction module to obtain feature decoding at multiple scales; multiple detection heads communicatively connected to the feature fusion module, and the detection heads are used to parse the feature decoding at different scales to obtain the edge coordinates of the parsed prediction boxes and output the detection results at multiple scales corresponding to the target categories.
[0063] PAN, namely PANet, adds a feature pyramid of downsampling fusion after the feature pyramid of upsampling fusion on FPN.
[0064] Among them, the improvement strategies of the feature extraction module include: a standard focus module, multiple C3 modules optimized by CSP, a ViT encoder, and a fast feature pyramid pooling layer. The standard focus module divides the input image into multiple sub-images, performs convolution operations on each sub-image, and connects the feature maps of all sub-images. The standard focus module can reduce the resolution of the feature map by half without losing information, thereby reducing the computational amount. In actual detection tasks, it can greatly improve the speed and accuracy of the model.
[0065] Multiple C3 modules optimized by CSP. The C3 module used for feature extraction comes from the optimized CrossStagePartialNetworks, which integrates the change of gradients into the feature map, reducing the number of model parameters and FLOPs, so the model is smaller. As Figure 5 shown, the C3 module adopts the optimized CSP Bottleneck, in which there are only 3 convolutions related to IO instead of the original 4 convolutions, which can make the model smaller without loss of accuracy, while ensuring the inference speed and accuracy of the model, and allowing the model to learn more features without changing the resolution size.
[0066] The ViT encoder uses the self-attention mechanism to globally model the input, which can help the model capture global information and improve the model's understanding and representation ability of the input. The ViT (Vision Transformer) encoder is a vision model that processes image data. Therefore, some special processing needs to be done on the image in the input part: VIT divides the input image into blocks and vectorizes it, so that the same encoding model as the word vector can be used.
[0067] The Spatial Pyramid Pooling Fast (SPPF) layer can improve the network's detection ability for objects of different scales. Compared with traditional network architectures, the SPPF module can process different scales of the input image and perform feature fusion at different levels, so as to obtain more context information and improve the network's object recognition and localization ability. It can process input images of different sizes, perform spatial pyramid pooling at different levels, and finally concatenate the feature maps of different levels to form a feature representation with better context information. In the SPPF module, the size of the pooling operation can be adjusted according to the size of different input images, so that the network can process input images of different sizes and has better generalization performance.
[0068] Feature fusion module: The feature fusion module is a series of networks that mix and combine image features and is used to transfer image features to the detection head. The Bi-FPN structure is introduced to optimize the transmission of information flow and accelerate the convergence of the model. The feature maps from multiple levels of the feature extraction module are fused and processed to enhance their expressive ability, and the processed feature maps with the same size are output for the detection head to use.
[0069] The Bi-FPN architecture proposed by EfficientDet is introduced and upgraded. Horizontal shortcuts are introduced on the PAN, but it is not pruned like EfficientDet. Compared with the PAN, the Bi-FPN adds horizontal connections, that is, the features of the feature extraction network are introduced into the subsequent same level for better feature fusion, and the nodes with only one input edge are removed to enhance the feature representation ability. Based on this idea, this component is introduced in the P4 and P5 levels of the improved YOLOv5 remote sensing image aircraft target detection model, and the [6,20,24][8,16,27] layers are spliced, and the detection weights of P3, P4, P5, and P6 are also adjusted to regulate the contribution of each scale detection.
[0070] Multiple detection heads, for example, 4 detection heads can be set. In all current detection models, the detection head part shares parameters, that is, P3, P4, and P5 share one detection head. This solves the problem of uneven numbers of multi-scale samples in multi-level detection. By sharing the detection head, the model parameters can be reduced, the accuracy can be improved, and the P6 layer is added to form four scales of small, medium, large, and extra-large for object detection to improve the model performance. To balance accuracy and speed, the present invention adopts the method of setting 4 detection heads. Through multi-scale detection, larger objects can be detected, and the entire model can benefit from training at a higher resolution.
[0071] RepVGG module reparameterization: The RepVGG module is mainly for accelerating the model while maintaining the original model accuracy. The specific approach is: adding parallel 1×1 convolution branches and identity mapping branches to each 3×3 convolution, thus constituting the basic RepVGG module. The RepVGG module adds branches at each layer.
[0072] Such as Figure 1 As shown, the feature extraction module can successively include: a convolutional layer, a RepVGG module, a C3 module, a RepVGG module, a C3 module, a RepVGG module, a C3 module, a RepVGG module, a C3TR module, a RepVGG module, a C3 module, and a fast feature pyramid pooling layer. The structure of the C3TR module is as shown in the appendix Figure 6 As shown, the C3TR module replaces the original C3 module for feature extraction, which can improve the feature expression ability of the model. Through experiments, it is found that placing the C3TR in the 9th layer of the feature extraction network can significantly improve the model performance because the feature maps after the 9th layer are more abstract and high-level, and are more suitable for global modeling. The standard focus module is set in the convolutional layer.
[0073] The feature fusion module includes: a convolutional layer, an Upsample layer, a C3 module, a RepVGG module, and a connection layer. The Bi-FPN structure is introduced to optimize the transmission of information flow and accelerate the model convergence. The feature maps from multiple levels of the feature extraction module are fused and processed to enhance their expression ability, and the processed feature maps with the same size are output for use by the detection head.
[0074] The improved YOLOv5 remote sensing image aircraft target detection model proposed by the present invention integrates effective deep learning technologies, can better utilize the characteristics of convolutional neural networks, introduces the ViT encoder module to model global information, can better utilize the global features in the image, and improves the detection accuracy. The improved YOLOv5 remote sensing image aircraft target detection model has excellent engineering performance and is flexible in aspects such as training, deployment, and optimization.
[0075] Refer to Figure 2and Figure 3 , Figure 2 is a flowchart of a method for detecting airplanes in remote sensing images by improving YOLOv5 provided by an embodiment of the present invention; Figure 3 is a flowchart of a method for detecting airplanes in remote sensing images by improving YOLOv5 provided by an embodiment of the present invention.
[0076] The method for detecting airplanes in remote sensing images by improving YOLOv5 includes but is not limited to steps S110 to S150.
[0077] Step S110, perform data cleaning and hierarchical annotation processing on the airplane image detection dataset;
[0078] Step S120, input the airplane image detection dataset into the trained remote sensing image airplane target detection model that improves YOLOv5, and use the feature extraction module to extract features from the airplane image detection dataset to obtain feature maps of multiple scales to meet the target detection requirements of multiple scales;
[0079] Step S130, use the feature fusion module with the Bi-FPN architecture to perform fusion of different-level semantic information transfer on the feature maps of multiple scales to obtain feature decodings of multiple scales;
[0080] Step S140, use the detection head to perform target detection at multiple scales, and perform parsing operations on the feature decodings of different scales to obtain the edge coordinates of the parsed prediction boxes, and output the detection results of multiple scales corresponding to the classes of the targets;
[0081] Step S150, perform calibration on the original image in the airplane image detection dataset according to the detection results, select the target positions, and obtain the class confidence levels.
[0082] In one embodiment, the method for detecting airplanes in remote sensing images by improving YOLOv5 includes the steps of: performing data cleaning and hierarchical annotation processing on the airplane image detection dataset; inputting the airplane image detection dataset into the trained remote sensing image airplane target detection model that improves YOLOv5 as described above, and using the feature extraction module to extract features from the airplane image detection dataset to obtain feature maps of multiple scales to meet the target detection requirements of multiple scales; using the feature fusion module with the Bi-FPN architecture to perform fusion of different-level semantic information transfer on the feature maps of multiple scales to obtain feature decodings of multiple scales; using the detection head to perform target detection at multiple scales, and perform parsing operations on the feature decodings of different scales to obtain the edge coordinates of the parsed prediction boxes, and output the detection results of multiple scales corresponding to the classes of the targets; performing calibration on the original image in the airplane image detection dataset according to the detection results, select the target positions, and obtain the class confidence levels.
[0083] Input the aircraft image detection dataset, and set the scale factor to meet the input requirements of the Yolox-Cr algorithm; input the aircraft image detection dataset into the trained remote sensing image aircraft target detection model with improved YOLOv5 for detection. Use CSPDarknet as the feature extraction network to extract rich features from the input image. After passing through the feature extraction block network, feature maps of four scales are output to meet the detection of small, medium, large, and extra-large targets; introduce the Bi-FPN idea for feature fusion. The four different-scale feature maps obtained after passing through the feature extraction module are fused through the feature pyramid fusion network for different-level semantic information transfer and fusion to strengthen the flow of information, and finally four different-scale feature decodings are obtained; the detection head performs target detection at four scales of small, medium, large, and extra-large, and decodes the feature maps of different scales to obtain the coordinates of the upper left and lower right corners of the parsed prediction box and the detection results of four different scales of the categories corresponding to the targets; based on the results of the detection head, the target position and category confidence are calibrated in the original image.
[0084] The attention mechanism introduced in the improved YOLOv5 remote sensing image aircraft detection method provided by the present invention can better extract the feature information of small object instances, thereby improving the detection accuracy and generalization ability; the proposed data augmentation strategy further increases the diversity and richness of the data and improves the robustness of the model. The method provided by the present invention can more effectively solve the problem of remote sensing aircraft target detection in complex environments, and has important practical application value and promotion significance.
[0085] In one embodiment, data cleaning and hierarchical annotation processing are performed on the aircraft image detection dataset, including the steps of: verifying whether the pictures in the aircraft image detection dataset contain aircraft targets, and removing the chaotic pictures in the aircraft image detection dataset; training a binary classification aircraft data recognition model using a noise data processing algorithm, and using the trained binary classification aircraft data recognition model to identify the picture data in the aircraft image detection dataset to clean the picture data that does not contain aircraft targets; controlling the proportion of noise picture data in the aircraft image detection dataset to be the first proportion; performing hierarchical annotation on each image in the aircraft image detection dataset, and the annotation information includes: category information, attribute information, bounding box, and object instance.
[0086] Among them, training a binary classification aircraft data recognition model using a noise data processing algorithm includes the steps of: selecting the first number of noise picture data and the second number of aircraft picture data in the aircraft image detection dataset to train the binary classification aircraft data recognition model, and the first number is much smaller than the second number; applying the binary classification aircraft data recognition model to the remaining aircraft picture data after cleaning to predict the noise picture data with a prediction threshold greater than 0.85.
[0087] The aircraft image detection dataset CrAP constructed by the present invention includes civil, military, and general aircraft. The aircraft image detection dataset CrAP provides a way of annotating horizontal bounding boxes, provides the Txt and Xml annotation formats used for object detection, includes images with interference situations such as too dark or too bright lighting, low contrast, and shadow occlusion, and the background environments in the images are also different, with rich scenes. The fully annotated CrAP dataset contains 20,000 pictures (19,800 aircraft images and 200 background images), includes 20 aircraft models, and more than 40,000 aircraft instances.
[0088] For remote sensing aircraft images with diverse categories and different resolutions, data cleaning and screening are required to construct an object detection dataset to ensure the quality and effectiveness of the dataset. The following is the construction process of the dataset:
[0089] Step 1: Manually review the original data, check whether the data contains aircraft targets, and eliminate the chaotic pictures among them;
[0090] Step 2: Train a binary classification model using a noise data processing algorithm:
[0091] (1) Train a binary classification aircraft data recognition model using about 100 noise pictures and about 5,000 aircraft pictures;
[0092] (2) Apply the binary classification model to the aircraft data after cleaning, and predict the noise data with a prediction threshold greater than 0.85;
[0093] (3) Iterate the above steps 2 times to complete the data cleaning that does not contain aircraft targets.
[0094] Step 3: Conduct a manual verification again. The proportion of the cleaned and artificially added noise data in the dataset finally accounts for 1%;
[0095] Step 4: Conduct hierarchical annotation: category, attribute, bounding box, and object instance. Fine-annotate each image. Use LabelImg in the Anaconda environment of Windows for data annotation. It can annotate and save the data commonly used in object detection in the PASCAL VOC format as an Xml file and the Txt format of YOLO. Create 20 categories according to the requirements, and then classify the collected and processed data into these categories. If non-aircraft images appear during annotation, these candidate images will be used as noise images.
[0096] The aircraft target detection dataset for remote sensing images established by the present invention has obtained a total of 20,000 sliced images containing aircraft targets, including aircraft objects with different proportions, orientations, and shapes, thus improving the quality of the dataset.
[0097] In one embodiment, the training process of the improved YOLOv5 remote sensing image aircraft target detection model includes the steps of: obtaining a training data set containing images and labels of target objects, where each label should include the category information and location information of the target; configuring the parameters of the improved YOLOv5 remote sensing image aircraft target detection model according to the size and number of categories of the training data set; using the RIFMosaic data augmentation strategy to splice multiple images containing target objects equally in the spatial dimension to complete the data augmentation process; setting the training size and number of training epochs to train the improved YOLOv5 remote sensing image aircraft target detection model to adjust the resolution and anchor box position; using the validation data set to verify the performance of the trained model and optimizing the model according to the verification results.
[0098] The training process of the improved YOLOv5 remote sensing image aircraft target detection model is as follows:
[0099] Data set preparation: The model is trained in the self-built remote sensing aircraft target detection data set CrAP, and a training data set containing images and labels of target objects is prepared. Each label should include the category and location information of the target;
[0100] Configure parameters: Configure the parameters of YOLOv5 according to the size and number of categories of the data set, such as batchsize and learning rate during training. In order to obtain better initial hyperparameters, we use a genetic algorithm for hyperparameter search, select the standard SGD optimizer for optimization, and use the gradient accumulation technique to set the NominalBatchSize to 64, while the batch size is set to the largest divisor divisible by 64 according to the specific model;
[0101] Data augmentation: The present invention introduces the RIFMosaic data augmentation strategy, that is, directly splicing four pictures equally in the spatial dimension, and each picture is scaled in an adaptive manner, and then randomly photometric jittering is used in the same way as Mosaic. Compared with Mosaic, it can save video memory, reduce information redundancy and improve accuracy.
[0102] Train the model: The training size starts from [640, 960, 1280] and increases by 100 epochs. When specifically adjusting the resolution, the rectangular inference method is used. Use K-Means clustering to obtain the results of the CrAP data set and then adjust the anchor box settings, conduct experiments and analyze the model receptive field to determine the appropriate anchor box settings. All experiments except the YOLOv5-Cr model are trained for 100 epochs, and the training time is shortened or extended according to the actual situation;
[0103] Verify the model: Use the validation data set to verify the performance of the trained model;
[0104] Model Optimization: Optimize the model according to the verification results, and feedback the loss after training to the model to adjust the prediction parameters, minimize the value of the loss function, and make the prediction results more accurate.
[0105] During the data augmentation process, a data augmentation model is applied. The training process of the data augmentation model is as follows:
[0106] Data Preparation: Randomly select multiple images from the dataset for stitching, and scale each image in an adaptive manner. The background is filled with the average value of the dataset pixels and the label boxes are retained;
[0107] Model Training: Perform data augmentation on each stitched image, use the augmented images for model training, and use Mosaic with a probability of 90%, RIFMosaic with a probability of 10%, and Mixup with a probability of 10% during training. Turn off RIFMosaic, Mosaic, and Mixup data augmentation in the last 10 rounds;
[0108] Model Evaluation: Evaluate the model using the validation set or the test set, and adjust the model according to the evaluation results.
[0109] The remote sensing image aircraft detection method based on the improved YOLOv5 proposed by the present invention can be applied to various complex scenarios. By constructing a dataset, improving the data augmentation strategy, optimizing the model structure and parameters, etc., it can effectively improve the accuracy and robustness of aircraft target detection, and has significant economic and social benefits, and has a certain promoting effect on other remote sensing image automatic interpretation systems.
[0110] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0111] In addition, an embodiment of the present invention also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor or a controller, for example, executed by a processor in the above terminal embodiment, the processor can execute the remote sensing image aircraft detection method for improving YOLOv5 in the above embodiment.
[0112] Those of ordinary skill in the art will understand that all or some of the steps and systems disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.
[0113] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.
[0114] The specific embodiments of the present invention described above do not constitute a limitation on the protection scope of the present invention. Any other corresponding changes and deformations made according to the technical concept of the present invention should be included within the protection scope of the claims of the present invention.
Claims
1. An improved remote sensing image aircraft target detection model for YOLOv5, characterized in that, Including: A feature extraction module, configured to extract features from an input image to obtain a feature map, including: A standard focusing module, configured to divide the input image into multiple sub-images, perform convolution operations on each of the sub-images, and concatenate the feature maps of all the sub-images to reduce the resolution of the feature map and improve the speed and accuracy of the model; Multiple C3 modules optimized by CSP, configured to integrate the change of gradients into the feature map to reduce the number of model parameters and FLOPs; A ViT encoder, using a self-attention mechanism to perform global modeling on the input image to help the model capture global information and improve the model's understanding and representation capabilities for the input image; A fast feature pyramid pooling layer, configured to process the input image at different scales and perform feature fusion at different levels to improve the model's object recognition and localization capabilities; A feature fusion module with a Bi-FPN architecture, introducing a lateral shortcut on the PAN and communicatively connected to the feature extraction module, the feature fusion module being configured to perform feature fusion processing on the feature maps of multiple levels output by the feature extraction module to obtain feature decodings of multiple scales; Multiple detection heads, communicatively connected to the feature fusion module, the detection heads being configured to parse the feature decodings of different scales to obtain the edge coordinates of the parsed prediction boxes and output the detection results of multiple scales corresponding to the categories of the targets.
2. A method for detecting aircraft targets in remote sensing images based on improved YOLOv5, characterized in that, Including the following steps: Performing data cleaning and hierarchical annotation processing on an aircraft image detection data set; Inputting the aircraft image detection data set into the trained remote sensing image aircraft target detection model of the improved YOLOv5 as claimed in claim 1, and using the feature extraction module to extract features from the aircraft image detection data set to obtain feature maps of multiple scales to meet the target detection requirements of multiple scales; Using the feature fusion module with a Bi-FPN architecture to perform fusion of different-level semantic information transmission on the feature maps of multiple scales to obtain feature decodings of multiple scales; Using the detection heads to perform target detection at multiple scales and performing parsing operations on the feature decodings of different scales to obtain the edge coordinates of the parsed prediction boxes and output the detection results of multiple scales corresponding to the categories of the targets; Calibrating the original image in the aircraft image detection data set according to the detection results, selecting the target positions, and obtaining the category confidence levels.
3. A remote sensing image aircraft target detection method based on improved YOLOv5 according to claim 2, characterized in that, Performing data cleaning and hierarchical annotation processing on an aircraft image detection data set, including the steps of: Verifying whether the pictures in the aircraft image detection data set contain aircraft targets, and removing the chaotic pictures in the aircraft image detection data set; Training a binary classification aircraft data recognition model using a noise data processing algorithm, and using the trained binary classification aircraft data recognition model to recognize the picture data in the aircraft image detection data set to clean the picture data that does not contain aircraft targets; Controlling the proportion of noise picture data in the aircraft image detection data set to be a first proportion; Performing hierarchical annotation on each image in the aircraft image detection data set, and the annotation information includes: category information, attribute information, bounding boxes, and object instances.
4. The aircraft target detection method for remote sensing images based on improved YOLOv5 according to claim 3, characterized in that, Train a binary classification aircraft data recognition model using a noise data processing algorithm, including the steps of: Select a first number of noise picture data and a second number of aircraft picture data from the aircraft image detection dataset to train the binary classification aircraft data recognition model, where the first number is much smaller than the second number; Apply the binary classification aircraft data recognition model to the remaining aircraft picture data for cleaning, and predict the noise picture data with a threshold greater than 0.
85.
5. The method for detecting aircraft targets in remote sensing images based on improved YOLOv5 according to claim 2, wherein, The training process of the improved YOLOv5 remote sensing image aircraft target detection model includes the steps of: Obtain a training dataset containing images and labels of target objects, where the labels contain category information and location information of the targets; Configure the parameters of the improved YOLOv5 remote sensing image aircraft target detection model according to the size and number of categories of the training dataset; Use the RIFMosaic data augmentation strategy to splice multiple images containing target objects equally in the spatial dimension to complete the data augmentation process; Set the training size and number of training rounds to train the improved YOLOv5 remote sensing image aircraft target detection model to adjust the resolution and anchor box positions; Use the validation dataset to verify the performance of the trained model and optimize the model according to the verification results.
6. The aircraft target detection method for remote sensing images based on improved YOLOv5 according to claim 5, wherein the data augmentation process applies a data augmentation model, characterized in that The training process of the data augmentation model is as follows: Data preparation: Randomly select multiple pictures from the dataset for splicing, and use an adaptive method to scale each picture, fill the background with the average value of the dataset pixels and retain the label boxes; Model training: Perform data augmentation on each spliced picture, use the augmented pictures for model training, use Mosaic with a 90% probability, RIFMosaic with a 10% probability, and Mixup with a 10% probability during training, and turn off RIFMosaic, Mosaic, and Mixup data augmentation in the last 10 rounds; Model evaluation: Use the validation set or test set to evaluate the model and adjust the model according to the evaluation results.
7. According to the method for detecting aircraft targets in remote sensing images based on improved YOLOv5 described in claim 5, set the training size and number of training rounds to train the improved YOLOv5 remote sensing image aircraft target detection model to adjust the resolution and anchor box positions, including the steps of: The training size starts from [640, 960, 1280] and increases in increments of 100 rounds; Adjust the resolution using the rectangular inference method, use K-Means clustering to obtain the results of the training dataset and then adjust the anchor box settings, conduct experiments and analyze the model receptive field to determine the appropriate anchor box settings.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to execute the method for detecting aircraft targets in remote sensing images based on improved YOLOv5 according to any one of claims 2 to 7.
Citation Information
Patent Citations
Detection method for identifying small target
CN114049572A
Method for rapidly and intelligently detecting house damage caused by remote sensing image of unmanned aerial vehicle after earthquake
CN114359756A