Remote sensing airport identification method

By improving the YOLOv8 model, combined with the CSPASPP module and the CBAM attention mechanism, the problem of low recognition accuracy in remote sensing airport recognition is solved, and higher recognition accuracy and detection accuracy are achieved, especially in complex backgrounds.

CN119919790APending Publication Date: 2025-05-02Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311557040.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-10-23
Filing Date
2023-11-21
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The prior art has the problem of low recognition accuracy in remote sensing airport recognition, especially in the lack of standardized large data volume, the model learning ability is insufficient, resulting in poor recognition effect.

Method used

The YOLOv8 object detection model is improved, by adding a CSPASPP module to the shallow feature output part of the feature extraction network, using the CBAM attention mechanism module to focus feature at the feature extraction and detection head input, and adding a convolution module to the feature extraction network to enhance learning ability.

Benefits of technology

By improving the YOLOv8 model, the context features are fully utilized, and the accuracy and detection accuracy of airport recognition are improved, especially on remote sensing images with complex backgrounds, insufficient images and small targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919790A_ABST
    Figure CN119919790A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of remote sensing target detection, and particularly relates to a remote sensing airport recognition method. Acquiring a remote sensing image and inputting the remote sensing image into the trained improved YOLOv8 target detection model to obtain an airport recognition result in the remote sensing image; wherein the improved YOLOv8 target detection model comprises a YOLOv8 model and two CSPASPP modules, and the two CSPASPP modules are respectively arranged on two shallow layer feature outputs of a feature extraction network included in the YOLOv8 model; the CSPASPP module is obtained by improving a traditional ASPP module, local multi-scale feature fusion can be achieved, and extraction and fusion of multi-range context features can be better met. According to the method, accurate identification of the remote sensing airport is realized, a good effect can be obtained on a relatively simple sample, and the method has more advantages in detection of a target on a relatively complex image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of remote sensing target detection, and in particular relates to a remote sensing airport recognition method. Background Art

[0002] With the widespread application of remote sensing technology, the demand for various intelligent target recognition on remote sensing images is also growing. The main purpose of remote sensing target detection is to detect and locate the target objects of interest on remote sensing images. There are many valuable targets worth detecting on remote sensing images, including ships, various types of small targets, etc. Airports are one of the important transportation hubs, which have great civil and military value. Airports usually contain a series of valuable target objects. Many researchers have developed detection algorithms for these targets, including YOLOv3, CGAN and other target detection algorithms.

[0003] However, to complete the detection of valuable targets inside the airport, it is necessary to first obtain high-resolution and high-definition remote sensing images of the airport. Therefore, it is of great significance to detect the entire airport as a target and determine the exact location of the airport. For airport target recognition on remote sensing images, traditional methods mainly use significant features, edge detection and other methods to detect airports. For example, the authors Chen Xuguang and Lin Hui published the "Identification Method of Airport Targets in Remote Sensing Images" in Volume 48 of "Computer Engineering and Applications" in 2012. The Canny operator is used for edge detection, and the Hough transform is used to extract straight line segments, and then the airport is located. Later, airport target recognition methods based on machine learning algorithms appeared. For example, the authors Zhu Dan, Wang Bin, and Zhang Liming published the "Remote Sensing Image Airport Target Detection Based on Line Neighboring Parallelism and GBVS Significance" in the 3rd issue of "Journal of Infrared and Millimeter Waves" in 2015. The SVM support vector machine realizes the target detection of the airport. The above algorithms often require human observation, discovery, and summary of some notable features of the airport. For example, the long runway of an airport is a striking long straight line segment, which makes feature extraction and algorithm design difficult.

[0004] In recent years, deep learning algorithms with deep convolutional neural networks as the core have been widely used in the field of computer vision. Important tasks in the field of computer vision include image recognition, image segmentation, and target recognition. For airport target detection, for example, the authors Zhang Yimin, Han Xianwei, and Zhang Shichao published "Rapid Detection of Airport Targets Based on Visual Saliency and Convolutional Neural Networks" in Volume 42 of "Space Return and Remote Sensing" in 2021, which realized airport target detection based on the Faster RCNN model architecture. However, remote sensing airport target recognition still lacks standardized, large-scale remote sensing airport data sets, which places high demands on the learning ability of the model. The model algorithm is not advanced enough. Using only these small remote sensing airport data sets makes these less advanced models unable to achieve effective and accurate learning effects, and there are problems of low accuracy and poor recognition effect. Summary of the invention

[0005] The purpose of the present invention is to provide a remote sensing airport identification method to solve the problem of low airport identification accuracy caused by the existing technical methods.

[0006] To solve the above technical problems, the present invention provides a remote sensing airport recognition method, which obtains a remote sensing image and inputs it into a trained improved YOLOv8 target detection model to obtain an airport recognition result in the remote sensing image; wherein the improved YOLOv8 target detection model includes a YOLOv8 model and two CSPASPP modules, and the two CSPASPPs are respectively set on two shallow feature outputs of a feature extraction network included in the YOLOv8 model; a CSPASPP module includes multiple convolution modules, which are used to make the input pass through a convolution module for feature map compression, and the compressed feature map passes through multiple convolution modules with different void rates for context feature extraction, and the context feature extraction results are spliced ​​and passed through a convolution module for context feature fusion and compression, and then spliced ​​with the compressed feature map, and then passed through a convolution module after splicing for feature fusion and compression.

[0007] The beneficial effects of the above technical solution are as follows: the present invention improves the traditional YOLOv8 model, and adds a CSPASPP module to the shallow feature output part of its feature extraction network. The CSPASPP module is obtained by improving the traditional ASPP module, and can realize local multi-scale feature fusion to better meet the extraction and fusion of multi-range context features. The visual receptive field of the shallow network is relatively small, and the context features cannot be fully utilized. Therefore, after adding the CSPASPP module to the shallow feature output of the feature extraction network, the CSPASPP module can be fully utilized to extract and fuse context features of different ranges on the feature map, and the context features are more fully applied to achieve accurate identification of remote sensing airports. Good results can be achieved on relatively simple samples, and experiments have found that the model of the present invention has more advantages for remote sensing images with complex backgrounds, unclear images, and small targets.

[0008] Furthermore, the improved YOLOv8 target detection model also includes two CBAM attention mechanism modules respectively set on the input of the deepest Conv layer and the input of the second deepest Conv layer in the feature extraction network.

[0009] The beneficial effects of the above technical solution are as follows: adding the CBAM attention mechanism module in the deep layer of the feature extraction network can focus on the features of key areas and channels, reduce the risk of feature information loss caused by the downsampling operation of the Conv module with a step size of 2 in the YOLOv8 model, thereby improving the detection accuracy of the improved YOLOv8 target detection model.

[0010] Furthermore, the improved YOLOv8 target detection model also includes three CBAM attention mechanism modules respectively set on three inputs of the detection head included in the YOLOv8 model.

[0011] The beneficial effects of the above technical solution are: adding a CBAM attention mechanism module to the input of the detection head, focusing on the features of important areas and channels again, and further improving the detection accuracy of the improved YOLOv8 target detection model.

[0012] Furthermore, the improved YOLOv8 target detection model also includes three Conv modules respectively arranged on the input of the second shallow C2f module, the input of the second deep C2f module and the input of the deepest C2f module in the feature extraction network.

[0013] The beneficial effects of the above technical solution are as follows: the improvement is designed based on the airport target features, and adding a Conv module to the input of the second shallow C2f module can not only condense and filter the shallow information and increase the visual receptive field of the model, but also reduce the computing power consumption of subsequent calculations; moreover, after the above improvement, the information of the deep feature map will become more macroscopic and complex, and a Conv module is added to the input of the second deep C2f module and the input of the deepest C2f module to enhance the learning ability of the improved YOLOv8 target detection model.

[0014] Furthermore, the step size of the Conv module set on the input of the next shallow C2f module is 2, and the step sizes of the Conv modules set on the input of the next deep C2f module and the input of the deepest C2f module are both 1.

[0015] Furthermore, the convolution module is a CBS module. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 It is a structural diagram of ASPP in the prior art;

[0017] Figure 2 It is a structural diagram of CBAM in the prior art;

[0018] Figure 3 It is a structural diagram of the CSPASPP module of the present invention;

[0019] Figure 4 It is a structural diagram of the improved YOLOv8 target detection model of the present invention;

[0020] Figure 5(1a), Figure 5(1b), and Figure 5(1c) are three different airport remote sensing images against an urban background;

[0021] Figure 5(2a), Figure 5(2b), and Figure 5(2c) are three different remote sensing images of airports against a seaside background;

[0022] Figure 5(3a), Figure 5(3b), and Figure 5(3c) are three different airport remote sensing images against the background of an island;

[0023] Figure 5(4a), Figure 5(4b), and Figure 5(4c) are three different remote sensing images of airports against the background of artificial islands;

[0024] Figure 5 (5a), Figure 5 (5b), and Figure 5 (5c) are three different remote sensing images of airports against a desert background;

[0025] Figure 5 (6a), Figure 5 (6b), and Figure 5 (6c) are three different remote sensing images of airports against a hilly background;

[0026] Figure 5 (7a), Figure 5 (7b), and Figure 5 (7c) are three different airport remote sensing images with farmland backgrounds;

[0027] Figure 6 (1a), Figure 6 (1b), and Figure 6 (1c) are three different airport remote sensing images with simple complexity;

[0028] Figure 6 (2a), Figure 6 (2b), and Figure 6 (2c) are three different airport remote sensing images with medium complexity;

[0029] Figure 6 (3a), Figure 6 (3b), and Figure 6 (3c) are three different airport remote sensing images with complex complexity;

[0030] Figure 7(1a), Figure 7(1b), and Figure 7(1c) are three high-resolution remote sensing images of different airports;

[0031] Figure 7(2a), Figure 7(2b), and Figure 7(2c) are three different airport remote sensing images with medium resolution;

[0032] Figure 7 (3a), Figure 7 (3b), and Figure 7 (3c) are three remote sensing images of different airports with low resolution;

[0033] Figure 8(a), Figure 8(b), Figure 8(c), and Figure 8(d) are the curve diagrams of the changes of the evaluation indicators Precision, Recall, mAP50, and mAP 50-95 of the model during the training process, respectively;

[0034] Figure 9 (1a) is the original image of the remote sensing image 1 which is clear and has a simple airport structure;

[0035] FIG9(1b), FIG9(1c), and FIG9(1d) are respectively the detection results of the airport detection of FIG9(1a) using YOLOv8s, the model of the present invention, and the HR model of the present invention;

[0036] Figure 9 (2a) is the original image of the remote sensing image 2 which is clear and has a simple airport structure;

[0037] FIG9(2b), FIG9(2c), and FIG9(2d) are respectively the detection results of the airport detection of FIG9(2a) using YOLOv8s, the model of the present invention, and the HR model of the present invention;

[0038] Figure 9 (3a) is the original image of the remote sensing image 3 which is clear and has a simple airport structure;

[0039] FIG9(3b), FIG9(3c), and FIG9(3d) are respectively the detection results of the airport detection of FIG9(3a) using YOLOv8s, the model of the present invention, and the HR model of the present invention;

[0040] Figure 10 (1a) is the original image of some remote sensing images 1 that are characteristic and difficult to obtain;

[0041] FIG10(1b), FIG10(1c), and FIG10(1d) are respectively the detection results of the airport detection of FIG10(1a) using YOLOv8s, the model of the present invention, and the HR model of the present invention;

[0042] Figure 10 (2a) is the original image of some remote sensing images 2 that are characteristic and difficult to obtain;

[0043] FIG10(2b), FIG10(2c), and FIG10(2d) are respectively the detection results of the airport detection of FIG10(2a) using YOLOv8s, the model of the present invention, and the HR model of the present invention;

[0044] Figure 10 (3a) is the original image of some remote sensing images 3 that are characteristic and difficult to obtain;

[0045] FIG10(3b), FIG10(3c), and FIG10(3d) are respectively the detection results of the airport detection of FIG10(3a) using YOLOv8s, the model of the present invention, and the HR model of the present invention;

[0046] Figure 10 (4a) is the original image of some remote sensing images 4 that are characteristic and difficult to obtain;

[0047] FIG10(4b), FIG10(4c), and FIG10(4d) are respectively the detection results of the airport detection of FIG10(4a) using YOLOv8s, the model of the present invention, and the HR model of the present invention;

[0048] Figure 10 (5a) is the original image of some remote sensing images 5 that are characteristic and difficult to obtain;

[0049] FIG10(5b), FIG10(5c), and FIG10(5d) are respectively the detection results of the airport detection of FIG10(5a) using YOLOv8s, the model of the present invention, and the HR model of the present invention;

[0050] Figure 10 (6a) is the original image of some remote sensing images 6 that are characteristic and difficult to obtain;

[0051] FIG10(6b), FIG10(6c), and FIG10(6d) are respectively the detection results of the airport detection of FIG10(6a) using YOLOv8s, the model of the present invention, and the HR model of the present invention;

[0052] Figure 11 (1a) is the original image of remote sensing image 1;

[0053] FIG11(1b), FIG11(1c), and FIG11(1d) are respectively detection results of the attention heat map of FIG11(1a) using YOLOv8s, the model of the present invention, and the HR model of the present invention for airport detection;

[0054] Figure 11 (2a) is the original image of remote sensing image 2;

[0055] FIG11(2b), FIG11(2c), and FIG11(2d) are respectively the detection results of the airport detection using YOLOv8s, the model of the present invention, and the HR model of the present invention on the attention heat map of FIG11(2a);

[0056] Figure 11 (3a) is the original image of remote sensing image 3;

[0057] FIG11(3b), FIG11(3c), and FIG11(3d) are respectively the detection results of the airport detection using YOLOv8s, the model of the present invention, and the HR model of the present invention on the attention heat map of FIG11(3a);

[0058] Figure 11 (4a) is the original image of remote sensing image 4;

[0059] Figure 11(4b), Figure 11(4c), and Figure 11(4d) are respectively the detection results of airport detection on the attention heat map of Figure 11(4a) using YOLOv8s, the model of the present invention, and the HR model of the present invention. DETAILED DESCRIPTION

[0060] The model used in the present invention for remote sensing airport identification is a model obtained by improving the YOLOv8 model. Before introducing the method of the present invention, the deep learning target recognition model, the dilated convolution and ASPP module, the Attention module, and the standard YOLOv8 model are first introduced.

[0061] 1) Deep learning target recognition model. Neural network target recognition algorithms are mainly divided into two categories: one-stage and two-stage. The two-stage algorithm is represented by the RCNN algorithm series, which focuses on the accuracy of the algorithm. The one-stage algorithm is represented by the YOLO series, which pays more attention to the real-time and efficiency of the algorithm. The YOLO algorithm series has always been a hot topic in the field of image target recognition due to its simple and efficient characteristics. The first proposed target detection idea, its model is called YOLOv1, and the later YOLO9000 has supported the detection of more than 9,000 types of targets through various improvements. Then YOLOv3 introduced multiple detection head branches and fused the structure with the feature pyramid to improve the detection effect of targets of multiple sizes in a picture. YOLOv3 only used the image pyramid structure when fusion, and YOLOv4 added the PAN structure. All detection head input feature maps are fused with multi-scale features. On this basis, YOLOv5 has made comprehensive improvements in data enhancement, model structure, IoU, etc., creating a complete and powerful target detection algorithm. The YOLOX algorithm proposed an anchor-free method architecture and decoupled the prediction head, making the algorithm simpler and clearer and more effective. The YOLOv8 used in the present invention is an anchor-free target detection algorithm optimized on the basis of YOLOv5.

[0062] 2) Dilated convolution and ASPP module. The ASPP module was proposed in the segmentation model DeepLabv2. Its general structure is as follows: Figure 1 As shown in the figure, it is composed of dilated convolutions with different dilation rates. Generally, if you want to improve the visual receptive field, you can use a larger convolution kernel, but it will cause a surge in the amount of computation and parameters of the model. You can also use a larger convolution step size to improve the visual receptive field, but it will lose the resolution of the feature map, which is not conducive to the subsequent fusion of multi-scale information. Using dilated convolutions can avoid the surge in parameters and computation caused by large kernel convolutions, and avoid the loss of resolution caused by increasing the step size. Dilated convolutions are achieved by adding holes to conventional convolutions, which is equivalent to expanding the sampling range. Sparse sampling is performed in a large sampling range, which can realize feature extraction within a large context range. The ASPP module extracts multi-scale context information through dilated convolutions with different dilation rates, and realizes the fusion of multi-scale context information through feature map splicing operations and a conventional convolution layer.

[0063] 3) Attention module. In computer vision tasks, adding an attention module to the basic model can effectively improve the model's performance. The structure of a conventional attention module is as follows: Figure 2As shown, its working mechanism is to calculate a weight tensor for a given feature map through the designed attention module, and then multiply the weight tensor with the given feature map to weight the feature map. According to the different final weighting effects, it is divided into spatial attention and channel attention. Spatial attention highlights the features of certain areas by weighting the values ​​at different positions of the feature map, and channel attention highlights the features of certain feature channels by weighting different feature channels of the feature map. Commonly used attention modules for computer vision tasks include SE module, CBAM module, etc. The SE module is a channel attention module, and the CBAM module combines channel attention and spatial attention. Usually, CBAM uses the maximum and average methods to extract the main features in both channel attention and spatial attention. The CBAM used in the present invention only uses the average method to extract the main features of the feature map.

[0064] 4) Standard YOLOv8 model. The standard YOLOv8 model mainly consists of three parts: the feature extraction network Backbone, the feature fusion network PAFPN structure and the detection head Head. The feature extraction network is mainly used to extract features of different scales from the input image, the feature fusion network is mainly used to complete the fusion of feature maps of different scales, and the final multiple detection heads are responsible for predicting the category and position of the target from feature maps of different scales. The YOLOv8 model replaces the C3 module in YOLOv5 with the C2f module. The C2f module is an optimization of the DenseBlock structure in the DenseNet model, which enriches the gradient information and improves the computational efficiency of the module. The final detection head uses the Decoupled-Head idea to separate the category prediction and position regression, which can improve the detection effect of the model.

[0065] On this basis, the improved YOLOv8 target detection model used in the present invention for remote sensing airport identification is introduced. The improvements of the improved YOLOv8 target detection model include:

[0066] Improvement 1: Add the CSPASPP module to the shallow output part of the feature extraction network. Because the visual receptive field of the shallow network is small and cannot fully utilize the context features, the CSPASPP module is used to extract and fuse context features of different ranges for the feature map to make fuller use of the context features. The network structure of the CSPASPP module is as follows: Figure 3 As shown in FIG, the module is obtained by improving the traditional ASPP module by drawing on the idea of ​​CSPNet. This module can realize local multi-scale feature fusion and better meet the extraction and fusion of multi-range context features. Figure 3The CBS module in the figure represents the combination of Conv, BatchNormal, and SiLU, and the letters below represent the number of input feature layers, the number of output feature layers, and the core voiding rate. The first CBS module of this module can complete the compression of the feature map according to the needs, and then through multiple CBS modules with different voiding rates, the extraction of context features in different ranges can be realized according to the configuration, and then the fusion and compression of context feature maps in different ranges are realized through splicing and a CBS module, and finally spliced ​​with the compressed input feature map, and the feature fusion and compression are realized through the last CBS module. The size and number of layers of the final output feature map are the same as the input feature map.

[0067] Improvement 2: Add the CBAM attention mechanism before the deepest Conv layer and the second deepest Conv layer of the feature extraction network, and add the CBAM attention mechanism before the detection head. YOLOv8 uses a Conv module with a step size of 2 to implement downsampling operations, and downsampling operations easily lead to the loss of feature information. Therefore, the present invention increases the CBAM attention mechanism to focus on the features of key areas and channels, thereby reducing the risk of feature information loss caused by downsampling. For the detection head, a CBAM attention mechanism is added before each of the three inputs to focus on the features of important areas and channels again. Among them, the YOLOv8 target detection model obtained by performing improvements to the YOLOv8 model including "Improvement 1" and "Improvement 2" is shown as follows Figure 4 As shown (the following experimental part will refer to this model as the “model of the present invention”).

[0068] Improvement 3: Add a Conv module with a step size of 2 before the second C2f module of the feature extraction network (i.e., the second shallow C2f module of the feature extraction network), and add a Conv module with a step size of 1 before the third and fourth C2f modules of the model (i.e., the second deep C2f module and the deepest C2f module of the feature extraction network). This improvement is Figure 4 Not drawn in the figure. This improvement is designed based on the airport target features in the data set. The image size of the data set is large and contains many airport targets with a large distribution range. Therefore, the local texture information extracted by the shallow network is not very important. Therefore, a downsampling Conv module is added before the second shallow C2f module. This can not only concentrate and filter the shallow information and increase the visual receptive field of the model, but also reduce the computing power consumption of subsequent calculations. This will cause the information of the deep feature map to be more macroscopic and complex. Therefore, a Conv module with a step size of 1 is added before the third and fourth C2f modules of the model to enhance the learning ability of the model. The model obtained by the improvements including "Improvement 1", "Improvement 2" and "Improvement 3" is hereinafter referred to as the "HR (high resolution image) model of the present invention".

[0069] The improved YOLOv8 target detection model introduced above is trained. After the training is completed, the trained improved YOLOv8 target detection model can be used to accurately identify the airport in the remote sensing image. That is, the remote sensing image is input into the trained improved YOLOv8 target detection model to obtain the airport recognition result in the remote sensing image.

[0070] The following experiments are conducted to verify the effectiveness of the method of the present invention.

[0071] 1) Preparation of airport dataset. Currently, there is a lack of large-scale, public, and high-quality datasets in the field of airport identification, which has greatly hindered research in this field. A high-quality dataset that conforms to the distribution rules of reality can often train a neural network model with better results, higher robustness, and more applicability. Therefore, this embodiment produces a remote sensing image airport identification dataset RSAD (remote sensing airport dataset). RSAD images are derived from Google remote sensing imagery. RSAD selects 1,000 airports worldwide, and mainly obtains 2,985 remote sensing images of airports at three typical resolutions, and finally annotates them. In terms of background richness, RSAD contains airport targets in almost all kinds of real backgrounds, such as cities, islands, artificial islands, seashores, deserts, hills, farmlands, etc. Figure 5(1a) to Figure 5(7c) In terms of the diversity of forms, RSAD includes many types of airports, from simple single-runway airports to complex multi-runway airports, such as Figure 6(1a) to Figure 6(3c) As shown in Figure 2, in terms of airport size, RSAD greatly improves the size richness of airports in RSAD by capturing airports at different resolutions. Figure 7(1a) to Figure 7(3c) Subsequent experiments were all completed based on RSAD.

[0072] 2) Experimental environment and parameter configuration. The hardware environment of the experiment is CPU: Intel I9-10900, GPU: NVIDIARTX 3090, the software environment is: Windows 10 operating system, Python version is 3.9.7, and pytorch version is 1.10.2. The dataset is divided into training validation set and test set in a ratio of 3:1, and the training validation set is divided into training set and validation set in a ratio of 9:1. All models are trained on the training set, the relevant parameters are adjusted and tested on the validation set, and the final test is performed on the test set. The input image size is 1024*1024. The data in the training process is augmented by Mosaic data, and the model is optimized using SGD stochastic gradient descent. The training batch is 200 epochs. After the final observation of the training loss, validation set test indicators and other data, such as Figure 8(a) to Figure 8(d) As shown in the figure, the model can basically reach a stable state after training for 200 epochs.

[0073] 3) Evaluation method. Commonly used evaluation criteria in target detection tasks include recall, precision, F1 score and mAP. The final detection results can be divided into four categories: TP, TN, FP, and FN. TP is a true positive example, and its value is the number of correctly identified positive samples; TN is a true negative example, and its value is the number of correctly identified negative samples; FP is a false positive example, and its value is the number of negative samples mistakenly identified as positive samples; FN is a false negative example, and its value is the number of positive samples identified as negative samples. The value of the recall rate Recall can reflect the ability of the model to detect the target, and the value of the precision rate Precision can reflect the ability of the model to correctly detect the target. The F1 score is the harmonic mean of Recall and Precision, and their formulas are shown below.

[0074]

[0075]

[0076]

[0077] Since the recall rate, precision rate and F1 score are all related to the confidence threshold, and the confidence threshold is a variable, it is difficult to fully reflect the effect of the model. Therefore, mAP is often used to reflect the effect of the model. AP is the value of the area enclosed by the PR (Precision-Recall) curve and the coordinate axis, and mAP is the average value of AP of all categories.

[0078] 4) Test results and analysis. The models used in this embodiment are finally tested on the test set. The input image size is the same as that in training, which is 1024*1024. The specific test results are shown in Table 1, where the F1 value is the maximum value that the model can achieve by adjusting the confidence threshold, Precision and Recall are the values ​​corresponding to this F1, mAP@0.5 is the mAP corresponding to the IoU threshold of 0.5, and mAP@0.5:0.95 is the average value of all mAPs obtained at different IoU thresholds (0.5 to 0.95, with a step size of 0.05).

[0079] Table 1

[0080]

[0081] According to the data in the table, the model of the present invention exceeds the original versions of the YOLOv8, YOLOv5, and YOLOv3-tiny algorithms in terms of Recall, F1, mAP@0.5, and mAP@0.5:0.95 indicators. The HR model of the present invention, which is improved on this basis, has achieved higher scores in multiple indicators. The analysis is that this is because the HR version of the algorithm adds new downsampling and convolution modules to the feature extraction network, which improves the visual receptive field of the model and can filter out more noise. The model's ability to detect large-scale airports and complex background airports has been enhanced, but because the early downsampling lost more detail information, the accuracy of the target frame and the effect of small targets are not as good as YOLOv8s and the original method of the present invention. YOLOv8s, the model of the present invention, and the HR model of the present invention were used to detect non-training pictures, and the results are as follows Figure 9 (1a) to Figure 10 (6d) As shown, it can be seen that for remote sensing images with simple background, clear image and simple target structure, the above three detection methods can complete the task well. However, for remote sensing images with complex background, unclear image and small target, the present invention has more advantages.

[0082] This embodiment also generates an attention heat map through the Grad-Cam method to analyze the focus position of the model. The results are as follows: Figure 11(1a) to Figure 11(4d) As shown, it can be seen that the present invention focuses more on the target.

[0083] For the remote sensing airport identification task, we first created an airport target detection dataset based on Google remote sensing images. This dataset has rich backgrounds, multi-scale airport targets, and multi-complexity airport targets, which effectively solves the problems of insufficient data volume and insufficient data samples in this task. Then, based on the latest YOLOv8 target detection algorithm and the corresponding data augmentation method, the function of remote sensing airport identification is realized, and better results are achieved on relatively simple samples. This embodiment also makes targeted improvements on the basis of YOLOv8 to obtain the model of the present invention and the HR model of the present invention. Experiments show that these two models have better effects on remote sensing airport target detection tasks, have more advantages in detecting targets on more complex images, and can better meet the needs of remote sensing image recognition of airports.

Claims

1. A remote sensing airport identification method, characterized in that: Acquire remote sensing images and input them into the trained improved YOLOv8 target detection model to obtain airport recognition results in remote sensing images; Among them, the improved YOLOv8 target detection model includes a YOLOv8 model and two CSPASPP modules, and the two CSPASPPs are respectively set on two shallow feature outputs of the feature extraction network included in the YOLOv8 model; a CSPASPP module includes multiple convolution modules, which are used to make the input pass through a convolution module for feature map compression, and the compressed feature map passes through multiple convolution modules with different void rates for context feature extraction, and the context feature extraction results are spliced ​​and passed through a convolution module for context feature fusion and compression, and then spliced ​​with the compressed feature map, and after splicing, they pass through a convolution module for feature fusion and compression.

2. The remote sensing airport identification method according to claim 1, characterized in that: The improved YOLOv8 target detection model also includes two CBAM attention mechanism modules respectively set on the input of the deepest Conv layer and the input of the second deepest Conv layer in the feature extraction network.

3. The remote sensing airport identification method according to claim 2, characterized in that: The improved YOLOv8 target detection model also includes three CBAM attention mechanism modules respectively set on three inputs of the detection head included in the YOLOv8 model.

4. The remote sensing airport identification method according to any one of claims 1 to 3, characterized in that: The improved YOLOv8 target detection model also includes three Conv modules respectively arranged on the input of the second shallow C2f module, the input of the second deep C2f module and the input of the deepest C2f module in the feature extraction network.

5. The remote sensing airport identification method according to claim 4, characterized in that: The step size of the Conv module set on the input of the second shallow C2f module is 2, and the step size of the Conv module set on the input of the second deep C2f module and the input of the deepest C2f module are both 1.

6. The remote sensing airport identification method according to any one of claims 1 to 3, characterized in that: The convolution module is a CBS module.