Remote sensing image aircraft target detection method based on improved Oriented R-CNN
By improving the Oriented R-CNN algorithm, adding a self-attention mechanism and a DBR deformable convolution module, and building a bottom-up information transmission channel, the problem of insufficient accuracy and robustness of remote sensing image aircraft target detection in complex backgrounds is solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510489772.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-06-24
AI Technical Summary
The existing remote sensing image aircraft target detection algorithms have problems with insufficient accuracy and robustness in complex backgrounds, especially in the multi-directional, multi-scale features and complex background interference of targets.
The Oriented R-CNN algorithm is improved, and the self-attention mechanism is added to the backbone network ResNet-50, replaced with ResNet-SA, and a DBR deformable convolution module is added to the neck network FPN to build a bottom-up information transmission channel, and at the same time, multiple detection head networks are used to improve detection accuracy.
In complex backgrounds, the accuracy and robustness of aircraft target detection are significantly improved, and the multi-directional and multi-scale characteristics of the target can be better handled, the interference of complex backgrounds is reduced, and the error and missed detection rates are reduced.
Smart Images

Figure CN120198651A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an improved remote sensing image aircraft target detection algorithm. Background Art
[0002] As an important means of obtaining geographical information, remote sensing images play an important role in fields such as aviation monitoring, national land surveying and mapping, and traffic management. Among them, aircraft target detection is one of the key tasks in remote sensing image analysis and is widely used in scenarios such as aviation safety, airport management, and UAV cruising. However, remote sensing images have the complex background characteristics of traditional optical images and UAV images, making the detection of aircraft targets face many challenges.
[0003] Existing target detection algorithms still have performance bottlenecks in remote sensing images. In order to improve the accuracy and robustness of aircraft target detection, there is an urgent need for a new method that can effectively process multi-directional and multi-scale features of targets while suppressing complex background interference. The improved algorithm based on Oriented R-CNN provides a new technical idea for solving the above problems by adding an attention mechanism, improving the direction perception mechanism, and multi-angle feature extraction strategy. Summary of the Invention
[0004] The purpose of the present invention is to provide an improved remote sensing image aircraft target detection method based on Oriented R-CNN, which has higher accuracy for aircraft target detection and better meets the current requirements for the accuracy of recognized targets in the case of complex backgrounds and the characteristics of different sizes and dense arrangements of aircraft targets.
[0005] Aiming at the deficiencies of the existing technology, the present invention provides an improved remote sensing image aircraft target detection method based on Oriented R-CNN, including obtaining a remote sensing data set, cleaning the original data to remove blurred and duplicate data, and finally annotating the obtained image data. The preprocessed remote sensing image data set is divided into a training set, a validation set, and a test set, where the remote sensing image data set includes aircraft images under different complex background conditions.
[0006] Replace the backbone network ResNet-50 of Oriented R-CNN with ResNet-SA to ensure that the model can pay more attention to and extract important information of the target area in a complex environmental background. At the same time, improve the neck FPN network of the original algorithm structure. On the basis of FPN, a bottom-up information transmission channel is built, effectively solving the problem that the information contained in the low-level feature map cannot be transmitted to the high-level feature map due to the one-way information transmission of the FPN network, further enriching the semantic information and target position information contained in the feature map, and being able to better meet the requirements of the actual remote sensing image aircraft target detection accuracy. At the same time, add the DBR (Deformable Convolutional Networks-BN-ReLU) deformable convolution module to enable the network to better capture the complex relationships between aircraft features. In addition, cascade multiple detection head networks to correct the prediction boxes, which can alleviate the noise detection caused by the low IoU threshold and avoid the overfitting problem caused by directly increasing the IoU value.
[0007] Use the new network model obtained by the improvement method as above, iteratively train the new model with the training set to obtain the target weights. Use the test set to test the target weights to obtain the remote sensing image aircraft target detection results.
[0008] The described ResNet-SA network structure is used as the backbone network of the improved Oriented R-CNN algorithm structure. By adding a self-attention mechanism (Self-Attention) to the original ResNet-50, specifically as Figure 2 shown, ResNet-SA aims to enhance the comprehensiveness and accuracy of feature extraction. The self-attention mechanism enables the model to better focus on important features by capturing global context information, thereby improving the detection accuracy of targets in any direction. This improvement not only optimizes the feature representation but also enhances the model's adaptability to complex scenes and rotated targets, ultimately improving the detection accuracy and robustness.
[0009] The described neck network FPN network of the Oriented R-CNN algorithm structure builds a low-to-up information transmission channel, so that the semantic information and target position information contained in the high-level feature map can be transmitted to the underlying feature map, so that the model can learn more target information. A DBR convolution block is added before the FPN network structure. This module adds a BN layer and a ReLU activation function after the output of the DCN (Deformable Convolutional Networks) deformable convolution module, so that the network structure is more robust to changes in remote sensing aircraft image target data. The ReLU activation function can ensure that the network can learn and express complex, nonlinear data features, and can better capture the complex relationship between aircraft features in remote sensing images.
[0010] The cascaded detection head network takes the features {P2, P3, P4, P5} and a series of rotated candidate boxes generated by Oriented RPN as input, first extracts the features corresponding to the selected boxes, obtains the features of each candidate box that are rotationally invariant and fixed in size, and then sends the features to the prediction network for classification and regression prediction, cascading three prediction heads to progressively correct the prediction box. Given a rotated box (x, y, w, h, θ), first project it to the corresponding feature layer according to the area w×h of the rotated candidate box to obtain the Rotated Region of Interest (R-RoI), assuming that the feature The stride is s, then the parameters of R-RoI (xr, yr, wr, hr, θ) are calculated as follows: x r =[x / s],y r =[y / s],w r =w / s,h r =h / s In the formula, [.] means rounding down.
[0011] Get fixed-size features through R-RoI Align After that, classification and regression prediction can be performed based on F. Specifically, the prediction head networks with the same structure are cascaded, and the IoU threshold for dividing positive and negative samples is set to increase step by step. During training, the sequential step-by-step training method is adopted. First, the regression result of RoL Head1 is decoded to obtain the prediction and regression prediction box, and then the prediction box is resampled (positive and negative samples are re-divided and sampled) using a higher IoU threshold than RoL Head1. Then, the new sample set is used to train RoL Head2, and RoL Head3 repeats the above steps, that is, the regression output of a certain stage is used to train the next stage network. The goal of each stage is to find a higher quality sample set for the next stage.
[0012] During testing, the same cascaded structure is used. All candidate boxes pass through the prediction head network in sequence, and the final classification score is the average of the classification predictions of the three prediction head networks. By adopting the above cascaded prediction head form to optimize the prediction boxes stage by stage, it can not only alleviate the noise detection caused by low IoU thresholds, but also avoid the overfitting problem caused by the exponential disappearance of positive samples during training due to directly increasing the IoU threshold, and the mismatch problem between the optimal IoU of the detector during testing and the input assumed IoU, thus achieving significant performance gains.
[0013] In summary, the present application provides a method for detecting aircraft targets in remote sensing images by improving Oriented R-CNN. The backbone network is replaced with the ResNet-SA structure designed in this application. A self-attention mechanism (Self-Attention) is added after the third and fourth layer STAGE modules in the original ResNet-50 structure to improve the expressive ability of global feature modeling, which is particularly important when dealing with aircraft target detection in high-resolution remote sensing images. At the same time, by improving the neck network and adding the designed DBR deformable convolutional network and building an information transmission channel from bottom to top, the semantic information and target position information contained in the high-level feature map can be transmitted to the low-level feature map, enabling the model to learn more target information. This is used to solve problems such as false detection and missed detection caused by different sizes and dense parking of aircraft in remote sensing images. Finally, through the designed cascaded detection head network, the rotated candidate boxes obtained after the above processing are sequentially detected, continuously optimizing the candidate boxes to alleviate noise and reducing the overfitting problem during training, further improving the accuracy of the model in aircraft target detection applications in remote sensing images. Brief Description of the Drawings
[0013] Figure 1 It is a technical roadmap of a method for detecting aircraft targets in remote sensing images provided by the present invention.
[0014] Figure 2 It is a structural diagram of the ResNet-SA backbone network provided by the present invention.
[0015] Figure 3 It is a structural diagram of an improved neck network provided by the present invention
[0016] Figure 4 It is a structural diagram of a DBR deformable convolutional block provided by the present invention.
[0017] Figure 5 It is a schematic diagram of the structure of a cascaded detection head network provided by the present invention.
[0018] Figure 6 It is a structural diagram of the improved Oriented R-CNN network provided by the present invention. Detailed Embodiments
[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings and implementation schemes in the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present application.
[0021] Dataset: The dataset is a large image dataset for object detection in remote sensing images, which can be used to discover and evaluate objects in aerial images. The dataset has more than 5,000 pictures, more than 15,000 instances, and a total of 2 categories.
[0022] Use Labellmg to perform object annotation on the collected images, annotating the category and location information of the objects.
[0023] Perform operations such as denoising and enhancing contrast on the collected images to improve the image quality; enhance the images by means of random cropping, rotation, scaling, etc. to increase the diversity of training samples; divide the dataset into a training set, a validation set, and a test set for model training, tuning, and evaluation. The training set, the validation set, and the test set are all composed of different remote sensing images under different disaster conditions.
[0024] In this example, Oriented R-CNN is selected as the basic model, and its model structure is improved to enhance the detection ability of aircraft targets in remote sensing images.
[0025] The content to be studied in the present invention includes changing the original backbone network ResNet-50. By adding a self-attention network after the end of the third and fourth layers of the original backbone network, a ResNet-SA backbone network is formed. After the input remote sensing image with aircraft targets is resized and normalized, the image first enters the first and second layers of ResNet-SA to extract low-level features. In the third layer, due to the existence of the attention mechanism, by capturing global context information, the output feature expression of this layer is strengthened, especially in the spatial dimension, making it more distinguishable; in the fourth layer, the extraction of high-level features is further optimized, the expression ability of high-level features is enhanced, and the detection accuracy under complex backgrounds is improved.
[0026] Improvements to the neck network include adding an enhanced deformable convolution module DBR, adding a BN layer and a ReLU activation function again in the deformable convolution network. The BN layer reduces internal covariate shift, speeds up the convergence of the network, and improves the generalization ability of the model by normalizing each batch of data during the training process of the neural network. Finally, it passes through the ReLU activation function, which is defined as a widely used non-linear activation function as follows: f(x) = max(0, x) This functional mechanism truncates the negative components in the input value x to zero while keeping the positive components unchanged, thus cleverly realizing the non-linear transformation of the feature map. The computational efficiency is significantly higher than traditional activation functions such as sigmoid and tanh, and the operation logic is simple and clear, greatly reducing the computational burden.
[0027] In the neck network, a bottom-up information transmission channel is built, effectively solving the problem that the information contained in the low-level feature map cannot be transmitted to the high-level feature map due to the one-way information transmission of the traditional FPN network. By adding DBR deformable convolution in the neck network and building a bottom-up information transmission channel, not only can the shape and size of the receptive field be automatically adjusted according to the irregularity of the target on the remote sensing image, but the bottom-up information transmission channel further enriches the semantic information and target location information contained in the feature map.
[0028] After obtaining fixed-size features through R-RoI Align , classification and regression prediction can be performed based on F. Specifically, cascade prediction head networks with the same structure, set the IoU threshold for dividing positive and negative samples to increase gradually, and adopt a sequential step-by-step training method during training. First, decode the regression results of RoL Head1 to obtain prediction and regression prediction boxes, then resample the prediction boxes (re-divide positive and negative samples and sample) using an IoU threshold higher than that of RoL Head1, and then use the new sample set to train RoL Head2. RoL Head3 repeats the above steps, that is, use the regression output of a certain stage to train the network of the next stage, and the goal of each stage is to find a higher-quality sample set for the next stage.
[0029] After the input first completes feature extraction through the backbone network and the neck network, a series of high-quality candidate boxes that can adapt to the tilt angle of the target are generated in the Oriented RPN network structure. The regression targets for this region are: x in the formula * , y * , w * , h * , θ *respectively represent the center coordinates, width, height, and angle of the ground truth box; x a , y a , w a , h a respectively represent the corresponding parameters of the anchor box; norm(θ) represents the normalization function of the angle, which is used to limit the angle range. In the present invention, the long-edge 90 definition method is selected to define the rotated box, and the angle range is [-π / 2, π / 2). Therefore, the normalization function can be used as: norm(θ) = (θ + π / 2) % π - π / 2 Here represents the offset of the ground truth box relative to the anchor box.
[0030] For the prediction detection head RoL Head, in order to make the offset rotation-invariant, the relative offset proposed by RoI Trandformer is used, that is, the local coordinate system determined by the rotated candidate box is used instead of the global coordinate system determined by the image. Its regression target is: In the formula, x r , y r , w r , h r , θ r respectively represent the center coordinates, width, height, and angle of the rotated candidate box; the meanings of the remaining variables are the same as those in the above formula. Here represents the offset of the ground truth box relative to the rotated candidate box. The classification loss and the bounding box regression loss are gradually optimized through the cascade detection head network design to ensure the prediction accuracy.
[0031] In summary, the remote sensing image aircraft target detection method based on the improved Oriented R-CNN proposed by the present invention has excellent detection accuracy for the characteristics that aircraft targets are densely arranged and have different sizes in complex backgrounds. The implementation method of the present invention application has important practical application value and can be applied to the remote sensing image aircraft target detection task.
[0032] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as these technical feature combinations do not conflict, they should all be considered as the scope recorded in this specification.
[0033] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
Claims
1. A remote sensing image aircraft target detection method based on improved Oriented R-CNN, characterized in that: Here are the steps: S1: Get the dataset; S2: Use LABELLMG to annotate the collected images and mark the category and location information of the target; S3: De-noise the collected images, increase the contrast, and other operations to improve the image quality; enhance the images by random cropping, rotation, scaling, and other methods to increase the diversity of training samples; S4: Divide the dataset into training set, validation set and test set for model training, tuning and evaluation; To test the excellent performance and generalization ability of remote sensing image aircraft targets in different environments; S5: For the path using self-attention, the intermediate features are aggregated into N groups, each group contains three feature maps and each group is a feature generated by a different 1*1 convolution, and then these three feature maps are input into the multi-head self-attention module as query key values; S6: ResNet-SA aims to enhance the comprehensiveness and accuracy of feature extraction. The self-attention mechanism enables the model to better focus on important features by capturing global context information, thereby improving the detection accuracy of targets in any direction; a bottom-up information transmission channel and a deformable convolution module are built on the basis of FPN, so that the network can better capture the complex relationship between aircraft features. In addition, the use of cascaded multiple detection head networks to correct the prediction box can alleviate the noise detection caused by low IoU thresholds and avoid overfitting problems caused by directly increasing the IoU value.
2. The method for detecting aircraft targets in remote sensing images based on an improved Oriented R-CNN according to claim 1, characterized in that Accuracy, recall, mean average precision (mAP) and average precision were used as evaluation indicators.
3. The method for detecting aircraft targets in remote sensing images based on improved Oriented R-CNN according to claim 2, characterized in that: Precision refers to the ratio of the number of samples correctly predicted as positive examples by the model to the number of all samples predicted as positive examples; recall refers to the ratio of the number of samples correctly identified as positive examples by the model to the number of samples that are actually positive examples; average precision is an indicator for evaluating the performance of the model in each category, which is obtained by calculating the average precision (AP) of each category and averaging the AP of all categories.
4. The method for detecting aircraft targets in remote sensing images based on improved Oriented R-CNN according to claim 1, characterized in that: The cascade detection head network corrects the prediction box by gradually increasing the IoU threshold, thereby improving the detection accuracy while ensuring the recall rate.
5. The method for detecting aircraft targets in remote sensing images based on improved Oriented R-CNN according to claim 1, characterized in that: The cascade detection head network corrects the prediction box by gradually increasing the IoU threshold, thereby improving the detection accuracy while ensuring the recall rate.
6. The method for detecting aircraft targets in remote sensing images based on improved Oriented R-CNN according to claim 1, characterized in that: The neck FNP network builds a top-down and bottom-up bidirectional information transmission channel and integrates multi-scale features to enhance the model's detection ability for small and dense targets.
7. The method for detecting aircraft targets in remote sensing images based on improved Oriented R-CNN according to claim 1, characterized in that: The method is applicable to a variety of remote sensing image scenes, including but not limited to aircraft target detection in complex environments such as airports, ports and military bases.