Remote sensing image aircraft target detection method based on regional importance perception attention
By introducing a regional importance perceived attention mechanism in remote sensing image aircraft target detection, the error and missed detection problems in complex backgrounds and target detection at different scales are solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510113008.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-24
AI Technical Summary
Existing remote sensing image aircraft target detection methods have problems of mis-detection and missed detection when dealing with complex backgrounds and targets of different scales, and traditional attention mechanisms are difficult to accurately identify aircraft targets and distinguish backgrounds.
Using a method based on regional importance perceived attention, through the combination of image preprocessing, feature extraction, regional importance perceived attention module and target detection network, key features closely related to aircraft targets are intelligently identified and enhanced, and background information is suppressed.
It significantly improves the accuracy and robustness of target detection, can effectively highlight target characteristics and suppress background information in complex scenarios, and reduce the possibility of false detection and missed detection.
Smart Images

Figure CN120047669A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target detection, and more particularly to a method for detecting aircraft targets in remote sensing images based on region importance-aware attention. Background Art
[0002] With the rapid development of remote sensing technology, remote sensing images have been widely used in many fields such as geographic information monitoring, environmental protection, resource management, disaster assessment, military reconnaissance, etc. In these applications, the detection of targets in remote sensing images, especially the detection of aircraft targets, has important practical significance. Aircraft targets in remote sensing images usually have small sizes, high deformations, low contrasts, and complex backgrounds. Therefore, traditional target detection methods face many challenges, such as fuzzy target features, large interference from background information, complex target scale changes, etc.
[0003] Traditional remote sensing image target detection methods mostly rely on manual feature extraction and classical machine learning models. Early detection methods used techniques such as edge detection, region growing, and morphological processing to extract target features in images, and then used classification algorithms such as support vector machines (SVM) and K-nearest neighbors (K-NN) for target recognition. However, these methods usually rely on manually designed features and have poor adaptability to image quality and environment, and perform poorly in dealing with complex backgrounds, targets of different scales, etc.
[0004] With the rise of deep learning technology, convolutional neural networks (CNNs) have been widely used in remote sensing image target detection, especially for the detection of aircraft targets. Deep learning methods can automatically learn high-level features from raw images, avoiding the cumbersome manual feature extraction. However, existing deep learning-based target detection methods still have some problems, such as: 1) The problem of background interference. In remote sensing images, aircraft targets are often mixed with complex backgrounds (such as clouds, ground buildings, etc.), and existing detection methods are difficult to effectively suppress background information, resulting in serious false detection and missed detection phenomena. 2) The problem of scale change. The sizes of targets in remote sensing images vary greatly, and aircraft targets from the ground to the air may present different sizes in the image. Existing methods may have poor detection effects on small targets due to the limitation of the receptive field when dealing with targets of different scales.
[0005] To address these problems, in recent years, researchers have proposed various optimization methods, such as multi-scale feature fusion, attention mechanism, region proposal network and other techniques, aiming to improve the accuracy and efficiency of target detection. In particular, the introduction of the attention mechanism has become one of the key technologies to enhance the target detection effect. The attention mechanism simulates the attention selectivity of human vision, gives the network more attention to key information, and suppresses unimportant information, thereby improving the performance of the model.
[0006] However, the existing attention mechanisms mainly focus on the weighting of global information or local information, and the way of assigning weights to the importance of different regions is relatively single. Especially in remote sensing images, the complexity of the background and the diversity of targets make it difficult for traditional attention mechanisms to accurately identify aircraft targets and distinguish them from the background. Therefore, how to better process the targets in remote sensing images through regional importance perception to improve the target detection performance is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0007] In view of this, the purpose of the present invention is to provide a method for detecting aircraft targets in remote sensing images based on regional importance perception attention.
[0008] To achieve the above purpose, the present invention adopts the following technical solutions:
[0009] In the first aspect, a method for detecting aircraft targets in remote sensing images based on regional importance perception attention is provided, including the following steps:
[0010] S1: Perform image preprocessing on the remote sensing image to obtain the preprocessed remote sensing image;
[0011] S2: Input the preprocessed remote sensing image into a feature extraction network to obtain a multi-scale fusion feature map;
[0012] S3: Perform preprocessing on the multi-scale fusion feature map to obtain an initial feature map;
[0013] S4: Input the initial feature map into a regional importance perception attention module to obtain an enhanced feature map;
[0014] S5: Input the enhanced feature map into a target detection network to obtain the aircraft target detection result.
[0015] Preferably, the image preprocessing in S1 includes normalization, contrast enhancement, brightness adjustment, denoising, cropping or scaling.
[0016] Preferably, S2 specifically includes:
[0017] S21: Input the preprocessed remote sensing image into a pre-trained improved ResNet-50 network;
[0018] S22: Extract the feature maps output at different stages of the pre-trained improved ResNet-50 network, and splice or weighted fuse the upsampled high-resolution feature maps with the low-resolution feature maps to obtain the multi-scale fusion feature map.
[0019] Preferably, the pre-trained improved ResNet-50 network removes the fully connected layer, and uses the second stage, the third stage, the fourth stage, and the fifth stage as the output layers for feature extraction.
[0020] Preferably, the preprocessing in S3 includes adjusting the size of the feature map and normalizing the feature map.
[0021] Preferably, the region importance-aware attention module includes a 1*1 convolutional layer, a first matrix multiplication unit, an exponentiation operation unit, a global average pooling layer, a second matrix multiplication unit, an inverse operation unit, a first 3*3 convolutional layer, a second 3*3 convolutional layer, a first activation function layer, a bilinear transformation layer, a second activation function layer, and a third matrix multiplication unit;
[0022] The input end of the 1*1 convolutional layer is used to input the initial feature map;
[0023] The output end of the 1*1 convolutional layer is respectively connected to the input end of the first matrix multiplication unit and the input end of the exponentiation operation unit;
[0024] The output end of the exponentiation operation unit is respectively connected to the input end of the first matrix multiplication unit and the input end of the global average pooling layer;
[0025] The output end of the first matrix multiplication unit is connected to the input end of the global average pooling layer;
[0026] The output end of the global average pooling layer is respectively connected to the input end of the second matrix multiplication unit and the input end of the inverse operation unit;
[0027] The output end of the inverse operation unit is connected to the input end of the second matrix multiplication unit;
[0028] The output end of the second matrix multiplication unit is sequentially connected to the input end of the third matrix multiplication unit through the first 3*3 convolutional layer, the second 3*3 convolutional layer, the first activation function layer, and the bilinear transformation layer;
[0029] The input end of the second activation function layer is used to input the first channel feature map in the initial feature map;
[0030] The output end of the second activation function layer is connected to the input end of the third matrix multiplication unit;
[0031] The input end of the third matrix multiplication unit is also used to input the initial feature map.
[0032] Preferably, S5 specifically includes: performing target frame regression and classification on the enhanced feature map through the detection head of the target detection network, and then removing redundant frames in combination with non-maximum suppression, and finally outputting an aircraft target detection result with category and confidence.
[0033] In a second aspect, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for detecting aircraft targets in remote sensing images based on regional importance-aware attention as described in any one of the above items is implemented.
[0034] According to a third aspect, a non-transitory computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for detecting aircraft targets in remote sensing images based on regional importance-aware attention is implemented as described in any one of the above items.
[0035] In a fourth aspect, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the remote sensing image aircraft target detection method based on regional importance-aware attention as described in any one of the above items.
[0036] It can be seen from the above technical solutions that, compared with the prior art, the present invention discloses a remote sensing image aircraft target detection method based on regional importance perception attention, which can achieve the following beneficial effects:
[0037] 1) By introducing a regional importance-aware attention mechanism, the present invention can intelligently identify and enhance key features in the image that are closely related to the aircraft target. This mechanism is particularly suitable for complex scenes, including remote sensing images with low contrast, low resolution, or large changes in target scale. By adaptively adjusting the focus on features, the present invention can effectively highlight target features while suppressing background information that is not related to the target, thereby significantly improving the accuracy of target detection in these challenging environments.
[0038] 2) The region importance-aware attention mechanism also greatly enhances the robustness of the model. In the face of different environmental conditions and image quality, the mechanism can maintain stable performance and reduce the possibility of false detection and missed detection. This robustness is particularly important for practical applications because it ensures that the system can reliably perform target detection tasks even under undesirable observation conditions and provide users with accurate and reliable results. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0040] Figure 1 It is a flowchart of a remote sensing image aircraft target detection method based on region importance-aware attention provided in an embodiment of the present invention;
[0041] Figure 2 It is a schematic structural diagram of a region importance-aware attention module provided in an embodiment of the present invention;
[0042] Figure 3 It is a schematic structural diagram of an electronic device provided in an embodiment of the present invention. Specific embodiments
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0044] As Figure 1 shown, in the first aspect, the embodiments of the present invention disclose a remote sensing image aircraft target detection method based on region importance-aware attention, including the following steps:
[0045] S1: Perform image preprocessing on the remote sensing image to obtain the preprocessed remote sensing image;
[0046] Further, the image preprocessing in S1 includes normalization, contrast enhancement, brightness adjustment, denoising, cropping or scaling.
[0047] It can be understood that: before extracting the initial feature map, the input remote sensing image is first preprocessed to ensure that the image quality is suitable for subsequent feature extraction. First, the pixel values of the input remote sensing image are normalized to the range of [0, 1] to eliminate the influence of factors such as uneven illumination on the image quality. Then, according to the characteristics of the remote sensing image, contrast enhancement, brightness adjustment or denoising processing can be adopted to improve the key information in the image and enhance the effect of feature extraction. Finally, the image will be cropped or scaled to a size suitable for the input size of the feature extraction network so as to enter the subsequent processing stage.
[0048] S2: Input the preprocessed remote sensing image into the feature extraction network to obtain a multi-scale fusion feature map;
[0049] Further, S2 specifically includes:
[0050] S21: Input the preprocessed remote sensing image into the pre-trained improved ResNet-50 network;
[0051] S22: Extract the feature maps output at different stages of the pre-trained improved ResNet-50 network, and perform upsampling on the high-resolution feature maps and then splice or weighted fuse them with the low-resolution feature maps to obtain the multi-scale fusion feature map.
[0052] Further, the fully connected layer is removed from the pre-trained improved ResNet-50 network, and the second stage, the third stage, the fourth stage, and the fifth stage are used as the output layers for feature extraction.
[0053] It can be understood that: In the feature extraction stage: First, a suitable deep learning network needs to be selected. Usually, a convolutional neural network (CNN) pre-trained on a large-scale dataset (such as ImageNet) is selected as the feature extraction network, such as ResNet, VGG, or DenseNet, etc.
[0054] To ensure that the extracted features contain rich semantic information, the present invention selects ResNet-50 as the feature extraction network, and the second stage, the third stage, the fourth stage, and the fifth stage of ResNet-50 are used as the output layers for feature extraction. Among them, the second stage (conv2_x) can extract relatively simple texture and edge information; the third stage (conv3_) can extract higher-level semantic information, but still retains certain spatial details, which helps to identify aircraft targets in remote sensing images; the fourth stage and the fifth stage (conv4_x and conv5_x) can extract more complex semantic information and abstract features.
[0055] That is, to adapt to the object detection task, the present invention removes the fully connected layer in ResNet-50, retains the feature maps output by the convolutional layers of multiple stages (the second stage, the third stage, the fourth stage, and the fifth stage), and uses them as the basic features for subsequent processing.
[0056] Multi-scale Feature Fusion: Considering the different scales of aircraft targets in remote sensing images, a multi-scale feature fusion strategy is adopted to enhance the model's detection ability for targets of different scales. In the present invention, by extracting feature maps from multiple different levels of ResNet-50 (such as conv2_x, conv3_x, conv4_x, and conv5_x), multi-level feature information from low-level to high-level can be captured. On this basis, the high-resolution feature maps are upsampled and concatenated or weighted-fused with the low-resolution feature maps, thereby obtaining multi-scale feature maps containing different scale information, ensuring that the model can process aircraft targets of various sizes in the image.
[0057] Specifically, in the present invention, the feature maps output by the second stage (conv2_x) are upsampled and concatenated or weighted-fused with the feature maps output by the third stage (conv3_x);
[0058] The feature maps output by the second stage (conv2_x) are upsampled and concatenated or weighted-fused with the feature maps output by the fourth stage (conv4_x);
[0059] The feature maps output by the second stage (conv2_x) are upsampled and concatenated or weighted-fused with the feature maps output by the fifth stage (conv5_x);
[0060] The feature maps output by the third stage (conv3_x) are upsampled and concatenated or weighted-fused with the feature maps output by the fourth stage (conv4_x);
[0061] The feature maps output by the third stage (conv3_x) are upsampled and concatenated or weighted-fused with the feature maps output by the fifth stage (conv5_x);
[0062] S3: Preprocess the multi-scale fusion feature maps to obtain the initial feature map φ o ;
[0063] Furthermore, the preprocessing in S3 includes adjusting the size of the feature maps and normalizing the feature maps.
[0064] It can be understood that after completing the multi-scale feature fusion in the present invention, the size of the multi-scale fusion feature maps is adjusted to ensure that it meets the input requirements of the subsequent network. In the present invention, through upsampling or downsampling operations, the multi-scale fusion feature maps are appropriately adjusted to ensure that their sizes and resolutions are adapted. Then, these multi-scale fusion feature maps are normalized to stabilize the training process of the network and ensure the balance of the information of each channel of the multi-scale fusion feature maps. Finally, these feature maps that have undergone multi-scale fusion and normalization processing are defined as the initial feature map φ o, providing basic features for subsequent regional importance perception and object detection tasks.
[0065] S4: Input the initial feature map φ o into the regional importance perception attention module to obtain an enhanced feature map φ';
[0066] Furthermore, as Figure 2 shown, the regional importance perception attention module includes a 1*1 convolutional layer, a first matrix multiplication unit, an exponentiation operation unit, a global average pooling layer, a second matrix multiplication unit, an inverse operation unit, a first 3*3 convolutional layer, a second 3*3 convolutional layer, a first activation function layer, a bilinear transformation layer, a second activation function layer, and a third matrix multiplication unit;
[0067] The input end of the 1*1 convolutional layer is used to input the initial feature map;
[0068] The output end of the 1*1 convolutional layer is respectively connected to the input end of the first matrix multiplication unit and the input end of the exponentiation operation unit;
[0069] The output end of the exponentiation operation unit is respectively connected to the input end of the first matrix multiplication unit and the input end of the global average pooling layer;
[0070] The output end of the first matrix multiplication unit is connected to the input end of the global average pooling layer;
[0071] The output end of the global average pooling layer is respectively connected to the input end of the second matrix multiplication unit and the input end of the inverse operation unit;
[0072] The output end of the inverse operation unit is connected to the input end of the second matrix multiplication unit;
[0073] The output end of the second matrix multiplication unit is sequentially connected to the input end of the third matrix multiplication unit through the first 3*3 convolutional layer, the second 3*3 convolutional layer, the first activation function layer, and the bilinear transformation layer;
[0074] The input end of the second activation function layer is used to input the first channel feature map in the initial feature map;
[0075] The output end of the second activation function layer is connected to the input end of the third matrix multiplication unit;
[0076] The input end of the third matrix multiplication unit is also used to input the initial feature map.
[0077] It can be understood that:
[0078] For the initialization and feature mapping of the gating mechanism:
[0079] The obtained initial feature map φ o , different from the existing gate units that apply additional networks, for simplicity, the present invention selects the first channel feature map φ[0] in the initial feature map φ o as the gate, that is, Sigmoid(φ[0]). In addition, φ o Through a 1×1 convolution operation, the input initial feature map φ o is adjusted in the channel dimension to obtain a new feature
[0080] feature The importance value within the surrounding R is instantiated by applying exponentiation operations, global average pooling, and inverse operations (adjusting the stride and compression convolution in these operations to reduce calculations and expand the receptive field): The importance value within the surrounding R:
[0081]
[0082] Wherein, represents the weight or importance measure at the feature location. R represents the set of regions of the neighborhood. R k is a specific region in R, k represents the index in the region R, and is used to traverse all regions within the neighborhood centered on the pixel x. i and j are the pixel indices within R k and are used to traverse all pixels within this region. and respectively represent the feature values of pixels i and j within the region R k . W k is the learnable weight associated with the region R k and is used to adjust the importance of this region. This formula is used to calculate the local importance of the feature at each region within its neighborhood, and integrates this local importance information through weighted summation. In the context of aircraft target detection in remote sensing images, this measure can help the model identify and emphasize the features related to aircraft targets in the image, thereby improving the detection accuracy.
[0083] Next, two 3×3 convolutional layers, a sigmoid activation function, and a bilinear transformation layer are used for activation and rescaling to obtain the region importance-aware attention A:
[0084]
[0085] Among them, Bilinear represents the bilinear transformation layer, and sigmoid represents the Sigmoid activation function. This step ensures that the output of the attention mechanism not only has non-linear characteristics but also can adapt to different feature scales.
[0086] Apply region importance-aware attention to the feature φ[0] and φ o , to obtain the enhanced feature map φ':
[0087]
[0088] Among them, represents the matrix multiplication operation. This process not only enhances the important features related to object detection but also suppresses the unimportant background information. The enhanced feature map φ' obtained through this operation enables the model to focus more on the aircraft targets in the remote sensing image, thereby improving the accuracy and robustness of object detection.
[0089] S5: Input the enhanced feature map into the object detection network to obtain the aircraft object detection result.
[0090] Furthermore, S5 specifically includes: performing object bounding box regression and classification on the enhanced feature map through the detection head of the object detection network, then removing redundant boxes by combining non-maximum suppression, and finally outputting the aircraft object detection result with category and confidence.
[0091] It can be understood that: after obtaining the enhanced feature map φ', in order to perform object detection, a dedicated object detection network is required to generate the output of object detection. This process generally includes the following steps: performing object bounding box regression and classification through the detection head (Detection Head), using the enhanced feature map φ' for object localization and classification, and removing redundant boxes through non-maximum suppression, and finally obtaining the detection result.
[0092] Specifically:
[0093] 1) Perform object bounding box regression and classification through the object detection head: After obtaining the enhanced feature map φ', it is necessary to perform object bounding box regression and object classification through the detection head of the object detection network. The detection head generally consists of a convolutional layer and a fully connected layer, and its main task is to generate candidate object bounding boxes and corresponding class probabilities for each position from the enhanced feature map φ'.
[0094] In the object bounding box regression task, the object detection network will generate the coordinates of the candidate boxes (such as the upper left and lower right coordinates of the bounding box). Usually, the object detection network will output 4 regression values (x min , y min , x max , y max ), where, (xmin , y min , x max , y max ) represents the coordinates of the candidate bounding box. The input enhanced feature map (where H and W are the height and width of the enhanced feature map, and C is the number of channels) through a series of convolutional operations, the object detection network will output 4 channels as the result of the target bounding box regression:
[0095]
[0096] The classification task is to assign a class label to each candidate bounding box. The number of classes is N, and each candidate bounding box will output an N-dimensional probability distribution through a fully connected layer, and use the Softmax function to convert it into class probabilities:
[0097]
[0098] Among them, is the class score output by the object detection network, and the Softmax function can convert the output into probabilities so that the sum of the probabilities of all classes is 1.
[0099] 2) Generate object detection results: Based on the regression bounding box and classification probability, the object detection network will assign a class to each candidate bounding box and calculate a confidence score. Each candidate bounding box will have a confidence score, usually expressed as:
[0100]
[0101] Among them, is the confidence of the candidate bounding box, is the predicted bounding box, is the intersection over union (IoU) value with the ground truth target, P class is the class probability of this box.
[0102] 3) After obtaining multiple candidate bounding boxes, it is usually necessary to use non-maximum suppression (NMS) to remove redundant detection boxes. First, calculate the intersection over union (IoU) between each pair of candidate bounding boxes. For each pair of candidate bounding boxes and calculate their IoU values. Then remove the overlapping boxes. For candidate bounding boxes and if their IoU exceeds a certain set threshold (such as 0.5), then keep the box with the higher confidence and suppress (remove) the box with the lower confidence. The commonly used strategy is to sort by confidence from high to low, select the box with the highest confidence, and remove the boxes with a higher overlap with it. Finally, the detection boxes are output. After NMS processing, the finally retained are a set of non-redundant target boxes, and each box has a corresponding class and confidence score.
[0103] 4) Finally, output the object detection results. After regression, classification, and NMS, the network finally outputs the bounding boxes of the objects and their categories. Each bounding box will contain the following information:
[0104]
[0105] Among them, are the coordinates of the detection box, is the probability of the category to which the box belongs. The final output is the set of these bounding boxes and their categories, representing all the aircraft targets detected in the remote sensing image.
[0106] In a second aspect, an embodiment of the present invention provides an electronic device, as Figure 3 shown. The electronic device may include: a processor 301, a communication interface 302, a memory 303, and a communication bus 304. Among them, the processor 301, the communication interface 302, and the memory 303 complete mutual communication through the communication bus 304. The processor 301 can call the logical instructions in the memory 303 to execute a method for detecting aircraft targets in remote sensing images based on region importance-aware attention.
[0107] In addition, when the logical instructions in the above-mentioned memory 303 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0108] In a third aspect, the present invention further provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a method for detecting aircraft targets in remote sensing images based on region importance-aware attention provided by the above-mentioned various methods.
[0109] Fourthly, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a method for detecting aircraft targets in remote sensing images based on region importance-aware attention provided by the above-mentioned various methods.
[0110] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0111] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0112] The various embodiments in this specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0113] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein,
[0114] but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for detecting aircraft targets in remote sensing images based on regional importance-aware attention, characterized in that: The following steps are involved: S1: performing image preprocessing on the remote sensing image to obtain a preprocessed remote sensing image; S2: inputting the preprocessed remote sensing image into a feature extraction network to obtain a multi-scale fusion feature map; S3: preprocessing the multi-scale fusion feature map to obtain an initial feature map; S4: inputting the initial feature map into a region importance-aware attention module to obtain an enhanced feature map; S5: Input the enhanced feature map into the target detection network to obtain the aircraft target detection result.
2. The method for detecting aircraft targets in remote sensing images based on regional importance perception attention according to claim 1, characterized in that: Image preprocessing in S1 includes normalization, contrast enhancement, brightness adjustment, denoising, cropping or scaling.
3. The method for detecting aircraft targets in remote sensing images based on regional importance perception attention according to claim 1, characterized in that: S2 specifically includes: S21: input the preprocessed remote sensing image into the pre-trained improved ResNet-50 network; S22: extract the feature maps output by the pre-trained improved ResNet-50 network at different stages, and upsample the high-resolution feature maps and perform splicing or weighted fusion with the low-resolution feature maps to obtain the multi-scale fused feature maps.
4. The method for detecting aircraft targets in remote sensing images based on regional importance perception attention according to claim 3 is characterized in that: The pre-trained improved ResNet-50 network removes the fully connected layer and uses the second stage, the third stage, the fourth stage and the fifth stage as the output layer for feature extraction.
5. The method for detecting aircraft targets in remote sensing images based on regional importance perception attention according to claim 1, characterized in that: The preprocessing in S3 includes feature map resizing and feature map normalization.
6. The method for detecting aircraft targets in remote sensing images based on regional importance-aware attention according to claim 1, characterized in that: The regional importance perception attention module includes a 1*1 convolution layer, a first matrix multiplication unit, an exponential operation unit, a global average pooling layer, a second matrix multiplication unit, an inverse operation unit, a first 3*3 convolution layer, a second 3*3 convolution layer, a first activation function layer, a bilinear transformation layer, a second activation function layer and a third matrix multiplication unit; The input end of the 1*1 convolutional layer is used to input the initial feature map; The output end of the 1*1 convolutional layer is connected to the input end of the first matrix multiplication unit and the input end of the exponential operation unit respectively; The output end of the exponential operation unit is connected to the input end of the first matrix multiplication unit and the input end of the global average pooling layer respectively; An output end of the first matrix multiplication unit is connected to an input end of the global average pooling layer; The output end of the global average pooling layer is connected to the input end of the second matrix multiplication unit and the input end of the inverse operation unit respectively; The output end of the inverse operation unit is connected to the input end of the second matrix multiplication unit; The output end of the second matrix multiplication unit is connected to the input end of the third matrix multiplication unit through the first 3*3 convolution layer, the second 3*3 convolution layer, the first activation function layer, and the bilinear transformation layer in sequence; The input end of the second activation function layer is used to input the first channel feature map in the initial feature map; The output end of the second activation function layer is connected to the input end of the third matrix multiplication unit; The input end of the third matrix multiplication unit is also used to input the initial feature map.
7. The method for detecting aircraft targets in remote sensing images based on regional importance-aware attention according to claim 1, characterized in that: S5 specifically includes: performing target frame regression and classification on the enhanced feature map through the detection head of the target detection network, and then removing redundant frames in combination with non-maximum suppression, and finally outputting an aircraft target detection result with category and confidence.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the remote sensing image aircraft target detection method based on regional importance perception attention as described in any one of claims 1 to 7 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for detecting aircraft targets in remote sensing images based on regional importance-aware attention is implemented as described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for detecting aircraft targets in remote sensing images based on regional importance-aware attention is implemented as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Face super-resolution method based on pre-training generative model
CN113379606A
Remote sensing image directional target detection method based on multi-feature aggregation and interaction
CN114926747A
SAR image change detection method based on multidirectional attention representation enhancement
CN115376012A
Explicit contour guidance and spatial variation context enhancement remote sensing target detection method
CN115830449A
Anchor-frame-free remote sensing image rotating target detection method under attention mechanism
CN118379617A