Small target detection method in satellite remote sensing images based on high-resolution feature self-attention
By using a detection method of high-resolution feature self-attention mechanism in satellite remote sensing images, the detection problem of extremely small targets in complex backgrounds is solved, and small target detection with high accuracy is achieved.
Patent Information
- Application Number
- CN202311108808.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-31
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-08-31
AI Technical Summary
In satellite remote sensing scenarios, it is difficult for the prior art to effectively detect extremely small targets because these targets are only low-resolution targets of a few pixel sizes and are easily confused with complex backgrounds, resulting in increased detection difficulty.
A small object detection method for satellite remote sensing images based on self-attention based on high-resolution features is adopted, which includes a backbone network, a high-resolution feature fusion network, a main detection structure, an auxiliary detection structure and a prediction head. Through high-resolution feature fusion network and self-attention mechanism, global information and local features are captured to improve the feature effectiveness of small goals.
This method can more effectively retain the semantic information of small targets, reduce the interference of complex backgrounds on the target, and achieve high-accuracy small target detection.
Smart Images

Figure CN117036980B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of satellite remote sensing image target detection, and in particular to a satellite remote sensing image small target detection method based on high-resolution feature self-attention. Background Art
[0002] At present, some research has been carried out on target detection technology for satellite remote sensing scenes, which can be widely used in tasks such as military target recognition, traffic management and environmental monitoring. However, achieving accurate target detection in remote sensing scenes is a challenging problem: on the one hand, in many remote sensing images, the target is only a low-resolution target of a few pixels in size; on the other hand, smaller targets will be confused with complex backgrounds, which will interfere with detection and further increase the difficulty of small target detection.
[0003] As a mainstream method in the field of image processing, deep learning technology plays an important role in the field of remote sensing image processing. In recent years, object detection methods for general scenes have achieved great breakthroughs, such as Faster R-CNN, SSD, YOLO and other convolutional neural networks have been proposed one after another. However, due to the characteristics of large data volume, rich information and complex images of remote sensing images, it is difficult to obtain satisfactory results by directly applying these algorithms to remote sensing target detection tasks, especially when facing extremely small targets. It will bring great challenges. Although the CNN-type methods that have been widely studied in general scene detection methods have good performance in spatial position representation, it is difficult to model global space and context information.
[0004] In the paper "ECAP-YOLO: Efficient Channel Attention Pyramid YOLO for Small Object Detection in Aerial Image", Kim et al. proposed an efficient channel attention pyramid module based on YOLOv5. By weighting the channel dimension of the feature map, the feature saliency of small targets in the channel dimension is improved, but it does not have the ability to obtain the global correlation of the feature map. In the paper "PAG-YOLO: A Portable Attention-Guided YOLO Network for Small Ship Detection", Hu et al. proposed an attention mechanism in the channel and spatial dimensions, which can guide the model to optimize the feature information representation of small targets, and combined with YOLOv5 to achieve fast and accurate target recognition. Similar to the paper by Kim et al., it also uses an attention mechanism to improve feature effectiveness, but has the same problem. Zhang et al. proposed a multi-stage feature enhancement pyramid network in the paper "Multi-Stage Feature Enhancement Pyramid Network for Detecting Objects in Optical Remote Sensing Images". In feature fusion, multi-level features are used to enhance feature visibility and solve the problem of fuzzy target recognition. However, when facing extremely small targets, the lack of detailed information in the image leads to poor recognition effect.
[0005] With the continuous advancement of remote sensing imaging methods, small target detection in satellite remote sensing images has developed into one of the important research directions in the field of computer vision. However, in the process of ultra-long-distance satellite observation, the image background information is highly complex, and the target is only a low-resolution target of a few pixels in size. To address these problems, the existing methods using attention mechanisms can only extract local image information and lack global information processing methods; while image details are very important for small target detection, existing methods cannot achieve high-resolution feature map forward transmission. Summary of the invention
[0006] The present invention aims to solve the technical problems in the prior art and provides a small target detection method for satellite remote sensing images based on high-resolution feature self-attention.
[0007] In order to solve the above technical problems, the technical solutions of the present invention are as follows:
[0008] A method for detecting small targets in satellite remote sensing images based on high-resolution feature self-attention, the method being applicable to a system comprising: a backbone network, a high-resolution feature fusion network, a main detection structure, an auxiliary detection structure and a prediction head;
[0009] The backbone network is used to perform convolution operations on the input image, downsample the image, and reduce the spatial resolution of the image;
[0010] The high-resolution feature network includes: a basic module, a fusion module and a channel attention space pyramid module;
[0011] The main detection structure is composed of a C3 module, which includes three convolution modules and a bottleneck structure; the convolution module includes batch normalization, Relu activation function and convolution operation; the bottleneck structure is composed of three convolution modules connected in series, where the number of channels output by the second convolution module is half of the input; in the C3 module, one branch passes through the first convolution module and the bottleneck module, and the other passes through the second convolution module, and the results of the two branches are spliced and passed through the third convolution module;
[0012] The auxiliary detection structure includes: a hybrid attention module and a self-attention module;
[0013] The prediction head is used to obtain a prediction result after the input feature map is processed by non-maximum suppression;
[0014] The method comprises the following steps:
[0015] (1) Obtain a data set and perform Mosaic data enhancement on the data set; send the enhanced data into the network for training;
[0016] (2) Constructing a network, including a backbone network, a high-resolution feature fusion network, a main detection structure, an auxiliary detection structure, and a prediction head;
[0017] (3) The optimization algorithm uses the stochastic gradient descent algorithm SGD as the optimizer, with 16 images as a training batch, the initial learning rate of the model is 1e-2, the weight decay parameter is 5e-4, and the momentum is 0.937, and the training is 300 epochs; in the initial stage of model training, 3 epochs are used for warm-up training;
[0018] (4) After training the model, predict the image and get the result.
[0019] In the above technical solution, the high-resolution feature network is divided into three stages:
[0020] In the first stage, the number of channels of the network feature map is 256;
[0021] In the second phase, the number of network channels present is 256 and 512;
[0022] In the third phase, the number of network channels present is 256, 512, and 1024;
[0023] The input feature map with 256 channels in the first stage and the input feature map with 512 channels in the second stage are concatenated with the output feature map in the third stage in the channel dimension using residual connection.
[0024] In the above technical solution, in the high-resolution feature fusion network:
[0025] The input feature map of the basic module is The input feature map will undergo two convolution operations with batch normalization and ReLU activation function, and then be added to the input feature map according to the corresponding bit elements to finally obtain the output feature map The calculation process is as follows:
[0026] X out =X in +f2(f1(X in ))
[0027] Where f(·) represents a convolutional layer with batch normalization and ReLU activation function;
[0028] The feature map input by the fusion module is stacked and then convolved with batch normalization and ReLU activation function.
[0029] The channel attention space pyramid module includes: a channel attention mechanism module and a feature pyramid module;
[0030] In the channel attention mechanism module, first, the input feature map is After that, the feature map is obtained by average pooling in the channel dimension. Then the feature map is obtained by Sigmoid activation function. Finally, the input I and the output P are multiplied element by element to get the output of the ECA module.
[0031] In the feature pyramid module, the output of ECA will be used as the input of the SPP module. After that, it will pass through three maximum average pooling operations with different convolution kernels, and the sizes of the convolution kernels are 5*5, 9*9, and 13*13 respectively; then, the three pooled feature maps will be spliced with the input Z of the SPP part, and finally pass through a convolution with batch normalization and Relu activation function, and the output feature map is The calculation process is as follows:
[0032] T=f(cat(Z5,Z9,Z 13 , Z)
[0033] Among them, Z5, Z9, Z 13They represent the feature maps after pooling operations with three different sizes of convolution kernels, cat represents the concatenation operation, and f(·) represents the convolution layer with batch normalization and ReLU activation function.
[0034] In the above technical solution, in the hybrid attention module in the auxiliary detection structure:
[0035] The input of the self-channel attention is After that, the feature map is obtained by average pooling in the channel dimension. Then the feature map is obtained by Sigmoid activation function. Finally, the input I and the output P are multiplied according to the corresponding elements to obtain the output of the channel attention
[0036] In the above technical solution, the working process of the hybrid attention module in the auxiliary detection structure is:
[0037] First, average pooling and maximum pooling operations are applied along the channel axis, and the generated feature map is and and stitch them together to produce a valid descriptive feature;
[0038] Then, the concatenated feature map is combined with the input feature map Multiply the corresponding elements;
[0039] Finally, the output of the spatial attention module is
[0040] In the above technical solution, the output process of the spatial attention module is expressed as:
[0041]
[0042] Among them, cat represents the operation of splicing the feature map after the pooling operation along the channel. represents a convolutional layer and a Relu activation function, and ⊙ represents the multiplication of corresponding elements.
[0043] In the above technical solution, the working process of the self-attention module in the auxiliary detection structure is:
[0044] First, input size feature map
[0045] Then, the feature map is windowed, and the input tensor is divided into groups of n pixels along the length and width, with each window size being n*n.
[0046] Afterwards, the three-dimensional tensor in the window is extended from H*W*C to *1*C;
[0047] Then, the linear layer expands the channel dimension to 3*C, and divides the matrix into matrix Q, matrix K and matrix V along the channel dimension; the calculation method of the self-attention module is given by the following formula:
[0048]
[0049] Among them, d k is the dimension of matrix K;
[0050] Finally, the results are rearranged into H*W*C as the output of the self-attention mechanism
[0051] In the above technical solution, during the working process of the auxiliary detection structure, the following steps are also included:
[0052] The output T of the self-attention module and the output M of the mixed attention module are added according to the corresponding elements to form the output of the mixed attention based on the self-attention mechanism, which is expressed as:
[0053] in, Indicates the addition of corresponding elements.
[0054] In the above technical solution, the output of the main detection structure and the output of the auxiliary detection structure are spliced along the channel dimension and then input into the prediction head.
[0055] The present invention has the following beneficial effects:
[0056] The present invention provides a method for detecting small targets in satellite remote sensing images based on high-resolution feature self-attention.
[0057] The present invention proposes a high-resolution feature fusion network, which can retain more semantic information when facing small-scale targets by maintaining a high-resolution feature extraction structure.
[0058] This paper proposes a STMA attention mechanism, which can capture global information and local features, improve the feature effectiveness of small targets, and thus suppress the impact of complex background on the target.
[0059] The algorithm of the present invention can be used not only in satellite remote sensing image target detection, but also in remote sensing video target tracking, taking into account both aerial remote sensing image target detection and tracking tasks.
[0060] The method of the present invention uses a high-resolution feature fusion network to maintain the features of small targets in the forward transmission, so that the global information encoding ability of self-attention and the ability of spatial attention mechanism to enhance local features complement each other, and then combines the detection information obtained by the convolution layer with the attention information obtained by the mixed attention through a parallel detection structure, thereby reducing the interference of complex background information on small targets. Compared with the existing remote sensing small target detection method, the present invention can better eliminate the interference background in the image and achieve high-accuracy small target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0062] Figure 1 The schematic diagram of the structure of the system applicable to the satellite remote sensing image small target detection method based on high-resolution feature self-attention of the present invention.
[0063] Figure 2 Schematic diagram of the structure of the high-resolution feature fusion network.
[0064] Figure 3 Schematic diagram of the structure of the channel attention space pyramid module.
[0065] Figure 4 Schematic diagram of the structure of the hybrid attention module based on the self-attention mechanism.
[0066] Figure 5 The present invention is a schematic diagram of the steps of the satellite remote sensing image small target detection method based on high-resolution feature self-attention. DETAILED DESCRIPTION
[0067] The present invention is described in detail below with reference to the accompanying drawings.
[0068] The structure of the system applicable to the satellite remote sensing image small target detection method based on high-resolution feature self-attention of the present invention is as follows: Figure 1 As shown in the figure, it mainly includes five parts: backbone network, high-resolution feature fusion network, main detection structure, auxiliary detection structure and prediction head.
[0069] Introduce the module structure of the network:
[0070] (1) Backbone network:
[0071] In the backbone network, the input image first undergoes three convolution operations to downsample the image to 1 / 8 of its original size with 256 channels, thereby reducing the spatial resolution of the image and reducing the computational pressure of the network.
[0072] (2) High-resolution feature fusion network:
[0073] The structure diagram of the high-resolution feature network is as follows Figure 2 As shown. The entire network is divided into three stages. In the first stage, the number of channels of the network feature map is 256. In the second stage, the number of network channels is 256 and 512. In the third stage, the number of network channels is 256, 512 and 1024. Among them, the input feature map with 256 channels in the first stage and the input feature map with 512 channels in the second stage are spliced with the output feature map in the third stage in the channel dimension using residual connection.
[0074] The high-resolution feature fusion network includes three modules, namely the basic module, the fusion module and the channel attention spatial pyramid module (E-SPP).
[0075] (a) Basic module: The input feature map of this module is The input feature map will undergo two convolution operations with batch normalization and ReLU activation function, and then be added to the input feature map according to the corresponding bit elements to finally obtain the output feature map The calculation process is as follows:
[0076] X out =X in +f2(f1(X in )) (1)
[0077] Here, f(·) represents a convolutional layer with batch normalization and ReLU activation function.
[0078] (b) Fusion module: In the fusion module, the input feature maps are stacked and then go through a convolution operation with batch normalization and ReLU activation function.
[0079] (c) Channel Attention Space Pyramid Module: The structure of this module is shown in the figure below. Figure 3 As shown in Figure 2, this module consists of two parts: channel attention mechanism (ECA) and feature pyramid (SPP). In ECA, first, the input feature map is After that, the feature map is obtained by average pooling in the channel dimension. Then the feature map is obtained by Sigmoid activation function. Finally, the input I and the output P are multiplied element by element to get the output of the ECA module. In the SPP module, the output of ECA will be used as the input of the SPP module. After that, it will go through three maximum average pooling operations with different convolution kernels, the sizes of which are 5*5, 9*9, and 13*13 respectively. Then, the three pooled feature maps will be concatenated with the input Z of the SPP part, and finally go through a convolution with batch normalization and Relu activation function to get the output feature map as The calculation process is as follows:
[0080] T=f(cat(Z5,Z9,Z 13 ,Z) (2)
[0081] Among them, Z5, Z9, Z 13 They represent the feature maps after pooling operations with three different sizes of convolution kernels, cat represents the concatenation operation, and f(·) represents the convolution layer with batch normalization and ReLU activation function.
[0082] (3) Main detection structure: The main detection structure consists of a C3 module, which includes three convolution modules and a bottleneck structure. The convolution module contains batch normalization, Relu activation function and convolution operation. The bottleneck structure consists of three convolution modules connected in series, where the number of channels output by the second convolution module is half of the input. In the C3 module, one branch passes through the convolution module and the bottleneck module, and the other passes through the convolution module. The results of the two branches are concatenated and passed through the last convolution module.
[0083] (4) Auxiliary detection structure: The auxiliary detection structure consists of three hybrid attentions based on the self-attention mechanism (STMA), such as Figure 4 As shown in Figure 3, STMA consists of two parts: a hybrid attention module and a self-attention module.
[0084] (a) Hybrid attention module: The hybrid attention module consists of two parts: channel attention and spatial attention. The input of the self-channel attention is After that, the feature map is obtained by average pooling in the channel dimension. Then the feature map is obtained by Sigmoid activation function. Finally, the input I and the output P are multiplied according to the corresponding elements to obtain the output of the channel attention When calculating spatial attention, we first apply average pooling and maximum pooling operations along the channel axis, and the resulting feature map is and And concatenate them to generate a valid description feature, and then compare the concatenated feature map with the input feature map Multiply the corresponding elements, and the output of the final spatial attention module is This process can be expressed by the following formula:
[0085]
[0086] Among them, cat represents the operation of splicing the feature map after the pooling operation along the channel. represents a convolutional layer and a Relu activation function, and ⊙ represents the multiplication of corresponding elements.
[0087] (b) Self-attention module: First input the size feature map Then the feature map is windowed, and the input tensor is divided into groups of n pixels along the length and width, and the size of each window is n*n. Finally, the three-dimensional tensor in the window is extended from H*W*C to (HW)*1*C. After that, the linear layer expands the channel dimension to 3*C, and the matrix is divided into matrix Q, matrix K and matrix V along the channel dimension. The calculation method of the self-attention module is given by the following formula:
[0088]
[0089] where d k is the dimension of matrix K.
[0090] Rearrange the results into H*W*C as the output of the self-attention mechanism
[0091] Finally, the output T of the self-attention module and the output M of the mixed attention module are added according to the corresponding elements to form the output of STMA, which can be expressed as:
[0092]
[0093] in, Indicates the addition of corresponding elements
[0094] (5) Prediction head: After the output of the main detection structure and the output of the auxiliary detection structure are spliced along the channel dimension, they are input into the prediction head, and the input feature map is processed by non-maximum suppression to obtain the prediction result.
[0095] The present invention provides a method for detecting small targets in satellite remote sensing images based on high-resolution feature self-attention. Figure 5 As shown, the complete running steps are as follows:
[0096] (1) Obtain a data set and perform Mosaic data enhancement on the data set; send the enhanced data into the network for training;
[0097] (2) The constructed network consists of five parts: backbone network, high-resolution feature fusion network, main detection structure, auxiliary detection structure and prediction head.
[0098] (3) The optimization algorithm uses the stochastic gradient descent algorithm SGD as the optimizer, with 16 images as a training batch, the initial learning rate of the model is 1e-2, the weight decay parameter is 5e-4, and the momentum is 0.937, and the training is 300 epochs. In the initial stage of model training, 3 epochs are used for warm-up training;
[0099] (4) After training the model, the image is predicted to obtain the result. The present invention is compared with the YOLOv5 model, and the test results of the two networks are compared, as shown in Table 1. The precision index means how many of the targets predicted by the model are actually the targets. The recall index indicates how many of the real targets are successfully predicted by the model. The average accuracy can balance the two indicators of precision and recall. The recall rate is used as the horizontal coordinate and the precision rate is used as the vertical coordinate to calculate the area under the curve (PR-curve) surrounded by the two parameters.
[0100] Table 1 Comparison of test results of two networks
[0101] Model Accuracy Recall Average precision YOLOv5 78.1 70.0 70.9 Method of the present invention 84.6 72.0 77.3
[0102] The present invention proposes a high-resolution feature fusion network, which can retain more semantic information when facing small-scale targets by maintaining a high-resolution feature extraction structure.
[0103] This paper proposes a STMA attention mechanism, which can capture global information and local features, improve the feature effectiveness of small targets, and thus suppress the impact of complex background on the target.
[0104] The algorithm of the present invention can be used not only in satellite remote sensing image target detection, but also in remote sensing video target tracking, taking into account both aerial remote sensing image target detection and tracking tasks.
[0105] The method of the present invention uses a high-resolution feature fusion network to maintain the features of small targets in the forward transmission, so that the global information encoding ability of self-attention and the ability of spatial attention mechanism to enhance local features complement each other, and then combines the detection information obtained by the convolution layer with the attention information obtained by the mixed attention through a parallel detection structure, thereby reducing the interference of complex background information on small targets. Compared with the existing remote sensing small target detection method, the present invention can better eliminate the interference background in the image and achieve high-accuracy small target detection.
[0106] Obviously, the above embodiments are merely examples for the purpose of clear explanation, and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived therefrom are still within the scope of protection of the invention.
Claims
1. A method for detecting small targets in satellite remote sensing images based on high-resolution feature self-attention, characterized in that: The system to which the method is applicable includes: a backbone network, a high-resolution feature fusion network, a main detection structure, an auxiliary detection structure and a prediction head; The backbone network is used to perform convolution operations on the input image, downsample the image, and reduce the spatial resolution of the image; The high-resolution feature fusion network includes: a basic module, a fusion module and a channel attention space pyramid module; The main detection structure is composed of a C3 module, which includes three convolution modules and a bottleneck structure; the convolution module includes batch normalization, Relu activation function and convolution operation; the bottleneck structure is composed of three convolution modules connected in series, where the number of channels output by the second convolution module is half of the input; in the C3 module, one branch passes through the first convolution module and the bottleneck module, and the other passes through the second convolution module, and the results of the two branches are spliced and passed through the third convolution module; The auxiliary detection structure includes: a hybrid attention module and a self-attention module; The prediction head is used to obtain a prediction result after the input feature map is processed by non-maximum suppression; The method comprises the following steps: (1) Obtain a data set and perform Mosaic data enhancement on the data set; send the enhanced data into the network for training; (2) Constructing a network, including a backbone network, a high-resolution feature fusion network, a main detection structure, an auxiliary detection structure, and a prediction head; (3) The optimization algorithm uses the stochastic gradient descent algorithm SGD as the optimizer, with 16 images as a training batch, the initial learning rate of the model is 1e-2, the weight decay parameter is 5e-4, and the momentum is 0.937, and the training is 300 epochs; in the initial stage of model training, 3 epochs are used for warm-up training; (4) After training the model, predict the image and get the result.
2. The method for detecting small targets in satellite remote sensing images based on high-resolution feature self-attention according to claim 1, characterized in that: The high-resolution feature fusion network is divided into three stages: In the first stage, the number of channels of the network feature map is 256; In the second phase, the number of network channels present is 256 and 512; In the third phase, the number of network channels present is 256, 512, and 1024; The input feature map with 256 channels in the first stage and the input feature map with 512 channels in the second stage are concatenated with the output feature map in the third stage in the channel dimension using residual connection.
3. The method for detecting small targets in satellite remote sensing images based on high-resolution feature self-attention according to claim 1, characterized in that: In the high-resolution feature fusion network: The input feature map of the basic module is The input feature map will undergo two convolution operations with batch normalization and ReLU activation function, and then be added to the input feature map according to the corresponding bit elements to finally obtain the output feature map The calculation process is as follows: X out =X in +f2(f1(X in )) Where f(·) represents a convolutional layer with batch normalization and ReLU activation function; The feature map input by the fusion module is first subjected to a stacking operation, and then to a convolution operation with batch normalization and ReLU activation function; The channel attention space pyramid module includes: a channel attention mechanism module and a feature pyramid module; In the channel attention mechanism module, first, the input feature map is After that, the feature map is obtained by average pooling in the channel dimension. Then the feature map is obtained by Sigmoid activation function. Finally, the input I and the output P are multiplied element by element to get the output of the ECA module. In the feature pyramid module, the output of ECA will be used as the input of the SPP module. After that, it will pass through three maximum average pooling operations with different convolution kernels, and the sizes of the convolution kernels are 5*5, 9*9, and 13*13 respectively; then, the three pooled feature maps will be spliced with the input Z of the SPP part, and finally pass through a convolution with batch normalization and Relu activation function, and the output feature map is The calculation process is as follows: T=f(cat(Z5,Z9,Z 13 ,Z) Among them, Z5, Z9, Z 13 They represent the feature maps after pooling operations with three different sizes of convolution kernels, cat represents the concatenation operation, and f(·) represents the convolution layer with batch normalization and ReLU activation function.
4. The method for detecting small targets in satellite remote sensing images based on high-resolution feature self-attention according to claim 1, characterized in that: In the hybrid attention module in the auxiliary detection structure: The input of the self-channel attention is After that, the feature map is obtained by average pooling in the channel dimension. Then the feature map is obtained by Sigmoid activation function. Finally, the input I and the output P are multiplied according to the corresponding elements to obtain the output of the channel attention 5. The method for detecting small targets in satellite remote sensing images based on high-resolution feature self-attention according to claim 4, characterized in that: The working process of the hybrid attention module in the auxiliary detection structure is: First, average pooling and maximum pooling operations are applied along the channel axis, and the generated feature map is and and stitch them together to produce a valid descriptive feature; Then, the concatenated feature map is combined with the input feature map Multiply the corresponding elements; Finally, the output of the spatial attention module is 6. The method for detecting small targets in satellite remote sensing images based on high-resolution feature self-attention according to claim 5, characterized in that: The output process of the spatial attention module is expressed as: Among them, cat represents the operation of splicing the feature map after the pooling operation along the channel. represents a convolutional layer and a Relu activation function, and ⊙ represents the multiplication of corresponding elements.
7. The method for detecting small targets in satellite remote sensing images based on high-resolution feature self-attention according to claim 1, characterized in that: The working process of the self-attention module in the auxiliary detection structure is: First, input size feature map Then, the feature map is windowed, and the input tensor is divided into groups of n pixels along the length and width, with each window size being n*n. Afterwards, the three-dimensional tensor in the window is extended from H*W*C to *1*C; Then, the linear layer expands the channel dimension to 3*C, and divides the matrix into matrix Q, matrix K and matrix V along the channel dimension; the calculation method of the self-attention module is given by the following formula: Among them, d k is the dimension of matrix K; Finally, the results are rearranged into H*W*C as the output of the self-attention mechanism 8. The method for detecting small targets in satellite remote sensing images based on high-resolution feature self-attention according to any one of claims 5 to 7, characterized in that: The working process of the auxiliary detection structure also includes the following steps: The output T of the self-attention module and the output M of the mixed attention module are added according to the corresponding elements to form the output of the mixed attention based on the self-attention mechanism, which is expressed as: in, Indicates the addition of corresponding elements.
9. The method for detecting small targets in satellite remote sensing images based on high-resolution feature self-attention according to claim 1, characterized in that: After the output of the main detection structure and the output of the auxiliary detection structure are spliced along the channel dimension, they are input into the prediction head.
Citation Information
Patent Citations
Remote sensing image target detection method based on multi-scale feature fusion and feature enhancement
CN114708511A
Method, device and equipment for detecting weak-intensity small-scale target in remote sensing image
CN115830470A