Remote sensing image small target detection method and device, equipment and storage medium

By improving the Backbone and Neck structure of the YOLOv8s network model, combining data enhancement and multiple attention mechanisms, the problem of low detection accuracy of remote sensing images is solved, and more efficient remote sensing small object detection is achieved.

CN120339782APending Publication Date: 2025-07-18SOUTH CENTRAL UNIVERSITY FOR NATIONALITIES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311584686.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-22
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the remote sensing image small object detection algorithm has poor detection effect and low detection accuracy.

Method used

The improved YOLOv8s network model is adopted, including the improved Backbone and Neck. The receptive field amplification module is set in Backbone, coordinate space attention is set in Neck, and the model is trained through the Mosaic data augmentation method, combining the two-layer routing attention, feature pyramid structure and adaptive non-maximum suppression algorithm for small object detection.

Benefits of technology

It improves the detection accuracy and robustness of small targets in remote sensing images, effectively solves the detection difficulties of dense small targets, and improves the detection effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339782A_ABST
    Figure CN120339782A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of object detection, and discloses a remote sensing image small target detection method, device and equipment and a storage medium, and the method comprises the steps: inputting a to-be-detected remote sensing image into an improved YOLOv8s network model for small target detection, and the improved YOLOv8s network model comprises an improved Backbone and an improved Neck, the improved Backbone is internally provided with a receptive field amplification module, and the improved Neck is internally provided with coordinate space attention; and obtaining detection information corresponding to the target small object in the to-be-detected remote sensing image according to the detection result. According to the invention, small target detection is carried out on the to-be-detected remote sensing image through the improved YOLOv8s network model, the detection information corresponding to the target small object is obtained, and the problems of poor detection effect and low detection precision of the target detection algorithm on the remote sensing small target in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and particularly to a method, device, equipment and storage medium for detecting small targets in remote sensing images. Background Art

[0002] Remote sensing images are important data sources for obtaining information on the Earth's surface through vehicles such as spacecraft or airplanes. These data have important application values in fields such as precision agriculture, urban planning, and disaster management. Since remote sensing images contain many small targets, their backgrounds are complex, their sizes are small, and they are densely arranged, resulting in the existing object detection algorithms being far less effective in detecting small targets in remote sensing than large and medium targets. Therefore, how to improve the detection accuracy of small targets in remote sensing has become an urgent problem to be solved.

[0003] The above content is only used to assist in understanding the technical solution of the present invention, and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of the present invention is to provide a method, device, equipment and storage medium for detecting small targets in remote sensing images, aiming to solve the technical problems of poor detection effect and low detection accuracy of the object detection algorithm for small targets in remote sensing in the prior art.

[0005] To achieve the above purpose, the present invention provides a method for detecting small targets in remote sensing images, and the method for detecting small targets in remote sensing images includes:

[0006] Input the remote sensing image to be detected into the improved YOLOv8s network model for small target detection. The improved YOLOv8s network model includes an improved Backbone and an improved Neck. A receptive field amplification module is set in the improved Backbone, and a coordinate spatial attention is set in the improved Neck;

[0007] Obtain the detection information corresponding to the target small object in the remote sensing image to be detected according to the detection result.

[0008] Optionally, before the step of inputting the remote sensing image to be detected into the improved YOLOv8s network model for small target detection, it further includes:

[0009] Perform data augmentation processing on the image data set by the Mosaic method to obtain an enhanced image data set;

[0010] Input the enhanced image data set into the initial YOLOv8s network model for model training to obtain the improved YOLOv8s network model.

[0011] Optionally, the step of inputting the remote sensing image to be detected into the improved YOLOv8s network model for small target detection includes:

[0012] Input the remote sensing image to be detected into the improved YOLOv8s network model;

[0013] Perform feature extraction processing on the remote sensing image to be detected through the double-layer routing attention set in the improved Backbone to obtain the remote sensing feature map corresponding to the remote sensing image to be detected;

[0014] The step of obtaining the detection information corresponding to the target small object in the remote sensing image to be detected according to the detection result includes:

[0015] Obtain the detection information corresponding to the target small object in the remote sensing image to be detected based on the remote sensing feature map to be detected.

[0016] Optionally, the step of performing feature extraction processing on the remote sensing image to be detected through the double-layer routing attention set in the improved Backbone to obtain the remote sensing feature map corresponding to the remote sensing image to be detected includes:

[0017] Obtain the global receptive field information of the remote sensing image to be detected through the double-layer routing attention set in the improved Backbone;

[0018] Obtain the remote sensing feature map corresponding to the remote sensing image to be detected based on the global receptive field information.

[0019] Optionally, the step of obtaining the detection information corresponding to the target small object in the remote sensing image to be detected based on the remote sensing feature map to be detected includes:

[0020] Obtain the coordinate attention weight and spatial attention weight corresponding to the remote sensing feature map to be detected through the improved feature pyramid structure;

[0021] Generate a target remote sensing feature map based on the coordinate attention weight, the spatial attention weight, and the remote sensing feature map to be detected, where the target remote sensing feature map is a feature map carrying coordinate space information;

[0022] Obtain the detection information corresponding to the target small object in the remote sensing image to be detected based on the target remote sensing feature map.

[0023] Optionally, the improved feature pyramid structure is constructed based on the coordinate space attention, and the coordinate space attention includes: coordinate attention and spatial attention; the step of obtaining the coordinate attention weight and spatial attention weight corresponding to the remote sensing feature map to be detected through the improved feature pyramid structure includes:

[0024] Perform one-dimensional average pooling on the to-be-detected remote sensing feature map through the coordinate attention to obtain the coordinate attention weight corresponding to the to-be-detected remote sensing feature map;

[0025] Perform channel dimension pooling on the to-be-detected remote sensing feature map through the spatial attention to obtain the spatial attention weight corresponding to the to-be-detected remote sensing feature map.

[0026] Optionally, the step of obtaining the detection information corresponding to the target small object in the to-be-detected remote sensing image based on the target remote sensing feature map includes:

[0027] Perform post-detection processing on the target remote sensing feature map through an adaptive non-maximum suppression algorithm to obtain the processed target remote sensing feature map;

[0028] Obtain the detection information corresponding to the target small object in the to-be-detected remote sensing image based on the processed target remote sensing feature map.

[0029] In addition, to achieve the above object, the present invention also proposes a remote sensing image small target detection device, and the device includes:

[0030] An object detection module, configured to input a to-be-detected remote sensing image into an improved YOLOv8s network model for small target detection. The improved YOLOv8s network model includes an improved Backbone and an improved Neck. A receptive field amplification module is arranged in the improved Backbone, and coordinate spatial attention is arranged in the improved Neck;

[0031] A detection information acquisition module, configured to obtain the detection information corresponding to the target small object in the to-be-detected remote sensing image according to the detection result.

[0032] In addition, to achieve the above object, the present invention also proposes a remote sensing image small target detection device, and the device includes: a memory, a processor, and a remote sensing image small target detection program stored on the memory and executable on the processor. The remote sensing image small target detection program is configured to implement the steps of the remote sensing image small target detection method as described above.

[0033] In addition, to achieve the above object, the present invention also proposes a storage medium, on which a remote sensing image small target detection program is stored. When the remote sensing image small target detection program is executed by a processor, the steps of the remote sensing image small target detection method as described above are implemented.

[0034] In the present invention, it is disclosed that the remote sensing image to be detected is input into the improved YOLOv8s network model for small target detection. The improved YOLOv8s network model includes: an improved Backbone and an improved Neck. A receptive field amplification module is set in the improved Backbone, and coordinate space attention is set in the improved Neck. Detection information corresponding to the small target objects in the remote sensing image to be detected is obtained according to the detection results. Compared with the prior art in which the small targets in the remote sensing image are detected by a target detection algorithm, the detection effect is not good. Since the present invention performs small target detection on the remote sensing image to be detected through the improved YOLOv8s network model and obtains the detection information corresponding to the small target objects in the remote sensing image to be detected according to the detection results, the technical problems of poor detection effect and low detection accuracy of the target detection algorithm for small remote sensing targets in the prior art are solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 FIG. is a schematic structural diagram of a remote sensing image small target detection device for the hardware operating environment involved in the embodiment solution of the present invention;

[0036] Figure 2 FIG. is a schematic flowchart of the first embodiment of the remote sensing image small target detection method of the present invention;

[0037] Figure 3 FIG. is a structural diagram of the improved YOLOv8s network model in the first embodiment of the remote sensing image small target detection method of the present invention;

[0038] Figure 4 FIG. is a schematic flowchart of the second embodiment of the remote sensing image small target detection method of the present invention;

[0039] Figure 5 FIG. is a comparison diagram of Bottleneck and BRA - Bottleneck in the second embodiment of the remote sensing image small target detection method of the present invention;

[0040] Figure 6 FIG. is a schematic structural diagram of the double - layer routing attention in the second embodiment of the remote sensing image small target detection method of the present invention;

[0041] Figure 7 FIG. is a schematic flowchart of the third embodiment of the remote sensing image small target detection method of the present invention;

[0042] Figure 8 FIG. is a schematic diagram of the improved feature pyramid structure in the third embodiment of the remote sensing image small target detection method of the present invention;

[0043] Figure 9 FIG. is a schematic structural diagram of the improved coordinate space attention in the third embodiment of the remote sensing image small target detection method of the present invention;

[0044] Figure 10 This is the structural block diagram of the first embodiment of the small target detection device for remote sensing images of the present invention.

[0045] The realization, functional features and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. Specific embodiments

[0046] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0047] Refer to Figure 1 , Figure 1 This is the structural schematic diagram of the small target detection device for remote sensing images of the hardware operating environment involved in the embodiment solution of the present invention.

[0048] As Figure 1 shown, the small target detection device for remote sensing images may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a wireless fidelity (Wi-Fi) interface). The memory 1005 may be a high-speed random access memory (Random Access Memory, RAM), or a stable non-volatile memory (Non-Volatile Memory, NVM), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0049] Those skilled in the art can understand that Figure 1 the structure shown in

[0050] As Figure 1 shown, in the memory 1005 as a storage medium, there may be included an operating system, a network communication module, a user interface module, and a small target detection program for remote sensing images.

[0051] In Figure 1In the small target detection device for remote sensing images shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the small target detection device for remote sensing images of the present invention can be arranged in the small target detection device for remote sensing images. The small target detection device for remote sensing images calls the small target detection program stored in the memory 1005 through the processor 1001 and executes the small target detection method provided by the embodiments of the present invention.

[0052] An embodiment of the present invention provides a method for detecting small targets in remote sensing images. Refer to Figure 2 , Figure 2 which is a schematic flowchart of the first embodiment of the small target detection method for remote sensing images of the present invention.

[0053] In this embodiment, the small target detection method for remote sensing images includes the following steps:

[0054] Step S10: Input the remote sensing image to be detected into the improved YOLOv8s network model for small target detection. The improved YOLOv8s network model includes an improved Backbone and an improved Neck. A receptive field amplification module is arranged in the improved Backbone, and a coordinate spatial attention is arranged in the improved Neck.

[0055] It should be noted that the execution subject of the method in this embodiment can be a small target detection device for remote sensing images that detects small targets in remote sensing images, or other remote sensing image small target detection systems that can achieve the same or similar functions and include this small target detection device for remote sensing images. Here, a remote sensing image small target detection system (hereinafter referred to as the system) is used to specifically illustrate the small target detection method for remote sensing images provided in this embodiment and the following embodiments.

[0056] It should be understood that the above remote sensing image to be detected can be any remote sensing image that needs to be subjected to small target detection obtained by vehicles such as spacecraft or airplanes.

[0057] It can be understood that the above improved YOLOv8s network model can be a model obtained by improving and optimizing the YOLOv8s network model. Among them, the YOLOv8s network model can be a SOTA (State-of-the-Art) model in the YOLO series.

[0058] Specifically, refer to Figure 3 , Figure 3 which is a structural diagram of the improved YOLOv8s network model in the first embodiment of the small target detection method for remote sensing images of the present invention. As Figure 3As shown in the figure, the network structure of the improved YOLOv8s network model in this embodiment may include: an input end, an improved Backbone, an improved Neck, and an improved Head, which are four parts.

[0059] In practical applications, the improved Backbone in this embodiment can extract features from the remote sensing image to be detected. It can downsample the original image by 32 times to obtain feature maps of different scales, efficiently capture global receptive field information through the added receptive field amplification module, and then effectively fuse it with local feature information. Finally, it performs spatial pyramid pooling through the SPPF module to further fuse the information of each feature layer. At the same time, the improved Neck in this embodiment amplifies larger-sized feature maps on the basis of the original YOLOv8 for feature fusion to obtain richer underlying feature information, and adds coordinate spatial attention between the horizontal fusion connections of the shallow feature maps, enabling the model to pay more attention to the position information of the object and enhancing the feature extraction ability for small targets. In addition, the detection head part of the improved Head in this embodiment adopts a decoupled head structure, calculating classification and regression separately. Among them, one Head is used for classification, and the classification loss is calculated using VFL (V-arifocal Loss); one Head is used for regression localization, and the localization loss is calculated using CIoU and DFL (Distribution Focal Loss), and finally the prediction result of the network is obtained. Among them, the receptive field amplification module can be a module for capturing global receptive field information; the coordinate spatial attention can be an attention mechanism with a spatial attention SA (Spatial Attention) branch added to the coordinate attention CA (Coordinate Attention).

[0060] Furthermore, in order to improve the stability of the model, before the step S10, the method further includes: performing data augmentation processing on the image dataset through the Mosaic method to obtain an augmented image dataset; inputting the augmented image dataset into the initial YOLOv8s network model for model training to obtain the improved YOLOv8s network model.

[0061] It should be noted that the above Mosaic method can be a method of splicing four pictures into one picture as a training sample for model training. Correspondingly, the above image dataset can be a dataset composed of several pictures for model training. Among them, the image dataset in this embodiment can be composed of four remote sensing images.

[0062] It can be understood that the above initial YOLOv8s network model can be the original YOLOv8s network model that has not been improved and optimized.

[0063] In a specific implementation, in this embodiment, the input image dataset can first be subjected to Mosaic data augmentation through the Mosaic method. Four arbitrary images are stitched together by means of random scaling, random cropping, and random arrangement, thereby effectively increasing the diversity of the dataset. Turn off Mosaic in the last 10 epochs of the model training stage. Since the YOLOv8s network model has learned good feature representations and weight parameters at this time, turning off the Mosaic operation can reduce the complexity of data augmentation, thereby improving the stability of the model.

[0064] Step S20: Obtain the detection information corresponding to the target small object in the to-be-detected remote sensing image according to the detection result.

[0065] It should be noted that the above-mentioned target small object can be a small object in the to-be-detected remote sensing image, where the small object can be an object with a pixel area less than 0.12% of the original image.

[0066] It can be understood that the above-mentioned detection information can be information such as the bounding box coordinates, object category, and confidence of the target small object. This embodiment does not limit this.

[0067] In a specific implementation, first, an image dataset composed of four remote sensing images can be input into the initial YOLOv8s network model. The initial YOLOv8s network model can perform Mosaic data augmentation on these input training data, and stitch these four remote sensing images by means of random scaling, random cropping, and random arrangement, so as to train the initial YOLOv8s network model based on these data to obtain an improved YOLOv8s network model. Then, any to-be-detected remote sensing image that needs to perform small target detection can be input into the improved YOLOv8s network model. At this time, the improved YOLOv8s network model can output detection information such as the bounding box coordinates, object category, and confidence of the small object in the to-be-detected remote sensing image.

[0068] This embodiment discloses inputting the remote sensing image to be detected into an improved YOLOv8s network model for small target detection. The improved YOLOv8s network model includes an improved Backbone and an improved Neck. A receptive field amplification module is set in the improved Backbone, and coordinate spatial attention is set in the improved Neck. Detection information corresponding to the target small objects in the remote sensing image to be detected is obtained according to the detection result. Compared with the prior art in which small targets in remote sensing images are detected by a target detection algorithm, the detection effect is not good. Since in this embodiment, the improved YOLOv8s network model is used to detect small targets in the remote sensing image to be detected, and detection information corresponding to the target small objects in the remote sensing image to be detected is obtained according to the detection result, the technical problems of poor detection effect and low detection accuracy of the target detection algorithm for remote sensing small targets in the prior art are solved.

[0069] Reference Figure 4 , Figure 4 is a schematic flowchart of the second embodiment of the method for detecting small targets in remote sensing images of the present invention.

[0070] Based on the above first embodiment, in order to enhance the detection accuracy of small target objects, in this embodiment, the step S20 includes:

[0071] Step S101: Input the remote sensing image to be detected into the improved YOLOv8s network model.

[0072] Step S102: Perform feature extraction processing on the remote sensing image to be detected through the double-layer routing attention set in the improved Backbone to obtain a to-be-detected remote sensing feature map corresponding to the remote sensing image to be detected.

[0073] It should be noted that the above double-layer routing attention can be an attention mechanism including two attention layers. Among them, the first attention layer is used for adaptive feature fusion of the to-be-detected remote sensing feature map to extract richer semantic information; the second attention layer is used for weighted processing of the fused features to highlight important target regions. In this embodiment, the double-layer routing attention mechanism can improve the accuracy and robustness of small target detection.

[0074] It should be understood that the above to-be-detected remote sensing feature map can be a feature map obtained after mapping the input remote sensing image to be detected to a high-dimensional space in a Convolutional Neural Network (CNN). Specifically, the to-be-detected remote sensing feature map is the result of performing convolution operations on the input remote sensing image to be detected by a series of convolution kernels in the convolutional neural network.

[0075] Further, the step S102 specifically includes: obtaining the global receptive field information of the remote sensing image to be detected through the double-layer routing attention set in the improved Backbone; and obtaining the remote sensing feature map to be detected corresponding to the remote sensing image to be detected based on the global receptive field information.

[0076] It should be noted that the above global receptive field information may be the size information of the mapping area of the pixel points on the feature map output by each layer of the convolutional neural network on the original input image. In practical applications, the system can obtain the remote sensing feature map to be detected corresponding to the remote sensing image to be detected according to the global receptive field information of the remote sensing image to be detected.

[0077] It should be noted that remote sensing images usually cover a vast geographical area, and the objects to be detected usually have high inter-class similarity and intra-class diversity. For example, "bridge" and "overpass", "track and field stadium" and "stadium", etc. These object information have high similarity, and the same class of objects also have different shapes and sizes in the image. If there is no rich global semantic information in the detection feature map, it is easy to cause misdetection. Therefore, this embodiment can analyze the feature extraction part of the YOLOv8 network structure. Since the receptive field in the highest layer feature map cannot fully cover the original image and the extracted features lack long-range dependence, from the perspective of receptive field amplification, the Bottleneck part in the C2f module can be improved, and a double-layer routing attention (Bi-Level Routing Attention, BRA) is added to its residual branch structure, and the original Bottleneck part in the four C2f modules in the backbone network feature extraction part is replaced with the improved BRA-Bottleneck module. Specifically, refer to Figure 5 , Figure 5 is a comparison diagram of Bottleneck and BRA-Bottleneck in the second embodiment of the remote sensing image small target detection method of the present invention. As Figure 5 shown, Figure 5 in (a) is the Bottleneck part in the four C2f modules; Figure 5 in (b) is the improved BRA-Bottlen-eck part.

[0078] It can be understood that, Figure 5 (b) The BRA module in is a double-layer routing attention module. To simplify the notation, a single-input single-head self-attention is used to represent its detailed structure. In the actual process, multi-head self-attention and batch input are adopted. The specific process may include (refer to Figure 6 , Figure 6 is the structural schematic diagram of the double-layer routing attention in the second embodiment of the remote sensing image small target detection method of the present invention):

[0079] Step 1: Divide the remote sensing features to be detected Figure X ∈R H×W×C into S×S non-overlapping regions, so that each region contains feature vectors. At this time, the feature map is reshaped from X to Then, tensors of query, key, and value are obtained through linear projection The formulas are as follows:

[0080] Q = X r W q ;

[0081] Q = X r W q ;

[0082] V = X r W v ;

[0083] where W q , W k , W v ∈R C×C are the projection weights of query, key, and value respectively.

[0084] Step 2: Find other regions that each given region should focus on by constructing a directed graph. First, apply the regional average to Q and K respectively to obtain the regional-level query and key Then, multiply Q r by the transpose of K r to obtain the region-to-region adjacency matrix The formula is:

[0085] A r = Q r (K r ) T ;

[0086] We measure the semantic correlation between every two regions through the entries in the adjacency matrix A γ , and then only keep the top k regions with strong correlation for each region, thus deriving a routing index matrix with a row-wise topk operator The expression is as follows:

[0087] I r = topkIndex(A r );

[0088] Therefore, the i-th row in I r contains the k indices most relevant to the i-th region.

[0089] Step 3: Utilize the region-to-region routing index matrix I obtained in Step 2r , construct a finer-grained token-to-token attention. For each query token in region i, it will attend to all key-value pairs in the union of k routing regions indexed by . However, since these routing regions are expected to be scattered across the entire feature map, and GPUs rely on coalesced memory operations, first, tensors of keys and values are collected, i.e.:

[0090] K g = gather(K, I r );

[0091] V g = gather(V, I r );

[0092] where are the tensors of the collected keys and values respectively. Then, self-attention is applied to the collected key-value pairs as follows:

[0093] O = Attention(Q, K g , V g ) + LCE(V);

[0094] In the formula, Attention(Q, K g , V g ) is calculated according to the multi-head self-attention calculation formula, and LCE(V) is the local context enhancement term, which is parameterized using depth convolution. In this scheme, the kernel size can be set to 3.

[0095] Correspondingly, the step S30 includes:

[0096] Step S20': Obtain the detection information corresponding to the target small objects in the to-be-detected remote sensing image based on the to-be-detected remote sensing feature map.

[0097] It should be understood that after obtaining the to-be-detected remote sensing feature map corresponding to the to-be-detected remote sensing image, the system can obtain the prediction boxes corresponding to the target small objects in the to-be-detected remote sensing feature map through the improved Backbone in the improved YOLOv8s network model, so as to obtain detection information such as the border coordinates, object category, and confidence of the target small objects.

[0098] In a specific implementation, the system can input the remote sensing image to be detected into the improved YOLOv8s network model, and perform feature extraction processing on the remote sensing image to be detected through the double-layer routing attention set in the improved Backbone in the improved YOLOv8s network model, obtain the remote sensing feature map to be detected corresponding to the remote sensing image to be detected, and then obtain the detection information corresponding to the small target object in the remote sensing image to be detected based on the remote sensing feature map to be detected. Since the double-layer routing attention module in this embodiment has dynamic query-aware sparsity, each query first passes through coarse-grained similarity screening to filter out most irrelevant key-value pairs, and then through finer-grained point-to-point fusion, so that the model can more efficiently focus on the feature information of the object. At the same time, the improved double-layer routing attention module effectively fuses local feature information and global semantic information in the feature extraction stage, improving the detection accuracy of small target objects in remote sensing images.

[0099] In this embodiment, the double-layer routing attention set in the improved Backbone is used to perform feature extraction processing on the remote sensing image to be detected, so as to obtain the remote sensing feature map to be detected corresponding to the remote sensing image to be detected, and obtain the detection information corresponding to the small target object in the remote sensing image to be detected based on the remote sensing feature map to be detected. Since this embodiment can efficiently capture global receptive field information through double-layer routing attention, the detection accuracy of small target objects is improved.

[0100] Reference Figure 7 , Figure 7 is a schematic flowchart of the third embodiment of the method for detecting small targets in remote sensing images of the present invention.

[0101] Based on the above embodiments, in order to make up for the lack of small target information volume and effectively enhance the extraction of small target information, in this embodiment, in step S20', the method further includes:

[0102] Step S301: Obtain the coordinate attention weight and spatial attention weight corresponding to the remote sensing feature map to be detected through the improved feature pyramid structure.

[0103] It should be noted that in the YOLOv8 network model, the feature fusion part fuses the feature maps obtained by downsampling the original image by 8 times, 16 times, and 32 times through the PAN network structure, and finally performs detection on the feature maps corresponding to three resolutions. This structure that combines the semantic information of the deep feature map with the position information of the shallow feature map is beneficial to multi-scale object detection. However, for small remote sensing targets, as the number of network layers deepens, the pixel information contained in small targets becomes less and less, resulting in much worse detection effects than large and medium targets. Therefore, this embodiment can start from the perspective of enhancing the feature information of small targets and improve the feature pyramid structure in the YOLOv8 network model. Specifically, in order to make full use of the feature information of small targets, this embodiment can utilize the shallow feature map downsampled by 4 times in the backbone network in the feature pyramid network to obtain the improved feature pyramid structure.

[0104] It should be understood that the above coordinate attention weights can be the coordinate weights used to represent each pixel position in the remote sensing feature map to be detected. Correspondingly, the above spatial attention weights can be the coordinate weights used to represent each pixel spatial position in the remote sensing feature map to be detected.

[0105] Furthermore, the improved feature pyramid structure is constructed based on the coordinate spatial attention, and the coordinate spatial attention includes: coordinate attention and spatial attention; the step S301 includes: performing one-dimensional average pooling processing on the remote sensing feature map to be detected through the coordinate attention to obtain the coordinate attention weights corresponding to the remote sensing feature map to be detected; performing channel dimension pooling processing on the remote sensing feature map to be detected through the spatial attention to obtain the spatial attention weights corresponding to the remote sensing feature map to be detected.

[0106] It can be understood that coordinate attention (Coordinate Attention, CA) can be a mechanism that encodes channel relationships and long-range dependencies through precise position information, enabling the network to focus on large important regions with relatively low computational costs. Correspondingly, spatial attention can be a mechanism that enables the network to focus on large important regions through spatial position information.

[0107] It should be noted that with reference to Figure 8 , Figure 8 is a schematic diagram of the improved feature pyramid structure in the third embodiment of the small target detection method for remote sensing images of the present invention, as Figure 8As shown in the figure, in order to make full use of the position information in the shallow feature map, this embodiment introduces coordinate attention CA and improves its structure by adding a spatial attention SA branch. The improved coordinate spatial attention CSA (Coordinate Spatial Attention) is added to the lateral connection part of the shallow feature map to obtain an improved feature pyramid structure.

[0108] It should be noted that, referring to Figure 9 , Figure 9 is the structural schematic diagram of the improved coordinate spatial attention in the third embodiment of the remote sensing image small target detection method of the present invention. As Figure 9 shown, for a given remote sensing feature to be detected Figure X ∈R C×H×W , through Figure 9 the CA on the left part, the coordinate weights (i.e., the above-mentioned coordinate attention weights) of each pixel position in the feature map are obtained. Considering that it ignores the interaction between complete spatial positions, this embodiment can add an SA branch to its structure (as shown in the middle and right parts of Figure 9 ). Finally, the obtained coordinate attention weights, spatial attention weights and the original remote sensing feature map to be detected can be aggregated to generate a target remote sensing feature map with coordinate spatial information.

[0109] It can be understood that the above one-dimensional average pooling process can be to divide the input remote sensing feature map to be detected into several regions of the same size, and take the average value of the feature values in each region to obtain a new feature map.

[0110] Specifically, the specific process of coordinate attention CA can be: first, the original remote sensing feature map to be detected is subjected to one-dimensional average pooling along the horizontal and vertical directions respectively, Figure 9 denoted by x Avg Pool and yAvg Pool in

[0111]

[0112] In the formula, x c (h,i) is the feature value corresponding to the pixel point with height h and width i on the c-th channel. Similarly, the output of the c-th channel with width w can be expressed as:

[0113]

[0114] In the formula, x c (j,w) is the feature value corresponding to the pixel point with width w and height j on the c-th channel. Then, the two output feature maps are concatenated along the channel dimension and transformed using the 1×1 convolution transformation function F1. The formula is as follows: f = δ(F1([zh , z w ));

[0115] In the formula, [·, ·] represents channel concatenation along the spatial dimension, δ is a non-linear activation function, and f ∈ R C / r×1×(W+H) is the generated intermediate feature map, where r is the reduction rate used to control the number of channels, which is set to 16 in this paper. Considering that the number of channels in the shallow feature map is relatively small itself, the computational cost of the model is also reduced. Finally, f is split into two independent tensors f h ∈ R C / r×1×H and f w ∈ R C / r×1×W . Two 1×1 convolution transformation functions F2 and F3 are used to transform f h and f w into tensors with the same number of channels as the input X, and the formula is as follows:

[0116] g h = σ(F h (f h ));

[0117] g w = σ(F w (f w ));

[0118] In the formula: σ is the Sigmoid activation function, and g h and g w are the transformed tensor outputs.

[0119] It should be understood that the above channel dimension pooling process performs averaging on the input remote sensing feature map to be detected in the channel dimension to obtain a new feature map.

[0120] Specifically, the specific process of spatial attention SA can be as follows: First, perform max pooling and average pooling on the input remote sensing feature map to be detected in the channel dimension respectively, then concatenate the two generated feature maps in the channel dimension, and use the 7×7 convolution transformation function F2 for transformation to reduce it to 1 channel while keeping W and H unchanged. Finally, use the Sigmoid activation function to generate a feature map with spatial weight information, and the formula is as follows:

[0121] s(i, j) = σ(F2(<AvgPool C (X), MaxPool C (X)>));

[0122] In the formula, AvgPool C (X) represents performing average pooling on the input remote sensing feature map to be detected in the channel dimension, and MaxPool C(X) represents performing max pooling on the remote sensing feature map to be detected in the channel dimension, and <·,·> represents channel concatenation along the channel dimension.

[0123] Step S302: Generate a target remote sensing feature map based on the coordinate attention weight, the spatial attention weight, and the remote sensing feature map to be detected, where the target remote sensing feature map is a feature map carrying coordinate space information.

[0124] It can be understood that the above target remote sensing feature map is a feature map carrying coordinate space information. In practical applications, in this embodiment, the coordinate attention weight, the spatial attention weight, and the original remote sensing feature map to be detected can be aggregated to generate a target remote sensing feature map with coordinate space information.

[0125] Step S303: Obtain detection information corresponding to the target small objects in the remote sensing image to be detected based on the target remote sensing feature map.

[0126] In a specific implementation, in this embodiment, the g h and g w generated by CA, the s(i,j) generated by SA, and the remote sensing feature map to be detected are aggregated, corresponding to Figure 9 the Re - weight part in, and the target remote sensing feature map Y is output. The formula is as follows:

[0127]

[0128] In the formula, x c (i,j) is the feature value at the position coordinate (i,j) in the c - th channel of the remote sensing feature Figure X to be detected, is the horizontal direction attention weight of height i in the c - th channel of the remote sensing feature Figure X to be detected, is the vertical direction attention weight of width j in the c - th channel of the remote sensing feature Figure X to be detected, s(i,j) is the spatial attention weight at the position coordinate (i,j) of the remote sensing feature Figure X to be detected, and y c (i,j) is the feature value at the position coordinate (i,j) in the c - th channel of the output feature map Y.

[0129] Furthermore, in order to effectively solve the problem of difficult detection of dense small targets and improve the detection accuracy, the above step S303 includes: performing post - detection processing on the target remote sensing feature map through an adaptive non - maximum suppression algorithm to obtain a processed target remote sensing feature map; obtaining detection information corresponding to the target small objects in the remote sensing image to be detected based on the processed target remote sensing feature map.

[0130] It should be noted that the above adaptive non-maximum suppression algorithm can be an algorithm that uses a higher threshold for areas where objects are densely distributed to impose a lighter penalty on or retain prediction boxes with lower confidence, and a lower threshold for areas where objects are sparsely distributed to eliminate more redundant boxes.

[0131] It should be understood that the above processed target remote sensing feature map can be a remote sensing feature map obtained after removing redundant boxes from the target remote sensing feature map.

[0132] It can be understood that in the post-processing of object detection by the traditional non-maximum suppression (NMS) algorithm, when the IoU of two prediction boxes for predicting the same type of object is greater than a preset threshold, the prediction box with lower confidence will be directly and violently deleted, resulting in many missed detections when detecting small and dense objects. Among them, objects such as ships and vehicles in remote sensing images belong to small and dense objects. To solve this problem, the NMS algorithm in the improved YOLOv8 network model in this embodiment can use adaptive NMS for post-processing, thereby effectively reducing the missed detection rate of small targets and further improving the detection accuracy.

[0133] In a specific implementation, the system can input B = {b1, b2,..., b n},S = {s1, s2,..., s n},D = {d1, d2,..., d n},N t ,S t into the adaptive non-maximum suppression algorithm, where B is a set composed of all prediction boxes b i (i = 1,..., n) of a certain type of object in the target remote sensing feature map, S is a set composed of the confidence score s i (i = 1,..., n) corresponding to each prediction box, D is a set composed of the density value score d i (i = 1,..., n) corresponding to each prediction box, N t is a preset NMS threshold, S t is a preset confidence threshold, and its specific calculation is: R ← {}; while do; S m ← max(S); M ← b m ; N M = max(N t , d m ); R ← R ∪ M; B ← B - M; for b i in B do; if iou(M, b i ) ≥ N Mthen; S m ← s i f(iou(M, b i )); if s i < S t then; B ← B - b i ; S ← S - S1. At this time, the adaptive non - maximum suppression algorithm can output: R, S.

[0134] Among them, f(iou(M, b i )) adopts the Gaussian weighting method, and the specific expression is as follows:

[0135]

[0136] The above formula satisfies that the greater the overlap degree between two prediction boxes, the greater the penalty for the prediction box with low confidence, because they are very likely to identify the same object.

[0137] At the same time, in this embodiment, the density value can be defined as the maximum value of the IoU of each prediction box and all true annotation boxes, and the formula is as follows:

[0138] d i = max(iou(b i , G t ));

[0139] In the formula, G t represents the true annotation box, that is, the position of the small object in the remote - sensing image.

[0140] In this embodiment, the improved feature pyramid structure is used to obtain the coordinate attention weight and spatial attention weight corresponding to the remote - sensing feature map to be detected, and the target remote - sensing feature map carrying coordinate - space information is generated based on the coordinate attention weight, spatial attention weight and the remote - sensing feature map to be detected. Then, the detection information corresponding to the target small object in the remote - sensing image to be detected is obtained based on the target remote - sensing feature map, so as to make up for the lack of small - target information volume and effectively enhance the extraction of small - target information. At the same time, in this embodiment, the adaptive non - maximum suppression algorithm is used to perform post - detection processing on the target remote - sensing feature map to obtain the processed target remote - sensing feature map, and the detection information corresponding to the target small object in the remote - sensing image to be detected is obtained based on the processed target remote - sensing feature map, so as to effectively solve the problem of difficult detection of dense small targets and improve the detection accuracy.

[0141] In addition, an embodiment of the present invention also proposes a storage medium, on which a remote - sensing image small - target detection program is stored. When the remote - sensing image small - target detection program is executed by a processor, the steps of the remote - sensing image small - target detection method described above are implemented.

[0142] Refer to Figure 10 ,Figure 10 This is the structural block diagram of the first embodiment of the small target detection device for remote sensing images of the present invention.

[0143] As Figure 10 shown, the small target detection device for remote sensing images proposed in the embodiment of the present invention includes:

[0144] An object detection module 501, configured to input the remote sensing image to be detected into the improved YOLOv8s network model for small target detection. The improved YOLOv8s network model includes an improved Backbone and an improved Neck. A receptive field amplification module is provided in the improved Backbone, and a coordinate spatial attention is provided in the improved Neck.

[0145] A detection information acquisition module 502, configured to obtain the detection information corresponding to the target small object in the remote sensing image to be detected according to the detection result.

[0146] Furthermore, the object detection module 501 is further configured to perform data augmentation processing on the image data set by the Mosaic method to obtain an enhanced image data set; input the enhanced image data set into the initial YOLOv8s network model for model training to obtain the improved YOLOv8s network model.

[0147] The small target detection device for remote sensing images in this embodiment discloses inputting the remote sensing image to be detected into the improved YOLOv8s network model for small target detection. The improved YOLOv8s network model includes an improved Backbone and an improved Neck. A receptive field amplification module is provided in the improved Backbone, and a coordinate spatial attention is provided in the improved Neck; obtaining the detection information corresponding to the target small object in the remote sensing image to be detected according to the detection result; compared with the prior art for detecting small targets in remote sensing images by a target detection algorithm, the detection effect is not good. Since in this embodiment, the improved YOLOv8s network model is used to detect small targets in the remote sensing image to be detected, and the detection information corresponding to the target small object in the remote sensing image to be detected is obtained according to the detection result, the technical problems of poor detection effect and low detection accuracy of the target detection algorithm for small remote sensing targets in the prior art are solved.

[0148] Based on the first embodiment of the small target detection device for remote sensing images of the present invention, a second embodiment of the small target detection device for remote sensing images of the present invention is proposed.

[0149] In this embodiment, the object detection module 501 is further configured to input the remote sensing image to be detected into the improved YOLOv8s network model; perform feature extraction processing on the remote sensing image to be detected through the double-layer routing attention set in the improved Backbone to obtain the remote sensing feature map to be detected corresponding to the remote sensing image to be detected.

[0150] The detection information acquisition module 502 is further configured to obtain the detection information corresponding to the small target object in the remote sensing image to be detected based on the remote sensing feature map to be detected.

[0151] Furthermore, the object detection module 501 is further configured to obtain the global receptive field information of the remote sensing image to be detected through the double-layer routing attention set in the improved Backbone; obtain the remote sensing feature map to be detected corresponding to the remote sensing image to be detected based on the global receptive field information.

[0152] In this embodiment, feature extraction processing is performed on the remote sensing image to be detected through the double-layer routing attention set in the improved Backbone to obtain the remote sensing feature map to be detected corresponding to the remote sensing image to be detected, and the detection information corresponding to the small target object in the remote sensing image to be detected is obtained based on the remote sensing feature map to be detected. Since this embodiment can efficiently capture the global receptive field information through the double-layer routing attention, the detection accuracy of small target objects is improved.

[0153] Based on the above device embodiments, a third embodiment of the small target detection device for remote sensing images of the present invention is proposed.

[0154] In this embodiment, the detection information acquisition module 502 is further configured to obtain the coordinate attention weight and the spatial attention weight corresponding to the remote sensing feature map to be detected through the improved feature pyramid structure; generate a target remote sensing feature map based on the coordinate attention weight, the spatial attention weight, and the remote sensing feature map to be detected, where the target remote sensing feature map is a feature map carrying coordinate space information; obtain the detection information corresponding to the small target object in the remote sensing image to be detected based on the target remote sensing feature map.

[0155] Furthermore, the improved feature pyramid structure is constructed based on the coordinate space attention, and the coordinate space attention includes: coordinate attention and spatial attention; the detection information acquisition module 502 is further configured to perform one-dimensional average pooling processing on the remote sensing feature map to be detected through the coordinate attention to obtain the coordinate attention weight corresponding to the remote sensing feature map to be detected; perform channel dimension pooling processing on the remote sensing feature map to be detected through the spatial attention to obtain the spatial attention weight corresponding to the remote sensing feature map to be detected.

[0156] Further, the detection information acquisition module 502 is further configured to perform post-detection processing on the target remote sensing feature map through an adaptive non-maximum suppression algorithm to obtain a processed target remote sensing feature map; and acquire detection information corresponding to the target small object in the remote sensing image to be detected based on the processed target remote sensing feature map.

[0157] In this embodiment, the improved feature pyramid structure is used to obtain the coordinate attention weight and the spatial attention weight corresponding to the remote sensing feature map to be detected, and the target remote sensing feature map carrying coordinate space information is generated based on the coordinate attention weight, the spatial attention weight, and the remote sensing feature map to be detected. Then, the detection information corresponding to the target small object in the remote sensing image to be detected is acquired based on the target remote sensing feature map, so as to make up for the lack of small target information volume and effectively enhance the extraction of small target information. At the same time, in this embodiment, the adaptive non-maximum suppression algorithm is used to perform post-detection processing on the target remote sensing feature map to obtain a processed target remote sensing feature map, and the detection information corresponding to the target small object in the remote sensing image to be detected is acquired based on the processed target remote sensing feature map, so as to effectively solve the problem of difficult detection of dense small targets and improve the detection accuracy.

[0158] It should be noted that in this article, the terms "include", "comprise" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or system. Without further limitation, an element defined by the phrase "including a..." does not exclude the existence of additional identical elements in the process, method, article or system including the element.

[0159] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.

[0160] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as a read-only memory / random access memory, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0161] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.

Claims

1. A method for detecting small targets in remote sensing images, characterized in that The small target detection method for remote sensing images includes: Input the remote sensing image to be detected into the improved YOLOv8s network model for small target detection. The improved YOLOv8s network model includes an improved Backbone and an improved Neck. A receptive field amplification module is set in the improved Backbone, and a coordinate spatial attention is set in the improved Neck. Obtain the detection information corresponding to the small target objects in the remote sensing image to be detected according to the detection results.

2. The remote sensing image small target detection method according to claim 1, characterized in that, Before the step of inputting the remote sensing image to be detected into the improved YOLOv8s network model for small target detection, it further includes: Perform data augmentation processing on the image dataset through the Mosaic method to obtain an augmented image dataset. Input the augmented image dataset into the initial YOLOv8s network model for model training to obtain the improved YOLOv8s network model.

3. The remote sensing image small target detection method according to claim 2, characterized in that, The step of inputting the remote sensing image to be detected into the improved YOLOv8s network model for small target detection includes: Input the remote sensing image to be detected into the improved YOLOv8s network model. Perform feature extraction processing on the remote sensing image to be detected through the double-layer routing attention set in the improved Backbone to obtain the to-be-detected remote sensing feature map corresponding to the remote sensing image to be detected. The step of obtaining the detection information corresponding to the small target objects in the remote sensing image to be detected according to the detection results includes: Obtain the detection information corresponding to the small target objects in the remote sensing image to be detected based on the to-be-detected remote sensing feature map.

4. The remote sensing image small target detection method according to claim 3, characterized in that, The step of performing feature extraction processing on the remote sensing image to be detected through the double-layer routing attention set in the improved Backbone to obtain the to-be-detected remote sensing feature map corresponding to the remote sensing image to be detected includes: Obtain the global receptive field information of the remote sensing image to be detected through the double-layer routing attention set in the improved Backbone. Obtain the to-be-detected remote sensing feature map corresponding to the remote sensing image to be detected based on the global receptive field information.

5. The remote sensing image small target detection method according to claim 3, characterized in that, The step of obtaining the detection information corresponding to the small target objects in the remote sensing image to be detected based on the to-be-detected remote sensing feature map includes: Obtain the coordinate attention weight and spatial attention weight corresponding to the to-be-detected remote sensing feature map through the improved feature pyramid structure. Generate a target remote sensing feature map based on the coordinate attention weight, the spatial attention weight, and the to-be-detected remote sensing feature map. The target remote sensing feature map is a feature map carrying coordinate spatial information. Obtain the detection information corresponding to the small target objects in the remote sensing image to be detected based on the target remote sensing feature map.

6. The remote sensing image small target detection method according to claim 5, wherein The improved feature pyramid structure is constructed based on the coordinate spatial attention. The coordinate spatial attention includes coordinate attention and spatial attention. The step of obtaining the coordinate attention weight and spatial attention weight corresponding to the to-be-detected remote sensing feature map through the improved feature pyramid structure includes: Perform one-dimensional average pooling on the remote sensing feature map to be detected through the coordinate attention to obtain the coordinate attention weight corresponding to the remote sensing feature map to be detected; Perform channel dimension pooling on the remote sensing feature map to be detected through the spatial attention to obtain the spatial attention weight corresponding to the remote sensing feature map to be detected.

7. The remote sensing image small target detection method according to claim 5, characterized in that The step of obtaining the detection information corresponding to the target small object in the to-be-detected remote sensing image based on the target remote sensing feature map includes: Perform post-processing detection on the target remote sensing feature map through an adaptive non-maximum suppression algorithm to obtain the processed target remote sensing feature map; Obtain the detection information corresponding to the target small object in the to-be-detected remote sensing image based on the processed target remote sensing feature map.

8. A small target detection device for remote sensing images, characterized in that, The device includes: An object detection module, configured to input a to-be-detected remote sensing image into an improved YOLOv8s network model for small target detection. The improved YOLOv8s network model includes an improved Backbone and an improved Neck. A receptive field amplification module is provided in the improved Backbone, and coordinate spatial attention is provided in the improved Neck; A detection information acquisition module, configured to obtain the detection information corresponding to the target small object in the to-be-detected remote sensing image according to the detection result.

9. A small target detection device for remote sensing images, characterized in that, The device includes: a memory, a processor, and a remote sensing image small target detection program stored on the memory and executable on the processor. The remote sensing image small target detection program is configured to implement the steps of the remote sensing image small target detection method according to any one of claims 1 to 7.

10. A storage medium, characterized in that, A remote sensing image small target detection program is stored on the storage medium. When the remote sensing image small target detection program is executed by a processor, the steps of the remote sensing image small target detection method according to any one of claims 1 to 7 are implemented.