Remote sensing image small target detection method and device

By introducing windmill convolution and adaptive receptive field space attention modules into the remote sensing image small object detection model, multi-scale feature extraction and fusion of small objects in remote sensing images is realized, solving the problem of insufficient detection accuracy of small objects in remote sensing images, and improving the accuracy and reliability of detection.

CN120339594APending Publication Date: 2025-07-18GUANGDONG POWER GRID CO LTD
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510742400.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively extract and maintain the fine-grained and high-resolution feature information of small objects in remote sensing images, resulting in insufficient detection accuracy of small objects.

Method used

A small object detection model is adopted, including a backbone network, a neck network and a detection head network. By introducing a windmill convolution module and an adaptive receptive field space attention module, feature extraction capabilities are enhanced, and multi-scale feature fusion and prediction are carried out.

Benefits of technology

It improves the detection accuracy and reliability of small targets in remote sensing images, can accurately identify and locate small targets, and retain high-resolution detailed information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339594A_ABST
    Figure CN120339594A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing image small target detection method and device, and belongs to the field of target detection, and the method comprises the steps: obtaining a remote sensing image; the image is processed through a small target detection model to generate multiple paths of original prediction tensors, the model extracts multi-scale features through a backbone network integrated with a windmill convolution module and an adaptive receptive field space attention module, and a neck network of the model further fuses the features and outputs at least three paths of enhanced feature maps with different resolutions; the detection head network generates the first original prediction tensor, the second original prediction tensor and the third original prediction tensor in parallel based on the three paths of enhanced feature maps; and finally generating a detection target information list according to the three original prediction tensors. Therefore, by implementing the method, the problem that fine-grained and high-resolution feature information sufficient to distinguish and accurately position small targets is difficult to effectively extract and maintain in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection, and particularly to a method and device for detecting small objects in remote sensing images. Background Art

[0002] Remote sensing images play an increasingly important role in key fields such as national defense security, resource census, and disaster monitoring. Accurately detecting small-sized objects (such as vehicles, small facilities, etc.) from them is of core value for obtaining fine information and supporting rapid response. Therefore, continuously improving the detection accuracy and reliability of small objects in remote sensing images is an urgent need for the technological development in this field.

[0003] However, when the current mainstream deep learning-based object detection methods are applied to the detection of small objects in remote sensing images, they generally face a core bottleneck: it is difficult to effectively extract and maintain fine-grained and high-resolution feature information sufficient to distinguish and accurately locate small objects. This is mainly because small objects themselves occupy very few pixels, their inherent visual features are weak and easily confused with complex backgrounds; at the same time, the successive downsampling operations of the deep convolutional network to expand the receptive field, although beneficial to the recognition of medium and large objects, inevitably lead to a significant loss of spatial details and texture information crucial for small objects. Summary of the Invention

[0004] Embodiments of the present invention provide a method and device for detecting small objects in remote sensing images, which can solve the problem in the prior art of being difficult to effectively extract and maintain fine-grained and high-resolution feature information sufficient to distinguish and accurately locate small objects.

[0005] An embodiment of the present invention provides a method for detecting small objects in remote sensing images, including:

[0006] Obtaining a remote sensing image to be detected;

[0007] Inputting the remote sensing image into a small object detection model, so that the small object detection model generates a first tensor, a second tensor, and a third tensor according to the remote sensing image;

[0008] Generating a detection target information list according to the first tensor, the second tensor, and the third tensor;

[0009] Wherein, the small object detection model includes a backbone network, a neck network, and a detection head network; the backbone network includes at least four successively connected backbone module groups, the output of each previous backbone module group is the input of the subsequent backbone module group, and the spatial scale of the feature map decreases along the connection direction; at least one of the backbone module groups includes a windmill convolution module, and at least one of the backbone module groups includes an adaptive receptive field spatial attention module; the backbone network outputs at least four-way scale feature maps to the neck network;

[0010] The neck network includes at least three sequentially connected neck module groups, the output of each previous neck module group is the input of the subsequent neck module group, and each neck module group also fuses at least one scale feature map as another input, and the spatial scale of the feature map increases along the connection direction; the neck network outputs at least three enhanced feature maps to the detection head network; the enhanced feature maps include a first enhanced feature map, a second enhanced feature map, and a third enhanced feature map, and the resolution of the first enhanced feature map is greater than that of the second enhanced feature map, and the resolution of the second enhanced feature map is greater than that of the third enhanced feature map;

[0011] The detection head network includes three prediction paths respectively for receiving the first enhanced feature map, the second enhanced feature map, and the third enhanced feature map, and each prediction path is used to generate the first tensor, the second tensor, and the third tensor according to the corresponding enhanced feature map.

[0012] Further, the backbone network includes a first backbone module group, a second backbone module group, a third backbone module group, and a fourth backbone module group; the neck network includes a first neck module group, a second neck module group, and a third neck module group; the detection head network includes a first prediction path, a second prediction path, and a third prediction path;

[0013] The first output end of the first backbone module group is connected to the input end of the second backbone module group;

[0014] The second output end of the first backbone module group is connected to the first input end of the first neck module group;

[0015] The first output end of the second backbone module group is connected to the input end of the third backbone module group;

[0016] The second output end of the second backbone module group is connected to the first input end of the second neck module group;

[0017] The first output end of the third backbone module group is connected to the input end of the fourth backbone module group;

[0018] The second output end of the third backbone module group is connected to the first input end of the third neck module group;

[0019] The output end of the fourth backbone module group is connected to the second input end of the third neck module group;

[0020] The first output end of the third neck module group is connected to the second input end of the second neck module group;

[0021] The second output end of the third neck module group is connected to the input end of the third prediction path;

[0022] The first output end of the second neck module group is connected to the second input end of the first neck module group;

[0023] The second output end of the second neck module group is connected to the input end of the second prediction path;

[0024] The third output end of the second neck module group is connected to the third input end of the third neck module group;

[0025] The first output end of the first neck module group is connected to the input end of the first prediction path;

[0026] The second output end of the first neck module group is connected to the third input end of the second neck module group.

[0027] Furthermore, the first backbone module group includes a first windmill convolution module, a second windmill convolution module, and a first adaptive receptive field spatial attention module;

[0028] The second backbone module group includes a third windmill convolution module and a second adaptive receptive field spatial attention module;

[0029] The third backbone module group includes a fourth windmill convolution module and a first region attention module; the fourth backbone module group includes a fifth windmill convolution module and a second region attention module;

[0030] The output end of the first windmill convolution module is connected to the input end of the second windmill convolution module;

[0031] The output end of the second windmill convolution module is connected to the input end of the first adaptive receptive field spatial attention module;

[0032] The first output end of the first adaptive receptive field spatial attention module is the first output end of the first backbone module group;

[0033] The second output end of the first adaptive receptive field spatial attention module is the second output end of the first backbone module group;

[0034] The input end of the third windmill convolution module is the input end of the second backbone module group;

[0035] The output end of the third windmill convolution module is connected to the input end of the second adaptive receptive field spatial attention module;

[0036] The first output end of the second adaptive receptive field spatial attention module is the first output end of the second backbone module group;

[0037] The second output end of the second adaptive receptive field spatial attention module is the second output end of the second backbone module group;

[0038] The input end of the fourth windmill convolution module is the input end of the third backbone module group;

[0039] The output end of the fourth windmill convolution module is connected to the input end of the first region attention module;

[0040] The first output end of the first region attention module is the first output end of the third backbone module group;

[0041] The second output end of the first region attention module is the second output end of the third backbone module group;

[0042] The input end of the fifth windmill convolution module is the input end of the fourth backbone module group;

[0043] The output end of the fifth windmill convolution module is connected to the input end of the second region attention module;

[0044] The output end of the second region attention module is the output end of the fourth backbone module group.

[0045] Further, the first neck module group includes a first upsampling module, a first splicing module, a third region attention module, and a first standard convolution module;

[0046] The second neck module group includes a second upsampling module, a second splicing module, a fourth region attention module, a third splicing module, a fifth region attention module, and a second standard convolution module;

[0047] The third neck module group includes a third upsampling module, a fourth splicing module, a sixth region attention module, a fifth splicing module, and a seventh region attention module;

[0048] The input end of the third upsampling module is the second input end of the third neck module group;

[0049] The output end of the third upsampling module is connected to the first input end of the fourth splicing module;

[0050] The second input end of the fourth splicing module is the first input end of the third neck module group;

[0051] The output end of the fourth splicing module is connected to the input end of the sixth region attention module;

[0052] The first output end of the sixth region attention module is the first output end of the third neck module group;

[0053] The second output end of the sixth region attention module is connected to the first input end of the fifth splicing module;

[0054] The second input end of the fifth splicing module is the third input end of the third neck module group;

[0055] The output end of the fifth splicing module is connected to the input end of the seventh region attention module;

[0056] The output end of the seventh region attention module is the second output end of the third neck module group;

[0057] The input end of the second upsampling module is the second input end of the second neck module group;

[0058] The output end of the second upsampling module is connected to the first input end of the second splicing module;

[0059] The second input end of the second splicing module is the first input end of the second neck module group;

[0060] The output end of the second splicing module is connected to the input end of the fourth region attention module;

[0061] The first output end of the fourth region attention module is the first output end of the second neck module group;

[0062] The second output end of the fourth region attention module is connected to the first input end of the third splicing module;

[0063] The output end of the third splicing module is connected to the input end of the fifth region attention module;

[0064] The first output end of the fifth region attention module is the second output end of the second neck module group;

[0065] The second output end of the fifth region attention module is connected to the input end of the second standard convolution module;

[0066] The output end of the second standard convolution module is the third output end of the second neck module group;

[0067] The input end of the first upsampling module is the second input end of the first neck module group;

[0068] The output end of the first upsampling module is connected to the first input end of the first splicing module;

[0069] The second input end of the first splicing module is the first input end of the first neck module group;

[0070] The output end of the first splicing module is connected to the input end of the third region attention module;

[0071] The first output end of the third region attention module is the second output end of the first neck module group;

[0072] The second output end of the third region attention module is connected to the input end of the first standard convolution module;

[0073] The output end of the first standard convolution module is the third output end of the first neck module group.

[0074] Furthermore, the adaptive receptive field spatial attention module includes a first windmill convolution sub-module, a segmentation sub-module, a first receptive field spatial attention sub-module, a second receptive field spatial attention sub-module, a splicing sub-module, and a second windmill convolution sub-module;

[0075] The first output end of the first windmill convolution sub-module is connected to the input end of the segmentation sub-module;

[0076] The second output end of the first windmill convolution sub-module is connected to the first input end of the splicing sub-module;

[0077] The output end of the segmentation sub-module is connected to the input end of the first receptive field spatial attention sub-module;

[0078] The output end of the first receptive field spatial attention sub-module is connected to the input end of the second receptive field spatial attention sub-module;

[0079] The output end of the second receptive field spatial attention sub-module is connected to the second input end of the splicing sub-module;

[0080] The output end of the splicing sub-module is connected to the input end of the second windmill convolution sub-module.

[0081] Furthermore, each receptive field spatial attention sub-module includes a first windmill convolution unit, a second windmill convolution unit, a third windmill convolution unit, a first residual bottleneck unit, a second residual bottleneck unit, and a splicing unit;

[0082] The first output end of the first windmill convolution unit is connected to the input end of the first residual bottleneck unit;

[0083] The second output end of the first windmill convolution unit is connected to the input end of the second windmill convolution unit;

[0084] The output end of the first residual bottleneck unit is connected to the input end of the second residual bottleneck unit;

[0085] The output end of the second residual bottleneck unit is connected to the first input end of the splicing unit;

[0086] The output end of the second windmill convolution unit is connected to the second input end of the splicing unit;

[0087] The output end of the splicing unit is connected to the input end of the third windmill convolution unit.

[0088] Furthermore, each residual bottleneck unit includes a first windmill convolution sub-unit, a second windmill convolution sub-unit, and a multi-region perception attention convolution sub-unit;

[0089] The output end of the first windmill convolution sub-unit is connected to the input end of the second windmill convolution sub-unit;

[0090] The output end of the second windmill convolution sub-unit is connected to the input end of the multi-region perception attention convolution sub-unit;

[0091] The output of the residual bottleneck unit is formed by element-wise addition of the output of the multi-region perception attention convolution sub-unit and the external input features received by the first windmill convolution sub-unit.

[0092] Furthermore, the training of the small target detection model includes:

[0093] Obtain a remote sensing image small target detection data set; wherein, the remote sensing image small target detection data set includes a number of remote sensing images and corresponding first target reference tensors, second target reference tensors, and third target reference tensors for each of the remote sensing images; the first target reference tensor, the second target reference tensor, and the third target reference tensor are generated by performing preset target assignment processing and parameter encoding processing according to the original true target bounding boxes and class labels of the corresponding remote sensing images;

[0094] Randomly divide the remote sensing image small target detection data set into several batches of training samples according to a preset batch size;

[0095] Input each batch of training samples into the small target detection model in turn, and perform iterative training on the small target detection model until a preset number of training rounds is reached; wherein, when the small target detection model receives each batch of training samples, it outputs a first prediction tensor, a second prediction tensor, and a third prediction tensor corresponding to the remote sensing images of that batch; through a preset scale-aware dynamic loss function, according to the first prediction tensor, the second prediction tensor, the third prediction tensor, the first target reference tensor, the second target reference tensor, and the third target reference tensor, calculate and generate a loss function value; use a preset optimizer to update the small target detection model according to the loss function value.

[0096] Furthermore, the scale-aware dynamic loss function is specifically:

[0097] L SBDL = ω IoU ·LIoU +ω mask ·L mask

[0098] In the formula, L SBDL is the scale-aware dynamic loss function; L IoU is the intersection over union loss function; ω IoU is the weight of the intersection over union loss function; L mask is the mask loss function; ω mask is the weight of the mask loss function.

[0099] Based on the above method item embodiments, the present invention correspondingly provides apparatus item embodiments.

[0100] An embodiment of the present invention provides a remote sensing image small target detection apparatus, including: a remote sensing image acquisition module, a small target detection model solving module, and a detection target list generation module;

[0101] The remote sensing image acquisition module is used to acquire the remote sensing image to be detected;

[0102] The small target detection model solving module is used to input the remote sensing image into the small target detection model, so that the small target detection model generates a first tensor, a second tensor, and a third tensor according to the remote sensing image; wherein, the small target detection model includes a backbone network, a neck network, and a detection head network; the backbone network includes at least four sequentially connected backbone module groups, the output of each previous backbone module group is the input of the subsequent backbone module group, and the spatial scale of the feature map decreases along the connection direction; at least one of the backbone module groups includes a windmill convolution module, and at least one of the backbone module groups includes an adaptive receptive field spatial attention module; the backbone network outputs at least four-way scale feature maps to the neck network; the neck network includes at least three sequentially connected neck module groups, the output of each previous neck module group is the input of the subsequent neck module group, and each neck module group also fuses at least one-way scale feature map as another input, and the spatial scale of the feature map increases along the connection direction; the neck network outputs at least three-way enhanced feature maps to the detection head network; the enhanced feature maps include a first enhanced feature map, a second enhanced feature map, and a third enhanced feature map, and the resolution of the first enhanced feature map is greater than that of the second enhanced feature map, and the resolution of the second enhanced feature map is greater than that of the third enhanced feature map; the detection head network includes three prediction paths respectively used to receive the first enhanced feature map, the second enhanced feature map, and the third enhanced feature map, and each prediction path is used to generate the first tensor, the second tensor, and the third tensor according to the corresponding enhanced feature map;

[0103] The detection target list generation module is used to generate a detection target information list according to the first tensor, the second tensor, and the third tensor.

[0104] Compared with the prior art, the present invention has the following beneficial effects:

[0105] An embodiment of the present invention provides a method and device for detecting small targets in remote sensing images. The method obtains a remote sensing image to be detected and inputs it into a small target detection model including a backbone network, a neck network, and a detection head network, sequentially generating multi-scale feature maps and prediction tensors, and finally extracting a list of detection target information. Among them, the backbone network enhances the small target feature extraction ability by introducing windmill convolution and an adaptive receptive field spatial attention module; the neck network realizes multi-scale feature fusion and enhancement; the detection head network performs multi-path prediction based on enhanced feature maps with different resolutions.

[0106] The present invention integrates an innovative windmill convolution module and an adaptive receptive field spatial attention module in the backbone network, strengthening the capture ability of weak but key identification information of small targets from the source of feature extraction; at the same time, on the basis of multi-scale feature fusion, the neck network generates and provides a first enhanced feature map with the highest spatial resolution and other enhanced feature maps with different resolutions to the detection head network. This structural design enables the features finally sent to the detection head network for predicting small targets to retain both the fine spatial structure information brought by high resolution and the enhanced identification features optimized and extracted by the backbone network, thus providing more accurate feature input for accurately identifying and positioning small targets in subsequent detection links, and further improving the overall performance of small target detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0107] Figure 1 is a schematic flowchart of a method for detecting small targets in remote sensing images provided by an embodiment of the present invention.

[0108] Figure 2 is a schematic diagram of the network structure of a small target detection model provided by an embodiment of the present invention.

[0109] Figure 3 is a schematic diagram of the module structure of an adaptive receptive field spatial attention module provided by an embodiment of the present invention.

[0110] Figure 4 is a schematic diagram of the sub-module structure of a receptive field spatial attention sub-module provided by an embodiment of the present invention.

[0111] Figure 5 is a schematic diagram of the unit structure of a residual bottleneck unit provided by an embodiment of the present invention.

[0112] Figure 6 is a schematic diagram of the structure of a device for detecting small targets in remote sensing images provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0113] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0114] It should be noted that the small target detection model mentioned in the present invention is an improvement based on the existing YOLO (You Only Look Once) series of models, especially the improvement of YOLOv12. The pinwheel convolution module, pinwheel convolution sub-module, pinwheel convolution unit, and pinwheel convolution sub-unit mentioned in the present invention are expressions set to better highlight the hierarchy, and their essence is the pinwheel convolution layer (Pinwheel-shaped Convolution, abbreviated as PConv) in the prior art; the adaptive receptive field spatial attention module (abbreviated as C3k2-R) in the present invention is based on the improvement of the C3k2 module in the YOLO series of models; among them, the receptive field spatial attention sub-module in the present invention is the C3k2 module, and in the present invention, it is only called the receptive field spatial attention sub-module for the requirement of writing naming; the region attention module in the present invention is essentially the A2C2f module in the YOLO series of models in the prior art; the receptive field spatial attention sub-module (abbreviated as C3k-R) in the present invention is based on the improvement of the C3k module in the YOLO series of models; in the present invention, the residual bottleneck unit can be abbreviated as R-Bottleneck; in the present invention, the multi-region perception attention convolution sub-unit is the module named RFCBAMConv in the prior art; in the accompanying drawings of the present invention, the splicing module, splicing sub-module, and splicing unit are expressions set to better highlight the hierarchy, and their essence is the splicing operation layer, which is used to splice two feature maps along a specified dimension into a larger and more information-rich feature map; the upsampling module mentioned in the present invention is essentially the upsampling operation layer, which is used to improve the spatial resolution of the feature map; the standard convolution module mentioned in the present invention is essentially the standard convolution layer. + means element accumulation operation. The segmentation sub-module mentioned in the present invention is essentially the segmentation operation layer, which is used to divide an input feature map into two or more independent output feature maps along a specified dimension.

[0115] As Figure 1 shown, to solve the problem in the prior art that it is difficult to effectively extract and maintain fine-grained and high-resolution feature information sufficient to distinguish and accurately locate small targets, an embodiment of the present invention provides a method for detecting small targets in remote sensing images, which at least includes the following steps:

[0116] Step S1: Obtain the remote sensing image to be detected;

[0117] Specifically, these remote sensing images usually appear as digital image data, which can be sourced from image records of the Earth's surface area periodically or on demand by various types of sensors, including but not limited to optical, synthetic aperture radar, or other sensors mounted on satellites, aircraft, or drones and other Earth observation platforms. The obtained remote sensing images carry the spatial, spectral, and texture information of ground objects and are the original data source for performing subsequent target detection tasks. This step aims to provide the necessary visual information input for the subsequent small target detection model for the model to analyze, process, and ultimately identify and locate the small targets that may be contained therein.

[0118] Step S2: Input the remote sensing image into the small target detection model so that the small target detection model generates a first tensor, a second tensor, and a third tensor based on the remote sensing image;

[0119] Specifically, after obtaining the remote sensing image to be detected, the next step is to transmit the remote sensing image (in some embodiments, it may first undergo conventional preprocessing operations such as size normalization and data augmentation) as input data to the small target detection model that has been constructed and loaded with pre-trained parameters. The small target detection model, as a deep neural network, usually sequentially includes a backbone network part for extracting hierarchical features from the input image, a neck network part for fusing different-level features to enhance the representation of multi-scale targets (especially small targets), and a detection head network part for finally performing dense prediction on the fused feature map. When the remote sensing image data flows through these network parts, the model abstracts and analyzes the image information layer by layer through a series of pre-learned convolution, non-linear activation, and other feature transformation operations, and on the output layers of multiple different perceptual scales of its detection head network, for the targets that may exist in the image, respectively generate multi-channel structured data containing original prediction values such as position offsets, size information, target existence confidence, and probabilities of each predetermined category. These structured data are the aforementioned first tensor, second tensor, and third tensor. The completion of this step provides a direct multi-scale data basis for decoding and screening specific target detection results from the original prediction tensors output by these models.

[0120] Step S3: Generate a list of detection target information based on the first tensor, the second tensor, and the third tensor;

[0121] Among them, the small target detection model includes a backbone network, a neck network, and a detection head network; the backbone network includes at least four successively connected backbone module groups, the output of each previous backbone module group is the input of the subsequent backbone module group, and the spatial scale of the feature map decreases along the connection direction; at least one of the backbone module groups includes a windmill convolution module, and at least one of the backbone module groups includes an adaptive receptive field spatial attention module; the backbone network outputs at least four-way scale feature maps to the neck network;

[0122] The neck network includes at least three successively connected neck module groups, the output of each previous neck module group is the input of the subsequent neck module group, and each neck module group also fuses at least one-way scale feature map as another input, and the spatial scale of the feature map increases along the connection direction; the neck network outputs at least three enhanced feature maps to the detection head network; the enhanced feature maps include a first enhanced feature map, a second enhanced feature map, and a third enhanced feature map, and the resolution of the first enhanced feature map is greater than that of the second enhanced feature map, and the resolution of the second enhanced feature map is greater than that of the third enhanced feature map;

[0123] The detection head network includes three prediction paths respectively used to receive the first enhanced feature map, the second enhanced feature map, and the third enhanced feature map, and each prediction path is used to generate the first tensor, the second tensor, and the third tensor according to the corresponding enhanced feature map.

[0124] Specifically, after the small target detection model generates the first tensor, the second tensor, and the third tensor corresponding to different perception scales based on the input remote sensing image, these tensors contain a large amount of original prediction information about potential targets, such as the position and size of the bounding box in encoded form, the confidence of the target's existence, and the class score to which the target may belong. To obtain the final available small target detection results from these original and usually dense prediction data, the method will next perform a series of necessary post-processing operations on the prediction information represented by these three-way tensors. These post-processing operations generally first include decoding the encoded bounding box parameters output by the model, converting them to the coordinate space of the input remote sensing image, and forming candidate bounding boxes with clear geometric meanings; immediately afterwards, these candidate bounding boxes will be preliminarily screened according to a preset confidence threshold to filter out most of the predictions with low credibility and reduce the computational complexity of subsequent processing. On this basis, to solve the problem that the same target may be detected by multiple overlapping candidate boxes, the non-maximum suppression (NMS) algorithm or its optimized version will be further implemented. This algorithm compares the spatial overlap (such as intersection over union IoU) and confidence scores between candidate boxes, eliminates redundant detection boxes, and retains the optimal prediction for each target. Through the above post-processing steps such as decoding, screening, and suppression, the information from the three-way original prediction tensors is finally integrated, and a structured and clear list of detected target information is generated. Each item in the list represents an identified small target and contains its exact bounding box coordinates, determined class, and corresponding detection confidence score in the remote sensing image. The completion of this step effectively converts the original numerical predictions output by the model into precise information about small targets in the remote sensing image that can be directly used by users or further analyzed by other systems.

[0125] Specifically, the small target detection model adopted in the present invention is mainly composed of three core parts, namely the backbone network, the neck network, and the detection head network, which are cascaded. They work together to complete the entire process of extracting hierarchical features from the original remote sensing image, performing multi-scale information fusion and enhancement, and finally outputting the target prediction tensor.

[0126] In this model, the backbone network serves as the core for initial feature extraction, and its internal structure consists of at least four consecutively connected backbone module groups. Data is sequentially passed between these backbone module groups, that is, the feature maps extracted by each previous backbone module group will serve as the input for its adjacent subsequent backbone module group. And as the data stream progresses deeper in the backbone network, the spatial scale of the feature maps (i.e., their width and height) will systematically decrease step by step. This design enables the network to learn rich hierarchical features from the underlying details to the high-level semantics of the image. To specifically enhance the feature representation ability for weak and complex small targets in remote sensing images, at least one backbone module group in this backbone network innovatively incorporates a windmill convolution module, which uses its unique convolution operation method to enhance the diversity and effectiveness of feature extraction. Moreover, at least one backbone module group also integrates an adaptive receptive field spatial attention module, which can further optimize and refine the discriminability of the extracted features by dynamically adjusting its receptive field range and focusing on the key information regions in the image. After the complete forward propagation of the backbone network, at least four feature maps with different spatial scales and semantic depths will be formed, and these feature maps will be provided as the basic input to the subsequent neck network for further processing and information integration.

[0127] Following the backbone network is the neck network. The core responsibility of this network is to receive the at least four different-scale feature maps from the backbone network and perform efficient multi-scale fusion and depth enhancement on these features. Its internal structure includes at least three consecutively connected neck module groups, and these module groups work together to jointly construct complex feature fusion paths. In a main top-down information transmission direction (i.e., from the module group processing deeper and smaller-scale features to the module group processing shallower and larger-scale features), the output features processed by the previous neck module group will serve as a key input for the subsequent neck module group. At the same time, in order to fully combine the context information and detail features at different levels, each neck module group will also fuse at least another scale feature map directly or indirectly from the corresponding level of the backbone network. During this series of fusion and processing processes, the spatial scale of the feature maps gradually increases along this main information transmission direction, which helps to effectively combine the strong semantic information extracted from the deep layer of the backbone network with the high-resolution spatial details retained in the shallow layer. Finally, the neck network outputs at least three significantly enhanced feature maps to the detection head network. These enhanced feature maps have clear hierarchical relationships and different spatial resolutions, specifically including the first enhanced feature map with the highest spatial resolution, the second enhanced feature map with the second-highest spatial resolution, and the third enhanced feature map with a relatively lower spatial resolution, providing an optimized and adapted feature basis for the subsequent detection of targets of different sizes.

[0128] Finally, the detection head network is responsible for making final object predictions based on the enhanced feature maps carefully prepared by the neck network. There are three independent prediction paths inside the detection head network, which are respectively used to receive and process the first enhanced feature map, the second enhanced feature map, and the third enhanced feature map from the neck network. Each prediction path performs intensive calculations on the enhanced feature map of a specific scale it receives through the built-in prediction layer to generate the first tensor, the second tensor, and the third tensor containing the original encoded information such as the position coordinates, size information, subordinate category, and existence confidence of the possible target objects in the image. Through the deep feature extraction and optimization by the backbone network, the effective fusion of multi-scale features and the preservation of high-resolution details by the neck network, and the multi-path parallel prediction of the detection head network, the small object detection model of the present invention can provide high-quality and information-rich original prediction data for accurately identifying and locating small objects in remote sensing images.

[0129] In a preferred embodiment, the backbone network includes a first backbone module group, a second backbone module group, a third backbone module group, and a fourth backbone module group; the neck network includes a first neck module group, a second neck module group, and a third neck module group; the detection head network includes a first prediction path, a second prediction path, and a third prediction path;

[0130] The first output end of the first backbone module group is connected to the input end of the second backbone module group;

[0131] The second output end of the first backbone module group is connected to the first input end of the first neck module group;

[0132] The first output end of the second backbone module group is connected to the input end of the third backbone module group;

[0133] The second output end of the second backbone module group is connected to the first input end of the second neck module group;

[0134] The first output end of the third backbone module group is connected to the input end of the fourth backbone module group;

[0135] The second output end of the third backbone module group is connected to the first input end of the third neck module group;

[0136] The output end of the fourth backbone module group is connected to the second input end of the third neck module group;

[0137] The first output end of the third neck module group is connected to the second input end of the second neck module group;

[0138] The second output end of the third neck module group is connected to the input end of the third prediction path;

[0139] The first output end of the second neck module group is connected to the second input end of the first neck module group;

[0140] The second output end of the second neck module group is connected to the input end of the second prediction path;

[0141] The third output end of the second neck module group is connected to the third input end of the third neck module group;

[0142] The first output end of the first neck module group is connected to the input end of the first prediction path;

[0143] The second output end of the first neck module group is connected to the third input end of the second neck module group.

[0144] As Figure 2 shown, in a preferred embodiment, the first backbone module group includes a first windmill convolution module, a second windmill convolution module, and a first adaptive receptive field spatial attention module;

[0145] The second backbone module group includes a third windmill convolution module and a second adaptive receptive field spatial attention module;

[0146] The third backbone module group includes a fourth windmill convolution module and a first region attention module; the fourth backbone module group includes a fifth windmill convolution module and a second region attention module;

[0147] The output end of the first windmill convolution module is connected to the input end of the second windmill convolution module;

[0148] The output end of the second windmill convolution module is connected to the input end of the first adaptive receptive field spatial attention module;

[0149] The first output end of the first adaptive receptive field spatial attention module is the first output end of the first backbone module group;

[0150] The second output end of the first adaptive receptive field spatial attention module is the second output end of the first backbone module group;

[0151] The input end of the third windmill convolution module is the input end of the second backbone module group;

[0152] The output end of the third windmill convolution module is connected to the input end of the second adaptive receptive field spatial attention module;

[0153] The first output end of the second adaptive receptive field spatial attention module is the first output end of the second backbone module group;

[0154] The second output end of the second adaptive receptive field spatial attention module is the second output end of the second backbone module group;

[0155] The input end of the fourth windmill convolution module is the input end of the third backbone module group;

[0156] The output end of the fourth windmill convolution module is connected to the input end of the first region attention module;

[0157] The first output end of the first region attention module is the first output end of the third backbone module group;

[0158] The second output end of the first region attention module is the second output end of the third backbone module group;

[0159] The input end of the fifth windmill convolution module is the input end of the fourth backbone module group;

[0160] The output end of the fifth windmill convolution module is connected to the input end of the second region attention module;

[0161] The output end of the second region attention module is the output end of the fourth backbone module group.

[0162] As Figure 2 shown, in a preferred embodiment, the first neck module group includes a first upsampling module, a first splicing module, a third region attention module, and a first standard convolution module;

[0163] The second neck module group includes a second upsampling module, a second splicing module, a fourth region attention module, a third splicing module, a fifth region attention module, and a second standard convolution module;

[0164] The third neck module group includes a third upsampling module, a fourth splicing module, a sixth region attention module, a fifth splicing module, and a seventh region attention module;

[0165] The input end of the third upsampling module is the second input end of the third neck module group;

[0166] The output end of the third upsampling module is connected to the first input end of the fourth splicing module;

[0167] The second input end of the fourth splicing module is the first input end of the third neck module group;

[0168] The output end of the fourth splicing module is connected to the input end of the sixth region attention module;

[0169] The first output end of the sixth region attention module is the first output end of the third neck module group;

[0170] The second output terminal of the sixth region attention module is connected to the first input terminal of the fifth splicing module;

[0171] The second input terminal of the fifth splicing module is the third input terminal of the third neck module group;

[0172] The output terminal of the fifth splicing module is connected to the input terminal of the seventh region attention module;

[0173] The output terminal of the seventh region attention module is the second output terminal of the third neck module group;

[0174] The input terminal of the second upsampling module is the second input terminal of the second neck module group;

[0175] The output terminal of the second upsampling module is connected to the first input terminal of the second splicing module;

[0176] The second input terminal of the second splicing module is the first input terminal of the second neck module group;

[0177] The output terminal of the second splicing module is connected to the input terminal of the fourth region attention module;

[0178] The first output terminal of the fourth region attention module is the first output terminal of the second neck module group;

[0179] The second output terminal of the fourth region attention module is connected to the first input terminal of the third splicing module;

[0180] The output terminal of the third splicing module is connected to the input terminal of the fifth region attention module;

[0181] The first output terminal of the fifth region attention module is the second output terminal of the second neck module group;

[0182] The second output terminal of the fifth region attention module is connected to the input terminal of the second standard convolution module;

[0183] The output terminal of the second standard convolution module is the third output terminal of the second neck module group;

[0184] The input terminal of the first upsampling module is the second input terminal of the first neck module group;

[0185] The output terminal of the first upsampling module is connected to the first input terminal of the first splicing module;

[0186] The second input terminal of the first splicing module is the first input terminal of the first neck module group;

[0187] The output end of the first splicing module is connected to the input end of the third region attention module;

[0188] The first output end of the third region attention module is the second output end of the first neck module group;

[0189] The second output end of the third region attention module is connected to the input end of the first standard convolution module;

[0190] The output end of the first standard convolution module is the third output end of the first neck module group.

[0191] As Figure 3 shown, in a preferred embodiment, the adaptive receptive field spatial attention module includes a first windmill convolution sub-module, a segmentation sub-module, a first receptive field spatial attention sub-module, a second receptive field spatial attention sub-module, a splicing sub-module, and a second windmill convolution sub-module;

[0192] The first output end of the first windmill convolution sub-module is connected to the input end of the segmentation sub-module;

[0193] The second output end of the first windmill convolution sub-module is connected to the first input end of the splicing sub-module;

[0194] The output end of the segmentation sub-module is connected to the input end of the first receptive field spatial attention sub-module;

[0195] The output end of the first receptive field spatial attention sub-module is connected to the input end of the second receptive field spatial attention sub-module;

[0196] The output end of the second receptive field spatial attention sub-module is connected to the second input end of the splicing sub-module;

[0197] The output end of the splicing sub-module is connected to the input end of the second windmill convolution sub-module.

[0198] As Figure 4 shown, in a preferred embodiment, each receptive field spatial attention sub-module includes a first windmill convolution unit, a second windmill convolution unit, a third windmill convolution unit, a first residual bottleneck unit, a second residual bottleneck unit, and a splicing unit;

[0199] The first output end of the first windmill convolution unit is connected to the input end of the first residual bottleneck unit;

[0200] The second output end of the first windmill convolution unit is connected to the input end of the second windmill convolution unit;

[0201] The output end of the first residual bottleneck unit is connected to the input end of the second residual bottleneck unit;

[0202] The output end of the second residual bottleneck unit is connected to the first input end of the splicing unit;

[0203] The output end of the second windmill convolution unit is connected to the second input end of the splicing unit;

[0204] The output end of the splicing unit is connected to the input end of the third windmill convolution unit.

[0205] As Figure 5 shown, in a preferred embodiment, each residual bottleneck unit includes a first windmill convolution sub-unit, a second windmill convolution sub-unit, and a multi-region perception attention convolution sub-unit;

[0206] The output end of the first windmill convolution sub-unit is connected to the input end of the second windmill convolution sub-unit;

[0207] The output end of the second windmill convolution sub-unit is connected to the input end of the multi-region perception attention convolution sub-unit;

[0208] The output of the residual bottleneck unit is formed by element-wise addition of the output of the multi-region perception attention convolution sub-unit and the external input features received by the first windmill convolution sub-unit.

[0209] In a preferred embodiment, the training of the small target detection model includes:

[0210] Obtaining a remote sensing image small target detection data set; wherein, the remote sensing image small target detection data set includes a plurality of remote sensing images and first target reference tensors, second target reference tensors, and third target reference tensors corresponding to the respective remote sensing images; the first target reference tensors, second target reference tensors, and third target reference tensors are generated by performing preset target assignment processing and parameter encoding processing according to the original true target bounding boxes and class labels of the corresponding remote sensing images;

[0211] Randomly dividing the remote sensing image small target detection data set into several batches of training samples according to a preset batch size;

[0212] Input the training samples of each batch into the small target detection model in sequence, and perform iterative training on the small target detection model until the preset number of training rounds is reached; wherein, when the small target detection model receives each batch of training samples, it outputs the first prediction tensor, the second prediction tensor, and the third prediction tensor corresponding to the remote sensing image of this batch; through a preset scale-aware dynamic loss function, calculate and generate a loss function value according to the first prediction tensor, the second prediction tensor, the third prediction tensor, the first target reference tensor, the second target reference tensor, and the third target reference tensor; use a preset optimizer to update the small target detection model according to the loss function value.

[0213] Specifically, to train the small target detection model, it is first necessary to obtain and prepare a remote sensing image small target detection dataset. This dataset usually contains a large number of remote sensing images covering different geographical environments, different time phases, and having diverse scene contents, as well as training labels that are precisely corresponding to each remote sensing image and have been specially processed to directly adapt to the multi-scale output structure of the model. These labels are specifically manifested as the first target reference tensor, the second target reference tensor, and the third target reference tensor. These target reference tensors themselves are processed and generated based on the original true target bounding box coordinates and class information manually annotated or obtained by other means in each remote sensing image through a preset target assignment strategy (such as intersection over union matching or center point matching rules) and parameter encoding method (such as converting absolute coordinates into offsets and size ratios relative to anchor points or grid cells) that aim to map the true target information to the model prediction space, so as to ensure the accuracy and effectiveness of the supervision signal during the training process. Before the formal training starts, the entire dataset is usually randomly divided into several batches of training samples according to a preset batch size. This batch processing helps to perform efficient learning with limited computing resources and introduces a certain degree of randomness into the optimization algorithm to facilitate model generalization.

[0214] The training process itself is carried out in an iterative optimization manner: the divided batches of training samples are fed into the small target detection model one by one and in sequence. After the model receives each batch of remote sensing image data, it will perform a complete forward propagation calculation. According to its internal network structure (including the backbone network, neck network, and detection head network) and the currently learned parameters, it deeply processes and analyzes the input image information, and correspondingly outputs its original prediction results at three different perception scales, namely the first prediction tensor, the second prediction tensor, and the third prediction tensor. These tensors densely contain preliminary numerical judgments on information such as the position, size, class attribution, and existence confidence of potential targets in the image. Subsequently, a preset scale-aware dynamic loss function (SBDL) is adopted. The core mechanism of this loss function is to comprehensively compare the above three-way prediction tensors output by the model with the three-way target reference tensors corresponding to the current batch of training samples, so as to calculate a scalar loss value that can quantify the deviation degree between the current prediction performance of the model and the real situation. And the design of this loss function particularly focuses on the learning balance and contribution of different scale targets (especially small targets) during the training process. Finally, with the help of a preset optimization algorithm (such as a variant of the stochastic gradient descent method SGD, such as Adam, AdamW, etc.), according to the calculated loss function value, the gradient of the loss with respect to the learnable parameters of each layer of the model (such as convolution kernel weights, bias terms, etc.) is calculated through the backpropagation algorithm, and these parameters are adjusted and updated slightly based on this gradient information. This iterative training process will continue until the performance of the model on an independent validation dataset reaches a preset satisfactory standard, or the preset number of training epochs is completed. Through such a systematic and end-to-end training process, it is aimed to enable the small target detection model to fully learn and internalize the deep patterns and key discriminant features for accurately identifying and precisely locating various small targets from complex remote sensing image backgrounds.

[0215] In a preferred embodiment, the scale-aware dynamic loss function is specifically:

[0216] L SBDL = ω IoU ·L IoU + ω mask ·L mask

[0217] In the formula, L SBDL is the scale-aware dynamic loss function; L IoU is the intersection over union loss function; ω IoU is the weight of the intersection over union loss function; L mask is the mask loss function; ω mask is the weight of the mask loss function.

[0218] In one embodiment, the weight ω of the intersection over union loss function is calculated by the following formula IoU :

[0219]

[0220] In the formula, A bbox is the area of the target bounding box; A image is the total area of the image; α is a preset first parameter factor;

[0221] In one embodiment, the weight ω of the mask loss function is calculated by the following formula mask :

[0222]

[0223] In the formula, β is a preset second parameter factor;

[0224] Based on the above method item embodiments, the present invention correspondingly provides apparatus item embodiments.

[0225] As Figure 6 shown, an embodiment of the present invention provides a remote sensing image small target detection apparatus, including: a remote sensing image acquisition module, a small target detection model solving module, and a detection target list generation module;

[0226] The remote sensing image acquisition module is used to acquire the remote sensing image to be detected;

[0227] The small target detection model solving module is used to input the remote sensing image into the small target detection model, so that the small target detection model generates a first tensor, a second tensor and a third tensor according to the remote sensing image; wherein, the small target detection model includes a backbone network, a neck network and a detection head network; the backbone network includes at least four successively connected backbone module groups, the output of each previous backbone module group is the input of the subsequent backbone module group, and the spatial scale of the feature map decreases along the connection direction; at least one of the backbone module groups includes a windmill convolution module, and at least one of the backbone module groups includes an adaptive receptive field spatial attention module; the backbone network outputs at least four-way scale feature maps to the neck network; the neck network includes at least three successively connected neck module groups, the output of each previous neck module group is the input of the subsequent neck module group, and each neck module group also fuses at least one-way scale feature map as another input, and the spatial scale of the feature map increases along the connection direction; the neck network outputs at least three enhanced feature maps to the detection head network; the enhanced feature maps include a first enhanced feature map, a second enhanced feature map and a third enhanced feature map, and the resolution of the first enhanced feature map is greater than that of the second enhanced feature map, and the resolution of the second enhanced feature map is greater than that of the third enhanced feature map; the detection head network includes three prediction paths respectively for receiving the first enhanced feature map, the second enhanced feature map and the third enhanced feature map, and each prediction path is used to generate the first tensor, the second tensor and the third tensor according to the corresponding enhanced feature map;

[0228] The detection target list generation module is used to generate a detection target information list according to the first tensor, the second tensor and the third tensor.

[0229] It should be noted that the embodiments of the device described above correspond to the above embodiments of the present invention and can implement any one of the above remote sensing image small target detection methods of the present invention. In addition, the embodiments of the above device are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement without creative work.

[0230] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0231] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A method for detecting small targets in remote sensing images, characterized in that, including: obtaining a remote sensing image to be detected; inputting the remote sensing image into a small target detection model, so that the small target detection model generates a first tensor, a second tensor, and a third tensor according to the remote sensing image; generating a list of detection target information according to the first tensor, the second tensor, and the third tensor; wherein, the small target detection model includes a backbone network, a neck network, and a detection head network; the backbone network includes at least four successively connected backbone module groups, the output of each previous backbone module group is the input of the subsequent backbone module group, and the spatial scale of the feature map decreases along the connection direction; at least one of the backbone module groups includes a windmill convolution module, and at least one of the backbone module groups includes an adaptive receptive field spatial attention module; the backbone network outputs at least four-way scale feature maps to the neck network; the neck network includes at least three successively connected neck module groups, the output of each previous neck module group is the input of the subsequent neck module group, and each neck module group also fuses at least one-way scale feature map as another input, and the spatial scale of the feature map increases along the connection direction; the neck network outputs at least three enhanced feature maps to the detection head network; the enhanced feature maps include a first enhanced feature map, a second enhanced feature map, and a third enhanced feature map, and the resolution of the first enhanced feature map is greater than that of the second enhanced feature map, and the resolution of the second enhanced feature map is greater than that of the third enhanced feature map; the detection head network includes three prediction paths for receiving the first enhanced feature map, the second enhanced feature map, and the third enhanced feature map respectively, and each prediction path is used to generate the first tensor, the second tensor, and the third tensor according to the corresponding enhanced feature map.

2. The small target detection method for remote sensing images according to claim 1, wherein The backbone network includes a first backbone module group, a second backbone module group, a third backbone module group, and a fourth backbone module group; the neck network includes a first neck module group, a second neck module group, and a third neck module group; the detection head network includes a first prediction path, a second prediction path, and a third prediction path; the first output end of the first backbone module group is connected to the input end of the second backbone module group; the second output end of the first backbone module group is connected to the first input end of the first neck module group; the first output end of the second backbone module group is connected to the input end of the third backbone module group; the second output end of the second backbone module group is connected to the first input end of the second neck module group; the first output end of the third backbone module group is connected to the input end of the fourth backbone module group; the second output end of the third backbone module group is connected to the first input end of the third neck module group; the output end of the fourth backbone module group is connected to the second input end of the third neck module group; the first output end of the third neck module group is connected to the second input end of the second neck module group; the second output end of the third neck module group is connected to the input end of the third prediction path; the first output end of the second neck module group is connected to the second input end of the first neck module group; the second output end of the second neck module group is connected to the input end of the second prediction path; The third output terminal of the second neck module group is connected to the third input terminal of the third neck module group; The first output terminal of the first neck module group is connected to the input terminal of the first prediction path; The second output terminal of the first neck module group is connected to the third input terminal of the second neck module group.

3. The remote sensing image small target detection method according to claim 2, characterized in that, The first backbone module group includes a first windmill convolution module, a second windmill convolution module, and a first adaptive receptive field spatial attention module; The second backbone module group includes a third windmill convolution module and a second adaptive receptive field spatial attention module; The third backbone module group includes a fourth windmill convolution module and a first region attention module; the fourth backbone module group includes a fifth windmill convolution module and a second region attention module; The output terminal of the first windmill convolution module is connected to the input terminal of the second windmill convolution module; The output terminal of the second windmill convolution module is connected to the input terminal of the first adaptive receptive field spatial attention module; The first output terminal of the first adaptive receptive field spatial attention module is the first output terminal of the first backbone module group; The second output terminal of the first adaptive receptive field spatial attention module is the second output terminal of the first backbone module group; The input terminal of the third windmill convolution module is the input terminal of the second backbone module group; The output terminal of the third windmill convolution module is connected to the input terminal of the second adaptive receptive field spatial attention module; The first output terminal of the second adaptive receptive field spatial attention module is the first output terminal of the second backbone module group; The second output terminal of the second adaptive receptive field spatial attention module is the second output terminal of the second backbone module group; The input terminal of the fourth windmill convolution module is the input terminal of the third backbone module group; The output terminal of the fourth windmill convolution module is connected to the input terminal of the first region attention module; The first output terminal of the first region attention module is the first output terminal of the third backbone module group; The second output terminal of the first region attention module is the second output terminal of the third backbone module group; The input terminal of the fifth windmill convolution module is the input terminal of the fourth backbone module group; The output terminal of the fifth windmill convolution module is connected to the input terminal of the second region attention module; The output terminal of the second region attention module is the output terminal of the fourth backbone module group.

4. The remote sensing image small target detection method according to claim 3, wherein, The first neck module group includes a first upsampling module, a first splicing module, a third region attention module, and a first standard convolution module; The second neck module group includes a second upsampling module, a second splicing module, a fourth region attention module, a third splicing module, a fifth region attention module, and a second standard convolution module; The third neck module group includes a third upsampling module, a fourth splicing module, a sixth region attention module, a fifth splicing module, and a seventh region attention module; The input terminal of the third upsampling module is the second input terminal of the third neck module group; The output terminal of the third upsampling module is connected to the first input terminal of the fourth splicing module; The second input terminal of the fourth splicing module is the first input terminal of the third neck module group; The output end of the fourth splicing module is connected to the input end of the sixth region attention module; The first output end of the sixth region attention module is the first output end of the third neck module group; The second output end of the sixth region attention module is connected to the first input end of the fifth splicing module; The second input end of the fifth splicing module is the third input end of the third neck module group; The output end of the fifth splicing module is connected to the input end of the seventh region attention module; The output end of the seventh region attention module is the second output end of the third neck module group; The input end of the second upsampling module is the second input end of the second neck module group; The output end of the second upsampling module is connected to the first input end of the second splicing module; The second input end of the second splicing module is the first input end of the second neck module group; The output end of the second splicing module is connected to the input end of the fourth region attention module; The first output end of the fourth region attention module is the first output end of the second neck module group; The second output end of the fourth region attention module is connected to the first input end of the third splicing module; The output end of the third splicing module is connected to the input end of the fifth region attention module; The first output end of the fifth region attention module is the second output end of the second neck module group; The second output end of the fifth region attention module is connected to the input end of the second standard convolution module; The output end of the second standard convolution module is the third output end of the second neck module group; The input end of the first upsampling module is the second input end of the first neck module group; The output end of the first upsampling module is connected to the first input end of the first splicing module; The second input end of the first splicing module is the first input end of the first neck module group; The output end of the first splicing module is connected to the input end of the third region attention module; The first output end of the third region attention module is the second output end of the first neck module group; The second output end of the third region attention module is connected to the input end of the first standard convolution module; The output end of the first standard convolution module is the third output end of the first neck module group.

5. The remote sensing image small target detection method according to claim 4, wherein The adaptive receptive field spatial attention module includes a first windmill convolution sub-module, a segmentation sub-module, a first receptive field spatial attention sub-module, a second receptive field spatial attention sub-module, a splicing sub-module, and a second windmill convolution sub-module; The first output end of the first windmill convolution sub-module is connected to the input end of the segmentation sub-module; The second output end of the first windmill convolution sub-module is connected to the first input end of the splicing sub-module; The output end of the segmentation sub-module is connected to the input end of the first receptive field spatial attention sub-module; The output end of the first receptive field spatial attention sub-module is connected to the input end of the second receptive field spatial attention sub-module; The output end of the second receptive field spatial attention sub-module is connected to the second input end of the splicing sub-module; The output end of the splicing sub-module is connected to the input end of the second windmill convolution sub-module.

6. The remote sensing image small target detection method according to claim 5, wherein, Each receptive field spatial attention sub-module includes a first windmill convolution unit, a second windmill convolution unit, a third windmill convolution unit, a first residual bottleneck unit, a second residual bottleneck unit, and a splicing unit; The first output end of the first windmill convolution unit is connected to the input end of the first residual bottleneck unit; The second output end of the first windmill convolution unit is connected to the input end of the second windmill convolution unit; The output end of the first residual bottleneck unit is connected to the input end of the second residual bottleneck unit; The output end of the second residual bottleneck unit is connected to the first input end of the splicing unit; The output end of the second windmill convolution unit is connected to the second input end of the splicing unit; The output end of the splicing unit is connected to the input end of the third windmill convolution unit.

7. The remote sensing image small target detection method according to claim 6, characterized in that Each residual bottleneck unit includes a first windmill convolution sub-unit, a second windmill convolution sub-unit, and a multi-region perception attention convolution sub-unit; The output end of the first windmill convolution sub-unit is connected to the input end of the second windmill convolution sub-unit; The output end of the second windmill convolution sub-unit is connected to the input end of the multi-region perception attention convolution sub-unit; The output of the residual bottleneck unit is formed by element-wise addition of the output of the multi-region perception attention convolution sub-unit and the external input features received by the first windmill convolution sub-unit.

8. The small target detection method for remote sensing images according to claim 7, characterized in that, The training of the small target detection model includes: Obtaining a remote sensing image small target detection data set; wherein, the remote sensing image small target detection data set includes a number of remote sensing images and corresponding first target reference tensors, second target reference tensors, and third target reference tensors for each of the remote sensing images; the first target reference tensor, the second target reference tensor, and the third target reference tensor are generated by performing preset target assignment processing and parameter encoding processing according to the original true target bounding boxes and class labels of the corresponding remote sensing images; Randomly dividing the remote sensing image small target detection data set into several batches of training samples according to a preset batch size; Sequentially inputting each batch of training samples into the small target detection model, and performing iterative training on the small target detection model until a preset number of training rounds is reached; wherein, when the small target detection model receives each batch of training samples, it outputs a first prediction tensor, a second prediction tensor, and a third prediction tensor corresponding to the remote sensing images of that batch; through a preset scale-aware dynamic loss function, according to the first prediction tensor, the second prediction tensor, the third prediction tensor, the first target reference tensor, the second target reference tensor, and the third target reference tensor, calculating and generating a loss function value; using a preset optimizer to update the small target detection model according to the loss function value.

9. The remote sensing image small target detection method according to claim 8, wherein The scale-aware dynamic loss function is specifically: L SBDL = ω IoU ·L IoU + ω mask ·L mask Where, L SBDL is the scale-aware dynamic loss function; L IoU is the intersection over union loss function; ω IoU is the weight of the intersection over union loss function; L mask is the masked loss function; ω mask is the weight of the masked loss function.

10. A small target detection device for remote sensing images, characterized in that Including: A remote sensing image acquisition module, a small target detection model solution module, and a detection target list generation module; The remote sensing image acquisition module is used to acquire a remote sensing image to be detected; The small target detection model solving module is used to input the remote sensing image into the small target detection model, so that the small target detection model generates a first tensor, a second tensor, and a third tensor according to the remote sensing image; wherein, the small target detection model includes a backbone network, a neck network, and a detection head network; the backbone network includes at least four successively connected backbone module groups, the output of each previous backbone module group is the input of the subsequent backbone module group, and the spatial scale of the feature map decreases along the connection direction; at least one of the backbone module groups includes a windmill convolution module, and at least one of the backbone module groups includes an adaptive receptive field spatial attention module; the backbone network outputs at least four-way scale feature maps to the neck network; the neck network includes at least three successively connected neck module groups, the output of each previous neck module group is the input of the subsequent neck module group, and each neck module group also fuses at least one-way scale feature map as another input, and the spatial scale of the feature map increases along the connection direction; the neck network outputs at least three enhanced feature maps to the detection head network; the enhanced feature maps include a first enhanced feature map, a second enhanced feature map, and a third enhanced feature map, and the resolution of the first enhanced feature map is greater than that of the second enhanced feature map, and the resolution of the second enhanced feature map is greater than that of the third enhanced feature map; the detection head network includes three prediction paths respectively for receiving the first enhanced feature map, the second enhanced feature map, and the third enhanced feature map, and each prediction path is used to generate the first tensor, the second tensor, and the third tensor according to the corresponding enhanced feature map; The detection target list generation module is used to generate a detection target information list according to the first tensor, the second tensor, and the third tensor.

Citation Information

Cited By

  • Aerial photography small target detection method combining windmill convolution and weighted multi-branch fusion

    CN120510540A

  • Small target detection method in aerial photography by combining windmill convolution with weighted multi-branch fusion

    CN120510540B

  • Small target detection method and system, and computer equipment

    CN121280706A

  • Small target detection method and system, and computer device

    CN121280706B

  • Target detection method based on dynamic weight and hierarchical query selection strategy

    CN121482373A