A two-stage cooperative remote sensing small target detection method and system

CN118521913BActive Publication Date: 2026-09-25WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410742361.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-11
Publication Date
2026-09-25
Estimated Expiration
2044-06-11

AI Technical Summary

Technical Problem

然而,小目标由于其本身像素数量的限制,导致其外观特征信息极度有限,具有辨别性的细节特征往往集中在最底层,不同层的特征图包含的小目标信息极不平衡,这给深度学习模型对小目标的特征学习带来了挑战

Benefits of technology

[0100]本发明在神经网络前向传播阶段使用基于通道注意力的重加权方法,在训练过程中可以在极大保留小目标细节信息的同时,自适应地融合低层和高层特征,缓解小目标不同尺度特征不平衡问题;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118521913B_ABST
    Figure CN118521913B_ABST
Patent Text Reader

Abstract

The application provides a two-stage cooperative remote sensing small target detection method and system. A plurality of three-channel optical remote sensing images with small target bounding box and small target category are collected; a remote sensing small target detection network is constructed; the images are input into the network in batches for small target detection; a loss function model is constructed in combination with the small target bounding box and the corresponding small target category of each small target in the image; the updated remote sensing small target detection network is obtained through batch-by-batch iterative optimization training; and the updated detection network is used for remote sensing small target detection to obtain the small target bounding box and the corresponding small target category of the three-channel optical remote sensing image. The application can effectively utilize the bottommost feature map, adaptively reweight the attention of the detail information and the semantic information, and better learn the characteristics of the extremely small target by balancing the attention degree of the deep neural network to different samples, so that the detection precision of the remote sensing small target is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision, and in particular relates to a two-stage collaborative remote sensing small target detection method and system. Background Technology

[0002] Remote sensing small target detection technology plays a crucial role in aerial image analysis, with irreplaceable importance for applications such as refined urban planning, efficient urban management, timely urban security monitoring, accurate population estimation, and rapid disaster relief. This technology can identify and locate targets as small as a few pixels in an image, providing valuable data support for decision-makers. Currently, deep learning-based target detection algorithms, with their ability to learn from massive amounts of image data, have become the mainstream technique for improving the accuracy of remote sensing small target detection. However, due to the limited number of pixels, small targets have extremely limited appearance feature information. Distinguishing details are often concentrated at the lowest layers, and the feature maps of different layers contain highly imbalanced information about small targets, posing a challenge to deep learning models in learning the features of small targets. Furthermore, existing small target detection datasets exhibit significant variations in target size, leading to a significant scale imbalance in the positive sample allocation process for small targets of different scales. Larger-scale targets tend to receive more positive samples, while smaller-scale targets receive fewer. This causes deep learning networks to tend to optimize the detection of larger targets during training, exacerbating the problem of missed detections of extremely small targets and thus affecting the overall performance of the detection system. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this invention provides a two-stage collaborative remote sensing small target detection method and system. This method can effectively utilize the lowest-level feature map, adaptively reweight the attention to detail information and semantic information, and better learn the features of extremely small targets by balancing the importance of deep neural networks to different samples, thereby further improving the detection accuracy of remote sensing small targets.

[0004] The method of this invention is a two-stage collaborative remote sensing small target detection method, comprising the following steps:

[0005] Step 1: Acquire multiple three-channel optical remote sensing images. Randomly sample each batch of remote sensing images from the multiple three-channel optical remote sensing images. Label the bounding rectangles of all small targets in all three-channel optical remote sensing images of each batch of remote sensing images and the category of the small targets in the bounding rectangles of all small targets.

[0006] Step 2: Construct a remote sensing small target detection network. Input each three-channel optical remote sensing image sample from each batch of remote sensing images into the remote sensing small target detection network for small target detection. Obtain the bounding rectangle of each predicted small target and the predicted small target category of each bounding rectangle in each three-channel optical remote sensing image sample from each batch of remote sensing images. Combine the bounding rectangle of each small target and the small target category of each bounding rectangle in the corresponding three-channel optical remote sensing image of each batch of remote sensing images to construct a loss function model. Iterate and optimize the training batch by batch to obtain the updated remote sensing small target detection network.

[0007] Step 3: Acquire multiple real-time three-channel optical remote sensing images, construct real-time batch remote sensing image samples, input the updated remote sensing small target detection network to perform small target detection, and obtain the bounding rectangle of each predicted small target in each real-time three-channel optical remote sensing image sample, and the predicted small target category of each predicted small target bounding rectangle.

[0008] As a preferred embodiment, step 1 involves randomly sampling each batch of remote sensing image samples from multiple three-channel optical remote sensing images, as detailed below:

[0009] In K sum In the three-channel optical remote sensing images, the _iters_th random sampling is:

[0010] In K sum K three-channel optical remote sensing images are randomly selected from -(iters-1)*K images as the iters-th batch of remote sensing image samples, where iters∈[1,S]. The images extracted in the iterth iteration will be trained in the iterth iteration. S represents the total number of batches or the total number of iterations. Except for the last batch, where the number of images is determined by the number of remaining images, the number of images K in each batch is 4.

[0011] Each batch of remote sensing image samples contains K images;

[0012] The bounding rectangle of each small target in all three-channel optical remote sensing images of each batch of remote sensing image samples described in step 1 is defined as follows:

[0013]

[0014] Among them, Box n This represents the bounding rectangle of the nth small target in all three-channel optical remote sensing images of each batch of remote sensing image samples. This represents the x-coordinate of the top-left corner of the bounding rectangle of the nth small target in all three-channel optical remote sensing images of each batch of remote sensing image samples. This represents the top-left ordinate of the bounding rectangle of the nth small target in all three-channel optical remote sensing images of each batch of remote sensing image samples. This represents the x-coordinate of the lower right corner of the bounding rectangle of the nth small target in all three-channel optical remote sensing images of each batch of remote sensing image samples. The ordinate represents the lower right corner of the bounding rectangle of the nth small target in all three-channel optical remote sensing images of each batch of remote sensing image samples; N represents the number of small targets in all three-channel optical remote sensing images of each batch of remote sensing image samples.

[0015] The small target category in all three-channel optical remote sensing images of the remote sensing image sample described in step 1, with bounding rectangles for all small targets, is defined as follows:

[0016] label n ,label n ∈[0,C-1]

[0017] Where, label n denoted by , represents the small target category of the bounding rectangle of the nth small target in all three-channel optical remote sensing images of each batch of remote sensing image samples, and C represents the total number of small target categories in the dataset;

[0018] Preferably, the remote sensing small target detection network in step 2 includes:

[0019] Feature extraction backbone network, feature fusion network, detection head;

[0020] The feature extraction backbone network, feature fusion network, and detection head are cascaded in sequence.

[0021] The feature extraction backbone network adopts the ResNet50 network, and the batch normalization (BN) layer in the ResNet50 network is frozen, that is, the parameter settings of the batch normalization (BN) layer in the ResNet50 network are fixed.

[0022] The feature extraction backbone network is used to input each batch of remote sensing image samples, extract the feature maps of each batch of remote sensing image samples to obtain the first feature map, second feature map, third feature map and fourth feature map of the three-channel optical remote sensing image in each batch of remote sensing image samples, and output the first feature map, second feature map, third feature map and fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples to the feature fusion network;

[0023] The feature fusion network includes:

[0024] First convolution module, second convolution module, third convolution module, fourth convolution module, first addition module, second addition module, third addition module, first upsampling module, second upsampling module, third upsampling module, first low-level feature enhancement module, second low-level feature enhancement module, third low-level feature enhancement module, first concatenation module, second concatenation module, third concatenation module, first adaptive feature selection module, second adaptive feature selection module, third adaptive feature selection module, fourth addition module, fifth addition module, sixth addition module, fifth convolution module, sixth convolution module;

[0025] The first, second, third, and fourth convolutional modules are all convolutional layers with a kernel size of 1×1 and a stride of 1.

[0026] The first convolution module is used to input the first feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to extract the features of the first feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples after unifying the number of channels through convolution feature extraction, and output it to the first addition module;

[0027] The second convolution module is used to input the second feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to extract the features of the second feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples after unifying the number of channels through convolution feature extraction, and output it to the second addition module;

[0028] The third convolution module is used to input the third feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to extract the features of the third feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples after unifying the number of channels through convolution feature extraction, and output it to the third addition module.

[0029] The fourth convolution module is used to input the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to extract the features of the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples after unifying the number of channels through convolution feature extraction, and output them to the third upsampling module, the third stitching module and the sixth addition module respectively.

[0030] The first addition module, the second addition module, and the third addition module all employ element-by-element addition operations;

[0031] The first upsampling module, the second upsampling module, and the third upsampling module all use nearest neighbor interpolation to perform upsampling operations on the features;

[0032] The third upsampling module is used to input the one-dimensional features of the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to extract the upsampled features of the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples through feature upsampling, and output them to the third addition module.

[0033] The third addition module is used to input the one-dimensional features of the third feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples and the upsampled features of the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through addition calculation, the first fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples is obtained, and then output to the second upsampling module, the second stitching module and the fifth addition module respectively.

[0034] The second upsampling module is used to input the first fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, extract the first upsampled fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples through feature upsampling, and output it to the second addition module;

[0035] The second addition module is used to input the first upsampling fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples and the one-dimensional feature of the second feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples, perform addition calculations to obtain the second fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and output them to the first upsampling module, the first stitching module and the fourth addition module respectively.

[0036] The first upsampling module is used to input the second fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to obtain the second upsampled fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples through feature upsampling extraction, and output it to the first addition module;

[0037] The first addition module is used to input the second upsampling fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples and the one-dimensional feature of the first feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through addition calculation, the third fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples is obtained and output to the first low-level feature enhancement module.

[0038] The first, second, and third bottom-level feature enhancement modules all employ average pooling with a window size of 2 and a step size of 2.

[0039] The first splicing module, the second splicing module, and the third splicing module all employ channel-based splicing operations.

[0040] The first adaptive feature selection module, the second adaptive feature selection module, and the third adaptive feature selection module are all composed of global average pooling, a first adaptive convolutional layer, a sigmoid activation function, a hadamard product module, an element-wise addition module, and a second adaptive convolutional layer cascaded together in sequence.

[0041] The fourth, fifth, and sixth addition modules all employ element-by-element addition operations.

[0042] The first low-level feature enhancement module is used to input the third fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and through low-level feature enhancement processing, obtain the third low-level enhanced feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and output it to the first stitching module.

[0043] The first stitching module is used to input the third bottom-level enhanced features of each three-channel optical remote sensing image of each batch of remote sensing image samples and the second fusion features of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through stitching processing, the first stitching features of each three-channel optical remote sensing image of each batch of remote sensing image samples are obtained and output to the first adaptive feature selection module.

[0044] The first adaptive feature selection module is used to input the first stitched feature of each three-channel optical remote sensing image of each batch of remote sensing image samples. The feature is then processed by the global average pooling layer, the first adaptive convolutional layer, and the Sigmoid activation function of the first adaptive feature selection module to obtain the fusion weight of each channel. The fusion weight of each channel is multiplied by the first stitched feature by the Hadamard product module to obtain the intermediate feature. The intermediate feature and the first stitched feature are then added element-wise to obtain the fusion feature. The fusion feature is processed by the second adaptive convolutional layer to obtain the first adaptive feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and then output to the fourth addition module.

[0045] The fourth addition module is used to input the first adaptive feature of each three-channel optical remote sensing image of each batch of remote sensing image samples and the second fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through addition calculation, the fourth fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples is obtained and output to the second low-level feature enhancement module and the detection head respectively.

[0046] The second low-level feature enhancement module is used to input the fourth fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and through low-level feature enhancement processing, obtain the fourth low-level enhanced feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and output it to the second stitching module.

[0047] The second stitching module is used to input the fourth bottom-level enhanced features of each three-channel optical remote sensing image of each batch of remote sensing image samples and the first fusion features of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through stitching processing, the second stitching features of each three-channel optical remote sensing image of each batch of remote sensing image samples are obtained and output to the second adaptive feature selection module.

[0048] The second adaptive feature selection module is used to input the second stitching feature of each three-channel optical remote sensing image of each batch of remote sensing image samples. The feature is then processed by the global average pooling layer, the first adaptive convolutional layer, and the Sigmoid activation function to obtain the fusion weights of each channel. The fusion weights of each channel are multiplied by the second stitching feature through the Hadamard product module to obtain the intermediate feature. The intermediate feature and the second stitching feature are then added element-wise to obtain the fusion feature. The fusion feature is processed by the second adaptive convolutional layer to obtain the second adaptive feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and then output to the fifth addition module.

[0049] The fifth addition module is used to input the second adaptive feature of each three-channel optical remote sensing image of each batch of remote sensing image samples and the first fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through addition calculation, the fifth fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples is obtained and output to the third low-level feature enhancement module and the detection head respectively.

[0050] The third low-level feature enhancement module is used to input the fifth fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and through low-level feature enhancement processing, obtain the fifth low-level enhanced feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and output it to the third stitching module.

[0051] The third stitching module is used to input the fifth bottom-level enhanced features of each three-channel optical remote sensing image of each batch of remote sensing image samples and the features of the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples after unifying the number of channels. Through stitching processing, the third stitching features of each three-channel optical remote sensing image of each batch of remote sensing image samples are obtained and output to the first adaptive feature selection module.

[0052] The third adaptive feature selection module is used to input the third stitching feature of each three-channel optical remote sensing image of each batch of remote sensing image samples. The feature is then processed by the global average pooling layer, the first adaptive convolutional layer, and the Sigmoid activation function to obtain the fusion weights of each channel. The fusion weights of each channel are multiplied by the third stitching feature through the Hadamard product module to obtain the intermediate feature. The intermediate feature and the third stitching feature are then added element-wise to obtain the fusion feature. The fusion feature is processed by the second adaptive convolutional layer to obtain the third adaptive feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and then output to the sixth addition module.

[0053] The sixth addition module is used to input the third adaptive feature of each three-channel optical remote sensing image of each batch of remote sensing image samples and the features of the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples after unifying the number of channels. Through addition calculation, the sixth fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples is obtained and output to the fifth convolution module and the detection head respectively.

[0054] The fifth and sixth convolutional modules are both convolutional layers with a kernel size of 3×3 and a stride of 2.

[0055] The fifth convolution module is used to input the sixth fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to extract the seventh fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples through convolution feature extraction, and output them to the sixth convolution module and the detection head respectively.

[0056] The sixth convolution module is used to input the seventh fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, extract the eighth feature of each three-channel optical remote sensing image of each batch of remote sensing image samples through convolution feature extraction, and output it to the detection head;

[0057] The detection head includes: a regression branch module, a classification branch module, a label allocation module, and a post-processing module;

[0058] The detection head is used as input for the fourth, fifth, sixth, seventh, and eighth fusion features of each three-channel optical remote sensing image sample in each batch of remote sensing images. Each fusion feature is passed through the detection head with shared weights and the regression branch module of the detection head for regression prediction to obtain multiple predicted bounding rectangles in each three-channel optical remote sensing image sample in each batch of remote sensing images. The classification branch module of the detection head is then used for classification prediction to obtain the category of each small target corresponding to the multiple predicted bounding rectangles in each three-channel optical remote sensing image sample in each batch of remote sensing images. During the training phase, positive and negative samples are constructed and the loss function is calculated through the label allocation module to optimize the model. During the testing phase, the post-processing module obtains the bounding rectangle of each detected small target in each three-channel optical remote sensing image sample in each batch of remote sensing images, as well as the category of the small target corresponding to the bounding rectangle of each detected small target in each three-channel optical remote sensing image sample in each batch of remote sensing images.

[0059] The regression branch module consists of four cascaded regression convolutional layers and one cascaded regression convolutional layer;

[0060] The four cascaded regression convolutional layers are used to input the fourth, fifth, sixth, seventh, and eighth fusion features of each three-channel optical remote sensing image of each batch of remote sensing image samples. The regression features are extracted through the four cascaded regression convolutional layers to obtain the regression features of each three-channel optical remote sensing image of each batch of remote sensing image samples, and then output to the regression convolutional layers.

[0061] The regression convolutional layer is used for the input regression features of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through regression convolution operation, multiple predicted bounding rectangles are obtained in each three-channel optical remote sensing image of each batch of remote sensing image samples. The bounding rectangles are output to the label assignment module during the training phase and to the post-processing module during the testing phase.

[0062] The classification branch module consists of four cascaded classification convolutional layers and one classification convolutional layer.

[0063] The four cascaded classification convolutional layers are used to input the fourth, fifth, sixth, seventh, and eighth fusion features of each three-channel optical remote sensing image of each batch of remote sensing image samples. The classification features are extracted through the four cascaded classification convolutional layers to obtain the classification features of each three-channel optical remote sensing image of each batch of remote sensing image samples, and then output to the classification convolutional layers.

[0064] The classification convolutional layer is used for the input classification features of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through the classification convolution operation, the small target category corresponding to multiple predicted bounding rectangles in each three-channel optical remote sensing image of each batch of remote sensing image samples is obtained. The data is output to the label assignment module during the training phase and to the post-processing module during the testing phase.

[0065] The label allocation module is used to assign positive samples to the detected small targets in each three-channel optical remote sensing image of each batch of remote sensing image samples. The label allocation method based on point prior of FCOS is used to obtain the positive samples of small targets in all three-channel optical remote sensing images of each batch of remote sensing image samples, as well as the remaining negative samples. The prediction results corresponding to the positive samples are used to calculate the regression loss and classification loss, while the prediction results of the negative samples are only used to calculate the classification loss.

[0066] The post-processing module is used to input regression branch and classification branch predictions to obtain multiple predicted bounding rectangles and their corresponding small target categories in each three-channel optical remote sensing image of each batch of remote sensing image samples. The NMS operation is used for post-processing to obtain the bounding rectangles and their corresponding small target categories in each three-channel optical remote sensing image of each batch of remote sensing image samples.

[0067] The loss function model described in step 2 is defined as follows:

[0068] L sum =L reg +L cls

[0069] Among them, L reg L represents the regression loss. cls Indicates classification loss;

[0070] The regression loss is defined as follows:

[0071]

[0072] Where, N pos z represents the total number of positive samples of all small targets in all three-channel optical remote sensing images of this batch of remote sensing image samples. + This represents the z-th image among all three-channel optical remote sensing images in this batch of remote sensing image samples. + The serial number of each positive sample;

[0073] Indicate z + The bounding rectangle of the predicted small target.

[0074]

[0075] in, The z-th image of each batch of remote sensing image samples represents the z-th of all three-channel optical remote sensing images. + The x-coordinate of the top-left corner, the y-coordinate of the top-left corner, the x-coordinate of the bottom-right corner, and the y-coordinate of the bottom-right corner of the bounding rectangle of the small target predicted by each positive sample. Represents positive sample z + The regression target is the index of all small targets in all three-channel optical remote sensing images of this batch of remote sensing image samples. The small goal is the zth... + The regression objective for each sample point (equivalent to z) + The first in this batch (Positive samples of small goals);

[0076] This represents the first of all three-channel optical remote sensing images in each batch of remote sensing image samples. The bounding rectangle of a small target.

[0077] in, These are the images from all three-channel optical remote sensing images in this batch of remote sensing image samples. The x-coordinate of the top left corner, the y-coordinate of the top left corner, the x-coordinate of the bottom right corner, and the y-coordinate of the bottom right corner of the bounding rectangle of the two small targets; DIoU(·,·) represents the DIoU value between the bounding rectangles of the two small targets;

[0078] The classification loss is defined as follows:

[0079]

[0080] Where α represents reducing the loss weight of negative samples, γ represents reducing the loss weight of simple samples, and N pos Z represents the total number of positive samples of all small targets in all three-channel optical remote sensing images of this batch of remote sensing image samples. sum This represents the total number of positive and negative samples of all small targets in all three-channel optical remote sensing images of this batch of remote sensing image samples, where z represents the sequence number of the z-th sample in all three-channel optical remote sensing images of this batch of remote sensing image samples. It is the predicted score of sample point z. C represents the number of small object categories. This indicates that sample point z predicts the nth point among all three-channel optical remote sensing images in this batch of remote sensing image samples. z The probability that a small target belongs to category c. This represents the target score for sample point z. This indicates that sample point z predicts the nth point among all three-channel optical remote sensing images in this batch of remote sensing image samples. zThe probability of a target of category c, where c∈[1,C], when z is the nth target among all three-channel optical remote sensing images in this batch of remote sensing image samples. z The nth positive sample of a small target, and the nth of all three-channel optical remote sensing images in this batch of remote sensing image samples. z The target categories for each small goal are: but Where iters represents the number of iterations. These represent the bounding rectangle of the small target predicted by z and the nth target in all three-channel optical remote sensing images of this batch of remote sensing image samples, respectively. z The bounding rectangle of a small target, if z is a negative sample, then

[0081] α m The amplitude balancing weighting factor is defined as follows:

[0082]

[0083] in, N represents the number of small targets in all three-channel optical remote sensing images of this batch of remote sensing image samples, z + This represents the z-th image among all three-channel optical remote sensing images in this batch of remote sensing image samples. + The index of each positive sample, z + The regression target is the first of all three-channel optical remote sensing images in this batch of remote sensing image samples. A small goal This refers to the first of all three-channel optical remote sensing images in this batch of remote sensing image samples. The number of positive samples assigned to each small objective, and the magnitude weighting factor α. m It is inversely proportional to the number of positive samples and directly proportional to the average number of samples, which makes the network pay more attention to small targets with few positive samples;

[0084] α s The quality balance weighting factor is defined as follows:

[0085]

[0086] Where s is the th image among all three-channel optical remote sensing images in this batch of remote sensing image samples. The sequence number of the positive sample of each small target. The z-th image in this batch of remote sensing image samples is the one with the three-channel optical remote sensing capabilities. + The positive sample in the th... All of the small goals ( The relative quality among positive samples;

[0087]

[0088] in, The overall quality of this positive sample is determined by a combination of its classification and regression scores. Indicates the next batch The classification score predicted for the s-th positive sample of each small target, while using... and These represent the next [number] in this batch. The bounding rectangle of a small target and the bounding rectangle of the small target predicted by the s-th positive sample of the target;

[0089]

[0090] in, They are respectively the batch of the first The x-coordinate of the top-left corner, the y-coordinate of the top-left corner, the x-coordinate of the bottom-right corner, and the y-coordinate of the bottom-right corner of the bounding rectangle of the small target. Indicates the next [number] in this batch The x-coordinate of the top-left corner, y-coordinate of the top-left corner, and y-coordinate of the bottom-right corner of the bounding rectangle of the predicted small target are used to measure the normalized KL divergence distance between the Gaussian distributions of the labeled box and the predicted box. Will As a regression score, the overall quality can be expressed as:

[0091]

[0092] Where λ = 0.2 represents the weight of the classification score in the measurement of sample quality; This refers to the first of all three-channel optical remote sensing images in this batch of remote sensing image samples. The classification prediction score of the s-th positive sample of each small target The calculated Varifocal Loss

[0093]

[0094] Where iters represents the number of iterations, and the quality balance weight factor α s It increases with the relative quality of positive samples, enhancing the network's attention to high-quality positive samples and promoting feature learning for small targets;

[0095] This invention also provides a two-stage collaborative remote sensing small target detection system, as detailed below:

[0096] The sample labeling module is used to acquire multiple three-channel optical remote sensing images. It randomly selects each batch of remote sensing image samples from multiple three-channel optical remote sensing images and labels the bounding rectangles of all small targets and the small target categories of all small target bounding rectangles in all three-channel optical remote sensing images of each batch of remote sensing image samples.

[0097] The remote sensing small target detection network training module is used to construct the remote sensing small target detection network. Each batch of remote sensing image samples and each three-channel optical remote sensing image are sequentially input into the remote sensing small target detection network for small target detection. The module obtains the bounding rectangle of each predicted small target and the predicted small target category of each bounding rectangle in each batch of remote sensing image samples. A loss function model is constructed by combining the bounding rectangle of each small target and the small target category of each bounding rectangle in the corresponding three-channel optical remote sensing image of each batch of remote sensing image samples. The updated remote sensing small target detection network is obtained through iterative optimization training batch by batch.

[0098] The real-time batch remote sensing image sample detection module is used to acquire multiple real-time three-channel optical remote sensing images, construct real-time batch remote sensing image samples, input the updated remote sensing small target detection network to perform small target detection, and obtain the bounding rectangle of each predicted small target in each real-time three-channel optical remote sensing image of the real-time batch remote sensing image sample, and the predicted small target category of each predicted small target bounding rectangle.

[0099] The beneficial effects of this invention are as follows:

[0100] This invention uses a channel attention-based reweighting method in the forward propagation stage of the neural network. During training, it can adaptively fuse low-level and high-level features while preserving a great deal of detail information of small targets, thus alleviating the problem of feature imbalance at different scales of small targets.

[0101] In the backpropagation stage, this invention calculates the number of small target samples and their relative quality, and adjusts the loss for small targets of different scales and the loss for positive samples of different qualities accordingly. This enhances the neural network's ability to learn from extremely small target samples. The two stages are optimized in tandem, which can improve the model's detection performance. Attached Figure Description

[0102] Figure 1 : Flowchart of the method according to an embodiment of the present invention.

[0103] Figure 2 : A diagram of a deep convolutional neural network model according to an embodiment of the present invention.

[0104] Figure 3 : Schematic diagram of the detection results of an embodiment of the present invention.

[0105] Figure 4 : A schematic diagram of the underlying classification feature map of an embodiment of the present invention. Detailed Implementation

[0106] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0107] In specific implementation, the method proposed in the technical solution of this invention can be automatically executed by those skilled in the art using computer software technology. System devices for implementing the method, such as computer-readable storage media storing the corresponding computer program of the technical solution of this invention and computer equipment including the computer program running the corresponding computer program, should also be within the protection scope of this invention.

[0108] The following is combined with Figure 1-4 This invention introduces a specific implementation method for remote sensing small target detection using a two-stage collaborative approach, comprising the following steps:

[0109] like Figure 1 The diagram shown is a flowchart of a method according to an embodiment of the present invention.

[0110] Step 1: Acquire multiple three-channel optical remote sensing images. Randomly sample each batch of remote sensing images from these multiple three-channel optical remote sensing images. Label the bounding rectangles of all small targets in all three-channel optical remote sensing images of each batch of remote sensing images, and label the category of each small target within those bounding rectangles.

[0111] Step 1 involves randomly sampling multiple batches of remote sensing image samples from multiple three-channel optical remote sensing images, as detailed below:

[0112] In K sum In the three-channel optical remote sensing images, the _iters_th random sampling is:

[0113] In K sum K = 4 three-channel optical remote sensing images are randomly selected from -(iters-1)*K three-channel optical remote sensing images as the iters-th batch of remote sensing image samples, where iters∈[1,S]. The images extracted in the iterth iteration will be trained in the iterth iteration. S represents the total number of batches or the total number of iterations. Except for the last batch, where the number of images is determined by the number of remaining images, the number of images K in each batch is 4.

[0114] Each batch of remote sensing image samples contains K images;

[0115] The bounding rectangle of each small target in all three-channel optical remote sensing images of each batch of remote sensing image samples described in step 1 is defined as follows:

[0116]

[0117] Among them, Box n This represents the bounding rectangle of the nth small target in all three-channel optical remote sensing images of each batch of remote sensing image samples. This represents the x-coordinate of the top-left corner of the bounding rectangle of the nth small target in all three-channel optical remote sensing images of each batch of remote sensing image samples. This represents the top-left ordinate of the bounding rectangle of the nth small target in all three-channel optical remote sensing images of each batch of remote sensing image samples. This represents the x-coordinate of the lower right corner of the bounding rectangle of the nth small target in all three-channel optical remote sensing images of each batch of remote sensing image samples. The ordinate represents the lower right corner of the bounding rectangle of the nth small target in all three-channel optical remote sensing images of each batch of remote sensing image samples; N represents the number of small targets in all three-channel optical remote sensing images of each batch of remote sensing image samples.

[0118] The small target category in all three-channel optical remote sensing images of the remote sensing image sample described in step 1, with bounding rectangles for all small targets, is defined as follows:

[0119] label n ,label n ∈[0,C-1]

[0120] Where, label n denoted by , represents the small target category of the bounding rectangle of the nth small target in all three-channel optical remote sensing images of each batch of remote sensing image samples, and C represents the total number of small target categories in the dataset;

[0121] Step 2: Construct a remote sensing small target detection network. Input each three-channel optical remote sensing image sample from each batch of remote sensing images into the remote sensing small target detection network for small target detection. Obtain the bounding rectangle of each predicted small target and the predicted small target category of each bounding rectangle in each three-channel optical remote sensing image sample from each batch of remote sensing images. Combine the bounding rectangle of each small target and the small target category of each bounding rectangle in the corresponding three-channel optical remote sensing image of each batch of remote sensing images to construct a loss function model. Iterate and optimize the training batch by batch to obtain the updated remote sensing small target detection network.

[0122] like Figure 2As shown, the remote sensing small target detection network described in step 2 includes:

[0123] Feature extraction backbone network, adaptive feature fusion network, detection head;

[0124] The feature extraction backbone network, adaptive feature fusion network, and detection head are cascaded in sequence.

[0125] The feature extraction backbone network adopts the ResNet50 network, and the batch normalization (BN) layer in the ResNet50 network is frozen, that is, the parameter settings of the batch normalization (BN) layer in the ResNet50 network are fixed.

[0126] The feature extraction backbone network is used to input each batch of remote sensing image samples, extract the feature maps of each batch of remote sensing image samples to obtain the first feature map, second feature map, third feature map and fourth feature map of the three-channel optical remote sensing image in each batch of remote sensing image samples, and output the first feature map, second feature map, third feature map and fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples to the adaptive feature fusion network;

[0127] The adaptive feature fusion network includes:

[0128] First convolution module, second convolution module, third convolution module, fourth convolution module, first addition module, second addition module, third addition module, first upsampling module, second upsampling module, third upsampling module, first low-level feature enhancement module, second low-level feature enhancement module, third low-level feature enhancement module, first concatenation module, second concatenation module, third concatenation module, first adaptive feature selection module, second adaptive feature selection module, third adaptive feature selection module, fourth addition module, fifth addition module, sixth addition module, fifth convolution module, sixth convolution module;

[0129] The first, second, third, and fourth convolutional modules are all convolutional layers with a kernel size of 1×1 and a stride of 1.

[0130] The first convolution module is used to input the first feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to extract the features of the first feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples after a unified number of channels C=256 through convolution feature extraction, and output it to the first addition module;

[0131] The second convolution module is used to input the second feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to extract the features of the second feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples with a unified number of channels C=256 through convolution feature extraction, and output it to the second addition module;

[0132] The third convolution module is used to input the third feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to extract the features of the third feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples with a unified number of channels C=256 through convolution feature extraction, and output them to the third addition module.

[0133] The fourth convolution module is used to input the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to extract the features of the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples with a unified number of channels C=256 through convolution feature extraction, and output them to the third upsampling module, the third stitching module and the sixth addition module respectively.

[0134] The first addition module, the second addition module, and the third addition module all employ element-by-element addition operations;

[0135] The first upsampling module, the second upsampling module, and the third upsampling module all use nearest neighbor interpolation with a scaling factor of 2 to perform upsampling operations on the features;

[0136] The third upsampling module is used to input the one-dimensional features of the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to extract the upsampled features of the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples through feature upsampling, and output them to the third addition module.

[0137] The third addition module is used to input the one-dimensional features of the third feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples and the upsampled features of the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through addition calculation, the first fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples is obtained, and then output to the second upsampling module, the second stitching module and the fifth addition module respectively.

[0138] The second upsampling module is used to input the first fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, extract the first upsampled fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples through feature upsampling, and output it to the second addition module;

[0139] The second addition module is used to input the first upsampling fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples and the one-dimensional feature of the second feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples, perform addition calculations to obtain the second fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and output them to the first upsampling module, the first stitching module and the fourth addition module respectively.

[0140] The first upsampling module is used to input the second fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to obtain the second upsampled fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples through feature upsampling extraction, and output it to the first addition module;

[0141] The first addition module is used to input the second upsampling fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples and the one-dimensional feature of the first feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through addition calculation, the third fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples is obtained and output to the first low-level feature enhancement module.

[0142] The first, second, and third bottom-level feature enhancement modules all employ average pooling with a window size of 2 and a step size of 2.

[0143] The first splicing module, the second splicing module, and the third splicing module all employ channel-based splicing operations.

[0144] The first adaptive feature selection module, the second adaptive feature selection module, and the third adaptive feature selection module are all composed of global average pooling, a first adaptive convolutional layer, a sigmoid activation function, a hadamard product module, an element-wise addition module, and a second adaptive convolutional layer cascaded together in sequence.

[0145] Both the first adaptive convolutional layer and the second adaptive convolutional layer are convolutional layers with a kernel size of 1×1 and a stride of 1.

[0146] The fourth, fifth, and sixth addition modules all employ element-by-element addition operations.

[0147] The first low-level feature enhancement module is used to input the third fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and through low-level feature enhancement processing, obtain the third low-level enhanced feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and output it to the first stitching module.

[0148] The first stitching module is used to input the third bottom-level enhanced features of each three-channel optical remote sensing image of each batch of remote sensing image samples and the second fusion features of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through stitching processing, the first stitching features of each three-channel optical remote sensing image of each batch of remote sensing image samples are obtained and output to the first adaptive feature selection module.

[0149] The first adaptive feature selection module is used to input the first stitched feature of each three-channel optical remote sensing image of each batch of remote sensing image samples. The feature is then processed by the global average pooling layer, the first adaptive convolutional layer, and the Sigmoid activation function of the first adaptive feature selection module to obtain the fusion weight of each channel. The fusion weight of each channel is multiplied by the first stitched feature by the Hadamard product module to obtain the intermediate feature. The intermediate feature and the first stitched feature are then added element-wise to obtain the fusion feature. The fusion feature is processed by the second adaptive convolutional layer to obtain the first adaptive feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and then output to the fourth addition module.

[0150] The fourth addition module is used to input the first adaptive feature of each three-channel optical remote sensing image of each batch of remote sensing image samples and the second fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through addition calculation, the fourth fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples is obtained and output to the second low-level feature enhancement module and the detection head respectively.

[0151] The second low-level feature enhancement module is used to input the fourth fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and through low-level feature enhancement processing, obtain the fourth low-level enhanced feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and output it to the second stitching module.

[0152] The second stitching module is used to input the fourth bottom-level enhanced features of each three-channel optical remote sensing image of each batch of remote sensing image samples and the first fusion features of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through stitching processing, the second stitching features of each three-channel optical remote sensing image of each batch of remote sensing image samples are obtained and output to the second adaptive feature selection module.

[0153] The second adaptive feature selection module is used to input the second stitching feature of each three-channel optical remote sensing image of each batch of remote sensing image samples. The feature is then processed by the global average pooling layer, the first adaptive convolutional layer, and the Sigmoid activation function to obtain the fusion weights of each channel. The fusion weights of each channel are multiplied by the second stitching feature through the Hadamard product module to obtain the intermediate feature. The intermediate feature and the second stitching feature are then added element-wise to obtain the fusion feature. The fusion feature is processed by the second adaptive convolutional layer to obtain the second adaptive feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and then output to the fifth addition module.

[0154] The fifth addition module is used to input the second adaptive feature of each three-channel optical remote sensing image of each batch of remote sensing image samples and the first fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through addition calculation, the fifth fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples is obtained and output to the third low-level feature enhancement module and the detection head respectively.

[0155] The third low-level feature enhancement module is used to input the fifth fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and through low-level feature enhancement processing, obtain the fifth low-level enhanced feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and output it to the third stitching module.

[0156] The third stitching module is used to input the fifth bottom-level enhanced features of each three-channel optical remote sensing image of each batch of remote sensing image samples and the features of the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples after unifying the number of channels. Through stitching processing, the third stitching features of each three-channel optical remote sensing image of each batch of remote sensing image samples are obtained and output to the first adaptive feature selection module.

[0157] The third adaptive feature selection module is used to input the third stitching feature of each three-channel optical remote sensing image of each batch of remote sensing image samples. The feature is then processed by the global average pooling layer, the first adaptive convolutional layer, and the Sigmoid activation function to obtain the fusion weights of each channel. The fusion weights of each channel are multiplied by the third stitching feature through the Hadamard product module to obtain the intermediate feature. The intermediate feature and the third stitching feature are then added element-wise to obtain the fusion feature. The fusion feature is processed by the second adaptive convolutional layer to obtain the third adaptive feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and then output to the sixth addition module.

[0158] The sixth addition module is used to input the third adaptive feature of each three-channel optical remote sensing image of each batch of remote sensing image samples and the features of the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples after unifying the number of channels. Through addition calculation, the sixth fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples is obtained and output to the fifth convolution module and the detection head respectively.

[0159] The fifth and sixth convolutional modules are both convolutional layers with a kernel size of 3×3 and a stride of 2.

[0160] The fifth convolution module is used to input the sixth fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to extract the seventh fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples through convolution feature extraction, and output them to the sixth convolution module and the detection head respectively.

[0161] The sixth convolution module is used to input the seventh fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, extract the eighth feature of each three-channel optical remote sensing image of each batch of remote sensing image samples through convolution feature extraction, and output it to the detection head;

[0162] The detection head includes: a regression branch module, a classification branch module, a label allocation module, and a post-processing module;

[0163] The detection head is used as input for the fourth, fifth, sixth, seventh, and eighth fusion features of each three-channel optical remote sensing image sample in each batch of remote sensing images. Each fusion feature is passed through the detection head with shared weights and the regression branch module of the detection head for regression prediction to obtain multiple predicted bounding rectangles in each three-channel optical remote sensing image sample in each batch of remote sensing images. The classification branch module of the detection head is then used for classification prediction to obtain the category of each small target corresponding to the multiple predicted bounding rectangles in each three-channel optical remote sensing image sample in each batch of remote sensing images. During the training phase, positive and negative samples are constructed and the loss function is calculated through the label allocation module to optimize the model. During the testing phase, the post-processing module obtains the bounding rectangle of each detected small target in each three-channel optical remote sensing image sample in each batch of remote sensing images, as well as the category of the small target corresponding to the bounding rectangle of each detected small target in each three-channel optical remote sensing image sample in each batch of remote sensing images.

[0164] The regression branch module consists of four cascaded regression convolutional layers and one cascaded regression convolutional layer;

[0165] The four cascaded regression convolutional layers are used to input the fourth, fifth, sixth, seventh, and eighth fusion features of each three-channel optical remote sensing image of each batch of remote sensing image samples. The regression features are extracted through the four cascaded regression convolutional layers to obtain the regression features of each three-channel optical remote sensing image of each batch of remote sensing image samples, and then output to the regression convolutional layers.

[0166] The regression convolutional layer is used for the input regression features of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through regression convolution operation, multiple predicted bounding rectangles are obtained in each three-channel optical remote sensing image of each batch of remote sensing image samples. The bounding rectangles are output to the label assignment module during the training phase and to the post-processing module during the testing phase.

[0167] The classification branch module consists of four cascaded classification convolutional layers and one classification convolutional layer.

[0168] The four cascaded classification convolutional layers are used to input the fourth, fifth, sixth, seventh, and eighth fusion features of each three-channel optical remote sensing image of each batch of remote sensing image samples. The classification features are extracted through the four cascaded classification convolutional layers to obtain the classification features of each three-channel optical remote sensing image of each batch of remote sensing image samples, and then output to the classification convolutional layers.

[0169] The classification convolutional layer is used for the input classification features of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through the classification convolution operation, the small target category corresponding to multiple predicted bounding rectangles in each three-channel optical remote sensing image of each batch of remote sensing image samples is obtained. The data is output to the label assignment module during the training phase and to the post-processing module during the testing phase.

[0170] The label allocation module is used to assign positive samples to the detected small targets in each three-channel optical remote sensing image of each batch of remote sensing image samples. The label allocation method based on point prior of FCOS is used to obtain the positive samples of small targets in all three-channel optical remote sensing images of each batch of remote sensing image samples, as well as the remaining negative samples. The prediction results corresponding to the positive samples are used to calculate the regression loss and classification loss, while the prediction results of the negative samples are only used to calculate the classification loss.

[0171] The post-processing module is used to input regression branch and classification branch predictions to obtain multiple predicted bounding rectangles and their corresponding small target categories in each three-channel optical remote sensing image of each batch of remote sensing image samples. The NMS operation is used for post-processing to obtain the bounding rectangles and their corresponding small target categories in each three-channel optical remote sensing image of each batch of remote sensing image samples.

[0172] The loss function model described in step 2 is defined as follows:

[0173] L sum =L reg +L cls

[0174] Among them, L reg L represents the regression loss. cls Indicates classification loss;

[0175] The regression loss is defined as follows:

[0176]

[0177] Where, N pos z represents the total number of positive samples of all small targets in all three-channel optical remote sensing images of this batch of remote sensing image samples. + This represents the z-th image among all three-channel optical remote sensing images in this batch of remote sensing image samples. + The serial number of each positive sample;

[0178] Indicate z + The bounding rectangle of the predicted small target.

[0179] in, The z-th image of each batch of remote sensing image samples represents the z-th of all three-channel optical remote sensing images. + The x-coordinate of the top-left corner, the y-coordinate of the top-left corner, the x-coordinate of the bottom-right corner, and the y-coordinate of the bottom-right corner of the bounding rectangle of the small target predicted by each positive sample. Represents positive sample z + The regression target is the index of all small targets in all three-channel optical remote sensing images of this batch of remote sensing image samples. The small goal is the zth... + The regression objective for each sample point (equivalent to z) + The first in this batch (Positive samples of small goals);

[0180] This represents the first of all three-channel optical remote sensing images in each batch of remote sensing image samples. The bounding rectangle of a small target.

[0181] in, These are the images from all three-channel optical remote sensing images in this batch of remote sensing image samples. The x-coordinate of the top left corner, the y-coordinate of the top left corner, the x-coordinate of the bottom right corner, and the y-coordinate of the bottom right corner of the bounding rectangle of the two small targets; DIoU(·,·) represents the DIoU size between the bounding rectangles of the two small targets;

[0182] The classification loss is defined as follows:

[0183]

[0184]

[0185] Where α = 0.25 represents reducing the loss weight of negative samples, γ = 2 represents reducing the loss weight of simple samples, and N pos Z represents the total number of positive samples of all small targets in all three-channel optical remote sensing images of this batch of remote sensing image samples. sum This represents the total number of positive and negative samples of all small targets in all three-channel optical remote sensing images of this batch of remote sensing image samples, where z represents the sequence number of the z-th sample in all three-channel optical remote sensing images of this batch of remote sensing image samples. It is the predicted score of sample point z. C represents the number of small object categories. This indicates that sample point z predicts the nth point among all three-channel optical remote sensing images in this batch of remote sensing image samples. z The probability that a small target belongs to category c. This represents the target score for sample point z. This indicates that sample point z predicts the nth point among all three-channel optical remote sensing images in this batch of remote sensing image samples. z The probability of a target of category c, where c∈[1,C], when z is the nth target among all three-channel optical remote sensing images in this batch of remote sensing image samples. z The nth positive sample of a small target, and the nth of all three-channel optical remote sensing images in this batch of remote sensing image samples. z The target categories for each small goal are: but Where iters represents the number of iterations. These represent the bounding rectangle of the small target predicted by z and the nth target in all three-channel optical remote sensing images of this batch of remote sensing image samples, respectively. z The bounding rectangle of a small target, if z is a negative sample, then

[0186] a m The amplitude balancing weighting factor is defined as follows:

[0187]

[0188] in, N represents the number of small targets in all three-channel optical remote sensing images of this batch of remote sensing image samples, z + This represents the z-th image among all three-channel optical remote sensing images in this batch of remote sensing image samples. + The index of each positive sample, z + The regression target is the first of all three-channel optical remote sensing images in this batch of remote sensing image samples. A small goal This refers to the first of all three-channel optical remote sensing images in this batch of remote sensing image samples. The number of positive samples assigned to each small objective, and the magnitude weighting factor α. m It is inversely proportional to the number of positive samples and directly proportional to the average number of samples, which makes the network pay more attention to small targets with few positive samples;

[0189] α s The quality balance weighting factor is defined as follows:

[0190]

[0191] Where s is the th image among all three-channel optical remote sensing images in this batch of remote sensing image samples. The sequence number of the positive sample of each small target. The z-th image in this batch of remote sensing image samples is the one with the three-channel optical remote sensing capabilities. + The positive sample in the th... All of the small goals ( The relative quality among positive samples;

[0192]

[0193] in, The overall quality of this positive sample is determined by a combination of its classification and regression scores. Indicates the next batch The classification score predicted for the s-th positive sample of each small target, while using... and These represent the next [number] in this batch. The bounding rectangle of a small target and the bounding rectangle of the small target predicted by the s-th positive sample of the target;

[0194]

[0195] in, They are respectively the batch of the first The x-coordinate of the top-left corner, the y-coordinate of the top-left corner, the x-coordinate of the bottom-right corner, and the y-coordinate of the bottom-right corner of the bounding rectangle of the small target. Indicates the next [number] in this batch The x-coordinate of the top-left corner, y-coordinate of the top-left corner, and y-coordinate of the bottom-right corner of the bounding rectangle of the predicted small target are used to measure the normalized KL divergence distance between the Gaussian distributions of the labeled box and the predicted box. Will As a regression score, the overall quality can be expressed as:

[0196]

[0197] Where λ = 0.2 represents the weight of the classification score in the measurement of sample quality; This refers to the first of all three-channel optical remote sensing images in this batch of remote sensing image samples. The classification prediction score of the s-th positive sample of each small target The calculated Varifocal Loss

[0198]

[0199] Where iters represents the number of iterations, and the quality balance weight factor α s It increases with the relative quality of positive samples, enhancing the network's attention to high-quality positive samples and promoting feature learning for small targets;

[0200] Step 3: Acquire multiple real-time three-channel optical remote sensing images, construct real-time batch remote sensing image samples, input the updated remote sensing small target detection network to perform small target detection, and obtain the bounding rectangle of each predicted small target in each real-time three-channel optical remote sensing image sample, and the predicted small target category of each predicted small target bounding rectangle.

[0201] This method was validated on the AI-TOD-v2 dataset. AI-TOD-v2 is a large-scale dataset for small target detection in high-resolution optical remote sensing images, containing 28,036 aerial images with 752,745 annotated target instances. Each image is 1024×1024 pixels in size. AI-TOD-v2 has eight common small target categories: aircraft (AI), bridges (BR), tanks (ST), boats (SH), swimming pools (SP), vehicles (VE), people (PE), and windmills (WM). The absolute size of targets in AI-TOD-v2 is mainly around 12 pixels, much smaller than other datasets. To analyze the scale of small objects more specifically, AI-TOD-v2 considers objects with an absolute size in the range of [2,8] as very small, [8,16] as small, [16,32] as small, and [32,64] as medium. In AI-TOD-v2, the proportions of very small, small, small, and medium-sized objects are 12.4%, 73.4%, 12.4%, and 1.8%, respectively, with most categories of objects falling within the very small range.

[0202] We visualized the detected bounding boxes to qualitatively evaluate the algorithm's performance, categorizing the detection boxes into true positive (TP), false negative (FN), and false positive (FP) detections, with a confidence threshold of 0.2. Furthermore, we used Average Precision (AP) as a quantitative metric to measure the algorithm's performance. 50 This indicates that the IoU threshold for TP is defined as 0.5, and AP 75 This indicates that the IoU threshold for TP is defined as 0.75, and AP represents AP. 50 ~AP 95 The average value is 0.05, with an IoU range of 0.05. Note the AP. 50 AP 75 AP considers objects at all scales. Furthermore, AP... vt AP t AP s and AP m Suitable for evaluating the detection performance of very small, tiny, small and medium-sized targets in AI-TOD-v2.

[0203] In terms of qualitative analysis, Figure 3 The results are visualized, with (a) showing the results of the baseline method and (b) showing the results of the method proposed in this invention. Green bounding boxes represent correct detections (TP), red bounding boxes represent false negatives (FN), and blue bounding boxes represent false positives (FP). Compared with the baseline method, the method proposed in this paper has significant advantages in both detection accuracy and localization precision. Figure 4The diagram shows a comparison between the adaptive feature fusion proposed in this invention and the traditional feature pyramid fusion. (a) is a visualization of the classification features after traditional feature pyramid fusion, and (b) is a visualization of the classification features after feature fusion proposed in this invention. It can be seen that the method proposed in this paper makes the small target features more obvious.

[0204] In terms of quantitative analysis, the experiment compared the FSANet and ESG_TODNet methods. Both are novel methods for small target detection in remote sensing and, like the baseline methods proposed in this paper, are single-stage FCOS detectors. The ESG_TODNet method is currently the state-of-the-art (SOTA) method for single-stage small target detection. Furthermore, other classic methods and other methods with similar ideas to this invention were also compared. The experimental results are shown in Table 1. It can be seen that among all the compared methods, the method proposed in this invention achieves the best results in both the comprehensive performance index (AP) and the small target performance index (AP). vt and AP t The proposed method outperforms other comparative methods in all aspects. Furthermore, when added to the ESG_TODNet method, the proposed method improves the above indicators by about 2 points. Qualitative and quantitative analysis of detection accuracy shows that the proposed method achieves a high level of accuracy on AI-TOD-v2. Compared with the benchmark algorithm and other detection algorithms, the average accuracy of the proposed algorithm has been significantly improved.

[0205] Table 1 Comparison of results from different methods

[0206]

[0207] This invention provides a two-stage collaborative remote sensing small target detection system, as detailed below:

[0208] The sample labeling module is used to acquire multiple three-channel optical remote sensing images. It randomly selects each batch of remote sensing image samples from multiple three-channel optical remote sensing images and labels the bounding rectangles of all small targets and the small target categories of all small target bounding rectangles in all three-channel optical remote sensing images of each batch of remote sensing image samples.

[0209] The remote sensing small target detection network training module is used to construct the remote sensing small target detection network. Each batch of remote sensing image samples and each three-channel optical remote sensing image are sequentially input into the remote sensing small target detection network for small target detection. The module obtains the bounding rectangle of each predicted small target and the predicted small target category of each bounding rectangle in each batch of remote sensing image samples. A loss function model is constructed by combining the bounding rectangle of each small target and the small target category of each bounding rectangle in the corresponding three-channel optical remote sensing image of each batch of remote sensing image samples. The updated remote sensing small target detection network is obtained through iterative optimization training batch by batch.

[0210] The real-time batch remote sensing image sample detection module is used to acquire multiple real-time three-channel optical remote sensing images, construct real-time batch remote sensing image samples, input the updated remote sensing small target detection network to perform small target detection, and obtain the bounding rectangle of each predicted small target in each real-time three-channel optical remote sensing image of the real-time batch remote sensing image sample, and the predicted small target category of each predicted small target bounding rectangle.

[0211] The sample label construction module, the remote sensing small target detection network training module, and the real-time batch remote sensing image sample detection module are all deployed on the server.

[0212] It should be understood that any parts not described in detail in this specification belong to the prior art.

[0213] It should be understood that the above description of the embodiments is quite detailed, but it should not be considered as a limitation on the scope of protection of this invention. Those skilled in the art can make substitutions or modifications under the guidance of this invention without departing from the scope of protection of the claims of this invention, and all such substitutions or modifications fall within the scope of protection of this invention. The scope of protection of this invention should be determined by the appended claims.

Claims

1. A two-stage collaborative remote sensing method for small target detection, characterized in that, Includes the following steps: Step 1: Acquire multiple three-channel optical remote sensing images. Randomly sample each batch of remote sensing images from the multiple three-channel optical remote sensing images. Label the bounding rectangles of all small targets in all three-channel optical remote sensing images of each batch of remote sensing images and the category of the small targets in the bounding rectangles of all small targets. Step 2: Construct a remote sensing small target detection network. Input each three-channel optical remote sensing image sample from each batch of remote sensing images into the remote sensing small target detection network for small target detection. Obtain the bounding rectangle of each predicted small target and the predicted small target category of each bounding rectangle in each three-channel optical remote sensing image sample from each batch of remote sensing images. Combine the bounding rectangle of each small target and the small target category of each bounding rectangle in the corresponding three-channel optical remote sensing image of each batch of remote sensing images to construct a loss function model. Iterate and optimize the training batch by batch to obtain the updated remote sensing small target detection network. Step 3: Acquire multiple real-time three-channel optical remote sensing images, construct real-time batch remote sensing image samples, input the updated remote sensing small target detection network to perform small target detection, and obtain the bounding rectangle of each predicted small target in each real-time three-channel optical remote sensing image sample and the predicted small target category of each predicted small target bounding rectangle. The remote sensing small target detection network described in step 2 includes: Feature extraction backbone network, feature fusion network, detection head; The feature extraction backbone network, feature fusion network, and detection head are cascaded in sequence. The feature fusion network includes a first adaptive feature selection module: The first adaptive feature selection module is used to input the first stitched feature of each three-channel optical remote sensing image of each batch of remote sensing image samples. The first feature selection module sequentially processes the global average pooling layer, the first adaptive convolutional layer, and the Sigmoid activation function to obtain the fusion weight of each channel. The fusion weight of each channel is multiplied by the first stitched feature through the Hadamard product module to obtain the intermediate feature. The intermediate feature and the first stitched feature are then added element-wise to obtain the fusion feature.

2. The dual-stage collaborative remote sensing small target detection method according to claim 1, characterized in that: Step 1 involves randomly sampling multiple batches of remote sensing image samples from multiple three-channel optical remote sensing images, as detailed below: exist In the three-channel optical remote sensing images, the _iters_th random sampling is: exist -(iters-1) K three-channel optical remote sensing images are randomly selected from K images as the iters-th batch of remote sensing image samples, where iters∈[1, S]. The images extracted in the iterth iteration will be trained in the iterth iteration. S represents the total number of batches or the total number of iterations. Except for the last batch, where the number of images is determined by the number of remaining images, the number of images K in each batch is 4. Each batch of remote sensing image samples contains K images; The bounding rectangle of each small target in all three-channel optical remote sensing images of each batch of remote sensing image samples described in step 1 is defined as follows: in, This represents the bounding rectangle of the nth small target in all three-channel optical remote sensing images of each batch of remote sensing image samples. This represents the x-coordinate of the top-left corner of the bounding rectangle of the nth small target in all three-channel optical remote sensing images of each batch of remote sensing image samples. This represents the top-left ordinate of the bounding rectangle of the nth small target in all three-channel optical remote sensing images of each batch of remote sensing image samples. This represents the x-coordinate of the lower right corner of the bounding rectangle of the nth small target in all three-channel optical remote sensing images of each batch of remote sensing image samples. This represents the ordinate of the lower right corner of the bounding rectangle of the nth small target in all three-channel optical remote sensing images of each batch of remote sensing image samples; This indicates the number of small targets in all three-channel optical remote sensing images of each batch of remote sensing image samples; The small target category in all three-channel optical remote sensing images of the remote sensing image sample described in step 1, with bounding rectangles for all small targets, is defined as follows: , ∈[0,C-1] in, Let C represent the small target category of the bounding rectangle of the nth small target in all three-channel optical remote sensing images of each batch of remote sensing image samples, and let C represent the total number of small target categories in the dataset.

3. The dual-stage collaborative remote sensing small target detection method according to claim 2, characterized in that: The feature extraction backbone network adopts the ResNet50 network, and the batch normalization (BN) layer in the ResNet50 network is frozen, that is, the parameter settings of the batch normalization (BN) layer in the ResNet50 network are fixed. The feature extraction backbone network is used to input each batch of remote sensing image samples, extract the feature maps of each batch of remote sensing image samples to obtain the first feature map, second feature map, third feature map and fourth feature map of the three-channel optical remote sensing image in each batch of remote sensing image samples, and output the first feature map, second feature map, third feature map and fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples to the feature fusion network.

4. The dual-stage collaborative remote sensing small target detection method according to claim 3, characterized in that: The feature fusion network includes: First convolution module, second convolution module, third convolution module, fourth convolution module, first addition module, second addition module, third addition module, first upsampling module, second upsampling module, third upsampling module, first low-level feature enhancement module, second low-level feature enhancement module, third low-level feature enhancement module, first concatenation module, second concatenation module, third concatenation module, second adaptive feature selection module, third adaptive feature selection module, fourth addition module, fifth addition module, sixth addition module, fifth convolution module, sixth convolution module; The first, second, third, and fourth convolutional modules are all convolutional layers with a kernel size of 1×1 and a stride of 1. The first convolution module is used to input the first feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to extract the features of the first feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples after unifying the number of channels through convolution feature extraction, and output it to the first addition module; The second convolution module is used to input the second feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to extract the features of the second feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples after unifying the number of channels through convolution feature extraction, and output it to the second addition module; The third convolution module is used to input the third feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to extract the features of the third feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples after unifying the number of channels through convolution feature extraction, and output it to the third addition module. The fourth convolution module is used to input the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to extract the features of the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples after unifying the number of channels through convolution feature extraction, and output them to the third upsampling module, the third stitching module and the sixth addition module respectively. The first addition module, the second addition module, and the third addition module all employ element-by-element addition operations; The first upsampling module, the second upsampling module, and the third upsampling module all use nearest neighbor interpolation to perform upsampling operations on the features; The third upsampling module is used to input the one-dimensional features of the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to extract the upsampled features of the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples through feature upsampling, and output them to the third addition module. The third addition module is used to input the one-dimensional features of the third feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples and the upsampled features of the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through addition calculation, the first fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples is obtained, and then output to the second upsampling module, the second stitching module and the fifth addition module respectively. The second upsampling module is used to input the first fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, extract the first upsampled fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples through feature upsampling, and output it to the second addition module; The second addition module is used to input the first upsampling fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples and the one-dimensional feature of the second feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples, perform addition calculations to obtain the second fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and output them to the first upsampling module, the first stitching module and the fourth addition module respectively. The first upsampling module is used to input the second fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to obtain the second upsampled fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples through feature upsampling extraction, and output it to the first addition module; The first addition module is used to input the second upsampling fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples and the one-dimensional feature of the first feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through addition calculation, the third fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples is obtained and output to the first low-level feature enhancement module. The first, second, and third bottom-level feature enhancement modules all employ average pooling with a window size of 2 and a step size of 2. The first splicing module, the second splicing module, and the third splicing module all employ channel-based splicing operations. The first adaptive feature selection module, the second adaptive feature selection module, and the third adaptive feature selection module are all composed of global average pooling, a first adaptive convolutional layer, a sigmoid activation function, a hadamard product module, an element-wise addition module, and a second adaptive convolutional layer cascaded together in sequence. The fourth, fifth, and sixth addition modules all employ element-by-element addition operations. The first low-level feature enhancement module is used to input the third fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and through low-level feature enhancement processing, obtain the third low-level enhanced feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and output it to the first stitching module. The first stitching module is used to input the third bottom-level enhanced features of each three-channel optical remote sensing image of each batch of remote sensing image samples and the second fusion features of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through stitching processing, the first stitching features of each three-channel optical remote sensing image of each batch of remote sensing image samples are obtained and output to the first adaptive feature selection module. In the first adaptive feature selection module, the fused features are processed by the second adaptive convolutional layer to obtain the first adaptive features of each three-channel optical remote sensing image of each batch of remote sensing image samples, and then output to the fourth addition module; The fourth addition module is used to input the first adaptive feature of each three-channel optical remote sensing image of each batch of remote sensing image samples and the second fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through addition calculation, the fourth fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples is obtained and output to the second low-level feature enhancement module and the detection head respectively. The second low-level feature enhancement module is used to input the fourth fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and through low-level feature enhancement processing, obtain the fourth low-level enhanced feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and output it to the second stitching module. The second stitching module is used to input the fourth bottom-level enhanced features of each three-channel optical remote sensing image of each batch of remote sensing image samples and the first fusion features of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through stitching processing, the second stitching features of each three-channel optical remote sensing image of each batch of remote sensing image samples are obtained and output to the second adaptive feature selection module. The second adaptive feature selection module is used to input the second stitching feature of each three-channel optical remote sensing image of each batch of remote sensing image samples. The feature is then processed by the global average pooling layer, the first adaptive convolutional layer, and the Sigmoid activation function to obtain the fusion weights of each channel. The fusion weights of each channel are multiplied by the second stitching feature through the Hadamard product module to obtain the intermediate feature. The intermediate feature and the second stitching feature are then added element-wise to obtain the fusion feature. The fusion feature is processed by the second adaptive convolutional layer to obtain the second adaptive feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and then output to the fifth addition module. The fifth addition module is used to input the second adaptive feature of each three-channel optical remote sensing image of each batch of remote sensing image samples and the first fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through addition calculation, the fifth fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples is obtained and output to the third low-level feature enhancement module and the detection head respectively. The third low-level feature enhancement module is used to input the fifth fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and through low-level feature enhancement processing, obtain the fifth low-level enhanced feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and output it to the third stitching module. The third stitching module is used to input the fifth bottom-level enhanced features of each three-channel optical remote sensing image of each batch of remote sensing image samples and the features of the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples after unifying the number of channels. Through stitching processing, the third stitching features of each three-channel optical remote sensing image of each batch of remote sensing image samples are obtained and output to the first adaptive feature selection module. The third adaptive feature selection module is used to input the third stitching feature of each three-channel optical remote sensing image of each batch of remote sensing image samples. The feature is then processed by the global average pooling layer, the first adaptive convolutional layer, and the Sigmoid activation function to obtain the fusion weights of each channel. The fusion weights of each channel are multiplied by the third stitching feature through the Hadamard product module to obtain the intermediate feature. The intermediate feature and the third stitching feature are then added element-wise to obtain the fusion feature. The fusion feature is processed by the second adaptive convolutional layer to obtain the third adaptive feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and then output to the sixth addition module. The sixth addition module is used to input the third adaptive feature of each three-channel optical remote sensing image of each batch of remote sensing image samples and the features of the fourth feature map of each three-channel optical remote sensing image of each batch of remote sensing image samples after unifying the number of channels. Through addition calculation, the sixth fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples is obtained and output to the fifth convolution module and the detection head respectively. The fifth and sixth convolutional modules are both convolutional layers with a kernel size of 3×3 and a stride of 2. The fifth convolution module is used to input the sixth fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, and to extract the seventh fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples through convolution feature extraction, and output them to the sixth convolution module and the detection head respectively. The sixth convolution module is used to input the seventh fusion feature of each three-channel optical remote sensing image of each batch of remote sensing image samples, extract the eighth feature of each three-channel optical remote sensing image of each batch of remote sensing image samples through convolution feature extraction, and output it to the detection head.

5. The two-stage collaborative remote sensing small target detection method according to claim 4, characterized in that: The detection head includes: a regression branch module, a classification branch module, a label allocation module, and a post-processing module; The detection head is used as input for the fourth, fifth, sixth, seventh, and eighth fusion features of each three-channel optical remote sensing image sample in each batch of remote sensing images. Each fusion feature layer is passed through the detection head with shared weights. The regression branch module of the detection head performs regression prediction to obtain multiple predicted bounding rectangles in each three-channel optical remote sensing image sample in each batch of remote sensing images. The classification branch module of the detection head performs classification prediction to obtain the category of each small target corresponding to the multiple predicted bounding rectangles in each three-channel optical remote sensing image sample in each batch of remote sensing images. During the training phase, positive and negative samples are constructed and the loss function is calculated through the label allocation module to optimize the model. During the testing phase, the post-processing module obtains the bounding rectangle of each detected small target in each three-channel optical remote sensing image sample in each batch of remote sensing images, as well as the category of the small target corresponding to the bounding rectangle of each detected small target in each three-channel optical remote sensing image sample in each batch of remote sensing images.

6. The two-stage collaborative remote sensing small target detection method according to claim 5, characterized in that: The regression branch module consists of four cascaded regression convolutional layers and one cascaded regression convolutional layer; The four cascaded regression convolutional layers are used to input the fourth, fifth, sixth, seventh, and eighth fusion features of each three-channel optical remote sensing image of each batch of remote sensing image samples. The regression features are extracted through the four cascaded regression convolutional layers to obtain the regression features of each three-channel optical remote sensing image of each batch of remote sensing image samples, and then output to the regression convolutional layers. The regression convolutional layer is used for the input regression features of each three-channel optical remote sensing image of each batch of remote sensing image samples. Through regression convolution operation, multiple predicted bounding rectangles are obtained in each three-channel optical remote sensing image of each batch of remote sensing image samples. The bounding rectangles are output to the label assignment module during the training phase and to the post-processing module during the testing phase. The classification branch module consists of four cascaded classification convolutional layers and one classification convolutional layer. The four cascaded classification convolutional layers are used to input the fourth, fifth, sixth, seventh, and eighth fusion features of each three-channel optical remote sensing image of each batch of remote sensing image samples. The classification features are extracted through the four cascaded classification convolutional layers to obtain the classification features of each three-channel optical remote sensing image of each batch of remote sensing image samples, and then output to the classification convolutional layers. The classification convolutional layer is used for the input classification features of each three-channel optical remote sensing image for each batch of remote sensing image samples. Through classification convolution operation, it obtains the category of each small target corresponding to multiple predicted bounding rectangles in each three-channel optical remote sensing image for each batch of remote sensing image samples. The results are output to the label assignment module during the training phase and to the post-processing module during the testing phase.

7. The dual-stage collaborative remote sensing small target detection method according to claim 6, characterized in that: The label allocation module is used to assign positive samples to the detected small targets in each three-channel optical remote sensing image of each batch of remote sensing image samples. The label allocation method based on point prior of FCOS is used to obtain the positive samples of small targets in all three-channel optical remote sensing images of each batch of remote sensing image samples, as well as the remaining negative samples. The prediction results corresponding to the positive samples are used to calculate the regression loss and classification loss, while the prediction results of the negative samples are only used to calculate the classification loss. The post-processing module is used to input regression branch and classification branch predictions to obtain multiple predicted bounding rectangles and their corresponding small target categories in each three-channel optical remote sensing image of each batch of remote sensing image samples. The module then performs post-processing using NMS operations to obtain the bounding rectangles and their corresponding small target categories in each three-channel optical remote sensing image of each batch of remote sensing image samples.

8. The two-stage collaborative remote sensing small target detection method according to claim 7, characterized in that: The loss function model described in step 2 is defined as follows: in, Indicates regression loss, Indicates classification loss; The regression loss is defined as follows: in, This represents the total number of positive samples of all small targets in all three-channel optical remote sensing images of this batch of remote sensing image samples. This indicates the first of all three-channel optical remote sensing images in this batch of remote sensing image samples. The serial number of each positive sample; express The bounding rectangle of the predicted small target. ; in, The first three-channel optical remote sensing images of each batch of remote sensing image samples are respectively the first three-channel optical remote sensing images of each batch of remote sensing image samples. The x-coordinate of the top-left corner, the y-coordinate of the top-left corner, the x-coordinate of the bottom-right corner, and the y-coordinate of the bottom-right corner of the bounding rectangle of the small target predicted by each positive sample. Indicates positive samples The regression target is the index of all small targets in all three-channel optical remote sensing images of this batch of remote sensing image samples. The first small goal is the The regression objective for each sample point (equivalent to...) The first in this batch (positive samples of small goals) This represents the first of all three-channel optical remote sensing images in each batch of remote sensing image samples. The bounding rectangle of a small target. ; in, These are the images from all three-channel optical remote sensing images in this batch of remote sensing image samples. The x-coordinate of the top left corner, the y-coordinate of the top left corner, the x-coordinate of the bottom right corner, and the y-coordinate of the bottom right corner of the bounding rectangle of the small target; This represents the DIoU size between the bounding rectangles of the two smaller targets; The classification loss is defined as follows: in, This indicates reducing the loss weight for negative samples. This indicates reducing the loss weight for simpler samples. This represents the total number of positive samples of all small targets in all three-channel optical remote sensing images of this batch of remote sensing image samples. Let z represent the total number of positive and negative samples of all small targets in all three-channel optical remote sensing images of this batch of remote sensing image samples, and let z represent the nth ... The serial number of each sample. Sample points The predicted score , The number of small target categories, Represents sample points Predict the first of all three-channel optical remote sensing images in this batch of remote sensing image samples. The probability that a small target belongs to category c. Represents sample points The target score, , Represents sample points Predict the first of all three-channel optical remote sensing images in this batch of remote sensing image samples. The probability of a target of category c, where c∈[1,C], when This refers to the first of all three-channel optical remote sensing images in this batch of remote sensing image samples. Positive samples of small targets, and the first of all three-channel optical remote sensing images in this batch of remote sensing image samples. The target categories for each small goal are: ,but , ,in, Indicates the number of iterations. They represent The predicted bounding rectangle of the small target and the first three-channel optical remote sensing images in this batch of remote sensing image samples. The bounding rectangle of a small target, if If it is a negative sample, then ; The amplitude balancing weighting factor is defined as follows: in, , This represents the number of small targets in all three-channel optical remote sensing images of this batch of remote sensing image samples. This indicates the first of all three-channel optical remote sensing images in this batch of remote sensing image samples. The sequence number of each positive sample, positive sample The regression target is the first of all three-channel optical remote sensing images in this batch of remote sensing image samples. A small goal This refers to the first of all three-channel optical remote sensing images in this batch of remote sensing image samples. The number of positive samples assigned to each small objective, and the magnitude weighting factor. It is inversely proportional to the number of positive samples and directly proportional to the average number of samples, which makes the network pay more attention to small targets with few positive samples; The quality balance weighting factor is defined as follows: Where s is the th image among all three-channel optical remote sensing images in this batch of remote sensing image samples. The sequence number of the positive samples of each small target, s= , This refers to the first of all three-channel optical remote sensing images in this batch of remote sensing image samples. The positive sample in the th... All of the small goals ( The relative quality among positive samples (number of samples); in, The overall quality of this positive sample is determined by a combination of its classification and regression scores. Indicates the next batch The first of the small goals The classification score predicted for each positive sample, and simultaneously using and These represent the next [number] in this batch. The bounding rectangle of the small target and the first target The bounding rectangle of the small target predicted by each positive sample; in, They are respectively the batch of the first The x-coordinate of the top-left corner, the y-coordinate of the top-left corner, the x-coordinate of the bottom-right corner, and the y-coordinate of the bottom-right corner of the bounding rectangle of the small target. Indicates the next [number] in this batch The x-coordinate of the top-left corner, y-coordinate of the top-left corner, and y-coordinate of the bottom-right corner of the bounding rectangle of the predicted small target are used to measure the normalized KL divergence distance between the Gaussian distributions of the labeled box and the predicted box. ,Will As a regression score, the overall quality can be expressed as: in, =0.2 indicates that the classification score has a weight in the measurement of sample quality; This refers to the first of all three-channel optical remote sensing images in this batch of remote sensing image samples. The classification prediction score of the s-th positive sample of each small target The calculated Varifocal Loss in, Indicates the number of iterations, quality balance weight factor It increases with the relative quality of positive samples, enhancing the network's attention to high-quality positive samples and promoting feature learning for small targets.

9. A two-stage collaborative remote sensing small target detection system, used to implement the two-stage collaborative remote sensing small target detection method according to any one of claims 1-8, characterized in that, include: The sample labeling module is used to acquire multiple three-channel optical remote sensing images. It randomly selects each batch of remote sensing image samples from multiple three-channel optical remote sensing images and labels the bounding rectangles of all small targets and the small target categories of all small target bounding rectangles in all three-channel optical remote sensing images of each batch of remote sensing image samples. The remote sensing small target detection network training module is used to construct the remote sensing small target detection network. Each batch of remote sensing image samples and each three-channel optical remote sensing image are sequentially input into the remote sensing small target detection network for small target detection. The module obtains the bounding rectangle of each predicted small target and the predicted small target category of each bounding rectangle in each batch of remote sensing image samples. A loss function model is constructed by combining the bounding rectangle of each small target and the small target category of each bounding rectangle in the corresponding three-channel optical remote sensing image of each batch of remote sensing image samples. The updated remote sensing small target detection network is obtained through iterative optimization training batch by batch. The real-time batch remote sensing image sample detection module is used to acquire multiple real-time three-channel optical remote sensing images, construct real-time batch remote sensing image samples, input the updated remote sensing small target detection network to perform small target detection, and obtain the bounding rectangle of each predicted small target in each real-time three-channel optical remote sensing image of the real-time batch remote sensing image sample, and the predicted small target category of each predicted small target bounding rectangle.

Citation Information

Patent Citations

  • Aircraft target detection method based on multi-parameter optimization YOLOV4 network

    CN114494861A

  • Remote sensing image target detection method fusing multi-scale context features and channel enhancement

    CN116246173A