Remote sensing small target detection method, device and equipment based on improved RT-DETR model

By improving the backbone network, decoder, and loss function of the RT-DETR model, the feature extraction capability and detection accuracy of remote sensing small target detection are enhanced, the shortcomings of existing models in remote sensing small target detection are addressed, and more efficient remote sensing small target detection is achieved.

CN120913083AActive Publication Date: 2025-11-07XIAMEN UNIV OF TECH

Patent Information

Application Number
CN202511438242.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2025-11-07
Estimated Expiration
2045-10-10

AI Technical Summary

Technical Problem

Existing RT-DETR models suffer from problems in remote sensing small target detection, such as insufficient backbone network for small target feature extraction, insufficient receptive field of the decoder, and loss function that does not highlight the importance of small targets, resulting in low detection accuracy.

Method used

The improved RT-DETR model enhances feature extraction capabilities through element-wise multiplication, introduces extraction branches for small target features and dilated convolution to expand the receptive field, uses the SPDConv module to preserve detailed information, and introduces the CSPOmniKernel module for feature splitting and multi-directional capture in the feature fusion stage. The loss function is optimized to balance the feature extraction of small targets with computational complexity.

Benefits of technology

It improves the accuracy and efficiency of remote sensing small target detection, enhances the ability to extract small target features, solves the sample imbalance problem, and improves the accuracy and real-time performance of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913083A_ABST
    Figure CN120913083A_ABST
Patent Text Reader

Abstract

The invention discloses a remote sensing small target detection method, device and equipment based on an improved RT-DETR model, and the method comprises the steps: S1, obtaining a remote sensing image data set as a training set, carrying out the multi-scale transformation of remote sensing images in the remote sensing image data set, obtaining remote sensing images with different resolutions, and carrying out the multi-scale transformation of the remote sensing images; performing enhancement processing on the remote sensing images with different resolutions to obtain processed remote sensing images; s2, inputting the processed remote sensing image into an improved RT-DETR model so as to train the improved RT-DETR model, and obtaining a trained remote sensing small target detection network model; and S3, performing remote sensing small target detection on an input remote sensing image by using the remote sensing small target detection network model. According to the method, through the improved RT-DETR model, the feature extraction capability of the backbone network on the small target can be improved, and the decoder can optimize the size of the small target, so that the detection accuracy of the remote sensing small target is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of deep learning and computer vision, in particular to a remote sensing small target detection method, device and equipment based on an improved RT-DETR model. BACKGROUND

[0002] Remote sensing target detection (RSOD) is an important branch of computer vision, aiming to identify and locate targets from remote sensing images obtained by satellite, unmanned aerial vehicle and other platforms. However, remote sensing images have significant complexity, which is reflected in the following aspects: the size of the same target in different resolution images varies greatly, small targets (such as vehicles and ships) often occupy only a few pixels, and the features are difficult to identify; cloud cover, vegetation cover, building shadows and other factors result in low target-to-background contrast and unclear edge features; traditional target detection algorithms (such as YOLO and Faster R-CNN) are prone to missing detection or false detection when processing low-quality remote sensing images due to insufficient feature extraction.

[0003] The RT-DETR (Real-Time Detection Transformer) algorithm, as an end-to-end detection framework, has been widely concerned and applied to the field of remote sensing target detection due to its excellent performance in real-time and accuracy. However, when it comes to small remote sensing targets, it has the following problems: the backbone network lacks the ability to extract small target features and lacks an effective dimension expansion mechanism; the decoder is not optimized for small target sizes, resulting in a lack of context information due to insufficient receptive field; the loss function does not highlight the importance of small targets, and the sample imbalance problem affects the detection accuracy. SUMMARY

[0004] The present application aims to provide a remote sensing small target detection method, device and equipment based on an improved RT-DETR model to improve the above problems.

[0005] To solve the above technical problems, the present application realizes the following technical scheme: A remote sensing small target detection method based on an improved RT-DETR model, the method comprising: S1, obtaining a remote sensing image dataset as a training set, performing multi-scale transformation on the remote sensing images in the remote sensing image dataset to obtain remote sensing images of different resolutions, and then performing enhancement processing on the remote sensing images of different resolutions to obtain processed remote sensing images; S2, input the processed remote sensing image into the improved RT-DETR model to train the improved RT-DETR model, and obtain a trained remote sensing small target detection network model; wherein, the backbone network of the improved RT-DETR model adopts element-wise multiplication operation to process the features of the input remote sensing image, and maps the input features to a high-dimensional nonlinear space through element-wise multiplication, thereby enhancing the extraction capability of small target features; the decoding process of the decoder of the improved RT-DETR model includes a feature extraction stage, a feature down-sampling stage and a feature fusion stage; in the feature extraction stage, an extraction branch for small target features is introduced and configured to adapt to the size of small targets by adjusting the convolution kernel size and step length of the extraction branch, and a hollow convolution is added in the extraction branch to expand the receptive field and obtain more rich context information; in the feature down-sampling stage, an SPDConv module is introduced, and through the multi-branch sub-region feature splicing and convolution fusion of the SPDConv module, the small target detail information is preserved when the feature resolution is reduced; in the feature fusion stage, a CSPOmniKernel module is introduced, and through the feature splitting, multi-directional feature capturing and attention enhancement of the CSPOmniKernel module, the small target feature extraction capability and the computational complexity are balanced; S3, using the remote sensing small target detection network model to detect small targets in the input remote sensing image.

[0006] Preferably, in step S2, the SPDConv module is implemented as follows: The input channel number inc and the output channel number ouc are set when the module is initialized, and a 3x3 convolution layer with an input channel number of incx4 is set internally; In the forward propagation process, the input feature map is non-overlapping sub-region sampled by a step of 2 to generate 4 complementary sub-features, including: taking the upper left corner sub-region of even row and even column pixels, taking the lower left corner sub-region of odd row and even column pixels, taking the upper right corner sub-region of even row and odd column pixels, and taking the lower right corner sub-region of odd row and odd column pixels; The 4 sub-features are spliced in the channel dimension to form an intermediate feature with a channel number of incx4, and then compressed to an output channel number ouc through a 3x3 convolution, and the output resolution is Figure 1 / 2 of the input feature.

[0007] Preferably, in step S2, the CSPOmniKernel module includes 1x1 convolution layers cv1, cv2 and an OmniKernel submodule, and is implemented as follows: The feature channel number dim and the splitting ratio e=0.25 are set when the module is initialized, the input channel number of the OmniKernel submodule is dimo, and dimo is the integer part of dimx e; In the forward propagation process, the input features are split into two branches after cv1 convolution: the features of the dimo channels enter the OmniKernel submodule for processing, and the features of the remaining channels are directly retained as the identity branch; The output features of the OmniKernel submodule and the features of the identity branch are spliced in the channel dimension, and the fused features are output after cv1 convolution.

[0008] Preferably, the OmniKernel submodule comprises an input convolutional layer with GELU activation, an output convolutional layer, a multi-directional depth convolution group, a frequency domain attention module, a spatial attention module, and a ReLU activation layer; wherein: The multi-directional depth convolution group includes four types of depth separable convolution: 1x31 convolution for long-distance feature capture in the horizontal direction, 31x1 convolution for long-distance feature capture in the vertical direction, 31x31 convolution for global spatial feature capture, and 1x1 convolution for local feature interaction within the channel; The frequency domain attention module generates attention weights through global average pooling and 1x1 convolution, and weights the features in the Fourier domain to suppress high-frequency noise; the spatial attention module generates a spatial attention map through global average pooling and 1x1 convolution, and weights the features processed in the frequency domain to enhance the response of small target regions; During forward propagation, the input features are processed by the multi-directional depth convolution group, the frequency domain attention module, and the spatial attention module after the input convolutional layer, and all processing results are added to the original input features, then output after ReLU activation and output convolutional layer.

[0009] Preferably, the high-dimensional feature generated by the element-wise multiplication operation in the backbone network is as follows: The single-layer element-wise multiplication operation is represented as:

[0010] wherein, , is a weight vector, , is , the transpose of , is an input feature vector, is the feature dimension; represents the i-th element of the weight vector represents the j-th element of the weight vector , respectively represent the i-th and j-th features of the input feature vector; after expansion, a nonlinear feature term, realizing dimension expansion without increasing the amount of calculation; through a layer of element-by-element multiplication operation, the feature dimension increases exponentially, and the feature dimension of the output of the layer is the output feature dimension of the layer is .

[0011] Preferably, a loss term for small targets is added in the loss function of the improved RT-DETR model, including a small target size loss term and a small target positioning loss term; wherein the small target size loss term adopts a combination of IoU loss and CIoU loss, and the calculation formula is as follows:

[0012] wherein, is a prediction box, is a real box, is the square of the Euclidean distance between the center points of the prediction box and the real box, is the diagonal length of the smallest closed region containing the prediction box and the real box, is a weight coefficient, is used to measure the difference in aspect ratio between the prediction box and the real box.

[0013] Preferably, in step S2, an adaptive learning rate adjustment strategy is adopted in the training process of the improved RT-DETR model, which is a cosine annealing learning rate adjustment method, and the formula is as follows:

[0014] wherein, is the current learning rate, is the minimum learning rate, is the maximum learning rate, is the current training round, is the maximum training round.

[0015] Preferably, in step S3, the working principle of the remote sensing small target network model is as follows: extract multi-scale visual features of the remote sensing image through the backbone network, and generate a feature pyramid containing spatial semantic information; initialize a set of learnable target query vectors through a target query vector generation module, which are used to encode prior knowledge of the target detection task, including target position, size and category; input the feature pyramid into the encoder for multi-layer feature enhancement, the encoder adopts a multi-head self-attention mechanism and a feedforward neural network to model the cross-regional context of different levels of visual features, and enhance the feature expression of small target regions; at the same time, during the construction of the feature pyramid, the SPDConv module is used for down-sampling on the high-resolution feature layer to preserve the details of small targets; The decoder receives the enhanced features and the target query vector output by the encoder, and establishes the correspondence between the target query and the visual features through a cross-attention mechanism; specifically: first, the adjusting attention mechanism of the visual features to the target query selects the region features related to the small target from the feature pyramid by calculating the similarity between the target query vector and the visual features, and suppresses the background noise interference; then, the guiding attention mechanism of the target query to the visual features uses the adjusted target query vector to guide the feature aggregation direction, so that the decoder focuses on the detailed features of the small target; after the bidirectional attention interaction, the features are input into the self-attention module for inter-target relationship modeling to avoid detection confusion of adjacent small targets; meanwhile, the CSPOmniKernel module is introduced in the decoder feature fusion stage to enhance the discriminability of the small target features; The detection results output by the decoder are processed by the Hungarian matching algorithm to optimally match the predicted boxes and the real labeled boxes, and the classification loss and the positioning loss are calculated: the focal loss is used for the classification loss to solve the imbalance problem of small target samples; the CIoU loss is used for the positioning loss to combine the overlap area, the center point distance and the aspect ratio difference of the predicted box and the real box, and to improve the positioning accuracy of the small target; The network parameters are optimized through end-to-end training, and the weights of the backbone network, the encoder and the decoder are updated synchronously during the training process, so that the model can adaptively learn the feature representation and detection logic of the remote sensing small target.

[0016] The embodiment of the present application also provides a remote sensing small target detection device based on an improved RT-DETR model, which comprises: An image processing unit is configured to obtain a remote sensing image dataset as a training set, perform multi-scale transformation on remote sensing images in the remote sensing image dataset to obtain remote sensing images with different resolutions, and perform enhancement processing on the remote sensing images with different resolutions to obtain processed remote sensing images. The model training unit is configured to input the processed remote sensing image into an improved RT-DETR model to train the improved RT-DETR model, so as to obtain a trained remote sensing small target detection network model. The detection unit is configured to perform remote sensing small target detection on an input remote sensing image by using the remote sensing small target detection network model.

[0017] The embodiment of the present application also provides a remote sensing small target detection device based on an improved RT-DETR model, which comprises a memory, a processor and computer program instructions stored in the memory and capable of being executed by the processor, and when the processor executes the computer program instructions, the remote sensing small target detection method based on the improved RT-DETR model can be realized.

[0018] In summary, the improved RT-DETR model has the following advantages compared with the existing RT-DETR model: 1. The backbone network adopts element-wise multiplication operation to process the features of the input remote sensing image, and maps the input features to a high-dimensional nonlinear space through element-wise multiplication, thereby enhancing the extraction capability of small target features; 2. In the feature extraction stage, an extraction branch for small target features is introduced, and is configured to adapt to the size of the small target by adjusting the convolution kernel size and step length of the extraction branch, and a dilated convolution is added in the extraction branch to expand the receptive field and obtain more rich context information; 3. In the feature down-sampling stage, an SPDConv module is introduced, and the multi-branch sub-region feature splicing and convolution fusion of the SPDConv module are used to retain small target detail information when reducing the feature resolution; 4. Introducing the CSPOmniKernel module in the feature fusion stage, balancing the small target feature extraction capability and the computational complexity through feature splitting, multi-directional feature capture and attention enhancement of the CSPOmniKernel module. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced as follows. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0020] Figure 1 Fig. 1 shows a flowchart of a remote sensing small target detection method based on an improved RT-DETR model according to an embodiment of the present application; Figure 2 Fig. 2 shows a structure diagram of an improved RT-DETR model according to an embodiment of the present application; Figure 3 Fig. 3 shows a structure diagram of a remote sensing small target detection device based on an improved RT-DETR model according to an embodiment of the present application. DETAILED DESCRIPTION

[0021] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0022] In order to better understand the technical solutions of the present application, the embodiments of the present application will be described in detail below with reference to the drawings.

[0023] It should be clear that the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0024] The terms used in the embodiments of the present application are only for the purpose of describing the specific embodiments, and are not intended to limit the present application. The singular forms "a", "an" and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0025] It should be understood that the term "and / or" as used herein merely describes an associated relationship between associated objects, and can represent three relationships, for example, A and / or B can represent three cases of A alone, A and B together, and B alone. In addition, the character " / " herein generally represents an "or" relationship between the front and rear associated objects.

[0026] Depending on the context, the word "if" as used herein can be interpreted as meaning "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if it is determined" or "if (a stated condition or event) is detected" can be interpreted as meaning "when it is determined" or "in response to determining" or "when (a stated condition or event) is detected" or "in response to detecting (a stated condition or event)".

[0027] The "first / second" mentioned in the embodiments is only to distinguish similar objects, and does not represent a specific order of the objects. Understandably, the "first / second" can be interchanged in a specific order or sequence as allowed. It should be understood that the objects distinguished by "first / second" can be interchanged under appropriate circumstances to enable the embodiments described herein to be implemented in an order other than those illustrated or described herein.

[0028] The application will be described in further detail below in conjunction with the accompanying drawings and specific embodiments: Embodiment one The embodiment one of the application provides a remote sensing small target detection method based on an improved RT-DETR model, which can be realized by a remote sensing small target detection device based on an improved RT-DETR model (hereinafter referred to as a detection device). In particular, it is executed by one or more processors in the detection device to realize the following method: S1, obtaining a remote sensing image dataset as a training set, performing multi-scale transformation on remote sensing images in the remote sensing image dataset to obtain remote sensing images of different resolutions, and then performing enhancement processing on the remote sensing images of different resolutions to obtain processed remote sensing images; In this embodiment, the detection device can be an electronic device equipped with a processor, the processor has a computer program of the detection method and the computer program can be executed, for example, a computer, a smart phone, a smart tablet, a workstation, etc., which are not limited herein.

[0029] In this embodiment, in particular, the remote sensing image dataset can be VisDrone2019 and DOTA-v1.0 dataset, but is not limited specifically.

[0030] S2. The processed remote sensing image is input into the improved RT-DETR model to train the improved RT-DETR model and obtain a trained remote sensing small target detection network model.

[0031] Among them, the RT-DETR (Real-Time Detection Transformer) model is a real-time object detection model based on the Transformer architecture. It combines the advantages of CNN and Transformer, aiming to improve the balance between detection accuracy and speed. In this embodiment, the improvements of the improved RT-DETR model over the existing RT-DETR model include the backbone network, decoder, loss function, and learning rate adjustment strategy, which are described in detail below.

[0032] 1. Backbone Network In this embodiment, the backbone network of the improved RT-DETR model uses element-wise multiplication to process the features of the input remote sensing image. By mapping the input features to a high-dimensional nonlinear space through element-wise multiplication, the ability to extract features of small targets is enhanced.

[0033] Single-level element-wise multiplication is represented as:

[0034] in, , For the weight vector, , for , transpose, For the input feature vector, For feature dimensions; Represents the weight vector The i-th element, Represents the weight vector The j-th element, These represent the i-th and j-th features of the input feature vector, respectively; after expansion, they generate... The nonlinear feature terms achieve dimensional expansion without increasing computational cost; through layer-by-layer element-wise multiplication, the feature dimension grows exponentially, with the 1st layer... The layer output feature dimension is .

[0035] 2. Decoder In the embodiment, the decoding process of the decoder of the improved RT-DETR model includes a feature extraction stage, a feature down-sampling stage, and a feature fusion stage; in the feature extraction stage, an extraction branch for small target features is introduced and is configured to be able to adapt to the size of small targets by adjusting the kernel size and step of the extraction branch, and a dilated convolution is added in the extraction branch to expand the receptive field and obtain more rich context information; in the feature down-sampling stage, an SPDConv module is introduced, and through the multi-branch sub-region feature splicing and convolution fusion of the SPDConv module, the small target detail information is preserved when the feature resolution is reduced; in the feature fusion stage, a CSPOmniKernel module is introduced, and through the feature splitting, multi-directional feature capturing and attention enhancement of the CSPOmniKernel module, the small target feature extraction capability and the computational complexity are balanced; Specifically, the implementation of the SPDConv module is as follows: The input channel number inc and the output channel number ouc are set when the module is initialized, and a 3*3 convolution layer with an input channel number of inc*4 is set internally; In the forward propagation process, the input feature map is non-overlapping sub-region sampled with a step of 2 to generate 4 complementary sub-features, including: taking the upper left corner sub-region of the even row and even column pixels, the lower left corner sub-region of the odd row and even column pixels, the upper right corner sub-region of the even row and odd column pixels, and the lower right corner sub-region of the odd row and odd column pixels; The 4 sub-features are spliced in the channel dimension to form an intermediate feature with a channel number of inc*4, and then compressed to an output channel number of ouc through a 3*3 convolution, and the output resolution is Figure 1 / 2 of the input feature.

[0036] The CSPOmniKernel module includes 1*1 convolution layers cv1, cv2 and an OmniKernel submodule, and the implementation is as follows: The feature channel number dim and the splitting ratio e=0.25 are set when the module is initialized, and the input channel number of the OmniKernel submodule is dimo, which is the integer part of dim*e; In the forward propagation process, the input feature is split into two branches in proportion after cv1 convolution: the dimo channel feature enters the OmniKernel submodule for processing, and the remaining channel feature is directly retained as an identity branch; The output feature of the OmniKernel submodule and the feature of the identity branch are spliced in the channel dimension, and the fusion feature is output after cv1 convolution.

[0037] The OmniKernel submodule comprises an input convolutional layer with a GELU activation, an output convolutional layer, a multi-directional depth convolution group, a frequency domain attention module, a spatial attention module, and a ReLU activation layer; wherein: The multi-directional depth convolution group comprises four kinds of depth separable convolution: a 1x31 convolution for long-distance feature capture in the horizontal direction, a 31x1 convolution for long-distance feature capture in the vertical direction, a 31x31 convolution for global spatial feature capture, and a 1x1 convolution for local feature interaction within a channel; The frequency domain attention module generates attention weights through global average pooling and a 1x1 convolution, and weights the features in the Fourier domain to suppress high-frequency noise; the spatial attention module generates a spatial attention map through global average pooling and a 1x1 convolution, and weights the features processed in the frequency domain to enhance the response of small target regions; During forward propagation, after the input features are processed by the input convolutional layer, they are processed by the multi-directional depth convolution group, the frequency domain attention module, and the spatial attention module, respectively. All processing results are added to the original input features, and then output through the output convolutional layer after ReLU activation.

[0038] 3. Loss function In this embodiment, a loss term for small targets is added in the loss function of the improved RT-DETR model, including a small target size loss term and a small target positioning loss term; wherein the small target size loss term adopts a combination of IoU loss and CIoU loss, and the calculation formula is as follows:

[0039] wherein, is the predicted box, is the real box, is the square of the Euclidean distance between the center points of the predicted box and the real box, is the diagonal length of the smallest closed region containing the predicted box and the real box, is the weight coefficient, is used to measure the aspect ratio difference between the predicted box and the real box.

[0040] In this embodiment, an adaptive learning rate adjustment strategy is adopted in the training process of the improved RT-DETR model, which is a cosine annealing learning rate adjustment method, and the formula is as follows:

[0041] wherein, is the current learning rate, is the minimum learning rate, is the maximum learning rate, is the current training round, is the maximum training round.

[0042] S3, detecting a remote sensing small target in the input remote sensing image by using the remote sensing small target detection network model.

[0043] In the small target detection process, the working principle of the remote sensing small target network model is as follows in the embodiment: The backbone network extracts multi-scale visual features of the remote sensing image to generate a feature pyramid containing spatial semantic information; A set of learnable target query vectors are initialized by a target query vector generation module, which are used to encode prior knowledge of the target detection task, including target position, size and category; The feature pyramid is input into an encoder for multi-layer feature enhancement. The encoder adopts a multi-head self-attention mechanism and a feedforward neural network to model the cross-region context of different levels of visual features and enhance the feature expression of small target regions. At the same time, during the feature pyramid construction stage, the SPDConv module is used for down-sampling on the high-resolution feature layer to preserve the details of small targets. The decoder receives the enhanced features output by the encoder and the target query vector, and establishes the corresponding relationship between the target query and the visual features through the cross-attention mechanism. Specifically, first, the visual feature adjusts the attention mechanism of the target query by calculating the similarity between the target query vector and the visual feature, and selects the region features related to the small target from the feature pyramid to suppress background noise interference. Then, the guided attention mechanism of the visual feature uses the adjusted target query vector to guide the feature aggregation direction, so that the decoder focuses on the detailed features of the small target. After the bidirectional attention interaction, the features are input into the self-attention module for inter-target relationship modeling to avoid detection confusion of adjacent small targets. At the same time, the CSPOmniKernel module is introduced in the decoder feature fusion stage to enhance the discriminability of small target features. The detection results output by the decoder are processed by the Hungarian matching algorithm to optimally match the predicted boxes with the real labeled boxes, and the classification loss and the positioning loss are calculated. The focal loss is used for the classification loss to solve the imbalance problem of small target samples. The CIoU loss is used for the positioning loss to improve the positioning accuracy of small targets by combining the overlap area, center point distance and aspect ratio difference between the predicted box and the real box. The network parameters are optimized through end-to-end training, and the weights of the backbone network, the encoder and the decoder are updated synchronously during the training process, so that the model can adaptively learn the feature representation and detection logic of the remote sensing small target.

[0044] In summary, the improved RT-DETR model of the embodiment has the following advantages compared with the existing RT-DETR model: 1. The backbone network uses element-wise multiplication operation to process the features of the input remote sensing image, and maps the input features to a high-dimensional nonlinear space through element-wise multiplication, thereby enhancing the extraction ability of small target features; 2. In the feature extraction stage, a small target feature extraction branch is introduced, and the convolution kernel size and step of the extraction branch are configured to adapt to the size of the small target, and a dilated convolution is added to the extraction branch to expand the receptive field and obtain more rich context information; 3. In the feature down-sampling stage, the SPDConv module is introduced, and through the multi-branch sub-region feature splicing and convolution fusion of the SPDConv module, the small target detail information is preserved when reducing the feature resolution; 4. In the feature fusion stage, the CSPOmniKernel module is introduced, and through the feature splitting, multi-directional feature capturing and attention enhancement of the CSPOmniKernel module, the small target feature extraction ability and the computational complexity are balanced.

[0045] 5. The detection result output by the decoder is processed by the Hungarian matching algorithm, the predicted box and the real labeled box are optimally matched, and the classification loss and the positioning loss are calculated: the focal loss (Focal Loss) is used for classification loss to solve the imbalance problem of small target samples; the CIoU loss is used for positioning loss, which combines the overlap area, center point distance and aspect ratio difference of the predicted box and the real box to improve the positioning accuracy of the small target; 6. The network parameters are optimized through end-to-end training, and the weights of the backbone network, the encoder and the decoder are updated synchronously during the training process, so that the model can adaptively learn the feature representation and detection logic of the remote sensing small target.

[0046] Please refer to Figure 3 The second embodiment of the present application also provides a remote sensing small target detection device based on the improved RT-DETR model, which comprises: An image processing unit 210 is configured to obtain a remote sensing image dataset as a training set, perform multi-scale transformation on the remote sensing images in the remote sensing image dataset to obtain remote sensing images with different resolutions, and then perform enhancement processing on the remote sensing images with different resolutions to obtain processed remote sensing images. The model training unit 220 is configured to input the processed remote sensing image into an improved RT-DETR model to train the improved RT-DETR model, and obtain a trained remote sensing small target detection network model; wherein the backbone network of the improved RT-DETR model adopts an element-wise multiplication operation to process the features of the input remote sensing image, and maps the input features to a high-dimensional nonlinear space through element-wise multiplication to enhance the extraction capability of small target features; the decoding process of the decoder of the improved RT-DETR model includes a feature extraction stage, a feature down-sampling stage and a feature fusion stage; in the feature extraction stage, an extraction branch for small target features is introduced and configured to adapt to the size of small targets by adjusting the convolution kernel size and step length of the extraction branch, and a dilated convolution is added in the extraction branch to expand the receptive field and obtain more rich context information; in the feature down-sampling stage, a SPDConv module is introduced, and the multi-branch sub-region feature splicing and convolution fusion of the SPDConv module are used to retain small target detail information when reducing the feature resolution; in the feature fusion stage, a CSPOmniKernel module is introduced, and the feature splitting, multi-directional feature capturing and attention enhancement of the CSPOmniKernel module are used to balance the small target feature extraction capability and the computational complexity; The detection unit 230 is configured to perform remote sensing small target detection on the input remote sensing image by using the remote sensing small target detection network model.

[0047] The third embodiment of the present application further provides a remote sensing small target detection device based on an improved RT-DETR model, which comprises a memory, a processor and computer program instructions stored in the memory and capable of being executed by the processor, and when the processor executes the computer program instructions, the remote sensing small target detection method based on the improved RT-DETR model as described above can be realized.

[0048] In several embodiments provided by the present application, it should be understood that the disclosed apparatus and method can also be implemented by other manners. The apparatus and method embodiments described above are merely exemplary, for example, the flowcharts and block diagrams in the drawings show the possible implementation architecture, function and operation of the apparatus, method and computer program product according to the embodiments of the present application. In this regard, each block in the flowcharts or block diagrams can represent a module, a program segment or a part of code, which includes one or more executable instructions for implementing the specified logic function. It should also be noted that in some alternative implementation manners, the functions noted in the blocks can also occur in different order from that noted in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can also be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and the combination of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system for executing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0049] In addition, the function modules in the embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0050] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, and the computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes. It should be noted that in this document, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the processes, methods, articles or devices that include a series of elements not only include those elements, but also include other elements not explicitly listed, or include elements inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or device that includes the element.

[0051] The above description is only the preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A remote sensing small target detection method based on an improved RT-DETR model, characterized in that, The method comprises: S1, obtaining a remote sensing image dataset as a training set, performing multi-scale transformation on remote sensing images in the remote sensing image dataset to obtain remote sensing images of different resolutions, and then performing enhancement processing on the remote sensing images of different resolutions to obtain processed remote sensing images; S2, inputting the processed remote sensing images into an improved RT-DETR model to train the improved RT-DETR model, and obtaining a trained remote sensing small target detection network model; wherein the backbone network of the improved RT-DETR model uses element-wise multiplication operation to process the features of the input remote sensing images to map the input features to a high-dimensional nonlinear space; the decoding process of the decoder of the improved RT-DETR model comprises a feature extraction stage, a feature down-sampling stage and a feature fusion stage; in the feature extraction stage, an extraction branch for small target features is introduced and configured to adapt to the size of small targets by adjusting the convolution kernel size and step length of the extraction branch, and a hollow convolution is added in the extraction branch to expand the receptive field and obtain more rich context information; in the feature down-sampling stage, a SPDConv module is introduced, and through the multi-branch sub-region feature splicing and convolution fusion of the SPDConv module, the small target detail information is preserved when the feature resolution is reduced; in the feature fusion stage, a CSPOmniKernel module is introduced, and through the feature splitting, multi-directional feature capturing and attention enhancement of the CSPOmniKernel module, the small target feature extraction capability and the computational complexity are balanced; S3, using the remote sensing small target detection network model to detect remote sensing small targets in the input remote sensing images.

2. The remote sensing small target detection method based on the improved RT-DETR model according to claim 1, characterized in that, In step S2, the SPDConv module is implemented as follows: Set the input channel number inc and the output channel number ouc when initializing the module, and set a 3*3 convolution layer with an input channel number of inc*4 inside the module; In the forward propagation process, the input feature map is non-overlapping sub-region sampled by a step of 2 to generate 4 complementary sub-features, including: taking the upper left corner sub-region of even row and even column pixels, the lower left corner sub-region of odd row and even column pixels, the upper right corner sub-region of even row and odd column pixels, and the lower right corner sub-region of odd row and odd column pixels; The 4 sub-features are spliced in the channel dimension to form an intermediate feature with a channel number of inc*4, and then compressed to an output channel number of ouc through a 3*3 convolution, and the output resolution is 1 / 2 of the input feature map.

3. The remote sensing small target detection method based on the improved RT-DETR model according to claim 1, characterized in that, In step S2, the CSPOmniKernel module comprises 1*1 convolution layers cv1, cv2 and an OmniKernel submodule, and is implemented as follows: Set the feature channel number dim and the splitting ratio e=0.25 when initializing the module, and the input channel number of the OmniKernel submodule is dimo, which is the integer part of dim*e; In the forward propagation process, the input feature is split into two branches in proportion after cv1 convolution: the dimo channel feature enters the OmniKernel submodule for processing, and the remaining channel feature is directly retained as an identity branch. The output features of the OmniKernel submodule are spliced with the features of the identity branch in the channel dimension, and a fusion feature is output after cv1 convolution.

4. The remote sensing small target detection method based on the improved RT-DETR model according to claim 3, characterized in that, The OmniKernel submodule comprises an input convolutional layer with a GELU activation, an output convolutional layer, a multi-directional depth convolution group, a frequency domain attention module, a spatial attention module, and a ReLU activation layer; wherein: The multi-directional depth convolution group comprises four kinds of depth separable convolution: a 1×31 convolution for capturing long-distance features in the horizontal direction, a 31×1 convolution for capturing long-distance features in the vertical direction, a 31×31 convolution for capturing global spatial features, and a 1×1 convolution for local feature interaction within a channel; The frequency domain attention module generates attention weights through global average pooling and 1×1 convolution, and weights the features in the Fourier domain to suppress high-frequency noise; the spatial attention module generates a spatial attention map through global average pooling and 1×1 convolution, and weights the features processed in the frequency domain to enhance the response of small target regions; During forward propagation, after the input features pass through the input convolutional layer, they are processed by the multi-directional depth convolution group, the frequency domain attention module, and the spatial attention module, respectively. All processing results are added to the original input features, and then output after ReLU activation and output convolution.

5. The remote sensing small target detection method based on the improved RT-DETR model according to claim 1, characterized in that, The way the element-wise multiplication operation in the backbone network generates high-dimensional features is as follows: The single-layer element-wise multiplication operation is represented as: in, , For the weight vector, , for , transpose, For the input feature vector, For feature dimensions; Represents the weight vector The i-th element, Represents the weight vector The j-th element, These represent the i-th and j-th features of the input feature vector, respectively; after expansion, they generate... The nonlinear feature terms achieve dimensional expansion without increasing computational cost; through layer-by-layer element-wise multiplication, the feature dimension grows exponentially, with the 1st layer... The layer output feature dimension is R represents the real number field.

6. The remote sensing small target detection method based on the improved RT-DETR model according to claim 1, characterized in that, In the loss function of the improved RT-DETR model, a loss term for small targets is added, including a small target size loss term and a small target positioning loss term; wherein the small target size loss term adopts a combination of IoU loss and CIoU loss, and the calculation formula is as follows: wherein, is a predicted bounding box, is a ground truth bounding box, is a square of Euclidean distance between the center points of the predicted and ground truth bounding boxes, is a diagonal length of the minimum closed region containing the predicted and ground truth bounding boxes, is a weight coefficient, is used to measure the aspect ratio difference between the predicted and ground truth bounding boxes.

7. The remote sensing small target detection method based on the improved RT-DETR model according to claim 1, characterized in that, In step S2, an adaptive learning rate adjustment strategy is used in the training process of the improved RT-DETR model, which is a cosine annealing learning rate adjustment method, and the formula is as follows: wherein, is the current learning rate, is the minimum learning rate, is the maximum learning rate, is the current training epoch, is the maximum training epoch.

8. The remote sensing small target detection method based on the improved RT-DETR model according to claim 7, characterized in that, In step S3, the working principle of the remote sensing small target network model is as follows: Multi-scale visual features of the remote sensing image are extracted through the backbone network to generate a feature pyramid containing spatial semantic information; A set of learnable target query vectors are initialized by the target query vector generation module, which are used to encode prior knowledge of the target detection task, including target position, size, and class; The feature pyramid is input into the encoder for multi-layer feature enhancement. The encoder adopts a multi-head self-attention mechanism and a feedforward neural network to model the cross-region context of different levels of visual features and enhance the feature expression of small target regions; at the same time, during the construction of the feature pyramid, the SPDConv module is used for down-sampling on the high-resolution feature layer to preserve small target details. The decoder receives the enhanced features and the target query vector output by the encoder, and establishes the correspondence between the target query and the visual features through a cross-attention mechanism. Specifically, first, the visual feature adjusts the attention mechanism of the target query by calculating the similarity between the target query vector and the visual feature, and filters out the region features related to the small target from the feature pyramid, thereby suppressing the background noise interference. Then, the guided attention mechanism of the target query to the visual feature uses the adjusted target query vector to guide the feature aggregation direction, so that the decoder focuses on the detailed features of the small target. After the bidirectional attention interaction, the features are input into a self-attention module for inter-target relationship modeling to avoid detection confusion of adjacent small targets. Meanwhile, a CSPOmniKernel module is introduced in the decoder feature fusion stage to enhance the discriminability of small target features. The detection results output by the decoder are processed by the Hungarian matching algorithm to optimally match the predicted boxes with the real labeled boxes, and the classification loss and the positioning loss are calculated. The focal loss is used for the classification loss to solve the imbalance problem of small target samples. The CIoU loss is used for the positioning loss to combine the overlap area, center point distance and aspect ratio difference between the predicted box and the real box to improve the positioning accuracy of small targets. The network parameters are optimized through end-to-end training, and the weights of the backbone network, the encoder and the decoder are updated synchronously during the training process, so that the model can adaptively learn the feature representation and detection logic of remote sensing small targets.

9. A remote sensing small target detection device based on an improved RT-DETR model, characterized in that, The image processing unit is configured to obtain a remote sensing image dataset as a training set, perform multi-scale transformation on the remote sensing images in the remote sensing image dataset to obtain remote sensing images with different resolutions, and then perform enhancement processing on the remote sensing images with different resolutions to obtain processed remote sensing images. The model training unit is configured to input the processed remote sensing images into an improved RT-DETR model to train the improved RT-DETR model and obtain a trained remote sensing small target detection network model. The backbone network of the improved RT-DETR model uses element-wise multiplication to process the features of the input remote sensing images to map the input features to a high-dimensional nonlinear space. The decoding process of the decoder of the improved RT-DETR model includes a feature extraction stage, a feature down-sampling stage, and a feature fusion stage. In the feature extraction stage, an extraction branch for small target features is introduced and configured to adapt to small target sizes by adjusting the convolution kernel size and step of the extraction branch. A dilated convolution is added to the extraction branch to expand the receptive field and obtain more rich context information. In the feature down-sampling stage, a SPDConv module is introduced to preserve small target detail information while reducing feature resolution through multi-branch sub-region feature splicing and convolution fusion of the SPDConv module. In the feature fusion stage, a CSPOmniKernel module is introduced to balance the small target feature extraction capability and computational complexity through feature splitting, multi-directional feature capturing, and attention enhancement of the CSPOmniKernel module. ​ A detection unit is configured to perform remote sensing small target detection on an input remote sensing image by using the remote sensing small target detection network model.

10. A remote sensing small target detection device based on an improved RT-DETR model, characterized in that, The application also provides a computer readable storage medium storing the computer program instructions of the remote sensing small target detection method based on the improved RT-DETR model.

Citation Information

Patent Citations

  • Multi-scale remote sensing image target detection method based on enhanced small target feature extraction

    CN117809200A

  • Unmanned aerial vehicle aerial photography target detection method and device based on deep learning

    CN120411820A

  • Unmanned aerial vehicle aerial photography small target detection method and system based on RT-DETR, medium and equipment

    CN120495641A

  • Improved lightweight diabetic retina fundus image small target detection method and system based on RT-DETR and application

    CN120612727A

  • Hyperspectral remote sensing image classification method based on self-attention context network

    US20230260279A1

Cited By

  • Unmanned aerial vehicle foggy day target detection method

    CN121170657A

  • Transformer oil particle identification method combining lensless reconstruction and depth detection model

    CN121686452A

  • A Transformer Oil Particle Identification Method Combining Lens-Free Reconstruction and Depth Detection Models

    CN121686452B

  • Aerial photography small target detection method and system based on YOLOv12n

    CN122200431A