Low-illumination small target detection method based on improved real-time detection Transform
By improving the backbone network and neck network structure of the RT-DETR model, combining feature optimization and attention mechanism, the accuracy and robustness of small object detection under low illumination conditions are solved, and higher detection accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202510450283.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-29
AI Technical Summary
The existing RT-DETR model is difficult to accurately distinguish low-illumination targets from complex backgrounds under low illumination conditions, and the detection accuracy and robustness of small targets are insufficient, making it prone to missed detection and missed detection.
A low-illumination small object detection model was constructed, the backbone network was improved as a feature extraction module based on residual attention, the neck network was improved as a bidirectional cascade feature fusion structure, and an attention mechanism based on channel feature optimization was introduced, and the NIS loss function was designed in combination with Shape-IOU and NWD loss functions, and iterative training was performed.
It significantly improves the detection performance of small targets under low illumination conditions, improves feature extraction ability and noise resistance, enhances the detection accuracy and adaptability of the model in complex backgrounds, and improves feature fusion ability and target positioning accuracy.
Smart Images

Figure CN120388265A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target detection technology, and in particular to a low-illumination small target detection method based on an improved real-time detection Transformer. Background Art
[0002] Object detection is crucial in the field of computer vision and is widely used in areas such as autonomous driving and face detection. Low-light images suffer from shortcomings such as low visibility, poor contrast, and color deviation. These shortcomings are caused by factors such as insufficient ambient light, limited scene illumination, and limitations in the performance of the camera. Traditional object detection methods are limited in accuracy in low-light conditions, leading to false detections and missed detections, and failing to meet precision requirements.
[0003] As an improved real-time detection Transformer model, the RT-DETR model has improved real-time performance; however, the RT-DETR model has many shortcomings in low-light target detection, which are specifically manifested as follows: in low-light environments with complex backgrounds, such as nighttime urban street scenes with light reflections and interweaving shadows, it is difficult to accurately distinguish the boundary between low-light targets and complex backgrounds, resulting in inaccurate target positioning, and it is easy to misjudge the background as a target or miss the target; for small targets with weak features under low-light conditions and easily submerged by noise, its general detection mechanism is difficult to fully capture subtle features, which greatly reduces the detection accuracy and results in many missed detections and false detections; when extracting features, it is easily affected by noise and lacks effective denoising or anti-noise mechanisms, resulting in inaccurate extracted target features, which in turn affects the accuracy of detection and classification.
[0004] To this end, a low-illumination small target detection method based on an improved real-time detection Transformer is proposed. Summary of the Invention
[0005] The purpose of the present invention is to provide a low-light small target detection method based on an improved real-time detection Transformer. The present invention relates to the technical field of target detection, and specifically provides a low-light small target detection method based on an improved real-time detection Transformer, including: constructing a low-light small target detection model, where the low-light small target detection model includes an improved backbone network, an improved neck network, and a decoder; the improved backbone network includes a feature extraction module based on residual attention, and the structure of the improved neck network is a bidirectional cascaded feature fusion structure; the bidirectional cascaded feature fusion structure includes an attention mechanism based on channel feature optimization; combining the Shape-IOU loss function and the NWD loss function to obtain the NIS loss function; using the trained low-light small target detection model to detect the target to be detected and obtaining the detection result; by improving the RT-DETR model, the present invention effectively improves the detection performance of small targets under low-light conditions at night.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A low-light small target detection method based on an improved real-time detection Transformer, including:
[0008] S1. Construct a low-light small target detection model. The low-light small target detection model uses the RT-DETR model as the benchmark model. The low-light small target detection model includes an improved backbone network, an improved neck network, and a decoder;
[0009] S2. The improved backbone network improves the feature extraction module of the backbone network of the RT-DETR model into a feature extraction module based on residual attention; the improved neck network improves the neck network of the RT-DETR model into a bidirectional cascaded feature fusion structure; the bidirectional cascaded feature fusion structure includes an attention mechanism based on channel feature optimization; the low-light small target detection model includes an NIS loss function; the acquisition process of the NIS loss function is: obtaining the Inner-ShapeIoU loss function based on the Shape-IOU loss function, and combining the Inner-ShapeIoU loss function and the NWD loss function to obtain the NIS loss function;
[0010] S3. Iteratively train the low-light small target detection model to obtain the trained low-light small target detection model;
[0011] S4. Use the trained low-light small target detection model to detect the target to be detected and obtain the detection result.
[0012] Preferably, the feature extraction module based on residual attention is used to extract features from the first input image and denoise it to obtain the neck input feature map.
[0013] Preferably, the specific process of obtaining the neck input feature map includes:
[0014] Generate multiple intermediate feature maps by performing multiple 3×3 convolutions on the first input image, and splice the first input image and the multiple intermediate feature maps to obtain the first spliced feature map;
[0015] Reduce the number of channels of the first spliced feature map according to 1×1 convolution to obtain the first feature map; process the first spliced feature map according to the 3×3 convolutional layer, batch normalization and rectified linear unit to obtain the second feature map;
[0016] Extract multi-scale features from the second feature map according to three convolutional layers with different dilation rates to obtain the first multi-scale feature map; process the first multi-scale feature map according to the 1×1 convolutional layer and batch normalization to obtain the third feature map; add the first feature map and the third feature map to obtain the first fused feature map; adjust the number of channels of the first fused feature map according to the ESE channel adjustment module to obtain the second fused feature map; add the second fused feature map and the first input image, and pass through the RAB denoising module to obtain the neck input feature map.
[0017] Preferably, the improved backbone network includes the first neck input feature map, the second neck input feature map, the third neck input feature map, the fourth neck input feature map, the fifth neck input feature map and the sixth neck input feature map.
[0018] Preferably, the fusion principle steps of the bidirectional cascaded feature fusion structure include:
[0019] S10. Reduce the number of channels of the first neck input feature map by 1×1 convolution, and perform downsampling operation using 3×3 convolution with a stride of 2;
[0020] S20. Reduce the number of channels of the small-scale feature map in the second neck input feature map by 1×1 convolution;
[0021] S30. Reduce the number of channels of the large-scale feature map in the second neck input feature map by 1×1 convolution, and perform downsampling operation using 3×3 convolution with a stride of 2;
[0022] S40. Reduce the number of channels of the fifth neck input feature map by 1×1 convolution;
[0023] S50. Perform 1×1 convolution channel reduction on the sixth neck input feature map, and respectively pass through the internal scale interaction module based on the attention mechanism, 1×1 convolution channel reduction, the attention mechanism based on channel feature optimization, and 2×2 transposed convolution operation;
[0024] S60. Perform feature fusion on the feature maps obtained in S30, S40, and S50, and respectively pass through 1×1 convolution channel reduction, feature refinement, 1×1 convolution channel reduction, the attention mechanism based on channel feature optimization, and 2×2 transposed convolution operation;
[0025] S70. Perform feature fusion on the three feature maps obtained in S10, S20, and S60, perform 1×1 convolution channel reduction and feature refinement, and input to the decoder;
[0026] S80. Perform downsampling operation on the feature map obtained in S70 according to 3×3 convolution, perform feature fusion with the feature map after the operation of the attention mechanism based on channel feature optimization in S60, perform feature refinement, and input to the decoder;
[0027] S90. Perform 3×3 convolution downsampling operation on the feature map obtained in S80, perform feature fusion with the feature map after the operation of the attention mechanism based on channel feature optimization in S50, and input to the decoder after refining the features.
[0028] Preferably, the bidirectional cascaded feature fusion structure includes the attention mechanism based on channel feature optimization; the attention mechanism based on channel feature optimization includes an input module, BH-FPN, a spatial attention module, and an output module.
[0029] Preferably, the specific process of the BH-FPN includes:
[0030] S100. Divide the input feature map F ∈ R H×W×C into N×N non-overlapping regions, and then perform linear projection on each regional feature map to obtain query Q, key K, and value V tensors;
[0031] S200. Obtain the adjacency matrix U by normalizing the query Q and the key K and performing matrix multiplication;
[0032] S300. Perform matrix multiplication on U and V and process them in combination with the attention mechanism to obtain the first output matrix F1;
[0033] S400. Perform scale division on the first output matrix F1 to obtain four first feature matrices of different scales; the first feature matrix includes multiple channels; input the first feature matrix into the feature selection module, and the feature selection module performs feature fusion on the four first feature matrices of different scales according to global average pooling and global maximum pooling respectively, and uses the Sigmoid activation function to determine the weight values of the multiple channels, and then processes the weight values according to the channel attention mechanism; obtain four second feature matrices; multiply the four second feature matrices with the four first feature matrices of different scales respectively, and perform 1×1 convolution for channel dimensionality reduction processing to obtain a first high-level feature matrix and a first low-level feature matrix;
[0034] S500. Perform an expansion operation on the first high-level feature matrix using transposed convolution to obtain a second high-level feature matrix; process the second high-level feature matrix through the feature selection module to obtain a third high-level feature matrix, and perform a filtering operation on the first low-level feature matrix through the third high-level feature matrix to obtain a second low-level feature matrix, and perform matrix addition according to the second low-level feature matrix and the second high-level feature matrix to obtain a first output matrix, and further obtain a first output feature map F2;
[0035] S600. Perform an element-wise multiplication operation on F2 and F to generate F3 and input it into the spatial attention module, process the F3 through the spatial attention module to generate a spatial attention map A, and perform an element-wise multiplication operation on A and F3 to obtain the final output feature map;
[0036] The Shape-IoU loss function is:
[0037] L ShapeIoU = 1 - IoU + distance shape + 0.5 * Ω shape ;
[0038]
[0039] where L ShapeIoU represents the Shape-IoU loss function; IoU represents the intersection over union; distance shape represents the distance metric in the shape space; Ω shape represents the similarity metric between shapes; B and B gt are the predicted bounding box and the ground truth bounding box respectively; ∩ represents the intersection operation; ∪ represents the union operation; ww and hh are the weight coefficients in the horizontal and vertical directions respectively; x c gt and y c gtare the abscissa and ordinate of the center point of the predicted bounding box; x c and y c are the abscissa and ordinate of the center point of the ground truth bounding box respectively, c is a normalization constant; scale is a scaling factor; w and h represent the width and height of the predicted bounding box respectively; w gt and h gt represent the width and height of the ground truth bounding box respectively; e is the natural constant; ω w and ω h are parameters used to measure the degree of difference in the width and height directions of the shape respectively; max() represents the operation of taking the maximum value.
[0040] Preferably, the NWD loss function is:
[0041] L NWD = 1 - NWD(N p , N g );
[0042]
[0043] where, L NWD represents the NWD loss function; NWD(N p , N g ) represents the normalized Wasserstein distance; N p and N g are the two-dimensional Gaussian distribution models of the predicted bounding box and the ground truth bounding box respectively, W2 2 is the second-order distribution distance, C is a constant; exp is the exponential function with the constant e as the base.
[0044] Preferably, the NIS loss function:
[0045] L NIS = λ * L NWD + (1 - λ) * L Inner-ShapeIoU ;
[0046] L Inner-ShapeIoU = L ShapeIoU + IoU - IoU inner ;
[0047] where, L NIS represents the NIS loss function; λ represents the weight coefficient; L NWD represents the NWD loss function; L Inner-ShapeIoU represents the Inner - ShapeIoU loss function; L ShapeIoU represents the Shape - IoU loss function; IoU represents the intersection over union; IoU inner represents the internal intersection over union.
[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0049] 1. The RT-DETR model of the present invention is used as the benchmark model. The feature extraction module of the backbone network of the RT-DETR model is improved into a feature extraction module based on residual attention, which effectively realizes the feature screening and analysis in the feature map, can focus on more critical and representative feature regions in the image, thereby significantly enhancing the overall feature extraction ability. And through the noise reduction module, the interference of irrelevant information can be effectively suppressed, making the extracted features more pure and effective, providing a high-quality feature basis for subsequent model processing. The feature extraction ability and noise reduction performance are significantly enhanced, and the accuracy of small target detection is further improved.
[0050] 2. The present invention improves the neck structure of the RT-DETR model, which is designed as a bidirectional cascaded feature fusion structure, enabling features at different levels to interact and fuse in multiple directions. During the process of transmitting from low-level features to high-level features, the rich detailed information of low-level features can be fully utilized, supplementing accurate position and texture information for high-level features; while during the process of feeding back from high-level features to low-level features, the abstract semantic information contained in high-level features can guide the optimization and integration of low-level features, improving the semantic consistency and integrity of the overall features, thereby effectively improving the feature fusion ability of RT-DETR, enabling the model to better understand and process complex information in the image, and further improving the accuracy of small target detection.
[0051] 3. The present invention introduces an attention mechanism based on channel feature optimization into the bidirectional cascaded feature fusion structure. This mechanism can more precisely enhance the expression of important features through channel and spatial attention mechanisms, enabling the network to more sensitively identify and utilize key feature information when facing complex and variable image data, thereby improving the performance of the entire network model, showing stronger adaptability and accuracy in target detection tasks, and further improving the accuracy of small target detection.
[0052] 4. The present invention realizes the precise focusing of the shape and scale of the bounding box through the improved NIS loss function, enabling the model to accurately learn relevant features to reduce the positioning error, and can use different scales of auxiliary bounding boxes for different samples to speed up the regression process, which is beneficial for dealing with complex scenarios such as occlusion and overlap. When detecting tiny objects, its smoothness can stabilize the position regression, and can handle non-overlapping or mutually inclusive bounding boxes to improve the recall rate and accuracy, further improving the accuracy of small target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a schematic flow chart of a low-light target detection method based on an improved real-time detection Transformer provided by an embodiment of the present invention;
[0054] Figure 2 It is a structural diagram of a low-light small target detection model provided by an embodiment of the present invention;
[0055] Figure 3 It is a structural diagram of a feature extraction module based on residual attention provided by an embodiment of the present invention;
[0056] Figure 4 It is a schematic diagram of an attention mechanism based on channel feature optimization provided by an embodiment of the present invention;
[0057] Figure 5 It is a flowchart of a calculation method for the ShapeIoU loss function provided by an embodiment of the present invention. Specific implementation manners
[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0059] Embodiment 1
[0060] In order to improve the detection accuracy of small target A, a low-light target detection method based on an improved real-time detection Transformer is applied;
[0061] Refer to Figure 1 , which is a schematic flowchart of a low-light target detection method based on an improved real-time detection Transformer provided by an embodiment of the present invention, including:
[0062] S1. Construct a low-light small target detection model. The low-light small target detection model is based on the RT-DETR model as the benchmark model. The low-light small target detection model includes an improved backbone network, an improved neck network, and a decoder;
[0063] S2. The improved backbone network improves the feature extraction module of the backbone network of the RT-DETR model into a feature extraction module based on residual attention; the improved neck network improves the neck network of the RT-DETR model into a bidirectional cascaded feature fusion structure; the bidirectional cascaded feature fusion structure includes an attention mechanism based on channel feature optimization; the low-light small target detection model includes an NIS loss function; the acquisition process of the NIS loss function is as follows: obtaining an Inner-ShapeIoU loss function based on the Shape-IOU loss function, and combining the Inner-ShapeIoU loss function and the NWD loss function to obtain the NIS loss function; referring to Figure 2 , which is a structural diagram of a low-light small target detection model provided by an embodiment of the present invention;
[0064] Further, the feature extraction module based on residual attention is used to extract and denoise the features of the first input image to obtain a neck input feature map; referring to Figure 3 , which is a structural diagram of a feature extraction module based on residual attention provided by an embodiment of the present invention;
[0065] Further, the specific acquisition process of the neck input feature map includes:
[0066] Generating a plurality of intermediate feature maps by performing multiple 3×3 convolutions on the first input image, and splicing the first input image and the plurality of intermediate feature maps to obtain a first spliced feature map;
[0067] Reducing the number of channels of the first spliced feature map according to 1×1 convolution to obtain a first feature map; processing the first spliced feature map according to a 3×3 convolutional layer, batch normalization, and a rectified linear unit function to obtain a second feature map;
[0068] Performing multi-scale feature extraction on the second feature map according to three convolutional layers with different dilation rates to obtain a first multi-scale feature map; processing the first multi-scale feature map according to a 1×1 convolutional layer and batch normalization to obtain a third feature map; adding the first feature map and the third feature map to obtain a first fused feature map; adjusting the number of channels of the first fused feature map according to the ESE channel adjustment module to obtain a second fused feature map; adding the second fused feature map and the first input image, and passing through the RAB denoising module to obtain a neck input feature map.
[0069] Further, the different dilation rates in this embodiment respectively include 1, 3, and 5.
[0070] In this embodiment, the RT-DETR model serves as the benchmark model. The feature extraction module of the backbone network of the RT-DETR model is improved to a feature extraction module based on residual attention, effectively realizing the feature screening and analysis in the feature map, being able to focus on more critical and representative feature regions in the image, thereby significantly enhancing the overall feature extraction ability. And through the noise reduction module, the interference of irrelevant information can be effectively suppressed, making the extracted features more pure and effective, providing a high-quality feature basis for subsequent model processing. The feature extraction ability and noise reduction performance are significantly enhanced, further improving the accuracy of small target detection.
[0071] Further, the improved backbone network includes the first neck input feature map, the second neck input feature map, the third neck input feature map, the fourth neck input feature map, the fifth neck input feature map, and the sixth neck input feature map; corresponding to Figure 2 1:FE-RA, 3:FE-RA, 5:FE-RA, 6:FE-RA, 7:FE-RA, and 9:FE-RA in
[0072] Further, the fusion principle steps of the bidirectional cascaded feature fusion structure include:
[0073] S10. Perform 1×1 convolution channel reduction on the first neck input feature map, and perform downsampling operation using 3×3 convolution with a stride of 2;
[0074] S20. Perform 1×1 convolution channel reduction on the small-scale feature map in the second neck input feature map;
[0075] S30. Perform 1×1 convolution channel reduction on the large-scale feature map in the second neck input feature map, and perform downsampling operation using 3×3 convolution with a stride of 2;
[0076] S40. Perform 1×1 convolution channel reduction on the fifth neck input feature map;
[0077] S50. Perform 1×1 convolution channel reduction on the sixth neck input feature map, and respectively pass through the internal scale interaction module based on the attention mechanism, 1×1 convolution channel reduction, the attention mechanism based on channel feature optimization, and 2×2 transposed convolution operation;
[0078] S60. Perform feature fusion on the feature maps obtained from S30, S40, and S50, and respectively pass through 1×1 convolution channel reduction, feature refinement, 1×1 convolution channel reduction, the attention mechanism based on channel feature optimization, and 2×2 transposed convolution operation;
[0079] S70. Perform feature fusion on the three feature maps obtained from S10, S20, and S60, perform 1×1 convolution channel reduction and feature refinement, and input them into the decoder;
[0080] S80. Downsample the feature map obtained in S70 through a 3×3 convolution, perform feature fusion with the feature map after the operation of the attention mechanism optimized based on channel features in S60, refine the features, and input them into the decoder;
[0081] S90. Perform a 3×3 convolution downsampling operation on the feature map obtained in S80, perform feature fusion with the feature map after the operation of the attention mechanism optimized based on channel features in S50, and input it into the decoder after refining the features;
[0082] Furthermore, the flow of the fusion principle steps of the bidirectional cascaded feature fusion structure is as Figure 2 shown in the improved neck network in
[0083] In this embodiment, the neck structure of the RT-DETR model is improved and designed as a bidirectional cascaded feature fusion structure, enabling features at different levels to interact and fuse in multiple directions. During the process of transmitting from low-level features to high-level features, the rich detailed information of low-level features can be fully utilized, supplementing accurate position and texture information for high-level features; while in the process of feedback from high-level features to low-level features, the abstract semantic information contained in high-level features can guide the optimization and integration of low-level features, enhancing the semantic consistency and integrity of the overall features, thereby effectively improving the feature fusion ability of RT-DETR, enabling the model to better understand and process complex information in images, and further improving the accuracy of small object detection.
[0084] Furthermore, the bidirectional cascaded feature fusion structure includes an attention mechanism optimized based on channel features; referring to Figure 4 , which is a schematic diagram of an attention mechanism optimized based on channel features provided by an embodiment of the present invention;
[0085] The attention mechanism optimized based on channel features includes an input module, BH-FPN, a spatial attention module, and an output module.
[0086] Furthermore, the specific process of BH-FPN includes:
[0087] S100. Divide the input feature map F∈R H×W×C into N×N non-overlapping regions, and represent the feature map of each region as Then perform linear projection on the feature map of each region to obtain query Q, key K, and value V tensors, H is the height; W is the width; C is the number of channels;
[0088] S200. Obtain the adjacency matrix U by normalizing and performing matrix multiplication on the query Q and the key K;
[0089] S300. Through matrix multiplication operations on U and V and processing in combination with the attention mechanism, the first output matrix F1 is obtained;
[0090] S400. Perform scale division on the first output matrix F1 to obtain four first feature matrices of different scales; the first feature matrix includes multiple channels; input the first feature matrix into the feature selection module, and the feature selection module respectively performs feature fusion on the four first feature matrices of different scales according to global average pooling and global max pooling, and uses the Sigmoid activation function to determine the weight values of the multiple channels, and then processes the weight values according to the channel attention mechanism; four second feature matrices are obtained; the four second feature matrices are respectively multiplied by the four first feature matrices of different scales in matrix multiplication, and 1×1 convolution is performed for channel dimensionality reduction processing to obtain the first high-level feature matrix and the first low-level feature matrix;
[0091] The four first feature matrices of different scales are as Figure 4 S5, S4, S3, and S2 in Figure 4 ; the feature selection module is as Figure 4 CA in
[0092] Analyzed by P4 and P5 in Figure 4 , P5 is the first high-level feature matrix, and P4 is the first low-level feature matrix; Figure 4 S500. Perform an expansion operation on the first high-level feature matrix using transposed convolution to obtain the second high-level feature matrix; the second high-level feature matrix is as Figure 4 N5 in Figure 4
[0093] ; the second high-level feature matrix is processed by the feature selection module to obtain the third high-level feature matrix, and the first low-level feature matrix is filtered by the third high-level feature matrix to obtain the second low-level feature matrix, and the second low-level feature matrix and the second high-level feature matrix are added in matrix addition to obtain the first output matrix (such as Figure 4 N4 in Figure 4 ), and then the first output feature map F2 is obtained;
[0093] S600. Perform an element-wise multiplication operation on F2 and F to generate F3 and input it into the spatial attention module. The spatial attention module processes the F3 to generate the spatial attention map A, and performs an element-wise multiplication operation on A and F3 to obtain the final output feature map;
[0094] The output module is used to output the final output feature map;
[0095] In this embodiment, an attention mechanism based on channel feature optimization (AM-CFO) is introduced into the bidirectional cascaded feature fusion structure. This mechanism can more accurately strengthen the expression of important features through channel and spatial attention mechanisms, enabling the network to more sensitively identify and utilize key feature information when facing complex and variable image data, thereby improving the performance of the entire network model, showing stronger adaptability and accuracy in object detection tasks, and further improving the accuracy of small object detection.
[0096] The Shape-IoU loss function is as follows:
[0097] L ShapeIoU = 1 - IoU + distance shape + 0.5 * Ω shape ;
[0098]
[0099] Among them, L ShapeIoU represents the Shape-IoU loss function; IoU represents the intersection over union; distance shape represents the distance metric in the shape space; Ω shape represents the similarity metric between shapes; B and B gt are the predicted bounding box and the ground truth bounding box respectively; ∩ represents the intersection operation; ∪ represents the union operation; ww and hh are the weight coefficients in the horizontal and vertical directions respectively; x c gt and y c gt are the abscissa and ordinate of the center point of the predicted bounding box respectively; x c and y c are the abscissa and ordinate of the center point of the ground truth bounding box respectively, c is the normalization constant; scale is the scale factor; w and h represent the width and height of the predicted bounding box respectively; w gt and h gt represent the width and height of the ground truth bounding box respectively; e is the natural constant; ω w and ω h are parameters used to measure the degree of difference in the width and height directions of the shape respectively; max() represents the maximum value operation.
[0100] Referring to Figure 5 , it is a flowchart of the calculation method of a ShapeIoU loss function provided by an embodiment of the present invention; among them, GT represents the ground truth bounding box; Anchor represents the predicted bounding box;
[0101] Furthermore, the NWD loss function is as follows:
[0102] L NWD = 1 - NWD(Np ,N g );
[0103]
[0104] Among them, L NWD represents the NWD loss function; NWD(N p ,N g ) represents the normalized Wasserstein distance; N p and N g are respectively the two-dimensional Gaussian distribution models of the predicted box and the ground truth box, W2 2 is the second-order distribution distance, C is a constant; exp is the exponential function with the constant e as the base.
[0105] Furthermore, the NIS loss function:
[0106] L NIS = λ * L NWD +(1 - λ) * L Inner-ShapeIoU ;
[0107] L Inner-ShapeIoU = L ShapeIoU + IoU - IoU inner ;
[0108] Among them, L NIS represents the NIS loss function; λ represents the weight coefficient; L NWD represents the NWD loss function; L Inner-ShapeIoU represents the Inner - ShapeIoU loss function; L ShapeIoU represents the Shape - IoU loss function; IoU represents the intersection over union; IoU inner represents the internal intersection over union.
[0109] S3. Iteratively train the low - illumination small target detection model to obtain the trained low - illumination small target detection model;
[0110] The low - illumination small target detection model is iteratively trained using the dark dataset, and the SGD (Stochastic Gradient Descent) optimizer is used to optimize the model. The batch normalization size is 16, and the total number of training epochs is 120 times to obtain the trained low - illumination small target detection model;
[0111] S4. Use the trained low - illumination small target detection model to detect the target to be detected to obtain the detection result.
[0112] To verify the effectiveness of a low-light target detection method based on an improved real-time detection Transformer provided in this embodiment for small target detection tasks, multiple small targets in a low-light environment are detected using Model 1, Model 2, Model 3, and Model 4 respectively. Model 1 is a low-light small target detection model provided in this embodiment; Model 2 is an improvement without considering the backbone network based on Model 1; Model 3 is an improvement without considering the neck network based on Model 1; Model 4 is an improvement without considering the loss function based on Model 1; Through comparative analysis of the average accuracy of multiple small target detections A, the specific results are shown in Table 1;
[0113] Table 1 Comparison table of the detection accuracy of multiple small targets A in a low-light environment for different models
[0114] Model Average accuracy Model 1 97% Model 2 90% Model 3 93% Model 4 92%
[0115] As can be seen from Table 1, a low-light small target detection model provided in this embodiment has certain effectiveness.
[0116] In this embodiment, an improved NIS loss function is used to achieve precise focusing on the shape and scale of the bounding box, enabling the model to accurately learn relevant features to reduce the positioning error. It can use different scales of auxiliary bounding boxes for different samples to speed up the regression process, which is beneficial for dealing with complex scenarios such as occlusion and overlap. When detecting tiny objects, its smoothness can stabilize the position regression, and it can handle non-overlapping or mutually inclusive bounding boxes to improve the recall rate and accuracy, further improving the accuracy of small target detection.
[0117] In this embodiment, a low-light small target detection model is constructed. The low-light small target detection model includes an improved backbone network, an improved neck network, and a decoder; The improved backbone network includes a feature extraction module based on residual attention, and the structure of the improved neck network is a bidirectional cascaded feature fusion structure; The bidirectional cascaded feature fusion structure includes an attention mechanism based on channel feature optimization; The Inner-ShapeIoU loss function is obtained based on the Shape-IOU loss function, and the Inner-ShapeIoU loss function and the NWD loss function are combined to obtain the NIS loss function; The trained low-light small target detection model is used to detect the target to be detected to obtain the detection result; The present invention improves the RT-DETR model to obtain a low-light small target detection model, effectively improving the detection performance of small targets under low-light conditions at night.
[0118] Example Two
[0119] To improve the detection accuracy of small target B, a low-light target detection method based on an improved real-time detection Transformer is applied;
[0120] Refer toFigure 1 , which is a schematic flow chart of a low-light target detection method based on an improved real-time detection Transformer provided by an embodiment of the present invention, including:
[0121] S1. Construct a low-light small target detection model. The low-light small target detection model uses the RT-DETR model as the benchmark model. The low-light small target detection model includes an improved backbone network, an improved neck network, and a decoder;
[0122] S2. The improved backbone network improves the feature extraction module of the backbone network of the RT-DETR model into a feature extraction module based on residual attention; the improved neck network improves the neck network of the RT-DETR model into a bidirectional cascaded feature fusion structure; the bidirectional cascaded feature fusion structure includes an attention mechanism based on channel feature optimization; the low-light small target detection model includes an NIS loss function; the acquisition process of the NIS loss function is: obtaining an Inner-ShapeIoU loss function based on the Shape-IOU loss function, and combining the Inner-ShapeIoU loss function and the NWD loss function to obtain the NIS loss function; refer to Figure 2 , which is a structural diagram of a low-light small target detection model provided by an embodiment of the present invention;
[0123] Further, the feature extraction module based on residual attention is used to extract and denoise the features of the first input image to obtain a neck input feature map; refer to Figure 3 , which is a structural diagram of a feature extraction module based on residual attention provided by an embodiment of the present invention;
[0124] Further, the specific acquisition process of the neck input feature map includes:
[0125] Generating a plurality of intermediate feature maps by performing multiple 3×3 convolutions on the first input image, and splicing the first input image and the plurality of intermediate feature maps to obtain a first spliced feature map;
[0126] Reducing the channels of the first spliced feature map according to 1×1 convolution to obtain a first feature map; processing the first spliced feature map according to a 3×3 convolutional layer, batch normalization, and a rectified linear unit function to obtain a second feature map;
[0127] Perform multi-scale feature extraction on the second feature map using three convolutional layers with different dilation rates respectively to obtain the first multi-scale feature map; process the first multi-scale feature map using a 1×1 convolutional layer and batch normalization to obtain the third feature map; add the first feature map and the third feature map to obtain the first fused feature map; adjust the number of channels of the first fused feature map using the ESE channel adjustment module to obtain the second fused feature map; add the second fused feature map and the first input image and pass through the RAB denoising module to obtain the neck input feature map.
[0128] Further, the different dilation rates in this embodiment respectively include 1, 3, and 5.
[0129] The RT-DETR model in this embodiment is a baseline model. The feature extraction module of the backbone network of the RT-DETR model is improved to a feature extraction module based on residual attention, effectively realizing the feature screening and analysis in the feature map, being able to focus on more critical and representative feature regions in the image, thereby significantly enhancing the overall feature extraction ability. And through the denoising module, the interference of irrelevant information can be effectively suppressed, making the extracted features more pure and effective, providing a high-quality feature basis for subsequent model processing. Significantly enhancing the feature extraction ability and denoising performance, further improving the accuracy of small target detection.
[0130] Further, the improved backbone network includes the first neck input feature map, the second neck input feature map, the third neck input feature map, the fourth neck input feature map, the fifth neck input feature map, and the sixth neck input feature map; corresponding respectively Figure 2 to 1:FE-RA, 3:FE-RA, 5:FE-RA, 6:FE-RA, 7:FE-RA, and 9:FE-RA in
[0131] Further, the fusion principle steps of the bidirectional cascaded feature fusion structure include:
[0132] S10. Perform 1×1 convolutional channel reduction on the first neck input feature map, and perform downsampling operation using a 3×3 convolutional layer with a stride of 2;
[0133] S20. Perform 1×1 convolutional channel reduction on the small-scale feature map in the second neck input feature map;
[0134] S30. Perform 1×1 convolutional channel reduction on the large-scale feature map in the second neck input feature map, and perform downsampling operation using a 3×3 convolutional layer with a stride of 2;
[0135] S40. Perform 1×1 convolutional channel reduction on the fifth neck input feature map;
[0136] S50. Perform 1×1 convolution channel reduction on the sixth neck input feature map, and respectively pass through the internal scale interaction module based on the attention mechanism, 1×1 convolution channel reduction, the attention mechanism based on channel feature optimization, and 2×2 transposed convolution operation;
[0137] S60. Perform feature fusion on the feature maps obtained in S30, S40, and S50, and respectively pass through 1×1 convolution channel reduction, feature refinement, 1×1 convolution channel reduction, the attention mechanism based on channel feature optimization, and 2×2 transposed convolution operation;
[0138] S70. Perform feature fusion on the three feature maps obtained in S10, S20, and S60, perform 1×1 convolution channel reduction and feature refinement, and input to the decoder;
[0139] S80. Perform downsampling operation on the feature map obtained in S70 according to 3×3 convolution, perform feature fusion with the feature map after the operation of the attention mechanism based on channel feature optimization in S60, perform feature refinement, and input to the decoder;
[0140] S90. Perform 3×3 convolution downsampling operation on the feature map obtained in S80, perform feature fusion with the feature map after the operation of the attention mechanism based on channel feature optimization in S50, and input to the decoder after refining the features;
[0141] Further, the flow of the fusion principle steps of the bidirectional cascaded feature fusion structure is as Figure 2 shown in the improved neck network;
[0142] In this embodiment, the neck structure of the RT-DETR model is improved and designed as a bidirectional cascaded feature fusion structure, enabling features at different levels to interact and fuse in multiple directions. During the process of transmitting from low-level features to high-level features, the rich detailed information of low-level features can be fully utilized, supplementing accurate position and texture information for high-level features; while in the process of feedback from high-level features to low-level features, the abstract semantic information contained in high-level features can guide the optimization and integration of low-level features, enhancing the semantic consistency and integrity of the overall features, thereby effectively improving the feature fusion ability of RT-DETR, enabling the model to better understand and process complex information in images, and further improving the accuracy of small target detection.
[0143] Further, the bidirectional cascaded feature fusion structure includes an attention mechanism based on channel feature optimization; refer to Figure 4 , which is a schematic diagram of an attention mechanism based on channel feature optimization provided by an embodiment of the present invention;
[0144] The attention mechanism optimized based on channel features includes an input module, BH-FPN, a spatial attention module, and an output module.
[0145] Furthermore, the specific process of the BH-FPN includes:
[0146] S100. Divide the input feature map F ∈ R H×W×C into N×N non-overlapping regions, and the feature maps of each region are expressed as Then, perform linear projection on the feature maps of each region to obtain query Q, key K, and value V tensors. H is the height; W is the width; C is the number of channels;
[0147] S200. Obtain the adjacency matrix U by normalizing the query Q and the key K and performing matrix multiplication.
[0148] S300. Perform matrix multiplication on U and V and process them in combination with the attention mechanism to obtain the first output matrix F1.
[0149] S400. Perform scale division on the first output matrix F1 to obtain four first feature matrices of different scales; the first feature matrices include multiple channels; input the first feature matrices into a feature selection module, and the feature selection module performs feature fusion on the four first feature matrices of different scales respectively according to global average pooling and global maximum pooling, and uses the Sigmoid activation function to determine the weight values of the multiple channels, and then processes the weight values according to the channel attention mechanism; obtain four second feature matrices; the four second feature matrices are respectively multiplied by the four first feature matrices of different scales in matrix multiplication, and 1×1 convolution is performed for channel dimensionality reduction processing to obtain a first high-level feature matrix and a first low-level feature matrix.
[0150] The four first feature matrices of different scales are as Figure 4 S5, S4, S3, and S2 in; the feature selection module is as Figure 4 CA in; taking Figure 4 P4 and P5 in as an analysis, P5 is the first high-level feature matrix, and P4 is the first low-level feature matrix.
[0151] S500. Perform an expansion operation on the first high-level feature matrix using transposed convolution to obtain a second high-level feature matrix; the second high-level feature matrix is as Figure 4N5 in; the second high-level feature matrix is processed by the feature selection module to obtain a third high-level feature matrix, and the first low-level feature matrix is filtered by the third high-level feature matrix to obtain a second low-level feature matrix, and matrix addition is performed according to the second low-level feature matrix and the second high-level feature matrix to obtain a first output matrix (such as Figure 4 N4 in), and then a first output feature map F2 is obtained;
[0152] S600. Perform an element-wise multiplication operation on F2 and F to generate F3 and input it into the spatial attention module. The F3 is processed by the spatial attention module to generate a spatial attention map A, and an element-wise multiplication operation is performed on A and F3 to obtain the final output feature map;
[0153] The output module is used to output the final output feature map;
[0154] In this embodiment, an attention mechanism based on channel feature optimization is introduced into the bidirectional cascaded feature fusion structure. This mechanism can more accurately strengthen the expression of important features through the channel and spatial attention mechanisms, enabling the network to more sensitively recognize and utilize key feature information when facing complex and variable image data, thereby improving the performance of the entire network model, showing stronger adaptability and accuracy in the object detection task, and further improving the accuracy of small object detection.
[0155] The Shape-IoU loss function is:
[0156] L ShapeIoU = 1 - IoU + distance shape + 0.5 * Ω shape ;
[0157]
[0158]
[0159] where L ShapeIoU represents the Shape-IoU loss function; IoU represents the intersection over union; distance shape represents the distance metric in the shape space; Ω shape represents the similarity metric between shapes; B and B gt are the predicted box and the ground truth box respectively; ∩ represents the intersection operation; ∪ represents the union operation; ww and hh are the weight coefficients in the horizontal and vertical directions respectively; x c gt and y c gt are the abscissa and ordinate of the center point of the predicted box respectively; x c and y care the abscissa and ordinate of the center point of the ground truth box respectively, c is a normalization constant; scale is a scaling factor; w and h represent the width and height of the predicted box respectively; w gt and h gt represent the width and height of the ground truth box respectively; e is the natural constant; ω w and ω h are parameters used to measure the degree of difference in the width and height directions of the shape respectively; max() represents the operation of taking the maximum value.
[0160] Refer to Figure 5 , which is a flowchart of a calculation method for a ShapeIoU loss function provided by an embodiment of the present invention; where GT represents the ground truth box; Anchor represents the predicted box;
[0161] Furthermore, the NWD loss function is:
[0162] L NWD = 1 - NWD(N p , N g );
[0163]
[0164] where L NWD represents the NWD loss function; NWD(N p , N g ) represents the normalized Wasserstein distance; N p and N g are the two-dimensional Gaussian distribution models of the predicted box and the ground truth box respectively, W2 2 is the second-order distribution distance, C is a constant; exp is the exponential function with the constant e as the base.
[0165] Furthermore, the NIS loss function:
[0166] L NIS = λ * L NWD + (1 - λ) * L Inner-ShapeIoU ;
[0167] L Inner-ShapeIoU = L ShapeIoU + IoU - IoU inner ;
[0168] where L NIS represents the NIS loss function; λ represents the weight coefficient; L NWD represents the NWD loss function; L Inner-ShapeIoU represents the Inner-ShapeIoU loss function; L ShapeIoU represents the Shape-IoU loss function; IoU represents the intersection over union; IoUinner Represents the internal intersection over union.
[0169] S3. Iteratively train the low-light small target detection model to obtain the trained low-light small target detection model.
[0170] The low-light small target detection model is iteratively trained using the dark dataset, and the SGD (Stochastic Gradient Descent) optimizer is used to optimize the model. The batch normalization size is 16, and the total number of training cycles is 120 times to obtain the trained low-light small target detection model.
[0171] S4. Use the trained low-light small target detection model to detect the target to be detected and obtain the detection result.
[0172] To verify the effectiveness of a low-light target detection method based on an improved real-time detection Transformer for small target detection tasks, multiple small targets in a low-light environment are detected using Model 1, Model 2, Model 3, and Model 4 respectively. Model 1 is a low-light small target detection model provided in this embodiment; Model 2 is an improvement without considering the backbone network based on Model 1; Model 3 is an improvement without considering the neck network based on Model 1; Model 4 is an improvement without considering the loss function based on Model 1; through comparative analysis of the average detection accuracy of multiple small targets B, the specific results are shown in Table 1.
[0173] Table 2 Comparison table of the detection accuracy of multiple small targets B in a low-light environment by different models
[0174] Model Average accuracy Model 1 96% Model 2 89% Model 3 91% Model 4 90%
[0175] As can be seen from Table 2, a low-light small target detection model provided in this embodiment has certain effectiveness.
[0176] In this embodiment, the improved NIS loss function is used to achieve precise focusing on the shape and scale of the bounding box, enabling the model to accurately learn relevant features to reduce the positioning error. It can use different scales of auxiliary bounding boxes for different samples to speed up the regression process, which is beneficial for dealing with complex scenarios such as occlusion and overlap. When detecting tiny objects, its smoothness can stabilize the position regression, and it can handle non-overlapping or mutually inclusive bounding boxes to improve the recall rate and accuracy, further improving the accuracy of small target detection.
[0177] In this embodiment, a low-light small target detection model is constructed. The low-light small target detection model includes an improved backbone network, an improved neck network, and a decoder. The improved backbone network includes a feature extraction module based on residual attention. The structure of the improved neck network is a bidirectional cascaded feature fusion structure. The bidirectional cascaded feature fusion structure includes an attention mechanism based on channel feature optimization. An Inner-ShapeIoU loss function is obtained based on the Shape-IOU loss function, and the Inner-ShapeIoU loss function and the NWD loss function are combined to obtain the NIS loss function. The trained low-light small target detection model is used to detect the target to be detected, and the detection result is obtained. By improving the RT-DETR model, the low-light small target detection model is obtained, effectively improving the detection performance of small targets under low-light conditions at night.
[0178] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A low-light small target detection method based on an improved real-time detection Transformer, characterized in that, Including: S1. Construct a low-light small target detection model, which is based on the RT-DETR model as the benchmark model. The low-light small target detection model includes an improved backbone network, an improved neck network, and a decoder; S2. The improved backbone network improves the feature extraction module of the backbone network of the RT-DETR model into a feature extraction module based on residual attention; The improved neck network improves the neck network of the RT-DETR model into a bidirectional cascaded feature fusion structure; the bidirectional cascaded feature fusion structure includes an attention mechanism based on channel feature optimization; the low-light small target detection model includes an NIS loss function; the acquisition process of the NIS loss function is: obtaining an Inner-ShapeIoU loss function based on the Shape-IOU loss function, and combining the Inner-ShapeIoU loss function and the NWD loss function to obtain the NIS loss function; S3. Iteratively train the low-light small target detection model to obtain a trained low-light small target detection model; S4. Use the trained low-light small target detection model to detect the target to be detected and obtain a detection result.
2. The method for detecting low-light small targets based on an improved real-time detection Transformer according to claim 1, wherein: The feature extraction module based on residual attention is used to extract and denoise the features of the first input image to obtain a neck input feature map.
3. The low-light small target detection method based on an improved real-time detection Transformer according to claim 2, wherein, The specific acquisition process of the neck input feature map includes: Generating multiple intermediate feature maps by performing multiple 3×3 convolutions on the first input image, and splicing the first input image and the multiple intermediate feature maps to obtain a first spliced feature map; Reducing the channels of the first spliced feature map according to 1×1 convolution to obtain a first feature map; processing the first spliced feature map according to a 3×3 convolutional layer, batch normalization, and a rectified linear unit function to obtain a second feature map; Performing multi-scale feature extraction on the second feature map according to three convolutional layers with different dilation rates to obtain a first multi-scale feature map; processing the first multi-scale feature map according to a 1×1 convolutional layer and batch normalization to obtain a third feature map; adding the first feature map and the third feature map to obtain a first fused feature map; adjusting the number of channels of the first fused feature map according to the ESE channel adjustment module to obtain a second fused feature map; adding the second fused feature map and the first input image, and passing through the RAB denoising module to obtain a neck input feature map.
4. The low-light small target detection method based on an improved real-time detection Transformer according to claim 1, characterized in that: The improved backbone network includes a first neck input feature map, a second neck input feature map, a third neck input feature map, a fourth neck input feature map, a fifth neck input feature map, and a sixth neck input feature map.
5. The low-light small target detection method based on an improved real-time detection Transformer according to claim 4, characterized in that: The fusion principle steps of the bidirectional cascaded feature fusion structure include: S10. Perform 1×1 convolution channel reduction on the first neck input feature map, and perform a downsampling operation using a 3×3 convolution with a stride of 2; S20. Perform 1×1 convolution channel reduction on the small-scale feature map in the second neck input feature map; S30. Perform 1×1 convolution channel reduction on the large-scale feature map in the second neck input feature map, and perform downsampling operation using 3×3 convolution with a stride of 2; S40. Perform 1×1 convolution channel reduction on the fifth neck input feature map; S50. Perform 1×1 convolution channel reduction on the sixth neck input feature map, and respectively pass through the internal scale interaction module based on the attention mechanism, 1×1 convolution channel reduction, the attention mechanism based on channel feature optimization, and 2×2 transposed convolution operation; S60. Perform feature fusion on the feature maps obtained in S30, S40, and S50, and respectively pass through 1×1 convolution channel reduction, feature refinement, 1×1 convolution channel reduction, the attention mechanism based on channel feature optimization, and 2×2 transposed convolution operation; S70. Perform feature fusion on the three feature maps obtained in S10, S20, and S60, perform 1×1 convolution channel reduction and feature refinement, and input to the decoder; S80. Perform downsampling operation on the feature map obtained in S70 according to 3×3 convolution, perform feature fusion with the feature map after the operation of the attention mechanism based on channel feature optimization in S60, perform feature refinement, and input to the decoder; S90. Perform 3×3 convolution downsampling operation on the feature map obtained in S80, perform feature fusion with the feature map after the operation of the attention mechanism based on channel feature optimization in S50, and input to the decoder after refining the features.
6. The low-light small target detection method based on an improved real-time detection Transformer according to claim 1, characterized in that: The bidirectional cascaded feature fusion structure includes the attention mechanism based on channel feature optimization; the attention mechanism based on channel feature optimization includes an input module, BH-FPN, a spatial attention module, and an output module.
7. A low-light small target detection method based on an improved real-time detection Transformer according to claim 6, characterized in that: The specific process of the BH-FPN includes: S100. Input feature map F∈R H×W×C Divide into N×N non-overlapping regions, and then perform linear projection on the feature maps of each region to obtain the query Q, key K and value V tensors; S200. Obtain the adjacency matrix U by performing normalization processing on the query Q and the key K and performing matrix multiplication; S300. Perform matrix multiplication on U and V and process in combination with the attention mechanism to obtain the first output matrix F1; S400. Perform scale segmentation on the first output matrix F1 to obtain four first feature matrices with different scales; the first feature matrix includes multiple channels; input the first feature matrix into the feature selection module, and the feature selection module performs feature fusion on the four first feature matrices with different scales according to global average pooling and global maximum pooling respectively, determines the weight values of the multiple channels using the Sigmoid activation function, and then processes the weight values according to the channel attention mechanism; obtain four second feature matrices; the four second feature matrices are respectively multiplied by the four first feature matrices with different scales, and 1×1 convolution is performed for channel reduction processing to obtain the first high-level feature matrix and the first low-level feature matrix; S500. Expand the first high-level feature matrix by using transposed convolution to obtain a second high-level feature matrix; process the second high-level feature matrix through a feature selection module to obtain a third high-level feature matrix, and filter the first low-level feature matrix by using the third high-level feature matrix to obtain a second low-level feature matrix, and perform matrix addition according to the second low-level feature matrix and the second high-level feature matrix to obtain a first output matrix, and further obtain a first output feature map F2; S600. Perform an element-wise multiplication operation on F2 and F to generate F3 and input it into the spatial attention module. Process F3 through the spatial attention module to generate a spatial attention map A, and perform an element-wise multiplication operation on A and F3 to obtain the final output feature map.
8. The low-light small target detection method based on an improved real-time detection Transformer according to claim 1, wherein: The Shape-IoU loss function is: L ShapeIoU = 1 - IoU + distance shape + 0.5 * Ω shape ; Among them, L ShapeIoU represents the Shape-IoU loss function; IoU represents the intersection over union; distance shape represents the distance metric in the shape space; Ω shape represents the similarity metric between shapes; B and B gt are the predicted bounding box and the ground truth bounding box respectively; ∩ represents the intersection operation; ∪ represents the union operation; ww and hh are the weight coefficients in the horizontal and vertical directions respectively; x c gt and y c gt are the abscissa and ordinate of the center point of the predicted bounding box respectively; x c and y c are the abscissa and ordinate of the center point of the ground truth bounding box respectively, c is the normalization constant; scale is the scale factor; w and h represent the width and height of the predicted bounding box respectively; w gt and h gt represent the width and height of the ground truth bounding box respectively; e is the natural constant; ω w and ω h are parameters used to measure the degree of difference in the width and height directions of the shape respectively; max() represents the maximum value operation.
9. The low-light small target detection method based on an improved real-time detection Transformer according to claim 8, characterized in that: The NWD loss function is: L NWD =1-NWD(N p ,OF g ); Among them, L NWD Represents the NWD loss function; NWD(N p ,N g ) represents the normalized Wasserstein distance; N p and N g are two-dimensional Gaussian distribution models for the predicted box and the true box, W2 2 is the second-order distribution distance, C is a constant; exp is an exponential function with a constant e as the base.
10. The low-light small target detection method based on an improved real-time detection Transformer according to claim 9, wherein: Obtain the NIS loss function: L NIS = λ * L NWD + (1 - λ) * L Inner-ShapeIoU ; L Inner-ShapeIoU = L ShapeIoU + IoU - IoU inner ; Among them, L NIS represents the NIS loss function; λ represents the weight coefficient; L NWD represents the NWD loss function; L Inner-ShapeIoU represents the Inner-ShapeIoU loss function; L ShapeIoU represents the Shape-IoU loss function; IoU represents the intersection over union; IoU inner represents the internal intersection over union.