ERT-DETR model for detecting small target object on water surface
Through the dynamic feature pyramid network and microscopic attention module, the ERT-DETR model solves the real-time and accuracy of small-surface object detection in complex environments, and realizes efficient small-surface object detection, which is suitable for marine management and environmental monitoring.
Patent Information
- Application Number
- CN202510306301.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-03-14
AI Technical Summary
Small surface object detection is difficult to take into account real-time and end-to-end performance in complex environments, and lacks effective feature expression capabilities and data sets, resulting in insufficient generalization capabilities of the model in practical applications.
The ERT-DETR model is adopted, combined with the dynamic feature pyramid network and the micro attention module, and the high-level semantic features are extracted through the AIFI module. The RepBi-PAN network integrates low-level features, the SPD-Conv module performs multi-scale fusion, the deep separable convolution module enhances feature capture, the channel attention module enhances feature representation, and uses the Inner-IoU loss function to accelerate the model convergence.
It improves the accuracy and efficiency of small target detection on water surface in complex environments, and achieves 92.6% AP and 118.84FPS detection performance, which is better than the existing technology and is suitable for marine management, environmental monitoring and ship tracking.
Smart Images

Figure CN120388159A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of small target detection, and in particular to an ERT-DETR model for detecting small targets on the water surface. Background Art
[0002] 85% of marine litter is plastic, and it is increasing at a rate of 23 million to 37 million tons per year. If such floating objects as plastic waste are not monitored and treated, it will not only affect the safety and efficiency of waterborne shipping traffic, but also have a severe impact on human health, the global economy, biodiversity, and climate issues once they are corroded and degraded into the aquatic ecosystem. Detection of floating targets is a key step in the monitoring and treatment of floating objects, and it is of great significance in aspects such as marine management, environmental protection, resource development, and rescue operations.
[0003] However, in the field of target detection, detecting small floating targets has always been a challenging research topic. First of all, small targets generally occupy a limited number of pixels in the image, which leads to low feature resolution and information loss, and lack of feature expression ability. At the same time, compared with ordinary small targets, the detection difficulties of small floating targets are reflected in the following two aspects: one is that the shapes, colors, and textures of small floating targets in the water surface environment change greatly, and they often present a distorted and blurred state. It is difficult to learn appropriate features from these distorted features itself, especially when the targets in a complex water surface environment are also affected by factors such as light changes, water ripples, and water surface reflections. The other is that large-scale datasets for detecting small floating objects on the water surface are scarce, and they are often widely distributed and difficult to obtain. The lack of datasets results in insufficient generalization ability of the trained models in practical applications and difficulty in dealing with complex situations in the water surface environment. Therefore, the algorithms and data for detecting small floating targets in complex environments are relatively insufficient, making it far behind the target detection technologies in other application scenarios.
[0004] Currently, there are many improved methods for small target detection based on deep learning, mainly including optimizing the AnchorBox, optimizing the backbone network, introducing attention mechanisms and other networks into the model, performing feature fusion on feature maps of different scales, optimizing the IoU and loss functions, etc. Water surface target detection is used to locate floating objects on the water surface and determine the categories of various targets on the water surface (such as ships, buoys, etc.), which is of great significance for USV environmental perception and safe navigation.
[0005] Traditional water surface target detection methods mainly include background subtraction, frame difference, Hough transform, etc. These methods require manually crafted models to extract specific features. For example, Szpak et al. calibrated moving ships in the ocean based on real-time approximation of background subtraction and level set curve evolution to track moving ships in dynamic backgrounds. Wu et al. proposed an adaptive Gaussian mixture model to simulate complex backgrounds and improve the accuracy of buoy detection. These methods can achieve high detection accuracy in static background scenarios. However, they are vulnerable to backgrounds such as sunlight changes, rain, and surface fluctuations, resulting in poor performance. Xu et al. constructed a three-dimensional feature detector based on the multi-polarization features of radar signals to improve the accuracy of small water surface target detection. Although traditional algorithms have made great progress, they require manually designed features and do not allow end-to-end detection. In addition, due to the lack of robustness of manually designed features to different inputs, they cannot be applied to complex water surface target detection tasks. For small water floating target detection, more detailed feature extraction and higher resolution are usually required, which increases the computational complexity of the model and thus affects real-time performance. Summary of the Invention
[0006] In view of the technical problems existing in the prior art, the present invention provides an ERT-DETR model for detecting small water surface targets, which is of great significance for improving the ability to detect small targets in complex environments and can be extended to applications in ocean management, environmental monitoring, ship tracking and other fields.
[0007] According to a first aspect of the present invention, there is provided an ERT-DETR model for detecting small water surface targets, comprising: a backbone network, a dynamic feature pyramid network, and a micro-attention module;
[0008] The input of the dynamic feature pyramid network is the features of the last four stages of the backbone network;
[0009] The dynamic feature pyramid network includes: an AIFI module, a RepBi-PAN network, and a spatial depth conversion convolution module; the AIFI module uses a self-attention mechanism to process high-level semantic features and extracts key features from the deepest layer of the backbone network; the RepBi-PAN network fuses the key features with low-level features extracted from the backbone network to obtain fused features; the SPD-Conv module performs multi-scale fusion convolution on the fused features to obtain convolution features;
[0010] The micro attention module includes: a depthwise separable convolution module, a stationary point convolution module, and a channel attention module; the depthwise separable convolution module performs depthwise convolution on the features after splitting them along the channel direction to obtain depthwise convolution features, the pointwise convolution module integrates the channels of each depthwise convolution feature and then refines them to obtain key features, and the channel attention module is used to enhance the channel features of the key features to obtain a weighted feature map.
[0011] On the basis of the above technical solutions, the present invention can also be improved as follows.
[0012] Optionally, the SPD-Conv module includes: a space-to-depth layer and a non-stride convolution layer;
[0013] The space-to-depth layer is used to transform the feature map inside the entire CNN into an intermediate feature map. The transformation process includes: downsampling the feature map to obtain each sub-feature map, and connecting each sub-feature map along the channel dimension to obtain the intermediate feature map.
[0014] Optionally, the process by which the depthwise separable convolution module obtains the depthwise convolution features includes:
[0015] The input feature X is split into four equal parts along the channel direction to obtain features X corresponding to different directions i = Split(X), where i represents the i-th equal part into which X is divided along the channel, and Split represents the splitting operation;
[0016] Depthwise convolution kernels of different sizes are used for different feature map branches. For the split features X i = Split(X), depthwise convolution operations are respectively performed: X i ' = DWConv2d(X i ), where DWconv2d represents the depthwise convolution operation, X1 uses a 3×3 depthwise convolution; X2 and X3 respectively use 1×11 and 11×1 depthwise convolutions; and X4 remains unchanged as an identity branch.
[0017] Optionally, the process by which the pointwise convolution module obtains the key features includes:
[0018] The features obtained after depthwise convolution are concatenated, and then processed through an activation function, a normalization layer, and a residual connection;
[0019] Pointwise convolution is used to integrate channel information to refine and obtain the key features.
[0020] Optionally, the process by which the channel attention module obtains the weighted feature map from the key features includes:
[0021] The global features of the key features are obtained using an adaptive average pooling layer, and the global features are subjected to channel compression and recalibration through a two-layer fully connected network. The weights of each channel are obtained using the Sigmoid function to obtain the output features;
[0022] The output features after being processed by the exponential function are multiplied element-wise with the original input features of the micro attention module to obtain the weighted feature map.
[0023] Optionally, the loss function of the ERT-DETR model is to apply the Inner-IoU loss to the IoU-based bounding box regression loss function.
[0024] Optionally, the Inner-IoU loss is:
[0025]
[0026] where, b l 、b r 、b t 、b t 、inter and union are intermediate variables, are the center point coordinates inside the GT box and the GT box, (x c ,y c ) are the center points of the anchor box and the internal anchor box, w gt and h gt are the width and height of the GT box respectively, w and h are the width and height of the anchor box respectively; the variable ratio is the scale factor; IoU inner is the Inner-IoU loss.
[0027] Optionally, the IoU-based bounding box regression loss function includes: Iou, SIou and CIou;
[0028] L Inner-IoU = 1 - IoU inner ;
[0029] L Inner-SIoU = L SIoU + IoU - IoU inner ;
[0030] L Inner-CIoU = L CIoU + IoU - IoU inner ;
[0031] where, IoU inner is the Inner-IoU loss, IoU, L SIoU and L CIoULoss functions representing Iou, SIou, and CIou, L Inner-IoU , L Inner-SIoU and L Inner-CIoU respectively represent the loss functions that apply IoU inner to Iou, CIou, and SIou, where IoU inner is the Inner-IoU loss.
[0032] An ERT-DETR model for detecting small targets on the water surface provided by the present invention has the following beneficial effects:
[0033] 1. Aiming at the problem of difficult detection of small floating targets on the water surface, the embodiment of the present invention proposes an ERT-DETR model. On the premise of having real-time performance and end-to-end performance, this model can also improve the detection accuracy and efficiency of small floating targets in complex environments;
[0034] 2. In order to improve the feature representation ability of the network model, the embodiment of the present invention proposes a new neck network DFPNet, which uses the AIFI module to extract the deepest features of the backbone network and introduces the SPDConv module to enhance the local detail representation of the feature map.
[0035] 3. In order to make the model focus on learning the significant features of small floating targets, the embodiment of the present invention introduces a new attention module and uses the Inner-IoU assisted bounding box loss function to accelerate the convergence speed of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a structural diagram of an embodiment of an ERT-DETR model for detecting small targets on the water surface provided by the present invention;
[0037] Figure 2 is a structural diagram of an embodiment of a dynamic feature pyramid network provided by the present invention;
[0038] Figure 3 is a structural diagram of an embodiment of an SPD-Conv module provided by the present invention;
[0039] Figure 4 is a structural diagram of an embodiment of a microscopic attention module provided by the present invention;
[0040] Figure 5 is a schematic diagram of the calculation process of Inner-IoU provided by the present invention;
[0041] Figure 6(a) is a schematic diagram of the number distribution of different category targets in the small floating target detection dataset provided by the present invention;
[0042] Figure 6(b) Schematic diagram of the distribution of the center point positions of the targets in the small water-floating target detection dataset provided by the present invention;
[0043] Figure 6(c) Schematic diagram of the distribution of the object sizes in the small water-floating target detection dataset provided by the present invention;
[0044] Figure 7 Schematic diagram of the comparison between Inner-CIoU and Inner-SIoU under different ratios provided by the embodiment of the present invention;
[0045] Figure 8 Schematic diagram of the loss curve of the ablation experiment provided by the embodiment of the present invention;
[0046] Figure 9(a) Schematic diagram of an embodiment of the input picture;
[0047] Figure 9(b) Attention effect diagram of the input picture detected by the RT-DETR model;
[0048] Figure 9(c) Attention effect diagram of the input picture detected by the ERT-DETR model provided by the embodiment of the present invention;
[0049] Figure 10(a) Schematic diagram of another embodiment of the input picture;
[0050] Figure 10(b) Detection effect diagram of the input picture detected by the RT-DETR model;
[0051] Figure 10(c) Detection effect diagram of the input picture detected by the ERT-DETR model provided by the embodiment of the present invention. Detailed implementation manners
[0052] The principles and features of the present invention will be described below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.
[0053] Figure 1 Structure of an embodiment of an ERT-DETR model for detecting small water surface targets provided by the present invention, as Figure 1 shown. The ERT-DETR (Enhancing Real-Time Detection Transformer) model includes: a backbone network, a Dynamic Feature Pyramid Network (DFPNet), and a Micro-Attention Module (MAM).
[0054] The input of the Dynamic Feature Pyramid Network is the features of the last four stages of the backbone network; the backbone network is a backbone network based on the DETR model.
[0055] The Dynamic Feature Pyramid Network includes: the AIFI module, the RepBi-PAN network, and the Spatially Pooled Dilated Convolution (SPD-Conv) module; the AIFI module uses the self-attention mechanism to process high-level semantic features and extracts key features from the deepest layer of the backbone network; the RepBi-PAN network fuses the key features with the low-level features extracted from the backbone network to obtain the fused features; the SPD-Conv module performs multi-scale fusion convolution on the fused features to obtain convolution features.
[0056] The Microscopic Attention Module includes: the Depthwise Separable Convolution Module, the Stationary Point Convolution Module, and the Channel Attention Module; the Depthwise Separable Convolution Module performs depth convolution on the features after splitting them along the channel direction to obtain the depth-convolved features, the Pointwise Convolution Module integrates the channels of each depth-convolved feature and then refines them to obtain the key features, and the Channel Attention (SE, Squeeze and Excitation) Module is used to enhance the channel features of the key features to obtain the weighted feature map.
[0057] The ERT-DETR model for small water surface target detection provided by the present invention first uses the features of the last four stages of the backbone network {C2, C3, C4, C5} as the input of the neck network. The efficient hybrid encoder converts multi-scale features into an image feature sequence through in-scale feature interaction (AIFI). The DFPNet combines the idea of the bidirectional feature pyramid network, strengthens the connection between features through the enhanced path aggregation network, effectively shortens the information transmission path between high-level and low-level features, and then reduces information loss and improves the accuracy of feature extraction through SPD-Conv. The Microscopic Attention Module is used to enhance the capture of small-size features. Finally, Inner-IoU uses the auxiliary bounding box to calculate the IoU loss to accelerate the sample convergence process and improve the accuracy of bounding box localization, and the decoder generates boxes and confidence scores by iteratively optimizing the object queries. Comparative and ablation experiments were carried out with the water surface floating object dataset FloW-Img as the experimental object. The results show that the ERT-DETR model achieves 92.6% AP and 118.84 FPS, which are 4.3% and 19.8% higher than the benchmark RT-DETR respectively. The ERT-DETR model is superior to the state-of-the-art YOLO and DETR series detectors in terms of the detection accuracy and real-time performance of small targets in a dynamic water surface environment. This model is of great significance for improving the small target detection ability in complex environments and can be extended to applications in ocean management, environmental monitoring, ship tracking and other fields.
[0058] Example 1
[0059] Example 1 provided by the present invention is an example of an ERT-DETR model for detecting small targets on the water surface. Combining Figure 1 it can be seen that this ERT-DETR model example includes: a backbone network, a dynamic feature pyramid network, and a micro attention module.
[0060] The input of the dynamic feature pyramid network is the features of the last four stages of the backbone network.
[0061] The dynamic feature pyramid network aims to improve the detection effect of small targets. As Figure 2 shown in the structural schematic diagram of an example of a dynamic feature pyramid network provided by the present invention. Combining Figure 1 and Figure 2 it can be seen that it includes: an AIFI module, a RepBi-PAN network, and a spatially pooled dilated convolution (SPD-Conv) module.
[0062] The AIFI module uses the self-attention mechanism to process high-level semantic features, extracts key features from the deepest layer of the backbone network, and helps to capture the connections between conceptual entities in the image.
[0063] The RepBi-PAN network fuses the key features with the low-level features extracted from the backbone network to obtain fused features, strengthens the connection between features, and effectively shortens the information transmission path.
[0064] The SPD-Conv (Spatially Pooled Dilated Convolution) module performs multi-scale fusion convolution on the fused features to obtain convolution features, improving the information capture of small targets and the accuracy of target localization.
[0065] As Figure 3 shown in the structural diagram of an example of an SPD-Conv module provided by the present invention. In a possible embodiment, combining Figure 3 it can be seen that the SPD-Conv module includes: a space-to-depth (SPD) layer and a non-strided convolution (Conv) layer.
[0066] The space-to-depth layer is used to transform the feature map inside the entire CNN into an intermediate feature map. The transformation process includes: downsampling the feature map to obtain each sub-feature map, and connecting each sub-feature map along the channel dimension to obtain the intermediate feature map.
[0067] The space-to-depth layer generalizes the original image conversion technique to downsample the feature maps inside and throughout the CNN. Given any (original) feature S(H, W, C), where H and W are the height and width of the feature map respectively, and C is the number of channels. After downsampling, these sub-feature maps are concatenated along the channel dimension to obtain the feature map S'. SPD transforms the feature map S(H, W, C) into an intermediate feature map
[0068] To improve the detection effect of small targets in complex backgrounds and enhance the capture and representation of these details, a microscopic attention module is proposed. The structure diagram of MAM is as Figure 4 shown. MAM can utilize its depthwise separable convolution to enhance the capture of small-size features, then further integrate channel information through pointwise convolution to refine key features, and finally strengthen the channel features of the input feature map through the SE module.
[0069] Combined with Figure 4 it can be seen that the microscopic attention module includes: a depthwise separable convolution module, a stationary point convolution module, and a channel attention module; the depthwise separable convolution module performs depth convolution on the features after splitting them along the channel direction to obtain the depth-convolved features, the pointwise convolution module integrates the channels of each depth-convolved feature and then refines to obtain the key features, and the channel attention (SE, Squeeze and Excitation) module is used to strengthen the channel features of the key features to obtain the weighted feature map.
[0070] In a possible embodiment, the process of the depthwise separable convolution module obtaining the depth-convolved features includes:
[0071] Assume the input feature of MAM is X ∈ R B×C×H×W , where B is the batch size. The input feature X is split into four equal parts along the channel direction to obtain the features X i = Split(X), where i represents the i-th equal part of X split along the channel, and Split represents the splitting operation.
[0072] The split features are passed through the Inception branch structure, which combines multi-scale depthwise separable convolution (DWConv) modules. This module uses depth convolution (DWConv) operations on the split feature maps respectively, and uses depth convolution kernels of different sizes for different feature map branches. The depth convolution operations are performed on the split features X i = Split(X) respectively: X i ' = DWConv2d(X i), where DWconv2d represents the depth convolution operation, X1 uses a 3×3 depth convolution; X2 and X3 use 1×11 and 11×1 depth convolutions respectively; while X4 remains unchanged as an identity branch to preserve the integrity of the original features.
[0073] In a possible embodiment, the process of the pointwise convolution module obtaining the key features includes:
[0074] Concatenate the features obtained after depth convolution, and then process through an activation function, a normalization layer, and a residual connection: X′ = σ(Concat(X i )) + X, where σ(·) is a non-linear activation function.
[0075] Use pointwise convolution (PointwiseConvolution) to integrate channel information to refine the key features: X″ = σ(Conv2d 1×1 (X′)) + X.
[0076] In a possible embodiment, the process of the channel attention module obtaining the weighted feature map for the key features includes:
[0077] Use an adaptive average pooling layer to obtain the global features of the key features, compress and recalibrate the channels of the global features through a two-layer fully connected network, use the Sigmoid function to obtain the weight of each channel, and obtain the output features.
[0078] Multiply the output features processed by the exponential function element-wise with the original input features of the micro attention module to obtain the weighted feature map.
[0079] Specifically, subsequently through the channel attention module, first use the adaptive average pooling layer (AvgPool) to reduce the spatial dimension of X” to 1×1; then through a two-layer fully connected network (FC), the first layer reduces the feature C′ to C′ / / reduction, and the second layer raises the feature from C′ / / reduction back to C′:
[0080] SE = f(σ(f(AvgPool(X″)))).
[0081] Among them, f(·) is a fully connected network. The SE module uses an adaptive average pooling layer to obtain global features, then performs channel compression and recalibration through a two-layer fully connected network, and then uses the Sigmoid function to obtain the weight of each channel. Then the output of the SE module is:
[0082] SE out = Sigmoid(SE).
[0083] The output of the SE module is processed by an exponential function and then multiplied element-wise with the original input feature X of the module to obtain the final weighted feature map:
[0084] Y = X ⊙ exp(SE out ).
[0085] Where ⊙ represents the Hadamard product (element-wise multiplication), and exp(·) is the exponential function used to convert the output of Sigmoid(·) into a multiplicative weight.
[0086] The IoU loss function has a wide range of applications in computer vision tasks. During the bounding box regression process, it can not only evaluate the quality of the regression state but also perform gradient propagation by calculating the regression loss to accelerate convergence. In one possible embodiment, the loss function of the ERT-DETR model is to apply the Inner-IoU loss to the IoU-based bounding box regression loss function. Selecting Inner-IoU as the loss function uses auxiliary bounding boxes to accelerate sample convergence for small object detection. As Figure 5 shown is a schematic diagram of the calculation process of Inner-IoU provided by the present invention. As Figure 5 shown, the GT box and the anchor box are respectively denoted as B gt and B. The Inner-IoU loss is:
[0087]
[0088] Where b l , b r , b t , b t , inter, and union are intermediate variables, are the center point coordinates inside the GT box and the GT box, (x c , y c ) are the center points of the anchor box and the inner anchor box, w gt and h gt are respectively the width and height of the GT box, and w and h are respectively the width and height of the anchor box; the variable ratio is a scale factor, usually taking a range of [0.5, 1.5]; IoU inner is the Inner-IoU loss. The width and height of the InnerGT box are respectively denoted as and The width and height of the Inner anchor box are respectively denoted as w inner and h inner .
[0089] The Inner-IoU loss inherits some characteristics of the IoU loss while having its own characteristics. Similar to the IoU loss, the value range of the Inner-IoU loss is [0, 1]. Since there are only scale differences between the auxiliary bounding box and the actual bounding box, and the change trend of the IoU value during the regression process is consistent with that of the IoU value of the actual bounding box, it can reflect the quality of the regression result of the actual bounding box.
[0090] In a possible embodiment, the IoU-based bounding box regression loss function includes: Iou (Intersection of Union), SIou (SCYLLA-Intersection of Union), and CIou (Complete-Intersection of Union).
[0091] L Inner-IoU = 1 - IoU inner 。
[0092] L Inner-SIoU = L SIoU + IoU - IoU inner 。
[0093] L Inner-CIoU = L CIoU + IoU - IoU inner 。
[0094] Among them, IoU inner is the Inner-IoU loss, IoU, L SIoU and L CIoU represent the loss functions of Iou, SIou, and CIou, and L Inner-IoU , L Inner-SIoU and L Inner-CIoU respectively represent the loss functions obtained by applying IoU inner to Iou, CIou, and SIou. IoU inner is the Inner-IoU loss.
[0095] Embodiment 2
[0096] Example 2 provided by the present invention is an application example of an ERT-DETR model for detecting small targets on the water surface. In the application example provided by the present invention, in order to verify the actual performance of ERT-DETR, experiments are carried out on the FloW-Img dataset. FloW-Img is mainly intercepted from videos recorded by unmanned boats, and finally 2000 pictures containing small floating targets are selected as the dataset. The training set, validation set and test set are divided according to the ratio of 6:2:2. Figure 6 shows the visualization results of the analysis of the small floating target detection dataset. Figure 6(a) shows the number distribution of different types of targets in the dataset. The horizontal and vertical coordinates of Figure 6(b) and Figure 6(c) are both normalized results. Among them, Figure 6(b) shows the distribution of the target center point positions. The darker the color, the more concentrated the center points of the target boxes, which is beneficial to improving the detection effect of the model on occluded targets. Figure 6(c) shows the distribution of object sizes. It can be seen from this figure that the proportion of small objects is very high.
[0097] The detector is trained using the AdamW optimizer, with learning_rate = 0.0001, weight_decay = 0.0001, warmup_epochs = 3.0, warmup_momentum = 0.8, warmup_bias_lr = 0.1. Data augmentation includes random {HSV augmentation, translation, scaling, flipping} operations, and the specific parameter settings are as follows: hsv_h = 0.015, hsv_s = 0.7, hsv_v = 0.4, translate = 0.1, scale = 0.5, fliplr = 0.5. The scaling factors used by the end-to-end class algorithm and the real-time end-to-end class algorithm are both: [depth, width, max_channels] = [1.00, 1.00, 1024]. When the performance of the end-to-end class algorithm model does not improve significantly, the early stopping number of epochs Epoch = 200, patience = 50, while the real-time end-to-end class algorithm uses Epoch = 300, patience = 100. Inner-CIoU is used for Inner-IoU in the comparative experiment and ablation experiment, and the ratio ratio is set to 1.15. The training strategy and hyperparameters of the decoder almost follow RT-DETR.
[0098] In the embodiments of the present invention, RT-DETR is used as the detector, the Inner-CIoU method and the Inner-SIoU method are compared, and it degrades to the IoU method when ratio = 1. To prove the superiority of the ERT-DETR model for detecting small targets on the water surface provided by the present invention, it is trained for 200 Epochs on the training set, and a comparative experiment is carried out on the test set. The experimental results are shown in Fig. 6. It can be seen that after using the Inner-IoU method, both the detection effect and the real-time performance are improved. The increase in AP and mAP exceeds 0.9%, and the increase in FPS exceeds 5. The best performance is achieved by the Inner-CIoU method with ratio = 0.80, and the FPS / AP / mAP are increased by 8.85 / 1.3% / 1.8% respectively.
[0099] In the FloW-Img dataset, the proportion of small targets (area < 32×32) reaches 56.81%, and there may be a large number of low-IoU samples. The experimental results show that the overall performance of the Inner-CIoU method is better than that of eInner-SIoU. Therefore, in the embodiments of the present invention, the Inner-CIoU is used to set ratio > 1 for experiments, and the experimental results are shown in Table 1. When ratio = 1.15, the Inner-CIoU method has the best performance, and is better than the performance of Inner-CIoU and Inner-SIoU when ratio < 1. The FPS / AP / mAP reach 109.11 / 90.9% / 49.9% respectively, which are improved by 11.58 / 2.4% / 2.3% compared with the CIoU method.
[0100] Table 1 Performance of various CIoU losses (ratio > 1, between 1.1 and 1.2)
[0101]
[0102] Figure 7 It is a comparison schematic diagram of Inner-CIoU and Inner-SIoU under different ratios provided by the embodiments of the present invention (ratio < 1, between 0.7 and 0.8). From Figure 7 it can be seen that by setting ratio < 1, an auxiliary bounding box smaller than the actual bounding box is generated, and a certain performance improvement can still be obtained in high-IoU samples, indicating that this method has strong generalization ability. As can be seen from Table 1, when ratio > 1, by generating a larger auxiliary bounding box to accelerate the convergence of low-IoU samples, it can detect and track floating objects on the water surface more real-time and accurately.
[0103] The ERT-DETR model was compared with the state-of-the-art real-time end-to-end object detector RT-DETR and real-time and end-to-end object detectors.
[0104] Under the condition of following the end-to-end setting of RT-DETR, the performance of ERTSWSO-DETR was compared with that of real-time detectors. In the embodiments of the present invention, ERT-DETR was compared with YOLOv5, YOLOv6 v3.0 (hereinafter referred to as YOLOv6) and YOLOv8 in Table 2. Compared with YOLOv5-L / YOLOv6-L / YOLOv8-L, ERTSWSO-DETR increased the AP by 4.4% / 7.0% / 2.8%, increased the FPS by 103.2% / 114.3% / 55.7%, and reduced the number of parameters by 25.5% / 64.2% / 9.5%. Compared with YOLOv6-L and YOLOv8-L, the computational cost of ERTSWSO-DETR decreased by 65.4% / 17.8%. Compared with YOLOv5-X / YOLOv6-X / YOLOv8-X, ERTSWSO-DETR increased the AP by 3.5% / 15.2% / 2.4%, increased the FPS by 238.7% / 174.5% / 104.4%, reduced the number of parameters by 59.1% / 77.0% / 41.5%, and reduced the computational cost by 44.9% / 350.3% / 90.0%.
[0105] Compared to end-to-end detectors. For fair comparison, the embodiments of the present invention only compare with end-to-end detectors using the same backbone. As can be seen from Table 2, ERTSWSO-DETR is superior to the current state-of-the-art end-to-end detectors. Compared with DINO-Deformable-DETR-R50, ERTSWSO-DETR-R50 significantly increased the AP by 4.8% (92.2% vs 87.4%) and the speed by 11 times (118.84 FPS vs 10.75 FPS).
[0106] As can be seen from Table 2, under the same backbone, both the real-time performance and accuracy of ERTSWSO-DETR are superior to those of the first known real-time end-to-end detector RT-DETR. Compared with RT-DETR, ERTSWSO-DETR significantly increased the AP by 4.3% (92.6% vs 88.3%) and the speed by 19.8% (118.84 FPS vs 99.18 FPS).
[0107] Table 2: Comparison of results with SOTA detector (.(The input size of real-time and real-time end-to-end detectors is 640, and the input size of end-to-end detectors is (800, 1333))
[0108]
[0109]
[0110] As can be seen from the experimental results, an ERT-DETR model provided by the present invention achieves 92.6% AP and 118.84 FPS, and is superior to RT-DETR and YOLO series detectors in terms of both speed and accuracy. In addition, the ERT-DETR-R50 proposed in the embodiments of the present invention reaches 92.2% AP and 113.26 FPS, and is superior to the state-of-the-art end-to-end detectors with the same backbone in terms of both speed and accuracy.
[0111] To verify the effectiveness of the proposed innovation, the embodiments of the present invention designed seven groups of ablation experiments for comparison with the baseline, aiming to evaluate the improvement of the detector performance. The embodiments of the present invention visualized the loss curves of the experimental process, as Figure 8 shown. When only the Inner-CIoU improvement method is retained, the number of Epochs required for the loss to drop to 2 / 5 of the initial value is only 1 / 7 of the baseline. The model has basically converged after 163 Epochs, while the baseline requires 200 Epochs. After the improvement strategy proposed in the embodiments of the present invention, the convergence speed of the detector in the experimental group is significantly faster than that of the baseline, and the final loss value is approximately reduced to 1 / 4 of the initial value, while the baseline is only 1 / 3. It should be noted that the speed of the loss value reduction directly reflects the training efficiency and performance of the model.
[0112] The quantitative evaluation results of the ablation experiments are shown in Table 3. Compared with the baseline, the improvement methods proposed in the embodiments of the present invention have significantly improved performance. After introducing the MAM and Inner-CIoU methods of the embodiments of the present invention into RT-DETR, the number of parameters and the amount of calculation are almost unchanged, but the FPS is increased by 16.2% / 10.0% respectively, and the AP is increased by 3.6% / 2.6%. After introducing the DFPNet method of the embodiments of the present invention, although the number of parameters and the amount of calculation increase, the 13.5% / 3.6% increase in FPS / AP makes the increase in the number of parameters and the amount of calculation worthwhile. Through continuous improvement, the detector of the embodiments of the present invention finally obtains 118.84 FPS and 92.6 AP, which are increased by 19.8% / 4.3% respectively compared with the baseline. The ablation experiment results show that in the complex river channel environment, the improvement strategy of the embodiments of the present invention can detect small floating targets on the water surface more real-time and accurately.
[0113] Table 3 Main results of the ablation experiments
[0114]
[0115] Finally, to clearly demonstrate the improved performance of the model, the embodiment of the present invention is compared with the baseline RT-DETR method in terms of attention effect and actual detection effect. Figures 9(a), (b), and (c) are respectively the input image, the attention effect of the input image detected by the RT-DETR model, and the attention effect of the input image detected by the ERT-DETR model provided by the embodiment of the present invention. The spherical yellow blocks in the figures are the areas of attention. As can be seen from Figures 9(a), (b), and (c), the ERT-DETR detector provided by the embodiment of the present invention can still better focus on the target area in dim conditions, when multiple targets are mutually occluded, and when the target is far away and features are sparse. As can be seen from Figures 10(a), (b), and (c), compared with the baseline algorithm, ERT-DETR's detection effect in complex environments has been significantly improved, with fewer false detections and missed detections. Therefore, the performance of the ERT-DETR detector proposed in the embodiment of the present invention has surpassed the current SOTA detector.
[0116] An embodiment of the present invention provides an ERT-DETR model for detecting small objects on the water surface. The ERT-DETR model strikes a balance between performance and speed, adopting a lightweight model structure or optimizing the model inference process to improve speed. Its beneficial effects include:
[0117] 1. To address the difficulty in detecting small floating targets on the water surface, an embodiment of the present invention proposes an ERT-DETR model. This model not only achieves real-time and end-to-end performance, but also improves the detection accuracy and efficiency of small floating targets in complex environments.
[0118] 2. In order to improve the feature representation capability of the network model, an embodiment of the present invention proposes a new neck network DFPNet, which uses the AIFI module to extract the deepest features of the backbone network and introduces the SPDConv module to enhance the local detail representation of the feature map.
[0119] 3. In order to make the model focus on learning the salient features of small floating targets, the embodiment of the present invention introduces a new attention module and uses the Inner-IoU auxiliary bounding box loss function to accelerate the convergence of the model.
[0120] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0121] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0122] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0123] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0124] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0125] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications that fall within the scope of the present invention.
[0126] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these modifications and variations.
Claims
1. An ERT-DETR model for detecting small targets on the water surface, characterized in that, The ERT-DETR model includes: a backbone network, a dynamic feature pyramid network, and a micro attention module; The input of the dynamic feature pyramid network is the features of the last four stages of the backbone network; The dynamic feature pyramid network includes: an AIFI module, a RepBi-PAN network, and a spatial depth conversion convolution module; the AIFI module uses a self-attention mechanism to process high-level semantic features and extracts key features from the deepest layer of the backbone network; the RepBi-PAN network fuses the key features with low-level features extracted from the backbone network to obtain fused features; the SPD-Conv module performs multi-scale fusion convolution on the fused features to obtain convolution features; The micro attention module includes: a depthwise separable convolution module, a stationary point convolution module, and a channel attention module; the depthwise separable convolution module performs depth convolution on the features after splitting them along the channel direction to obtain depth-convolved features, the pointwise convolution module integrates the channels of each of the depth-convolved features and then refines them to obtain key features, and the channel attention module is used to enhance the channel features of the key features to obtain a weighted feature map.
2. The ERT-DETR model according to claim 1, wherein The SPD-Conv module includes: a space-to-depth layer and a non-stride convolution layer; The space-to-depth layer is used to transform the feature map inside the entire CNN into an intermediate feature map. The transformation process includes: downsampling the feature map to obtain each sub-feature map, and connecting each of the sub-feature maps along the channel dimension to obtain the intermediate feature map.
3. The ERT-DETR model according to claim 1, wherein The process by which the depthwise separable convolution module obtains the depth-convolved features includes: The input feature X is split into four equal parts along the channel direction to obtain features X corresponding to different directions i = Split(X), where i represents the i-th equal part into which X is divided along the channel, and Split represents the splitting operation; Use depth convolution kernels of different sizes for different feature map branches, and perform depth convolution operations on the segmented feature X i = Split(X) respectively: X' i = DWConv2d(X i ), where DWconv2d represents the depth convolution operation, X1 uses a 3×3 depth convolution; X2 and X3 use 1×11 and 11×1 depth convolutions respectively; and X4 remains unchanged as the identity branch.
4. The ERT-DETR model according to claim 1, wherein The process by which the pointwise convolution module obtains the key features includes: Concatenating the features obtained after depth convolution, and then passing through an activation function, a normalization layer, and a residual connection; Using pointwise convolution to integrate channel information to refine and obtain the key features.
5. The ERT-DETR model according to claim 1, characterized in that, The process by which the channel attention module obtains the weighted feature map from the key features includes: Using an adaptive average pooling layer to obtain the global feature of the key features, performing channel compression and recalibration on the global feature through a two-layer fully connected network, using the Sigmoid function to obtain the weight of each channel, and obtaining the output feature; Performing element-wise multiplication on the output feature after being processed by the exponential function and the original input feature of the micro attention module to obtain the weighted feature map.
6. The ERT-DETR model according to claim 1, wherein The loss function of the ERT-DETR model is to apply the Inner-IoU loss to the IoU-based bounding box regression loss function.
7. The ERT-DETR model according to claim 6, wherein The Inner-IoU loss is: Among them, b l , b r , b t , b t , inter and union are intermediate variables, is the center point coordinates of the GT box and inside the GT box, (x c , y c ) is the center point of the anchor box and the internal anchor box, w gt and h gt are the width and height of the GT box respectively, w and h are the width and height of the anchor box respectively; the variable ratio is the scale factor; IoU inner is the Inner-IoU loss.
8. The ERT-DETR model according to claim 6, characterized in that, The IoU-based bounding box regression loss function includes: Iou, SIou, and CIou; L Inner-IoU = 1 - IoU inner ; L Inner-SIoU = L SIoU + IoU - IoU inner ; L Inner-CIoU = L CIoU + IoU - IoU inner ; Among them, IoU inner is the Inner-IoU loss, and IoU, L SIoU and L CIoU represent the loss functions of Iou, SIou, and CIou, respectively. L Inner-IoU and L Inner-SIoU and L Inner-CIoU represent the loss functions obtained by applying IoU inner to Iou, CIou, and SIou, respectively. IoU inner is the Inner-IoU loss.
Citation Information
Patent Citations
Marine ship detection method based on improved YOLOv3 algorithm
CN113743322A
Lightweight convolutional neural network training method and system for water surface target detection
CN117151186A
High body seriola quinqueradiata detection method based on YOLOv8 network structure
CN118279935A
One-stage detection method for water surface floating garbage based on deep learning
CN118351345A
Small ship detection method based on convolution block attention mechanism and explicit visual center
CN119540760A