Underwater benthos detection method based on improved DETR model

By improving the DETR model, combining the compact inverted bottleneck module, channel shuffling module and multi-scale expansion convolution, the HIFI and EDCM modules are designed, which solves the efficiency and accuracy of benthic biological detection in underwater environments and achieves efficient and accurate detection results.

CN120236189APending Publication Date: 2025-07-01CHANGZHOU UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510389490.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The prior art is difficult to detect benthic organisms efficiently and accurately in underwater environments, which are limited by problems such as insufficient light, color distortion and blurred targets, and the computing resources are limited, resulting in bottlenecks in the performance of detection algorithms.

Method used

Using the improved DETR model, the HIFI and EDCM modules are designed by building a backbone network, encoder and decoder, combining compact inverted bottleneck module, channel shuffling module, HiLo high and low frequency attention mechanism and multi-scale expansion convolution, to reduce the amount of model parameters and calculations and improve detection accuracy.

Benefits of technology

It realizes high-precision and efficient benthic biological detection in underwater environments, which is better than other target detection models, especially in complex underwater environments, and reduces the calculation cost and parameter quantity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236189A_ABST
    Figure CN120236189A_ABST
Patent Text Reader

Abstract

The invention relates to the field of underwater benthos detection, in particular to an underwater benthos detection method based on an improved DETR model. The method comprises the following steps: constructing an improved DETR model, and training by using a marked underwater image set to obtain a trained improved DETR model; using the trained improved DETR model to detect benthic organisms in the underwater image; wherein the improved DETR model comprises a backbone network, an encoder and a decoder with an auxiliary prediction head; in the backbone network, extracting image features by adopting a convolution gating linear unit which introduces a compact inverted bottleneck module and a channel shuffling module; in an encoder, a HIFI module combining a HiLo high and low frequency attention mechanism and multi-scale expansion convolution is adopted to enhance local detail preservation and global environment understanding, and an EDCM module is adopted to enhance feature representation and calculation efficiency of an underwater scene. According to the invention, underwater benthos detection can be realized with high precision and high efficiency, and the method is suitable for an underwater environment with limited resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of underwater benthic organism detection, and particularly to an underwater benthic organism detection method based on an improved DETR model. Background Art

[0002] With the continuous in-depth research of marine science, benthic organisms, as an important part of the marine ecosystem, their research and detection have become increasingly important. These organisms inhabit the seabed, are key links in the marine food chain, and play an indispensable role in maintaining the marine ecological balance. At the same time, the diversity and quantity of benthic organisms also have an important indicative effect on the health status of the marine environment. For example, the population changes of sea urchins, sea cucumbers, scallops and starfish usually reflect the ecological pressure or pollution degree of the water body. Therefore, the detection and monitoring of benthic organisms have important scientific value and practical significance, and involve fields such as marine ecological protection, environmental governance and fishery management.

[0003] Traditional benthic organism investigation methods usually rely on manual operations, such as diving sampling or image observation, and identify them through manual classification. These methods not only consume a large amount of manpower and time, but are also easily affected by human experience and subjective factors, resulting in low efficiency and high error rates. In addition, with the expansion of monitoring requirements and the development of automation technology, there is an urgent need for more efficient and accurate automated benthic organism detection methods. However, the complexity of the underwater environment poses significant challenges to the automated detection of benthic organisms. First of all, due to the scattering and absorption of light, underwater images often show phenomena such as insufficient light, color distortion and blurred targets. The interference of suspended particles and plankton further exacerbates the difficulty of extracting target features. Secondly, underwater embedded devices are usually limited by storage space and computing power, making detection algorithms with large computational amounts face serious performance bottlenecks in practical applications.

[0004] Therefore, developing a lightweight object detection model suitable for the underwater environment has become an urgent problem to be solved. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the defects of the prior art and provide an underwater benthic organism detection method based on an improved DETR model, which can achieve underwater benthic organism detection with high precision and efficiency and is suitable for resource-constrained underwater environments.

[0006] To solve the above technical problems, the technical solution of the present invention is: an underwater benthic organism detection method based on an improved DETR model, including:

[0007] Build an improved DETR model, train it using the annotated underwater image set, and obtain the trained improved DETR model; use the trained improved DETR model to detect benthic organisms in underwater images; among them,

[0008] The improved DETR model includes a backbone network, an encoder, and a decoder with an auxiliary prediction head;

[0009] In the backbone network, a convolutional gated linear unit introducing a compact inverted bottleneck module and a channel shuffle module is used to extract image features;

[0010] In the encoder, a HIFI module combining HiLo high-low frequency attention mechanism and multi-scale dilated convolution is used to enhance local detail preservation and global environment understanding, and an EDCM module is used to enhance the feature representation and computational efficiency of underwater scenes.

[0011] Furthermore, a convolutional gated linear unit introducing a compact inverted bottleneck module and a channel shuffle module is used to extract image features; specifically including:

[0012] First, perform a 1×1 convolution on the input image and split it into two branches;

[0013] Then, after passing one branch through linear transformation, CIB convolution operation, and channel shuffle operation in sequence, multiply it element-wise with the other branch after linear transformation;

[0014] Finally, output through a 1×1 convolution kernel residual connection.

[0015] Furthermore, the CIB convolution operation is specifically:

[0016] Perform 1×1 convolution, 3×3 convolution, and 1×1 convolution on the input features in sequence; the formula is:

[0017] Y1 = X * W1

[0018] Y2 = Y1 * W2

[0019] Y3 = Y2 * W3

[0020] In the formula, X represents the input feature of the CIB convolution operation; respectively represent the convolution kernel weight parameter matrices corresponding to the three convolutions.

[0021] Furthermore, the HIFI module adopts a dual-path structure, namely a high-frequency path and a low-frequency path, and both the high-frequency path and the low-frequency path include a scaled dot-product attention module and a multi-scale dilated convolution module set in sequence;

[0022] The working process of the HIFI module is:

[0023] The feature map of the input HIFI module is window-divided to form high-frequency features; the high-frequency features are downsampled to obtain low-frequency features;

[0024] The query vector Q, key vector K, and value vector V of the scaled dot-product attention module in the high-frequency path take the high-frequency features as input;

[0025] The key vector K and value vector V of the scaled dot-product attention module in the low-frequency path take the low-frequency features as input, and the query vector Q takes the feature map of the input HIFI module as input;

[0026] The outputs of the high-frequency path and the low-frequency path are concatenated as the output of the HIFI module.

[0027] Furthermore, the EDCM module combines a dynamic group convolution module, a ghost convolution module, and an efficient channel attention ECA module;

[0028] The working process of the EDCM module is as follows:

[0029] The number of channels of the input feature map is adjusted by a 1×1 convolution, and then the feature map is divided into two parts along the number of channels;

[0030] One part of the features is input into the dynamic group convolution module, and the other part of the features is input into the ghost convolution module;

[0031] The outputs of the dynamic group convolution module and the ghost convolution module are concatenated and then input into the efficient channel attention ECA module;

[0032] Two convolution operations are performed on the output of the efficient channel attention ECA module, and then the output of the efficient channel attention ECA module is connected as the output of the EDCM module.

[0033] Furthermore, the dynamic group convolution module includes pointwise convolution, depthwise convolution, and channel shuffle operations; the formula is:

[0034] X1' = Conv 1×1 (X1)

[0035] X 11 ,X 12 = Split(X1')

[0036] X DGSM = Conv 1×1 (Concat(X 12 ,Shuffle(DWConv 3×3 (X 11 ))))

[0037] In the formula, X1 is the input of the dynamic group convolution module; Conv 1×1( ) represents a 1×1 convolution; Split( ) represents splitting along the channels; DWConv 3×3 ( ) represents a depth convolution; Shuffle( ) represents a channel shuffle operation; Concat( ) represents a concatenation operation; X DGSM represents the output of the dynamic group convolution module.

[0038] Furthermore, the ghost convolution module includes a main convolution and a cheap convolution operation; the formula is:

[0039] X Ghost = Concat(Conv primary (X2), Conv cheap (Conv primary (X2)))

[0040] In the formula, X2 is the input of the ghost convolution module; Conv primary ( ) represents the main convolution operation; Conv cheap ( ) represents the cheap convolution operation; Concat( ) represents the concatenation operation.

[0041] Furthermore, before the input image enters the convolutional gated linear unit that introduces the compact inverted bottleneck module and the channel shuffle module in the first layer of the backbone network, it passes through three convolutional layers and one max pooling layer in sequence to obtain a preliminary low-level feature map; among them,

[0042] The first convolutional layer is used to achieve downsampling, and the latter two convolutional layers are used to extract local textures.

[0043] Furthermore, the convolutional gated linear unit that introduces the compact inverted bottleneck module and the channel shuffle module is denoted as the CGLU module;

[0044] The backbone network extracts multi-scale features of the image through multiple cascaded CGLU modules;

[0045] The encoder processes different-scale features through multiple cascaded EDCM modules in series. After the outputs of each path are fused, the final multi-scale fusion features are formed and output to the Query selector; among them,

[0046] The last-scale feature is processed by the HIFI module and then input into the corresponding EDCM module.

[0047] After adopting the above technical solutions, the present invention first improves the backbone of RT-DETR using the lightweight structure CIB-ConvShuffle GLU, reducing the model's parameter quantity and computational complexity. Secondly, by combining the HiLo high-low frequency attention mechanism and multi-scale dilated convolution, the HIFI module is designed to significantly improve the model's accuracy. Finally, the EDCM module is designed to further reduce the model's parameter quantity and computational complexity. Experiments show that the improved DETR model (LUW-DETR model) in the present invention achieves an mAP of 83.1% on the URPC 2020 dataset. At the same time, the FLOPs of the LUW-DETR model in the present invention are 40.1G, a 29.6% reduction compared to the original model, and the parameter quantity is reduced by 28.3%. Therefore, the method in the present invention has certain advantages in terms of accuracy and computational complexity, achieving high-precision and low-computation underwater benthic detection, achieving a good balance in underwater target detection, and outperforming other target detection models in complex underwater environments. Description of the Drawings

[0048] Figure 1 It is a flowchart of the underwater benthic detection method based on the improved DETR model of the present invention;

[0049] Figure 2 It is a schematic diagram of the framework of the improved DETR model of the present invention;

[0050] Figure 3 It is a detailed diagram of the framework of the improved DETR model of the present invention;

[0051] Figure 4 It is a structural diagram of the CIB-ConvShuffle GLU module (CGUL module) of the present invention;

[0052] Figure 5 It is a structural diagram of the HIFI module of the present invention;

[0053] Figure 6 It is a structural diagram of the EDCM module of the present invention;

[0054] Figure 7 It is a P-R curve of the test results of the improved DETR model of the present invention on the dataset URPC 2020;

[0055] Figure 8 It is a confusion matrix of the test results of the improved DETR model of the present invention on the dataset URPC 2020;

[0056] Figure 9 It is a detection result diagram of the traditional model and the improved DETR model of the present invention for the test set. Detailed Embodiments

[0057] To make the content of the present invention more clearly understood, the following further details the present invention according to specific embodiments in conjunction with the accompanying drawings.

[0058] As Figure 1 shown, an underwater benthic organism detection method based on an improved DETR model includes:

[0059] Construct an improved DETR model, train it using the annotated underwater image set to obtain a trained improved DETR model; use the trained improved DETR model to detect benthic organisms in underwater images.

[0060] As Figure 2 and Figure 3 shown, the improved DETR model (LUW-DETR) includes a backbone network, an encoder, and a decoder with an auxiliary prediction head;

[0061] In the backbone network, a convolutional gated linear unit introducing a compact inverted bottleneck module and a channel shuffle module is used to extract image features;

[0062] In the encoder, a HIFI module combining HiLo high-low frequency attention mechanism and multi-scale dilated convolution is used to enhance local detail preservation and global environment understanding, and an EDCM module is used to enhance the feature representation and computational efficiency of the underwater scene.

[0063] In this embodiment, the convolutional gated linear unit introducing a compact inverted bottleneck module and a channel shuffle module is denoted as the CIB-ConvShuffle GLU module, abbreviated as the CGLU module. As Figure 2 and Figure 3 shown, the backbone network extracts multi-scale image features through a series of CGLU modules; the encoder processes features of different scales through a series of stacked EDCM modules; among them, the features of the last scale are processed by the HIFI module and then input into the corresponding EDCM module; after feature fusion through two paths of bottom-up and top-down, the final multi-scale fusion features are formed and output to the Query selector; the decoder first selects a fixed number of image features from the encoder through the IoU-aware query module as the initial object query, and then generates prediction boxes and confidence scores through iterative optimization.

[0064] In this embodiment, preferably, as Figure 2 and Figure 3 shown, before the input image enters the first convolutional gated linear unit introducing a compact inverted bottleneck module and a channel shuffle module of the backbone network, it sequentially passes through three convolutional layers and one max pooling layer to obtain a preliminary low-level feature map; among them, the first convolutional layer is used for downsampling, and the latter two convolutional layers are used for extracting local textures.

[0065] This combination can effectively reduce the image size and extract the preliminary feature representation, providing an efficient and representative initial input for the subsequent backbone network CIB-ConvShuffle GLU module.

[0066] In this embodiment, preferably, as Figure 4 shown, a convolutional gated linear unit that introduces a compact inverted bottleneck module and a channel shuffle module (i.e., the CIB-ConvShuffle GLU module, abbreviated as the CGLU module) is used to extract image features; specifically, it includes:

[0067] First, perform a 1×1 convolution on the input image and split it into two branches;

[0068] Then, after passing one branch through a linear transformation, a CIB convolution operation, and a channel shuffle operation in sequence, multiply it element-wise with the other branch after passing through a linear transformation;

[0069] Finally, output through a 1×1 convolution kernel residual connection.

[0070] Among them, the CIB convolution operation is specifically as follows: given the input feature map dimension as and the output feature map as First, perform a pointwise convolution to compress the input channels to the intermediate channel number C mid , then, the depth convolution only acts on the spatial dimension to further capture fine-grained features, and finally, restore the dimension through a pointwise convolution. That is,

[0071] Perform 1×1 convolution, 3×3 convolution, and 1×1 convolution on the input features in sequence; the formula is:

[0072] Y1 = X * W1

[0073] Y2 = Y1 * W2

[0074] Y3 = Y2 * W3

[0075] In the formula, X represents the input feature of the CIB convolution operation; respectively represent the convolution kernel weight parameter matrices corresponding to the three convolutions.

[0076] This combined CIB module design can significantly reduce the computational cost, and the total computational amount (FLOPs CIBmid ) can be expressed as

[0077] FLOPs CIB = H × W × (C in × C mid + k 2 × C mid + C mid × C out)

[0078] Among them, k = 3 is the size of the convolution kernel, and C mid = αC out is the intermediate number of channels determined by the compression ratio α.

[0079] Among them, the core idea of the channel shuffle operation is to group the channels and exchange information between groups. Assume that the number of input channels C is divided into g groups, and the size of each group is The calculation process of the channel shuffle can be expressed as:

[0080]

[0081] In the formula, Y3 represents the input feature map of the channel shuffle operation, with dimensions (B, C, H, W); g represents the number of groups, that is, the channels are divided into several groups and then shuffled; .reshape() and.permute() are dimension rearrangement functions provided by the pytorch framework, and their function is to realize the rearrangement and combination of data in different dimensions.

[0082] This channel shuffle significantly improves the efficiency of component information interaction, alleviates the problem of channel isolation in deep convolution, and enables the model to better capture typical complex background features in the underwater environment.

[0083] Among them, the traditional GLU consists of two parallel linear transformations. After one transformation is gated, it is multiplied element-wise with the result of the other linear transformation. This structure enables the network to control the complexity of the information flow and effectively select information. In this embodiment, in order to better adapt to the characteristics of underwater target detection, by adding a CIB convolutional layer before the gated part of the GLU, the compression and screening process of the input features can be optimized, thereby effectively improving the feature expression ability of the model.

[0084] Through the above analysis, the CGLU module in this embodiment introduces the ConvShuffle operation, i.e., the channel shuffle operation, into the compact inverted bottleneck (CIB) structure. The core idea of this operation is to group channels and exchange information between groups, effectively achieving cross-channel information mixing, significantly enhancing the efficiency of component information interaction, enabling the model to better capture typical complex background features in the underwater environment, and using the convolutional gated linear unit (GLU) to enhance the non-linear feature selection ability. These designs cooperate with each other to enhance feature representation. While effectively maintaining a low computational complexity, it can extract more robust and discriminative features from degraded underwater images. Therefore, the CIB-ConvShuffle GLU module achieves a balance between lightweight design and robust feature extraction. By combining CIB, ConvShuffle, and GLU, it effectively solves the challenges of underwater object detection, including color distortion, low contrast, and noisy background, and shows good performance in underwater vision scenarios with color distortion, noise interference, and blurred object boundaries, significantly improving the detection accuracy and computational efficiency, making it suitable for deployment on embedded underwater devices.

[0085] In this embodiment, preferably, as Figure 5 shown, the HIFI module adopts a dual-path structure, namely a high-frequency path and a low-frequency path. Both the high-frequency path and the low-frequency path include a scaled dot product attention module and a multi-scale dilated convolution module arranged in sequence;

[0086] The working process of the HIFI module is as follows:

[0087] The feature map input to the HIFI module is windowed to form high-frequency features; the high-frequency features are downsampled to obtain low-frequency features;

[0088] The query vector Q, key vector K, and value vector V of the scaled dot product attention module in the high-frequency path take the high-frequency features as input;

[0089] The key vector K and value vector V of the scaled dot product attention module in the low-frequency path take the low-frequency features as input, and the query vector Q takes the feature map input to the HIFI module as input;

[0090] The outputs of the high-frequency path and the low-frequency path are concatenated as the output of the HIFI module.

[0091] Specifically, in the HIFI module, multiple heads are assigned to high-frequency attention. Through multi-scale dilated convolution and local window self-attention mechanism, fine-grained high-frequency features are extracted to capture local details, which is obviously more effective than the standard MSA. For low-frequency attention, first, average pooling is performed on each window to obtain low-frequency signals, and the remaining heads are assigned to low-frequency attention to simulate the relationship between each query position in the input feature map and the low-frequency keys and values averaged and pooled in each window. Then, the features output by high-frequency and low-frequency attention are further enhanced by multi-scale dilated convolution to obtain the local feature X Hi-Fi / Lo-Fi,Conv = Concat(DilatedConv(X split , r i ))), where r i is the dilation rate. Finally, the high-frequency and low-frequency path features are concatenated to form the comprehensive feature representation X HIFI = Concat(X Hi-Fi , X Lo-Fi ). After linear transformation, the final feature representation X out = W out X HIFI is obtained, where W out ∈R C×C is the linear transformation matrix.

[0092] Through the above analysis, the HIFI module in this embodiment introduces a dual-path structure, separating the processing of high-frequency and low-frequency feature components. The high-frequency path captures fine-grained spatial details by combining window-based local self-attention and multi-scale dilated convolution (MSDC). This design effectively expands the receptive field while maintaining high spatial resolution, which is crucial for detecting small targets and objects with blurred boundaries in the underwater environment. On the other hand, the low-frequency path effectively captures global context information using downsampled features, ensuring consistency in complex underwater scenes. Therefore, this dual-path design enables HIFI to better balance local and global feature extraction. The high-frequency path enhances the model's sensitivity to fine details, which is crucial for detecting small and poorly contrasted underwater objects. At the same time, the low-frequency path ensures global semantic understanding and can better identify objects from complex underwater backgrounds.

[0093] In addition, by decoupling high-frequency and low-frequency attention in the HIFI module, compared with the standard multi-head self-attention, the HIFI module significantly reduces the computational overhead. In the high-frequency path, the calculation is performed for localization within a window of size s, while the low-frequency path operates on the downsampled representation. Assuming that both paths use half of the total attention heads, the total complexity can be approximated as:

[0094]

[0095] This is lower than the standard attention complexity o(4ND2 +2N 2 D) More effective, enabling HIFI to be applicable to high-resolution underwater detection tasks on devices with limited resources.

[0096] In underwater target detection tasks, images usually have problems such as low contrast, severe color distortion, and blurred boundaries. Therefore, achieving robust feature representation with limited computing resources remains a major challenge. In this embodiment, the combination of multi-scale dilated convolution and frequency division attention enables the HIFI module to adapt to the unique challenges of the underwater environment. MSDC enhances the local feature aggregation at different spatial scales, effectively captures targets of different sizes, and processes the distortion caused by underwater imaging conditions. Low-frequency global attention ensures the retention of context information, which helps to accurately identify objects in chaotic scenes.

[0097] In this embodiment, preferably, as Figure 6 shown, the EDCM module combines a dynamic group convolution module, a ghost convolution module, and an efficient channel attention ECA module;

[0098] The working process of the EDCM module is as follows:

[0099] Adjust the number of channels of the input feature map through a 1×1 convolution, and then divide the features into two parts along the number of channels;

[0100] One part of the features is input into the dynamic group convolution module, and the other part of the features is input into the ghost convolution module;

[0101] The outputs of the dynamic group convolution module and the ghost convolution module are concatenated and then input into the efficient channel attention ECA module;

[0102] Perform two convolution operations on the output of the efficient channel attention ECA module, and then connect the output of the efficient channel attention ECA module as the output of the EDCM module.

[0103] Among them, the dynamic group convolution module includes pointwise convolution, depthwise convolution, and channel shuffle operations; the formula is:

[0104] X1' = Conv 1×1 (X1)

[0105] X 11 ,X 12 = Split(X1')

[0106] X DGSM = Conv 1×1 (Concat(X 12 ,Shuffle(DWConv 3×3 (X 11 ))))

[0107] Wherein, X1 is the input of the dynamic grouped convolution module; Conv 1×1 () represents a 1×1 convolution; Split() represents splitting along the channels; DWConv 3×3 () represents a depth convolution; Shuffle() represents a channel shuffle operation; Concat() represents a concatenation operation; X DGSM represents the output of the dynamic grouped convolution module.

[0108] Specifically, the core of the EDCM module is a dynamic grouped convolution module (DGSM). For dynamic grouped convolution, since it divides the features into g groups and performs convolution calculations on each group separately, compared with the parameter quantity P of the traditional standard convolution s = k 2 ·C in ·C out , the parameter quantity of the grouped convolution is where g is the number of groups, k is the convolution kernel size, and C in and C out are the number of input and output channels respectively. It can be seen that the parameter quantity of the grouped convolution is reduced compared with the standard convolution by Therefore, the EDCM adopts dynamic group convolution and channel shuffle operations, ensuring efficient exchange of information between groups while reducing redundant calculations. The channel shuffle strategy realizes effective feature communication between groups, thus alleviating the feature isolation problem, which is a key issue for maintaining robust feature extraction in an underwater environment with color distortion and low visibility.

[0109] Among them, the ghost convolution module includes a main convolution and a cheap convolution operation; the formula is:

[0110] X Ghost = Concat(Conv primary (X2), Conv cheap (Conv primary (X2)))

[0111] Wherein, X2 is the input of the ghost convolution module; Conv primary () represents the main convolution operation; Conv cheap () represents the cheap convolution operation; Concat() represents the concatenation operation.

[0112] Specifically, the ghost convolution module decomposes the traditional convolution into a main convolution and a cheap convolution. The main convolution generates initial features, while the cheap operation generates additional redundant features through simple linear transformation. This decomposition design avoids redundant calculations while enhancing the diversity of feature expression.

[0113] In addition, EDCM incorporates the Efficient Channel Attention mechanism. Through adaptive pooling and convolutional operations, each channel is weighted, enhancing the selectivity and representational ability of features. After the high-frequency and low-frequency features are concatenated, channel compression is performed through linear transformation, and finally, through the ECA module, the selectivity and representational ability of features are further enhanced.

[0114] Through the above analysis, EDCM enhances the feature representation of underwater scenes and improves computational efficiency by combining Dynamic Group Convolution, Ghost Convolution, and the Efficient Channel Attention (ECA) mechanism. Moreover, while reducing the number of parameters and computational volume, it maintains the diversity and richness of features.

[0115] Next, through specific experiments, the advantages of the solutions involved in the above embodiments will be introduced in detail.

[0116] 1. Dataset

[0117] The dataset used in this paper is the dataset URPC 2020 of the underwater target detection algorithm competition in the National Underwater Robot Competition. The URPC2020 dataset contains multi-scene underwater optical images of five types of underwater benthic biological targets, namely holothurian, echinus, scallop, starfish, and a small amount of waterweeds.

[0118] Since the number of seagrass-related samples is too small, in order to ensure that the training data can comprehensively capture the diversity and complexity of this category to improve the detection accuracy, the seagrass-related samples are excluded in this paper. The final dataset retains 5454 underwater optical images in JPG format, which are divided into a training set, a validation set, and a test set according to the ratio of 6:2:2. Among them, the training set contains 3272 images, and the validation set and the test set each contain 1091 images, and the size of the input images is uniformly processed to 640×640.

[0119] 2. Experimental Details and Evaluation Metrics

[0120] The experiment is based on Python 3.8 and Pytorch 1.13.1, and the GPU used for training is NVIDIA GeForce RTX3090. All comparative experiments are carried out under the same parameter settings. The Adam optimizer is used for training, the batch size is set to 8, and the number of epochs is set to 200. The learning rate decay strategy is set. The initial value of the learning rate is 0.01, the momentum coefficient is set to 0.937, and the weight decay coefficient is set to 0.0005 to prevent overfitting.

[0121] Each image in the object detection task may contain some objects of different classes, and it is necessary to evaluate the object classification and localization performance of the model. In object detection, we use precision and recall as evaluation metrics for detection accuracy. Precision measures the accuracy of the model in classifying samples as positive samples, and recall measures the ability of the model to detect positive samples. Their definition formulas are as follows

[0122]

[0123] Among them, TP represents the number of detection boxes with IoU > IoU threshold , FP represents the number of detection boxes with IOU ≤ IOUthreshold, and FN represents the number of undetected targets.

[0124] Precision and recall interact with each other. To combine these two metrics, the Average Precision (AP) is introduced to measure the detection accuracy. The definition formula is as follows, where N is the number of object classes in the dataset. Since the model needs to consider both the accuracy and the number of model parameters, the computational complexity (FLOPs) and the number of parameters (Parameters) are also important evaluation metrics.

[0125]

[0126] 3. Comparative Experiment Results

[0127] To further verify the performance of LUW-DETR, this paper compares it with several classic algorithms such as YOLOv7, YOLOv7-tiny, YOLOv8l, Deformable-DETR, DINO, and RT-DETR. In addition, we included recent advanced detectors such as YOLOv9s in the evaluation. Comparative experiments were carried out using the same dataset and training method. The results of the comparative experiments are shown in Table 1:

[0128] Table 1 Comparative Experiments

[0129]

[0130] As shown in Table 1, compared with the YOLO series models, LUW-DETR shows superior performance with an mAP of 83.1%. Although its accuracy is slightly lower than that of the DINO model (83.4%), LUW-DETR has significant advantages in terms of computational efficiency and inference speed. Specifically, DINO has the highest accuracy, but its computational cost is extremely high, with 45.15M parameters and 178.6 GFLOPs, resulting in a low inference speed of only 5.1 FPS, which brings a heavy burden to underwater platforms with limited resources. Such high computational and storage requirements limit its practical deployment in real-time underwater applications.

[0131] In contrast, LUW-DETR reduces the computational cost by 30% compared to the original RT-DETR, reducing FLOPs from 57.0G to 40.1G, while still maintaining competitive accuracy and exceeding RT-DETR in terms of recall (76.4% vs. 76.1%) and mAP (83.1% vs. 82.7%). Although its FLOPs are still higher than those of lightweight YOLO models - such as YOLOv7-tiny (13.2G) and YOLOv9s (26.7G), LUW-DETR achieves a better balance between detection performance and computational efficiency.

[0132] Figure 7 Figure [8] shows the P-R curve of the test results of LUW-DETR on URPC2020. The value of the area under the curve for each category is the AP value for that category. The higher the AP value, the better the detection performance for that category. From Figure 7 it can be seen that the AP values of LUW-DETR for the sea cucumber and scallop categories are slightly lower than the average AP value for all categories, while the AP values for the starfish and sea urchin categories are higher than the average AP value for all categories.

[0133] Figure 8 Figure

[14] shows the confusion matrix of the test results of LUW-DETR on URPC 2020. In this confusion matrix, each row represents the predicted category, each column represents the actual category, and the values on the diagonal represent the proportion of correctly classified categories. The accuracy rates for the predicted categories "sea cucumber", "sea urchin", "starfish", and "scallop" are 80%, 96%, 93%, and 83% respectively, indicating that the model has a high detection accuracy.

[0134] Figure 9Detection result graphs of YOLOv7-tiny, YOLOv9s, DINO, RT-DETR, and the proposed LUW-DETR for the test set. In terms of detection performance, YOLOv7t and YOLOv9s exhibit varying degrees of false positives and missed detections, especially in images with blurred, complex backgrounds, or underwater low-light conditions. For example, in images of dense and cluttered scenes, YOLOv7t often fails to detect small objects such as scallops and misidentifies background regions as sea urchins. Compared with YOLOv7t, YOLOv9s has better detection performance and can better recall small targets, but still has occasional false positives and missed detections, especially for objects located at the edges of images or with low visibility. The DINO model has high detection accuracy in identifying small targets such as sea urchins and starfish. However, due to its high model complexity and large computational cost, its inference speed is slow and it is not suitable for real-time underwater applications. In addition, in some cluttered scenes, DINO may generate redundant detection boxes, potentially increasing the false positive rate. RT-DETR shows a balance between detection accuracy and inference speed and is superior to YOLOv7t and YOLOv9s in complex scenes. However, under challenging underwater conditions, RT-DETR still encounters difficulties in accurately detecting small targets, such as blurred images and low contrast.

[0135] In contrast, the proposed LUW-DETR algorithm demonstrates superior detection performance in various underwater scenes. It effectively addresses issues such as image blur and complex backgrounds, significantly reducing false positives and missed detections. LUW-DETR has strong detection capabilities for small targets and low-contrast targets, maintaining high accuracy and recall even under challenging conditions. The qualitative results clearly demonstrate the advantages of LUW-DETR in underwater object detection tasks, verifying its applicability for deployment in complex and resource-constrained underwater environments.

[0136] Inspired by the ideal embodiments of the present invention described above, through the above description, relevant staff can make various changes and modifications without departing from the technical idea of the present invention. The technical scope of the present invention is not limited to the content in the specification and must be determined according to the scope of the claims.

Claims

1. A method for detecting underwater benthic organisms based on an improved DETR model, characterized in that: include: An improved DETR model is constructed and trained using a labeled underwater image set to obtain a trained improved DETR model; the trained improved DETR model is used to detect benthic organisms in underwater images; wherein, The improved DETR model includes a backbone network, an encoder, and a decoder with an auxiliary prediction head; In the backbone network, convolutional gated linear units with compact inverted bottleneck modules and channel shuffle modules are used to extract image features. In the encoder, the HIFI module that combines the HiLo high- and low-frequency attention mechanism and multi-scale dilated convolution is used to enhance local detail preservation and global environment understanding, and the EDCM module is used to enhance the feature representation and computational efficiency of underwater scenes.

2. The underwater benthic organism detection method based on the improved DETR model according to claim 1, characterized in that: The convolutional gated linear unit with a compact inverted bottleneck module and a channel shuffle module is used to extract image features; specifically, it includes: First, perform a 1×1 convolution on the input image and split it into two branches; Then one branch is sequentially subjected to linear transformation, CIB convolution operation and channel shuffle operation, and then multiplied element-by-element with the other branch that has undergone linear transformation; Then the output is connected through a 1×1 convolution kernel residual connection.

3. The underwater benthic organism detection method based on the improved DETR model according to claim 2 is characterized in that: The CIB convolution operation is as follows: The input features are sequentially subjected to 1×1 convolution, 3×3 convolution, and 1×1 convolution; the formula is: Y1=X*W1 Y2=Y1*W2 Y3=Y2*W3 Where X represents the input feature of the CIB convolution operation; They represent the convolution kernel weight parameter matrices corresponding to the three convolutions respectively.

4. The underwater benthic organism detection method based on the improved DETR model according to claim 1, characterized in that: The HIFI module adopts a dual-path structure, namely a high-frequency path and a low-frequency path. Both the high-frequency path and the low-frequency path include a scaled dot-product attention module and a multi-scale dilated convolution module set in sequence. The working process of the HIFI module is: The feature map of the input HIFI module is divided into windows to form high-frequency features; the high-frequency features are downsampled to obtain low-frequency features; The query vector Q, key vector K, and value vector V of the scaled dot product attention module of the high-frequency path take high-frequency features as input; The key vector K and value vector V of the scaled dot product attention module of the low-frequency path take the low-frequency features as input, and the query vector Q takes the feature map of the input HIFI module as input; The outputs of the high-frequency path and the low-frequency path are spliced ​​as the output of the HIFI module.

5. The underwater benthic organism detection method based on the improved DETR model according to claim 1, characterized in that: The EDCM module combines a dynamic grouping convolution module, a ghost convolution module and an efficient channel attention ECA module; The working process of the EDCM module is: The number of channels of the input feature map is adjusted through a 1×1 convolution, and then the features are divided into two parts along the number of channels; Some features are input into the dynamic group convolution module, and the other features are input into the ghost convolution module; The outputs of the dynamic grouped convolution module and the ghost convolution module are concatenated and input into the efficient channel attention ECA module; The output of the efficient channel attention ECA module is convolved twice and then concatenated as the output of the EDCM module.

6. The underwater benthic organism detection method based on the improved DETR model according to claim 5, characterized in that: The dynamic group convolution module includes point-by-point convolution, depth-wise convolution, and channel shuffle operations; the formula is: X1'=Conv 1×1 (X1) X 11 ,X 12 =Split(X1') X DGSM =Conv 1×1 (Concat(X 12 ,Shuffle(DWConv 3×3 (X 11 )))) Where X1 is the input of the dynamic group convolution module; Conv 1×1 () indicates 1×1 convolution; Split() indicates splitting along the channel; DWConv 3×3 () represents depthwise convolution; Shuffle() represents channel shuffling operation; Concat() represents concatenation operation; X DGSM Represents the output of the dynamic grouped convolution module.

7. The underwater benthic organism detection method based on the improved DETR model according to claim 5, characterized in that: The ghost convolution module includes the main convolution and cheap convolution operations; the formula is: X Ghost =Concat(Conv primary (X2),Conv cheap (Conv primary (X2))) Where X2 is the input of the ghost convolution module; Conv primary () indicates the main convolution operation; Conv cheap () represents a cheap convolution operation; Concat() represents a concatenation operation.

8. The underwater benthic organism detection method based on the improved DETR model according to claim 1, characterized in that: Before entering the first convolutional gated linear unit of the backbone network that introduces a compact inverted bottleneck module and a channel shuffle module, the input image passes through three convolutional layers and one maximum pooling layer in sequence to obtain a preliminary low-level feature map; among them, The first convolutional layer is used to achieve downsampling, and the next two convolutional layers are used to extract local textures.

9. The underwater benthic organism detection method based on the improved DETR model according to claim 1, characterized in that: The convolutional gated linear unit that introduces the compact inverted bottleneck module and the channel shuffle module is denoted as the CGLU module; The backbone network extracts multi-scale features of the image through multiple CGLU modules connected in series; The encoder processes different scale features through multiple layers of EDCM modules connected in series. After the outputs of each path are fused, the final multi-scale fusion feature is formed and output to the query selector. The last scale feature is processed by the HIFI module and then input into the corresponding EDCM module.

Citation Information

Cited By

  • Rapid dam crack identification method

    CN121837647A

  • A fast dam crack identification method

    CN121837647B