Improved underwater optical target detection method based on RT-DETR
By using the fusion method of RT-DETR model and multi-scale adaptive weighted feature in underwater recognition technology, a lightweight underwater target detection model is built, which solves the problem of difficulty in achieving accurate target detection under limited hardware resources, and achieves efficient and accurate underwater target detection.
Patent Information
- Application Number
- CN202510131169.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-06-10
AI Technical Summary
Existing underwater identification technology is difficult to achieve accurate target detection under limited hardware resources.
A lightweight underwater object detection model is constructed using a method based on the RT-DETR model, SepNPFCSP_ELAN module, Hilo attention mechanism and multi-scale adaptive weighted feature fusion.
It achieves high detection accuracy under less computing resources, and solves the problem of accurate object detection by underwater recognition technology under limited hardware resources.
Smart Images

Figure CN120125980A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of underwater intelligent recognition, and particularly to an improved underwater optical target detection method based on RT-DETR. Background Art
[0002] With the development of computer technology, deep learning technology has been more widely applied in various fields, and the marine field has broad application prospects. Compared with traditional algorithms, deep learning has the characteristics of higher accuracy and faster recognition speed, and its accuracy can be continuously improved with the continuous collection of data.
[0003] However, traditional machine learning methods face great challenges in the recognition tasks of underwater complex environments. For example, the amount of calculation for underwater recognition is large, and higher-performance computing devices are required to maintain recognition accuracy, but it is difficult to carry too many hardware devices in the underwater complex environment. Summary of the Invention
[0004] The purpose of the present invention is to provide an improved underwater optical target detection method based on RT-DETR to solve the problem that it is difficult to achieve accurate target detection with limited hardware resources in existing underwater recognition technologies.
[0005] The above object of the present invention is achieved through the following technical solutions:
[0006] S1: Obtain the DUO underwater dataset and underwater images and perform preprocessing to generate a training set;
[0007] S2: Based on the RT-DETR model, SepNPFCSP_ELAN module, Hilo attention mechanism, and multi-scale adaptive weighted feature fusion method, construct a lightweight underwater target detection model;
[0008] S3: Train the underwater target detection model with the training set;
[0009] S4: Obtain the underwater images to be classified; perform underwater object target detection on the underwater images to be classified through the trained underwater target detection model to obtain the target detection results of underwater objects.
[0010] Optionally, step S1 includes:
[0011] Perform preprocessing on the images in the DUO underwater dataset and underwater images, and the preprocessing includes: color balance, image denoising, and data annotation.
[0012] Optionally, step S2 includes:
[0013] S2-1: Improve the backbone structure of the underwater target detection model through the lightweight SepNPFCSP_ELAN module;
[0014] S2-2: Improve the SepNPFCSP_ELAN module of the underwater target detection model;
[0015] S2-3: Improve the AIFI module of the underwater target detection model through the Hilo attention mechanism to obtain the AIFI-Hilo module;
[0016] S2-4: Adopt the improved SepNPFCSP_ELAN module, AIFI-Hilo module and EUCB module, and combine the method of multi-scale adaptive weighted feature fusion to design the CCFM module.
[0017] Optionally, step S2-1 includes:
[0018] The improved backbone structure includes: a convolutional block, four SepNPFCSP_ELAN modules, and an Adown module;
[0019] The Adown module is used to downsample the input features;
[0020] The SepNPFCSP_ELAN module is used to perform multi-scale feature extraction and fusion on the input features.
[0021] Optionally, step S2-2 includes:
[0022] The improved SepNPFCSP_ELAN module includes: the first DSConv depthwise separable convolution, a split unit, the first PFCSP module, the second DSConv depthwise separable convolution, the third DSConv depthwise separable convolution, the second PFCSP module, a fully connected unit, and the fourth DSConv depthwise separable convolution;
[0023] The first DSConv depthwise separable convolution is connected to the split unit; the split unit is connected to the fully connected unit;
[0024] The split unit is connected to the first PFCSP module; the first PFCSP module is connected to the second DSConv depthwise separable convolution; the second DSConv depthwise separable convolution is connected to the fully connected unit;
[0025] The first PFCSP module is connected to the third DSConv depthwise separable convolution; the third DSConv depthwise separable convolution is connected to the fully connected unit;
[0026] The third DSConv depthwise separable convolution is connected to the second PFCSP module;
[0027] The second PFCSP module is connected to the second DSConv depthwise separable convolution;
[0028] The fully connected unit is connected to the fourth DSConv depthwise separable convolution connection.
[0029] Optionally, the structures of the first PFCSP module and the second PFCSP module are the same. The first PFCSP module includes: a DSConv depthwise separable convolution, a PFBottleneck module, and a fully connected layer; the PFBottleneck module includes: a PFConv module and a DSConv depthwise separable convolution;
[0030] The PFConv module includes: a convolution block, a silu activation unit, and a sigmod activation unit.
[0031] Optionally, step S2-4 includes:
[0032] The CCFM module includes: an EUCB module, a SepNPFCSP_ELAN module, an AIFI-Hilo module, and a Fusion module;
[0033] The input features are upsampled by the EUCB module; the feature maps of different scales are subjected to multi-scale feature fusion by the Fusion module.
[0034] An improved underwater optical target detection system based on RT-DETR, the system includes: a training set construction module, a model construction module, a model training module, and a target detection module;
[0035] The training set construction module, the model construction module, the model training module, and the target detection module are sequentially connected;
[0036] The training set construction module is used to obtain the DUO underwater data set and underwater pictures and perform preprocessing to generate a training set;
[0037] The model construction module is used to construct a lightweight underwater target detection model based on the RT-DETR model, the SepNPFCSP_ELAN module, the Hilo attention mechanism, and the method of multi-scale adaptive weighted feature fusion;
[0038] The model training module is used to train the underwater target detection model through the training set;
[0039] The target detection module is used to obtain the underwater pictures to be classified; the trained underwater target detection model is used to perform underwater object target detection on the underwater pictures to be classified, and obtain the target detection result of the underwater object.
[0040] The beneficial effects brought by the technical solution provided by the present invention are as follows:
[0041] Based on the RT-DETR model, SepNPFCSP_ELAN module, Hilo attention mechanism, and multi-scale adaptive weighted feature fusion method, a lightweight underwater target detection model is constructed. The lightweight underwater target detection model is used to detect underwater object targets in the underwater pictures to be classified, and the target detection results of underwater objects are obtained, realizing high detection accuracy with less computational resources consumed. It solves the problem that the existing underwater recognition technology is difficult to achieve accurate target detection under limited hardware resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:
[0043] Figure 1 is the step diagram in the embodiment of the present invention;
[0044] Figure 2 is the structural diagram of the improved RT DETR underwater target detection model provided by the present invention;
[0045] Figure 3 is the structural diagram of the SepNPFCSP_ELAN module provided by the present invention;
[0046] Figure 4 is the structural diagram of the PFSCP module provided by the present invention;
[0047] Figure 5 is the structural diagram of the PFBottleneck module provided by the present invention;
[0048] Figure 6 is the structural diagram of the PFConv module provided by the present invention;
[0049] Figure 7 is the structural diagram of the AiFi-Hilo module provided by the present invention;
[0050] Figure 8 is the structural diagram of the EUCB module provided by the present invention;
[0051] Figure 9 is the structural diagram of the module provided by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0052] In order to have a clearer understanding of the technical features, objectives, and effects of the present invention, the specific embodiments of the present invention will now be described in detail with reference to the drawings.
[0053] The embodiment of the present invention provides an improved underwater optical target detection method based on RT-DETR.
[0054] Please refer to Figure 1 , Figure 1 which is a step diagram of an improved underwater optical target detection method based on RT-DETR in an embodiment of the present invention, including:
[0055] S1: Obtain the DUO underwater dataset and underwater pictures and perform preprocessing to generate a training set;
[0056] S2: Based on the RT-DETR model, SepNPFCSP_ELAN module, Hilo attention mechanism, and multi-scale adaptive weighted feature fusion method, construct a lightweight underwater target detection model;
[0057] S3: Train the underwater target detection model with the training set;
[0058] S4: Obtain the underwater pictures to be classified; use the trained underwater target detection model to perform underwater object target detection on the underwater pictures to be classified, and obtain the target detection results of the underwater objects.
[0059] Step S1 includes:
[0060] Perform preprocessing on the pictures in the DUO underwater dataset and underwater pictures. The preprocessing includes: color balance, image denoising, and data annotation.
[0061] As an embodiment, first collect Internet resources to obtain common underwater objects in web pages, such as pictures of divers, fish, sea urchins, jellyfish, etc., and construct a dataset by combining the publicly available DUO (Detecting Underwater Objects) underwater data. And uniformly set the format and name all pictures for subsequent operations. Through the labelme software, perform data annotation on the collected pictures, select the positions of ten common underwater target objects such as divers, fish, sea urchins, jellyfish, etc., and extract the coordinates to construct a training set in yolo format.
[0062] Step S2 includes:
[0063] S2-1: Improve the backbone structure of the underwater target detection model through the lightweight SepNPFCSP_ELAN module;
[0064] S2-2: Improve the SepNPFCSP_ELAN module of the underwater target detection model;
[0065] S2-3: Improve the AIFI module of the underwater target detection model through the Hilo attention mechanism to obtain the AIFI-Hilo module;
[0066] As an embodiment, in the AIFI structure, a high-low frequency fusion mechanism is adopted to focus on the high-frequency and low-frequency information in the image. The multi-head self-attention mechanism is divided into two groups, one of which focuses on high-frequency information and encodes the high-frequency information, and the other group focuses on low-frequency information and encodes the low-frequency information. Finally, they are fused to achieve the effect of capturing local information with high frequency and global information with low frequency, and finally fusion to improve the detection accuracy. First, the input feature map is average pooled to reduce the spatial resolution and capture the low-frequency spatial information. Then the low-frequency l_q,
[0067] l_kv and output projection l_proj, these linear layers reduce the input dimension and focus on the low-frequency part. Then operations are performed on the processed features, where the l_q query is compared with the l_kv key value to calculate the attention weight, focusing on capturing global relationships. The high-frequency part does not directly use the input feature map without downsampling, which allows it to focus on high-frequency local details. Similarly, the high-frequency queries h_q and h_kv are also calculated through the first pass, and the output is projected to the high-frequency space through h_proj. After calculating the high-frequency and low-frequency attention weights, the two outputs hifi_out and lofi_out are spliced along the feature dimension, and the final output is projected again through the first pass. Finally, the effect of multi-scale feature extraction that takes into account both global and local is achieved.
[0068] As an example, Figure 7As shown in the figure, in AIFI-Holo, the HiLo attention mechanism encodes specific feature information by dividing multi-head self-attention into high-frequency and low-frequency parts, respectively, thereby significantly improving computational efficiency and model performance. In the high-frequency part (Hi-Fi), the mechanism focuses on encoding high-frequency interaction information in the image, such as details such as edges and textures. In specific implementation, by combining local self-attention with high-resolution feature maps, a 2×2 local window is used to capture fine-grained high-frequency features, and a non-overlapping window partitioning strategy is used to reduce the calculation scope and redundant operations, thereby effectively reducing the computational complexity. This processing method reduces the need for long-distance dependencies and is crucial for building fast and efficient visual transformers. In the low-frequency part (Lo-Fi), the mechanism encodes low-frequency interaction information, such as overall features such as shape and color distribution, through global self-attention and downsampled feature maps. Specifically, the feature map is average-pooled to generate keys (Key, K) and values (Value, V), while retaining the original feature map as the query (Query, Q), and then rich low-frequency information is captured through global self-attention. Since this part mainly acts on the downsampled feature map, it effectively reduces the number of features to be processed, further improving the computational efficiency. In order to achieve efficient allocation, HiLo divides the attention heads of the multi-head self-attention into two groups and allocates them according to the ratio a, where (1-a)N h heads are used for high-frequency attention, and the remaining aN h The first one is used for low-frequency attention. This allocation strategy ensures that both high-frequency and low-frequency features can be properly paid attention to and processed, avoiding resource waste. Finally, the improved high-frequency and low-frequency features are integrated as the output of the module to form the final result of the AIFI-HiLo module.
[0069] The formula is as follows:
[0070] Q = K = V = Flatten (input)
[0071] AIFI-HiLo(x)=[HiFi(x),LoFi(x)]
[0072] output=reshape(AIFI-HiLo(Q,K,V))
[0073] Among them, the Flatten operation flattens the feature map into a forward vector, which converts the input feature input into a shape suitable for the attention mechanism. The HiFi operation encodes high-frequency details in the feature, such as edges and textures, through local window operations. The low-frequency branch (LoFi) uses a global attention mechanism to encode low-frequency information of the feature, such as shape and color distribution. Finally, the high-frequency and low-frequency features are fused together, and the shape output of the feature map is adjusted through the reshape operation.
[0074] S2-4: The CCFM module is designed by adopting the improved SepNPFCSP_ELAN module, AIFI-Hilo module and EUCB module, and combining the method of multi-scale adaptive weighted feature fusion.
[0075] As an embodiment, on the basis of the RT-DETR model (RT-DETR, Real-Time Detection Transformer), the backbone part is redesigned to improve its feature extraction part, and the lightweight SepNPFCSP_ELAN module is introduced to improve the feature extraction part, so as to reduce the model size and calculation amount. At the same time, Hilo is used to improve the AIFI part to improve the model detection accuracy. Finally, the structure of the CCFM part is redesigned, the idea of multi-scale adaptive weighted feature fusion is adopted, and an efficient upsampling module is introduced and redesigned into the AMSFPN structure.
[0076] As an embodiment, the present invention constructs an underwater target detection model adapted to the underwater environment by improving the RT-DETR model, effectively reducing the calculation amount and the number of parameters of the underwater target detection model, and improving the detection accuracy of underwater targets; it is adapted to the complex underwater environment and limited hardware devices, and can complete the task of underwater target detection under limited hardware resources.
[0077] Step S2-1 includes:
[0078] The improved backbone structure includes: a convolutional block, four SepNPFCSP_ELAN modules and an Adown module;
[0079] The Adown module is used to downsample the input features;
[0080] The SepNPFCSP_ELAN module is used to perform multi-scale feature extraction and fusion on the input features.
[0081] As an embodiment, such as Figure 2As shown in the figure, the backbone part of the RT-DETR network is redesigned. First, the original input image is subjected to feature extraction using two consecutive convolutions. Then, the Adown module is used for downsampling, reducing the resolution of the feature map, expanding the receptive field, and at the same time reducing the computational amount of the model. Then, the SepNPFCSP_ELAN module is used to perform multi-scale feature extraction and fusion on the input features. After each feature extraction, it passes through a convolutional layer. Finally, four feature layers, namely P2, P3, P4, and P5, are generated. These feature layers will perform multi-scale feature fusion in the CCFM structure later. Effective information transmission and integration are realized in multiple layers of features, enhancing the feature extraction ability for underwater targets and reducing redundant data to lighten the computational burden.
[0082] Step S2-2 includes:
[0083] The improved SepNPFCSP_ELAN module includes: the first DSConv depthwise separable convolution, the split unit, the first PFCSP module, the second DSConv depthwise separable convolution, the third DSConv depthwise separable convolution, the second PFCSP module, the fully connected unit, and the fourth DSConv depthwise separable convolution;
[0084] The first DSConv depthwise separable convolution is connected to the split unit; the split unit is connected to the fully connected unit;
[0085] The split unit is connected to the first PFCSP module; the first PFCSP module is connected to the second DSConv depthwise separable convolution; the second DSConv depthwise separable convolution is connected to the fully connected unit;
[0086] The first PFCSP module is connected to the third DSConv depthwise separable convolution; the third DSConv depthwise separable convolution is connected to the fully connected unit;
[0087] The third DSConv depthwise separable convolution is connected to the second PFCSP module;
[0088] The second PFCSP module is connected to the second DSConv depthwise separable convolution;
[0089] The fully connected unit is connected to the fourth DSConv depthwise separable convolution connection.
[0090] As an embodiment, the SepNPFCSP_ELAN feature extraction module is adopted in the feature extraction part to improve the original RepC3 module. The SepNPFCSP_ELAN module adopts a parameter-free attention mechanism to enhance the resolution of the feature map. Six consecutive LPFBlocks are used for feature extraction. Each LPFBlock sequentially extracts higher-level image features through three depthwise separable convolutions, then fuses the extracted image features with the input through residual fusion, and then obtains the attention map through an activation function. Finally, the elements in the feature map are rearranged through the Pixel Shuffle upsampling operation to improve the spatial resolution, and finally a super-resolution feature map is obtained. This structure can reduce the model computation while enhancing the feature extraction ability, adapting to the computing device environment with underwater priority.
[0091] As an embodiment, as Figures 3 - 6 shown, in the feature extraction module, due to the existence of a large amount of redundant computation in the original RepC3 module, which consumes a lot of computing resources, the feature extraction module (SepNPFCSP_ELAN module) is redesigned. As Figure 3 shown, this module first performs feature extraction on the input feature map through a 3*3 depthwise separable convolution, then divides the channels of the feature map into two parts. One part is further subjected to feature extraction, and multi-level feature fusion is performed through multiple PFCSP modules to further extract features, and then a depthwise separable convolution is performed again. Finally, the results of multiple feature extractions are fused together through Concat, and then the channel number and resolution are adjusted through a depthwise separable convolution for output. The PFCSP module here, as Figure 4 shown, divides the input feature map into two parts. One part is subjected to feature extraction through multiple PFBottleneck modules, and finally a concat operation is performed with the feature map directly subjected to depthwise separable convolution of the other part to fuse and obtain.
[0092] The structures of the first PFCSP module and the second PFCSP module are the same. The first PFCSP module includes: DSConv depthwise separable convolution, PFBottleneck module, and fully connected layer; the PFBottleneck module includes: PFConv module and DSConv depthwise separable convolution;
[0093] The PFConv module includes: convolutional block, silu activation unit, and sigmod activation unit.
[0094] As an embodiment, the PFBottleneck module here is as Figure 5As shown, a residual separation is performed on the input. One part directly performs the final residual connection, and the other part undergoes feature extraction through the PFConv module and finally performs the residual connection. The most crucial PFConv module is as Figure 6 shown. The input feature map goes through two consecutive convolutions and the silu activation function, and then through another convolution to align the number of channels and resolution to obtain H i . Then, H i is element-wise added to the original input image to obtain U i, . And after performing the sigmod activation operation on H i , the feature map V i is obtained. Finally, the final feature map is output after element-wise multiplication of V i and U i .
[0095] Its formula is as follows:
[0096]
[0097] Oi = Ui ο Vi
[0098]
[0099] Vi = σ(Hi)
[0100] Among them, and respectively represent the element-wise sum between feature extraction and residual connection and the convolution operation. ο represents element-wise multiplication, and σ represents the activation function operation after the convolution layer. It is precisely because of this special parameter-free Vi attention map that this structure greatly reduces the number of parameters while enhancing the feature extraction ability.
[0101] Step S2-4 includes:
[0102] The CCFM module includes: the EUCB module, the SepNPFCSP_ELAN module, the AIFI-Hilo module, and the Fusion module;
[0103] The input features are upsampled through the EUCB module; the feature maps of different scales are multi-scale feature fused through the Fusion module.
[0104] As an embodiment, on the redesigned CCFM structure, the idea of multi-scale weighted fusion is adopted for adaptive weighted fusion. By assigning self-learning weights to the inputs of the P3, P4, and P5 feature layers respectively, and then performing multiple feature extractions through the SepNPFCSP_ELAN module and fusing deeply with the context multiple times. After each feature extraction, it is fused with the feature extraction layers of different scales and different stages, achieving the detection effect that takes into account both high-dimensional and low-dimensional targets. At the same time, the efficient EUCB (Efficient Upsampling Convolutional Block) module is used to optimize the upsampling and improve the upsampling efficiency.
[0105] As an embodiment, a new CCFM module is designed to perform in-depth processing on the P2, P3, P4, and P5 layers output by the backbone structure. First, the feature map of the input P5 layer is adjusted in size through convolution operations after the AIFI-Hilo operation, and a Fusion feature fusion operation is performed with the input of the P4 layer. This Fusion feature fusion operation receives feature maps from different scales as inputs and can perform multi-scale feature fusion.
[0106] O i = w 1 ·P i-1 + w 2 ·P i + w 3 ·P i+1
[0107] Where O i represents the output feature map obtained after feature fusion; w i represents the learnable weight, and normalization is used to make the sum of w i equal to 1; P i represents the input feature maps of different levels.
[0108] Since features of different scales contribute differently to the object detection task. The weighted fusion mechanism can dynamically adjust the importance of each input feature through the learnable weight, thus more flexibly fusing features. Weighted fusion can more effectively combine low-level features (detail information) and high-level features (semantic information) to generate more expressive multi-scale features. At the same time, compared with the original structure, this feature fusion and extraction structure has a smaller computational amount and can accelerate the training process through fast normalization, with higher computational efficiency and is suitable for real-time detection applications.
[0109] After feature fusion through Fusion and feature extraction through SepNPFCSP_ELAN, upsampling is performed through the EUCB module, as Figure 8As shown in the figure. This module first doubles the feature map through double upsampling, and then passes through a 3×3 depthwise convolution DWC to preserve the spatial relationship and reduce the computational complexity. Next, non-linearity is introduced through batch normalization BN and the ReLU activation function, and the vanishing gradient is avoided. Finally, the number of channels of the feature map is adjusted through a 1×1 convolution for subsequent information interaction and feature fusion.
[0110] Meanwhile, as Figure 2 shown in the figure, multi-scale feature fusion is performed on the inputs of P2, P3, P4, and P5 respectively, and their outputs are subjected to re-feature extraction and upsampling through the EUCB module, and then multi-level cross-scale fusion is performed multiple times to finally obtain three feature map outputs of different scales.
[0111] As an embodiment, the three feature maps of different scales output by the CCFM module are input into the Decoder decoder, and after being selected by IoU and decoded by the predictor, the final object detection result is obtained, and the coordinate position and size of the object in the image are obtained.
[0112] As an embodiment, three input feature maps are extracted, and a fixed number of object queries are selected from these three inputs. Since in traditional IoU, the top k features are directly selected through the classification score (the confidence level output by the classifier) for subsequent processing, but this method has the following problems:
[0113] The classification score is used to represent the probability that the object belongs to a certain category, while the IoU score represents the degree of overlap between the predicted bounding box and the ground truth bounding box, and the distributions of the two are often inconsistent, resulting in possible misdetection and missed detection by the model, and ultimately leading to a deterioration in the performance of the model.
[0114] Therefore, a new IoU query method is adopted: the IoU score is introduced as a weighted term in the classification loss when selecting object queries, so that it takes into account both the classification score and the IoU score: it can reduce the classification score of low-IoU bounding boxes and increase the classification score of high-IoU bounding boxes. In this way, it can select high-IoU bounding boxes and ultimately improve the detection performance. Its IoU objective function can be expressed as
[0115]
[0116] Where: represents the regression loss of the bounding box, measuring the difference between the predicted bounding box and the ground truth bounding box b;
[0117] represents the classification loss, in which the IoU score is added to adjust the classification scores of different bounding boxes
[0118]
[0119] where IoU i represents the IoU score of the i-th predicted bounding box, and represents the classification probability of the i-th predicted bounding box. In this way, bounding boxes with high IoU will get higher weights to further optimize their classification scores; while bounding boxes with low IoU will be weakened to reduce their impact on the final loss.
[0120] Finally, the above output is passed through multiple Transformer Decoder Layers to finally obtain the prediction result. Each layer includes the following parts:
[0121] 1. Cross-attention: Calculate the correlation between object queries and the global features output by the Encoder. For each query, extract relevant information from the global features.
[0122]
[0123] Among them: Q: query, from the input object queries; K, V: keys and values, from the global features of the Encoder, which can make each query focus on the context information related to the target.
[0124] 2. Self-attention: Calculate the correlation between queries, capture the interaction information between different queries, can enhance the feature representation between queries, and ensure that there is no conflict between different targets.
[0125] These modules will be stacked multiple times, and the features of object queries will be gradually optimized in each layer. And in each layer, queries will continuously extract information from the Encoder features through cross-attention and self-attention, and interact with other queries to gradually enhance the features. At the same time, the output of each layer will be used as the input of the next layer, and finally a set of optimized queries will be output to obtain the final object detection result.
[0126] To better demonstrate the effectiveness of the present invention, the improved RT DETR model of the present invention is compared with other mainstream object detection models (including the original RT DETR-r18, yolov8, yolo11, Deformable DETR) on this dataset. It can be seen that compared with the original RT DETR model, the number of parameters of the improved RT DETR model is reduced from 20.1M to 12.1M, a reduction of 40%. The amount of computation is reduced from 58.3G to 31.2G, a reduction of 47%. At the same time, the detection accuracy (map50) is hardly reduced, and it can well complete the lightweight underwater object detection task. At the same time, compared with other similar object detection models, the model of the present invention has the advantages of lower parameters, less amount of computation, and the same or even higher detection accuracy.
[0127]
[0128]
[0129] Please refer to Figure 9 , Figure 9 which is a module diagram of an improved underwater optical object detection system based on RT-DETR in an embodiment of the present invention, including: the system includes: a training set construction module, a model construction module, a model training module, and an object detection module;
[0130] The training set construction module, the model construction module, the model training module, and the object detection module are sequentially connected in series;
[0131] The training set construction module is used to obtain the DUO underwater dataset and underwater pictures and perform preprocessing to generate a training set;
[0132] The model construction module is used to construct a lightweight underwater object detection model based on the RT-DETR model, the SepNPFCSP_ELAN module, the Hilo attention mechanism, and the method of multi-scale adaptive weighted feature fusion;
[0133] The model training module is used to train the underwater object detection model through the training set; the object detection module is used to obtain the underwater pictures to be classified; and the trained underwater object detection model is used to perform underwater object target detection on the underwater pictures to be classified to obtain the target detection result of the underwater object.
[0134] The above are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, all equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure.
[0135] This invention is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not recorded in the present disclosure. The description and examples are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. An improved underwater optical target detection method based on RT-DETR, characterized in that: The method comprises the following steps: S1: Obtain the DUO underwater dataset and underwater images and preprocess them to generate a training set; S2: A lightweight underwater target detection model is constructed based on the RT-DETR model, SepNPFCSP_ELAN module, Hilo attention mechanism, and multi-scale adaptive weighted feature fusion method; S3: Train the underwater target detection model using the training set; S4: Obtain underwater images to be classified; The trained underwater target detection model is used to perform underwater object detection on the underwater images to be classified, and the target detection results of the underwater objects are obtained.
2. The improved underwater optical target detection method based on RT-DETR as claimed in claim 1, characterized in that: Step S1 includes: The images in the DUO underwater dataset and underwater images are preprocessed, including color balancing, image denoising and data labeling.
3. The improved underwater optical target detection method based on RT-DETR as claimed in claim 1, characterized in that: Step S2 includes: S2-1: Improve the backbone structure of the underwater target detection model by lightweighting the SepNPFCSP_ELAN module; S2-2: SepNPFCSP_ELAN module for improving underwater target detection model; S2-3: Improve the AIFI module of the underwater target detection model through the Hilo attention mechanism to obtain the AIFI-Hilo module; S2-4: The CCFM module is designed by using the improved SepNPFCSP_ELAN module, AIFI-Hilo module and EUCB module, combined with the multi-scale adaptive weighted feature fusion method.
4. The improved underwater optical target detection method based on RT-DETR as claimed in claim 3, characterized in that: Step S2-1 includes: The improved backbone structure includes: convolutional blocks, four SepNPFCSP_ELAN modules and Adown modules; The Adown module is used to downsample the input features; The SepNPFCSP_ELAN module is used to perform multi-scale feature extraction and fusion on the input features.
5. The improved underwater optical target detection method based on RT-DETR as claimed in claim 3, characterized in that: Step S2-2 includes: The improved SepNPFCSP_ELAN module includes: a first DSConv depthwise separable convolution, a split unit, a first PFCSP module, a second DSConv depthwise separable convolution, a third DSConv depthwise separable convolution, a second PFCSP module, a fully connected unit, and a fourth DSConv depthwise separable convolution; The first DSConv depthwise separable convolution connects the split unit; the split unit connects the fully connected unit; The split unit is connected to the first PFCSP module; the first PFCSP module is connected to the second DSConv depthwise separable convolution; the second DSConv depthwise separable convolution is connected to the fully connected unit; The first PFCSP module is connected to the third DSConv depthwise separable convolution; the third DSConv depthwise separable convolution is connected to the fully connected unit; The third DSConv depthwise separable convolution is connected to the second PFCSP module; The second PFCSP module is connected to the second DSConv depthwise separable convolution; The fully connected unit is connected to the fourth DSConv depthwise separable convolution connection.
6. An improved underwater optical target detection method based on RT-DETR as claimed in claim 5, characterized in that: The first PFCSP module and the second PFCSP module have the same structure. The first PFCSP module includes: DSConv deep separable convolution, PFBottleneck module and fully connected layer; the PFBottleneck module includes: PFConv module and DSConv deep separable convolution; The PFConv module includes: convolution block, silu activation unit and sigmod activation unit.
7. The improved underwater optical target detection method based on RT-DETR as claimed in claim 3, characterized in that: Step S2-4 includes: CCFM modules include: EUCB module, SepNPFCSP_ELAN module, AIFI-Hilo module and Fusion module; The input features are upsampled through the EUCB module; the feature maps of different scales are fused into multi-scale features through the Fusion module.
8. An improved underwater optical target detection system based on RT-DETR, used to implement an improved underwater optical target detection method based on RT-DETR as described in claims 1-7, characterized in that: The system comprises: a training set construction module, a model construction module, a model training module and a target detection module; The training set construction module, the model construction module, the model training module and the target detection module are connected in sequence; The training set construction module is used to obtain the DUO underwater data set and underwater pictures and perform preprocessing to generate a training set; The model building module is used to build a lightweight underwater target detection model based on the RT-DETR model, SepNPFCSP_ELAN module, Hilo attention mechanism and multi-scale adaptive weighted feature fusion method; The model training module is used to train the underwater target detection model through a training set; The target detection module is used to obtain underwater pictures to be classified; and to perform underwater object detection on the underwater pictures to be classified using the trained underwater target detection model to obtain underwater object target detection results.