Lightweight underwater small target detection method based on YOLOv11 improvement

By improving the SMDown downsampling module, UDA attention module and wavelet dual-path enhancement mechanism of the YOLOv11 model, the accuracy and real-time problems of underwater target detection in complex environments are solved, and efficient and accurate underwater small target detection is achieved.

CN120833547APending Publication Date: 2025-10-24CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510915381.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Existing underwater target detection technologies face problems such as insufficient detection accuracy, high computational complexity, and poor real-time performance in complex ocean environments. In particular, false detection and missed detection are serious in the detection of small underwater targets.

Method used

An improved YOLOv11 model is adopted to enhance feature extraction and fusion through the SMDown downsampling module, UDA attention module and dual-path wavelet enhancement mechanism, including the coordinated processing of frequency domain analysis and spatial domain enhancement paths, dynamic gating and multi-scale feature fusion, to optimize small target recognition and positioning.

Benefits of technology

It improves the accuracy and real-time performance of underwater target detection, reduces false detections and missed detections, is suitable for underwater equipment with limited computing resources, and achieves efficient detection in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833547A_ABST
    Figure CN120833547A_ABST
Patent Text Reader

Abstract

The invention provides an underwater small target detection method based on YOLOv11 improvement, and belongs to the crossing field of computer vision and marine science research, and the method comprises the steps: carrying out the improvement with YOLOv11 as a reference model, replacing a down-sampling convolution layer in a backbone network with an improved SMDown down-sampling module, so as to enhance the multi-scale feature expression capability, and carrying out the detection of a small target. Fine feature information of the small target is effectively reserved; a UDA attention module is designed to replace a C2PSA module in the YOLOv11 backbone network and is arranged in front of an SPPF module, the module realizes cross-branch information complementation through a bidirectional feature interaction mechanism, and the problem of small target detail loss caused by traditional pooling operation can be effectively solved; a dual-path wavelet enhancement mechanism is introduced into the YOLOv11 feature fusion network, edge details are enhanced for shallow features, frequency domain feature calibration is carried out on deep features, and the feature extraction capacity of the model for small targets is comprehensively improved; and applying the lightweight underwater small target detection model to an actual environment to realize real-time detection of the underwater small target. According to the method, the problems that the traditional method is low in detection precision in a complex underwater environment and is difficult to deploy to edge equipment are solved, and the method has good universality and practicability.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of underwater target detection, and specifically relates to an improved lightweight underwater small target detection method based on YOLOv11. BACKGROUND

[0002] Underwater marine product fishing is a key link of modern fishing. In recent years, the progress of underwater robots and automation technology is driving the evolution of traditional fishing methods towards high efficiency, automation and intelligence. Among them, underwater target detection technology based on machine vision shows the potential to improve fishing accuracy and efficiency. However, this technology still faces serious challenges in practical application, mainly affected by the following three factors:

[0003] Complex underwater environment: the imaging quality of visual sensors is significantly disturbed by the changeable underwater lighting conditions, fluctuation of water turbidity and complex water flow, resulting in decreased target detection accuracy.

[0004] Algorithm robustness is insufficient: in the face of diverse and complex marine product targets in the marine environment, existing visual-based target detection algorithms are prone to false detection or missed detection.

[0005] Real-time performance and deployment performance bottleneck: the real-time processing capability of the target detection algorithm and its efficient deployment capability on underwater devices with limited computing resources still need to be improved.

[0006] The above factors result in high computational complexity and insufficient detection accuracy of existing underwater small target detection in complex marine environments. Therefore, how to accurately, quickly and reliably locate and identify underwater targets from low-quality images remains a major challenge. SUMMARY

[0007] To solve the above problems, the application provides an improved lightweight underwater small target detection method based on YOLOv11, which can achieve more accurate feature extraction in complex underwater environments and limited computing conditions, effectively improve detection accuracy, and reduce false detection and missed detection. The specific steps are as follows:

[0008] S1: Use YOLOv11 as the baseline model, improve the down-sampling module in the YOLOv11 backbone network, and use SMDown down-sampling module instead of the convolution down-sampling layer in the backbone network;

[0009] S2: Replace the C2PSA attention module in the backbone network with the UDA attention module and place it before the SPPF module to improve the accuracy of small target recognition and positioning;

[0010] S3: Introducing two different wavelet double-path enhancement mechanisms to the YOLOv11 network feature fusion part, enhancing high-frequency details at shallow features and optimizing cross-scale semantic fusion at deep features;

[0011] S4: Applying the underwater small target detection model to the actual environment to realize real-time detection of underwater small targets.

[0012] Further, the SMDown downsampling module in S1 uses a double-branch parallel structure. The frequency domain analysis path extracts high-level features of the foreground target through an attention module, and the spatial domain enhancement path focuses on detailed information through deformable convolution. After channel splicing, Conv2d is used for downsampling.

[0013] Further, the SMDown in S1 performs the following operations:

[0014] S11: The input feature map F in is divided into a frequency domain analysis path F in,1 and a spatial domain enhancement path F in,2 according to the channel dimension. in,1 After discrete wavelet transform, the channel attention and spatial attention modules are used to capture long-range dependencies between channels and highlight target dense areas. After inverse wavelet transform, Conv2d convolution is used to further fuse features.

[0015] S12: F in,2 uses deformable convolution to expand the receptive field range and further optimize the extraction efficiency of non-uniformly distributed features. Subsequently, normalization and SiLU activation functions are used to enhance feature distribution stability.

[0016] S13: The output feature maps of the double-branch are spliced and downsampled using Conv2d convolution and SiLU function to obtain the output F out of the downsampling module.

[0017] Further, the UDA module in S2 is mainly based on dynamic gating and multi-scale feature fusion and is a double-branch attention module, which includes the following steps:

[0018] S21: The input feature map to the UDA module is divided into two parts according to the channel dimension, namely the local branch and the global branch, with input feature maps F l and F g . Both branches use dynamic gating units and DynamicReLU-B to adaptively adjust the features, and then the local branch uses Conv2d and DCNv3 convolution layers to enhance the ability to extract features and obtain the output F l,out ; the global branch uses a 7x7 receptive field of dilated convolution and an efficient attention mechanism to filter noise to obtain the output feature map Fg,out .

[0019] S22: The information of the global branch and the local branch is interacted, specifically, F l,out is up-sampled to F g,out is added element by element, and the output F l is obtained; the global branch is first used for global average pooling and then added element by element with F l,out to obtain the global branch output F g ′ of local information injection.

[0020] S23: The global branch and the local branch output are spliced and connected in residual with the original input, and the features are integrated through the FFN.

[0021] Further, the dynamic gating unit dynamically fuses the double-branch features through local-global joint attention, the local attention captures the detail features through deep convolution, and the global attention extracts the context through global pooling, and finally outputs the gating weight G = σ(Conv 3×3 (local_attn,global_attn)).

[0022] Among them, the channel attention in the efficient attention module uses the compression-excitation structure, and the spatial attention uses the lightweight convolution, which reduces the calculation amount while ensuring the detection accuracy.

[0023] Further, the two different wavelet double-path enhancement mechanisms in S3 enhance high-frequency details at shallow features and optimize cross-scale semantic fusion at deep features, and the two wavelet mechanisms are as follows:

[0024] S31: The Wavelet Enhance_1 module is inserted after the channel splicing in the feature fusion network top-down path (FPN), which aims to enhance high-frequency details at shallow features and improve the feature expression ability of edges and small targets;

[0025] The input feature map F is decomposed into low-frequency (LL) and high-frequency (LH, HL, HH) components after discrete wavelet transform, wherein the low-frequency component feature map is F1, and the high-frequency component feature map is F2.

[0026] F1 compresses spatial information through global average pooling to generate a channel-level statistical vector wherein, 1×1 convolution is adopted for cross-channel information fusion, and GELU activation function is adopted, outputting F1′ = GELU(Conv2d(GAP(F1))), and finally a single-layer convolution and bilinear interpolation are used to restore the spatial size, wherein the interpolation ratio is consistent with the down-sampling rate during wavelet decomposition, ensuring that the features are aligned with the high-frequency branch.

[0027] F2 compresses the channel dimension through group convolution (GConv) and outputs F 2,g *W g ; Then it is activated by the GELU function, and the depth-separable convolution is used to enhance the local details to obtain the enhanced high-frequency feature F2′=Depthwise(GELU(F 2,g *W g ))*W 1×1 Among them, W g is the weight of the g-th group, * is the convolution operation, W 1×1 It is the convolution kernel of point-by-point convolution.

[0028] After concatenating F1′ and F2′, the inverse transform is used to reconstruct the features and obtain the feature map F′ after frequency domain enhancement. The GRN channel attention module then enhances the response of key channels and suppresses redundant channels, outputting the enhanced feature F out After being residually connected with the original input feature map F, a single-layer convolution is used for feature fusion.

[0029] S32: The WaveletEnhance_2 module is placed after the C3K2 module in the bottom-up (PAN) path of the feature fusion network. The input feature map F is also divided into two parts F1 and F2 after DWT. F1 also passes through the GRN channel attention to achieve efficient feature selection while maintaining the orthogonality between channels, and obtain the channel-enhanced feature F 1,ca ; F2 compresses the input channel into a single-channel spatial mask through a single-layer 3×3 convolution, and uses the sigmoid activation function to generate weights, outputting the spatial region enhancement feature F 2,sa =σ(Conv 3×3 (F2))⊙F2, where σ represents the sigmoid function and ⊙ refers to element-wise multiplication.

[0030] F 1,ca With F 2,sa The images are spliced ​​according to the channel dimension, and deformable convolution and inverse wavelet transform are used to reconstruct the high-resolution image. The residual connection is also performed with the original input feature map F to obtain features calibrated in the frequency domain at the deep features.

[0031] Compared with the existing technology, the present invention provides an improved lightweight underwater small target detection method based on YOLOv11. The method combines dynamic deformable convolution enhancement, dual attention feature fusion and lightweight design technology, and constructs a complete feature extraction-fusion-enhancement processing flow through the innovative SMDown downsampling module, UDA attention module and dual-path wavelet enhancement mechanism.

[0032] By introducing a downsampling module that collaboratively processes the frequency domain analysis path and the spatial domain enhancement path, rich feature information is retained; the UDA module achieves the optimized integration of local details and global context through the dynamic gating unit and multi-scale feature fusion mechanism; the Wavelet Enhance_1 module enhances high-frequency details at shallow features, significantly improving the feature expression ability of edges and small targets; the Wavelet Enhance_2 module optimizes cross-scale semantic fusion at deep features, achieving precise calibration of frequency domain features.

[0033] Compared with the existing technology, the present invention not only improves the detection accuracy in complex underwater environments, but also provides a reliable technical solution for real-time detection of edge devices such as underwater robots, effectively solving the problems of low target detection accuracy and large computational complexity in complex underwater environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 Flowchart of the method of the present invention.

[0035] Figure 2 This is the structure diagram of the underwater target detection network improved based on the YOLOv11 algorithm.

[0036] Figure 3 This is the SMDown structure diagram in the improved YOLOv11 algorithm.

[0037] Figure 4 This is the structure diagram of the UDA attention module in the improved YOLOv11 algorithm.

[0038] Figure 5 This is the CrossGate and EfficientAttention structure diagram in the UDA attention module.

[0039] Figure 6 This is a structural diagram of the two wavelet dual-path enhancement mechanisms in the improved YOLOv11 algorithm.

[0040] Figure 7 Two sets of pictures for testing the improved YOLOv11 algorithm on the DUO dataset, with real images on the left and predicted results on the right. DETAILED DESCRIPTION

[0041] In order to more clearly illustrate the principles and features of the present invention, the present invention will be further introduced below with reference to the accompanying drawings and implementation examples.

[0042] like Figure 1 As shown, the present invention provides an improved lightweight underwater small target detection method based on YOLOv11, which specifically includes the following steps:

[0043] S1: Taking YOLOv11 as a benchmark model, improving the down-sampling module in the YOLOv11 backbone network, and using SMDown down-sampling module instead of the convolution down-sampling layer in the backbone network;

[0044] S2: Replacing the C2PSA attention module in the backbone network with the UDA attention module and placing it before the SPPF module to improve the accuracy of small target recognition and positioning;

[0045] S3: Introducing two different wavelet double-path enhancement mechanisms to the YOLOv11 network feature fusion part, enhancing high-frequency details at shallow features and optimizing cross-scale semantic fusion at deep features;

[0046] S4: Applying the underwater small target detection model to the actual environment to realize real-time detection of underwater small targets.

[0047] The specific implementation of the above steps is described in detail below.

[0048] The SMDown down-sampling module in S1 adopts a dual-branch parallel structure, realizes efficient feature extraction through the cooperative processing of frequency domain analysis and spatial domain enhancement, and the module structure diagram is as shown in Figure 3 The main steps are as follows:

[0049] S11: The input feature map F in is split into a frequency domain analysis path F in,1 and a spatial domain enhancement path F in,2 , where F in,1 is split into a low-frequency component X l and a high-frequency component X h after discrete wavelet transform, where X l = DWT LL (F in ), X h = Concat(DWT LH (F in ), DWT HL (F in ), DWT HH (F in )). The low-frequency component X l enhances the key channels through a channel attention module to obtain the enhanced output F ca = σ(GAP(X l )·W1+b1)·W2+b2. Where σ is a sigmoid function, and W1 / W2 are learnable weights. The high-frequency component X h generates a mask M spatial = σ(Conv 3×3 (X h)) and the spatial features are reconstructed by IWT after space mask processing, and 3x3 convolution is used for feature fusion.

[0050] S12: Spatial enhancement path F in,2 The geometric deformation features are processed by deformable convolution, and the dynamic constraint of the offset is Δp=tanh(Conv(F in,2 ))·offset_scale, and then the BatchNorm and SiLU function are used to retain the nonlinear features, and the output F 2,out of the spatial domain path is obtained. in,2 .

[0051] S13: After splicing the outputs of the two paths, Conv2d convolution and SiLU activation function are used for down-sampling to obtain the output F out of the entire down-sampling module.

[0052] The UDA attention module in S2 has a structure diagram as shown in Figure 4 The UDA module is a hybrid attention mechanism combining dynamic gating and multi-scale feature fusion, which can realize precise capture of small target features by combining local detail enhancement and global context modeling. In addition, placing the UDA module before the SPPF module can strengthen high-frequency features and suppress irrelevant background noise through the gating mechanism, ensuring that the features input to the SPPF contain more detailed information. The specific steps are as follows.

[0053] S21: The UDA module contains two parallel paths, namely the local branch F l and the global branch F g . The feature maps of the two branches are adjusted by dynamic gating units and DynamicReLU-B adaptive regional weights, and then the local branch passes through 3x3 convolution layers and DCNv3 convolution in turn to enhance the geometric deformation adaptation ability of underwater targets and obtain the output F l,out of the local branch; the global branch captures large-scale features through dilated convolution with an expansion rate of 2 and an efficient attention module to enhance the extraction ability of key semantic information, and outputs the feature map F g,out of the global branch.

[0054] As shown in Figure 5 , in the dynamic gating unit, the local branch extracts local features localAttn=DepthwiseConv 3×3 (F l ) through deep convolution, which reduces the amount of calculation while retaining spatial detail classification; the global branch compresses the spatial dimension through GAP, and then generates channel weights through a fully connected layer, outputting glocalAttn=FC(GAP(Fg )),where GAP is the formula of The two-branch features are concatenated and then passed through a convolution and sigmoid function to generate gating weights G, where G = sigma(Conv 3×3 (local_attn,global_attn)),sigma is the sigmoid function, and the output G e [0,1] 2×H×W , which is divided into G1 (local weight) and G2 (global weight). Then the input features are weighted by channel, and the final output is the sum of the weighted local and global features.

[0055] DynamicReLU-B is a dynamic piecewise linear function whose parameters depend on the input. By dynamically adjusting the slope and intercept of ReLU, it achieves differentiated enhancement of different channel features. It neither increases the depth of the network nor the width of the network, but can effectively increase the model capacity. For input features The calculation formula is: F dy-ReLU = ReLU (aF in +b),

[0056] The efficient attention module is a hybrid attention mechanism that combines channel attention and lightweight spatial attention, aiming to enhance feature expression with lower computational cost. Channel attention compresses the input feature map through global average pooling and generates channel weights through a bottleneck structure, which are then multiplied with the input channel by channel to enhance important feature channels. Lightweight spatial attention simplifies the traditional spatial attention to a 3x3 convolution. First, the input feature map is compressed into a single channel along the channel dimension, and then a 3x3 convolution is used to generate spatial weights, which are multiplied with the channel-weighted features by position.

[0057] S22: F l,out is upsampled to the same size of the global branch by bilinear interpolation and added element by element to the global branch, which supplements the spatial details of the global branch by dynamically adjusting the size of the local features; the global branch F g,out generates a channel-level global context vector through GAP and adjusts it to the same spatial size as the local features. Similarly, element-wise addition is used to inject global semantics into local features. Through the bidirectional feature interaction mechanism, local details and global semantics are optimized cooperatively.

[0058] S23: The outputs of the two branches are concatenated and then connected in residual with the original input feature map to ensure the stability of the training, and the features are integrated through the FFN feedforward network.

[0059] FFN is a lightweight feature transformation module that enhances the model expression ability through nonlinear mapping of channel dimension. The feature is mapped to a high-dimensional space through the first fully connected layer to enhance the separability of the feature, the nonlinearity is introduced by using the GELU activation function, and the second fully connected layer is used to compress to the original dimension to complete the feature fusion. The overall formula is FFN(F in )=W2·GELU(W1·F in ), wherein F in is the input feature, and W1 and W2 are the weight matrices of the fully connected layer.

[0060] The two wavelet double-path enhancement mechanisms in S3 are Wavelet Enhance_1 module and Wavelet Enhance_2 module. Wavelet Enhance_1 module is inserted after channel splicing in the top-down path (FPN) of the feature fusion network, aiming to enhance high-frequency details at shallow features, significantly improving the feature expression ability of edges and small targets; Wavelet Enhance_2 module is inserted after the C3K2 module in the bottom-up (PAN) path of the feature fusion network, optimizing cross-scale semantic fusion at deep features, and through the synergistic effect of GRN channel attention and spatial attention, precise calibration of frequency domain features is achieved.

[0061] The specific implementation of step S3 is the same as the foregoing, and will not be described in detail here. The structural schematic diagram is shown in Figure 6 .

[0062] Further, the underwater small target detection model in S4 is applied to the actual environment to realize real-time detection of underwater small targets. The specific steps are as follows:

[0063] The weight parameters with the best performance in the test set are deployed to the corresponding hardware devices, and the underwater images captured by the camera are uploaded to the detection model for target detection.

[0064] Obviously, the above embodiments are only one embodiment for clearly describing the technical solutions of the present application, and the protection scope of the present application is not limited thereto. All other equivalent alternative solutions that can be thought of by those skilled in the art without creative labor after reading the embodiments of the present application belong to the protection scope of the present application.

Claims

1. A lightweight underwater small target detection method based on YOLOv11 improvement, characterized in that, The method comprises the following steps: S1: taking YOLOv11 as a reference model, improving a down-sampling module in the YOLOv11 backbone network, and using an SMDown down-sampling module to replace a convolution down-sampling layer in the backbone network; S2: adopting an UDA attention module to replace a C2PSA attention module in the backbone network and placing the UDA attention module before an SPPF module, so as to improve the precision of small target identification and positioning; S3: introducing two different wavelet double-path enhancement mechanisms to a network feature fusion part of the YOLOv11, enhancing high-frequency details at a shallow feature, and optimizing cross-scale semantic fusion at a deep feature; S4: applying the underwater small target detection model to an actual environment to realize real-time detection of underwater small targets; The two wavelet double-path enhancement modules in the S3 are a Wavelet Enhance_1 module and a Wavelet Enhance_2 module, forming a synergistic optimization architecture of 'high-frequency enhancement-low-frequency calibration'; S31: the Wavelet Enhance_1 module decomposes the input feature map F into a low-frequency component F1 and a high-frequency component F2 after discrete wavelet transform (DWT); a low-frequency optimization branch generates a channel-level statistical vector after global average pooling (GAP) The enhanced low-frequency features are then fused by 1x1 convolution across channels, and a GELU activation function is used to realize nonlinear transformation and output F1'. A single-layer convolution and bilinear interpolation (Interpolate) are used to restore the spatial size and maintain the smoothness of the low-frequency component. S32: the high-frequency component in the high-frequency enhancement branch is compressed in the channel dimension through group convolution (GConv), and F is output 2,g *W g , W g is the weight of the gth group of group convolution, * is a convolution operation; then, the enhanced high-frequency feature F2' is obtained through GELU function activation and depth separable convolution. S33: splice F1' and F2', reconstruct the feature using inverse transform (IDWT), then enhance the response of key channels and suppress redundant channels through the GRN (global response normalization) channel attention module, obtain the channel-level descriptor gx using the L2 norm, and use simple division to calculate the weight coefficient F of each channel n , multiply the input channel by channel using the normalized weight, and output the enhanced feature F out , F out is connected in residual with the original input F to realize cross-channel feature fusion through Conv2d, and output the optimized feature S34: The original input feature map F of the Wavelet Enhance_2 module is also divided into two parts F1 and F2 after DWT transformation according to the channel dimension, F1 obtains the channel enhanced feature F through GRN channel attention 1,ca ; F2 obtains the output spatial region enhanced feature F 2,sa through the optimized light spatial attention module S35: F 1,ca With F 2,sa After stitching, the high-resolution image is reconstructed by deformable convolution (DCNv3) and inverse wavelet transform (IWT), and is connected with the original input feature map F in residual connection.

2. The improved lightweight underwater small target detection method based on YOLOv11 according to claim 1, characterized in that, The SMDown down-sampling module in the S1 adopts a double-branch parallel structure to perform the following operations: S11: input feature map F in Splitting along channel dimension into frequency domain analysis path F in,1 With spatial domain enhancement path F in,2 ; F in,1 After discrete wavelet transform, pass through channel attention and spatial attention modules in turn. Channel attention uses global average pooling and fully connected layer to calculate channel weight. Spatial attention module uses single-layer 3x3 convolution to generate single-channel spatial mask. Then, reconstruct spatial features through IWT and further fuse features using Conv2d convolution. S12: F in,2 The deformable convolution focuses on the key features of the foreground target, and the batch normalization and SiLU activation function are used to accelerate the model convergence. S13: The double-branch feature map is spliced along the channel dimension, and a Conv2d convolution and a SiLU function are used for down-sampling to output F out .

3. The improved lightweight underwater small target detection method based on YOLOv11 according to claim 1, characterized in that, The S2 double-branch underwater attention module UDA based on dynamic gating and multi-scale feature fusion is split into local branches F in the channel dimension l and global branches F g , both of which can adaptively adjust the feature importance at different positions through a dynamic gating unit. The local branch adaptively enhances features through DynamicReLU-B, and uses 3x3 convolution to maintain a local receptive field, and adaptively adjusts the sampling position through DCNv3 deformable convolution to obtain output F l,out ; The global branch also passes through the dynamic activation function, and then captures a larger range of receptive fields through the dilated convolution with a dilation of 2, and uses an efficient attention mechanism to obtain a feature map F g,out ; F l,out 、F g,out The two branches interact with each other and use bilinear interpolation to convert F l,out Upsample to F g,out Same size, and F g,out Take element-by-element addition to output F l ′; Global branch F g,out Then, feature expansion is performed after global average pooling, and F l,out Add element by element to get the output F g '; F l ′ and F g ′After concatenation, a residual connection is taken with the original input, and then the features are integrated through FFN.

4. The improved lightweight underwater small target detection method based on YOLOv11 according to claim 3, characterized in that, A dynamic gating unit dynamically fuses double-branch features through local-global joint attention and calibrates features of each branch.

5. The improved lightweight underwater small target detection method based on YOLOv11 according to claim 3, characterized in that, DynamicReLU-B uses the same activation function for different channels at the same position.

Citation Information

Cited By

  • Underwater target detection and tracking system and method

    CN121600389A

  • A system and method for underwater target detection and tracking

    CN121600389B

  • Improved underwater target detection method based on joint image enhancement and YOLO26

    CN122368750A

  • Improved underwater target detection method based on joint image enhancement and YOLO26

    CN122368750B