Underwater small target detection method based on YOLOv8
By combining StarNet's StarsBlock module and YOLOv8's C2F module, the CARAFE_Enhanced module and gradient-guided multi-scale convolutional attention mechanism are introduced, which solves the problem of insufficient difficulty in distinguishing between targets and backgrounds and positioning accuracy in underwater small target detection, and realizes high-precision small target detection in complex underwater environments.
Patent Information
- Application Number
- CN202411891003.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art is difficult to effectively distinguish between targets and backgrounds in underwater small target detection, and traditional target detection methods lack positioning accuracy when dealing with underwater small targets, especially in complex backgrounds and low contrast environments.
By combining the star convolution StarsBlock module in the StarNet backbone network with the C2F module of YOLOv8, a new C2F_StarsBlock module is formed to enhance the diversity and robustness of feature extraction. At the same time, the CARAFE_Enhanced module is introduced instead of the traditional sampling method and integrates a gradient-guided multi-scale convolutional attention mechanism (GMSAA) to improve the detection accuracy and robustness of the model in complex underwater environments.
It significantly improves the accuracy and robustness of underwater small target detection, and can effectively identify and locate small targets in complex underwater environments, reduce background interference, and improve detection effect.
Smart Images

Figure CN120088629A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an underwater small target detection method based on YOLOv8, belonging to the technical fields of computer vision, deep learning, target detection, underwater image processing, etc. Background Technique
[0002] Underwater target detection, as one of the important applications in the field of computer vision, is widely used in fields such as marine biology monitoring, environmental protection, fishery resources investigation, military reconnaissance, etc. Different from ground or aerial target detection, underwater target detection faces unique challenges. First of all, due to the changes in lighting conditions, water turbidity, and complex backgrounds in the underwater environment, the contrast between the target and the background is relatively low, and traditional visual detection methods are difficult to effectively distinguish the target and the background. Secondly, factors such as noise, light refraction, and attenuation in underwater images result in poor image quality, further increasing the difficulty of target detection. Especially when facing underwater small targets (such as marine organisms, etc.), due to their small size and unclear features, traditional target detection methods often cannot accurately identify and locate these targets.
[0003] In recent years, target detection methods based on deep learning, especially the YOLO series algorithms, have made significant progress in various computer vision tasks. The YOLO series algorithms have become important methods in the field of target detection due to their excellent real-time detection capabilities and relatively high detection accuracy. However, although YOLO performs excellently in general target detection tasks, its application in underwater small target detection still faces many challenges. The background in underwater images is complex and disorderly, and the YOLO algorithm is easily interfered by the background during the feature extraction process, resulting in a decrease in target recognition accuracy. In addition, underwater small targets usually have small sizes and lack texture information, which makes traditional target detection algorithms such as YOLO insufficient in terms of positioning accuracy and small target detection capabilities.
[0004] To address these problems, researchers have proposed various improvement methods, including convolutional modules that enhance feature extraction capabilities, optimized upsampling methods, and the introduction of attention mechanisms. Although these methods have improved the accuracy of small target detection to a certain extent, in the underwater environment, they still cannot completely solve problems such as background interference, scale changes, and detail loss. Therefore, how to achieve high-precision small target detection in the underwater environment, especially for tiny underwater targets such as marine organisms, remains an important challenge faced by current technologies.
[0005] In response to these challenges, the present invention proposes an improved YOLOv8 model specifically for underwater small target detection. First, the present invention combines the star-shaped convolution StarsBlock module in the StarNet backbone network with the C2F module of YOLOv8 to form a new C2F_StarsBlock module. This module enhances the diversity and robustness of feature extraction and can more effectively capture the detailed information in underwater images, especially for small-sized targets. Second, the present invention proposes a CARAFE_Enhanced module to replace the traditional upsampling method in the Neck part of YOLOv8. By introducing dynamic convolution, grouped convolution, and non-linear feature enhancement strategies, the CARAFE_Enhanced module improves the computational efficiency and adaptability of the model based on the CARAFE module and is particularly suitable for dealing with complex lighting changes and background noise in the underwater environment. Finally, in response to the special requirements of underwater small target detection, the present invention also integrates a gradient-guided multi-scale convolutional attention mechanism (Gradient-guided Multi-Scale Attention Mechanism, GMSAA), which introduces a gradient-guided attention mechanism in the backpropagation stage to help the network more precisely focus on the target area during training, reduce background interference, and improve the detection accuracy of small targets. Summary of the Invention
[0006] In view of this, the purpose of the present invention is to provide an improved underwater small target detection method based on YOLOv8, which can improve the detection accuracy and robustness of the YOLOv8 model for small targets in a complex underwater environment, especially in underwater scenarios with occlusion, low contrast, and complex backgrounds.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] An improved underwater small target detection method based on YOLOv8, which specifically includes the following steps:
[0009] Step 1) Construct a dataset for underwater small target detection in a complex underwater environment and preprocess the images in the dataset;
[0010] Step 2) Input the preprocessed image to be detected into the improved underwater target detection model, extract features from the image through the improved backbone network, and generate a feature map of the image to be detected to capture the multi-level feature information of the target;
[0011] Step 3) Use the above-extracted feature map as the input of the feature fusion module to fuse the feature map information of different scales;
[0012] Step 4) Input the fused high-quality feature map into the detection head. The detection head processes the feature map, locates the position of the target, and determines the category, thus completing the target detection task;
[0013] Furthermore, step 1) specifically includes the following steps:
[0014] Step 11) Obtain diverse underwater scene image data to construct a dataset, which contains various typical underwater scenes, such as uneven illumination, complex background, target occlusion, color deviation under different water depth conditions, etc.;
[0015] Step 12) Preprocess the images in the dataset, including image normalization, size adjustment, color restoration, and data augmentation operations, such as adding random noise, enhancing contrast, gamma transformation, etc., to improve the robustness and generalization ability of the model in complex underwater environments.
[0016] Furthermore, step 2) specifically includes the following steps:
[0017] Step 21) Feed the preprocessed dataset into the improved backbone feature extraction network YOLOv8 for feature extraction. In the backbone network, use the improved C2F_StarsBlock module to replace the standard C2F module. C2F_StarsBlock combines star-shaped convolution, which can improve the small target detection ability while maintaining a low computational cost and enhancing the feature extraction ability for targets in complex underwater scenes;
[0018] Step 22) Output multi-scale feature maps in the backbone network. These feature maps contain local detail information of the target and global semantic information of the underwater scene, serving as the input for subsequent modules.
[0019] Furthermore, step 3) specifically includes the following steps:
[0020] Step 31) Replace the upsampling module in the neck network with the CARAFE_Enhanced module. CARAFE_Enhanced is a lightweight upsampling module that combines dynamic convolution operations and grouped convolution, which can dynamically adjust according to features at different levels, achieving efficient resampling and precise expression of feature information in underwater scenes;
[0021] Step 32) Add the Gradient-guided Multi-Scale Attention Mechanism (GMSAA) to the neck network. This module combines the channel attention mechanism (CA) and the spatial attention mechanism (SA). By performing feature weighting on feature maps of different scales, it enhances the model's detection ability for targets of different sizes. The additionally introduced Gradient-guided Attention Mechanism dynamically adjusts the attention weights using gradient information during the backpropagation process, enabling the model to focus on the target regions that have the greatest impact on the loss function during training, thereby reducing interference from irrelevant background information.
[0022] Furthermore, step 4) specifically includes the following steps:
[0023] Step 41) Input the fused feature map into the object detection head. The detection head predicts the bounding box coordinates and classes of each target through regression and classification operations, and performs non-maximum suppression (NMS) to remove redundant boxes, ensuring the accuracy and uniqueness of the final output result.
[0024] Compared with existing underwater object detection methods, the advantages of the present invention are as follows:
[0025] (1) Existing methods usually focus on improving the object detection performance in a specific environment. However, the present invention proposes a small object detection method suitable for complex underwater environments by making multi-level improvements to YOLOv8, enabling the model to perform excellently in complex scenarios such as insufficient light, severe occlusion, and noise interference, significantly improving the comprehensiveness and robustness of underwater small object detection.
[0026] (2) Based on the YOLOv8 framework, the present invention optimizes the backbone network by introducing the C2F_StarsBlock module and innovatively improves the CARAFE module in the neck network to form the CARAFE_Enhanced module. On the basis of these two improvements, the number of parameters and computational complexity of the model remain basically unchanged, significantly enhancing the detection effect of the model in underwater small object detection tasks. This design not only meets the efficiency requirements of lightweight models but also ensures excellent performance in complex underwater scenarios, successfully balancing the model efficiency and resource consumption.
[0027] (3) Through the combination of the C2F_StarsBlock module and the CARAFE_Enhanced module, the present invention realizes multi-level feature extraction and expression from small-scale to large-scale targets. In addition, the MSAA module introduces a gradient-guided attention mechanism, which can focus on the target area in a complex underwater environment, effectively suppress background interference, further enhance the target features, and improve the occlusion problem and detail capture ability of small target detection.
[0028] (4) The present invention not only conducts in-depth optimization for underwater target detection, but also provides new ideas for other small target detection tasks. By means of the gradient-guided attention mechanism and multi-module fusion, the semantic information between different feature layers is fully integrated, forming a general framework suitable for small target detection in complex environments. This method has achieved remarkable results in the field of underwater detection, and also provides a reference for the model improvement of similar detection tasks. Brief Description of the Drawings
[0029] Figure 1 Overall flowchart of the invention
[0030] Figure 2 Overall network framework of the improved YOLOv8
[0031] Figure 3 Feature extraction module C2F_StarsBlock
[0032] Figure 4 Lightweight upsampling module CARAFE_Enhanced
[0033] Figure 5 Module GMSAA introducing gradient-guided multi-scale convolutional attention mechanism Detailed Description of the Invention
[0034] To make the technical solution of the present invention clearer, the present invention will be further described below with reference to the accompanying drawings.
[0035] An underwater small target detection method based on YOLOv8 includes the following steps:
[0036] Step 1) Construct an underwater small target detection dataset
[0037] Step 2) Construct a backbone feature extraction network
[0038] Step 3) Construct the CARAFE_Enhanced module
[0039] Step 4) Construct the module GMSAA based on the gradient-guided multi-scale convolutional attention mechanism
[0040] Step 5) Construct an underwater small target detection network based on YOLOv8
[0041] Step 6) Train the object detection network
[0042] Step 7) Test the object detection network
[0043] Furthermore, the specific content of Step 1) includes the following steps:
[0044] Step 11) Utilize the publicly available underwater image dataset (URPC dataset). This dataset contains small objects in various underwater scenarios, such as sea cucumbers, shellfish, etc., with different degrees of occlusion, lighting variations, and background complexities. Use the LabelImg image annotation software to annotate the objects in each image, clarify the object categories and their true positions, and generate corresponding annotation files. Combine the original images with the annotation files to form a dataset.
[0045] Step 12) Preprocess the dataset, including operations such as image size adjustment and image normalization, to ensure that the dataset can meet the input requirements of the YOLOv8 model and improve the robustness and generalization ability of the model.
[0046] Step 13) Divide the processed dataset into a training set, a validation set, and a test set. The dataset is divided in a ratio of 8:1:1. The training set is used to train the model, the validation set is used to verify the performance of the model, and the test set is used to evaluate the detection ability and accuracy of the model in practical applications.
[0047] Furthermore, the specific content of Step 2) includes the following steps:
[0048] Step 21) As Figure 3As shown, the C2F_StarsBlock module is improved by replacing the Bottleneck structure in the original C2F module with the StarsBlock module. First, the input feature map passes through a convolutional layer with a size of 1×1 to complete the compression or expansion of the number of channels, adjusting the number of channels of the feature map to half of the output channels. Then, the feature map is divided into two parts, one part is directly retained, and the other part enters the subsequent StarsBlock module for feature extraction. A new type of computational unit is adopted inside the StarsBlock module, combining depthwise separable convolution (DW-Conv), fully connected layer (FC), and element-wise multiplication. The features first extract local spatial information through DW-Conv convolution, then pass through the BN normalization and activation function layer (ReLU) respectively, and then use the FC layer to expand the feature dimension, thereby enhancing the non-linear expression ability. Subsequently, the features generated by the two FC layers are fused through pointwise multiplication to highlight the contribution of the key feature regions and finally restored to the original number of channels. In the entire StarsBlock module, multi-scale convolutional kernels are used to enhance the fusion ability of local and global features, while lightweight operations are used to ensure efficiency. After the processing of the StarsBlock module, it is concatenated (Concat) with the unprocessed feature map of the previous part, and the number of channels of the feature map is restored through a 1×1 convolutional operation, and finally the output features are obtained. The entire C2F_StarsBlock module uses the above improvements to significantly enhance the diversity and expressiveness of feature extraction, especially suitable for complex underwater target detection scenarios.
[0049] Step 22) The improved YOLOv8 feature extraction backbone network significantly enhances the feature extraction ability by introducing the C2F_StarsBlock module, while retaining the characteristics of lightweight and high efficiency in the architecture design. First, the input image passes through a standard convolutional layer (Conv) to extract low-level features and adjust the number of channels simultaneously. Next, it goes through the CSP structure (CrossStagePartialNetwork) to reduce computational complexity and increase feature reuse. In the CSP structure, each residual block (ResidualBlock) replaces the original C2F module with the C2F_StarsBlock, thereby improving the fineness of feature extraction and the non-linear expression ability. The feature extraction backbone network consists of multiple stages, which process feature maps of different scales respectively. In the shallower stages, the network focuses on extracting texture information and edge features, which is particularly crucial for small object detection; in the middle and deeper stages, the network gradually focuses on more abstract semantic information and global features. These feature layers are downsampled through convolutional or pooling operations with a stride of 2. After each downsampling, the resolution of the feature map is halved while the number of channels increases, thus maintaining the expressive ability of the network. In the final stage, the feature map passes through the Spatial Pyramid Pooling (SPPF) module to aggregate multi-scale feature information. The SPPF module uses pooling kernels of various different sizes to extract local and global information simultaneously and integrates this information through concatenation operations. Finally, the backbone network outputs multi-scale feature maps, which are passed to the subsequent detection head part to complete the specific object detection task. Through the above improvements, the YOLOv8 backbone network achieves efficient and accurate feature extraction in underwater object detection tasks, and has significant advantages for small objects and complex scenes.
[0050] Furthermore, the specific content of step 3) includes the following steps:
[0051] Step 31) As Figure 4 shown, the CARAFE_Enhanced module is an improved version of the CARAFE module, which has been enhanced in the following three aspects for the special requirements of underwater small object detection: dynamic convolution operation, efficiency optimization, and non-linear feature recombination. Through these improvements, the CARAFE_Enhanced module achieves more accurate upsampling and enhancement of features while taking into account computational efficiency. The overall process of the module is as follows: First, the input feature map X comp = R C×H×W passes through the channel compression layer to generate the intermediate feature map X comp .
[0052] X comp = ReLU(BN(Conv(X, C mid , kernel_size = 1, groups = G)))
[0053] Among them, RELU is an activation function used to introduce non-linearity, BN is a normalization operation, and C mid represents the number of channels after compression, G is the number of groups of grouped convolutions used to improve efficiency.
[0054] Subsequently, X comp generates dynamic weight W for feature recombination through an encoding convolutional layer. Specifically, first, a convolutional operation is performed on the feature map, and the output channel number of the convolution is jointly determined by the upsampling scale and the size of the recombination convolution kernel. Then, the convolution result is normalized through a batch normalization layer (BatchNormalization) to ensure a more stable data distribution. Finally, the GELU activation function is used to perform a non-linear transformation on the normalized result to generate dynamic weight W. The dynamic weight W is decoded into two-dimensional recombination weights and rearranged to the high-resolution space through the PixelShuffle operation. Finally, the Softmax normalization is used to ensure the effectiveness of the weights to obtain W normolized .
[0055] The input feature map is initially upsampled through nearest neighbor interpolation:
[0056] X up = NearestUpsample(X, scale_factor = S)
[0057] The upsampled feature map is extended to the recombination convolution kernel dimension through the Unfold operation to facilitate subsequent feature fusion:
[0058] X unfold = Unfold(X up , kernel_size = K up , dilation = S, padding = (K up / 2)·S)
[0059] X represents the input feature map of the module, with dimensions (B, C, H, W), where B is the batch size, C is the number of channels, and H and W are the height and width of the feature map respectively. X up represents the feature map after initial upsampling through nearest neighbor interpolation. X unfold represents the feature map extended through the Unfold operation, which is for adapting to subsequent feature recombination operations. K up represents the size of the recombination convolution kernel. S represents the multiple of the dilation convolution used to expand the receptive field so that each convolution kernel can capture a wider range of features.
[0060] Finally, in the feature fusion stage, the dynamic weight Wnormolized and the expanded feature X unfold They are fused in an element-wise multiplication manner. The fused features are further enhanced in global representation through dynamic convolution. Finally, the module enhances the non-linear representation ability of the features through a Tanh non-linear activation layer.
[0061] The CARAFE_Enhanced module has been significantly optimized in dynamic weight generation, non-linear feature fusion, and efficient computation, and is particularly suitable for the refined feature modeling and upsampling tasks of underwater small targets. Its lightweight design ensures computational efficiency, while significantly improving the accuracy of feature reconstruction through dynamic convolution and non-linear activation.
[0062] Furthermore, the specific content of step 4) includes the following steps:
[0063] Step 41) As Figure 5 shown, the Gradient-guided Multi-scale Convolutional Attention Mechanism module (GMSAA) is a highly flexible and powerful feature extraction module, aiming to combine multi-scale feature extraction, channel and spatial attention mechanisms, and gradient-based dynamic adjustment capabilities to enhance the model's perception ability of key target regions in complex scenarios. Its construction process includes the following core parts: multi-scale feature extraction, channel attention mechanism, spatial attention mechanism, and gradient-guided dynamic adjustment. The GMSAA module first processes the input feature map through a dimensionality reduction convolutional layer to reduce the number of input channels and thus the computational complexity, and projects the features into a low-dimensional space to ensure efficient subsequent operations. The dimensionality-reduced feature map is fed into convolutional layers with three different convolutional kernel sizes (3×3, 5×5, and 7×7). These convolutional layers have different receptive fields for capturing multi-scale information.
[0064] The channel attention mechanism generates two global feature representations through global average pooling and max pooling, respectively capturing the saliency information between channels. These two representations are processed through a shared two-layer convolutional network and then added together to obtain the channel weights. Finally, the weights are restricted to the range [0,1] through the Sigmoid activation function for adjusting the channel response of the input feature map. The formula is described as follows:
[0065] F avg = W 2 ·ReLU(W 1 ·AvgPool(X)), F max = W 2 ·ReLU(W 1 ·MaxPool(X))
[0066] Channel attention weights:
[0067] M channel= σ(F avg + F max )
[0068] Adjusted feature:
[0069] X channel = X·M channel
[0070] X is the input feature map. AvgPool(X) is the average pooling operation. MaxPool(X) is the max pooling operation. W 1 is the weight of the first fully connected layer. W 2 is the weight of the second fully connected layer. F avg is the output of the average pooling branch. F max is the output of the max pooling branch. σ is the Sigmoid function, aiming to map the sum of F avg and F max to the interval [0, 1]. M channel is the channel attention weight. X channel is the feature map after channel enhancement.
[0071] The spatial attention mechanism generates two single-channel feature maps through average pooling and max pooling of the feature map in the spatial dimension. After these feature maps are concatenated in the channel dimension, a 7×7 convolutional layer is used to generate the spatial attention weight. The weight map is restricted to the range [0, 1] through the Sigmoid activation function and is used to adjust the spatial distribution of the input feature map. The formula is described as follows:
[0072]
[0073] After concatenation, the generated spatial weight is M spatial , and finally the adjusted feature is as follows:
[0074] X spatial = X·M spatial
[0075] X is the input feature map. Mean = (X, dim = 1) is the average feature map, which can capture the global statistical information of the input feature in the spatial dimension. Max = (X, dim = 1) is the max feature map, which can capture the strongest activation value of the input feature in the spatial dimension.
[0076] Subsequently, the gradient guidance mechanism extracts the saliency information of the key regions by calculating the gradients of the loss function with respect to the multi-scale features. First, the L2 norm of the gradients is calculated to obtain the gradient magnitude map; then the gradient magnitude map is normalized to generate the gradient attention map. This attention map is used to dynamically adjust the weights of the feature map, enabling the model to pay more attention to the regions that have a greater impact on the loss function. The formula is described as follows:
[0077] And the adjusted feature is X grad = X ms ·G
[0078] X ms is the input multi-scale feature map. is about X ms 's loss gradient. G is the gradient weight. X grad is the feature map after gradient guidance.
[0079] Finally, the features after combining the multi-scale features and the three attention mechanisms are restored to the original number of channels through the upsampling convolutional layer. To better fuse the input and output features, the initial features are also processed by dimensionality reduction and then added to the final features for fusion, ensuring the preservation of the global information of the input features while enhancing the expression ability of the key regions. The finally output features have both global context information and can highlight the key target regions, thus adapting to complex scenarios and small target detection tasks.
[0080] Furthermore, the specific content of step 5) includes the following steps:
[0081] Step 51) Using the backbone feature extraction network constructed by the above steps, and the feature fusion network including the lightweight enhanced upsampling module and the gradient-guided multi-scale convolutional attention module and combining with the detection head, a YOLOv8-based underwater small target detection network is formed.
[0082] Furthermore, the specific content of step 6) includes the following steps:
[0083] Step 61) Input the dataset into the target detection network model, configure the corresponding training environment, set the corresponding training parameters, and carry out the training task of the model;
[0084] Step 62) This invention is carried out on the operating system version of Ubuntu20.04, with GPU being NVIDIA RTX 4060 (16GB), CPU being Intel(R) Core(TM) i7-13650HX, the host memory being 16GB, the programming language being Python3.8, and training is carried out based on the deep learning framework PyTorch1.8.1;
[0085] In the configuration of training parameters, the optimizer uses SGD, the learning rate is set to 0.01, the weight decay rate is set to 0.0005, the number of training epochs is set to 300, and the batch size is set to 4.
[0086] Furthermore, the specific content of step 7) includes the following steps:
[0087] Step 71) Load the trained weight file into the underwater target detection network, randomly select pictures from the test dataset for detection, and return the labeled pictures with the position and category information of the target to be measured.
[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for detecting underwater small targets based on YOLOv8, which is characterized by comprising the following steps: Step 1: Collect and preprocess the datasets related to small target detection in underwater environments: Download the open source datasets for small target detection in complex underwater environments from public websites, and divide them into training set, validation set, and test set in the ratio of 8:1:
1. There are four categories of detection targets, namely sea cucumber, sea urchin, starfish, and scallop. Use annotation software to mark the target images of the initial dataset, mark various targets in the image with specific boxes, generate corresponding annotation files, and use them as training datasets together with the original image files. Step 2: Construct a backbone feature extraction network: The C2F_StarsBlock module is based on the original C2F module, replacing the Bottleneck structure with the StarsBlock module. The StarsBlock module uses deep separable convolution (DW-Conv) to extract local spatial information, and combines BN normalization, activation function (ReLU), fully connected layer (FC) and point-by-point multiplication to fuse key feature areas to enhance nonlinear expression capabilities and multi-scale feature fusion. The backbone network consists of multiple stages, each of which is combined with the C2F_StarsBlock module through the CSP structure to achieve efficient feature reuse and fine feature extraction. The shallow network focuses on texture information and edge features, and the intermediate and deep networks gradually focus on abstract semantics and global features. Convolution with a stride of 2 is used for downsampling to maintain a balance between resolution and number of channels. Finally, the spatial pyramid pooling (SPPF) module aggregates multi-scale information to provide high-quality feature maps for the detection head. This improved design achieves accurate detection of small targets in complex underwater scenes. Step 3, build the CARAFE_Enhanced module: The core process of the module consists of three parts. The first is the generation of channel compression and dynamic weights. The input feature map first passes through the channel compression layer to generate intermediate features, and generates dynamic weights through encoded convolution. After the weights are processed by batch normalization and GELU activation function, they are decoded into two-dimensional reorganized weights, assigned to the high-resolution space through the PixelShuffle operation, and normalized by Softmax to ensure effectiveness. Then there is upsampling and feature expansion. The feature map is first preliminarily upsampled through nearest neighbor interpolation, and then expanded to a format that adapts to the reorganized convolution kernel through the Unfold operation to facilitate detailed feature fusion. Finally, there is feature fusion and nonlinear enhancement. The dynamic weights and the expanded features are multiplied element by element to complete feature fusion, and the global expression ability is improved through dynamic convolution. Finally, the nonlinear expression ability is enhanced through Tanh activation, so as to reconstruct the features more accurately. Step 4: Construct a gradient-guided multi-scale convolutional attention mechanism module GMSAA: The input feature map first reduces the number of channels through a dimensionality reduction convolution layer to reduce computational complexity, and at the same time projects the features into a low-dimensional space to improve processing efficiency. The reduced-dimensional feature map is fed into three convolutional layers, which use different convolution kernel sizes (such as 3×3, 5×5, and 7×7) to capture multi-scale features with different receptive fields. This process enhances the network's ability to express local details and global semantic information. Next, the channel attention mechanism extracts two global feature representations by performing global average pooling and maximum pooling operations on the feature map to capture salient information between channels. The two features are processed by a shared convolutional network to generate channel weights, which are mapped to the [0,1] interval through an activation function to dynamically adjust the channel response of the feature map, thereby highlighting key channel information. Subsequently, the spatial attention mechanism performs average pooling and maximum pooling on the feature map in the spatial dimension to generate two single-channel feature maps, which are concatenated and then convolved to generate spatial weights to optimize the spatial distribution of the feature map again. On this basis, the module introduces a gradient-guided dynamic adjustment mechanism. By calculating the gradient of the loss function for multi-scale features, a gradient amplitude map is generated and normalized into a gradient attention map. The attention map dynamically adjusts the weights of the feature map so that the model pays more attention to the areas that have a significant impact on the target detection task. Finally, the features adjusted by combining multi-scale feature extraction, channel and spatial attention mechanisms, and gradient-guided are restored to the original number of channels through upsampling convolution. To further improve the feature fusion effect, the initial input features are processed by dimensionality reduction and then added and fused with the final features, which not only retains the global information of the input features, but also significantly enhances the expression ability of key areas. Step 5, construct an underwater small target detection network based on YOLOv8: use the backbone feature extraction network constructed in the above steps, and the feature fusion network including a lightweight enhanced upsampling module and a gradient-guided multi-scale convolutional attention module and combine it with the detection head to form the underwater small target detection network based on YOLOv8 of the present invention. Step 6: Train the target detection network model: Use the training set to fully train the YOLOv8-based underwater small target detection network constructed in step 5 to obtain the trained network model weights. Step 7: Test the target detection network model: load the trained weight file into the target detection network, randomly select images from the test set for detection, return labeled images with the location and category information of the target to be tested, and obtain the detection accuracy and detection speed of each category in the small target data set in a complex underwater environment, as well as the overall average detection accuracy.
2. The underwater small target detection method based on YOLOv8 according to claim 1, characterized in that: The backbone feature extraction network in step 2 achieves efficient feature reuse and fine feature extraction of underwater small targets by introducing the C2F_StarsBlock module. First, the C2F_StarsBlock module combines the star convolution characteristics of the StarNet network with the structural advantages of the original C2F module of YOLOv8. Through the design of star convolution, the module introduces a wider receptive field on the basis of traditional convolution operations, which can effectively capture the local details and global structural information of underwater targets. This feature is particularly suitable for complex underwater scenes with blurred target edges or severe background interference. Secondly, the C2F_StarsBlock module adopts a multi-branch feature extraction architecture, which divides the feature channels into multiple groups. Each group of features is processed by star convolution and conventional convolution respectively, and the complementarity of deep and shallow features is used in the fusion stage to enhance the expression ability of small targets. This design not only retains the advantages of multi-scale feature expression, but also avoids the waste of computing resources, achieving a balance between efficiency and precision. Finally, C2F_StarsBlock further reduces the computational complexity and improves the real-time performance of the network by integrating group convolution and depth-wise separable convolution, laying the foundation for the deployment of the model in underwater small target detection tasks.
3. The underwater small target detection method based on YOLOv8 according to claim 1, characterized in that: The CARAFE_Enhanced module in step 3 achieves accurate upsampling and enhancement of features through multiple innovations, providing excellent support for underwater small target detection. First, the CARAFE_Enhanced module optimizes the traditional CARAFE in three aspects based on the characteristics of underwater small targets: dynamic convolution operation, efficiency improvement, and nonlinear feature enhancement. Through these improvements, the module can reconstruct the edge and detail features of the target in a complex underwater environment more carefully while maintaining high computational efficiency. Specifically, the input feature map first passes through a channel compression layer to reduce the computational complexity and generate an intermediate feature map. Next, a coded convolution layer generates dynamic weights. After nonlinear activation and batch normalization, the weights form dynamic convolution kernel weights that adapt to the specific feature distribution, thereby realizing adaptive modeling of the input features. In the upsampling stage, the input feature map is initially expanded using nearest neighbor interpolation, and then the features are flattened to the reorganized convolution kernel dimension through the Unfold operation. The dynamic weights and the expanded features are fused element by element multiplication, and the global feature expression capability is further improved through dynamic convolution. Finally, the Tanh nonlinear activation layer further improves the nonlinear expression of features. Compared with the traditional upsampling method, the CARAFE_Enhanced module can accurately restore the spatial distribution and detailed features of the target in complex scenes while retaining the global context information by integrating dynamic weight generation and efficient feature fusion strategy. This design greatly improves the detection accuracy and model performance in underwater small target detection tasks.
4. The underwater small target detection method based on YOLOv8 according to claim 1, characterized in that: The GMSAA module in step 4 effectively improves the model's target perception ability in complex underwater environments by combining multi-scale feature extraction, channel and spatial attention mechanisms, and gradient-guided dynamic adjustment strategies. The gradient-guided mechanism is the key innovation of the GMSAA module. The gradient-guided mechanism calculates the gradient of the loss function for multi-scale features to extract regional information in the feature map that has an important impact on the detection task. Specifically, it first calculates the gradient amplitude map, extracts the significant features of the key area, and then normalizes the gradient amplitude to generate a gradient attention map. This attention map can dynamically adjust the weight of the feature map, so that the model pays more attention to the area related to the target detection task and reduces the interference of complex backgrounds. Compared with the traditional attention mechanism, the innovation of this mechanism is that it dynamically adjusts the sensitive areas of model loss based on gradient information directly, thereby achieving accurate optimization of features. This method is particularly suitable for scenes with complex backgrounds and low significance of small targets in underwater environments. It can effectively enhance the feature expression of the target area while suppressing the interference of irrelevant background on the model. Finally, the features generated by the gradient-guided mechanism are fused with the outputs of multi-scale convolution, channel and spatial attention mechanisms to form a more discriminative feature expression, providing strong technical support for underwater small target detection.
Citation Information
Cited By
Image segmentation method based on space attention mechanism
CN120374992A
Soldering tin defect detection method and device, electronic equipment and storage medium
CN120411109A
Improved YOLO11-based water hyacinth target rapid detection method
CN120635395A
Industrial camera operation optimization method and system based on complex scene
CN120894662A
A method and system for optimizing the operation of industrial cameras in complex scenarios
CN120894662B