Marine ranch benthos detection method based on improved YOLO11n

The improved YOLO11n model addresses the challenges of detecting small, densely distributed marine organisms by incorporating a modified neck network and occlusion-aware attention module, resulting in enhanced detection precision and robustness.

CN120318571AActive Publication Date: 2025-07-15NORTHEAST DIANLI UNIVERSITY
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510387526.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-15
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

The existing single-stage object detection algorithm is difficult to effectively extract feature information and distinguish occlusion targets in marine ranch benthic biological detection, resulting in poor detection accuracy.

Method used

Improved YOLO11n model by introducing C3k2-MSEE module and FPSConv module into the backbone network, adding P2 detection layer to the neck network and using SBA module, and adding Multi-SEAM module to the head network to improve feature extraction and occlusion perception capabilities.

Benefits of technology

It improves the accuracy and efficiency of benthic biological detection in marine ranch, enhances the detection ability of small targets and occlusion targets, and reduces model complexity and parameter quantity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318571A_ABST
    Figure CN120318571A_ABST
Patent Text Reader

Abstract

The invention discloses a marine ranch benthos detection method based on improved YOLO11n. The marine ranch benthos detection method specifically comprises the following steps: (1) establishing a marine ranch benthos image data set; (2) adding annotation information to the images in the data set; (3) constructing a marine ranch benthic organism detection model based on the improved YOLO11n; (4) training the model by adopting the training set and the verification set, and storing the trained model; and (5) testing the model by adopting the test set to meet the requirements of precision and generalization, thereby obtaining the final benthic organism detection model. Compared with the prior art, the improved YOLO11n-based marine ranch benthos detection method disclosed by the invention has the advantages that the detection capability of sheltered targets can be enhanced, the distinction degree between benthos and a background is enhanced, and the detection capability of small-target benthos is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to a method for detecting benthic organisms in a marine ranch by improving YOLO11n. Background Art

[0002] Marine benthic organisms are widely distributed, with a wide variety of species and extremely rich resources. They are an indispensable and important part of the marine ecosystem. These organisms, such as starfish, sea urchins, sea cucumbers, scallops, etc., not only play an important role in maintaining the marine ecological balance, but also become the targets of fishing and aquaculture due to their significant economic value. In the past, the aquaculture of various benthic organisms relied heavily on manpower, mainly including manual diving fishing and underwater monitoring. These tasks not only have high labor costs but also pose high safety risks. In recent years, with the development of technologies such as underwater robots, the automated fishing of underwater benthic organisms has been realized. And underwater biological target detection is an important part of the automated fishing of underwater benthic organisms and a key technology for realizing the automated fishing of underwater robots.

[0003] Traditional underwater target detection algorithms usually refer to methods based on feature extraction and classifiers. First, preprocessing such as noise reduction, enhancement, and segmentation is performed on the target image, then the target is extracted from features such as the shape, color, and texture of the target, and finally the target is recognized through a classifier. With the rise and development of artificial intelligence technology, the introduction of target detection and classification methods based on deep learning technology has made it possible to accurately and quickly detect targets in complex environments.

[0004] Among the target detection algorithms based on deep learning, single-stage target detection algorithms have become a research and application hotspot due to their high efficiency and real-time performance. These algorithms directly predict the category and location of the target on the input image in an end-to-end manner, omitting the step of generating candidate regions in the traditional two-stage method. Typical representatives include the YOLO (You Only Look Once) series and SSD (Single Shot MultiBox Detector). The core idea of YOLO is to divide the image into grid cells, and each cell directly predicts multiple bounding boxes and the corresponding category probabilities, and global inference can be completed through a single forward propagation. This design greatly improves the detection speed. And SSD balances speed and accuracy by fusing and predicting multi-scale feature maps and using features at different levels to enhance the ability to capture small targets.

[0005] However, in the task of detecting benthic organisms in marine pastures, single-stage algorithms still face challenges. First, benthic organisms are usually small in size and densely distributed, occupying only a very small number of pixels in the image, making it difficult for the model to extract effective feature information. Second, benthic organisms are often occluded due to population aggregation or environmental interference, resulting in blurred target boundaries and making it difficult for the model to distinguish individual organisms, leading to poor detection results for benthic organisms in marine pastures. Summary of the Invention

[0006] In view of the problems existing in the prior art, the present invention provides a method for detecting benthic organisms in marine pastures based on improved YOLO11n. By improving the YOLO11n object detection model and using the improved neck network structure, shared convolution module, and detection head with a multi-branch occlusion-aware attention module, the problems of low detection accuracy caused by small benthic organisms in marine pastures and mutual occlusion between benthic organisms in marine pastures are solved.

[0007] The technical solution provided by the present invention includes the following steps:

[0008] Step 1: Obtain images of benthic organisms in marine pastures to form a first dataset;

[0009] Step 2: Add annotation information to the images in the first dataset to form a second dataset, and divide the second dataset into a training set, a validation set, and a test set;

[0010] Step 3: Build a detection model for benthic organisms in marine pastures based on improved YOLO11n, where the model includes a backbone network, a neck network, and a head network;

[0011] Step 4: Use the training set and the validation set to train the detection model for benthic organisms in marine pastures based on improved YOLO11n;

[0012] Step 5: Use the test set of benthic organisms in marine pastures to test the optimal model obtained by training, and obtain the final detection model for benthic organisms in marine pastures;

[0013] Further, in step 1, the images of benthic organisms in the marine pasture in the first dataset can be taken by an underwater camera, taken by an autonomous underwater vehicle, or collected from the network;

[0014] Preferably, in step 2, the LableImg annotation tool can be used to add annotation information to road defects; the training set, validation set, and test set can be divided according to a ratio of 7:2:1;

[0015] Further, step 3 specifically includes steps 3.1 to 3.3:

[0016] Step 3.1: In the backbone network, replace the Bottleneck module inside C3k2 with the newly designed MSEE (MutilScale Edge Enhance) module to form the new feature extraction module C3k2-MSEE, and replace the original SPPF module in the backbone network of the YOLO11n model with the parameter-sharing convolution module FPSConv to form a new backbone structure;

[0017] Furthermore, the C3k2-MSEE module has two cases, namely C3k = False and C3k = True. Among the above 4 C3k2-MSEE modules, the first two are C3k2-MSEE modules in the case of C3k = False, and the last two are C3k2-MSEE modules in the case of C3k = True; the C3k2-MSEE module in the case of C3k = False is composed of a 1×1 Conv module, a Split module, 2 MSEE modules, a Concat module, and a 1×1 Conv module connected in sequence; the C3k2-MSEE module in the case of C3k = True is composed of a 1×1 Conv module, a Split module, 2 C3k-MSEE N = 2 modules, a Concat module, and a 1×1 Conv module connected in sequence;

[0018] Furthermore, the MSEE module divides the input into five branches. The first four branches all first go through an AdaptiveAvg Pool with different-sized pooling windows, and then go through two convolution modules. The first is a 1×1 convolution module for channel compression, and the second is a 3×3 convolution module for local feature extraction. Then, it goes through an upsampling module and an Edge Enhance module. Through the above four parallel branches of different scales, the model can obtain different information from different scales, improving the model's feature extraction ability. The fifth branch only has a 3×3 convolution module for retaining the spatial information of the picture; among them, the Edge Enhance module is composed of an Avg Pool layer, an edge calculation module, a convolution module, and a fusion module. Among them, the 3×3 average pooling can retain low-frequency information on a larger scale, and the edge calculation module obtains the edge information of the image by comparing the differences before and after average pooling. The above operations can enable the model to extract more detailed features of the model;

[0019] Furthermore, the FPSConv module is composed of a 1×1 Conv module, three 3×3 Conv modules, a Concat module, and a 1×1 Conv module connected in sequence. Among them, the outputs of the first 1×1 Conv module, the first 3×3 Conv module, and the second 3×3 Conv module are cross-connected to the Concat module;

[0020] The role of the first 1×1 Conv module passed through in the FPSConv module is to adjust the number of channels of the module. Then, the three 3×3 Conv modules passed through are parameter-sharing convolution modules with different dilation rates, which are used to extract features of different scales. This module captures local details through low dilation rates and global context through high dilation rates. Using a convolution module with shared parameters can greatly reduce the number of parameters and improve the model efficiency;

[0021] Step 3.2: In the neck network, first, a new P2 small object detection layer is added on the basis of the original three detection layers to improve the model's detection ability for small objects. Secondly, the SBA (SpecificBlockAttention) attention module is introduced to replace the Concat module for feature fusion, forming a new neck network;

[0022] Furthermore, the SBA module is used to fuse the boundary information of the underlying features and the semantic information of the high-level features to obtain a finer-grained object contour and re-calibrate the position of the object. The fusion method in the SBA module uses the Re-calibration attention unit (RAU) module. This module adaptively extracts the mutual representation of the two inputs (T1, T2) before fusion, and the shallow and deep information is transmitted to the two RAU modules in different ways to make up for the missing spatial boundary information in the deep features and the missing semantic information in the shallow features. Finally, the outputs of the two RUA modules are connected and then input into a 1×1 Conv module. The improved aggregation strategy realizes the robust combination of different features and refines the rough features. The processing process of the RUA module in the model is as follows:

[0023] T′1 = W θ (T1)(1)

[0024]

[0025] Among them, T1 and T2 are input features, and two linear mappings and the Sigmoid function W θ 、 Applied to the input features to reduce the channel dimension to 32, obtaining feature maps T′1 and T′2. ⊙ is point-wise multiplication. The reverse operation is achieved by subtracting feature T′1, optimizing the inaccurate and rough estimates into an accurate and complete prediction map. A convolutional operation with a kernel size of 1×1 is used as the linear mapping process. Therefore, the process of SBA is shown in Equation (4):

[0026] Y = C 3×3 (Concat(RAU(X a , X b ), RAU(X b , X a ))) (4)

[0027] Among them, C 3×3 (·) is a convolution with a kernel size of 3×3 with batch normalization and ReLU activation layers. Contains the deep features of the image. Contains the rich shallow features of the image. Concat(·) is a concatenation operation along the channel dimension. Is the output of the SBA module;

[0028] Step 3.3: In the head network, add the multi-branch occlusion-aware attention module Multi-SEAM (Multi-Branch Separated and Enhancement Attention Module) to the detection head of the original YOLO11n model to form a new detection head;

[0029] Furthermore, the head network is composed of 4 Multi-SEAM Head detection heads. The input of each detection head corresponds to the 4 outputs of the above-mentioned neck network. Each detection head consists of 4 Conv modules, 2 DWConv modules, 2 Multi-SEAM modules, 2 Conv2d modules, 1 CIoU module, and 1 CLSLoss module. Each detection head is divided into two branches. The first branch passes through 2 3×3 Conv modules, 1 Multi-SEAM module, 1 1×1 Conv2d module, and 1 CIoU module in sequence. The second branch passes through 1 3×3 DWConv module, 1 1×1 Conv module, 1 Multi-SEAM module, 1 3×3 DWConv module, 1 1×1 Conv module, 1 1×1 Conv2d module, and 1 CLSLoss module in sequence;

[0030] Furthermore, the core of the Multi-SEAM module consists of multiple components. First, the input features are input into the Channel and Spatial Mixing Module (CSMM). The CSMM module extracts deeper features through depthwise convolution and pointwise convolution. Meanwhile, an activation function (GELU) and batch normalization (Batch Norm) are introduced to enhance the learning ability and accelerate convergence. Depthwise convolution can significantly reduce the number of parameters through per-channel separable operations and effectively distinguish the importance of each channel. However, this operation will ignore the information correlation between channels. To make up for this deficiency, after convolution on each channel, the module combines the outputs of different-depth convolutions through a pointwise convolution layer to restore the correlation between channels. Second, all-channel information is further fused through a two-layer fully connected network to strengthen the connection between channels and make the feature representation more complete and accurate. Finally, the Multi-SEAM module performs exponential normalization on the output of the fully connected layer, expanding the range of output values from [0, 1] to [1, e]. This exponential normalization process provides the model with stronger robustness, making it more stable in the face of position errors and further improving the detection accuracy of the model;

[0031] Furthermore, step 4 specifically includes steps 4.1 to 4.4:

[0032] Step 4.1: Set the training parameters of the marine ranch benthic organism detection model based on the improved YOLO11n. The model training parameters specifically include: learning rate, momentum, weight decay, optimizer, number of iteration rounds, batch size, and label smoothing coefficient;

[0033] Step 4.2: Input the training set and validation set images and their corresponding labels into the marine ranch benthic organism detection model based on the improved YOLO11n. Use the backpropagation algorithm to calculate the gradient of the loss function with respect to the model parameters, and adjust the model parameters by minimizing the loss function to gradually approach the optimal solution;

[0034] Step 4.3: Use the optimizer to update the model parameters in the direction of gradient descent; until the loss functions of the training set and validation set no longer decrease, and at the same time, evaluation metrics such as accuracy P, recall R, and mAP no longer improve;

[0035] Step 4.4: Save the trained model parameters as the optimal model;

[0036] Furthermore, step 5 specifically includes steps 5.1 to 5.3:

[0037] Step 5.1: Input the test set into the improved optimal model described in Step 5;

[0038] Step 5.2: Calculate the model performance metrics: The performance metrics specifically include accuracy P, recall R, mAP@0.5, mAP@0.5:0.95, and model size. The specific calculation formulas are as follows:

[0039]

[0040] Among them, P is the accuracy, R is the recall, mAP is the mean average precision of all classes, AP is the average precision, m is the total number of benthic organism categories in the marine ranch, TP represents the number of positive samples correctly identified as positive samples, FP represents the number of negative samples misidentified as positive samples, and FN represents the number of positive samples misidentified as negative samples;

[0041] mAP@0.5 represents the average precision when the IoU threshold is fixed at 0.5, and mAP@0.5:0.95 represents calculating the average precision every 0.05 from the IoU threshold of 0.5 to 0.95, and then taking the average of these average precisions;

[0042] Step 5.3: When the performance metrics meet the accuracy requirements, obtain the final marine ranch benthic organism detection model based on the improved YOLO11n.

[0043] Compared with the prior art, the beneficial effects of the present invention are:

[0044] (1) In the backbone network, first, design a new feature extraction module, the C3k2-MSEE module, to replace the C3k2 module in the original model, enabling the model to extract more edge feature information of the marine ranch benthic organism images, effectively improving the detection accuracy of the marine ranch benthic organism detection model; second, use the FPSConv module to replace the SPPF module in the original model, which can not only enable the model to extract finer-grained features but also significantly reduce the training parameters, effectively reducing redundancy and improving the detection efficiency of the marine ranch benthic organism detection model;

[0045] (2) In the neck network, design a new neck network structure, add a P2 detection layer to the original neck network to increase the detection ability of small targets, and introduce the SBA module to replace the Concat module for feature fusion in the new neck network structure, improving the feature sensitivity of the marine ranch benthic organism detection model to the key features of benthic organisms and the robustness of detection;

[0046] (3) In the head network, a new detection head, Multi-SEAMHead, is designed. In this detection head, the Multi-SEAM module is used to improve the detection ability of the benthic organism detection model in the marine ranch for occluded organisms among benthic organisms, and at the same time, it also reduces the impact of complex marine backgrounds on benthic organism detection. Description of the Drawings

[0047] Figure 1 This is the flowchart of the method for detecting benthic organisms in a marine ranch based on the improved YOLO11n of the present invention;

[0048] Figure 2 This is the schematic diagram of the model structure of the method for detecting benthic organisms in a marine ranch based on the improved YOLO11n of the present invention;

[0049] Figure 3 This is the schematic diagram of the C3k2-MSEE module structure;

[0050] Figure 4 This is the schematic diagram of the FPSConv module structure;

[0051] Figure 5 This is the schematic diagram of the SBA module structure;

[0052] Figure 6 This is the schematic diagram of the head network structure

[0053] Figure 7 This is the schematic diagram of the Multi-SEAM module structure; Detailed Embodiment

[0054] In order to make the technical solution, structural features, achieved objectives and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in combination with specific embodiments and with reference to the accompanying drawings. It should be noted that the specific embodiments described herein are only used to more clearly explain the present invention and are not used to limit the present invention.

[0055] Figure 1 This is the flowchart of a method for detecting benthic organisms in a marine ranch based on the improved YOLO11n disclosed by the present invention, and its implementation process is as follows:

[0056] Step 1: Obtain images of benthic organisms in the marine ranch to form a first data set; the images of benthic organisms in the first data set can be taken by an underwater camera, taken by an autonomous underwater vehicle, or collected from the network;

[0057] In this embodiment, in order to better evaluate the detection effect of an underwater benthic organism detection method based on improved YOLO11n disclosed by the present invention, the public dataset DUO (Detecting Underwater Objects) underwater organism dataset is adopted.

[0058] Step 2: Add annotation information to the images in the first dataset using the lableme annotation tool.

[0059] In this example, since there is already annotation information in the adopted public dataset DUO, this step is skipped. The label files in json format in the first dataset described in this embodiment are converted into txt format label files required by YOLO11n to form a second dataset. The DUO dataset contains 7,782 precisely labeled images, among which 6,671 are used for training and 1,111 are used for testing, and the 1,111 images used for testing are also used as the validation set.

[0060] Step 3: Construct an underwater benthic organism detection model based on improved YOLO11n. The model includes a backbone network, a neck network, and a head network. The structure of the improved YOLO11n model is as Figure 2 shown. The construction of the model further includes steps 3.1 to 3.3:

[0061] Step 3.1: In the backbone network, replace the Bottleneck module inside C3k2 with a newly designed MSEE module to form a new feature extraction module C3k2-MSEE, and replace the original SPPF module in the backbone network of the YOLO11n model with a parameter-sharing convolution module FPSConv to form a new backbone structure;

[0062] Furthermore, the structure diagram of the C3k2-MSEE module is as Figure 3 shown. The C3k2-MSEE module has two cases, namely C3k = False and C3k = True. Among the above 4 C3k2-MSEE modules, the first two are C3k2-MSEE modules in the case of C3k = False, and the last two are C3k2-MSEE modules in the case of C3k = True; the C3k2-MSEE module in the case of C3k = False is composed of a 1×1 Conv module, a Split module, 2 MSEE modules, a Concat module, and a 1×1 Conv module connected in sequence; the C3k2-MSEE module in the case of C3k = True is composed of a 1×1 Conv module, a Split module, 2 C3k-MSEE N = 2 modules, a Concat module, and a 1×1 Conv module connected in sequence;

[0063] Furthermore, the structural diagram of the MSEE module is as follows Figure 3 shown. The input is divided into five branches. The first four branches all first go through AdaptiveAvg Pool with different-sized pooling windows, and then go through two convolutional modules. The first one is a 1×1 convolutional module for channel compression, and the second one is a 3×3 convolutional module for local feature extraction. Then, it goes through an upsampling module and an Edge Enhance module. Through these four parallel branches with different scales, the model can obtain different information from different scales, improving the model's feature extraction ability. The fifth branch only has a 3×3 convolutional module for retaining the spatial information of the picture. Among them, the Edge Enhance module is composed of an Avg Pool layer, an edge calculation module, a convolutional module, and a fusion module. The 3×3 Avg Pool can retain low-frequency information on a larger scale, and the edge calculation module obtains the edge information of the image by comparing the differences before and after Avg Pool. The above operations can enable the model to extract more detailed features of the model;

[0064] Furthermore, the structure of the FPSConv module is as follows Figure 4 shown. The FPSConv module is sequentially connected by a 1×1 Conv module, three 3×3 Conv modules, a Concat module, and a 1×1 Conv module. Among them, the outputs of the first 1×1 Conv module, the first 3×3 Conv module, and the second 3×3 Conv module are skip-connected to the Concat module;

[0065] The role of the first 1×1 Conv module passed through in the FPSConv module is to adjust the number of channels of the module. Then, the three 3×3 Conv modules passed through are parameter-sharing convolutional modules with different dilation rates, used to extract features of different scales. This module captures local details through low dilation rates and global context through high dilation rates. Using a convolutional module with shared parameters can greatly reduce the number of parameters and improve the model efficiency;

[0066] Step 3.2: In the neck network, first, a new P2 small object detection layer is added on the basis of the original three detection layers to improve the model's detection ability for small objects. Secondly, the SBA (Specific BlockAttention) attention module is introduced to replace the Concat module for feature fusion, forming a new neck network;

[0067] Furthermore, the structure of the SBA module is as follows Figure 5As shown, the SBA module is used to fuse the boundary information of low-level features and the semantic information of high-level features to obtain a finer-grained object contour and re-calibrate the object's position. The fusion method in the SBA module uses a re-calibration attention unit (RAU) module, which adaptively extracts the mutual representation of two inputs (T1, T2) before fusion, as Figure 5 shown, the shallow and deep information is transmitted to the two RAU modules in different ways to make up for the missing spatial boundary information of high-level semantic features and the missing semantic new information of low-level features. Finally, the outputs of the two RUA modules are connected and input into a 1×1 Conv module. The improved aggregation strategy realizes the robust combination of different features and refines the rough features, Figure 5 The processing process of the RUA module in

[0068] is as follows: T′1 = W θ (T1) (1)

[0069]

[0070] where T1 and T2 are input features. Two linear mappings and the Sigmoid function W θ 、 are applied to the input features to reduce the channel dimension to 32, obtaining the feature maps T′1 and T′2. ⊙ is point-wise multiplication. is the reverse operation achieved by subtracting the feature T′1, optimizing the inaccurate and rough estimation into an accurate and complete prediction map. A convolution operation with a kernel size of 1×1 is used as the linear mapping process. Therefore, the process of SBA is as shown in formula (4):

[0071] Y = C 3×3 (Concat(RAU(X a , X b ), RAU(X b , X a ))) (4)

[0072] where C 3×3 (·) is a convolution with a kernel size of 3×3 with batch normalization and ReLU activation layers. contains the deep features of the image. contains the rich shallow features of the image. Concat(·) is a concatenation operation along the channel dimension. is the output of the SBA module;

[0073] Step 3.3: In the head network, add the multi-branch occlusion-aware attention module Multi-SEAM (Multi-Branch Separated and Enhancement Attention Module) to the detection head of the original YOLO11n model to form a new head network;

[0074] Furthermore, the new head network is as Figure 6 shown, and it is composed of 4 Multi-SEAMHead detection heads. The input of each detection head corresponds to the 4 outputs of the above-mentioned neck network. Each detection head consists of 4 Conv modules, 2 DWConv modules, 2 Multi-SEAM modules, 2 Conv2d modules, 1 CIoU module and 1 CLS Loss module. Each detection head is divided into two branches. The first branch passes through 2 3×3 Conv modules, 1 Multi-SEAM module, 1 1×1 Conv2d module and 1 CIoU module in sequence. The second branch passes through 1 3×3 DWConv module, 1 1×1 Conv module, 1 Multi-SEAM module, 1 3×3 DWConv module, 1 1×1 Conv module, 1 1×1 Conv2d module and 1 CLS Loss module in sequence;

[0075] Furthermore, the structure of the Multi-SEAM module is as Figure 7 shown. The core of the Multi-SEAM module consists of multiple components. First, input the input features into the channel and spatial mixing module (CSMM). The CSMM module extracts deeper features through depthwise convolution and pointwise convolution. At the same time, an activation function (GELU) and batch normalization (Batch Norm) are introduced to enhance the learning ability and accelerate convergence. Depthwise convolution can significantly reduce the number of parameters through per-channel separated operations and effectively distinguish the importance of each channel. However, this operation will ignore the information correlation between channels. To make up for this deficiency, after convolving each channel, the module combines the outputs of different depth convolutions through a pointwise convolution layer to restore the correlation between channels. Secondly, further fuse the information of all channels through a two-layer fully connected network to strengthen the connection between channels and make the feature representation more complete and accurate. Finally, the Multi-SEAM module performs exponential normalization on the output of the fully connected layer, expanding the range of the output value from [0,1] to [1,e]. This exponential normalization process provides the model with stronger robustness, making it more stable in the face of position errors, thereby further improving the detection accuracy of the model;

[0076] Step 4: Input the training set and the validation set into the marine ranch benthic organism detection model based on the improved YOLO11n described in Step 3 for training, which specifically includes Steps 4.1 to 4.4:

[0077] Step 4.1: Set the training parameters of the marine ranch benthic organism detection model based on the improved YOLO11n. The model training parameters include: learning rate, momentum, weight decay, optimizer, number of epochs, batch size, and label smoothing coefficient.

[0078] In this embodiment, the optimizer is SGD, the initial learning rate lr0 is 0.01, the momentum is 0.937, the weight decay is 0.0005, the batch size is 16, the number of epochs Epoch is 300, and the label smoothing coefficient label_smoothing is 0.1.

[0079] Step 4.2: Input the training set and validation set images and their corresponding labels into the marine ranch benthic organism detection model based on the improved YOLO11n, and use the backpropagation algorithm to calculate the gradient of the loss function with respect to the model parameters. The backpropagation algorithm is an effective method for calculating gradients. It uses the chain rule to calculate the gradient of each parameter with respect to the loss function. Specifically, backpropagation propagates the loss function backward from the output layer and calculates the gradient of each parameter layer by layer. In this process, the gradient of each parameter represents the rate of change of the loss function with respect to that parameter, that is, how the loss function changes as the parameter changes. Adjust the model parameters by minimizing the loss function to gradually approach the optimal solution.

[0080] Step 4.3: After calculating the gradients of the model parameters, use the optimizer to update these parameters. The optimizer updates the parameters in the opposite direction of the gradient according to the gradient information of the parameters. Parameters with larger gradients will be updated with larger steps, while parameters with smaller gradients will be updated with smaller steps. By continuously iteratively updating the model parameters, the value of the loss function can be gradually reduced. By minimizing the loss function, adjust the values of the model parameters to gradually approach the optimal solution, that is, the parameter values at which the loss function reaches the minimum value, so that the difference between the prediction results of the model and the true values is minimized, and at the same time, evaluation metrics such as mAP, recall rate R, and accuracy P no longer improve.

[0081] Step 4.4: Save the trained model parameters as the optimal model.

[0082] Step 5: Use the test set to test the optimal model described in Step 4, evaluate the test results of the test set, and if the accuracy requirement is met, the final benthic organism detection model based on the improved YOLO11n is obtained. Specifically, Step 5 further includes Steps 5.1 to 5.3:

[0083] Step 5.1: Input the test set into the optimal model described in Step 4;

[0084] Step 5.2: Calculate the model performance metrics: The performance metrics specifically include accuracy P, recall R, mAP@0.5, mAP@0.5:0.95, and model size. The specific calculation formulas are as follows:

[0085]

[0086] Among them, P is the accuracy, R is the recall, mAP is the mean average precision of all classes, AP is the average precision, m is the total number of benthic organism categories in the marine ranch, TP represents the number of positive samples correctly identified as positive samples, FP represents the number of negative samples misidentified as positive samples, and FN represents the number of positive samples misidentified as negative samples;

[0087] mAP@0.5 represents the average precision when the IoU threshold is fixed at 0.5, and mAP@0.5:0.95 represents calculating the average precision every 0.05 between the IoU thresholds from 0.5 to 0.95, and then taking the average of these average precisions;

[0088] Step 5.3: When the performance metrics meet the accuracy requirements, obtain the final benthic organism detection model based on the improved YOLO11n.

[0089] In this embodiment, in order to verify the effect of the improved model disclosed in the present invention, the present invention uses the YOLOv8n model, YOLOv10n model, YOLO11n model, YOLO11s model, and the detection model disclosed in this patent to conduct tests on the DUO dataset. Ours in the table is the model disclosed in the present invention. The evaluation index data is shown in Table 1;

[0090] Table 1 Comparative experiment results

[0091]

[0092] As can be seen from Table 1, compared with the original YOLO11n model, the benthic organism detection model disclosed in the present invention has an improvement of 1.7% in terms of accuracy P, an improvement of 2.4% in terms of recall rate R, an improvement of 2.3% in terms of mAP@0.5, and an improvement of 3.7% in terms of mAP@0.5:0.95. At the same time, while the indicators are comparable to those of the YOLO11s model, the model size is 9.7 MB smaller than that of the YOLO11s model, which is convenient for subsequent deployment on underwater edge devices.

[0093] The above description is only one embodiment of the present invention and does not limit the patent scope of the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An improved YOLO11n-based method for detecting benthic organisms in marine ranch, characterized in that, Specifically, it includes the following steps: Step 1: Obtain benthic organism images of the marine ranch to form a first dataset; in the first dataset, the benthic organism images of the marine ranch can be taken by an underwater camera, taken by an autonomous underwater vehicle, or collected from the network; Step 2: Add annotation information to the images in the first dataset to form a second dataset, and divide the second dataset into a training set, a validation set, and a test set; Step 3: Build a benthic organism detection model for the marine ranch based on the improved YOLO11n. The model includes a backbone network, a neck network, and a head network. The construction of the model further includes steps 3.1 to 3.3: Step 3.1: The backbone network is composed of 2 consecutive 3×3 Conv modules, a C3k2-MSEE module, a 3×3 Conv module, a C3k2-MSEE module, a 3×3 Conv module, a C3k2-MSEE module, a 3×3 Conv module, a C3k2-MSEE module, an FPSConv module, and a C2PSA module, thus constituting a new backbone network structure; The C3k2-MSEE module has two cases, namely C3k = False and C3k = True. For the 4 C3k2-MSEE modules, the first two are C3k2-MSEE modules in the case of C3k = False, and the last two are C3k2-MSEE modules in the case of C3k = True; the C3k2-MSEE module in the case of C3k = False is composed of 1 consecutive 1×1 Conv module, 1 Split module, 2 MSEE modules, 1 Concat module, and 1 consecutive 1×1 Conv module; the C3k2-MSEE module in the case of C3k = True is composed of 1 consecutive 1×1 Conv module, 1 Split module, 2 C3k-MSEEN = 2 modules, 1 Concat module, and 1 consecutive 1×1 Conv module; The C3k-MSEE N = 2 module is composed of 3 Conv modules, 2 MSEE modules, and 1 Concat module. After the input passes through 1 1×1 Conv module, the output is divided into two branches. The first branch passes through 2 MSEE modules and 1 Concat module; the second branch inputs the output of the 1×1 Conv module into 1 3×3 Conv module and then inputs it into the Concat module of the first branch, and then inputs the output of the Concat module into 1 1×1 Conv module; The MSEE module is composed of five branches. Each of the first four branches sequentially passes through one AdaptiveAvgPoul module, one 1×1 Conv module, one 3×3 Conv module, one upsample module, and one Edge Enhance module; the fifth branch only has one 3×3 Conv module; finally, the outputs of the five branches are input into one Concat module and then pass through one 1×1 Conv module; The Edge Enhance module is sequentially composed of one AvgPoul module, one Edge Computing module, one 1×1 Conv module, and one FusionAddition module; the output of the upsample module in the MSEE module is respectively cross-connected and input into the Edge Computing module and the FusionAddition module; The improved backbone network outputs four different scales of feature information through the first three C3k2-MSEE modules and the C2PSA module respectively; Step 3.2: The component modules of the neck network include four Conv modules, six SBA modules, and six C3k2 modules. The neck network is divided into four branches. The input of the first branch is the output of the first C3k2-MSEE module in the backbone network, which sequentially passes through one 1×1 Conv module, one SBA module, and one C3k2 module, and then the output of the C3k2 module is input into the SBA module of the second branch; the input of the second branch is the output of the second C3k2-MSEE module of the backbone network, which sequentially passes through one 1×1 Conv module, one SBA module, and one C3k2 module, and then the output of the C3k2 module is input into the SBA module of the third branch; the input of the third branch is the third C3k2-MSEE module of the backbone network, which sequentially passes through one 1×1 Conv module, one SBA module, and one C3k2 module, and then the output of the C3k2 module is used as the first output of the neck network while continuing to be input into one SBA module and then passing through one C3k2 module. The output of this C3k2 module is used as the second output of the neck network while continuing to be input into one SBA module and then passing through one C3k2 module. The output of this C3k2 module is input into the SBA module of the fourth layer; the input of the fourth branch is the output of the C2PSA module in the backbone network, which sequentially passes through one 1×1 Conv module, one SBA module, and one C3k2 module, and then the output of the C3k2 module is used as the fourth output of the neck network, thus constituting a new neck network structure; Step 3.3: The head network consists of 4 Multi-SEAMHead detection heads. The input of each detection head corresponds to the 4 outputs of the above-mentioned neck network. Each detection head is composed of 4 Conv modules, 2 DWConv modules, 2 Multi-SEAM modules, 2 Conv2d modules, 1 CIoU module and 1 CLSLoss module; Each detection head is divided into two branches. The first branch passes through two 3×3 Conv modules, 1 Multi-SEAM module, 1 1×1 Conv2d module and 1 CIoU module in sequence; the second branch passes through 1 3×3 DWConv module, 1 1×1 Conv module, 1 Multi-SEAM module, 1 3×3 DWConv module, 1 1×1 Conv module, 1 1×1 Conv2d module and 1 CLSLoss module in sequence; Step 4: Use the training set and the validation set to train the benthic organism detection model based on the improved YOLO11n, and save the trained model as the optimal model, which further includes steps 4.1 to 4.4: Step 4.1: Set the training parameters of the benthic organism detection model based on the improved YOLO11n. The model training parameters include: learning rate, momentum, weight decay, optimizer, number of iterations, batch size and label smoothing coefficient; Step 4.2: Input the training set and validation set images and their corresponding labels into the benthic organism detection model based on the improved YOLO11n. Use the backpropagation algorithm to calculate the gradient of the loss function with respect to the model parameters, and adjust the model parameters by minimizing the loss function to gradually approach the optimal solution; Step 4.3: Use the optimizer SGD to update the model parameters, so that the model parameters are updated in the direction of gradient descent until the loss functions of the training set and the validation set no longer decrease, and at the same time the evaluation metrics mAP, recall rate R, and accuracy P no longer increase, then stop training to avoid overfitting of the model; Step 4.4: Save the trained model parameters. The model at this time is the optimal model; Step 5: Use the test set to test the optimal model, evaluate the test results of the test set, and if the accuracy requirement is met, the final benthic organism detection model based on the improved YOLO11n is obtained.

2. The method for detecting benthic organisms in a marine ranch based on improved YOLO11n according to claim 1, wherein In step 2, the training set, validation set and test set are divided in the ratio of 7:2:

1.

3. A method for detecting benthic organisms in a marine ranch based on improved YOLO11n according to claim 1, characterized in that, Step 5 further includes steps 5.1 to 5.3: Step 5.1: Input the test set into the optimal model described in step 5; Step 5.2: Calculate the model performance metrics: The performance metrics specifically include accuracy P, recall rate R, mAP@0.5, mAP@0.5:0.95, and model size. The specific calculation formulas are as follows: Among them, P is the accuracy rate, R is the recall rate, mAP is the mean average precision of all categories, AP is the average precision, m is the total number of categories of benthic organisms in the marine ranch, TP represents the number of positive samples correctly identified as positive samples, FP represents the number of negative samples misidentified as positive samples, and FN represents the number of positive samples misidentified as negative samples; mAP@0.5 represents the mean average precision when the IoU threshold is fixed at 0.5, and mAP@0.5:0.95 means that the mean average precision is calculated every 0.05 between the IoU thresholds from 0.5 to 0.95, and then the average of these mean average precisions is taken; Step 5.3: When the performance indicators meet the accuracy requirements, obtain the final marine ranch benthic organism detection model based on the improved YOLO11n.

Citation Information

Patent Citations

  • Marine organism intelligent detection method based on deep learning

    CN114782982A

  • Weld joint X-ray image defect detection and identification method based on deep neural network

    CN116630263A

  • Steel defect detection method based on SBA cross-scale feature fusion

    CN119180785A

  • Road well lid disease detection method based on edge enhanced feature aggregation

    CN119693927A

  • Improving geo-registration using machine-learning based object identification

    US20240020968A1