Water surface algae target detection method based on deep learning

By building an improved YOLOv8 object detection model, adding CBAM attention module, deformable convolution module and CARAFE upsampling module, and improving the loss function, the problem of insufficient accuracy and reliability of surface algae target detection in complex environments is solved, and efficient and accurate algae target detection is achieved.

CN120125966AInactive Publication Date: 2025-06-10ZHEJIANG UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510195299.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing water surface algae target detection methods have problems with insufficient detection accuracy and reliability in complex environments, especially when facing complex environmental factors such as water surface waves, ripples, light changes, reflections and ripple.

Method used

Using deep learning-based water algae object detection method, an improved YOLOv8 object detection model is built, a CBAM attention module, a deformable convolution module and a CARAFE upsampling module are added, and the loss function is improved to WIoUv3 to improve detection accuracy and efficiency.

Benefits of technology

It significantly improves the accuracy and efficiency of algae target detection, can accurately identify algae targets in complex water surface environments, reduce missed and misdetection conditions, improve overall detection efficiency, and enhance the generalization ability and anti-overfit ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125966A_ABST
    Figure CN120125966A_ABST
Patent Text Reader

Abstract

The invention discloses a water surface algae target detection method based on deep learning, and the method specifically comprises the following steps: 1, collecting an image data sample, generating a data set, and carrying out the data preprocessing and data enhancement; 2, constructing a YOLOv8 target detection model, and adding a CBAM attention module, a deformable convolution module and a CARAFE up-sampling module; improving a loss function of YOLOv8 target detection on the basis of distribution characteristics of algae targets on the water surface; 3, inputting the data set after data preprocessing and data enhancement into a YOLOv8 target detection model for model training; and 4, inputting image data acquired in real time, and carrying out model detection on algae through the trained YOLOv8 target detection model. Interference of complex environmental factors can be overcome, and the water surface algae detection precision and reliability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The field to which the present invention belongs is the field of target detection technology, and specifically relates to a method for detecting water surface algae targets based on deep learning. Background Art

[0002] With the frequent occurrence of water disasters such as algal blooms and red tides, their causes and prevention methods have attracted people's attention. The treatment mainly includes two aspects: the first is to control wastewater discharge, reduce the concentration of nitrogen and phosphorus elements, and treat the eutrophication problem of water bodies to avoid excessive proliferation of microalgae; the second is to clean up algal pollution manually. However, the traditional manual cleaning method relies on professional personnel to drive ships for fixed-point and regular operations, which not only has a large workload, low efficiency, high labor costs, but also has potential safety hazards in operation. Aquatic algae cleaning robots have gradually come into the public eye. Because they can adapt to diverse inland waters, have higher work efficiency, better safety, and lower costs, they have become an important development direction for cleaning algal pollution in inland waters in the future. Compared with the development of intelligent technology for driverless cars, the development of aquatic robots in the field of target perception lags behind, especially the research on aquatic robots for algae cleaning is still in its infancy. The target perception task of algae cleaning aquatic robots is mainly to identify algae floating on the water surface. The Chinese invention patent with the publication number CN115601559A provides a construction method of a lightweight algae target detection algorithm Algae-YOLO based on YOLOv5, including: using ShuffleNetV2 to replace the original backbone network of YOLOv5; adding an ECA attention mechanism to ShuffleNetV2 to construct a ShuffleNetV2-ECA backbone network; in the neck structure of YOLOv5, designing a lightweight structure combined with Ghost convolution. The present invention uses ShuffleNetV2 to replace the original backbone network, greatly reducing the model parameters and calculation costs, and adding an ECA attention mechanism to ShuffleNetV2 to greatly improve the detection accuracy without adding calculation costs. However, compared with YOLOv8, YOLOv5 has problems such as lower accuracy, slightly weaker detection capabilities for small objects and complex scenes, and due to the interference of complex environmental factors such as water surface waves, ripples, light changes, reflections, and ripples, the existing methods still have deficiencies in the accuracy and reliability of target perception. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for detecting water surface algae targets based on deep learning. The present invention can overcome the interference of complex environmental factors and improve the detection accuracy and reliability of water surface algae.

[0004] The technical solution of the present invention: A method for detecting water surface algae targets based on deep learning specifically includes the following steps:

[0005] Step 1: Collect image data samples to generate a dataset, and perform data preprocessing and data augmentation;

[0006] Step 2: Build a YOLOv8 object detection model, add a CBAM attention module to strengthen the attention distribution of YOLOv8 object detection and improve the ability to extract algae features; introduce a deformable convolution module into the YOLOv8 object detection model to enhance the adaptability of the convolutional neural network to irregular and small algae targets and the ability to extract algae features; introduce a CARAFE upsampling module into the YOLOv8 object detection model to improve the resolution of the feature map and enhance the detection effect of small algae targets; based on the distribution characteristics of water surface algae targets, improve the loss function of YOLOv8 object detection;

[0007] Step 3: Input the dataset after data preprocessing and data augmentation into the YOLOv8 object detection model for model training;

[0008] Step 4: Input the real-time collected image data, and use the trained YOLOv8 object detection model to detect algae.

[0009] In the above-mentioned method for detecting water surface algae targets based on deep learning, in Step 1, the data preprocessing and data augmentation include dataset annotation, dataset division, color transformation, geometric transformation, image cropping, and size unification.

[0010] In the above-mentioned method for detecting water surface algae targets based on deep learning, in Step 2, the CBAM attention module includes a channel attention module and a spatial attention module;

[0011] For the feature map F extracted by the YOLOv8 object detection model, the channel attention module uses average pooling operation and max pooling operation respectively to aggregate spatial information, and outputs a C-dimensional pooled feature map and Input and into a multi-layer perceptron MLP with hidden layers to obtain two channel attention maps with parameters of 1×1×C. The number of neurons in the hidden layer compresses the input information neurons to C / r, where r is the compression ratio; add the two channel attention maps to obtain a channel feature map F' and the corresponding channel attention vector M c , to obtain the channel attention vector M c The formula for

[0012]

[0013] Among them, MLP(AvgPool(F)) means performing average pooling operation on the feature map F and then inputting it into the multi-layer perceptron MLP. MLP(MaxPool(F)) means performing max pooling operation on the feature map F and then inputting it into the multi-layer perceptron MLP, and W 0 (·) means performing global pooling and average pooling, and W 1 (·) means learning using the multi-layer perceptron MLP, and σ(·) means the activation function;

[0014] The spatial attention module performs average pooling operation and max pooling operation on the channel feature map F' according to the channel dimension, and outputs the feature map and The output parameter is 1×H×W; concatenate the feature maps and in the channel dimension to obtain the spatial feature map F"; use a convolutional layer with a size of 7×7 to further generate the spatial attention vector M s , and the formula is as follows:

[0015]

[0016] Among them, f 7×7 represents a convolutional layer with a size of 7×7. AvgPool(F') means performing average pooling operation on the channel feature map F', and MaxPool(F') means performing max pooling operation on the channel feature map F'.

[0017] In the aforementioned water surface algae target detection method based on deep learning, the process of introducing the deformable convolution module is to replace the last, second-to-last, and third C2f modules in the backbone network of the YOLOv8 target detection model with deformable convolution modules; the convolution description formula of the deformable convolution module is as follows:

[0018]

[0019] In the formula, G represents the total number of convolution groups, K represents the number of sampling points, w g represents the shared projection weight within each convolution group, m gk represents the modulation factor of the kth sampling point in the gth group and is normalized by softmax along the dimension K, x g represents the input feature map of the slice, p 0 represents the current pixel, p k represents the kth position sampled by the predefined network in the regular convolution, (g,k) represents the network sampling position at the gth group and the kth position, and Δp gk represents the network sampling position (g,k) in the gth group.

[0020] In the above-mentioned water surface algae target detection method based on deep learning, the CARAFE upsampling module includes an upsampling kernel prediction module and a feature recombination module; in the CARAFE upsampling module, an input feature map is input, and the input size is expressed as H×W×C, and the upsampling multiple is σ; first, a 1×1 convolution is used to reduce the number of channels to C m , then content encoding is performed on the input feature map, and the size of the upsampling kernel is expressed as k up ×k up , a convolutional layer is used to predict the upsampling kernel, the number of input channels is C m , and the number of output channels is σ 2 ×k 2 up ; then the channel dimension is expanded to the spatial dimension to obtain an upsampling kernel with a shape of σH×σW×k 2 up , and then the weights of each item of the upsampling kernel are normalized; finally, the position of each pixel point of the upsampling kernel is mapped back to the input feature map, and a k up ×k up area centered on the mapped point is found, and the channel dimension of the upsampling kernel at this point is expanded into a k up ×k up spatial area is dot-producted, and finally an output feature map with a size of σH×σW×C is obtained.

[0021] In the above-mentioned water surface algae target detection method based on deep learning, the loss function for improving the YOLOv8 target detection model is to use the WIoUv3 loss function L WIoUv3 ; the formula of the WIoUv3 loss function L WIoUv3 is as follows:

[0022] L WIoUv3 =rR WIoU L IoU ;

[0023]

[0024] In the formula, L IoU represents the loss of the ratio of the intersection area to the union area between the predicted box and the true box of the YOLOv8 target detection model, L IoU ∈[0,1], is the monotonic focusing coefficient, α and δ are hyperparameters, β is the outlier degree, x is the x coordinate of the center point of the predicted box, x gt is the x coordinate of the center point of the true box, y is the y coordinate of the center point of the predicted box, y gt is the y coordinate of the center point of the true box, and the corresponding position of (x,y) in the target box is (x gt ,y gt ), W gand H g represent the height and width of the minimum bounding rectangle jointly formed by the predicted bounding box and the ground truth bounding box.

[0025] In the above-mentioned deep learning-based water surface algae target detection method, the process of model training is to first set the model training parameters, input the data set in Step 1 into the YOLOv8 target detection model for iterative training until the model training is completed, and generate the weight parameters of the YOLOv8 target detection model.

[0026] In the above-mentioned deep learning-based water surface algae target detection method, the process of model detection is to load the weight parameters in the YOLOv8 target detection model, input the real-time collected image data into the YOLOv8 target detection model, calculate various indicators, output the class position boxes and confidence levels in the image data, and finally calculate the positions where each class is located.

[0027] Compared with the prior art, the present invention has the following beneficial effects:

[0028] This method greatly improves the accuracy and efficiency of algae target detection. By collecting multi-scene image data and performing comprehensive data preprocessing and enhancement, it provides rich and diverse data for the model training of the YOLOv8 target detection model, enabling the YOLOv8 target detection model to learn the characteristics of algae in different environments. This method constructs an improved YOLOv8 target detection model, introducing the CBAM attention module, deformable convolution module, and CARAFE upsampling module, which enhance the algae feature extraction ability of the YOLOv8 target detection model from aspects such as attention allocation, adapting to irregular and small algae targets, and improving the resolution of the feature map. Even in the face of a complex water surface environment, it can accurately identify algae targets, reduce missed detections and false detections, and thus improve the overall detection efficiency. This method uses the improved loss function WIoUv3 to replace the CIou loss function, which can better reflect the matching situation between the predicted bounding box and the ground truth bounding box of the YOLOv8 target detection model. Through an intelligent gradient gain allocation strategy, the YOLOv8 target detection model can more effectively learn the target detection task, accelerate the convergence speed, reduce the training time, and at the same time improve the generalization ability of the model, avoid overfitting, and ensure that the YOLOv8 target detection model can show good detection performance on different data sets. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is the flow chart of the method of the present invention;

[0030] Figure 2 is the network structure diagram of the YOLOv8 target detection model of the present invention;

[0031] Figure 3It is the network structure diagram of the improved YOLOv8 object detection model of the present invention;

[0032] Figure 4 It is the network structure diagram of the CBAM of the present invention;

[0033] Figure 5 It is the structure diagram of the channel attention module of the present invention;

[0034] Figure 6 It is the structure diagram of the spatial attention module of the present invention;

[0035] Figure 7 It is the implementation process diagram of the deformable convolution of the present invention;

[0036] Figure 8 It is the network structure diagram of the CARAFE of the present invention. Detailed implementation manners

[0037] The present invention will be further described below in conjunction with the accompanying drawings and embodiments, but it shall not be used as a basis for limiting the present invention.

[0038] Embodiment

[0039] A method for detecting water surface algae targets based on deep learning, as Figure 1 shown, specifically includes the following steps:

[0040] Step 1: Collect image data samples and generate a data set, and perform data preprocessing and data augmentation; the data preprocessing and data augmentation include data set annotation, data set division, color transformation, geometric transformation, picture cropping and size unification.

[0041] In this embodiment, first, multi-scene image data samples containing water surface algae are collected through a camera and a sonar module to ensure that the image data samples cover a variety of water environments (such as different lighting, algae distribution density, color, etc.). Then, a labeling tool (such as labelImg) is used to accurately label the algae targets in the collected image data samples to generate corresponding labeling files (such as PASCAL VOC or COCO format). Then, the data set is divided into a training set, a validation set and a test set to ensure that the distribution of the training and test data is representative to improve the generalization ability. Next, the picture data samples are standardized to improve the training effect and convergence speed of the YOLOv8 object detection model, which specifically includes the following steps:

[0042] (1) Size unification: Adjust the size of the picture data samples to the fixed resolution required by the YOLOv8 object detection model (such as 640×640);

[0043] (2) Color transformation: Adjust the brightness and contrast of the picture, and perform color normalization to eliminate the influence of environmental lighting differences;

[0044] (3) Geometric transformation: including data augmentation operations such as random flipping, rotation, and scaling to improve the robustness of the YOLOv8 object detection model.

[0045] (4) Image cropping: Crop out the image segments containing the algae area to reduce background interference and highlight the algae features.

[0046] Step 2: Build the YOLOv8 object detection model and add the CBAM attention module to strengthen the attention allocation of YOLOv8 object detection to key target areas and improve the ability to extract algae features; introduce the deformable convolution module into the YOLOv8 object detection model to enhance the adaptability of the convolutional neural network to irregular algae targets and small algae targets and the ability to extract algae features; introduce the CARAFE upsampling module into the YOLOv8 object detection model to improve the resolution of the feature map and enhance the detection effect of small algae targets; based on the distribution characteristics of water surface algae targets, improve the loss function of YOLOv8 object detection;

[0047] The YOLOv8 object detection model is divided into 5 types according to network depth and network width: YOLOv8n, YOLOv8s, YOLOv8m, YOLOv8l, and YOLOv8x. Among them, the network depth and width control the number of sub-modules and the number of convolutional kernels respectively. The network structure of the YOLOv8 object detection model is divided into 4 parts: input end (input), backbone network (Backbone), neck (Neck), and detection head (Head), as shown in the appendix Figure 2 as shown.

[0048] In this embodiment, the YOLOv8 object detection model is mainly improved by adding the CBAM attention module, deformable convolution (DCNv3), optimizing the loss function, and introducing the CARAFE upsampling module. The network structure of the improved YOLOv8 object detection model is as shown in Figure 3 as shown, and the specific improvements are as follows:

[0049] The CBAM attention module includes a channel attention module and a spatial attention module, which are serially composed and use both global average pooling and global max pooling operations, so as to effectively prevent information loss. The CBAM attention module not only pays attention to channels but also attaches importance to spatial information, so better results can be obtained. The network structure of the CBAM attention module is as shown in Figure 4 as shown;

[0050] As Figure 5 shown, the channel attention module uses average pooling operation and max pooling operation respectively for the feature map F extracted by the YOLOv8 object detection model to aggregate spatial information and output a C-dimensional pooled feature map and Will and Input into the multilayer perceptron MLP containing the hidden layer, and obtain two channel attention maps with parameters of 1×1×C. The number of neurons in the hidden layer compresses the input information as C / r, where r is the compression ratio; add the two channel attention maps to obtain the channel feature map F' and the corresponding channel attention vector M c , get the channel attention vector M c The formula is as follows:

[0051]

[0052] Among them, MLP(AvgPool(F)) means that the feature map F is averaged and then input into the multilayer perceptron MLP, MLP(MaxPool(F)) means that the feature map F is max-pooled and then input into the multilayer perceptron MLP, W 0 (·) indicates global pooling and average pooling, W 1 (·) indicates learning using multilayer perceptron MLP, σ(·) indicates activation function;

[0053] like Figure 6 As shown, the spatial attention module performs average pooling and maximum pooling operations on the channel feature map F' according to the channel dimension, and outputs the feature map and The output parameters are 1×H×W; the feature maps are concatenated in the channel dimension and Get the spatial feature map F"; use a convolutional layer of size 7×7 to further generate the spatial attention vector M s , the formula is as follows:

[0054]

[0055] Among them, f 7×7 represents a convolutional layer of size 7×7, AvgPool(F') represents the average pooling operation on the channel feature map F', and MaxPool(F') represents the maximum pooling operation on the channel feature map F'.

[0056] During the target detection process, bounding boxes are used to describe the location of algal targets. Although their calculation is simple and intuitive, they can only provide a rough localization of algal targets and are difficult to accurately fit the shape and pose of algal targets. This limitation leads to the features extracted from bounding boxes being easily interfered by background noise or invalid information in the foreground area, thereby reducing the feature quality and further affecting the classification performance of the YOLOv8 target detection model. Especially for small targets, due to the fixed receptive field and convolution kernel size of traditional standard convolutions, it is difficult to effectively capture sufficient information, showing insufficient adaptability to the complex shapes and poses of algal targets. To solve the above problems, in this embodiment, a deformable convolution module (Deformable Convolution, DCN) is introduced into the backbone network of the YOLOv8 target detection model, aiming to improve the adaptability of the convolutional neural network to irregular algal targets and the feature extraction ability by dynamically adjusting the shape and size of the receptive field. Deformable convolution enables the convolution operation to dynamically fit the shape and size of algal targets by introducing learnable offsets, especially significantly improving the detection accuracy for small and irregular algal targets. Specifically:

[0057] (1) Due to the fixed convolution kernel weights, the receptive field of the network is consistent when processing different positions of the image, making it difficult to adapt to targets with different scales and deformations. However, deformable convolution overcomes this limitation by adaptively adjusting the shape and size of the receptive field.

[0058] (2) Deformable convolution is closer to the true size and shape of algal targets during sampling, improving the robustness to complex algal targets, which cannot be achieved by ordinary convolution.

[0059] (3) For small algal targets with small sizes and variable shapes, traditional convolution may lead to incomplete feature extraction due to insufficient receptive fields, ultimately affecting the detection performance. In contrast, deformable convolution can focus on the key feature information in the area where small algal targets are located, thereby improving the detection accuracy.

[0060] For the above reasons, in this paper, the C2f module in the backbone network of the YOLOv8 target detection model is improved. The process of introducing the deformable convolution module is to replace the last, second-to-last, and third C2f modules in the backbone network of the YOLOv8 target detection model with deformable convolution modules. Its improved network structure is as Figure 3As shown. The backbone network, as the core part of feature extraction, extracts features at different levels through stacking multiple convolutional operations and fuses these features to construct a more comprehensive image description. After introducing the deformable convolution module, the receptive field of the area where small algae targets are located is adaptively adjusted, enabling the YOLOv8 object detection model to more precisely optimize the prediction box regression parameters, enhancing the ability to focus on small targets in algae, and at the same time improving the overall performance. It is worth noting that while significantly improving the quality of feature extraction, deformable convolution is more efficient in terms of computational amount and memory consumption, further meeting the resource limitation requirements in practical applications.

[0061] The convolution description formula of the deformable convolution module is as follows:

[0062]

[0063] In the formula, G represents the total number of convolution groups, K represents the number of sampling points, w g represents the shared projection weight within each convolution group, m gk represents the modulation factor at the k-th sampling point in the g-th group and is normalized along dimension K through softmax, x g represents the input feature map of the slice, p 0 represents the current pixel, p k represents the k-th position sampled by the predefined network in the regular convolution, (g, k) represents the network sampling position at the g-th group and k-th position, Δp gk represents the network sampling position (g, k) in the g-th group.

[0064] Such as Figure 7As shown in the figure, after the feature map is processed by Deformable Convolution, a set of offsets is first generated through the mechanism unique to the Deformable Convolution layer. These offsets are adaptively learned by the network based on the input features and can dynamically adjust the positions of the sampling points. Subsequently, according to the calculated offsets, the feature map is fused with the feature map generated by the conventional convolution layer to form a new feature map with spatial adjustment. Through bilinear interpolation, this new feature map is further refined, thereby reducing the feature loss in object detection and improving the detection performance. In this embodiment, DCNv3, that is, Deformable Convolution v3, is adopted. It is an improved module in a Convolutional Neural Network (CNN) developed based on the Deformable Convolution series technology and has excellent performance in computer vision tasks such as object detection and image segmentation. It is especially suitable for processing scenarios where the shapes and postures of objects are diverse. The main innovation of DCNv3 lies in the introduction of a more advanced deformable convolution mechanism, as well as unique grouped independent sampling offsets and modulation factors. This mechanism allows the convolution process to adopt diverse spatial aggregation modes, significantly enhancing the network's adaptability to the shapes and postures of target objects and improving the flexibility and accuracy of feature expression. In addition, DCNv3 combines Depthwise Separable Convolution and Pointwise Convolution to calculate the mask and offsets respectively. This design not only reduces the computational complexity but also effectively improves the efficiency and performance of the network, especially showing significant advantages when processing irregular-shaped algae targets and small targets.

[0065] As Figure 8 shown, the CARAFE upsampling module includes an upsampling kernel prediction module and a feature recombination module; in the CARAFE upsampling module, the input feature map is input, the input size is represented as H×W×C, and the upsampling factor is σ; first, a 1×1 convolution is used to reduce the number of channels to C m , and then the input feature map is content-encoded. The size of the upsampling kernel is represented as k up ×k up , and a convolutional layer is used to predict the upsampling kernel. The number of input channels is C m , and the number of output channels is σ 2 ×k 2 up ; then the channel dimension is expanded to the spatial dimension to obtain an upsampling kernel with a shape of σH×σW×k 2 up , and then the weights of each item of the upsampling kernel are normalized; finally, the position of each pixel point of the upsampling kernel is mapped back to the input feature map, and a k centered on the mapped point is foundup × k up region, and take the dot product with the spatial region of k up × k up to finally obtain an output feature map of size σH × σW × C.

[0066] The CARAFE upsampling module can enhance the YOLOv8 object detection model's ability to perceive detailed information, thereby improving detection accuracy. The CARAFE upsampling module captures a wider range of context information through adaptive receptive fields and feature recombination, and maintains the consistency and continuity of features during the upsampling process, thus improving the quality of features. Therefore, in the improved YOLOv8, introducing the CARAFE upsampling module to replace the traditional upsampling operation can, to a certain extent, enhance the detection accuracy of the YOLOv8 object detection model, especially in dealing with detailed information. For example, in the case of small object detection or complex visual scenes, this improvement can significantly enhance the performance of the YOLOv8 object detection model.

[0067] Although the CIoU loss function in the bounding box regression task of the YOLOv8 object detection model considers three geometric factors: the aspect ratio, overlap area, and center point distance between the predicted box and the ground truth box, it still belongs to a static focusing mechanism and there is a possibility of over-punishment. For example, when the predicted box can match the ground truth box well, the punishment for geometric metrics should be weakened, thereby reducing the interference to network training. The predicted box is the result output by the YOLOv8 object detection model during the inference process of the input image data sample, and it represents the position and range of the algae target that the YOLOv8 object detection model believes exists in the image. The ground truth box, also known as the annotation box, is obtained by manual annotation during the dataset production stage and represents the true position and range of the algae target in the image data sample. In this embodiment, the WIoUv3 loss function is used to replace the CIou loss function in YOLOv8. The WIoU loss function adopts a novel method, that is, using "outlier degree" to evaluate the quality of anchor boxes. This outlier degree evaluation can better reflect the matching situation between the anchor box and the ground truth bounding box, thus more accurately measuring the quality of the anchor box. Secondly, the WIoU loss function designs an intelligent gradient gain allocation strategy based on the outlier degree. This strategy aims to effectively allocate gradients to the anchor boxes so that the YOLOv8 object detection model can better learn the object detection task. Compared with the traditional IoU loss function, the gradient gain allocation strategy of the WIoU loss function is more intelligent and flexible, and can allocate different gradients according to the quality of different anchor boxes, thereby better guiding the learning process of the YOLOv8 object detection model. The WIoU loss function not only considers the azimuth angle, centroid distance, and overlap area, but also introduces a dynamic non-monotonic focusing mechanism r. Assume that (x, y) corresponds to the position (xgt , y gt ), R WIoU represents the loss of high-quality anchor boxes. The process of improving the loss function of the YOLOv8 object detection model is to use the WIoUv3 loss function to replace the CIou loss function of the YOLOv8 object detection model; the WIoUv3 loss function L WIoUv3 has the following formula:

[0068] L WIoUv3 = rR WIoU L IoU ;

[0069]

[0070] In the formula, L IoU represents the loss of the ratio of the intersection area to the union area between the predicted box and the ground truth box of the YOLOv8 object detection model. L IoU ∈[0, 1], is the monotonic focusing coefficient, α and δ are hyperparameters, β is the outlier degree, x is the x coordinate of the center point of the predicted box, x gt is the x coordinate of the center point of the ground truth box, y is the y coordinate of the center point of the predicted box, y gt is the y coordinate of the center point of the ground truth box, and the corresponding position of (x, y) in the target box is (x gt , y gt ), W g and H g represent the height and width of the smallest bounding rectangle jointly formed by the predicted box and the ground truth box.

[0071] Step 3: Input the dataset after data preprocessing and data augmentation into the YOLOv8 object detection model for model training;

[0072] The process of the model training is to first set the model training parameters, such as learning rate, batch size, number of training epochs, etc., input the training set in Step 1 into the YOLOv8 object detection model for iterative training, and evaluate the detection performance of the model (such as mAP value, F1 score) in real time during the training process until the model training is completed, generating the weight parameters of the YOLOv8 object detection model.

[0073] Step 4: Input the real-time collected image data, and perform model detection on algae through the trained YOLOv8 object detection model.

[0074] The process of the model detection is to load the weight parameters in the YOLOv8 object detection model, input the real-time collected image data into the YOLOv8 object detection model, calculate various metrics, output the category position boxes and confidence levels in the image data, and finally calculate the positions where each category is located.

[0075] This method significantly improves the accuracy and efficiency of algae target detection. By collecting multi-scenario image data and conducting comprehensive data preprocessing and enhancement, it provides rich and diverse data for the model training of the YOLOv8 target detection model, enabling the YOLOv8 target detection model to learn the characteristics of algae in different environments. This method constructs an improved YOLOv8 target detection model, introducing the CBAM attention module, deformable convolution module, and CARAFE upsampling module, which enhance the algae feature extraction ability of the YOLOv8 target detection model from aspects such as attention allocation, adapting to the irregularity and small targets of algae, and improving the resolution of feature maps. Even in the face of a complex water surface environment, it can accurately identify algae targets, reduce missed detections and false detections, and thus improve the overall detection efficiency. This method uses the improved loss function WIoUv3 to replace the CIou loss function, which can better reflect the matching situation between the prediction box and the ground truth box of the YOLOv8 target detection model. Through an intelligent gradient gain allocation strategy, the YOLOv8 target detection model can learn the target detection task more effectively, accelerate the convergence speed, reduce the training time, and at the same time improve the generalization ability of the model, avoid overfitting, and ensure that the YOLOv8 target detection model can exhibit good detection performance on different datasets.

[0076] In summary, the present invention can overcome the interference of complex environmental factors and improve the detection accuracy and reliability of water surface algae.

Claims

1. A method for detecting algae targets on a water surface based on deep learning, characterized in that: The specific steps include: Step 1: Collect image data samples and generate data sets, and perform data preprocessing and data enhancement; Step 2: Build a YOLOv8 target detection model and add the CBAM attention module to strengthen the attention allocation of YOLOv8 target detection and improve the ability to extract algae features; introduce a deformable convolution module into the YOLOv8 target detection model to improve the adaptability of the convolutional neural network to irregular algae targets and small algae targets and the ability to extract algae features; introduce the CARAFE upsampling module into the YOLOv8 target detection model to improve the resolution of the feature map and enhance the detection effect of small algae targets; based on the distribution characteristics of algae targets on the water surface, improve the loss function of YOLOv8 target detection; Step 3: Input the data preprocessed and enhanced data set into the YOLOv8 target detection model for model training; Step 4: Input the real-time collected image data and perform model detection on algae using the trained YOLOv8 target detection model.

2. The method for detecting algae on a water surface based on deep learning according to claim 1, characterized in that: In step one, the data preprocessing and data enhancement include data set annotation, data set partitioning, color transformation, geometric transformation, image cropping and size unification.

3. The method for detecting algae on water surface based on deep learning according to claim 1, characterized in that: In step 2, the CBAM attention module includes a channel attention module and a spatial attention module; The channel attention module uses average pooling and maximum pooling operations to aggregate spatial information for the feature map F extracted by the YOLOv8 target detection model, and outputs a C-dimensional pooled feature map and Will and Input into the multilayer perceptron MLP containing the hidden layer, and obtain two channel attention maps with parameters of 1×1×C. The number of neurons in the hidden layer compresses the input information as C / r, where r is the compression ratio; add the two channel attention maps to obtain the channel feature map F' and the corresponding channel attention vector M c , get the channel attention vector M c The formula is as follows: Wherein, MLP(AvgPool(F)) means performing average pooling operation on feature map F and then inputting it into multilayer perceptron MLP, MLP(MaxPool(F)) means performing maximum pooling operation on feature map F and then inputting it into multilayer perceptron MLP, W0(·) means performing global pooling and average pooling, W1(·) means learning using multilayer perceptron MLP, and σ(·) means activation function; The spatial attention module performs average pooling and maximum pooling operations on the channel feature map F' according to the channel dimension, and outputs the feature map and The output parameters are 1×H×W; the feature maps are concatenated in the channel dimension and Get the spatial feature map F"; use a convolutional layer of size 7×7 to further generate the spatial attention vector M s , the formula is as follows: Among them, f 7×7 represents a convolutional layer of size 7×7, AvgPool(F') represents the average pooling operation on the channel feature map F', and MaxPool(F') represents the maximum pooling operation on the channel feature map F'.

4. The method for detecting algae on a water surface based on deep learning according to claim 1, characterized in that: The process of introducing the deformable convolution module is to replace the penultimate, second and third C2f modules in the backbone network of the YOLOv8 target detection model with the deformable convolution module; the convolution description formula of the deformable convolution module is as follows: In the formula, G represents the total number of convolution groups, K represents the number of sampling points, and w g represents the shared projection weight within each convolution group, m gk represents the modulation factor of the kth sample point in the gth group and is normalized by softmax along dimension K, x g represents the input feature map of the slice, p0 represents the current pixel, p k represents the k-th position of the predefined network sampling in the regular convolution, (g,k) represents the network sampling position of the k-th position of the g-th group, Δp gk represents the network sampling position (g,k) in the g-th group.

5. The method for detecting algae on water surface based on deep learning according to claim 1, characterized in that: The CARAFE upsampling module includes an upsampling kernel prediction module and a feature recombination module; The feature map is input into the CARAFE upsampling module. The input size is expressed as H×W×C, and the upsampling factor is σ. First, a 1×1 convolution is used to reduce the channel dimension to C. m , then the input feature map is content encoded, and the size of the upsampling kernel is represented as k up ×k up , using a convolutional layer to predict the upsampling kernel, the number of input channels is C m , the number of output channels is σ 2 ×k 2 up ; Then expand the channel dimension to the spatial dimension, and get the shape of σH×σW×k 2 up The upsampling kernel is then normalized for each weight of the upsampling kernel; finally, the position of each pixel of the upsampling kernel is mapped back to the input feature map, and k with the mapping point as the center point is found. up ×k up area, and the channel dimension of the sampling kernel at this point is expanded into k up ×k up The dot product is performed on the spatial area, and finally the output feature map of size σH×σW×C is obtained.

6. The method for detecting algae on water surface based on deep learning according to claim 5, characterized in that: The loss function of the improved YOLOv8 target detection model is to use the WIoUv3 loss function L WIoUv3 ; The WIoUv3 loss function L WIoUv3 The formula is as follows: L WIoUv3 =rR WIoU L IoU ; Where, L IoU The loss of the ratio of the intersection area to the union area between the predicted box and the true box of the YOLOv8 target detection model, L IoU ∈[0,1], is the monotone focusing coefficient, α and δ are hyperparameters, β is the outlier degree, x is the x coordinate of the center point of the prediction box, and x gt is the x-coordinate of the center point of the real box, y is the y-coordinate of the center point of the predicted box, gt is the y coordinate of the center point of the real frame, and the corresponding position of (x, y) in the target frame is (x gt ,y gt ), W g and H g Represents the height and width of the minimum bounding rectangle formed by the predicted box and the true box.

7. The method for detecting algae on a water surface based on deep learning according to claim 1, characterized in that: The model training process is to first set the model training parameters, input the data set in step 1 into the YOLOv8 target detection model for iterative training until the model training is completed, and generate the weight parameters of the YOLOv8 target detection model.

8. The method for detecting algae on water surface based on deep learning according to claim 7, characterized in that: The model detection process is to load weight parameters in the YOLOv8 target detection model, input the real-time collected image data into the YOLOv8 target detection model, calculate various indicators, output the category position box and confidence in the image data, and finally calculate the location of each category.

Citation Information

Patent Citations

  • Construction method of lightweight algae target detection algorithm Alga-YOLO based on YOLOv5

    CN115601559A