Water surface cleaning unmanned ship garbage recognition algorithm based on deep learning
Through the deep learning-based surface garbage recognition algorithm, the problems of high error detection rate and slow inference speed in the existing technology are solved, and high-precision and real-time surface garbage detection are achieved, which is suitable for the deployment of low-computing equipment for unmanned ships.
Patent Information
- Application Number
- CN202510288683.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
The existing surface waste detection technology has the problems of high error detection rate, slow inference speed, complex model difficulty in deploying, and insufficient detection accuracy of small targets, which cannot meet the real-time and high-precision detection requirements.
A deep learning-based waste recognition algorithm for surface clean unmanned ships is proposed. By establishing a target data set, optimizing anchor frame size, improving the YOLOv5 network model, adopting the EIoU loss function and introducing the ECA attention mechanism, the detection speed and accuracy are improved.
It significantly improves the detection speed and accuracy of surface garbage detection, meets real-time requirements, reduces the missed detection rate, adapts to mobile devices with low computing power, and improves the garbage capture rate and operating efficiency in actual applications.
Smart Images

Figure CN120219831A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of surface target recognition, and specifically to a garbage recognition algorithm for a surface cleaning unmanned ship based on deep learning. Background Technique
[0002] With the acceleration of the urbanization process and the increase in population density, water pollution and ecological damage have become core issues restricting sustainable development. According to the "China Ecological Environment Bulletin (2023)", about 35% of the key lakes and reservoirs in the country have eutrophication problems, among which the contribution rate of floating garbage on the water surface reaches 21%, mainly from domestic sewage, tourism, and construction waste. Although the country has implemented the "River Chief System" and strengthened the garbage classification policy, traditional governance means still face the dual challenges of difficult dynamic pollution control (the coverage rate of manual inspections < 30%) and insufficient performance of intelligent equipment (the false detection rate of existing algorithms is as high as 40%), and there is an urgent need for high-precision and low-latency detection technology support.
[0003] Current water surface garbage detection technologies mainly rely on classical image processing and deep learning algorithms, but both have significant defects. Classical methods (such as HSV threshold segmentation) have a false detection rate exceeding 25% in strong reflection scenarios. In deep learning solutions, although two-stage algorithms (such as Faster RCNN) have relatively high accuracy (mAP 78.5%), their inference speed is only 57 FPS, which cannot meet real-time requirements; single-stage algorithms (such as YOLO series) have a speed increased to 20 - 30 FPS, but due to the complex model (the number of parameters reaches 7.2M), it is difficult to deploy to low-computing power terminals, and the detection accuracy of small targets is less than 45%. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the existing defects, provide a garbage recognition algorithm for a surface cleaning unmanned ship based on deep learning, improve the detection speed of water surface garbage detection, reduce the missed detection of overlapping targets, and at the same time improve the accuracy of garbage detection, which can effectively solve the problems in the background technique.
[0005] To achieve the above object, the present invention proposes:
[0006] As a preferred technical solution of the present invention:
[0007] The present invention also provides a garbage recognition algorithm for a surface cleaning unmanned ship based on deep learning, including the following steps:
[0008] Step S1: Establishment of the target data set:
[0009] Construct a preliminary sample set through water surface garbage pictures actively collected by the cleaning ship in the real scene, existing open source data sets, and other Internet resources, and use the LabelImg tool to annotate and draw bounding boxes.
[0010] S11: Data preprocessing: Use the median filtering algorithm to denoise the image and reduce interference elements such as water surface illumination ripples.
[0011] S12: Data augmentation: Expand the sample size through methods such as random rotation, linear mixing, translation, scaling, and brightness adjustment to improve the generalization ability of the model. Finally, establish a data set, where the ratio of the training set, validation set, and test set is 7:2:1.
[0012] Step S2: Anchor box size optimization:
[0013] Use the Kmeans clustering algorithm to analyze the labeled real target boxes, calculate the IoU value of each bounding box with all cluster centers, and assign it to the cluster with the largest IoU value. Continuously update the cluster centers until the data points within the clusters are as similar as possible and the data points between the clusters are as different as possible, thereby optimizing the anchor box size.
[0014] Step S3: Improve the YOLOv5 network model:
[0015] Use the EfficientNetv2 backbone network to replace the backbone network of YOLOv5.
[0016] Introduce the ECA (EfficientChannelAttention) attention mechanism in the backbone part of the model. Perform global average pooling on the input feature map, keep the number of channels unchanged, and adaptively calculate the size of the one-dimensional convolution kernel according to the number of channels to generate the weight vector for each channel, and multiply it with the original feature map to achieve channel-wise weighting.
[0017] Add an auxiliary training head in the middle layer of the model to enhance gradient information and assist in model training.
[0018] Step S4: Loss function improvement and model training:
[0019] Adopt the EIoU (EfficientIntersectionoverUnion) loss function to replace the original CIoU loss function. Split the loss of the width-height ratio into the loss of width and the loss of height and calculate them separately. The loss function formula is:
[0020]
[0021] Where: IoU represents the ratio of the area of the intersection to the union of the true box and the predicted box; ρ represents the Euclidean distance between two points; b and bgt represent the center points of the predicted box and the true box respectively; wc and hc represent the width and height of the minimum bounding rectangle of the predicted box and the true box respectively; w and h represent the width and height of the predicted box. wgt and hgt represent the width and height of the true box.
[0022] Step S5: Model evaluation and optimization:
[0023] Calculate the gradient based on the loss value, update the model parameters using an optimizer, calculate the loss value and accuracy metrics on the validation set, evaluate the precision P, recall R, and mean average precision mAP of the model, and adjust the model parameters according to the evaluation results.
[0024] Preferably, in step S1, a water surface garbage dataset is made, and the pictures in the dataset are divided into the following four categories: meshes, small organic garbage, large organic garbage, and plastic garbage.
[0025] Preferably, in step S2, the specific implementation of the Kmeans clustering algorithm includes: dividing the data points into K clusters; calculating the IoU value between each bounding box and all cluster centers, and assigning it to the cluster with the largest IoU value; continuously updating the cluster centers, and finally making the data points within the clusters as similar as possible and the data points between the clusters as different as possible.
[0026] Preferably, in step S3, the specific implementation of the ECA attention mechanism includes: performing global average pooling on the input feature map while keeping the number of channels unchanged; adaptively calculating the size of the one-dimensional convolution kernel according to the number of channels; calculating the weights using the weight sharing strategy and applying a one-dimensional convolution operation to generate the weight vector for each channel; multiplying the generated weight vector with each channel of the original feature map to achieve channel-wise weighting.
[0027] Preferably, in step S5, the specific steps of model evaluation include: Precision (P): Calculate the proportion of samples that are actually positive among the samples predicted as positive by the model. The formula is:
[0028]
[0029] where TP represents the number of positive samples correctly predicted as positive by the model, and FP represents the number of negative samples wrongly predicted as positive by the model;
[0030] Recall (R): Calculate the proportion of samples correctly predicted as positive by the model among all actual positive samples. The formula is:
[0031]
[0032] where FN represents the number of positive samples wrongly predicted as negative by the model;
[0033] Average Precision (AP): Draw a curve with recall as the horizontal axis and precision as the vertical axis, and calculate the area enclosed by the curve and the coordinate axes;
[0034] Mean Average Precision (mAP): Calculate the mean of the average precisions of all classes. The formula is:
[0035]
[0036] Among them, m represents the number of dataset categories.
[0037] Preferably, in step S5, the improved YOLOv5s model is compared with the original YOLOv5s model and other mainstream models, and the performance differences of different models in terms of accuracy, recall rate, mAP, and FPS are analyzed to provide data support for subsequent model optimization and practical applications.
[0038] Preferably, in step S3, the improved YOLOv5 network model achieves an inference speed of 23.8 FPS on the embedded platform, with a power consumption ≤ 5W, adapting to low-computing-power mobile devices.
[0039] Preferably, in step S1, the data augmentation methods include random rotation, linear mixing, translation, scaling, and brightness adjustment to improve the generalization ability of the model.
[0040] Preferably, in step S3, the introduction of the auxiliary training head enhances the feature information of the intermediate layer and improves the detection accuracy of the model for small target garbage.
[0041] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0042] 1. The detection accuracy is significantly improved: By introducing the EfficientNetv2 backbone network and the ECA (EfficientChannel Attention) attention mechanism, the improved YOLOv5 model reaches 89.2% in the mean average precision (mAP), which is 12.4% higher than that of the traditional YOLOv5s model (76.8%), and is significantly better than other mainstream models (such as 78.5% of Faster R-CNN and 72.3% of SSD); the detection accuracy for small target garbage (such as small pieces of organic garbage and plastic garbage) is particularly improved significantly, and the missed detection rate is reduced to less than 5%;
[0043] 2. The inference speed is optimized to meet real-time requirements: The improved model achieves an inference speed of 23.8 FPS on the embedded platform (such as the Rockchip RK3588 chip), which is more than 3 times higher than that of the traditional two-stage algorithm (such as 5 - 7 FPS of Faster R-CNN), meeting the real-time detection requirements of the unmanned surface cleaning ship; through the lightweight design of the model (the number of parameters is compressed to 2.0M), the consumption of computing resources is significantly reduced, adapting to low-computing-power mobile devices;
[0044] 3. Enhanced environmental adaptability: By adopting the median filtering algorithm and data augmentation techniques (such as random rotation and brightness adjustment), interference factors such as water surface reflection and ripples are effectively suppressed. The false detection rate in strong reflection scenarios is reduced to less than 10%, which is significantly improved compared with the traditional HSV threshold segmentation method (false detection rate > 25%); by introducing the EIoU loss function, the calculation method of the width-to-height ratio is optimized, further improving the detection stability of the model in complex environments;
[0045] 4. Model lightweight and low power consumption: The number of parameters of the improved YOLOv5 network model is only 2.0M, which is compressed by 72% compared with the original YOLOv5s model (7.2M), significantly reducing the storage and computing resource requirements; when running on an embedded platform, the power consumption ≤ 5W, which is suitable for the use of unmanned surface cleaning boats for long-term operations;
[0046] 5. Significant actual application effects: In the actual test in Taihu Lake waters, the garbage capture rate of the unmanned boat equipped with the improved algorithm is increased to 92%, which is 22% higher than the traditional scheme (70%); the average daily operation area is expanded by 3 times, covering a water area of 5000 square meters per day, and the operation and maintenance cost is reduced by 41%, saving about 500,000 yuan annually;
[0047] 6. Versatility and expandability: Support the detection of multiple types of garbage (such as meshes, small pieces of organic garbage, large pieces of organic garbage, plastic garbage), and adapt to the governance needs of different water environments; by introducing an auxiliary training head, the intermediate layer feature information is enhanced, further improving the model's detection ability for small target garbage. Description of the Drawings
[0048] Figure 1 is the overall experimental flowchart of the present invention;
[0049] Figure 2 is the dataset K-means clustering flowchart of the present invention;
[0050] Figure 3 is the ECA structure diagram of the present invention;
[0051] Figure 4 is the auxiliary detection head structure diagram of the present invention. Detailed Embodiments
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0053] Please refer to Figure 1, the present invention provides the following technical solutions:
[0054] Construction and preprocessing of the water surface garbage dataset:
[0055] 1. Dataset construction:
[0056] Data source: 1000 water surface garbage pictures collected by the unmanned surface cleaning ship in the Taihu Lake waters, combined with relevant pictures in the public datasets (such as COCO, Pascal VOC), to construct a preliminary sample set containing 5000 pictures in total.
[0057] Data annotation: Use the LabelImg tool to annotate the garbage targets in the pictures. The annotation categories include: mesh, small pieces of organic garbage (such as grass, leaves), large pieces of organic garbage (such as branches, wood), plastic garbage.
[0058] Generate corresponding bounding boxes and class labels for each picture.
[0059] 2. Data preprocessing:
[0060] Denoising processing: Use the median filtering algorithm to denoise the image, eliminate interference elements such as water surface illumination ripples, and improve the image quality.
[0061] Data augmentation: Expand the sample size through the following methods:
[0062] Random rotation (angle range: 30° to 30°);
[0063] Linear mixing (Alpha = 0.5);
[0064] Translation (maximum offset: 10% of the image width and height);
[0065] Scaling (scale range: 0.8 to 1.2);
[0066] Brightness adjustment (brightness change range: 20% to 20%).
[0067] The final dataset is divided into a training set (3500 pictures), a validation set (1000 pictures), and a test set (500 pictures) according to the ratio of 7:2:1.
[0068] Optimization of anchor box sizes and Kmeans clustering:
[0069] 1. Optimization of anchor box sizes
[0070] Kmeans clustering: Conduct Kmeans clustering analysis on the labeled real target boxes, and set K = 9 (that is, generate 9 anchor box sizes).
[0071] Clustering process:
[0072] Initialize 9 cluster centers;
[0073] Calculate the IoU value of each bounding box with all cluster centers and assign it to the cluster with the largest IoU value;
[0074] Update the cluster centers and repeat the iteration until convergence.
[0075] Result: The optimized anchor box size fits the distribution of actual garbage targets better, significantly improving the model's detection ability for small target garbage.
[0076] Improve the YOLOv5 network model:
[0077] 1. Backbone network replacement
[0078] Replace the CSPDarknet53 backbone network of YOLOv5 with the EfficientNetv2 backbone network to improve the feature extraction ability.
[0079] 2. Introduction of ECA attention mechanism
[0080] Insert the ECA module after the convolutional layer of the backbone network. The specific implementation steps are as follows:
[0081] Perform global average pooling on the input feature map while keeping the number of channels unchanged;
[0082] Adaptive calculate the one-dimensional convolutional kernel size according to the number of channels;
[0083] Formula: where C is the number of channels, and y and b are hyperparameters;
[0084] Calculate the weights using the weight sharing strategy to generate the weight vector for each channel;
[0085] Multiply the weight vector by the original feature map to achieve channel-wise weighting.
[0086] 3. Addition of auxiliary training heads
[0087] Add auxiliary training heads in the middle layer of the model to enhance the gradient information, assist the model training, and improve the detection accuracy for small target garbage.
[0088] EIoU loss function and model training:
[0089] 1. Improvement of the loss function
[0090] Adopt the EIoU loss function to replace the original CIoU loss function. Split the loss of the width-height ratio into the loss of width and the loss of height and calculate them separately. The formula is:
[0091]
[0092] Among them: IoU represents the ratio of the area of the intersection of the ground truth box and the predicted box to the area of their union; ρ represents the Euclidean distance between two points; b and bgt represent the center points of the predicted box and the ground truth box respectively; wc and hc represent the width and height of the minimum bounding rectangle of the predicted box and the ground truth box respectively; w and h represent the width and height of the predicted box; wgt and hgt represent the width and height of the ground truth box respectively.
[0093] 2. Model Training
[0094] Use the Adam optimizer with an initial learning rate of 0.001 and a batch size of 16;
[0095] Iteratively train on the training set for 100 epochs, and calculate the loss value and accuracy metrics on the validation set every 10 epochs;
[0096] When the value of the loss function reaches the minimum, save the optimal model parameters.
[0097] Model Evaluation and Performance Comparison:
[0098] 1. Evaluation Metrics
[0099] Precision (P):
[0100]
[0101] Among them, TP represents the number of positive samples correctly predicted as positive by the model, and FP represents the number of negative samples wrongly predicted as positive by the model.
[0102] Recall (R):
[0103]
[0104] Among them, FN represents the number of positive samples wrongly predicted as negative by the model.
[0105] Mean Average Precision (mAP):
[0106]
[0107] Among them, m represents the number of dataset categories.
[0108] 2. Performance Comparison
[0109] Compare the improved YOLOv5s model with the original YOLOv5s model and other mainstream models (such as FasterRCNN, SSD), and the results are as follows:
[0110] Model mAP (%) FPS Number of parameters (M) Faster R-CNN 78.5 5-7 137.1 SSD 72.3 20-25 26.8 YOLOv5s (Original) 76.8 20-30 7.2 YOLOv5s (Improved) 89.2 23.8 2.0
[0111] 3. Result Analysis
[0112] The improved YOLOv5s model has a 12.4% increase in mAP, the inference speed remains at 23.8 FPS, and the number of parameters is compressed to 2.0M, significantly outperforming other models;
[0113] Real-time detection is achieved on an embedded platform (Rockchip RK3588 chip), with power consumption ≤ 5W, suitable for low-computing mobile devices.
[0114] Practical application and effect verification:
[0115] 1. Application scenarios
[0116] Deploy a surface cleaning unmanned boat equipped with the improved algorithm in the Taihu Lake waters for a 30-day garbage cleaning test.
[0117] 2. Effect verification
[0118] Garbage capture rate: increased to 92%, a 22% increase compared to the traditional solution (70%);
[0119] Average daily operation area: expanded by 3 times, covering a water area of 5000 square meters per day;
[0120] Operation and maintenance cost: reduced by 41%, saving about 500,000 yuan annually.
[0121] 3. User feedback
[0122] Feedback from operators: The improved algorithm significantly improves the accuracy and real-time performance of garbage recognition, reducing the frequency of manual intervention;
[0123] Evaluation by the environmental protection department: This technology provides efficient and reliable technical support for water area ecological governance, contributing to the realization of the "dual carbon" goal.
[0124] The specific implementation process of the garbage recognition algorithm for the surface cleaning unmanned boat based on deep learning is described in detail through the above embodiments, covering key steps such as dataset construction, anchor box optimization, network model improvement, loss function design, model evaluation and optimization, and the effectiveness and practicality of the technology are verified through practical applications.
[0125] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A deep learning-based garbage recognition algorithm for surface cleaning unmanned boats, characterized by: The following steps are involved: Step S1: Establishment of target data set: A preliminary sample set was constructed using images of surface garbage collected by cleaning ships in real scenarios, existing open source datasets, and other Internet resources. The LabelImg tool was used to annotate and draw bounding boxes. S11: Data preprocessing: Use the median filter algorithm to denoise the image and reduce interference elements such as water surface illumination ripples; S12: Data enhancement: The sample size is expanded through random rotation, linear mixing, translation, scaling, brightness adjustment and other methods to improve the generalization ability of the model, and finally a data set is established, in which the ratio of training set, validation set and test set is 7:2:1; Step S2: Anchor box size optimization: The Kmeans clustering algorithm is used to analyze the annotated real target boxes, calculate the IoU value between each bounding box and all cluster centers, and assign it to the cluster with the largest IoU value. The cluster center is continuously updated to ultimately make the data points within the cluster as similar as possible and the data points between clusters as different as possible, thereby optimizing the anchor box size. Step S3: Improve the YOLOv5 network model: Use EfficientNetv2 backbone network to replace the YOLOv5 backbone network; The ECA attention mechanism is introduced into the main part of the model to perform global average pooling on the input feature map, keep the number of channels unchanged, and adaptively calculate the one-dimensional convolution kernel size according to the number of channels to generate a weight vector for each channel, which is multiplied with the original feature map to achieve inter-channel weighting; Add auxiliary training heads to the middle layer of the model to enhance gradient information and assist model training; Step S4: Loss function improvement and model training: The EIoU loss function is used to replace the original CIoU loss function, and the loss of the aspect ratio is split into the width loss and the height loss, and the loss function formula is: Where: IoU represents the ratio of the intersection and union of the real box and the predicted box; ρ represents the Euclidean distance between two points; b and bgt represent the center points of the predicted box and the real box respectively; wc and hc represent the width and height of the minimum bounding rectangle of the predicted box and the real box respectively; w and h represent the width and height of the predicted box respectively. wgt and hgt represent the width and height of the real box respectively; Step S5: Model evaluation and optimization: Calculate the gradient according to the loss value, use the optimizer to update the model parameters, calculate the loss value and accuracy index on the validation set, evaluate the model's precision P, recall rate R and average precision mAP, and adjust the model parameters according to the evaluation results.
2. The deep learning-based water surface cleaning unmanned ship garbage recognition algorithm according to claim 1 is characterized by: In step S1, a water surface garbage dataset is created, and the images in the dataset are divided into the following four categories: webs, small pieces of organic garbage, large pieces of organic garbage, and plastic garbage.
3. The deep learning-based water surface cleaning unmanned ship garbage recognition algorithm according to claim 1 is characterized by: In step S2, the specific implementation of the Kmeans clustering algorithm includes: dividing the data points into K clusters; calculating the IoU value between each bounding box and all cluster centers, and assigning it to the cluster with the largest IoU value; continuously updating the cluster center, and finally achieving that the data points within the cluster are as similar as possible and the data points between clusters are as different as possible.
4. The deep learning-based water surface cleaning unmanned ship garbage recognition algorithm according to claim 1 is characterized by: In step S3, the specific implementation of the ECA attention mechanism includes: performing global average pooling on the input feature map, keeping the number of channels unchanged; adaptively calculating the one-dimensional convolution kernel size according to the number of channels; calculating the weight using the weight sharing strategy, and applying a one-dimensional convolution operation to generate a weight vector for each channel; multiplying the generated weight vector with each channel of the original feature map to achieve inter-channel weighting.
5. The deep learning-based water surface cleaning unmanned ship garbage recognition algorithm according to claim 1 is characterized by: In step S5, the specific steps of model evaluation include: Precision (P): Calculate the proportion of samples predicted by the model to be positive that are actually positive, the formula is: Among them, TP represents the number of positive samples correctly predicted as positive by the model, and FP represents the number of negative samples incorrectly predicted as positive by the model; Recall (R): The ratio of samples correctly predicted by the model as positive to all actual positive samples. The formula is: Among them, FN represents the number of positive samples that are incorrectly predicted as negative by the model; Average precision (AP): Draw a curve with recall as the horizontal axis and precision as the vertical axis, and calculate the area enclosed by the curve and the coordinate axis; Mean Average Precision (mAP): Calculate the mean of the average precision of all categories, the formula is: Among them, m represents the number of categories in the dataset.
6. The deep learning-based water surface cleaning unmanned ship garbage recognition algorithm according to claim 1 is characterized by: In step S5, the improved YOLOv5s model is compared with the original YOLOv5s model and other mainstream models to analyze the performance differences of different models in terms of accuracy, recall, mAP and FPS, so as to provide data support for subsequent model optimization and practical application.
7. The deep learning-based garbage recognition algorithm for surface cleaning unmanned boats according to claim 1 is characterized by: In step S3, the improved YOLOv5 network model achieves an inference speed of 23.8FPS on the embedded platform with a power consumption of ≤5W, which is suitable for low-computing mobile devices.
8. The deep learning-based garbage recognition algorithm for surface cleaning unmanned boats according to claim 1 is characterized by: In step S1, data augmentation methods include random rotation, linear mixing, translation, scaling, and brightness adjustment to improve the generalization ability of the model.
9. The deep learning-based water surface cleaning unmanned ship garbage recognition algorithm according to claim 1 is characterized by: In step S3, the introduction of the auxiliary training head enhances the feature information of the middle layer and improves the detection accuracy of the model for small target garbage.