Improved YOLOv8-based inland water area plastic garbage detection method
By improving the YOLOv8 framework, the problem of insufficient accuracy and adaptability in plastic waste detection in inland waters has been solved, and efficient and accurate plastic waste detection is achieved, which is suitable for real-time monitoring and cleaning of unmanned surface vehicles.
Patent Information
- Application Number
- CN202510679804.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-08-19
AI Technical Summary
The existing deep learning models have problems in the detection of floating plastic waste in inland waters that are not high in detection accuracy and high leakage detection rate, especially in the detection of small-target plastic waste, and the detection effect is poor in complex environments.
The improved YOLOv8 framework is adopted to improve the detection accuracy and adaptability of the model to small-target plastic waste by using lightweight receptive field coordinate attention convolution operation, adding high-resolution prediction heads and coordinate attention modules, building SimSPPFCSPC blocks, and optimizing the loss function. Combining data augmentation technology, the model's detection accuracy and adaptability to small-target plastic waste is improved.
It significantly improves the model's detection accuracy of plastic waste in inland waters, especially the detection accuracy of small targets, enhances the model's adaptability in complex environments, and improves the detection efficiency to meet the real-time monitoring needs.
Smart Images

Figure CN120510362A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of environmental monitoring and protection, and in particular to a method for detecting plastic waste in inland waters based on deep learning technology. Background Art
[0002] Marine plastic pollution has become a major environmental issue that has attracted global attention. Its potential risks to marine ecosystems and human health cannot be underestimated. According to authoritative data, global plastic production has increased rapidly from 234 million tons to 460 million tons between 2000 and 2019. At the same time, the amount of plastic waste has also increased dramatically, from 156 million tons to 353 million tons. 1 There are many ways for plastic waste to enter the ocean, including wind, frequent ship transportation, the impact of various activities in coastal areas, and illegal discharge of sewage, all of which have become channels for plastic to invade the ocean. [2] Inland waters, including canals, rivers, lakes and bays, continuously transport large amounts of plastic waste to the ocean, becoming the main source of marine plastic waste. 3 ]. Relevant studies estimate that approximately 0.8 to 2.7 million tons of river plastic are discharged into the ocean through inland waters each year.
[0003] Faced with the proliferation of floating plastic waste in inland waters, traditional cleaning strategies mainly rely on manual methods. In actual operations, staff need to drive boats or use tools on the shore, such as long-pole net bags, garbage salvage boats, etc., to salvage plastic waste floating on the water one by one. This method may be feasible when the amount of garbage is small, but when faced with large-scale floating plastic waste, its disadvantages are obvious. Manual cleaning is not only time-consuming and labor-intensive, but also costly. [4] For example, cleaning up large rivers or lakes requires a lot of manpower and time, and is inefficient.
[0004] In recent years, with the rapid development of computer vision technology and deep learning algorithms, new hope has been brought to the detection and cleaning of plastic waste in inland waters. 5 ] [6] Deep learning methods, especially object detection models based on convolutional neural networks (CNNs), have achieved remarkable results in the fields of image recognition and object detection. Among them, the You Only Look Once (YOLO) series of models has occupied an important position in the field of computer vision with its excellent performance and efficient detection speed. 7 ]. These models can automatically learn feature representations in images by building multi-layer neural networks, thereby achieving accurate detection and classification of target objects.
[0005] However, despite the excellent performance of the YOLO series of models in general target detection tasks, they still face many severe challenges when applied to the detection of floating plastic waste in inland waters. The inland water environment is complex and changeable. Factors such as the color, transparency, lighting conditions, and fluctuations of the water surface will greatly interfere with image acquisition and target detection. Moreover, floating plastic waste varies in shape, size, and color, especially small-target plastic waste, which is difficult to accurately identify and locate against a complex background. In addition, when dealing with a large number of small target objects, existing models often have problems such as low detection accuracy and high missed detection rates. For example, in actual inland water scenarios, a large number of tiny plastic fragments or small plastic packaging may be missed by the model, resulting in incomplete cleaning.
[0006] In summary, developing a deep learning model specifically for the complex environment of inland waters that can detect floating plastic waste with high precision and applying it to autonomous monitoring and cleaning systems has become a key issue that needs to be urgently addressed in the current field of environmental protection.
[0007] References:
[0008] [1] Kumar, R., Verma, A., Shome, A., Sinha, R., Sinha, S., Jha, PK, Kumar, R., Kumar, P., Shubham, Das, S., 2021. Impacts of plastic pollution on ecosystem services, sustainable development goals, and need to focus on circular economy and policy interventions. Sustainability 13, 9963.
[0009] [2]Woodall,LC,Sanchez-Vidal,A.,Canals,M.,Paterson,GL,Coppock,R.,Sleight,V.,Calafat,A.,Rogers,AD,Narayanaswamy,BE,Thompson,RC,2014.The deep sea is a major sink for microplastic debris.Royal Society Open Science1,140317.
[0010] [3]Cheng,Y.,Zhu,J.,Jiang,M.,Fu,J.,Pang,C.,Wang,P.,Sankaran,K.,Onabola,O.,Liu,Y.,Liu,D.,2021.Flow:A dataset and benchmark for floating wastedetection in inland waters,in:Proceedings of the IEEE / CVF InternationalConference on Computer Vision,pp.10953–10962.
[0011] [4]Ruangpayoongsak,N.,Sumroengrit,J.,Leanglum,M.,2017.A floatingwaste scooper robot on water surface,in:201717th International Conference onControl,Automation and Systems(ICCAS),IEEE.pp.1543–1548.
[0012] [5]Fallati,L.,Polidori,A.,Salvatore,C.,Saponari,L.,Savini,A.,Galli,P.,2019.Anthropogenic marine debris assessment with unmanned aerial vehicleimagery and deep learning:A case study along the beaches ofthe republicofmaldives.Science ofThe Total Environment 693,133581.
[0013] [6]Jakovljevic,G.,Govedarica,M.,Alvarez-Taboada,F.,2020.A deeplearning model for automatic plastic mapping using unmanned aerial vehicle(UAV)data.Remote Sensing 12,1515.
[0014] [7]Jocher, G., Chaurasia, A., Qiu, J., 2023. YOLO by Ultralytics (Version8.0.0) [computer software]. YOLO by Ultralytics (Version 8.0.0) [Computer software]. Summary of the Invention
[0015] The core purpose of this invention is to build a deep learning-based inland water plastic waste detection and cleaning system to effectively overcome the many difficulties faced by existing technologies in handling inland water plastic waste detection and cleaning tasks, significantly improve detection accuracy and efficiency, and can be used for unmanned surface vehicles (USVs) to autonomously monitor and efficiently clean inland water plastic waste. The technical solution of this invention is as follows:
[0016] A method for detecting plastic waste in inland waters based on improved YOLOv8 is characterized by adopting the YOLOv8 basic framework and performing the following network modifications on it:
[0017] In the convolutional layer of the model, the lightweight receptive field coordinate attention convolution operation LRFCAConv replaces the standard convolution operation. LRFCAConv first aggregates the global information of the receptive field features through global average pooling (AvgPool), compresses the two-dimensional feature map in the spatial dimension, and obtains a one-dimensional feature vector containing global information. Then, it uses depthwise separable convolution (DSC) for information interaction, decomposing the traditional convolution into depthwise convolution and point-by-point convolution. Then, the Softmax function is used to generate attention weights that reflect the importance of features at different positions. Finally, a fixed-step 3×3 convolution operation is used to extract feature information.
[0018] Improved prediction heads: The original model had three-scale prediction heads of 80×80px, 40×40px, and 20×20px. The improved model added a 160px×160px prediction head for detecting higher-resolution feature maps, which is used to detect small plastic waste targets. At the same time, a coordinate attention (CA) module was added before each prediction head to form a coordinate attention head (CAH), enabling the model to better capture local and global relationships in space.
[0019] The SimSPPFCSPC block is constructed based on the SPPCSPC block: First, the two CBS layers before the pooling layer are trimmed, retaining only one CBS layer for feature smoothing. Second, the SimAM attention mechanism is introduced before the pooling layer. The SimAM attention mechanism automatically distinguishes target pixels from other pixels by calculating attention weights in the channel and spatial dimensions, highlighting the features of small target areas and enhancing attention to dense small target areas. Finally, the pooling kernel setting is reduced from (5,5,5) to (3,3,3).
[0020] Replace the bounding box loss function of the network model from Complete IoU (CIoU) to Inner-WIoU loss function which is a combination of Wise-IoU and Inner-IoU;
[0021] Use the prepared training set to train the improved YOLOv8 model.
[0022] Furthermore, the coordinate attention (CA) module can take into account both input feature information and pixel position information, and enable the model to better capture local and global relationships in space by aggregating and recalibrating the features of the feature map in the horizontal and vertical directions.
[0023] Furthermore, data augmentation and processing are also included: randomly flipping the image horizontally with a probability of 50% to simulate the appearance of plastic waste from different perspectives; randomly rotating the image within a certain angle range to increase image diversity; randomly cropping the image according to a certain ratio to simulate the partial occlusion of plastic waste in real scenes; and randomly adjusting the brightness, contrast, and saturation of the image to enhance the model's adaptability to plastic waste under different lighting conditions. Oversampling technology is used to generate new samples by duplicating small object sample images or performing local transformations on them, increasing the proportion of small object samples in the training set.
[0024] Furthermore, during the training process of the improved YOLOv8 model, the training set images are input into the model. The model predicts the plastic waste in the image based on the current parameters, and calculates the Inner-WIoU loss value between the predicted result and the true annotation; the gradient of the loss value with respect to each model parameter is calculated through the back propagation algorithm, and the model parameters are updated using the gradient descent method.
[0025] The beneficial effects of the present invention are:
[0026] 1. Improving Detection Accuracy: By introducing the LRFCAConv operation, improving the prediction head, using the SimSPPFCSPC block, and optimizing the loss function, the model effectively improves the detection accuracy of plastic waste in inland waters, especially small plastic objects. Tested on the FloW dataset, the improved ENS-YOLO model achieved mAP50 and mAP50-95 performance of 91.8% and 51.7%, respectively. This is a significant improvement over the traditional YOLOv8 model, and can more accurately identify and locate plastic waste in inland waters.
[0027] 2. Enhanced Model Adaptability: This method takes into account the complex environmental factors of inland waters. Through data augmentation and optimized model structure, the model is able to better learn the characteristics of plastic waste in different environments. For example, data augmentation simulates environmental conditions such as varying lighting intensity and water turbidity, exposing the model to diverse data during training and enhancing its adaptability to complex environments. The improved model maintains high detection accuracy under varying conditions such as light intensity and water turbidity, reducing the impact of environmental factors on detection results.
[0028] 3. Improved detection efficiency: While ensuring high detection accuracy, the optimized model has a relatively reasonable number of parameters and computational complexity. Fewer parameters make the model more advantageous in storage and deployment, while lower computational complexity enables it to run smoothly in resource-constrained environments, such as enabling real-time detection on unmanned surface vehicles (USVs). The model achieves a frame rate of 120 frames per second (FPS), enabling rapid image processing and meeting the needs of real-time monitoring of plastic waste in inland waters, thereby improving detection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to illustrate the reliability and accuracy of the technical solutions in the embodiments of the present invention in more detail and intuitively, the following briefly introduces the drawings required for use in the embodiments. The drawings described below are only some examples of the present invention. Through some simple modifications, the present invention can be applied to other examples without paying any creative work.
[0030] Figure 1 FloW dataset image samples
[0031] Figure 2 YOLOv8 network structure diagram
[0032] Figure 3 ENS-YOLO network structure diagram proposed by this invention
[0033] Figure 4 Comparison of detection results of YOLOv8, ENS-YOLO and real annotations on FloW images DETAILED DESCRIPTION
[0034] First, the basic solution of the present invention is described with an eye on the modification of the existing network.
[0035] LRFCAConv Operation: In the convolutional layers of the model, the standard convolution operation is replaced with the lightweight receptive field coordinate attention convolution operation (LRFCAConv). LRFCAConv first aggregates the global information of the receptive field features through global average pooling (AvgPool), compressing the two-dimensional feature map in the spatial dimension to obtain a one-dimensional feature vector containing global information. Next, depthwise separable convolution (DSC) is used for information exchange. Depthwise separable convolution decomposes the traditional convolution into depthwise convolution and pointwise convolution, reducing computational complexity while preserving the spatial information of the features. The Softmax function is then used to generate attention weights. The Softmax function converts the feature vector into a probability distribution of attention weights that reflects the importance of features at different locations. Finally, feature information is extracted through a fixed-step 3×3 convolution operation. In this process, the receptive field spatial features are dynamically generated based on the convolution kernel size. In other words, the generation of the receptive field spatial features depends on parameters such as the convolution kernel size and stride, and adapts to changes in these parameters. In this way, LRFCAConv can effectively improve the feature extraction capability while reducing the computational burden and number of parameters of the model, enabling the model to better capture the characteristics of plastic waste.
[0036] Prediction Head Improvement: The low-resolution feature maps corresponding to the original model's three-scale prediction heads (80×80px, 40×40px, and 20×20px) struggle to maintain the geometric details of small objects. Therefore, a prediction head for detecting higher-resolution feature maps (160px×160px) was added to the model head, specifically for detecting small plastic waste objects. Since small objects are easily lost in low-resolution feature maps after multiple layers of convolution, the newly added high-resolution prediction head better preserves their feature information. Furthermore, a coordinate attention (CA) module was added before each prediction head, forming a coordinate attention head (CAH). The CA module takes into account both input feature information and pixel position information. By aggregating and recalibrating the feature maps in the horizontal and vertical directions, it enables the model to better capture local and global relationships in space. This improves the accuracy of small object detection while reducing the number of model parameters and computational complexity, thereby improving the model's operational efficiency.
[0037] SimSPPFCSPC block: The SimSPPFCSPC block is constructed based on the SPPCSPC (Spatial Pyramid Pooling Cross Stage PartialChannel) block. First, the two CBS layers preceding the pooling layer are pruned, retaining only one CBS layer for feature smoothing. CBS layers (Convolution+BatchNormalization+SiLU) are used in traditional models for feature extraction and normalization. However, excessive CBS layers may over-filter edge information of small objects, affecting small object detection performance. Pruning CBS layers can reduce this over-filtering, reducing computational complexity and parameter count. Second, the SimAM (Similarity-based Attention Module) attention mechanism is introduced before the pooling layer. By calculating attention weights in both the channel and spatial dimensions, the SimAM attention mechanism automatically distinguishes target pixels from other pixels, highlighting features in areas containing small objects, enhancing attention to densely packed small objects, reducing the loss of valid features, and suppressing the expression of interfering features. Finally, the pooling kernel is reduced from (5,5,5) to (3,3,3). Larger pooling kernels may lose detailed information about small objects when aggregating features. However, smaller pooling kernels have a receptive field that better matches the scale of small objects, helping to extract more small object features and further improve the localization accuracy of small object detection. Through these improvements, the SimSPPFCSPC block can better adapt to the detection needs of densely packed small objects in inland water plastic waste detection.
[0038] Optimized loss function
[0039] The network model's bounding box loss function was replaced from Complete IoU (CIoU) to the Inner-WIoU loss, a combination of Wise-IoU and Inner-IoU. The quality of the anchor boxes annotated in the dataset varies. Low-quality data annotations, even with the support of an efficient fitting loss function, can interfere with model convergence and hinder effective feature learning. Wise-IoU, based on boundary regression using a dynamic non-monotonic focusing mechanism, dynamically allocates gradient gains based on anchor box outliers. In the early stages of training, it prioritizes the regression of high-quality anchor box boundaries to mitigate the impact of harmful gradients from low-quality anchor boxes. In the middle and late stages of training, it reduces the gradient gains for both high- and low-quality anchor boxes, concentrating the loss calculation on anchor boxes of average quality. Furthermore, Wise-IoU decouples the minimum bounding box width and height, which hinder convergence, from its calculation, improving the convergence speed of the network model. However, networks trained using the IoU loss cannot adapt to different detectors and detection tasks, resulting in poor generalization. Therefore, based on Wise-IoU, Inner-IoU is introduced. It uses auxiliary bounding boxes of different sizes to calculate the loss according to different IoU samples to accelerate convergence. For high IoU samples, smaller auxiliary bounding boxes are used to calculate the loss; for low IoU samples, larger auxiliary bounding boxes are more appropriate. By combining Wise-IoU and Inner-IoU to obtain the Inner-WIoU loss function, the generalization ability and convergence speed of the network are improved, further enhancing the detection accuracy of the model. The formula of the Inner-WIoU loss function is defined as:
[0040]
[0041] Data processing and training strategies
[0042] FloW and DeepTrash were selected as datasets for model training and validation. FloW is the first dataset for detecting floating garbage from the perspective of unmanned ships. It includes the image sub-dataset FloW-Img and the multimodal sub-dataset FloW-RI. This study selected FloW-Img, which contains 2,000 images and 5,271 annotated floating garbage. Small objects (area <32px×32px) account for the largest proportion, which increases the difficulty of detection. The DeepTrash dataset is an important resource in the field of marine plastic waste detection. It contains 3,200 images covering marine environments of different quality, depth and visibility. Before training, the image size is resized to a fixed 640px×640px, and the pixel values are normalized to the range of [0,1] to speed up the convergence of the model training. At the same time, data enhancement techniques such as random horizontal flipping, random cropping and color jittering are applied to increase the diversity of training data and prevent model overfitting. To address the high proportion of small objects in FloW-Img and the possible environmental bias in DeepTrash, we oversampled small objects in FloW-Img and adopted a weighted sampling method for DeepTrash, giving higher weights to less common environmental data. This ensures that the model can effectively learn from various types of data and improves its generalization ability.
[0043] Experimental training was conducted using the YOLOv8s model as the base model on an NVIDIA TITAN V GPU, CUDA v11.0.194, and cuDNN v11.0.194. During training, the number of training epochs was set to 200, with an initial learning rate of 0.01 and a final learning rate of 0.0001. A cosine annealing learning rate adjustment strategy was used to balance convergence speed and prevent overfitting. Weight decay was set to 0.0003 to regularize the model and avoid overfitting. Preliminary experiments determined that a batch size of 16 ensured training stability and efficient use of GPU resources.
[0044] The present invention will be described below with reference to the embodiments.
[0045] 1. Data Collection and Preprocessing
[0046] Multiple representative inland water areas were selected, such as different river sections and lake areas. Image collection was conducted at various times of day and under varying weather conditions. Images containing plastic litter were collected using devices equipped with high-definition cameras, such as drones, unmanned surface vehicles, and fixed shore surveillance cameras. While collecting images, a positioning system was used to obtain the geographic location of the image acquisition, a light sensor was used to record light intensity, and water quality testing equipment was used to measure environmental parameters such as water turbidity to construct an initial image dataset. For example, three collection points were selected in the upper, middle, and lower reaches of a particular river. Under varying weather conditions, including sunny, cloudy, and light rainy days, 100 images were collected daily at each collection point, for a total of 2,700 images. These 2,700 images were annotated using professional image annotation tools, such as LabelImg. Annotators carefully observed the outlines of the plastic litter in the images, drew bounding boxes that accurately covered each piece of plastic litter, and labeled the plastic litter according to its shape and characteristics, such as plastic bottles, plastic bags, and plastic fragments. After labeling, the labeled dataset was divided into training, validation, and test sets in a ratio of 80%, 10%, and 10%. During this division process, the distribution of data in different regions and environments was fully considered to ensure representative data in each subset. For example, for areas with large differences in light intensity, appropriate data samples were appropriately allocated to each subset. At the same time, the test set was ensured to contain 270 images, with each image containing at least one plastic waste sample.
[0047] 2. Data augmentation and processing
[0048] Various data augmentation operations were performed on the 2,160 images in the training set. Images were randomly flipped horizontally with a probability of 50% to simulate plastic waste from different perspectives; images were randomly rotated within a certain range of angles to increase image diversity; images were randomly cropped at a certain ratio to simulate the partial occlusion of plastic waste in real-world scenarios; and image brightness, contrast, and saturation were randomly adjusted to enhance the model's adaptability to plastic waste under varying lighting conditions. The number of small plastic waste samples (smaller than 32px × 32px) in the training set was counted and found to account for approximately 20%. Oversampling techniques were used to generate new samples by replicating small object samples or performing local transformations (such as cropping, scaling, and adding a small amount of noise) to increase the proportion of small object samples in the training set to 30%. For data with unique environmental parameters, a weighted sampling method was used to increase the sampling probability of these special environmental data, increasing their proportion in the training set by 10%, enabling the model to better learn the characteristics of plastic waste in these environments.
[0049] 3. Model construction and training
[0050] In the convolutional layer part of the model, the lightweight receptive field coordinate attention convolution operation (LRFCAConv) is used to replace the standard convolution operation. In a specific embodiment, the convolution kernel size k is set to 3. First, the global information of the receptive field features is aggregated through global average pooling (AvgPool), and the input feature map is compressed in the spatial dimension. Next, depth-wise separable convolution (DSC) is used for information interaction. The depth-wise convolution performs a convolution operation on the feature map of each channel, and the point-by-point convolution fuses the results of the depth-wise convolution in the channel dimension. Then, the Softmax function is used to generate attention weights, which reflect the importance of features at different positions. Finally, feature information is extracted through a 3×3 convolution operation with a fixed step size of 1. In this process, the receptive field spatial features are dynamically generated according to the convolution kernel size, which effectively improves the feature extraction capability. A prediction head for detecting 160px×160px feature maps is added to the model head, which is specifically used to detect small plastic waste targets. A coordinate attention (CA) module is added before each prediction head to form a coordinate attention head (CAH). The CA module enables the model to better capture local and global relationships in space by aggregating and recalibrating the features of the feature map in the horizontal and vertical directions. Experimental verification shows that this improvement reduces the number of model parameters by 10% and the computational complexity by 15% while improving the accuracy of small target detection. The SimSPPFCSPC block is constructed based on the SPPCSPC block. The two CBS layers before the pooling layer are trimmed, and only one CBS layer is retained for smoothing features. The SimAM attention mechanism is introduced before the pooling layer to highlight the features of the area where the small target is located by calculating the attention weights of the channel dimension and the spatial dimension. The setting of the pooling kernel is reduced from (5,5,5) to (3,3,3). In the small target detection experiment, the improved SimSPPFCSPC block improves the small target detection accuracy by 8% compared with the original SPPF block.
[0051] YOLOv8s was selected as the base model, with the number of training rounds set to N = 200, the initial learning rate LR1 = 0.01, and the final learning rate LR2 = 0.0001. A cosine annealing learning rate adjustment strategy was used. Weight decay was set to WD = 0.0003, and the batch size was set to BS = 16. The bounding box loss function of the network model was replaced from CIoU to Inner-WIoU. During training, by adjusting relevant parameters of Wise-IoU and Inner-IoU, such as the threshold parameter for dynamically allocating gradient gain in Wise-IoU and the scaling factor for the auxiliary bounding boxes in Inner-IoU, the model's convergence speed was increased by 30%, significantly enhancing generalization.
[0052] The improved YOLOv8 model (ENS-YOLO) was trained using the prepared training set. The training process was performed on a platform equipped with an NVIDIA TITAN V GPU, CUDA v11.0.194, and cuDNN v11.0.194. During training, the training set images were fed into the model, which then predicted the plastic waste in the images based on the current parameters and calculated the Inner-Width-of-Union (IWoU) loss between the predicted results and the ground-truth annotations. The gradient of the loss with respect to each model parameter was calculated using the backpropagation algorithm, and the model parameters were updated using gradient descent. During training, the validation set loss function was continuously monitored. When the validation set loss function did not significantly decrease over 10 consecutive rounds of training, training was stopped and the model weights at that point were saved. After training, the model achieved an mAP50 score of 91.8% on the validation set, laying a solid foundation for subsequent testing and application.
[0053] 4. Model testing and evaluation
[0054] The 270 images in the test set were fed one by one into the trained ENS-YOLO model, which detected plastic waste in the images and obtained information such as its location and category. Model performance was evaluated using metrics such as precision, recall, average precision (AP), mean average precision (mAP), number of model parameters, floating-point operations per second (GFLOPs), and frames per second (FPS). Test results on the FloW dataset showed an mAP50 of 91.8%, an mAP50-95 of 51.7%, 17.45M model parameters, 43.0 GFLOPs, and 120 FPS. Compared to other common object detection models, such as SSD and YOLOv5s, ENS-YOLO has significant advantages in small object detection accuracy and overall detection performance, meeting the accuracy and efficiency requirements for practical applications of inland water plastic waste detection.
[0055] In summary, the inland water plastic waste detection method based on improved YOLOv8 proposed in the present invention has significant advantages in improving detection accuracy, enhancing model adaptability, improving detection efficiency and facilitating practical application through comprehensive data processing, innovative model improvement and reasonable application deployment. It provides an efficient and reliable solution for the monitoring and cleaning of inland water plastic waste, and has important practical significance for protecting the ecological environment of inland waters.
[0056] Although the present invention has been described in detail with reference to specific embodiments, it will be understood by those skilled in the art that various modifications and variations can be made therein without departing from the spirit and scope of the invention.
Claims
1. A method for detecting plastic waste in inland waters based on improved YOLOv8, characterized in that: The YOLOv8 basic framework is used, and the following network modifications are made to it: In the convolutional layer of the model, the lightweight receptive field coordinate attention convolution operation LRFCAConv replaces the standard convolution operation. LRFCAConv first aggregates the global information of the receptive field features through global average pooling (AvgPool), compresses the two-dimensional feature map in the spatial dimension, and obtains a one-dimensional feature vector containing global information. Then, it uses depthwise separable convolution (DSC) for information interaction, decomposing the traditional convolution into depthwise convolution and point-by-point convolution. Then, the Softmax function is used to generate attention weights that reflect the importance of features at different positions. Finally, a fixed-step 3×3 convolution operation is used to extract feature information. Improved prediction heads: The original model had three-scale prediction heads of 80×80px, 40×40px, and 20×20px. The improved model added a 160px×160px prediction head for detecting higher-resolution feature maps, which is used to detect small plastic waste targets. At the same time, a coordinate attention (CA) module was added before each prediction head to form a coordinate attention head (CAH), enabling the model to better capture local and global relationships in space. The SimSPPFCSPC block is constructed based on the SPPCSPC block: First, the two CBS layers before the pooling layer are trimmed, retaining only one CBS layer for feature smoothing. Second, the SimAM attention mechanism is introduced before the pooling layer. The SimAM attention mechanism automatically distinguishes target pixels from other pixels by calculating attention weights in the channel and spatial dimensions, highlighting the features of small target areas and enhancing attention to dense small target areas. Finally, the pooling kernel setting is reduced from (5,5,5) to (3,3,3). The bounding box loss function of the network model is determined to be the Inner-WIoU loss function that combines Wise-IoU and Inner-IoU; Use the prepared training set to train the improved YOLOv8 model.
2. The method for detecting plastic waste in inland waters according to claim 1, characterized in that: The coordinate attention (CA) module can take into account both input feature information and pixel position information. By aggregating and recalibrating the features of the feature map in the horizontal and vertical directions, the model can better capture the local and global relationships in space.
3. The method for detecting plastic waste in inland waters according to claim 1, characterized in that: It also includes data augmentation and processing: randomly flipping the image horizontally to simulate the plastic waste situation under different perspectives; randomly rotating the image within a certain angle range to increase the diversity of the image; randomly cropping the image according to a certain proportion to simulate the situation in actual scenes where plastic waste may be partially blocked; randomly adjusting the brightness, contrast and saturation of the image to adjust and enhance the model's adaptability to plastic waste under different lighting conditions.
4. The method for detecting plastic waste in inland waters according to claim 3, characterized in that: Oversampling technology is used to generate new samples by copying small target sample images or performing local transformations on them, thereby increasing the proportion of small target samples in the training set.
5. The method for detecting plastic waste in inland waters according to claim 1, characterized in that: During the training process of the improved YOLOv8 model, the training set images are input into the model. The model predicts the plastic waste in the image based on the current parameters and calculates the Inner-WIoU loss value between the predicted result and the true annotation. The gradient of the loss value with respect to each model parameter is calculated through the back-propagation algorithm, and the model parameters are updated using the gradient descent method.