Device and method for automatically monitoring garbage floating on water surface
Through the technology combined with drone and deep learning model, efficient and accurate monitoring and cleaning of floating garbage on the water surface is achieved, solving the problems of limited monitoring range and poor identification accuracy in the existing technology, and significantly improving the efficiency and quality of environmental protection work.
Patent Information
- Application Number
- CN202510525635.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has problems with limited monitoring range, poor identification accuracy and sensitivity to environmental factors in monitoring and cleaning of floating garbage on water surfaces.
The drone is equipped with a high-resolution camera, combined with deep learning models (such as DeepLabv3+) for image preprocessing and semantic segmentation, realizing dynamic monitoring of large-area waters and high-precision garbage recognition. At the same time, the cleaning ship is automatically navigated and operated, and the cleaning path and task planning are optimized.
It significantly improves the monitoring and cleaning efficiency and effect of floating garbage on the water surface, achieves efficient coverage and high-precision identification of large-area water areas, and reduces the dependence on environmental factors.
Smart Images

Figure CN120047756A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of environmental monitoring, and particularly to an automatic monitoring device and method for floating garbage on the water surface. Background Art
[0002] With the increasing severity of global environmental problems, water pollution, especially the accumulation of floating garbage on the water surface, has become one of the major environmental challenges faced around the world. The floating garbage on the water surface not only affects the water landscape and ecological environment, but also may pose a threat to aquatic organisms and human health. Traditional manual cleaning methods are inefficient and difficult to implement in large areas of water or complex environments. Therefore, developing a technology that can automatically detect, identify, and clean floating garbage on the water surface has become an urgent need in the field of environmental protection.
[0003] In response to these problems, some improvements have been made in the prior art, but there are still certain limitations. For example, the patent with the publication number CN111950357A discloses a method for quickly identifying ship-borne water surface garbage based on multi-feature YOLOV3, which monitors the floating garbage on the water surface through a camera installed on the cleaning ship. Although this method can provide certain monitoring capabilities, there are still many limitations: 1. Limited monitoring range: This method mainly relies on the camera fixed on the cleaning ship, and its monitoring ability in large areas of water is limited. The fixed camera can only cover a limited monitoring area and it is difficult to effectively identify small floating garbage. 2. Poor recognition accuracy: It is greatly affected by weather conditions. Environmental factors such as water surface haze and raindrops will interfere with the image quality and affect the accuracy of garbage recognition. At the same time, there are various types of floating garbage on the water surface, with different shapes and colors, and how to effectively extract and classify garbage features is also a difficult point.
[0004] In order to overcome the above deficiencies, the present invention proposes an automatic monitoring method for floating garbage on the water surface, which combines an unmanned aerial vehicle, a deep learning model, and an automatic cleaning ship. Through efficient image processing and analysis technologies, it realizes the accurate recognition and classification of water surface garbage, and automatically generates cleaning tasks and optimizes cleaning paths based on this, guiding the cleaning ship to perform autonomous navigation and operation, thereby greatly improving the cleaning efficiency and effect of floating garbage on the water surface. Summary of the Invention
[0005] (I) Technical Problems to be Solved In view of the deficiencies of the prior art, the purpose of the present invention is to provide an automatic monitoring device and method for floating garbage on the water surface, which solves the problems existing in the prior art. The method of this patent uses a drone equipped with a high-resolution camera to achieve dynamic monitoring of large areas of water. The drone can fly flexibly, cover a vast water area, and collect images at different angles and heights, thus overcoming the limitations of fixed cameras and satellite remote sensing in terms of monitoring range. And by advanced image preprocessing techniques (such as noise removal, fog removal, and rain removal) to improve the image quality, combined with the DeepLabv3+ model for semantic segmentation, greatly improving the accuracy of garbage recognition. This method is specifically optimized for complex environmental conditions and can maintain high-efficiency garbage recognition capabilities in situations such as light changes, haze, and rainy days.
[0006] (2) Technical solution To achieve the above object, the present invention provides the following technical solution: An automatic monitoring method for floating garbage on the water surface, comprising the following steps: Data collection: The host controls the drone equipped with a high-resolution camera to collect and store the data of the target water area image; Image preprocessing: The host preprocesses the collected images, including noise removal, fog removal, rain removal, and image enhancement; Feature extraction: The host inputs the preprocessed image data into the DeepLabv3+ model, extracts the basic features in the image through the backbone network, including edges, textures, shapes, and object features, expands the receptive field of the convolutional kernel through dilated convolution, performs feature fusion and upsampling, combines multi-scale features, and restores the image resolution through upsampling to generate a high-precision segmentation mask; Model training: Input training data into the DeepLabv3+ model in the host, transform the training data, use the cross-entropy loss method to calculate the difference between the model prediction and the true annotation, adjust the model parameters with the Adam optimizer, and train on the dataset; Garbage recognition and classification: Parse the semantic segmentation mask in the host, classify each pixel as garbage or non-garbage; use the connected component labeling algorithm to merge adjacent garbage pixels into continuous garbage regions and assign a unique label to each independent garbage region; based on the features extracted in step feature extraction, use a classifier to classify different types of garbage, and count the identified garbage regions to generate the final recognition result, including the type, location, quantity, and area information of each type of garbage; Task planning: After receiving the data, the host parses the structured data, extracts the garbage type, location, quantity, and area information; according to the location and quantity of the garbage, uses a path planning algorithm to generate the optimal route for the cleaning task; Garbage cleaning: The host sends the generated cleaning tasks and routes to the garbage cleaning ship. After receiving the task data, the cleaning ship automatically navigates to the garbage area according to the planned route and activates the cleaning device to clean the floating garbage on the water surface. During the execution of the task, the cleaning ship feeds back the cleaning progress and status to the host in real time.
[0007] Preferably, in the image preprocessing step, the dehazing process includes: calculating the dark channel image, selecting the maximum value of the pixel values in the original image I corresponding to the top 0.1% brightest pixel points in the dark channel image as the atmospheric light A, calculating the transmission rate t based on the dark channel image and the atmospheric light, and using the estimated atmospheric light and transmission rate to restore the haze-free image J.
[0008] Preferably, in the image preprocessing step, the rainfall removal includes: Motion detection: Identifying the positions and movement trajectories of raindrops in the image; Raindrop feature extraction: Using morphological operations to extract raindrop features and identify raindrop positions, where the morphological operations include opening and closing operations; Raindrop removal: After identifying the raindrop positions, the raindrop areas are patched through image inpainting technology to restore the original appearance of the image.
[0009] Preferably, in the feature extraction step, the backbone network extracts features including the following steps: Input image preprocessing: Resizing and normalizing the input image; Feature extraction by convolutional layers: Extracting low-level features through early convolutional layers, including: edge and texture features; extracting more advanced features through middle convolutional layers, including: object shape and contour features; extracting high-level abstract features through late convolutional layers, including: complex object and scene information; Output: The backbone network outputs a feature map.
[0010] Preferably, in the feature extraction step, the dilated convolution includes the following steps: Multi-scale dilated convolution: Using dilated convolutions with different dilation rates in the convolutional layers of the backbone network to form multi-scale feature representations and capture local and global context information; Feature pyramid: Combining the output features through dilated convolutions with different dilation rates to form a comprehensive multi-scale feature map; Output: The ASPP module outputs a high-level feature map containing multi-scale context information; Feature fusion: Combining the multi-scale features and low-level features output by the encoder, retaining high-level semantic information and low-level detail information; Upsampling: Gradually enlarging the feature map to the same resolution as the input image through bilinear interpolation to generate a semantic segmentation mask, where each pixel represents garbage or non-garbage.
[0011] Preferably, in the model training step, the expression of the cross-entropy loss function is as follows: where N is the total number of pixels, C is the number of classes, is the true annotation of the i-th pixel in the c-th class, is the predicted probability of the i-th pixel in the c-th class.
[0012] Preferably, the training process in the model training step includes: Data loading: Use a data loader to read the training data and annotation data and perform batch processing; Forward propagation: Input the training image into the DeepLabv3+ model, and generate a segmentation mask through the backbone network, dilated convolution, and decoder; Loss calculation: Use the cross-entropy loss function to calculate the difference between the model prediction result and the true annotation. Traverse all pixels, calculate the cross-entropy between the predicted value and the true value of each pixel, and sum them to obtain an overall loss value; Backward propagation: Through the backward propagation algorithm, calculate the gradient of the loss function with respect to the model parameters, and pass the calculated gradient information back layer by layer to update the model parameters; Parameter optimization: Adopt the Adam optimizer and use the optimization algorithm to update the model parameters, so that the model gradually reduces the loss value in each iteration and improves the model performance; Iterative training: Perform multiple iterations on the entire training dataset until the number of iterations exceeds the set threshold S. Each iteration includes several processes of forward propagation, loss calculation, backward propagation, and parameter optimization; Model saving: Save the trained model parameters for subsequent use and further optimization.
[0013] Preferably, the connected component labeling algorithm in the garbage recognition and classification step includes: Select a starting point: Select an unlabeled garbage pixel as the starting point; Check adjacent pixels: Check the pixels around the starting point. If these pixels belong to garbage, label them with the same label as the starting point; Recursive labeling: Repeat the step of checking adjacent pixels for the newly labeled pixels until all adjacent garbage pixels are labeled; Repeat the process: Select the next unlabeled garbage pixel as the new starting point and repeat the above process until all garbage pixels are labeled.
[0014] Preferably, using a classifier to classify different types of garbage in the garbage recognition and classification step includes: Feature integration: Integrate the color features, texture features, and shape features of each garbage area into a feature vector; Classification prediction: Input the feature vector of the new garbage area into the trained classifier, and the classifier will assign each garbage area to a specific category according to the learned patterns; Classification statistics: Count the quantity of each type of garbage and calculate the total area of each type of garbage; Analysis of statistical results: Generate a final report, summarize the classification and statistical results, and generate a detailed report containing information such as the quantity, total area, and distribution of each type of garbage.
[0015] An automatic monitoring device for floating garbage on the water surface, comprising: Drone: It is used to fly above the water surface, cruise according to a preset path, and carry a high-resolution camera; High-resolution camera: It is used to obtain high-definition images of floating garbage on the water surface; Host: It is used to receive and process the image data obtained from the high-resolution camera; Cleaning ship: It is used to collect and process floating garbage on the water surface.
[0016] (III) Beneficial effects The purpose of the present invention is to provide an automatic monitoring device and method for floating garbage on the water surface, which has significant beneficial effects. First, by carrying a high-resolution camera on a drone, it can quickly and comprehensively cover a large area of water, collect high-quality image data, and avoid the limitations of low efficiency and environmental factors in the traditional manual inspection method. Second, advanced image preprocessing techniques, including noise reduction, gray-scale transformation, dehazing, rain removal, and image enhancement, are adopted to effectively improve the image quality and reduce the interference of environmental factors such as water surface reflection, haze, and raindrops on image recognition. Combining all steps, the present invention significantly improves the automation and intelligence level of floating garbage monitoring and cleaning on the water surface, has significant advantages such as high efficiency, accuracy, wide coverage, and no interference from environmental factors, greatly improves the efficiency and quality of environmental protection work, and provides strong technical support for the continuous monitoring and treatment of water area environment. Description of the drawings
[0017] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference symbols are used to represent the same components. In the drawings: Figure 1 is the overall flowchart of an automatic monitoring method for floating garbage on the water surface in an embodiment of the present application.
[0018] Figure 2 This is the flowchart of image preprocessing in an automatic monitoring method for floating garbage on the water surface in an embodiment of the present application.
[0019] Figure 3 This is the flowchart of feature extraction in an automatic monitoring method for floating garbage on the water surface in an embodiment of the present application.
[0020] Figure 4 This is the flowchart of model training in an automatic monitoring method for floating garbage on the water surface in an embodiment of the present application.
[0021] Figure 5 This is the flowchart of garbage classification and recognition in an automatic monitoring method for floating garbage on the water surface in an embodiment of the present application. Detailed implementation manners
[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the appended Figures 1 - 5 For a clear and complete description of the technical solutions in the embodiments of the present invention, it is obvious that the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0023] An automatic monitoring device for floating garbage on the water surface includes: An unmanned aerial vehicle (UAV): responsible for cruising above the water surface. It can fly along a preset path to cover a specified water area. The flexibility of the UAV enables it to quickly reach the target area for real-time monitoring, especially suitable for large water areas or areas that are difficult to access. It performs a cruising task above the water surface according to the preset flight path to ensure comprehensive monitoring of the water area situation.
[0024] A high-resolution camera: used to capture high-definition images of floating garbage on the water surface. The camera captures images on the water surface in real time to ensure high clarity and can clearly distinguish the floating objects on the water surface. The acquired image data is transmitted to the host in real time for subsequent processing and analysis.
[0025] A host: The host is the data processing center of the system, responsible for receiving, storing, and processing the image data from the UAV camera. The host is built-in with image processing algorithms and classification models for analyzing and identifying the types and positions of garbage on the water surface.
[0026] A cleaning ship: The cleaning ship is a device for performing cleaning tasks, responsible for collecting and processing the floating garbage identified by the host from the water surface. The cleaning ship automatically goes to the identified garbage area according to the instructions of the host and uses tools such as robotic arms and nets to collect the garbage onto the ship.
[0027] This system realizes the efficient and automatic monitoring of floating garbage on the water surface through the combination of drones and high-resolution cameras. The host, as the data processing center, is responsible for analyzing and identifying the types of garbage, and then commanding the cleaning ship to conduct precise cleaning. The overall goal of the system is to improve the monitoring and cleaning efficiency of the water area environment, reduce the need for manual intervention, and enhance the intelligent level of environmental protection.
[0028] An automatic monitoring method for floating garbage on the water surface includes the following steps: I. Data collection: Drone flight planning: Design the flight path of the drone to ensure coverage of the entire target water area. The drone is equipped with a high-resolution camera to capture clear images of the water surface.
[0029] Image capture: Regularly capture images along the set flight path to ensure coverage of all monitoring areas.
[0030] II. Image preprocessing 1. Denoising: Use a filter (Gaussian filter) to remove random noise in the image and improve the initial clarity of the image.
[0031] 2. Dehazing processing: The dehazing processing steps aim to weaken or eliminate the impact of haze on the image through a series of image processing techniques, thereby restoring the clarity and details of the image. The following are the specific steps: a. Calculate the dark channel image: For the input image I, within a local window ω(x), calculate the minimum value of each pixel, and then take the minimum value among the three color channels to obtain the dark channel image , and the expression is: represents for each pixel x in the input image I, within its local window , take the minimum value of the R, G, and B color channels of all pixels within this window to obtain the dark channel image .
[0032] b. Estimate the atmospheric light A: Select the maximum value of the pixel values in the original image I corresponding to the top 0.1% brightest pixel points in the dark channel image as the atmospheric light A.
[0033] c. Estimate the transmittance t: Calculate the transmittance t based on the dark channel image and the atmospheric light.
[0034] where is the transmittance at pixel x, is the adjustment parameter, with a value of 0.95 here, and A is the atmospheric light value.
[0035] d. Restore the fog-free image: Using the estimated atmospheric light and transmittance, restore the fog-free image J.
[0036] where is the pixel value of the restored fog-free image.
[0037] 3. Raindrop removal: By detecting, identifying, and repairing the areas in the image affected by raindrops, the clarity of the image is restored.
[0038] a. Motion detection: Identify the positions and motion trajectories of raindrops in the image. See the following steps for details: ① Obtain consecutive frames: Obtain consecutive image frames from the video stream.
[0039] ② Calculate the difference image: Perform pixel difference on adjacent frames to obtain the difference image.
[0040] represents calculating the absolute value of the gray value difference of the same pixel and in two consecutive image frames to obtain the difference image .
[0041] ③ Threshold processing: Perform threshold processing on the difference image to mark the regions with large changes as motion regions.
[0042] where T is the set threshold (the critical value for distinguishing motion regions from the background), is the binarization result (1 for motion regions, 0 for the background).
[0043] ④ Raindrop motion trajectory: Track the motion trajectory of raindrops through the differences of consecutive frames.
[0044] b. Raindrop feature extraction: Use morphological operations (opening operation, closing operation) to extract raindrop features and identify raindrop positions.
[0045] ① Opening operation: Use a structuring element to perform an opening operation on the binarized motion detection result to remove small noises. The expression is: where M is the binarized motion region mask, S is the structuring element, represents the opening operation.
[0046] ② Closing operation: Perform a closing operation on the result of the opening operation to fill the small holes in the raindrop regions. The expression is: where Indicates a closing operation, which smooths the boundary of the raindrop region through the structuring element S.
[0047] ③ Raindrop recognition: After morphological operations, a clear raindrop region is obtained.
[0048] c. Raindrop removal: After identifying the raindrop positions, the raindrop regions are patched through image inpainting technology to restore the original appearance of the image.
[0049] Among them, the image inpainting steps include: Image interpolation: Interpolate the pixels in the raindrop region, and use the pixel values in the surrounding non - raindrop regions to fill the raindrop region.
[0050] Linear interpolation: Use the linear combination of surrounding pixels to calculate the pixel values in the raindrop region.
[0051] Inpainting through the PatchMatch algorithm: ① Initial matching: Randomly select a set of matching patches in the whole image for initial inpainting.
[0052] ② Propagation step: For each pixel, check the matching situation of its neighboring pixels and update the best match.
[0053] ③ Random search: During the matching process, randomly select some pixel patches for matching to find the optimal matching patch.
[0054] ④ Inpaint the raindrop region: Replace the pixel values in the raindrop region with the pixel values of the best matching patch to restore the integrity of the image.
[0055] 4. Image enhancement: The host enhances the details in the image by adjusting the contrast, brightness and other parameters of the image to ensure that the garbage regions are more obvious. It includes the following steps: a. Contrast adjustment: By histogram equalization, redistribute the image gray values so that the gray value distribution is more uniform, thereby enhancing the overall contrast and clarity of the image. It includes the following steps: ① Calculate the gray histogram: Count the number of pixels at each gray level in the image to obtain the gray histogram.
[0056] H(i)=Number of pixels with gray level i ② Calculate the cumulative distribution function (CDF): According to the gray histogram, calculate the cumulative distribution function.
[0057] Among them, j is a variable used to traverse the gray levels. It starts from 0 and increases sequentially to the current gray level i for which the cumulative distribution is to be calculated. Indicates the number of pixels with gray level j. Represents the cumulative sum of the number of all pixels from gray level 0 to gray level i, reflecting the cumulative situation of pixels with gray levels less than or equal to i in the image.
[0058] ③ Normalized cumulative distribution function: Normalize the cumulative distribution function to the range of gray values (0 to 255).
[0059] Where N is the total number of pixels in the image.
[0060] ④ Remap gray values: According to the normalized cumulative distribution function, remap the gray values of the original image to obtain the equalized image.
[0061] Where is the new gray value of the pixel at coordinates after histogram equalization.
[0062] b. Sharpening processing: Increase the contrast of edges in the image through the Laplacian filter to further enhance the details of the image. The sharpening processing makes the contours of the garbage areas clearer.
[0063] Through these enhancement steps, the garbage areas in the image will become more obvious and the details will be clearer, which helps to improve the recognition and classification accuracy of the entire system.
[0064] III. Feature extraction: Aims to extract various features from the preprocessed image that can effectively distinguish between garbage and non-garbage areas, including edge, texture, shape and other information.
[0065] Input the preprocessed image data into the DeepLabv3+ model. In the DeepLabv3+ model, feature extraction is mainly achieved through the backbone network and atrous convolution.
[0066] Backbone network: The backbone network is the basic part of the entire model, responsible for extracting multi-level features from the input image. Use the pre-trained network ResNet-101 on a large-scale dataset to obtain better generalization ability and feature extraction effect. The specific steps include: 1. Input image preprocessing: Image size adjustment: Adjust the image to the standard size of 224x224 to ensure that the input image meets the model requirements.
[0067] Normalization: Normalize the pixel values of the image and scale the pixel values to the range of [0, 1].
[0068] Color space conversion: Convert the image from the RGB space to the HSV or Lab space. This is done to reduce the impact of lighting changes and highlight color information. Especially for garbage types with obvious color contrasts (such as plastic bottles, paper scraps, etc.), it can help the model better perform garbage classification.
[0069] 2. Convolutional layer for feature extraction: a. Early convolutional layer (low-level features), which is used for color feature extraction, such as color distribution, saturation, etc. information in the garbage area. A small convolutional kernel (3×3) is used to extract low-level features from the input image. The convolution operation formula is: Among them, is the convolutional kernel, is the normalized input image.
[0070] b. Intermediate convolutional layer (mid-level features), which is used for texture feature extraction, such as the detailed differences between the garbage surface and the background. The convolutional layer can capture the texture changes of the image through local feature extraction. A larger-sized convolutional kernel (5×5) is used to extract more complex mid-level features from the image. The convolution operation formula: Among them, is the intermediate layer convolutional kernel, is the output after convolution of the previous layer.
[0071] c. Late convolutional layer (high-level features), which is used for shape feature extraction. High-level feature maps containing information such as the contours, boundaries, and shapes of objects are extracted through the late convolutional layer (7×7 convolutional kernel): Among them, is the convolutional kernel of the late convolutional layer.
[0072] 3. Output: The backbone network finally outputs a low-resolution but feature map with rich semantic information for subsequent dilated convolution and feature fusion.
[0073] Furthermore, pooling layers are connected after the early and intermediate convolutional layers to reduce the spatial resolution of the feature map while retaining important feature information.
[0074] Dilated convolution: Dilated convolution expands the receptive field by inserting holes (i.e., increasing the spacing between convolutional kernels) in the convolutional kernel to capture a larger range of context information without increasing the computational amount. It includes the following specific steps: 1. Multi-scale dilated convolution Convolution process: In the convolutional layer of the backbone network, convolutional kernels with different dilation rates (e.g., d = 1, 2, 3) are used to capture features at different scales. The operation of dilated convolution can be expressed as: where d is the dilation rate, is the convolutional kernel of dilated convolution.
[0075] 2. Atrous Spatial Pyramid Pooling (ASPP) Multi-scale feature fusion: The Atrous Spatial Pyramid Pooling module (ASPP) combines the convolutional outputs with different dilation rates and fuses them into a comprehensive multi-scale feature map. The formula is as follows: where concat represents concatenating the convolutional outputs with different dilation rates.
[0076] 3. Output high-level feature map: The ASPP module fuses these multi-scale feature maps and outputs a high-level feature map containing rich context information: 4. Feature fusion: Fuse the multi-scale features output by the encoder with the low-level features to ensure that both detailed information and high-level semantic information are retained: where, is the low-level feature, is the high-level feature.
[0077] 5. Upsampling: Restore the feature map to the same resolution as the input image through bilinear interpolation. Assuming the upsampling factor is α, the upsampling operation formula is: where, is the fused feature map, α is the upsampling factor, usually the ratio of the input image size to the feature map size.
[0078] 6. Output semantic segmentation mask Finally, the model generates a semantic segmentation mask through per-pixel classification. For each pixel point (x, y), the model predicts its class (garbage or non-garbage).
[0079] Through the above detailed steps, the DeepLabv3+ model can efficiently extract multi-scale features from the input image, capture rich context information, and generate a high-precision segmentation mask, providing strong support for the recognition of floating garbage on the water surface.
[0080] IV. Model training, including the following steps: 1. Dataset Construction: Prepare a labeled training dataset, including a large number of images and their corresponding semantic segmentation masks, ensuring that the dataset has sufficient diversity.
[0081] 2. Data Augmentation: Increase the diversity of data by performing various transformations on the training data, improve the generalization ability of the model, and avoid overfitting.
[0082] Rotation: Randomly rotate the image by a certain angle (e.g., ±15 degrees) to increase the directional diversity of the image.
[0083] Flipping: Randomly flip the image horizontally or vertically to increase symmetry.
[0084] Cropping: Randomly crop a part of the image and then resize it to a fixed size to simulate different viewing perspectives.
[0085] Scaling: Randomly scale the image to increase information at different scales.
[0086] Color Adjustment: Randomly adjust the brightness, contrast, saturation, etc. of the image to increase the diversity of lighting conditions.
[0087] Noise Addition: Add random noise to the image to enhance the noise resistance of the model.
[0088] 3. Data Loading: Create a data loader to load the augmented training data and annotation data into the model in batches. This step includes data preprocessing, such as image size adjustment and normalization. By setting a reasonable batch size, in each training iteration, the data is divided into several batches for processing to improve training efficiency and model stability.
[0089] 4. Forward Propagation: Input a batch of training images into the DeepLabv3+ model. The model generates predicted semantic segmentation masks through a series of convolutional operations, dilated convolutions, feature fusion, etc. Through the encoder and decoder networks of the model, image features are extracted and the segmentation results of the predicted garbage and non-garbage regions are output.
[0090] 5. Loss Calculation Loss Function Selection: Use the cross-entropy loss function to calculate the difference between the model prediction results and the true annotations. This function can measure the matching degree between the predicted class of each pixel and the true class. Its expression is: where N is the total number of pixels, C is the number of classes, is the true annotation of the i-th pixel in the c-th class, is the predicted probability of the i-th pixel in the c-th class.
[0091] Loss value calculation: Traverse all pixels, calculate the cross-entropy between the predicted value and the true value of each pixel, and sum them up to obtain an overall loss value.
[0092] 6. Backpropagation Gradient calculation: Through the backpropagation algorithm, calculate the gradient of the loss function with respect to the model parameters to clarify the adjustment direction of the model parameters.
[0093] Gradient propagation: Pass the calculated gradient information back layer by layer to update the model parameters.
[0094] 7. Parameter optimization Optimization algorithm selection: Use the Adam optimizer, which combines the advantages of the momentum method and RMSProp, and can dynamically adjust the learning rate according to the first-order and second-order momentum estimates to optimize the model parameters.
[0095] Adam optimizer (Adaptive Moment Estimation): Combines the momentum method and RMSProp, with the ability to adaptively adjust the learning rate. Update formula: where is the learning rate, is the first-order moment estimate of the gradient, that is, the mean of the gradient, is the second-order moment estimate of the gradient, that is, the mean of the square of the gradient. The role is to prevent the denominator from being zero and ensure the numerical stability of the formula calculation.
[0096] Parameter update: Use the optimization algorithm to update the model parameters, so that the model gradually reduces the loss value in each iteration and improves the model performance.
[0097] 8. Iterative training Multiple iterations: Conduct multiple iterations on the entire training dataset until the number of iterations exceeds the set threshold S. Each iteration includes several processes of forward propagation, loss calculation, backpropagation, and parameter optimization.
[0098] Validation set evaluation: After each iteration, use the validation set to evaluate the performance of the model, record the loss value and accuracy on the validation set to detect whether the model is overfitting or underfitting.
[0099] Hyperparameter tuning: According to the performance of the validation set, adjust the hyperparameters (such as learning rate, batch size, etc.) to further improve the performance of the model.
[0100] 9. Model saving Regular saving: During the training process, regularly save the parameters and configurations of the model so that training can continue in case of training interruption or for subsequent optimization and application.
[0101] Optimal Model Selection: Based on the performance on the validation set, save the optimal model parameters to ensure that the finally deployed model can achieve the best effect in actual applications.
[0102] Through the above detailed steps, the DeepLabv3+ model can be effectively trained to improve its segmentation accuracy and recognition performance in the task of identifying floating garbage on the water surface.
[0103] V. Garbage Recognition and Classification: Analyze the semantic segmentation mask extracted in the feature extraction step, classify each pixel as garbage or non-garbage. Then merge adjacent garbage pixels into continuous regions to identify each garbage block. Finally, classify and count according to the extracted features. The specific steps are as follows: 1. Analyze the semantic segmentation mask: In the feature extraction stage, the DeepLabv3+ model has generated a semantic segmentation mask through techniques such as multi-layer convolution and dilated convolution. The value of each pixel in the mask represents the probability that the pixel belongs to a certain class (garbage or non-garbage).
[0104] Mask Parsing: Extract the classification result of each pixel in the semantic segmentation mask to determine whether each pixel is garbage or non-garbage: where represents selecting the class (garbage or non-garbage) corresponding to the maximum value from multiple channels of the feature map at each pixel position.
[0105] 2. Region Merging: The main purpose of region merging is to merge adjacent garbage pixels into complete garbage regions by analyzing the connected regions in the image. This process mainly relies on the Flood Fill algorithm to identify and merge adjacent garbage pixel regions. In this way, the scattered garbage pixels are classified into individual garbage regions (blocks). The specific steps are as follows: a. Select the starting point: Select an unlabeled garbage pixel as the starting point. In the image, each pixel may belong to garbage or non-garbage. Initially, the pixels of the garbage region are not labeled. Start from an unlabeled garbage pixel as the starting point of the Flood Fill algorithm.
[0106] b. Check adjacent pixels: Starting from the current starting point check the neighboring pixels around this pixel. Specifically, check the adjacent pixels in four directions (up, down, left, right). If the values of these adjacent pixels are 1 (indicating garbage), mark them with the same region label as the starting point. In this way, all adjacent garbage pixels can be extended.
[0107] c. Recursive labeling: All adjacent garbage pixels that meet the criteria are labeled recursively or iteratively. For example, if an adjacent pixel also belongs to the garbage and has not been labeled yet, then this pixel is further labeled.
[0108] This labeling method recurses continuously until there are no more adjacent garbage pixels to be labeled, ensuring that all garbage pixels in the same area are assigned the same label.
[0109] d. Repeating the process: After completing one Flood Fill labeling, select the next unlabeled garbage pixel as the new starting point and repeat the above process. In this way, the Flood Fill algorithm traverses the entire image and merges all garbage pixels into several connected regions.
[0110] e. Label assignment: When the Flood Fill algorithm finishes labeling a region, a merged garbage region is obtained and this region is assigned a unique label , indicating that this region belongs to the same garbage block. The specific steps are as follows: ① Assign a unique label to each region: In the Flood Fill algorithm, whenever a new garbage region is identified, a unique label is assigned to this region , and this label value is an integer representing the unique identifier of this garbage region.
[0111] For example, assume that the first garbage region we label is region 1, the second is region 2, and so on.
[0112] Among them, unique_id is the unique identifier assigned to each garbage region. Each garbage region will have an independent label for subsequent feature extraction, classification, statistics, etc.
[0113] ② Store the labels: After the labeling process is completed, we store the label of each garbage region and the corresponding pixel positions in a list or data structure. The label of each garbage region indicates that all garbage pixels in this region share the same label.
[0114] For example, assume there are several garbage regions in the image, and each region corresponds to a unique label, which can be represented as follows: Each represents the pixel position of the i-th garbage region in the image corresponding to the label.
[0115] ③Merge all adjacent garbage pixels: By repeatedly performing the above marking operation, all garbage pixels will be merged into several independent regions, each with a unique label. In this way, we can distinguish different garbage regions by labels, facilitating subsequent processing such as feature extraction, classification, and statistics.
[0116] 3. Classification and statistics: The goal of garbage classification is to accurately classify garbage based on the extracted features (color, texture, shape). The goal of statistics is to summarize the classification results and calculate the quantity and total area of each garbage type.
[0117] a. Classification: The goal of classification is to accurately classify different types of garbage based on the extracted features such as color, texture, and shape. Use machine learning algorithms to train the classification model. The specific steps of using support vector machine (SVM) for garbage classification.
[0118] ①Feature vector preparation: First, extract the features for each garbage region and integrate these features into a feature vector. Each feature vector includes: Color features: Such as the average color (R, G, B values) of the pixels within the region and the color histogram.
[0119] Texture features: Such as features describing texture like gray-level co-occurrence matrix (GLCM), calculating contrast, homogeneity, etc.
[0120] Shape features: Such as the area, perimeter, aspect ratio, bounding box, etc. of the region.
[0121] These features are integrated into a complete feature vector: Among them, is the color feature, is the texture feature, is the shape feature.
[0122] ②Classifier training: Use machine learning methods (support vector machine SVM) to train the classifier. First, a set of sample data of garbage regions with known categories needs to be prepared, including the feature vectors of each garbage region and their corresponding labels (such as plastic, paper, wood, etc.).
[0123] Data collection: Prepare a labeled dataset of garbage regions, and each data point includes a feature vector and the corresponding label.
[0124] Train the classifier: Use this data to train the support vector machine (SVM) classifier. SVM classifies the sample points in the feature space and finds an optimal hyperplane to separate the sample points of different categories.
[0125] The core of the training process is to optimize the decision boundary of SVM: Among them, is the normal vector of the decision hyperplane, is the bias, is the slack variable, is the label, is the feature vector.
[0126] ③ Classification prediction Classifier application: Input the feature vectors of new garbage areas into the trained classifier. The classifier will assign each garbage area to a specific category according to the learned pattern.
[0127] Result output: Each garbage area will be labeled with a category, such as "plastic", "wood", or "paper", etc.
[0128] b. Statistics The goal of statistics is to summarize and analyze the classification results, including calculating information such as the quantity and total area of each type of garbage.
[0129] ① Classification statistics Quantity statistics: Count the quantity of each type of garbage. For example, count the number of garbage areas classified as "plastic". The expression is: Among them, is the indicator function, indicating whether the garbage area i belongs to a specific type.
[0130] Area statistics: Calculate the total area of each type of garbage. For each classification label, accumulate the areas of the corresponding garbage areas: Among them, is the area of the garbage area i.
[0131] ② Statistical result analysis and output of statistical results: Generate a report containing the following content: The quantity of each garbage type. The total area of each garbage type. The distribution and proportion of the garbage ③ Result visualization: Use bar charts or pie charts, etc., to visually display the quantity and area proportion of each garbage type. Visualization helps to quickly understand the garbage distribution.
[0132] ④ Output report: Generate a final report, summarize the classification and statistical results, and generate a detailed report containing information such as the quantity, total area, and distribution of each garbage type. This report can be used for environmental monitoring analysis and decision support.
[0133] Through the above detailed steps, the efficient identification and classification of floating garbage on the water surface can be achieved, improving the automation and intelligence level of environmental monitoring.
[0134] VI. Task Planning: After the host receives the data, it parses the structured data and extracts information such as the type, location, quantity, and area of the garbage. Based on the location and quantity of the garbage, a route plan for the cleaning task is generated. The route plan can consider multiple factors, such as garbage concentration, distance, and the capacity of the cleaning ship. The optimal route is generated using a path planning algorithm (Dijkstra algorithm).
[0135] VII. Garbage Cleaning: The host sends the generated cleaning task and route to the garbage cleaning ship. The task data is transmitted to the control system of the cleaning ship through a wireless communication module (such as Wi-Fi or a dedicated communication link). After receiving the task data, the cleaning ship automatically navigates to the garbage area according to the planned route. Precise positioning and navigation are carried out using GPS and an inertial navigation system (INS). After reaching the target area, the cleaning ship activates the cleaning device (such as a robotic arm or a suction device) to clean the floating garbage on the water surface.
[0136] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for automatically monitoring floating garbage on a water surface, characterized in that: The following steps are involved: Data collection: The host controls the drone to collect and store data of the target water area image; Image preprocessing: The host preprocesses the collected images, including noise removal, defogging, rain removal and image enhancement; Feature extraction: The host inputs the preprocessed image data into the DeepLabv3+ model, extracts edge, texture, and shape features in the image through the backbone network, and generates a segmentation mask through dilated convolution; Model training: Input training data to the DeepLabv3+ model, use the cross entropy loss method to calculate the difference between the model prediction and the true annotation, use the Adam optimizer to adjust the model parameters and perform training; Garbage identification and classification: parse the semantic segmentation mask, use the connected component labeling algorithm to merge adjacent garbage pixels into continuous garbage areas, use the classifier to classify the garbage, and generate identification results; Task planning: The host generates the optimal route for the cleaning task; Garbage cleaning: The host controls the cleaning device to clean up floating garbage on the water surface.
2. The method for automatically monitoring floating garbage on a water surface according to claim 1, characterized in that: In the image preprocessing step, the defogging process includes: calculating the dark channel image , select the maximum pixel value in the original image I corresponding to the first 0.1% brightest pixels in the dark channel image as the atmospheric light A, calculate the transmittance t based on the dark channel image and the atmospheric light, and use the estimated atmospheric light and transmittance to restore the fog-free image J.
3. The method for automatically monitoring floating garbage on a water surface according to claim 1, characterized in that: In the image preprocessing step, rain removal includes: Motion detection: Identify the location and movement of raindrops in the image; Raindrop feature extraction: extract raindrop features and identify raindrop positions using morphological operations, including opening and closing operations; Raindrop removal: After identifying the location of raindrops, the raindrop area is repaired through image repair technology to restore the original appearance of the image.
4. The method for automatically monitoring floating garbage on a water surface according to claim 1, characterized in that: In the feature extraction step, the backbone network extracts features including the following steps: Input image preprocessing: resize and normalize the input image; Convolutional layers extract features: shape features through early convolutional layers; texture features through intermediate convolutional layers; shape features through late convolutional layers; Output: backbone network output feature map.
5. The method for automatically monitoring floating garbage on a water surface according to claim 1, characterized in that: In the feature extraction step, the dilated convolution includes the following steps: Multi-scale dilated convolution: dilated convolutions with different dilation rates are used in the convolutional layers of the backbone network to form multi-scale feature representations and capture local and global context information; Feature pyramid: Combine output features through dilated convolutions with different dilation rates to form a comprehensive multi-scale feature map; Output: The ASPP module outputs a high-level feature map containing multi-scale contextual information; Feature fusion: Combine the multi-scale features and low-level features output by the encoder to retain high-level semantic information and low-level detail information; Upsampling: restore the feature map to the same resolution as the input image through bilinear interpolation; Output Semantic Segmentation Mask: A semantic segmentation mask is generated by pixel-by-pixel classification, where each pixel represents garbage or non-garbage.
6. The method for automatically monitoring floating garbage on a water surface according to claim 1, characterized in that: In the model training step, the expression of the cross entropy loss function is: In the formula, N is the total number of pixels, C is the number of categories, is the true label of the i-th pixel in the c-th category, is the predicted probability of the i-th pixel in the c-th class.
7. The method for automatically monitoring floating garbage on a water surface according to claim 1, characterized in that: The model training adopts the DeepLabv3+ model, and the training steps include data loading, forward propagation, loss calculation, back propagation, parameter optimization, iterative training and model saving.
8. The method for automatically monitoring floating garbage on a water surface according to claim 1, characterized in that: The connected component labeling algorithm in the garbage identification and classification step includes: Select starting point: select an unlabeled garbage pixel as the starting point; Check neighboring pixels: Check the pixels around the starting point, if these pixels are garbage, mark them with the same label as the starting point; Recursive labeling: Repeat the steps to check neighboring pixels for the newly labeled pixel until all neighboring junk pixels are labeled; Repeat the process: Select the next unmarked junk pixel as a new starting point and repeat the above process until all junk pixels are marked.
9. The method for automatically monitoring floating garbage on a water surface according to claim 1, characterized in that: The garbage identification and classification step uses a classifier to classify different types of garbage, including: Feature integration: Integrate the color features, texture features, and shape features of each garbage area into a feature vector; Classification prediction: The new garbage area feature vector is input into the trained classifier, and the classifier will assign each garbage area to a specific category based on the learned pattern; Classification statistics: count the quantity of each type of garbage and calculate the total area of each type of garbage; Statistical results analysis: Generate a final report, summarize the classification and statistical results, and generate a detailed report containing information on the quantity, total area, and distribution of each type of garbage.
10. An automatic monitoring device for floating garbage on a water surface according to any one of claims 1 to 9, characterized in that: include: Drones: These are used to fly above the water, cruise along a preset path and carry high-resolution cameras; High-resolution camera: used to obtain high-definition images of floating garbage on the water surface; Host: It is used to receive and process image data obtained from the high-resolution camera; Cleaning boat: It is used to collect and dispose of floating garbage on the water surface.
Citation Information
Patent Citations
Marine water surface garbage rapid identification method based on multi-feature YOLOV3
CN111950357A
Video raindrop removing method and system based on combination of morphology and fuzzy C clustering
CN105139358A
Channel foreign matter intelligent detection and classification method based on aerial image superpixel texture
CN112241692A
Water environment pollution condition detection and evaluation device based on unmanned aerial vehicle aerial photography data
CN112488020A
Water surface floating object identification method based on semantic segmentation and image anomaly detection
CN116824352A
Cited By
Decoration garbage identification and classification method and system
CN121259458A
Method and system for intelligently generating reservoir dispatching scheme fused with AI Agent
CN122114504A
Water surface floating object monitoring method based on multi-scale feature filtering and recombination
CN122493342A