Underwater garbage detection method and system

Through the improved YOLOX network model, combined with CBAM and ASPP modules, the automatic identification and capture of underwater garbage is realized, and the noise interference and complex environmental problems of underwater target detection are solved, and the accuracy and practicality of detection are improved.

CN115797658BActive Publication Date: 2025-07-22WUHAN POLYTECHNIC UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211669055.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-23
Publication Date
2025-07-22
Estimated Expiration
2042-12-23

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively and automatically detect underwater garbage. Underwater target detection faces problems such as high noise interference, low contrast, and blurred texture characteristics. It is dangerous and unpractical to artificial diving.

Method used

The improved YOLOX network model is adopted to collect underwater garbage image data sets, pre-process and label labels, and set up the YOLOX underwater garbage image object detection model, use the CBAM module and the ASPP module to perform feature fusion, and train and carry underwater robots for real-time detection.

Benefits of technology

It realizes automatic identification and capture of underwater garbage, has good accuracy and detection speed, covers existing garbage types, and solves the problem of low practicality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115797658B_ABST
    Figure CN115797658B_ABST
Patent Text Reader

Abstract

The present invention provides an underwater garbage detection method and system, including collecting an original image dataset of underwater garbage, preprocessing the original image dataset and annotating labels; setting a YOLOX underwater garbage image target detection model, in which the CBAM module in the YOLOX underwater garbage image target detection model adopts an improved structure, including two sub-modules: a channel attention module and a spatial attention module; training the YOLOX underwater garbage image target detection model; mounting the trained YOLOX underwater garbage image target detection model on an underwater robot; placing the underwater robot underwater for garbage detection and autonomous grasping, including the underwater robot capturing video images through a camera and using the trained YOLOX underwater garbage image target detection model to perform real-time detection on objects. The technical solution proposed by the present invention has good accuracy and detection speed, and can cover existing types of underwater garbage, providing a reference for underwater garbage detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of underwater object target detection, and particularly to an underwater garbage detection method and system. Background Technique

[0002] Underwater target detection is to accurately locate and identify targets in an underwater scene. This research has attracted continuous attention due to its wide applications in fields such as oceanography, underwater navigation, and fish farming. With the continuous development of the economic society, the degree of pollution of oceans and lakes has become increasingly serious, and reports of underwater garbage threatening biodiversity are common. These garbage pose long-term economic and environmental threats to the marine and lake ecosystems. However, due to the complex underwater environment and the limitations of the special underwater imaging environment, underwater optical images often have problems such as a lot of noise interference, low contrast, and blurred texture features, and underwater target detection still faces huge challenges.

[0003] Currently, the most commonly used method for underwater operations is manual diving. However, this operation method has a high risk factor, a short operation time, and great harm to the body. With the development of deep learning and the gradual maturity and application of underwater robot technology, applying the deep learning method to underwater robot devices to achieve underwater garbage target detection is a popular direction in the field of underwater garbage recognition today. However, due to the limitations of the difficulty of underwater target detection, no practical technology that can be actually applied has emerged yet. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide an improved underwater garbage detection method based on the YOLOX network model, which is used to effectively and automatically detect the garbage existing in the underwater environment and provide technical support for the underwater operations of underwater robots.

[0005] To achieve the above purpose, the present invention proposes an underwater garbage detection method, including the following steps.

[0006] Step S1: Collect the original image dataset of underwater garbage;

[0007] Step S2: Preprocess the original image dataset and label the tags;

[0008] Step S3: Set up the YOLOX underwater garbage image target detection model. The YOLOX underwater garbage image target detection model is set up based on the YOLOX network framework, including the backbone feature extraction network CSPDarknet, the CBAM module, the feature fusion network PAFPN, and the decoupled head. Three preliminary feature layers of different scales are respectively output from the CSPDarknet structure of the backbone feature extraction network. The three enhanced feature layers after feature fusion by the PAFPN are input into the three decoupled heads for classification prediction to achieve the detection of underwater garbage.

[0009] Among them, the CBAM module set between the backbone feature extraction network CSPDarknet and the feature fusion network PAFPN adopts an improved structure, including two sub-modules: the channel attention module and the spatial attention module. The input feature layer is divided into two branches. One branch passes through the attention mechanism module, and the other branch is directly multiplied element-wise with the output of the attention mechanism module without any processing. Then it is divided into two branches again. One branch passes through the spatial attention module, and the other branch is directly multiplied element-wise with the output of the spatial attention module without any processing. Finally, the feature layer is output.

[0010] Step S4: Train the YOLOX underwater garbage image target detection model.

[0011] Step S5: Install the trained YOLOX underwater garbage image target detection model on the underwater robot.

[0012] Step S6: Place the underwater robot underwater for garbage detection and autonomous grasping, including the underwater robot capturing video images through a camera and using the trained YOLOX underwater garbage image target detection model to perform real-time detection on objects.

[0013] Moreover, in step S2, first, the image convolution kernel is filtered and denoised by the Gaussian function, and then the gray level of the pixels is changed by histogram equalization to adjust the image exposure for image feature enhancement.

[0014] Moreover, in step S4, the Mosaic data augmentation method is adopted. The images are stitched together by randomly scaling, randomly cropping, and randomly arranging them. Then the MixUp augmentation method is used to fill the two sides, top and bottom of the images and perform weighted fusion. Finally, through repeated iterative training, a network model that meets the requirements of target detection performance is obtained.

[0015] Moreover, when using the trained YOLOX underwater garbage image target detection model to perform real-time detection on objects, the detection content includes plastics, plants, metals, woods, fabrics, and fish.

[0016] Moreover, the CBS modules used in the YOLOX underwater garbage image target detection model all adopt an improved structure, including a convolutional layer, a BN layer, and a Swish activation layer connected in sequence. The convolutional layer is used to extract the RGB three-channel features of the original image, and form a new feature map through convolution operation with the convolutional kernel. The BN layer prevents gradient disappearance and overfitting through data normalization operation. The Swish activation layer compresses the feature part with a value less than 0 in a smaller non-zero interval to retain the feature information.

[0017] Moreover, the improved ASPP module is adopted in the YOLOX underwater garbage image target detection model to replace the SPP module in the YOLOX network. The improved ASPP module includes 1 ordinary convolutional layer, 3 dilated convolutional layers, and a global average pooling layer. The feature layers of these 5 channels are stacked together, and the channels are adjusted through a convolutional operation to finally obtain the output.

[0018] On the other hand, the present invention provides an underwater garbage detection system for implementing an underwater garbage detection method as described above.

[0019] Moreover, it includes the following modules

[0020] The first module is used to collect the original image dataset of underwater garbage;

[0021] The second module is used to preprocess the original image dataset and label the tags;

[0022] The third module is used to set the YOLOX underwater garbage image target detection model. The YOLOX underwater garbage image target detection model is set based on the YOLOX network framework, including the backbone feature extraction network CSPDarknet, the CBAM module, the feature fusion network PAFPN, and the decoupled detection head decoupled head. Three preliminary feature layers of different scales are respectively output from the structure of the backbone feature extraction network CSPDarknet, and the 3 enhanced feature layers after feature fusion by the feature fusion network PAFPN are input into the 3 decoupled heads for classification prediction to achieve the detection of underwater garbage;

[0023] Among them, the CBAM module set between the backbone feature extraction network CSPDarknet and the feature fusion network PAFPN adopts an improved structure, which includes two sub-modules: a channel attention module and a spatial attention module. The input feature layer is divided into two branches. One branch passes through the attention mechanism module, and the other branch directly performs an element-wise multiplication operation with the output of the attention mechanism module without any processing. Then it is divided into two branches again. One branch passes through the spatial attention module, and the other branch directly performs an element-wise multiplication operation with the output of the spatial attention module without any processing. Finally, the output feature layer is obtained;

[0024] The fourth module is used to train the YOLOX underwater garbage image target detection model;

[0025] The fifth module is used to carry the trained YOLOX underwater garbage image target detection model on an underwater robot;

[0026] The sixth module is used to place the underwater robot underwater for garbage detection and autonomous grasping, including the underwater robot capturing video images through a camera and using the trained YOLOX underwater garbage image target detection model to perform real-time detection on objects.

[0027] Alternatively, it includes a processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute an underwater garbage detection method as described above.

[0028] Alternatively, it includes a readable storage medium, on which a computer program is stored. When the computer program is executed, it implements an underwater garbage detection method as described above.

[0029] In the present invention, the publicly available underwater garbage image dataset is preprocessed by denoising, defogging, and histogram equalization, which significantly improves the color, clarity, and contrast of the original image. Then, when constructing the YOLOX underwater garbage image target detection model, an improved CBAM is added between the backbone feature extraction network and PAFPN to achieve attention to features in space and channels. Finally, the trained target detection model is carried on the underwater robot system to automatically identify underwater garbage. The technical solution proposed by the present invention has good accuracy and detection speed, and can cover existing underwater garbage types, providing a reference for underwater garbage detection.

[0030] The implementation of the solution of the present invention is simple and convenient, with strong practicability. It solves the problems of low practicability and inconvenience in actual application existing in the related technologies, can improve the user experience, and has important market value. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a flowchart of the method according to the embodiment of the present invention;

[0032] Figure 2 This is the structural diagram of the underwater robot system in the embodiment of the present invention;

[0033] Figure 3 This is the framework diagram of the improved YOLOX underwater garbage image target detection model constructed in the embodiment of the present invention;

[0034] Figure 4 This is the schematic structural diagram of the improved CBAM attention mechanism module proposed in the embodiment of the present invention;

[0035] Figure 5 This is the schematic structural diagram of the improved CBS module and the improved ASPP module proposed in the embodiment of the present invention.

[0036] Figure 6 This is the segmentation schematic diagram of the embodiment of the present invention. Detailed implementation manners

[0037] The technical solution of the present invention will be specifically described below in conjunction with the accompanying drawings and embodiments.

[0038] In the embodiment of the present invention, an underwater garbage detection method based on the YOLOX network model is proposed. As Figure 1 shown, it includes the following steps:

[0039] S1. Collect the original underwater garbage image dataset: Preferably, use the marine debris dataset from J-EDI, which contains images of many different types of marine debris, all captured from real-world environments, providing various object images in different states of decay, occlusion, and overgrowth for the detection task.

[0040] In this embodiment, 6540 images of biological objects such as marine garbage, plants, and animals are selected from the publicly available J-EDI marine debris dataset. These pictures are divided into six categories: plastic, plant, metal, wood, fabric, and fish, and are divided into a training set, a validation set, and a test set according to a ratio of 8:1:1.

[0041] S2. Preprocess and label the original image dataset: Due to the complex and variable underwater environment, the captured images have problems such as color cast, low contrast, noise, and blurring. It is necessary to preprocess the images by denoising, dehazing, and histogram equalization; then mark the dataset through the dataset marking software Labelimg to obtain a set of files in xml format. The xml files are used to store the annotation data corresponding to the jpg format pictures, mainly including the picture name, picture size, and the bounding box positions of target detection.

[0042] Further, in step S2, it is preferred to first filter and denoise the image convolution kernel through a Gaussian function, and then change the gray scale of the pixels by histogram equalization to adjust the image exposure to achieve the effect of image feature enhancement.

[0043] S3. Set up the YOLOX underwater garbage image target detection model:

[0044] The present invention proposes to improve based on the traditional YOLOX network model framework in the prior art to provide a YOLOX underwater garbage image target detection model.

[0045] Further, the present invention proposes to use CSPDarknet as the backbone feature extraction network of the model, input the underwater garbage image for feature extraction to obtain three effective feature layers; use the improved convolutional block attention module (CBAM) to respectively perform channel and spatial attention on the feature layers on the three channels, effectively screen out meaningful features, and improve the representation ability of the convolutional neural network; use PAFPN to fuse the three effective feature layers, and PAFPN is composed of FPN (feature pyramid network) and PAN (path aggregation network); use 3 decoupled heads to realize the classification and prediction of underwater garbage;

[0046] Therefore, the YOLOX underwater garbage image target detection model proposed by the present invention includes the backbone feature extraction network CSPDarknet, the CBAM module, the feature fusion network PAFPN, and the decoupled head; the backbone feature extraction network CSPDarknet is composed of the Focus network, the CBS module, the CSP_X module, the CSP2_X module, and the ASPP module; an improved CBAM attention mechanism is added between the backbone feature extraction network CSPDarknet and the feature fusion network PAFPN; three preliminary feature layers feat1, feat2, and feat3 of different scale sizes are respectively output from the structure of the backbone feature extraction network CSPDarknet, and the three effective feature layers are upsampled and downsampled in PAFPN to achieve feature fusion; the 3 enhanced feature layers after feature fusion are input into 3 decoupled heads for classification prediction to achieve the detection of underwater garbage.

[0047] The Focus network expands the channels of the image input from the original three channels to twelve channels, thereby reducing the parameter calculation of the model and improving the operation speed; the CBS module consists of a convolutional layer, a BN layer, and a Swish activation layer; the CSP_X module is stacked by the CBS convolutional module and X Resunit residual blocks. The main part of the Resunit residual block consists of a 1X1 CBS convolutional module and a 3X3 CBS convolutional module, and no processing is done on the residual side. Finally, the two are added together; the difference between the CSP2_X module and the CSP_X module is that the CSP2_X module replaces the X Resunit residual blocks in the CSP_X module with 2X CBS modules; the improved CBAM is located between the backbone network and the PAFPN and includes two sub-modules, the channel attention module and the spatial attention module, which perform attention on the channel and space respectively.

[0048] The framework diagram of the improved YOLOX underwater garbage image target detection model constructed in this embodiment is as Figure 3 shown. The structure of the YOLOX underwater garbage image target detection model consists of four parts. CSPDarknet is the backbone network of the model, and the original image input size is set to 640*640*3.

[0049] Step 1: The original image is input into the Focus network for initial feature extraction. In the Focus network, first, slicing operations are performed on the three channels respectively. Each channel obtains 4 independent pictures Slice, and then these 4 pictures are concatenated by the concat operation and stacked together, so that the width and height information of the picture is concentrated in the channel space. At this time, the input channels are expanded by 4 times, that is, the concatenated picture becomes 12 channels relative to the original RGB three-channel mode. Finally, a convolution operation is performed through the CBS module to finally become a feature map of 320*320*64;

[0050] The slicing operation that can be adopted in specific implementation can be seen in Figure 6 , divide a picture into 4×4 blocks, then divide it into 4 2×2 units, and then concatenate the same pixel points of each unit.

[0051] The Focus network framework adopted in the present invention is the prior art, and the CBS module included therein is improved on this basis. The other CBS modules subsequently adopted in the YOLOX underwater garbage image target detection model of the present invention are all the same improvement results.

[0052] See Figure 5(1) Part. The improved CBS module includes a convolutional layer Conv, a BN layer BatchNorm, and a Swish activation layer (activation function) connected in sequence. The convolutional layer is used to extract the RGB three-channel features of the original image and form a new feature map through convolution operations with the convolution kernel. The BN layer prevents gradient vanishing and overfitting through data normalization operations. The Swish activation layer replaces the LeakyReLU activation layer in the original CBl module. The Swish activation layer compresses the feature parts with values less than 0 into a smaller non-zero interval, which helps prevent the gradient from gradually approaching 0 and causing saturation during slow training and can relatively completely retain feature information. The Swish function is calculated as follows:

[0053] f(x) = x * sigmoid(βx)

[0054] where x is the data processed by the convolutional layer and the BN layer, sigmoid is the activation function, and β is an adjustable parameter. The Swish function can be regarded as a smooth function between a linear function and a relu activation function, which is superior to the traditional relu in deep models and has the characteristics of having no upper bound, having a lower bound, being smooth, and being non-monotonic.

[0055] Step 2: The output of the Focus network enters the dark2 module. The dark2 module includes 1 CBS module and a CSP_1 module. Based on the output result of the Focus network obtained in Step 1, first use an improved CBS module to extract features. See Figure 5 (1).

[0056] Step 3: Input the feature layer processed by the CBS module obtained in Step 2 into the CSP_1 module. The internal implementation of the CSP_1 module can be seen in the CSP_X module structure part in Figure 3 . The feature layer entering the CSP_X module is divided into two branches.

[0057] One branch first passes through the CBS module and then passes through X Resunit residual blocks connected in sequence (X = 1 for the CSP_1 module, passing through 1 Resunit residual block, X = 3 for the CSP_3 module, passing through 3 Resunit residual blocks). The Resunit residual block is divided into two sub-branches. One sub-branch passes through a 1*1 convolution and a 3*3 convolution in sequence, and the other sub-branch does not perform any processing and directly performs an ADD addition operation with the output of the previous sub-branch to increase the residual structure to increase the gradient value of backpropagation between layers and avoid gradient vanishing caused by deepening.

[0058] The other branch directly passes through the CBS module, and then the two branches are stacked (concat), and finally enter a CBS module to integrate the stacked feature layers.

[0059] Step 4: Input the feature layers processed by the CSP_1 module into two groups of CBS modules and the CSP_3 module in sequence. These two groups of CBS modules and the CSP_3 module form the dark3 module and the dark4 module respectively, and at the same time, a feature layer feat1 and feat2 are respectively led out. The sizes of the feature layers are: 80*80*256 and 40*40*512.

[0060] Step 5: Input the feature layer processed by the dark4 module into the dark5 module. The dark5 module sequentially includes a CBS module, an improved ASPP module, and a CSP2_1 module, and finally leads out the feature layer feat3 with a size of 20*20*1024.

[0061] The improved ASPP module proposed by the present invention replaces the SPP module in the YOLOX network. This is because the SPP module cannot fully integrate local information and global information, which is prone to information loss. While the ASPP introduces dilated convolution, and different receptive fields can be obtained by setting different dilation coefficients, thereby extracting multi-scale context information and improving the ability of the network model to identify targets of different scales. The improved ASPP module uses the LeakyReLU activation function to replace the ReLU activation function used in the original ASPP module to solve the problem that the function gradient is zero and the weights cannot be updated. See Figure 5 (2) Part, the ASPP module includes 1 ordinary convolutional layer, 3 dilated convolutional layers, and a global average pooling layer. The difference between the dilated convolutional layer and the general convolution lies in the dilation (dilation rate). By different padding and dilation, different receptive fields can be obtained to extract multi-scale information. The convolutional kernels of the 3 dilated convolutional layers are 3*3. The global average pooling layer compresses the feature maps of each channel to 1*1 respectively, so as to extract the features of each channel, and then obtain the global features. The feature layers of these 5 channels are stacked together (concat), and the channels are adjusted through a 1*1 convolutional operation to finally obtain the output Output.

[0062] The implementation of the CSP2_X module is similar to that of the CSP_X. The difference is that the CSP2_X replaces the Resunit residual block with 2X CBS modules.

[0063] Step 6: Add the three feature layers feat1, feat2, and feat3 into the improved CBAM attention mechanism module respectively.

[0064] See Figure 4, the improved CBAM attention mechanism module proposed by the present invention includes two sub-modules: a channel attention module and a spatial attention module. The input feature layer is divided into two branches. One branch passes through the attention mechanism module, and the other branch directly performs an element-wise multiplication operation with the output of the attention mechanism module without any processing. Then it is divided into two branches again. One branch passes through the spatial attention module, and the other branch directly performs an element-wise multiplication operation with the output of the spatial attention module without any processing. Finally, the output feature layer is obtained.

[0065] In the channel attention module, first, the input feature map is divided into two branches, and average pooling (avgpooling) and max pooling operations are respectively performed. Then, it passes through two fully connected layers FC with a relu activation function layer inserted in the middle, first reducing the dimension of the feature map and then increasing the dimension to fit the correlation between channels. Finally, the two feature maps are added together, and normalized using the Sigmoid function to re-allocate the feature weights of each channel, which is beneficial to learning the channel information corresponding to the target area. In the spatial attention module, it is divided into two branches. One branch performs an average pooling operation on the feature layer to extract the spatial information of the feature map, then connects it with a 7*7 convolutional layer and adds non-linear features through the Sigmoid function to obtain the feature map m1. The other branch first uses dilated convolutions with receptive field sizes of 7*7 and 3*3 to increase the receptive field and fuse context information, then uses a 1*1 convolutional layer to reduce the dimension, and finally obtains the feature map m2 through the Sigmoid function. Finally, the two feature maps m1 and m2 are added together.

[0066] Step 7: Input the three feature layers feat1, feat2, and feat3 output by the CBAM attention mechanism module into the feature fusion network PAFPN. The feature layer feat3 (20*20*1024) passes through the CBS module. In the CBS module, a 1*1 convolution is performed once to adjust the channels to obtain a new feature layer P5 (20*20*512), and then upsampling (40*40*512) is performed, that is, the size of the image is enlarged. Subsequently, it is concatenated with the feature layer feat2 (40*40*512) to obtain a new feature layer C3_p4 (40*40*1024);

[0067] C3_p4 passes through two branches of the CSP2_1 module, and 1*1 convolution operations are respectively performed to adjust the channels. At this time, the size of the feature layer is 40*40*512. Then it enters the CBS module to perform a 1*1 convolution to adjust the channels to obtain a new feature layer P4 (40*40*256), and then upsampling (80*80*256) is performed and concatenated with the feature layer feat1 (80*80*512);

[0068] Then it enters the CSP2_1 module, and the two branches in the CSP2_1 module respectively perform 1*1 convolution operations to obtain the feature C3_p3 (80*80*256).

[0069] The feature C3_p3 passes through the CBS module to perform a 3*3 convolution operation to adjust the channels (40*40*256), then is concatenated with P4 (40*40*512), and then passes through the CSP2_1 module to obtain the feature C3_n3 (40*40*512). C3_n3 passes through the CBS module to perform a 3X3 convolution to adjust the channels (20*20*512), and then is concatenated with P5 (20*20*1024). It enters the CSP2_1 module, and the two branches of the CSP2_1 module respectively perform 1*1 convolution operations to obtain the feature C3_n4

[0070] (20*20*1024).

[0071] Step 8: Input the enhanced feature layers C3_p3, C3_n3, and C3_n4 output from the feature fusion network PAFPN into decoupled heads of three different sizes, 80*80*256, 40*40*512, and 20*20*1024, for classification prediction. In the YOLOX object detection model, there are three decoupled heads, and their structures are the same. First, the feature layer is input into the CBS module, and in the CBS module, a 1*1 convolutional kernel is used for channel integration. Then it is divided into two branches, and each branch is input into 2 successively connected CBS modules, and 3*3 convolutional kernels are used for feature extraction. Then a convolutional layer conv is used to obtain the classification prediction result. One branch is used to judge the category to which the feature layer belongs, and the other branch is used to judge the regression parameters of the feature layer and to judge whether there is a corresponding object, that is, to predict the four-dimensional coordinates and confidence of the prediction box.

[0072] S4. Train the YOLOX underwater garbage image object detection model: Load the dataset in xml file format, adopt the Mosaic data augmentation method, splice the images by randomly scaling, randomly cropping, and randomly arranging them, and then adopt the MixUp augmentation method on the basis of Mosaic to fill the two sides of the image, fill the top and bottom, and perform weighted fusion. Finally, through repeated iterative training, a network model with good performance is obtained. This model can generate stable parameters and meet the performance requirements of object detection.

[0073] Using the Mosaic data augmentation method and through multiple trainings, training weights with smaller loss values can be obtained. The smaller the loss value of the weights, the better the training effect of the model. Using this weight can better perform classification prediction.

[0074] S5. Apply the trained object detection model to the underwater robot system: Refer to Figure 2 , the underwater robot is equipped with an underwater camera, a sonar system, and a short baseline positioning system; the position information of the robot is obtained through the short baseline positioning system. The robot captures video images through the camera within its area, supports real-time detection of objects using the object detection model, obtains the category of the target in the image and the coordinate information of its bounding box in the image, and feeds back the detection result to the robot. If there is no target in the current area, the robot is controlled to move, and the sonar system is used to emit pulse signals to the seabed. The depth information from the seabed is calculated based on the corresponding time to determine whether the robot continues to dive.

[0075] S6. Use the underwater robot equipped with the object detection model to detect garbage: Place the underwater robot underwater for garbage detection and autonomous grasping. The detection contents include: plastic, plants, metal, wood, fabric, and fish.

[0076] Specifically, when implemented, the method proposed by the technical solution of the present invention can be automatically run by those skilled in the art using computer software technology. The system device for implementing the method, such as a computer-readable storage medium storing the corresponding computer program of the technical solution of the present invention and a computer device including running the corresponding computer program, should also be within the protection scope of the present invention.

[0077] In some possible embodiments, an underwater garbage detection system is provided, including the following modules,

[0078] The first module is used to collect the original image dataset of underwater garbage;

[0079] The second module is used to preprocess the original image dataset and label the tags;

[0080] The third module is used to set up the YOLOX underwater garbage image object detection model. The YOLOX underwater garbage image object detection model is set based on the YOLOX network framework, including the backbone feature extraction network CSPDarknet, the CBAM module, the feature fusion network PAFPN, and the decoupled head; three preliminary feature layers of different scales are respectively output from the structure of the backbone feature extraction network CSPDarknet, and the three enhanced feature layers after feature fusion by the feature fusion network PAFPN are input into the three decoupled heads for classification prediction to achieve the detection of underwater garbage;

[0081] Among them, the CBAM module set between the backbone feature extraction network CSPDarknet and the feature fusion network PAFPN adopts an improved structure, which includes two sub-modules: a channel attention module and a spatial attention module. The input feature layer is divided into two branches. One branch passes through the attention mechanism module, and the other branch directly performs an element-wise multiplication operation with the output of the attention mechanism module without any processing. Then it is divided into two branches again. One branch passes through the spatial attention module, and the other branch directly performs an element-wise multiplication operation with the output of the spatial attention module without any processing. Finally, the output feature layer is obtained;

[0082] The fourth module is used to train the YOLOX underwater garbage image target detection model;

[0083] The fifth module is used to carry the trained YOLOX underwater garbage image target detection model on the underwater robot;

[0084] The sixth module is used to place the underwater robot underwater for garbage detection and autonomous grasping, including the underwater robot capturing video images through a camera and using the trained YOLOX underwater garbage image target detection model to perform real-time detection on objects.

[0085] In some possible embodiments, an underwater garbage detection system is provided, which includes a processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute an underwater garbage detection method as described above.

[0086] In some possible embodiments, a readable storage medium is provided, on which a computer program is stored. When the computer program is executed, an underwater garbage detection method as described above is implemented.

[0087] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar methods to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

Claims

1. An underwater garbage detection method, characterized in that: Including the following steps, Step S1: Collect the original image dataset of underwater garbage; Step S2: Preprocess the original image dataset and label the tags; Step S3: Set up the YOLOX underwater garbage image target detection model. The YOLOX underwater garbage image target detection model is set up based on the YOLOX network framework, including the backbone feature extraction network CSPDarknet, the CBAM module, the feature fusion network PAFPN, and the decoupled head. Three preliminary feature layers of different scales are respectively output from the CSPDarknet structure of the backbone feature extraction network. The three enhanced feature layers after feature fusion by the PAFPN are input into the three decoupled heads for classification prediction to achieve the detection of underwater garbage; Among them, the CBAM module set between the backbone feature extraction network CSPDarknet and the feature fusion network PAFPN adopts an improved structure, including two sub-modules: a channel attention module and a spatial attention module. The input feature layer is divided into two branches. One branch passes through the attention mechanism module, and the other branch is directly multiplied by the output of the attention mechanism module without any processing. Then it is divided into two branches again. One branch passes through the spatial attention module, and the other branch is directly multiplied by the output of the spatial attention module without any processing. Finally, the output feature layer is obtained; The CBS modules used in the YOLOX underwater garbage image target detection model all adopt improved structures, including a convolutional layer, a BN layer, and a Swish activation layer connected in sequence. The convolutional layer is used to extract the RGB three-channel features of the original image and form a new feature map through convolution operations with the convolutional kernel. The BN layer prevents gradient disappearance and overfitting through data normalization operations. The Swish activation layer compresses the feature parts with values less than 0 in a smaller non-zero interval to retain feature information; The improved ASPP module is used to replace the SPP module in the YOLOX network in the YOLOX underwater garbage image target detection model. The improved ASPP module includes 1 ordinary convolutional layer, 3 dilated convolutional layers, and a global average pooling layer. The feature layers of these 5 channels are stacked together, and the channels are adjusted through a convolutional operation to finally obtain the output; Step S4: Train the YOLOX underwater garbage image target detection model; Step S5: Mount the trained YOLOX underwater garbage image target detection model on the underwater robot; Step S6: Place the underwater robot underwater for garbage detection and autonomous grasping, including the underwater robot capturing video images through the camera and using the trained YOLOX underwater garbage image target detection model to perform real-time detection on the object.

2. The underwater garbage detection method according to claim 1, wherein: In step S2, first filter and denoise the image convolution kernel through the Gaussian function, and then change the pixel grayscale by histogram equalization to adjust the image exposure for image feature enhancement.

3. The underwater garbage detection method according to claim 1, wherein: In step S4, the Mosaic data augmentation method is adopted. The images are stitched by means of random scaling, random cropping and random arrangement. Then, the MixUp augmentation method is used to perform padding on both sides and top and bottom of the images and weighted fusion. Finally, through repeated iterative training, a network model that meets the requirements of object detection performance is obtained.

4. The underwater garbage detection method according to claim 1, wherein: When using the trained YOLOX underwater garbage image object detection model to perform real-time detection on objects, the detection contents include plastics, plants, metals, woods, fabrics and fish.

5. An underwater garbage detection system, characterized in that: It includes the following modules: The first module is used to collect the original image dataset of underwater garbage. The second module is used to preprocess the original image dataset and label the tags. The third module is used to set up the YOLOX underwater garbage image object detection model. The YOLOX underwater garbage image object detection model is set based on the YOLOX network framework, and includes a backbone feature extraction network CSPDarknet, a CBAM module, a feature fusion network PAFPN, and a decoupled head; three preliminary feature layers of different scales are respectively output from the structure of the backbone feature extraction network CSPDarknet, and the three enhanced feature layers after feature fusion by the feature fusion network PAFPN are input into the three decoupled heads for classification prediction to achieve the detection of underwater garbage. Among them, the CBAM module set between the backbone feature extraction network CSPDarknet and the feature fusion network PAFPN adopts an improved structure, which includes two sub-modules: a channel attention module and a spatial attention module. The input feature layer is divided into two branches. One branch passes through the attention mechanism module, and the other branch directly performs an element-wise multiplication operation with the output of the attention mechanism module without any processing, and then is divided into two branches again. One branch passes through the spatial attention module, and the other branch directly performs an element-wise multiplication operation with the output of the spatial attention module without any processing, and finally the output feature layer is obtained. All CBS modules used in the YOLOX underwater garbage image object detection model adopt an improved structure, which includes a convolutional layer, a BN layer and a Swish activation layer connected in sequence. The convolutional layer is used to extract the RGB three-channel features of the original image and form a new feature map through convolution operations with the convolutional kernel. The BN layer prevents gradient disappearance and overfitting through data normalization operations. The Swish activation layer compresses the feature parts with values less than 0 in a smaller non-zero interval to retain the feature information. The improved ASPP module is adopted to replace the SPP module in the YOLOX network in the YOLOX underwater garbage image object detection model. The improved ASPP module includes 1 ordinary convolutional layer, 3 dilated convolutional layers and a global average pooling layer. The feature layers of these 5 channels are stacked together, and the channels are adjusted through a convolutional operation, and finally the output is obtained. The fourth module is used to train the YOLOX underwater garbage image object detection model. The fifth module is used to carry the trained YOLOX underwater garbage image target detection model on the underwater robot; The sixth module is used to place the underwater robot underwater for garbage detection and autonomous grasping, including the underwater robot capturing video images through a camera and using the trained YOLOX underwater garbage image target detection model to perform real-time detection on objects.

6. An underwater garbage detection device, characterized in that: It includes a processor and a memory. The memory is used to store program instructions, and the processor is used to call the stored instructions in the memory to execute an underwater garbage detection method according to any one of claims 1-4.

7. A readable storage medium, characterized in that: A computer program is stored on the readable storage medium, and when the computer program is executed, an underwater garbage detection method according to any one of claims 1-4 is implemented.

Citation Information

Patent Citations

  • Shielding object detection method and device

    CN114187491A

  • Light-weighted target detection method and device, and storage medium

    WO2022213395A1