An underwater target recognition method based on YOLOv4

By improving the YOLOv4 model, combining Mosaic and e-Mosaic data augmentation methods, combining lightweight networks and decoupled detection heads, the problem of insufficient recognition speed, accuracy and stability in underwater target recognition technology is solved, and more efficient underwater target recognition is achieved.

CN115909039BActive Publication Date: 2025-07-04FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211201950.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2025-07-04
Estimated Expiration
2042-09-29

AI Technical Summary

Technical Problem

The existing underwater target recognition technology is not effective in underwater environments, cannot effectively improve the recognition speed and accuracy, and is insufficient instability.

Method used

The improved YOLOv4 model is adopted, combining Mosaic image augmentation, e-Mosaic data augmentation, GW and RGHS and ICM combined image augmentation methods, and combining lightweight Mobilenetv1 backbone network and decoupled detection head to build an underwater target recognition network.

Benefits of technology

It improves the speed, accuracy and stability of underwater target recognition, and improves detection accuracy and time efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115909039B_ABST
    Figure CN115909039B_ABST
Patent Text Reader

Abstract

The present invention provides an underwater target recognition method based on YOLOv4, comprising the following steps: Step S1: Collect relevant underwater images; Step S2: Apply the Mosaic image augmentation technique for data augmentation; Step S3: Combine the underwater image enhancement method and the Mosaic data augmentation method; Step S4: Combine GW and RGHS, ICM combination with the Mosaic data augmentation method; Step S5: Obtain an improved YOLOv4 model; Step S6: Use the training set of the dataset described in S1 and apply the Mosaic data augmentation method to train the improved YOLO V4 model; Step S7: Use the test set in the dataset to test the model trained in S6 and then use it for underwater target recognition; Applying this technical solution can improve the speed and accuracy of underwater target recognition and enhance the stability of recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target recognition, and particularly to an underwater target recognition method based on YOLOv4. Background Technique

[0002] In the exploration of the ocean by humans, underwater target recognition technology has played an important role. In view of the chaotic underwater environment, with problems such as color deviation, low contrast, and blurriness, some excellent target recognition technologies have not played their true roles.

[0003] With the continuous development of neural networks, some excellent neural network structures have been added to existing target recognition models, greatly improving the accuracy and efficiency of target recognition.

[0004] The YOLO network is the first single-stage detector in the field of deep learning. YOLO adopts a different idea from two-stage object detection. Instead of generating a candidate region and then extracting features and performing classification regression in the candidate region, it directly segments the target image and directly performs feature extraction, classification, and prediction box regression on the segmented regions. The characteristic of the YOLO network is its very high speed. Subsequently, a series of improvements were made based on YOLO, and versions v2 and v3 were proposed. YOLOv3 is widely used because of its fast training speed and fast detection speed. It uses excellent structures such as the Darknet53 network, anchor boxes, and the FPN network. And YOLOv4 is an improvement based on YOLOv3. Compared with version v3, the improvements in version v4 include: changing the backbone network Darknet53 to CSPDarknet53, adding the SPP and PANet network structures, using CIoU as the loss function, and using Mosaic image augmentation, etc.

[0005] However, currently, underwater target recognition technology often directly transplants target recognition technology for land environments, and some excellent target recognition technologies have not played their true roles. Therefore, the present invention designs an underwater target recognition method based on YOLOv4. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide an underwater target recognition method based on YOLOv4, which improves the speed and accuracy of recognizing underwater targets and also improves the stability of recognition.

[0007] To achieve the above purpose, the present invention adopts the following technical scheme: An underwater target recognition technology based on improved YOLOV4, including the following steps:

[0008] Step S1: Collect relevant underwater images and then make them into an underwater dataset;

[0009] Step S2: Apply the Mosaic image augmentation technique to the established underwater dataset for data augmentation;

[0010] Step S3: In view of the low background complexity of the Mosaic data augmentation method in the underwater environment, design an e-Mosaic data augmentation method, that is, combine the underwater image enhancement method and the Mosaic data augmentation method;

[0011] Step S4: After validating the effectiveness in the underwater dataset, further combine GW, RGHS, and ICM with the Mosaic data augmentation method;

[0012] Step S5: Build an improved YOLOv4 model based on the YOLOv4 algorithm to obtain the improved YOLOv4 model;

[0013] Step S6: Use the training set of the dataset described in S1 to train the improved YOLOV4 model using the Mosaic data augmentation method, and then load the trained weight file into the underwater target recognition network obtained by the improved YOLOv4 algorithm;

[0014] Step S7: Use the test set in the dataset to test the model trained in S6, and then use it for underwater target recognition;

[0015] The YOLOv4 network structure improves the detection head part in the YOLOv4 model by using a decoupled detection head structure; further apply this structure to a lightweight network model, which is the YOLOv4-lite neural network based on the Mobilenetv1 backbone network.

[0016] In a preferred embodiment, the step S1 includes the following steps: The dataset contains a total of 4,757 pictures, 3,805 of which are used as the training set, 952 pictures are used as the test set, and 381 pictures are drawn from the training set as the validation set during training; the target categories are divided into four categories: sea urchin, starfish, sea cucumber, and shell.

[0017] In a preferred embodiment, in the Mosaic image augmentation in step S2, first randomly select 4 pictures from the picture set to generate a new picture, input the new picture into the target recognition model, and then continue to randomly select 4 new pictures from the picture set for Mosaic image augmentation until all the pictures in the picture set are taken.

[0018] In a preferred embodiment, in step S3, e-Mosaic is a Mosaic augmentation fused with image enhancement, which is a new augmentation method formed by combining six underwater image enhancement methods with the ordinary Mosaic augmentation method; e-Mosaic augmentation selects four images. Before stitching, e-Mosaic augmentation first uses an image enhancement algorithm to enhance the two diagonal images.

[0019] In a preferred embodiment, by setting a random value r, where r is a random number within the range of [0, 1], a judgment of the r value is performed for both the No. 1 image and the No. 3 image; when the r value is less than 0.4, the GW algorithm is used for enhancement.

[0020] In a preferred embodiment, the first 14 convolutional layers of the Mobilenetv1 lightweight network are used as the backbone part in the YOLOv4 network. According to the changes in the width and height dimensions of the feature layers in the Mobilenetv1 network, the 6th, 12th, and 14th layers in the Monilenetv1 network are used as the outputs of the backbone network; meanwhile, all 3×3 convolutional blocks in the YOLOv4 network are replaced with 3×3 depthwise separable convolutions; the size of the input layer is changed from 224×224 to 416×416.

[0021] In a preferred embodiment, the coupled detection head is changed to a decoupled detection head. After the decoupled detection head performs a 1×1 convolution on the input feature layer, it is divided into two parts, one containing classification information and the other containing regression information. The classification information part passes through two 3×3 convolutions and one 1×1 convolution to obtain the final classification output layer, while the regression information part is divided into two parts again after two 3×3 convolutions. One part contains the position information of the regression box, and the other part is the probability information that the regression box contains the target. Finally, the three types of information respectively pass through a 1×1 convolution to obtain the final regression box output layer, target class output layer, and target confidence output layer.

[0022] Compared with the prior art, the present invention has the following beneficial effects: BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is the preprocessing process of using Mosaic image augmentation in the preferred embodiment of the present invention;

[0024] Figure 2 It is the splicing process of applying Mosaic image augmentation in the preferred embodiment of the present invention;

[0025] Figure 3 It is the preprocessing process of e-Mosaic image augmentation in the preferred embodiment of the present invention;

[0026] Figure 4Schematic diagram of the YOLOv4-lite network structure based on Mobilenetv1 according to the preferred embodiment of the present invention;

[0027] Figure 5 Schematic diagrams of the coupled detection head and the decoupled detection head of YOLO according to the preferred embodiment of the present invention;

[0028] Figure 6 Schematic diagram of the prediction process of the feature layer of the decoupled detection head according to the preferred embodiment of the present invention;

[0029] Figure 7 Schematic diagram of the improved YOLOv4 network structure based on Mobilenetv1 and the decoupled detection head according to the preferred embodiment of the present invention. Detailed implementation manners

[0030] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0031] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0032] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0033] An underwater target recognition method based on YOLOv4 includes the following steps:

[0034] Step S1, the data set comes from the target recognition group of the 2020 National Underwater Robot Competition in China sponsored by the National Natural Science Foundation of China. The data set contains a total of 4,757 pictures. Among them, 3,805 pictures are used as the training set, 952 pictures are used as the test set, and 381 pictures are drawn from the training set as the validation set during training. The target types are divided into four categories: sea urchin, starfish, sea cucumber, and shell.

[0035] Step S2, as Figure 1 shown, Mosaic image augmentation first randomly selects 4 pictures from the picture set to generate a new picture, inputs the new picture into the target recognition model, and then continues to randomly select 4 new pictures from the picture set for Mosaic image augmentation until all the pictures in the picture set are taken.

[0036] The specific steps are asFigure 2 As shown in the figure, first select four pictures, and then randomly reduce and flip the four pictures respectively; then splice them. When splicing, first generate a blank picture (the part of the black frame in the figure), and then put the four pictures processed before together on the blank picture for splicing. The positions of the splicing points are respectively at the four vertices of a rectangle on the blank picture, where w and h are generally taken as 0.4 times the width and height of the blank picture; finally, crop the overlapping or exceeding parts, and thus a new picture processed by Mosaic image augmentation is obtained.

[0037] Step S3, Mosaic augmentation integrated with image enhancement. Select six underwater image enhancement methods (GW algorithm, CLAHE algorithm, MSRCR algorithm, RGHS algorithm, ICM algorithm and UCM algorithm) and combine them with the ordinary Mosaic augmentation method to form a new augmentation method (named e-Mosaic image augmentation). Similar to the ordinary Mosaic augmentation, e-Mosaic augmentation selects four images. The difference is that before stitching, e-Mosaic augmentation first uses the image enhancement algorithm to enhance the two images on the diagonal. The purpose of doing this is to increase the complexity of the background, which is beneficial to the training of the target recognition neural network.

[0038] As Figure 3 shown is the preprocessing process of e-Mosaic image augmentation. First, perform image enhancement on the original picture set through the underwater image enhancement algorithm to form an underwater enhanced picture set, where the picture numbers in the enhanced picture dataset correspond one-to-one with the picture numbers in the original picture set. Then randomly select 4 pictures from the original picture set, and replace two of them with enhanced pictures according to the picture numbers. Generate a new picture from the 4 pictures through Mosaic image augmentation, and then input the new synthesized picture into the target recognition model. Then continue to randomly select 4 new pictures from the picture set for e-Mosaic image augmentation, and repeat this process until all the pictures in the picture set are taken.

[0039] The present invention respectively studies the applications of several preprocessing methods such as e-Mosaic(CLAHE) augmentation, e-Mosaic(GW) augmentation, e-Mosaic(ICM) augmentation, e-Mosaic(RGHS) augmentation, e-Mosaic(MSRCR) augmentation and e-Mosaic(UCM) augmentation in the YOLOv4 target detection network. The four methods of GW, RGHS, ICM and CLAHE have relatively good recognition accuracies (the performance decreases in order), so these four methods are combined for result comparison.

[0040] For the comparison results, the method of the present invention enhances the image on the diagonal. By setting a random value r to select the image enhancement method, the three enhancement methods with the best performance, including GW, RGHS, and ICM, are selected as alternative methods. r is a random number with a value range in [0, 1]. A judgment of the r value is performed for both Picture No. 1 and Picture No. 3. When the r value is less than 0.4, the GW algorithm is used for enhancement; the probabilities of using the other two methods are both 0.3. When the r value is greater than 0.4 and less than 0.7, the ICM algorithm is used for enhancement; when the r value is greater than 0.7, the RGHS algorithm is used for enhancement. It can be found that the probability of using the GW algorithm for enhancement is 0.4, while the probabilities of using the other two algorithms are both 0.3. The method based on the GW algorithm has the best effect, so the probability of using the GW algorithm is set higher.

[0041] As Figure 4 shown, the schematic diagram of the YOLOv4-lite network structure based on Mobilenetv1.

[0042] The first 14 convolutional layers of the Mobilenetv1 network, that is, all operations before the pooling operation, are used as the backbone part in the YOLOv4 network. The Mobilenetv1 network is a lightweight classification network, and its core technology is the depthwise separable convolution block. Different from ordinary three-dimensional convolution, the depthwise separable convolution divides the convolution operation into two steps, reducing the amount of calculation and improving the training and detection speed.

[0043] According to the changes in the width and height dimensions of the feature layers in the Mobilenetv1 network, the 6th, 12th, and 14th layers in the Monilenetv1 network are used as the outputs of the backbone network. At the same time, in order to further increase the speed of the overall network, the 3×3 convolution blocks in the YOLOv4 network are all replaced with 3×3 depthwise separable convolutions. To correspond to the original YOLOv4 network, the size of the input layer is changed from 224×224 to 416×416.

[0044] As Figure 5 shown, the final coupled detection head and decoupled detection head of the YOLO neural network.

[0045] The coupled detection head will perform a 1×1 convolution on each of the three input feature layers to form the final output layer, which includes three types of information: the type of the target, the position of the regression box, and whether there is an object.

[0046] After the decoupled detection head performs a 1×1 convolution on the input feature layer, it is divided into two parts, one containing classification information and the other containing regression information. The classification information part goes through two 3×3 convolutions and one 1×1 convolution to obtain the final classification output layer. The regression information part is divided into two parts again after two 3×3 convolutions. One part contains the position information of the regression box, and the other part is the probability information that the regression box contains the target. Finally, the three types of information go through a 1×1 convolution respectively to obtain the final regression box output layer, target class output layer, and target confidence output layer.

[0047] As Figure 6 shown, in the prediction process of the decoupled detection head feature layer, the feature layer is predicted in a way that does not use prior boxes.

[0048] The feature layer is divided into n×n blue feature points, where n is the size of the feature layer (the size of the feature layer shown in the figure is 13×13). Then, the detection head predicts six pieces of information, a1, a2, a3, a4, c, and p, at each feature point (for ease of observation, only the prediction of a certain feature point is shown in the figure, and other feature points are hidden). Among them, a1, a2, a3, and a4 are the position parameters of the prediction box. a1 and a2 are the offset distances from the blue feature point. The green point is the center of the prediction box. ea3 and ea4 are the width and height of the prediction box. The green feature point and the red feature box are the final prediction results of the detection head. c and p are the predicted class and the confidence of the presence of the target, respectively.

[0049] As Figure 7 shown, the schematic diagram of the improved YOLOv4 network structure based on Mobilenetv1 and the decoupled detection head.

[0050] The decoupled detection head is added to the YOLOv4 network. There are three convolutional layers at the input end of the detection head, corresponding to the outputs of three different sizes in the Neck. In the detection head, to improve the accuracy, first, the CBM convolutional block is selected as the initial 1×1 convolution part and the two 3×3 convolution parts for separating features in the decoupled detection head. Finally, ordinary convolution is used as the last 1×1 convolution part. For the recognition of each different target size, a tensor splicing layer is used to splice the feature layers containing three types of information together to obtain three outputs of different sizes. The SPP layer represents the Spatial Pyramid Pooling layer; the US layer represents the Up Sampling layer; the DS layer represents the Down Sampling layer.

[0051] In the lightweight network, the 3×3 convolution block of the decoupled detection head uses the depthwise separable convolution block, continuing the lightweight and fast characteristics of the lightweight network.

[0052] The improved YOLOv4 network structure of the present invention has a significant improvement in both detection accuracy and detection time.

[0053] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An underwater target recognition method based on YOLOv4, characterized in that, It includes the following steps: Step S1: Collect relevant underwater images and then make them into an underwater dataset; Step S2: Apply the Mosaic image augmentation technique to the established underwater dataset for data augmentation; Step S3: In view of the situation that the background complexity is not high in the underwater environment for the Mosaic data augmentation method, design an e-Mosaic data augmentation method, that is, combine the underwater image enhancement method and the Mosaic data augmentation method; Step S4: After conducting validity verification in the underwater dataset, further combine GW and RGHS, ICM combinations with the Mosaic data augmentation method; Step S5: Build an improved YOLOv4 model based on the YOLOv4 algorithm to obtain an improved YOLOv4 model; Step S6: Use the training set of the dataset described in S1 to train the improved YOLO V4 model using the Mosaic data augmentation method, and then load the obtained weight file into the underwater target recognition network obtained by the improved YOLOv4 algorithm; Step S7: Use the test set in the dataset to test the model trained in S6, and then use it for underwater target recognition; The improved YOLOv4 network structure improves the detection head part in the YOLOv4 model by using a decoupled detection head structure; further apply this structure to a lightweight network model, which is the YOLOv4-lite neural network based on the Mobilenetv1 backbone network; In step S3, e-Mosaic is a Mosaic augmentation fused with image enhancement, which is a new augmentation method formed by combining six underwater image enhancement methods with the ordinary Mosaic augmentation method; e-Mosaic augmentation selects four images. Before stitching, e-Mosaic augmentation first uses an image enhancement algorithm to enhance the two diagonal images; the six underwater image enhancement methods include the GW algorithm, the CLAHE algorithm, the MSRCR algorithm, the RGHS algorithm, the ICM algorithm, and the UCM algorithm.

2. The underwater target recognition method based on YOLOv4 according to claim 1, characterized in that, The step S1 includes the following steps: The dataset contains a total of 4757 pictures, 3805 of which are used as the training set, 952 pictures are used as the test set, and 381 pictures are selected from the training set as the validation set during training; the target categories are divided into four categories: sea urchins, starfish, sea cucumbers, and shells.

3. The underwater target recognition method based on YOLOv4 according to claim 1, characterized in that, In step S2, Mosaic image augmentation first randomly selects 4 pictures from the picture set to generate a new picture, inputs the new picture into the target recognition model, and then continues to randomly select 4 new pictures from the picture set for Mosaic image augmentation until all the pictures in the picture set are taken.

4. An underwater target recognition method based on YOLOv4 according to claim 1, characterized in that, Take the first 14 convolutional layers of the Mobilenetv1 lightweight network as the backbone part in the YOLOv4 network. According to the changes in the width and height dimensions of the feature layers in the Mobilenetv1 network, take the 6th, 12th, and 14th layers in the Monilenetv1 network as the outputs of the backbone network. At the same time, replace all 3×3 convolutional blocks in the YOLOv4 network with 3×3 depthwise separable convolutions. Change the size of the input layer from 224×224 to 416×416.

5. The underwater target recognition method based on YOLOv4 according to claim 1, characterized in that, Change the coupled detection head to a decoupled detection head. After performing a 1×1 convolution on the input feature layer, the decoupled detection head divides it into two parts, one containing classification information and the other containing regression information. The classification information part goes through two 3×3 convolutions and one 1×1 convolution to obtain the final classification output layer. The regression information part is divided into two parts again after two 3×3 convolutions. One part contains the position information of the regression box, and the other part is the probability information that the regression box contains the target. Finally, the three types of information go through a 1×1 convolution respectively to obtain the final regression box output layer, target class output layer, and target confidence output layer.

Citation Information

Patent Citations

  • Handheld call detection method based on lightweight target detection network

    AU2020103494A4

  • Smoking behavior identification method based on optimized YOLOv4 model

    CN113807276A