Underwater target detection method based on improved YOLOv5s

By improving the YOLOv5s method, designing a new attention module and reconstructing the pyramid pooling structure, combining the GFPN mechanism to improve BiFPN and establishing Focaler-Shape loss function, the problems of low detection accuracy and high error detection rate in underwater environments are solved, and higher detection accuracy and recall rate are achieved.

CN120125983APending Publication Date: 2025-06-10SHENYANG UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510203787.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Traditional target detection methods are difficult to achieve high accuracy and low error detection rate target detection due to factors such as light refraction, edge blur and target occlusion in underwater environments.

Method used

Improve the YOLOv5s method, design a new attention module MSDW to integrate it into the C3 module, reconstruct the SPPF pyramid pooling, improve BiFPN based on the GFPN mechanism, weight the bidirectional pyramid network, and establish the Focaler-Shape loss function to optimize the training process.

Benefits of technology

It effectively improves the comprehensive performance of underwater blurred target image detection tasks, significantly improves the accuracy and recall of underwater target detection, and is suitable for target detection in complex underwater environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125983A_ABST
    Figure CN120125983A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of target detection methods, and discloses an underwater target detection method based on improved YOLOv5s, and the method comprises the steps: S1, building an underwater target data set, and labeling the underwater target data set; s2, a new attention module MSDW is designed and fused into the C3 module to form a new C3MSDW module, and the C3MSDW module replaces the trunk C3 module; s3, reconstructing SPPF pyramid pooling; s4, improving the BiFPN weighted bidirectional pyramid network based on a GFPN mechanism idea; s5, a Focaler-Shape loss function is established, and an improved YOLOv5s underwater target detection network is obtained; s6, inputting the established data set into the improved model for training, and obtaining an underwater target detection model of the improved YOLOv5s; and S7, inputting a to-be-detected image or video into the trained improved YOLOv5s underwater target detection model for detection, and outputting a detection result with confidence, an interest frame and category information. According to the method, the comprehensive performance in an underwater fuzzy target image detection task can be effectively improved, and the underwater target detection accuracy and recall rate are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target detection methods, and specifically to an underwater target detection method based on improved YOLOv5s. Background Technique

[0002] In the field of underwater marine treasure fishing, the traditional manual diving fishing method exposes operators to a relatively high risk factor and poses a serious threat to their physical health; while trawl fishing will cause great damage to the underwater ecological environment. With the development of technology, the underwater robot vision grasping technology has gradually emerged. Among them, underwater target detection, as a key link in the grasping operation, is of great importance. In recent years, deep learning technology has made remarkable progress in the field of target detection. Compared with traditional detection methods based on color, texture, etc., deep learning can automatically learn high-level feature representations and improve the accuracy and robustness of target detection. The classic deep learning detection algorithms are divided into two types: single-stage and two-stage. The single-stage YOLO algorithm shows advantages in actual target detection and promotes the rapid development of target detection technology.

[0003] However, the underwater environment is extremely complex, and the phenomenon of light refraction is widespread, which makes the collected underwater images have problems such as blurred edges and color distortion. At the same time, the gregarious characteristics of marine treasures cause them to often block each other in the images. These factors seriously affect the performance of traditional target detection methods and pose huge challenges in underwater target detection applications. Traditional methods perform poorly in identifying underwater target categories and detection accuracy, and it is difficult to meet the actual needs, greatly limiting the effective application and development of underwater target detection technology in related fields.

[0004] To solve the above problems, we propose an underwater target detection method based on improved YOLOv5s. Summary of the Invention

[0005] The purpose of the present invention is to provide an underwater target detection method based on improved YOLOv5s to solve the problems raised in the above background technique.

[0006] To achieve the above purpose, the present invention provides the following technical solution: An underwater target detection method based on improved YOLOv5s, specifically including the following steps:

[0007] S1: Establish an underwater target data set and label it;

[0008] S2: Design a new attention module MSDW and integrate it into the C3 module to form a new C3MSDW module, and the C3MSDW module replaces the C3 module of the backbone;

[0009] S3: Reconstruct the SPPF pyramid pooling;

[0010] S4: Improve the BiFPN weighted bidirectional pyramid network based on the GFPN mechanism idea;

[0011] S5: Establish the Focaler-Shape loss function to obtain an improved YOLOv5s underwater target detection network;

[0012] S6: Input the established dataset into the improved model for training to obtain an improved YOLOv5s underwater target detection model;

[0013] S7: Input the image or video to be detected into the trained improved YOLOv5s underwater target detection model for detection, and output the detection results with confidence, bounding boxes, and class information.

[0014] Preferably, S2 further includes S2-1 and S2-2;

[0015] S2-1: Design a new attention module MSDW; input the input feature map into the channel attention module MSCA, and the channel attention module MSCA will output the output feature map after channel attention processing. Then multiply the input feature map by the output feature map after channel attention, and input the obtained output result into the spatial attention module DW;

[0016] S2-2: Combine the attention module MSDW generated in S2-1 with the C3 module of the backbone to generate the C3MSDW module. The specific process includes: the feature map is subjected to preliminary feature extraction and transformation in two parallel CBS module branches. The output of one branch enters the BottleNeck structure for output, and the outputs of the two branches are concatenated through the Concat operation in the channel dimension to fuse the feature information to obtain the concatenated feature map; the concatenated feature map enters the attention module MSDW, and finally is further processed and transformed by one of the CBS modules to obtain the transformed output feature map.

[0017] Preferably, S3 includes S3-1;

[0018] S3-1: First, the transformed output feature map is preliminarily transformed by the Conv convolutional layer, and then sequentially passes through three MaxPooling maximum pooling layers to obtain features of different scales; subsequently, two branches respectively perform adaptive average pooling and maximum pooling to adjust the feature map size and extract features; the feature maps processed by different paths are concatenated in the channel dimension through Concat, and then transformed by the final Conv convolutional layer to output the feature map.

[0019] Preferably, improve the BiFPN weighted bidirectional pyramid network based on the GFPN mechanism idea, and the content is as follows:

[0020] S4-1: Replace the FPN in the original YOLOv5s network model with the weighted bidirectional pyramid network BiFPN;

[0021] S4-2: Retain the original cross-stage skip connection fusion method of BiFPN. On this basis, improve the BiFPN pyramid structure based on the GFPN mechanism idea, set multi-sided input nodes, and add cross-scale connections to enhance feature transmission between different scales.

[0022] Preferably, combine the FocalerIOU loss function and the ShapeIOU loss function to generate Focaler-ShapeIOU, and the formula is as follows:

[0023] F shape-IOU = 1 - F IOU + λ shape + 0.5 × Ω shape ;

[0024]

[0025] F focaier-IOU = 1 - F focaier ;

[0026] F focaier-shape = F shape + F IOU - F focaier 。

[0027] Compared with the prior art, the beneficial effects of the present invention are as follows: The underwater target detection method based on the improved YOLOv5s improves the problems of low detection accuracy and high false detection rate caused by environmental interference and target blur encountered in underwater target detection by designing a new attention module MSDW and integrating it into C3 to form a new module C3MSDW, reconstructing the SPPF module structure, improving BiFPN based on the GFPN mechanism, and establishing a Focaler-Shape loss function to optimize the training process. This method can effectively improve the comprehensive performance in the underwater blurred target image detection task, greatly improve the accuracy and recall rate of underwater target detection, and has important significance for the development of the underwater aquaculture industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is the flow chart of the underwater target detection method based on the improved YOLOv5s in the present invention;

[0029] Figure 2 is the network structure diagram of the underwater target detection based on the improved YOLOv5s in the present invention;

[0030] Figure 3 is the structure diagram of the MSDW module in the present invention;

[0031] Figure 4 This is the structural diagram of the MSCA module in the present invention;

[0032] Figure 5 This is the Depth-Wise structural diagram of the spatial attention module in the present invention;

[0033] Figure 6 This is the schematic diagram of the C3MSDW structure in the present invention;

[0034] Figure 7 This is the schematic diagram of the Improve-SPPF structure in the present invention;

[0035] Figure 8 This is the schematic diagram of the original BiFPN structure in the present invention;

[0036] Figure 9 This is the schematic diagram of the improved BiFPN structure in the present invention. Detailed implementation manners

[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0038] Please refer to Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6 , Figure 7 , Figure 8 , Figure 9 , the present invention provides a technical solution: an underwater target detection method based on improved YOLOv5s. The following explains the present invention in conjunction with the accompanying drawings:

[0039] As Figure 1 shown in the flowchart of the method of the present invention, an underwater target detection algorithm based on improved YOLOv5s is characterized in that the method steps are as follows:

[0040] S1: Establish an underwater target data set and label it;

[0041] S2: Design a new attention module MSDW and integrate it into the C3 module to form a new C3MSDW module, and the C3MSDW module replaces the C3 module of the backbone;

[0042] S3: Reconstruct the SPPF pyramid pooling;

[0043] S4: Improve the BiFPN weighted bidirectional pyramid network based on the idea of the GFPN mechanism;

[0044] S5: Establish the Focaler-Shape loss function to obtain an improved YOLOv5s underwater target detection network;

[0045] S6: Input the established dataset into the improved model for training to obtain an improved YOLOv5s underwater target detection model;

[0046] S7: Input the image or video to be detected into the trained improved YOLOv5s underwater target detection model for detection, and output the detection results with confidence, bounding boxes, and class information.

[0047] S1 uses the URPC underwater dataset and accurately annotates the aquatic treasure targets in it using annotation tools. The annotation content involves information such as the category, precise location, and size range of the targets, providing rich and high-quality samples for subsequent model training.

[0048] In the research fields of underwater robots and environmental monitoring, etc., the URPC (Underwater Robotics and Perception Challenge) underwater dataset is a very useful resource. This dataset aims to help researchers and developers develop underwater robot technologies, including underwater positioning, navigation, perception, and communication, etc. URPC provides various types of tasks, such as underwater object recognition, underwater terrain mapping, and environmental monitoring, etc.

[0049] The main characteristics of the URPC dataset include:

[0050] Diverse scenarios: The dataset contains various underwater scenarios, such as ports, rivers, lakes, etc., to simulate different working environments.

[0051] High-quality images and videos: Provide high-resolution image and video data, which is crucial for training and testing machine vision algorithms.

[0052] Annotation information: The dataset contains annotation information, such as the position, shape, size of objects, etc., which is very helpful for training and validating machine learning models.

[0053] Challenge tasks: The dataset aims to solve practical underwater robot challenges, such as underwater obstacle detection, underwater terrain mapping, etc.

[0054] How to specifically use the URPC dataset belongs to the prior art and will not be elaborated here.

[0055] Such as Figure 3As shown, S2 uses the multi-scale convolutional attention module MSCA as the channel attention module MSCA in this application in the design of the new attention module MSDW. Among them, the channel attention module MSCA of this application is as follows Figure 4 As shown in the figure, the channel attention module MSCA uses deep convolution to deeply mine the local information of the feature map, captures contextual information from multiple scales through multi-branch deep strip convolution, and then uses 1×1 convolution to sort out the relationship between different channels, thereby achieving effective modeling of channel attention.

[0056] It should be noted here that the feature map in S2-2 is the union of the input feature map and the output feature map in S2-1.

[0057] For the spatial attention multi-scale depth-wise separable convolution module, Figure 5 As shown in the figure, after receiving the input feature map, the spatial attention module Depth-Wise uses convolution kernels of different sizes (such as 5×5, 1×7, etc.) to extract spatial features from multiple scales, adds and fuses them after unifying the number of channels through 1×1 convolution, and finally generates an attention map that can highlight the key spatial area through the Sigmoid activation function. In actual operation, Figure 3 As shown, after the feature map is input into the channel attention module MSCA, the output feature map obtained after channel attention processing is multiplied with the input feature map, and the product result is input into the channel attention module MSCA again. After another multiplication operation, the output feature map processed by the attention module MSDW is finally obtained.

[0058] C3MSDW module structure is as follows Figure 6 As shown in the figure, when constructing the C3MSDW module, the feature map enters two parallel CBS module branches, where preliminary feature extraction and transformation operations are performed. The output of one of the branches enters the BottleNeck structure for further processing and output, and the outputs of the two branches are spliced ​​through the Concat operation in the channel dimension to achieve the fusion of feature information. The fused feature map enters the attention module MSDW for deep processing, and finally passes through a CBS module for final processing and transformation, and outputs the final feature map, thereby completing the replacement of the original C3 module and enhancing the model's perception and fusion capabilities of marine treasures of different scales.

[0059] The original SPPF in S3 only relies on three 5×5 maximum pooling layers to extract features, which has limitations. The Improve-SPPF module reconstructed by the present invention is as follows: Figure 7As shown in the figure, first let the feature map pass through the Conv convolution layer for preliminary feature transformation to adjust the distribution and expression of the features. Then, it passes through three MaxPool2d maximum pooling layers in sequence to extract multi-scale features from different levels. Subsequently, the two branches perform adaptive average pooling and maximum pooling operations respectively. Adaptive average pooling can adjust the pooling window size according to the local statistical information of the feature map, and maximum pooling highlights the local maximum value features. The two work together to adjust the size of the feature map and extract more representative features. The feature maps processed by different paths are spliced ​​in the channel dimension through the Concat operation to integrate the global background and edge information. Finally, they are transformed by the final Conv convolution layer to output a feature map that integrates rich information, effectively improving the detection accuracy and reducing the detection difficulty brought by targets of different scales.

[0060] The underwater target detection method based on improved YOLOv5s is characterized in that S4 improves the BiFPN weighted bidirectional pyramid structure based on the GFPN mechanism idea:

[0061] The FPN in the original YOLOv5s network model is replaced with a weighted bidirectional pyramid network. This network structure can establish a more efficient connection and fusion mechanism between feature maps of different scales, achieving full transmission and effective utilization of features. Figure 8 The original BiFPN cross-stage jump connection fusion method shown in the figure ensures that important feature information can be quickly transmitted between different stages of the network, maintaining the coherence and integrity of the features.

[0062] The BiFPN pyramid structure is improved based on the GFPN mechanism idea. The improved structure is as follows Figure 9 As shown in the figure. By setting up multilateral input nodes, the single input limitation of the traditional structure is broken, and the source and diversity of features are increased. At the same time, cross-scale connections are added, so that features of different scales can interact and merge more directly and frequently, further enhancing the model's adaptability to targets of different scales, especially significantly improving the detection performance of small targets, so that the model can more accurately detect marine treasures of various sizes in complex underwater environments.

[0063] In S5, the FocalerIOU loss function and the ShapeIOU loss function are cleverly combined to generate the Focaler-ShapeIOU loss function. The ShapeIOU loss function and the FocalerIOU loss function are combined for underwater marine treasure detection, which can improve the detection accuracy, comprehensively consider the shape information and sample difficulty, accurately distinguish targets and avoid missed detections; enhance the robustness of the model and better adapt to complex underwater environments; improve training efficiency, guide the model to focus on difficult samples and adjust parameters in a targeted manner; and better handle small targets and occlusion problems to reduce missed detections.

[0064] The formula of the Focaler-Shape IOU loss function is as follows:

[0065] F shape-IOU = 1 - F IOU + λ shape + 0.5 × Ω shape ;

[0066]

[0067] F focaier-IOU = 1 - F focaier ;

[0068] F focaier-shape = F shape + F IOU - F focaier .

[0069] Input the dataset established by S1 into the improved model for training. During the training process, reasonably select training parameters according to the characteristics of the dataset and the structure of the model.

[0070] Input the image or video to be detected into the trained underwater target detection model for detection: The model will analyze and process the input image or video according to the knowledge and rules learned during the training process, and output detection results with information such as confidence, interest box, and category. The confidence represents the reliability evaluation of the model for the detection results. The interest box accurately defines the position and range of the target. The category information clarifies the type to which the target belongs, providing accurate and reliable target positioning and recognition information for practical applications such as underwater sea treasure fishing, and strongly supporting the efficient development of underwater operations.

[0071] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.

[0072] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. An underwater target detection method based on improved YOLOv5s, characterized in that The specific steps include: S1: Establish and label underwater target dataset; S2: Design a new attention module MSDW to integrate it into the C3 module to form a new C3MSDW module, and the C3MSDW module replaces the C3 module of the backbone; S3: Reconstruct SPPF pyramid pooling; S4: Improve BiFPN weighted bidirectional pyramid network based on the idea of ​​GFPN mechanism; S5: Establish a YOLOv5s underwater target detection network with improved Focaler-Shape loss function; S6: Enter the established data set into the improved model for training, and obtain the improved YOLOv5s underwater object detection model; S7: The images or videos to be detected are input into the trained and improved YOLOv5s underwater object detection model for detection, and the detection results with confidence, interest boxes, and category information are output.

2. An underwater target detection method based on improved YOLOv5s according to claim 1, characterized in that Said S2 also includes S2-1 and S2-2; S2-1: Design a new attention module MSDW; pass the input feature map to the channel attention module MSCA, which will output the output feature map after channel attention processing, and then multiply the input feature map with the output feature map after channel attention, and input the output result into the spatial attention module DW; S2-2: The attention module MSDW generated in S2-1 is combined with the C3 module of the backbone to generate the C3MSDW module. The specific process includes: the feature map is initially extracted and transformed in two parallel CBS module branches, one of which outputs enters the BottleNeck structure output, and the outputs of the two branches are spliced ​​in the channel dimension through the Concat operation to fuse the feature information to obtain the spliced ​​feature map; the spliced ​​feature map enters the attention module MSDW, and finally is further processed and transformed by one of the CBS modules to obtain the transformed output feature map.

3. An underwater target detection method based on improved YOLOv5s according to claim 1, characterized in that S3 includes S3-1; S3-1: The transformed output feature map is first transformed by the Conv convolution layer, and then passes through three MaxPooling layers in sequence to obtain features of different scales; then, the two branches perform adaptive average pooling and maximum pooling respectively to adjust the feature map size and extract features; The feature maps processed by different paths are spliced ​​in the channel dimension through Concat, and then transformed through the Conv convolution layer at the end to output the feature map.

4. An underwater target detection method based on improved YOLOv5s according to claim 1, characterized in that Based on the idea of ​​GFPN mechanism, the BiFPN weighted bidirectional pyramid network is improved, and its content is as follows: S4-1: Replace FPN in YOLOv5s original network model with weighted bidirectional pyramid network BiFPN; S4-2: The original BiFPN cross-stage jump connection fusion method is retained. On this basis, the BiFPN pyramid structure is improved based on the GFPN mechanism idea, multilateral input nodes are set, and cross-scale connections are added to enhance the feature transfer between different scales.

5. An underwater target detection method based on improved YOLOv5s according to claim 1, characterized in that Combine FocalerIOU loss function with ShapeIOU loss function to generate Focaler-ShapeIOU. The formula is as follows: F shape-IOU =1-F IOU +λ shape +0.5×Ω shape ; F focaier-IOU =1-F focaier ; F focaier-shape =F shape +F IOU -F focaier 。