Side-scan sonar shipwreck target identification method based on improved Yolov5 network model
By improving the Yolov5 network model, increasing the CBAM attention mechanism and using the GhostNet network, the problem of lateral sweep sonar image target recognition in the prior art is solved, and more efficient and accurate detection of subsea wreck targets is achieved.
Patent Information
- Application Number
- CN202510067989.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-16
AI Technical Summary
The existing side-scan sonar image target recognition methods rely on personnel experience, are susceptible to subjective factors, and are inefficient, especially in small or no samples, which is difficult to effectively identify subsea shipwreck targets.
Using the improved Yolov5 network model, by increasing the CBAM attention mechanism and replacing the backbone network with the GhostNet network, the model's ability to detect targets is improved, and anti-interference ability and recognition efficiency are improved.
It effectively improves the detection ability of small and medium-sized shipwreck targets in side-sweep sonar images, improves the impact of complex seabed backgrounds on target detection, and improves the anti-interference ability and recognition accuracy of the model.
Smart Images

Figure CN120014425A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of side-scan sonar image target recognition, and in particular to side-scan sonar image sunken ship target recognition using a deep learning algorithm. Background Art
[0002] As the main device for acquiring underwater topographic images, side scan sonar has the advantages of low cost, high efficiency, and high imaging resolution. The side scan sonar converts the strength, number, and time interval of the echo reflected by the seabed target into the length and grayscale change of the instrument recording line, and transmits and receives multiple times to form a side scan sonar sound map. Side scan sonar imaging has the advantages of long range and strong penetration ability, and is particularly suitable for muddy waters. Therefore, it has been widely used in underwater geological and geomorphological surveys, underwater lost object searches, dam foundation inspections and other fields.
[0003] The traditional sonar image target recognition process is to first extract the target slices of interest (ROI) from the sonar image; then, segment the target in the slice, and then further extract the features of the segmented image; finally, use the classifier to classify and identify the extracted features. Since the targets to be identified are mostly metal structures, they will produce strong backscattered echoes, showing a highlight area in the sound image; at the same time, the sound waves cannot be received due to the occlusion of the target, forming a shadow area behind it. Usually, the sound wave irradiation causes the conjugation of highlights and shadows, and the position of the shadow projection is geometrically related to the highlight of the target and the height of the sonar. Based on the sonar acquisition principle and the prior knowledge of the target, scholars have designed a series of target judgment rules and templates to achieve underwater target recognition. The development of deep learning provides a new solution to the traditional target recognition process that is limited by the experience of the staff. Considering the two-dimensional sonar image target recognition from the perspective of deep learning, it can usually be decomposed into two subtasks: classification recognition and target detection. Classification recognition is to assign ROI to predefined classes; target detection is to integrate the location, classification and range estimation of the object into one step. In order to improve the recognition performance, some scholars have combined the characteristics of acoustic image targets to improve and optimize the traditional deep convolution model. In the field of traditional underwater target recognition, the concept of detection is mostly used to describe whether there is a target in the image. It is often used as a pre-process for classification and recognition to locate the ROI of the target. However, deep learning target detection directly outputs the location and type of the object in the form of a bounding box.
[0004] Traditional underwater target detection and recognition methods have proposed solutions for underwater target recognition in the case of small or no samples. However, traditional machine learning methods are limited by the processing capacity of acoustic image data, and require professional and in-depth understanding of acoustic image data to design appropriate feature representation methods. Compared with traditional machine learning, deep learning does not require manual extraction or manual creation of features. It builds cognition of underwater environment and underwater targets through data-driven, which is suitable for underwater detection fields where human cognition of the environment is relatively scarce. In the work of underwater target recognition in sonar images, the research method has transitioned from traditional feature extraction to deep learning; the research object has expanded from small underwater targets to multi-scale targets; and the research scene has changed from flat seabed to complex landforms. Improving the accuracy of underwater target recognition in two-dimensional sonar images by designing more reliable underwater target features, generating high-quality sample data, and adopting more advanced network structures has become a hot direction for side-scan sonar acoustic image target recognition. Summary of the invention
[0005] The present invention mainly focuses on the problem that side-scan sonar image target recognition mainly relies on personnel experience, is affected by subjective factors and has low efficiency. Considering making full use of existing side-scan sonar acoustic image data, while taking into account the seabed background field and the characteristics of the sunken ship target, a side-scan sonar sunken ship target recognition method based on the improved Yolov5 network model is proposed. The improved method consists of adding a CBAM attention mechanism (Attention Mechanism) and replacing the backbone network with a GhostNet network. The CBAM attention mechanism is used to improve the model's detection ability for targets from two dimensions: space and channel. The GhostNet network effectively utilizes redundant feature maps while achieving lightweight network structure.
[0006] In order to achieve the above object, the technical solution of the present invention is:
[0007] A side-scan sonar shipwreck target recognition method based on an improved Yolov5 network model specifically comprises the following steps:
[0008] The first step is data preprocessing and data set division
[0009] Data preprocessing mainly uses data enhancement methods, including image cropping (cutout), image noise addition, changing image brightness, image translation, image rotation, and image mirroring.
[0010] After data preprocessing, the dataset is randomly divided into training set, evaluation set and test set according to the proportion.
[0011] Step 2: Improvement of network model
[0012] Improvement method 1: The backbone network Backbone in Yolov5 is used to extract features of the input image. The Backbone is replaced with the GhostNet network, that is, the original convolution process is replaced with a GhostNet network composed of n stacked Ghost bottleneck layers; the input data of the GhostNet network first passes through a 3×3 convolution layer, then through n Ghost Bottleneck layers, and then passes through a 1×1 convolution kernel, a global average pooling, a 1×1 convolution kernel, and finally flattened through the fully connected layer output.
[0013] Furthermore, in the replaced Backbone part, the input layer (Input) includes 2 convolutional layers (CBS) and 15 Ghost bottleneck layers, and the entire Backbone will get three interfaces at the 5th, 9th, and 16th layers respectively connected to the Neck part.
[0014] Improvement method 2: In the backbone network Backbone in Yolov5, make improvements by any of the following methods:
[0015] (1) The CBAM module is inserted into the C3 (CSP) module in Backbone, and the attention mechanism is introduced after the feature fusion layer (Concat), and then output through the convolution layer.
[0016] (2) Insert the attention mechanism layer directly before the SPPF (spatial pyramid pooling layer) of Backbone.
[0017] The improvement of the network model is to select one of the methods or to adopt both methods.
[0018] In the third step, the preprocessed training sets are input into the two improved network models for training respectively, and the evaluation sets are used to obtain the best results which are retained as the training results.
[0019] Step 4: Network model prediction
[0020] The test set after data preprocessing is input into the trained model to obtain the prediction results of the network model, and the effect of the improved network model is evaluated based on multiple indicators.
[0021] The beneficial effects of the present invention are as follows: a side-scan sonar shipwreck target recognition method based on an improved Yolov5 network model provided by the present invention integrates a convolutional neural network, a GhostNet network and an attention mechanism. While effectively improving the detection capability of small and medium-sized shipwreck targets in side-scan sonar images, it focuses on the impact of complex seabed background fields on target detection, effectively improving the model's anti-interference capability, and can be used as a means of quickly and accurately processing large quantities of data when performing side-scan sonar image shipwreck target recognition, and has practical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is the structure diagram of the Yolov5 network model;
[0023] Figure 2 This is the structure diagram of the Yolov5m_GhostNet network model;
[0024] Figure 3 This is the CBAM module structure diagram;
[0025] Figure 4 The CSP layer before and after adding the CBAM module;
[0026] Figure 5 PR curve of the model after introducing the CBAM attention mechanism;
[0027] Figure 6 The PR curves of the model detection after replacing the backbone network, where (a) is the PR curve of the Yolov5m_GhostNet model, and (b) is the PR curve of the Yolov5l_GhostNet model. DETAILED DESCRIPTION
[0028] In order to make the method problems solved by the present invention, the method solutions adopted and the method effects achieved clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and experiments. It is understood that the specific experiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for the convenience of description, only the parts related to the present invention are shown in the accompanying drawings, rather than all the contents.
[0029] A side-scan sonar shipwreck target recognition method based on an improved Yolov5 network model specifically comprises the following steps:
[0030] The first step is data preprocessing and data set division
[0031] The side-scan sonar images used for training and testing of the shipwreck targets in this experiment were mainly obtained by mainstream side-scan sonar equipment used at home and abroad, including Klein3000, Klein3900, etc. Most of them were side-scan sonar shipwreck images collected on the Internet. Finally, 1,114 side-scan sonar images with seabed shipwreck targets were selected.
[0032] Due to the small number of submarine shipwreck images, complex backgrounds, and a small proportion of the target area, if only this dataset is used to train the network, the model may be overfitted. Therefore, the dataset needs to be preprocessed accordingly.
[0033] Data preprocessing mainly uses data enhancement methods, including image cropping (cutout), image noise, image brightness change, image translation, image rotation, image mirroring, etc. Image cropping refers to the method of randomly deleting a rectangular area in the image to force the neural network to extract image features more accurately and finely. By adding noise such as Gaussian noise and salt and pepper noise, the feature image can be made more consistent with the real side-scan sonar sound image, so that the model can be more adaptable to the learning and prediction of complex seabed acoustic images. Because the seabed environment itself has complex and changeable characteristics, the original data obtained from it has different or even considerable noise. Changing the image brightness is a kind of color jittering method. Color jittering is to generate new images by randomly adjusting the saturation, brightness, and contrast of the original image, so as to achieve the purpose of increasing the data set. Image translation usually refers to moving the image in two directions of the X or Y axis. Appropriate translation operation can make the shipwreck target almost anywhere in the image, thus forcing the network to retrieve all corners of the image and learn the characteristics of the target image under different distribution conditions.
[0034] After data preprocessing, the data set is randomly divided into training set, evaluation set and test set in a ratio of 7:1:2. According to the characteristics of the data, this experiment expands the 1114 input images to 5570 through data enhancement, and expands 779 images in the training set to 3895 images, 112 images in the evaluation set to 560 images, and 223 images in the test set to 1115 images through random division.
[0035] Step 2: Improvement of network model
[0036] The structure of Yolov5 neural network before improvement is as follows Figure 1As shown in the figure, it includes input, Backbone, Neck and Head. The input reads the dataset information and performs image scale transformation, unifies the image data of different sizes and sends them to the improved network model; Backbone is the backbone network part, which is used to extract features from the input image; Neck is used for feature fusion; Head is used for model detection. In the Backbone part, there are 5 convolutional layers (CBS) and 4 C3 layers (CSP) between the input layer (input) and the pooling layer (spatial pyramid pooling SSPF). The entire Backbone will get three interfaces at the 4th, 6th and 9th layers.
[0037] Improvement method 1. The CBAM attention mechanism consists of a channel attention mechanism (channel) and a spatial attention mechanism (spatial). The channel attention mechanism is used to improve the model's ability to allocate feature map channels, while the spatial attention mechanism is used to make the model pay more attention to important pixel areas on the feature map (areas that play a decisive role in classification) and ignore other low-value areas. In order to improve the performance of the model from the two dimensions of space and channel, the CBAM attention module is divided into two parts: the channel attention mechanism module (CAM) and the spatial attention mechanism module (SAM). The CBAM attention mechanism has a significant effect on improving the model's ability to detect small targets. There are generally two ways to add the CBAM attention mechanism: directly insert the attention mechanism layer before the SPPF (spatial pyramid pooling layer) of the backbone network or insert the attention mechanism module into the C3 module. This experiment uses the second method, that is, inserting the CBAM module into the C3 module in the Backbone, performing the attention mechanism after Concat, i.e. feature fusion, and then outputting through the convolutional layer to realize the attention mechanism.
[0038] Improvement method 2: The GhostNet network treats redundant feature maps in a different way from other lightweight networks. Since redundant feature maps can enhance a network model's ability to understand graphic features, redundant feature maps are also an indispensable part of a successful model. Most lightweight networks remove redundant feature maps, while the GhostNet network chooses a low-cost method to retain them. The Ghost basic unit uses a series of linear transformations to generate feature maps instead of using convolution to generate feature maps, which reduces the computational complexity and the number of parameters while generating feature maps of the same size as standard convolution. The GhostNet network is mainly composed of n GhostBottleneck stacks. The input data first passes through a 3×3 convolution layer, then passes through n Ghost Bottlenecks, and then passes through a 1×1 convolution kernel, a global average pooling, a 1×1 convolution kernel, and finally is flattened and output through the fully connected layer. In this embodiment, in the replaced Backbone part, the input layer (Input) includes 2 convolution layers (CBS, one after the input layer and one at the end of all Ghost bottleneck layers) and 15 Ghost bottleneck layers. The entire Backbone will obtain three interfaces at the 5th, 9th, and 16th layers, which are respectively connected to the Neck part.
[0039] Step 3: Model training
[0040] The preprocessed training sets are input into the two improved network models for training respectively, and the evaluation sets are used to obtain the best results which are retained as the training results.
[0041] The input end reads the dataset information and performs image scale transformation, unifies the image data of different sizes into a unified format and sends them to the improved network model;
[0042] Backbone is the backbone network part, which is used to extract features from the input image;
[0043] The Neck part adopts the FPN+PAN structure, which mainly includes the convolution layer (CBS), upsampling layer (Upsample), C3 layer and Concat (scale feature fusion layer), etc.
[0044] The Head part is mainly used for model detection and contains three detection layers of different scales.
[0045] Step 4: Network model prediction
[0046] The test set after data preprocessing is input into the trained model to obtain the model's prediction results, and the effect of the improved model is evaluated based on a variety of indicators.
[0047] In order to comprehensively and objectively evaluate the prediction effects of different models, this paper selects Precision, Recall and AP as model quality evaluation indicators. Precision is the ratio of the number of correct positive samples identified by the model to the number of all positive samples identified by the model, and recall is the ratio of the number of correct positive samples identified by the model to the number of true positive samples. In experiments, we always hope that both precision and recall can achieve the highest results, but in fact the two are contradictory. Excessive pursuit of precision will lead to a decrease in recall, and vice versa. Therefore, a comprehensive analysis is usually made with the help of the PR curve. mAP is the average precision, which refers to the area between the bottom of the PR curve and the coordinate axis. The calculation formulas for the two are as follows:
[0048]
[0049] Among them, TP (True Positive) is the number of samples whose true category is positive samples and the model predicts that the samples are also positive samples. FP (False Positive) is the number of samples whose true category is negative samples and the model predicts that the samples are positive samples. FN (False Negative) is the number of samples whose true category is positive samples and the model predicts that the samples are negative samples.
[0050] In order to verify the effectiveness and superiority of the present invention, when detecting shipwreck targets in complex seabed background, the above experimental process is synchronously executed using Yolov5m, Yolov5m_CBAM, Yolov5m_GhostNet, Yolov5l_GhostNet, and models, and the prediction results of different models are compared. The results are shown in Table 1:
[0051] Table 1 Prediction evaluation results of different models.
[0052]
[0053] From Table 1, we can see that the introduction of the CBAM attention mechanism effectively improves the recall rate of the model to 94.8%; replacing the backbone network with the GhostNet network makes the model mAP reach the highest 0.974 and the precision rate reaches 97%. Although the two improvements have decreased in precision and recall rate, in general, no matter which evaluation indicator is used, the performance of the improved model is better than the original Yolov5 model.
[0054] The prediction results of the improved model should also be analyzed through the trend and distribution of the PR curve.
[0055] Figure 5This is the PR curve of model detection after the introduction of the CBAM attention mechanism. Overall, the PR curve of the model detection results shows a trend approaching the right, which indicates that the improved model has more successful detections of positive samples than the original model.
[0056] Figure 6 This is the PR curve of model detection after replacing the backbone network. Overall, the trend of the PR curve of the model detection result is already quite close to the upper right corner. This result shows that the precision and recall rates of the Yolov5l_GhostNet model are both stable at a very high level.
[0057] Finally, it should be noted that the above experiments are only used to illustrate the method scheme of the present invention, rather than to limit it. Although the present invention has been described in detail, ordinary method personnel in the field should understand that modifying the aforementioned method scheme, or replacing part or all of the method features therein by equivalents, does not deviate the essence of the corresponding method scheme from the scope of the method scheme of the present invention.
Claims
1. A side-scan sonar shipwreck target recognition method based on an improved Yolov5 network model, characterized in that: The specific steps include: The first step is data preprocessing and data set division Step 2: Improvement of network model Improvement method 1: The backbone network Backbone in Yolov5 is used to extract features from the input image. Backbone is replaced with the GhostNet network, that is, the original convolution process is replaced with a GhostNet network consisting of n stacked Ghost bottleneck layers; the input data of the GhostNet network first passes through a 3×3 convolution layer, then passes through n Ghost Bottleneck layers, and then passes through a 1×1 convolution kernel, a global average pooling, a 1×1 convolution kernel, and finally flattens through the fully connected layer output; Improvement method 2: In the backbone network Backbone in Yolov5, make improvements by any of the following methods: (1) Insert the CBAM module into the C3 module in Backbone, introduce the attention mechanism after the feature fusion layer, and then output it through the convolution layer; (2) Insert the attention mechanism layer directly before Backbone’s SPPF; In the third step, the preprocessed training set is input into the two improved network models for training, and the best result is obtained by using the evaluation set and retained as the training result; Step 4: Network model prediction The test set after data preprocessing is input into the trained model to obtain the prediction results of the network model, and the effect of the improved network model is evaluated based on multiple indicators.
2. The side-scan sonar shipwreck target recognition method based on the improved Yolov5 network model according to claim 1 is characterized in that: Data preprocessing mainly uses data enhancement methods, including image cropping, image noise addition, changing image brightness, image translation, image rotation, and image mirroring.
3. The side-scan sonar shipwreck target recognition method based on the improved Yolov5 network model according to claim 1 is characterized in that: After data preprocessing, the dataset is randomly divided into training set, evaluation set and test set according to the proportion.
4. The side-scan sonar shipwreck target recognition method based on the improved Yolov5 network model according to claim 1 is characterized in that: In the replaced Backbone part, the input layer includes 2 convolutional layers and 15 Ghostbottleneck layers. The entire Backbone will get three interfaces at the 5th, 9th, and 16th layers respectively connected to the Neck part.
5. The side-scan sonar shipwreck target recognition method based on the improved Yolov5 network model according to claim 1 is characterized in that: The improvement of the network model is to select one of the methods or to adopt both methods.