Seabed target detection method based on double-branch lightweight trunk

By adopting dual-branch lightweight backbone and sliding window slicing technology in the side-sweep sonar image subsea target detection model, the problem of large number of model parameters and low detection efficiency is solved, and efficient and real-time subsea target detection capabilities are achieved, which is suitable for mobile platforms such as small AUVs.

CN119992304AActive Publication Date: 2025-05-13PLA DALIAN NAVAL ACADEMY

Patent Information

Application Number
CN202510068385.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13
Estimated Expiration
2045-01-16

AI Technical Summary

Technical Problem

The existing side-scan sonar image subsea target detection model has large parameters and low detection efficiency, making it difficult to deploy to mobile platforms such as small AUVs, and the detection accuracy and strategies are difficult to meet the needs of engineering deployment.

Method used

A two-branch lightweight backbone subsea object detection model DBnet is proposed. It adopts a lightweight network backbone and a streamlined feature fusion structure, and features are extracted through a dual backbone network, and a sliding window slicing technology is introduced to improve the target detection capability at small and medium-sized scales.

Benefits of technology

It effectively reduces the amount of model parameters and calculations, improves detection efficiency and accuracy, enhances the detection ability of targets at different scales, and can be deployed to underwater mobile platforms such as AUV to achieve real-time intelligent detection of subsea targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992304A_ABST
    Figure CN119992304A_ABST
Patent Text Reader

Abstract

The invention discloses a seabed target detection method based on a double-branch lightweight trunk, and belongs to the field of side-scan sonar image seabed target detection research. A double-branch lightweight trunk model DBnet is designed, two lightweight trunk networks, namely PP-LCNet and GhostNet, with excellent performance are used in a feature extraction stage to extract features of an input image, and features of all levels extracted by double trunks can still be well fused while parameter quantity is reduced through a simplified Neck structure. And the detection capability of different-scale targets is enhanced. In the reasoning stage, an SAHI method is adopted, according to the actual underwater operation of the AUV, the whole waterfall plot is cut into a plurality of slices with a certain overlapping ratio and a fixed size, and then the slices are input into the network, so that the problems that target detection models of unmanned platforms such as the AUV are difficult to deploy, and the detection efficiency and the detection precision are low are effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the research field of seabed target detection in side-scan sonar images, and in particular relates to a lightweight method for seabed target detection models. Background Art

[0002] The detection and identification of seabed targets plays an extremely important role in underwater search and rescue, marine engineering construction, marine topography and geomorphology survey, marine resource survey and other fields. However, due to the complex marine environment, imaging conditions and measurement methods, its accuracy and efficiency are relatively low, which is difficult to meet the needs. It has become a hot spot and difficulty in current research. Side-scan sonar can detect a large range of seabed areas in a short time. At the same time, due to its acoustic detection method, it is not affected by factors such as underwater visibility and has low cost. Therefore, it is widely used in seabed target detection and plays an important role in searching for crashed aircraft, ships and crashed personnel, locating seabed pipelines, etc. However, the current recognition of seabed targets in side-scan sonar images is still mainly manual, which is overly dependent on subjective experience and inefficient, which seriously affects its wide application in seabed target detection. It is difficult to be installed on unmanned equipment such as cable-free underwater robots (AUVs) with limited computing power and high requirements for detection efficiency and real-time performance. Therefore, how to improve detection accuracy and efficiency is an important task in side-scan sonar measurement.

[0003] With the rapid development of computer vision and the emergence of convolutional neural networks, deep learning-based methods have been widely used in the field of seabed target detection in side-scan sonar images. Existing target detection models are generally divided into two categories: two-stage detection models and one-stage detection models. However, because the two-stage model needs to generate a region proposal network for candidate target boxes, and then classify and regress these candidate boxes, target detection is achieved through two processes. There are problems such as complex structure, huge amount of calculation, and long calculation time. Therefore, it is rarely used in actual engineering applications. The core advantage of the single-stage target detection model is first reflected in speed. By simplifying the detection process and reducing the consumption of computing resources, efficient and instant detection is achieved, which is especially suitable for application scenarios with high real-time requirements. Secondly, its structure is relatively simple, easy to implement and deploy, while reducing the hyperparameters that need to be adjusted and the complexity of model tuning. In addition, the single-stage model also shows strong adaptability and flexibility, can be easily adjusted to meet different task requirements, and is easy to integrate with other image processing systems or platforms.

[0004] Although existing research has improved the detection efficiency and accuracy of target detection models to a certain extent, it is still difficult to meet actual needs in engineering deployment. Taking the side-scan sonar carried by AUV as an example, in order to meet the detection efficiency, image quality and carrier endurance, the side-scan sonar is usually used for seabed detection at an economic speed of 2-3 knots. This requires the target detection model to improve the detection efficiency on the basis of high accuracy to meet the real-time detection requirements. Moreover, due to the limitations of acoustic detection principles, working modes and other factors, it is difficult to obtain different modal data in the same area for seabed target detection compared with remote sensing target detection. That is, the non-repeatability of data acquisition and the singleness of the source lead to the inability to integrate multimodal information like the remote sensing detection field during feature extraction learning, which restricts the improvement of the accuracy of the target detection model. Moreover, since the images measured by the side-scan sonar are superimposed by the data of each frame (ping), they are usually large in size and unbalanced in length and width. If the entire image is input into the network, the image will be compressed, making it difficult to detect small targets in the image. However, small seabed targets often arranged on the seabed are of great significance for emergency search and rescue. Summary of the invention

[0005] The present invention mainly targets the side-scan sonar image target detection model, which has a large number of parameters and is not easy to deploy to small AUV and other mobile platforms, and the detection accuracy and detection strategy are difficult to meet the engineering deployment requirements. A submarine target detection model named DBnet is proposed. First, in order to solve the problem of large model parameter detection, low efficiency and difficulty in deployment, we use a lightweight network backbone and a streamlined feature fusion structure to greatly reduce model parameters and calculations. Secondly, in order to alleviate the problem of reduced accuracy caused by lightweight model and the single source of detection data seriously limiting the performance of the model, we designed a dual-branch structure, that is, feature extraction through a dual-trunk network, effectively obtaining richer semantic information, and finally, because the size of the original waterfall image of the side-scan sonar is usually large, the resampling (reshape) operation in the input network will cause some small and medium-scale targets to be difficult to detect. We introduce sliding window slicing to segment the original image and input the network, set the window size according to the model detection efficiency, and fuse all slices and global detection results, effectively improving the detection capability of small and medium-scale targets. The proposed algorithm can be deployed to underwater mobile platforms such as AUV to realize real-time intelligent detection of submarine targets.

[0006] In order to achieve the above object, the technical solution of the present invention is:

[0007] A method for detecting seabed targets based on a dual-branch lightweight trunk specifically comprises the following steps:

[0008] The first step is data set preprocessing

[0009] Collect and annotate side-scan sonar image data. For categories with a small number of samples in the dataset, use data augmentation methods to generate more sample images. Finally, divide the dataset into training set, validation set, and test set in proportion.

[0010] Step 2: Network model construction and training

[0011] The network model includes input layer, Backbone, Neck, and output layer. The network model is trained and verified using the training set and the verification set to obtain a trained network model.

[0012] (2.1) Input layer

[0013] The input layer is the first layer of the network model to receive image data. It is responsible for scaling the original image to the fixed size required by the network and inputting it into the network. It then normalizes the pixel values ​​to provide raw data for subsequent feature extraction and processing.

[0014] Before the image data is input, the SAHI algorithm (Slicing Aided Hyper Inference) is used to slice the side scan sonar raw data waterfall chart into multiple slices, and target detection is performed independently on these slices.

[0015] (2.2)Backbone

[0016] Backbone extracts image features through a series of convolutional layers and activation functions, gradually reducing the spatial dimension of the image while increasing the number of channels. PP-LCNet and GhostNet are used to form a dual-branch lightweight backbone network, which extracts high-level features of the image from the input image through multiple layers of convolutional neural network layers in parallel.

[0017] Furthermore, the GhostNet is used as the first feature extraction backbone network. GhostNet is built based on the Ghost module and consists of 5 layers of GhostConv, 4 layers of C3Ghost and 1 SPPF module.

[0018] Furthermore, the PP-LCNet is used as the second feature extraction backbone network. PP-LCNet is a feature extraction backbone based on depthwise separable convolution (DepthSepConv), which consists of 1 layer of standard convolution, 6 consecutive layers of 3×3 depthwise separable convolution and 3 layers of 5×5 depthwise separable convolution.

[0019] (2.3)Neck

[0020] The Neck part is located between the Backbone and the output layer. It fuses the high-level features of the image extracted by the dual-branch lightweight backbone network, and fuses the high-resolution but less information feature map output by the same backbone network with the low-resolution but rich information feature map. The Neck structure fuses feature maps of different scales to enhance the network model's detection ability for targets of different sizes.

[0021] Furthermore, the 5th, 7th and 10th layers in each dual-branch lightweight backbone network are extracted as the input of the Neck part. In the Neck part, it is only necessary to fuse the feature maps extracted from each layer.

[0022] (2.4) Output layer

[0023] The output layer is responsible for generating the final target detection results and predicting the target location, category, and confidence in the image based on the fused feature maps.

[0024] Step 3: Model Evaluation

[0025] By putting the test set into the trained network model for detection, the target information in the side-scan sonar image is obtained, and the model performance is evaluated through indicators such as precision P (precision), recall R (recall), and mean average precision.

[0026] The beneficial effects of the present invention are as follows: in view of the problems that the existing target detection model has low detection efficiency and large number of parameters, which are difficult to be deployed on small unmanned platforms such as AUV, we designed a dual-branch lightweight backbone model DBnet. In the feature extraction stage, we used two lightweight backbone networks with superior performance, PP-LCNet and GhostNet, to extract input image features, and through the streamlined Neck structure, while reducing the number of parameters, it can still better integrate the features of each level extracted by the dual backbone, thereby enhancing the detection capability of targets of different scales. In the inference stage, since the original image of the side-scan sonar is large in size, directly inputting it into the network will compress the actual size of the target image after resampling. Therefore, we adopted the SAHI method, and according to the actual underwater operation of AUV, the entire waterfall image was cut into several slices with a certain overlap rate and fixed size, and then input into the network, which effectively solved the problem of difficult deployment of target detection models for unmanned platforms such as AUV, low detection efficiency and detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is the DBnet network structure diagram;

[0028] Figure 2 This is the SAHI schematic diagram;

[0029] Figure 3 This is a comparison chart of the mAP@0.5 curves of DBnet and YOLOv8;

[0030] Figure 4 Figure 2 is the model prediction effect diagram, (a) is the prediction effect diagram before improvement, and (b) is the prediction effect diagram after improvement. DETAILED DESCRIPTION

[0031] In order to make the method problems solved by the present invention, the method solutions adopted and the method effects achieved clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and experiments. It is understood that the specific experiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for the convenience of description, only the parts related to the present invention are shown in the accompanying drawings, rather than all the contents.

[0032] The model of the present invention is built based on the YOLOv8 algorithm framework, and the DBnet network structure is as follows Figure 1 As shown, it includes four parts: input layer, Backbone, Neck and output layer. In the present invention, we first use PP-LCNet and GhostNet as the dual backbone networks of the model, which enables the model to fuse feature information extracted from different backbone networks when there is only one modality data. Using different backbone networks can extract complementary feature information. By fusing these features, more comprehensive target information can be captured, thereby improving the detection accuracy of the model. Data from a single source is easily affected by noise, occlusion, etc., and this fusion of similar multimodal data can enhance the robustness of the model through the redundant information between features extracted from different backbone networks. In addition, by extracting features through a dual backbone network, the model can learn more generalized feature representations, thereby showing better performance under different background environments. Specifically, we use GhostNet as the first feature extraction backbone network, which is built based on the Ghost module and consists of 5 layers of GhostConv, 4 layers of C3Ghost and SPPF modules. Then PP-LCNet is used as the second feature extraction backbone network. PP-LCNet is a feature extraction backbone based on depthwise separable convolution (DepthSepConv), which consists of six consecutive layers of 3×3 depthwise separable convolution and three layers of 5×5 depthwise separable convolution.

[0033] In the Neck part, multiple upsampling steps of the original structure often bring adverse effects such as noise amplification. Therefore, we avoid using upsampling in the standard YOLOv8 model to merge small-scale high-level convolutional features and large-scale low-level convolutional features. Instead, we optimize the Neck part by fusion using sufficiently rich feature maps from dual-stem feature extraction. This approach reduces the computational load of the model, simplifies the complexity of the model, and helps to make the model lightweight. After feature extraction, we extracted the 5th, 7th, and 10th layers in each backbone network respectively as input to the neck part. In the Neck part, we only need to fuse the feature maps extracted from each layer. It is worth mentioning that we not only simplified the structure, but also used C3Ghost to replace the C2f module in the original YOLOv8 network, further reducing the parameters and computation of the model.

[0034] The function of the output layer is to complete the output of target detection results.

[0035] We collected two SSS datasets, SSUTD (Side-Scan Sonar Undersea Target Dataset) and Sonar Common Target Detection Dataset (SCTD), for the training and testing of detectors. The SSUTD dataset is mainly obtained by various hydrographic survey departments and companies using mainstream side-scan sonar equipment, and a small part of the data is collected from the Internet, totaling 1584 images. The SCTD dataset mainly contains side-scan sonar images of aircraft wreckage, shipwrecks, etc.

[0036] The first step is data set preprocessing

[0037] Collect and annotate side scan sonar image data. Due to the data-driven nature of target detection models, the performance of deep learning algorithms is easily affected by the number and distribution of training datasets. There are two problems with underwater target recognition model training: insufficient number of images and imbalanced number of different targets will limit the performance of deep learning algorithms. In order to enrich the number of samples in the dataset and improve the generalization of the model, data augmentation technology is introduced to generate more SSS images for the smaller number of drowning targets. Commonly used data augmentation algorithms include Gaussian noise, brightness change, image translation, rotation, flipping, scaling, shearing, etc.

[0038] The dataset SSUTD consists of 980 images of shipwrecks, 36 images of drowning people, and 568 images of aircraft wreckage. The dataset SCTD consists of 266 images of shipwrecks, 34 images of drowning people, and 57 images of aircraft wreckage. The dataset A is based on the dataset SSUTD and performs data enhancement on each drowning person image. It is divided into three subsets in a ratio of 8:1:1, namely, the training set, the validation set, and the test set. The detailed division is shown in Table 1.

[0039] Table 1 Dataset division

[0040]

[0041] Step 2: Model training

[0042] The model training process is carried out according to the technical solution

[0043] The specific principles of predictive reasoning are as follows: Figure 2 As shown, the original image I (blue frame) is first cut into x M×N slices (red frames) with a certain degree of overlap, denoted as P1, P2...P X , and then resize each slice while maintaining the aspect ratio. After that, the content contained in each slice is predicted. As the slice size becomes smaller, the model's detection performance for larger targets will decrease. Therefore, in order to detect larger targets more accurately, the full-inference (FI) result of the original image is used, and the NMS (non-maximum suppression) algorithm merges the predicted results and FI results of the slice back to the original size. During the NMS process, the boxes with an intersection over union (IoU) higher than the pre-set matching threshold Tm are matched, and the detection probability lower than Td is deleted for each match. This parallel processing method can significantly improve detection efficiency, especially when processing side-scan sonar images or in scenarios with high computing resource requirements. And by performing reasoning on smaller slices, the SAHI algorithm can reduce the consumption of memory and computing resources. This is especially important for devices with limited computing resources or real-time detection systems.

[0044] Step 3: Model Evaluation

[0045] In order to comprehensively and objectively evaluate the prediction effects of different models, the present invention evaluates and optimizes model performance through the following coefficients: average precision AP (average precision) and mean AP (mean AP).

[0046] Precision and recall: In the classification task of predicting whether an image contains a bag, the four elements of precision and recall can be explained as follows: TP (true positive): the positive sample is correctly marked as a positive sample in the prediction result; TN (true negative): the negative sample is correctly marked as a negative sample in the prediction result; FP (false positive): the positive sample is incorrectly marked as a negative sample in the prediction result; FN (false negative): the negative sample is incorrectly marked as a positive sample in the prediction result. The calculation relationship is:

[0047]

[0048] Average Precision: The geometric meaning of average precision AP is the area corresponding to the PR curve as shown in formula (3), where the interpolation summation method can be used to approximate the integral.

[0049]

[0050] Average precision: The mAP types used in this paper are: mAP@0.5 and mAP@0.5:0.95. mAP@0.5 (average precision coefficient of all categories) represents the precision and average of all images in each category when the IoU is set to 0.5; mAP@0.5:0.95 means that the IoU threshold of mAP is in the range of 0.5 to 0.95, and the calculation method is as follows:

[0051]

[0052] Where N represents the number of target categories;

[0053] In addition, the following parameters can reflect the complexity and lightweight effect of the model. FPS is an important indicator to measure the detection speed of an algorithm; it is the ratio of the number of frames to the time consumed.

[0054] Params and FLOPs represent the computational space complexity and computational complexity of the model, respectively.

[0055] FLOPs (floating point of operations) is the number of floating point operations. For the convolutional layer, the calculation formula for FLOPs is as follows:

[0056] FLOPs = 2 × H × W × (C in ×K 2 +1)×C out (5)

[0057] Among them, H and W represent the height and width of the convolutional layer output, respectively, and Cin Indicates the number of input channels, C out Represents the number of output channels, and K represents the convolution kernel size;

[0058] In order to verify the impact of different backbone network selections on model performance and the generalization of the model on different data sets, we conducted multiple sets of ablation experiments. First, in terms of the selection of the backbone network, the dual-backbone detector model proposed by the model of the present invention makes up for the shortcomings of a single means of detecting submarine obstacles and the lack of multi-source multi-modal data fusion. At the same time, in order to meet the lightweight implementation of the model for engineering deployment, we selected the current mainstream lightweight backbone network for combined experiments. The experimental results are shown in Table 2.

[0059] Table 2 Experimental results of different data sets. The bold fonts are the methods proposed by the present invention.

[0060] Table 2 Prediction evaluation results of different models.

[0061]

[0062] From the table data, it can be seen that DBnet has a significant improvement in mAP value compared with most single-backbone networks. In the dual-backbone network, although the number of parameters and GFLOPs of the method of the present invention are not the smallest, in terms of the comprehensive detection performance index mAP, the algorithm we proposed not only has better detection performance, but also has a great advantage in the lightweight degree of the model compared with other networks, which is sufficient to meet the needs of real-time target detection and engineering deployment.

[0063] The analysis of the experimental results is mainly divided into two aspects: one is to analyze the evaluation indicators of the original model and the improved model, and the other is to analyze the detection effect of the improved model.

[0064] Figure 3 This is a comparison of the mAP@0.5 curves of the improved model and the baseline model. It can be seen that the improved curve is higher than the baseline model. Not only is the accuracy improved, but the convergence speed is also faster and more stable. From the statistical data, it can be found that the improved model has better prediction performance.

[0065] Figure 4 The prediction results of the YOLOv8 model before and after the improvement are shown. When the improved model performs target detection, the model can accurately identify the shipwreck in the image and avoids the situation where the terrain undulations in the background are mistakenly detected as a shipwreck as in (a). It can be seen that the use of the present invention can effectively improve the accuracy of target detection, especially for large-size images such as waterfall charts of side-scan sonar raw data, avoiding the situation where small targets are missed and large targets are misdetected due to image compression when input into the network.

[0066] Finally, it should be noted that the above experiments are only used to illustrate the method scheme of the present invention, rather than to limit it. Although the present invention has been described in detail, ordinary method personnel in the field should understand that modifying the aforementioned method scheme, or replacing part or all of the method features therein by equivalents, does not deviate the essence of the corresponding method scheme from the scope of the method scheme of the present invention.

Claims

1. A method for detecting submarine targets based on a dual-branch lightweight trunk, characterized in that: The specific steps include: The first step is data set preprocessing Collect and annotate side-scan sonar image data; use data augmentation methods to generate more sample images for categories with a small number of samples in the dataset; finally, divide the dataset into training set, validation set, and test set in proportion; Step 2: Network model construction and training The network model includes input layer, Backbone, Neck, and output layer. The network model is trained and verified using the training set and the verification set to obtain a trained network model. (2.1) Input layer The input layer is the first layer of the network model to receive image data. It is responsible for scaling the original image to the fixed size required by the network and inputting it into the network. After normalization of the pixel values, it provides raw data for subsequent feature extraction and processing. Before the image data is input, the SAHI algorithm is used to divide the side scan sonar raw data waterfall chart into multiple slices, and target detection is performed independently on these slices; (2.2)Backbone Backbone extracts image features through a series of convolutional layers and activation functions, gradually reducing the spatial dimension of the image while increasing the number of channels. PP-LCNet and GhostNet are used to form a dual-branch lightweight backbone network, which extracts high-level features of the image from the input image through multiple layers of convolutional neural network layers in parallel. (2.3)Neck The Neck part is located between the Backbone and the output layer. It fuses the high-level features extracted from the dual-branch lightweight backbone network to the image, and fuses the high-resolution but less-information feature map output by the same backbone network with the low-resolution but rich-information feature map. The Neck structure fuses feature maps of different scales to enhance the network model's ability to detect objects of different sizes. (2.4) Output layer The output layer is responsible for generating the final target detection results and predicting the target location, category, and confidence in the image based on the fused feature map; Step 3: Model Evaluation By putting the test set into the trained network model for detection, the target information in the side-scan sonar image is obtained, and the model performance is evaluated by the precision rate P, recall rate R, and average precision mean indicators.

2. A method for detecting submarine targets based on a dual-branch lightweight trunk according to claim 1, characterized in that: The GhostNet is the first feature extraction backbone network. GhostNet is built based on the Ghost module and consists of 5 layers of GhostConv, 4 layers of C3Ghost and 1 SPPF module.

3. The method for detecting submarine targets based on a dual-branch lightweight trunk according to claim 1, characterized in that: The PP-LCNet is used as the second feature extraction backbone network. PP-LCNet is a feature extraction backbone based on depthwise separable convolution, which consists of 1 layer of standard convolution, 6 consecutive layers of 3×3 depthwise separable convolution and 3 layers of 5×5 depthwise separable convolution.

4. The method for detecting submarine targets based on a dual-branch lightweight trunk according to claim 1, characterized in that: The 5th, 7th and 10th layers in each dual-branch lightweight backbone network are extracted as the input of the Neck part; in the Neck part, it is only necessary to fuse the feature maps extracted from each layer.

Citation Information

Patent Citations

  • Improved number identification method

    CN114419415A

  • Side-scan sonar target detection method combining accurate image segmentation and target shadow information

    CN115240058A

  • Forward-looking sonar image small target identification method based on SSE-YOLO deep learning model

    CN116863321A

  • Underwater target detection method and system based on spatial deep convolutional network

    CN117392524A

  • Target detection method, and moving-target tracking method using same

    WO2023138300A1

Cited By

  • Traffic cone real-time target detection method, system and equipment based on YOLOv8 identification model

    CN121861595A