A Side-Scan Sonar Target Detection Method and Device
By building a diverse training sample set and using the LANConvNeXtv2 module, the accuracy and robustness of the traditional side-swept sonar object detection method in low contrast and fuzzy environments is solved, and higher detection accuracy and model robustness are achieved.
Patent Information
- Application Number
- CN202510368266.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-26
AI Technical Summary
When traditional side-sweep sonar object detection methods deal with targets with low contrast, blurred boundary, obstructed and background noise, there are problems with poor detection accuracy and robustness.
By constructing a training sample set, including multiple subsea background sonar images and target object sonar images, and extracting features based on the backbone network of the LANConvNeXtv2 module, combining the attention mechanism module and the neck network, the model's perception ability of low contrast and fuzzy targets is improved.
The target detection accuracy is significantly improved, especially in the case of low contrast, fuzzy and small targets, and the robustness and generalization ability of the model are improved.
Smart Images

Figure CN119888451B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of underwater target detection, and particularly to a side-scan sonar target detection method and device. Background Art
[0002] By transmitting and receiving acoustic wave signals, a side-scan sonar generates high-resolution seabed images and is widely used in underwater target detection. However, the targets in the images generated by the side-scan sonar usually have the characteristics of low contrast, blurred boundaries, being occluded, and having a lot of background noise. As a result, when traditional target detection methods are used to detect targets in side-scan sonar images, there are problems such as poor noise robustness, difficulty in dealing with multi-scale and multi-angle targets, and false detection and missed detection of targets in the case of target occlusion and low contrast between the target and the background. Summary of the Invention
[0003] In view of this, the present application provides a side-scan sonar target detection method and device to solve the problems of poor detection accuracy and robustness of traditional target detection methods when detecting side-scan sonar images.
[0004] Specifically, the present application is implemented through the following technical solutions:
[0005] The first aspect of the present application provides a side-scan sonar target detection method, and the method includes:
[0006] Construct a training sample set, where the training sample set includes multiple seabed background sonar images and target object sonar images;
[0007] Train a target detection model based on the training sample set;
[0008] Wherein, the target detection model includes a backbone network, a neck network, and a detection network connected in sequence. The backbone network extracts features from the sonar image to be detected based on the LANConvNeXtv2 module. The LANConvNeXtv2 module includes a convolution module and an attention mechanism module. The output of the convolution module is connected to the input of the attention mechanism module, and the output of the attention mechanism module is connected to the input of the neck network;
[0009] The convolution module is used to extract the feature map of the sonar image to be detected. The attention mechanism module calculates the attention value of each patch in the feature map based on the relationship between each patch and the target object, and generates a first feature map carrying weight information according to the attention value, and outputs the first feature map to the neck network. The neck network is used to enhance the first feature map to obtain a second feature map. The detection network is used to detect the second feature map to obtain a detection result;
[0010] Use the trained object detection model to perform object detection on the sonar image to be detected, and obtain the object detection result.
[0011] The second aspect of this application provides a side-scan sonar object detection device, which includes a construction module, a training module, and a detection module; among them,
[0012] The construction module is used to construct a training sample set, which includes multiple sonar images of the seabed background and sonar images of target objects;
[0013] The training module is used to train an object detection model based on the training sample set;
[0014] Among them, the object detection model includes a backbone network, a neck network, and a detection network connected in sequence. The backbone network extracts features from the sonar image to be detected based on the LANConvNeXtv2 module. The LANConvNeXtv2 module includes a convolution module and an attention mechanism module. The output of the convolution module is connected to the input of the attention mechanism module, and the output of the attention mechanism module is connected to the input of the neck network;
[0015] The convolution module is used to extract the feature map of the sonar image to be detected. The attention mechanism module calculates the attention value of each patch in the feature map based on the relationship between each patch and the target object, and generates a first feature map carrying weight information according to the attention value, and outputs the first feature map to the neck network. The neck network is used to enhance the first feature map to obtain a second feature map. The detection network is used to detect the second feature map to obtain the detection result;
[0016] The detection module is used to perform object detection on the sonar image to be detected by using the trained object detection model, and obtain the object detection result.
[0017] The side-scan sonar object detection method and device provided by this application effectively solve the problems of uneven target distribution and insufficient samples in side-scan sonar images by constructing a training sample set. Further, the backbone network extracts features from the sonar image to be detected based on the LANConvNeXtv2 module. The convolution module in the LANConvNeXtv2 module is responsible for extracting the feature map of the sonar image to be detected. Its attention mechanism module calculates the attention value based on the relationship between each patch in the feature map and the target object, and generates a first feature map carrying weight information. This design based on the local attention mechanism can pay more attention to the details of the key regions in the image, especially in the case of low contrast or blurriness, significantly improving the model's perception ability for low-contrast, blurred, and small targets, and is particularly suitable for the case where the target boundary in underwater side-scan sonar images is unclear or the contrast is low, thereby improving the object detection accuracy. Description of the Drawings
[0018] Figure 1 It is a flowchart of the first embodiment of the side-scan sonar target detection method provided by this application;
[0019] Figure 2 It is a schematic diagram of a side-scan sonar image exemplarily shown by this application;
[0020] Figure 3 It is a schematic structural diagram of the target detection model exemplarily shown by this application;
[0021] Figure 4 It is a schematic structural diagram of the first embodiment of the side-scan sonar target detection device provided by this application. Detailed Embodiments
[0022] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.
[0023] The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms "a", "the" and "said" used in this application are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0024] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0025] Specific embodiments are given below to introduce the technical solutions of this application in detail.
[0026] Figure 1 It is a flowchart of the first embodiment of the side-scan sonar target detection method provided by this application. Please refer to Figure 1 , the method provided in this embodiment may include:
[0027] S101. Construct a training sample set, where the training sample set includes multiple sonar images of the seabed background and sonar images of target objects.
[0028] Specifically, the steps for constructing the training sample set include:
[0029] (1) Obtain the original seabed side-scan sonar image.
[0030] Specifically, the original seabed side-scan sonar image can be obtained through a side-scan sonar device, or from existing public datasets. The original seabed side-scan sonar image contains different seabed substrates (such as sandy, rocky, etc.) and target objects (such as shipwrecks, airplanes, etc.).
[0031] (2) Construct a background processing dataset and a target object processing dataset.
[0032] Specifically, to prepare for subsequent separate processing of the background and target objects, empty background processing and target object processing datasets are constructed respectively. The background processing dataset is used to store sample data for processing the background area, and the target object processing dataset is used to store sample data for processing the target object.
[0033] (3) Identify the background area pixels in the original seabed side-scan sonar image.
[0034] Specifically, using image recognition technology, identify the target object area pixels and background area pixels in the original seabed side-scan sonar image. In specific implementation, the background and target objects in the original image can be distinguished through an image segmentation algorithm, and then the positions and attributes of the background area pixels can be determined. For the specific implementation process of using image recognition technology to identify background area pixels, please refer to the descriptions in related technologies and will not be elaborated here.
[0035] (4) Generate random noise based on the background area pixels, and use the random noise to process the target objects in the original seabed side-scan sonar image to generate background processing samples and add them to the background processing dataset.
[0036] Specifically, random noise with similar characteristics can be generated according to the characteristics of the background area pixels, such as color, grayscale, etc. distribution. These noises are randomly added to the target object areas in the original seabed side-scan sonar image to simulate the interference of noise on the target in a complex seabed environment, thereby obtaining background processing samples, and then adding the background processing samples to the background processing dataset.
[0037] (5) Identify the boundaries where the target objects are located in the original seabed side-scan sonar image, process the pixels within the boundaries, generate target object samples, and add them to the target object processing dataset.
[0038] Specifically, the boundaries of target objects (such as aircraft and sunken ships) in the original seabed side-scan sonar images can be identified through image edge detection algorithms or based on specific features of the target objects. For the pixels within the boundaries, operations such as cropping, scaling, and rotation are performed to generate target object samples of various forms. The generated target object samples are then added to the target object processing data set to enrich the learning samples of the target objects during model training and improve the model's ability to recognize targets of different postures and scales.
[0039] (6) Mix the background processed data set and the target object processed data set to obtain a training sample set.
[0040] Specifically, the background processing data set and the target object processing data set that have been processed as above are mixed and their order is disrupted. The training sample set obtained in this way includes samples under various background interferences and target object samples with different processing methods, which can better simulate the complex actual seabed environment and be used to train the seabed side-scan sonar image target detection model, thereby improving the detection accuracy and robustness of the model in complex environments.
[0041] Furthermore, the step of constructing a training sample set may also include:
[0042] (1) Constructing an initial data set, wherein the initial data set includes a plurality of seabed background sonar images and target object sonar images.
[0043] Specifically, the initial data set is the SeabedObjects-KLSG side-scan sonar data set built with the support of several sonar equipment suppliers such as Lcocean, Hydro-tech Marine, Klein Marine, Tritecand EdgeTech, etc. This data set constitutes the initial data set, which contains 578 seabed background sonar images, 385 shipwrecks and 62 aircraft sonar image data. The shipwreck and aircraft sonar images are the target object sonar images. It can be understood that the shipwreck and the aircraft are the target objects that the target detection model needs to learn. Furthermore, the seabed background sonar image is a seabed sonar image that does not contain the target object, and the target object sonar image is a seabed sonar image that contains the target object. Figure 2 For a schematic diagram of a side scan sonar image exemplified in this application, please refer to Figure 2 , Figure 2 Parts (a) and (b) are schematic diagrams of sonar images of target objects. Figure 2 Part (c) is a schematic diagram of the seabed background sonar image.
[0044] Extract the target object from the sonar image of the target object, and combine the extracted target object with the sonar image of the seabed background to form multiple first sonar images of the target object.
[0045] Specifically, the specific implementation steps for forming multiple first sonar images of the target object include:
[0046] 2.1. Determine the target object pixels according to the binary mask of each pixel in the sonar image of the target object, retain the target object pixels, and determine the target object according to the target object pixels.
[0047] Specifically, the target object in the sonar image of the target object can be extracted by the following formula:
[0048] ;
[0049] where is the pixel coordinate in the sonar image of the target object;
[0050] is the binary mask of the target object in the sonar image of the target object. If the pixel coordinate (x, y) belongs to the target object, then = 1. If the pixel coordinate (x, y) does not belong to the target object, then = 0;
[0051] is the pixel coordinate of the target object.
[0052] 2.2. Paste the target object into the sonar image of the seabed background to obtain the multiple first sonar images of the target object.
[0053] Specifically, after the target object is extracted, the target object and the sonar image of the seabed background can be synthesized by the following formula:
[0054] ;
[0055] where is the pixel coordinate of the first sonar image of the target object;
[0056] is the pixel coordinate of the sonar image of the seabed background;
[0057] is the pixel coordinate of the target object;
[0058] is the binary mask of the target object in the sonar image of the target object. If the pixel coordinate (x, y) belongs to the target object, then = 1. If the pixel coordinate (x, y) does not belong to the target object, then = 0。
[0059] For the sonar image of the target object, by retaining the background area of the non-target object in the image and covering the extracted target object in the corresponding area, for the sonar image of the seabed background, the extracted target object is directly randomly covered in the image, so as to obtain multiple first target object sonar images.
[0060] Furthermore, after extracting the target object, the extracted target object can also be randomly rotated and scaled to ensure that the target object is not static when synthesized into the new background image, thereby expanding the diversity of the target object in the dataset. By processing the sonar images of the target object in the initial dataset, the targets in 385 sunken ship and 62 aircraft sonar images are extracted and pasted into 578 seabed background sonar images, obtaining 5700 first target object sonar images.
[0061] (3) Perform stitching processing on the target object sonar image and the first target object sonar image to form multiple second target object sonar images.
[0062] Specifically, in the side-scan sonar image, the target object may have different sizes and shapes, and it is difficult for the single-scale image data to enable the model to learn comprehensive feature representations. By constructing multi-scale images, the model can learn the feature differences of the target at different scales and improve the detection ability for targets of different sizes.
[0063] Specifically, the specific implementation steps of performing stitching processing on the target object sonar image and the first target object sonar image to form multiple second target object sonar images include:
[0064] 3.1. Determine the target object sonar image and the first target object sonar image as the first-scale images.
[0065] Specifically, the target object sonar image and the first target object sonar image are uniformly determined as the first-scale images. This step provides the basic data source for the subsequent construction of multi-scale images. And in the first-scale image, only one target object is included in one image.
[0066] 3.2. Randomly select M images from the first-scale images, and splice the M images to form a second-scale target image.
[0067] Specifically, the value of M is set according to actual needs and is not limited in this embodiment. For example, in one embodiment, M = 4. Four images are randomly selected from the first-scale images, and these four images are randomly stitched together to obtain the second-scale target image. It should be noted that when randomly selecting images from the first-scale images, each image in the first-scale images can only be selected once.
[0068] 3.3 Randomly select N images from the first-scale images and stitch the N images together to form a third-scale target image.
[0069] Specifically, the value of N is set according to actual needs and is not limited in this embodiment. For example, in one embodiment, N = 9. Referring to the above description, similarly, when randomly selecting 9 images from the first-scale images and stitching the 9 images together, a third-scale target image is formed. It should be noted that the value of N is greater than M.
[0070] 3.4 Use the first-scale images, the second-scale target images, and the third-scale target images as the sonar images of the second target object.
[0071] Specifically, the first-scale images, the second-scale target images, and the third-scale target images are jointly used as the sonar images of the second target object, thus forming a sonar image of the second target object containing multi-scale information. In the subsequent model training process, the target detection model can learn the features of the target at different feature levels, improving the detection accuracy and robustness of the target detection model for targets of different scales. For example, in a complex underwater environment, there may be small target objects (such as small sunken ship components) and large target objects (such as a complete sunken ship or large underwater facilities). The target detection model trained with such multi-scale images can more effectively detect these targets of different sizes, reducing the situations of missed detection and false detection.
[0072] (4) Perform enhancement processing on the sonar images of the target object, the sonar images of the first target object, and the sonar images of the second target object to form multiple sonar images of the third target object.
[0073] Specifically, the enhancement processing of the sonar image of the target object, the sonar image of the first target object, and the sonar image of the second target object includes Cutout processing and noise addition processing. Cutout processing refers to randomly selecting and generating a square occlusion area in the image and setting its pixel values to zero, thereby simulating the partial occlusion of the target object by sediments, seabed structures, or ocean clutter in the seabed environment. Noise addition processing refers to adding different types of noises such as Gaussian noise and salt-and-pepper noise to the image to simulate the actual seabed noise interference scenario, enabling the target detection model to learn to recognize and adapt to the target features under various noise conditions during the training process. The enhancement processing of the above images can enhance the ability of the target detection model to infer the complete form of the target under information loss conditions, and at the same time improve the generalization performance of the target detection model and the robustness to complex occlusion scenarios. Further, it improves the anti-noise ability of the target detection model, reduces the negative impact of noise on detection, and reduces the risk of missed detection and false detection of the target object by the target detection model.
[0074] Further, in addition to the method of globally enhancing the image given above, the implementation steps for forming multiple sonar images of the third target object further include:
[0075] 4.1. Determine the target object in the sonar image of the target object, the sonar image of the first target object, and the sonar image of the second target object.
[0076] Specifically, image processing algorithms can be used to locate and determine it based on the gray-scale difference between the target object and the background, or based on the unique geometric shape of the target object. And for different sonar images, different feature extraction and target localization methods can be used. For example, for some simple target objects, a threshold-based segmentation method can be used to regard the area with gray-scale higher or lower than the threshold as the target object; for complex target objects, more complex edge detection and contour analysis algorithms, such as the Canny edge detection algorithm, can be used to find the contour of the target object, and then determine its position and range. The specific implementation process of determining the target object in the image can refer to the description in the related technology and will not be elaborated here.
[0077] 4.2. Preprocess the target object to obtain the sonar image of the third target object.
[0078] Specifically, after determining the position and size of the target object in the image, the target object is enhanced to obtain the sonar image of the third target object. For the implementation steps of enhancing the target object, please refer to the above description and will not be elaborated here. This method only enhances the target object in the image to obtain the sonar image of the third target object, which helps the target detection model to pay more attention to the features of the target object and enhance the detection accuracy in the subsequent model training process.
[0079] (5) Update the sonar image of the target object, the sonar image of the first target object, the sonar image of the second target object, and the sonar image of the third target object to the initial data set to obtain a training sample set.
[0080] Specifically, after a series of processes in the above steps, the initial data set has been greatly expanded and enriched, forming a more comprehensive and diverse training sample set. This training sample set contains sonar images of target objects in different scenarios, different combinations, and different conditions, providing rich sample data for the subsequent training of the target detection model, which helps to train a target detection model with stronger generalization ability and better robustness.
[0081] S102. Train a target detection model based on the training sample set;
[0082] Among them, the target detection model includes a backbone network, a neck network, and a detection network connected in sequence. The backbone network extracts features from the sonar image to be detected based on the LANConvNeXtv2 module. The LANConvNeXtv2 module includes a convolution module and an attention mechanism module. The output of the convolution module is connected to the input of the attention mechanism module, and the output of the attention mechanism module is connected to the input of the neck network;
[0083] The convolution module is used to extract the feature map of the sonar image to be detected. The attention mechanism module calculates the attention value of each patch in the feature map based on the relationship between each patch and the target object, and generates a first feature map carrying weight information according to the attention value, and outputs the first feature map to the neck network. The neck network is used to enhance the first feature map to obtain a second feature map. The detection network is used to detect the second feature map to obtain a detection result.
[0084] Specifically, the YOLOv8 model is selected as the basic model for object detection in side-scan sonar images in complex environments. YOLOv8 is one version of the YOLO (You Only Look Once) series, which has the advantages of high speed and high precision of the YOLO series and has been significantly optimized in terms of architecture, detection ability, and efficiency. However, in the underwater environment, the objects in sonar images usually have extremely low contrast, which poses significant challenges to object detection. Different from traditional optical images, sonar images are usually acquired under complex underwater conditions due to the particularity of their generation method, and the contrast between the target object and the background in the image is often very low. This low-contrast phenomenon greatly increases the difficulty of distinguishing the target from the background, making it difficult for traditional computer vision techniques to effectively identify the target object. In addition to the low-contrast problem, there are also a large number of complex noise sources and interference factors in the underwater environment. These noises include seabed reflections, ocean clutter, and sensor noise of the sonar device itself. Although the YOLOv8 model, as a powerful object detection model, shows excellent results in general object detection tasks, it still has problems with insufficient feature extraction ability when facing sonar images with low contrast and complex noise.
[0085] Figure 3 The structural schematic diagram of the object detection model exemplarily shown in this application is shown in Figure 3 In this embodiment, the proposed object detection model, based on the initial structure of the YOLOv8 model, includes a backbone network, a neck network, and a detection network. The backbone network extracts features from the sonar image to be detected based on the LANConvNeXtv2 module. The LANConvNeXtv2 module includes a convolution module and an attention mechanism module. The output of the convolution module is connected to the input of the attention mechanism module, and the output of the attention mechanism module is connected to the input of the neck network.
[0086] Among them, the backbone network includes multiple convolution modules, CSPLayer_2Conv (C2f), LANConvNeXtV2, etc. The ConvModule is used to perform preliminary feature extraction on the input image and extract the basic features of the image through convolution operations; CSPLayer_2Conv (C2f) is a cross-stage local network structure that can effectively reduce the computational amount and enhance the feature extraction ability; LANConvNeXtV2 is a convolution module combined with specific improvements for further extracting and processing image features. Starting from the input image, the backbone network gradually performs feature extraction and downsampling. Through the combination of different modules, it continuously extracts higher-level and more abstract features, providing a basis for subsequent feature fusion and object detection. As the number of network layers deepens, the image size gradually decreases, and the feature dimension gradually increases to capture richer semantic information in the image.
[0087] The neck network includes multiple ConvModule, CSPLayer_2Conv (C2f), SPPF (Spatial Pyramid Pooling Fusion) and other modules. The SPPF module is used to perform multi-scale pooling operations on the feature map, fuse features of different scales, and enhance the adaptability of the model to targets of different sizes; ConvModule and CSPLayer_2Conv (C2f) continue to process and fuse the features; the neck network mainly plays the role of feature fusion and adjustment. It receives feature maps of different levels output by the backbone network, and through the processing of various modules, fuses features of different scales, enabling the network to better process multi-scale targets, while reducing information loss and improving the feature expression ability.
[0088] The detection network includes multiple SequentiaL feature processing modules, DEL optimization modules, and ImplicitHead implicit learning modules. The SequentiaL feature processing module is used to output detection results of different scales, the DEL optimization module is used to further optimize the detection results, and the ImplicitHead implicit learning module is used to further process the feature map to generate the final detection output; the detection network performs object detection based on the features fused by the neck network. Through different module connections, feature maps of different scales are processed to detect targets of different sizes and positions in the image. Feature output layers of different scales can detect targets of different sizes, and finally, the detection results are optimized through modules such as DEL to generate the final detection output.
[0089] Further, the convolution module includes a two-dimensional convolution layer, a multi-scale convolution layer, a fusion layer, and a dimensionality reduction layer; the output of the two-dimensional convolution layer is connected to the input of the multi-scale convolution layer, the output of the multi-scale convolution layer is connected to the input of the fusion layer, and the output of the fusion layer is connected to the input of the dimensionality reduction layer;
[0090] The two-dimensional convolution layer is used to extract local features of the sonar image to be detected, the multi-scale convolution layer extracts features of different scales of the local features through convolution kernels of different sizes to obtain the features of different scales, the fusion layer is used to fuse the features of different scales to obtain fused features, and the dimensionality reduction layer is used to perform channel dimensionality reduction on the fused features to obtain dimensionality-reduced features, and the dimensionality-reduced features are used as the output of the convolution module.
[0091] Specifically, the two-dimensional convolutional layer slides a convolutional kernel over an image, performs a convolution operation on a local region of the image, and sums up the weighted local pixel information to generate a new feature map. During this process, the weights of the convolutional kernel are continuously adjusted during training to learn the pattern that best represents the local features of the image. The local features of the image to be detected can be extracted using the following formula:
[0092] ;
[0093] where, is the convolutional kernel in the two-dimensional convolutional layer;
[0094] is the input image to be detected;
[0095] is the bias term.
[0096] The multi-scale convolutional layer extracts features of different scales of local features through convolutional kernels of different sizes. Convolutional kernels of different sizes can capture information in different ranges. Larger convolutional kernels can cover larger regions and obtain more macroscopic features, while smaller convolutional kernels focus more on local details. The features of different scales can be extracted using the following formula:
[0097] ;
[0098] where, is the set of scales;
[0099] is the weight of the convolutional kernel at scale s;
[0100] is the local feature;
[0101] is the bias term corresponding to scale s.
[0102] Furthermore, the fusion layer fuses features of different scales to obtain fused features. Additive fusion can be performed using the following formula:
[0103] ;
[0104] where, is the feature map at scale s.
[0105] Concatenation fusion can also be performed using the following formula:
[0106] ;
[0107] where, is the feature map at scale s.
[0108] Furthermore, in deep learning, feature maps usually have a high dimension, which may lead to problems such as excessive computational complexity and overfitting. The dimensionality reduction layer can reduce the number of channels of the fused features by performing a convolution operation on the fused features using a 1x1 convolutional kernel, while retaining important information. This can reduce the number of parameters in the model, improve computational efficiency, and also make the fused features more compact, which is helpful for subsequent processing and the final object detection task.
[0109] Furthermore, after the convolutional module performs a convolution operation on the image to be detected and outputs the dimensionality-reduced features, the dimensionality-reduced features are input into the attention mechanism module. The attention mechanism module calculates the attention values of each patch in the dimensionality-reduced features through the sigmoid function, and the attention values will be used as weight information for the dimensionality-reduced features to generate the first feature map. The first feature map carrying weight information can be generated through the following formula:
[0110] ;
[0111] where, is the weight information of the attention value corresponding to patch a;
[0112] is the input dimensionality-reduced feature;
[0113] is the sigmoid function.
[0114] Specifically, the attention mechanism module enables the object detection model to focus on processing key object features and reduce the attention to noise regions, thereby improving the detection accuracy. Furthermore, the neck network is used to enhance the first feature map to improve the effect of subsequent processing and provide higher-quality feature information for the final tasks (such as object detection, classification, etc.). The specific implementation steps include:
[0115] (1) Calculate the sampling points of the target object according to the scale information of the first feature map, and form a sampling point set according to the sampling points.
[0116] Specifically, the specific implementation steps for calculating the sampling points of the target object include:
[0117] 1.1. Determine the range factor and offset. The range factor includes a static range factor and a dynamic range factor.
[0118] Specifically, the static range factor is usually a preset fixed value, which may be determined based on experience or general understanding of the data. Its purpose is to provide a basic range reference for the determination of sampling points. Specifically, the static range factor can be determined according to the size information of the target object, and the size of the target object can be used as the static range factor. The dynamic range factor changes according to the real-time information of the first feature map or the current state of the target object. It can dynamically adjust the sampling range according to different characteristics of the target object in the feature map (such as the size, position, intensity, etc. of the target object) to better adapt to various different situations. It can be understood that in this step, the static range factor and the dynamic range factor can be set according to the scale information of the first feature map according to empirical values.
[0119] 1.2. Calculate the static range factor offset according to the static range factor, and calculate the dynamic range factor offset according to the dynamic range factor.
[0120] Specifically, the static range factor offset can be calculated by the following formula:
[0121] ;
[0122] where is the first feature map;
[0123] is a linear transformation of the first feature map.
[0124] In the static range factor offset calculation formula, the first feature map undergoes a linear transformation and outputs a size of H×W×2gs 2 ; after being scaled by a factor of 0.25, it is rearranged through the Pixel Shuffle operation to a size of sH×sW×2g.
[0125] The dynamic range factor offset can be calculated by the following formula:
[0126] .
[0127] In the dynamic range factor offset calculation formula, the first feature map obtains two sets of tensors through two linear transformations. These two sets of tensors perform an element-wise addition operation, are multiplied by a factor of 0.25, and then are converted to the target size of sH×sW×2g through Pixel_Shuffle.
[0128] 1.3. Determine the sampling points according to the sum of the static range factor and the static range factor offset and the sum of the dynamic range factor and the dynamic range factor offset.
[0129] Specifically, add the static range factor and the static range factor offset to obtain the range of an adjusted static range; add the dynamic range factor and the dynamic range factor offset to obtain the range of an adjusted dynamic range. According to these two adjusted ranges, select corresponding sampling points in the first feature map.
[0130] Furthermore, for the first feature map with only background, the range factor provides the initial sampling distribution, and the offset introduces dynamic adjustment. The static range factor fixes the offset generation strategy, while the dynamic range factor combines two linear transformations to further enhance the flexibility and adaptability of sampling.
[0131] Furthermore, dynamic sampling can focus on the target area for sampling. In the first feature map with a target object, the neck network can generate more sampling points in and around the target area according to the position and scale information of the target, enhancing the ability to extract features of the target object. In the case where the target object is partially occluded, the sampling points will be concentrated in the unoccluded target area and the occlusion edge to obtain the effective features of the target as much as possible and improve the detection accuracy.
[0132] Specifically, determine the feature differences of each region in the first feature map, and determine the classification of the first feature map based on the feature differences. The categories of the first feature map include the seabed background and the target object. In the seabed background category, the feature differences of each region in the first feature map are less than the threshold. In the target object category, the feature differences of each region in the first feature map are greater than the threshold. Determine the change rate of the dynamic range factor according to the classification result. The change rates of the seabed background category and the target object category are different, and the change rate of the seabed background category is less than that of the target object category. Among them, the method for determining each region is: determine the size information of the target object, and divide the first feature map based on the size information.
[0133] (2) Use the Grid Sample operation to process the sampling point set and the first feature map to obtain the sampled second feature map.
[0134] Specifically, Grid Sample can extract information from the first feature map according to the sampling point set. This operation can perform flexible resampling on the feature map, obtain the pixel values at the corresponding positions from the first feature map according to the positions of the sampling points, and generate a new feature map according to a certain interpolation rule. For the implementation steps of using the Grid Sample operation to process the sampling point set and the first feature map, please refer to the description in the related technology and will not be elaborated here.
[0135] Further, the detection network performs implicit processing on the input second feature map through the ImplicitHead module, thereby more efficiently capturing key information in the image and reducing noise interference. For the implementation steps of using ImplicitHead to perform implicit processing on the second feature map, please refer to the description in the related technology and will not be elaborated here.
[0136] S103. Use the trained object detection model to perform object detection on the sonar image to be detected, and obtain the object detection result.
[0137] Specifically, the sonar image to be detected has the characteristics of low contrast between the target object and the background, blurred boundaries of the target object, noise interference, and multi-scale targets. Since the side-scan sonar image is generated by emitting and receiving sound waves, it is greatly affected by the underwater environment. The absorption and scattering of sound waves by water bodies and the reflection of the complex seabed sediments make the contrast between the target object and the background in the image extremely low. Targets such as sunken ships and airplanes are difficult to distinguish from the background by gray scale or color difference in the sonar image. In addition, the underwater environment is complex, and the target is often attached with sediments and marine organisms, or due to the characteristics of sound wave propagation, the boundaries of the target object in the sonar image are not clear. When part of the target is blocked, the boundary is more difficult to determine, which makes the object detection method based on boundary features ineffective, and it is difficult for the model to accurately outline the target contour and determine its position.
[0138] Specifically, the object detection model is trained based on the constructed training sample set. During the training process, the sonar images in the training sample set are input into the object detection model, and the model will process the images according to its internal structure and parameters to predict the results. Then, the predicted results are compared with the real annotation results, and the error is calculated through a loss function (such as cross-entropy loss, mean square error, etc.).
[0139] Use an optimization algorithm (such as stochastic gradient descent, Adam, etc.) to update the parameters of the model according to the error, so that the model is gradually optimized during continuous iteration, reducing the gap between the predicted result and the real result, and improving the accuracy and performance of object detection.
[0140] Further, input the sonar image to be detected into the trained object detection model. The object detection model will sequentially extract features through the backbone network, enhance features through the neck network, and detect features through the detection network, and finally output the detection results, such as the position information (possibly represented in the form of a bounding box) and category information of the target object.
[0141] The side-scan sonar target detection method provided in this embodiment effectively solves the problems of uneven target distribution and insufficient samples in side-scan sonar images by constructing a training sample set. By copying and pasting the target objects in the sonar images of the target objects to different positions, backgrounds or scales, it also simulates situations such as target occlusion and uneven target distribution that may be encountered in the underwater environment, solves the problems of small amount of side-scan sonar target image data and insufficient underwater target detection in complex underwater environments, enables the target detection model to be trained and tested in various complex environments, and enhances the robustness of the target detection model during training. Further, the backbone network extracts features from the sonar image to be detected based on the LANConvNeXtv2 module. The convolutional module in the LANConvNeXtv2 module is responsible for extracting the feature map of the sonar image to be detected, and its attention mechanism module calculates the attention value based on the relationship between each patch in the feature map and the target object, and generates a first feature map carrying weight information. This design based on the local attention mechanism can pay more attention to the details of the key regions in the image, especially in the case of low contrast or blurriness, significantly improving the model's perception ability for low-contrast, blurred and small targets, and is particularly suitable for the situation where the target boundaries in underwater side-scan sonar images are unclear or the contrast is low, thus improving the target detection accuracy.
[0142] Corresponding to the foregoing embodiment of a side-scan sonar target detection method, the present application also provides an embodiment of a side-scan sonar target detection device.
[0143] Figure 4 It is a schematic structural diagram of Embodiment 1 of the side-scan sonar target detection device provided by the present application. Please refer to Figure 4 , the device provided in this embodiment includes a construction module 410, a training module 420, and a detection module 430; wherein,
[0144] The construction module 410 is used to construct a training sample set, and the training sample set includes multiple sonar images of the seabed background and sonar images of target objects;
[0145] The training module 420 is used to train a target detection model based on the training sample set;
[0146] Among them, the target detection model includes a backbone network, a neck network, and a detection network connected in sequence. The backbone network extracts features from the sonar image to be detected based on the LANConvNeXtv2 module. The LANConvNeXtv2 module includes a convolutional module and an attention mechanism module. The output of the convolutional module is connected to the input of the attention mechanism module, and the output of the attention mechanism module is connected to the input of the neck network;
[0147] The convolution module is used to extract the feature map of the sonar image to be detected. The attention mechanism module calculates the attention value of each patch in the feature map based on the relationship between each patch and the target object, and generates a first feature map carrying weight information according to the attention value, and outputs the first feature map to the neck network. The neck network is used to enhance the first feature map to obtain a second feature map. The detection network is used to detect the second feature map to obtain a detection result;
[0148] The detection module 430 is used to perform target detection on the sonar image to be detected by using a trained target detection model to obtain a target detection result.
[0149] The device of this embodiment can be used to execute Figure 1 the steps of the method embodiment shown. The specific implementation principle and process are similar and will not be elaborated here.
[0150] For the implementation process of the functions and roles of each unit in the above device, please refer to the implementation process of the corresponding steps in the above method for details, which will not be elaborated here.
[0151] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative work.
[0152] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the scope of protection of the present application.
Claims
1. A side scan sonar target detection method, characterized in that: The method comprises: Constructing a training sample set, wherein the training sample set includes a plurality of seabed background sonar images and target object sonar images; Training a target detection model based on the training sample set; The target detection model is constructed based on the YOLOV8 model. The target detection model includes a backbone network, a neck network and a detection network connected in sequence. The backbone network extracts features from the sonar image to be detected based on the LANConvNeXtv2 module. The LANConvNeXtv2 module includes a convolution module and an attention mechanism module. The output of the convolution module is connected to the input of the attention mechanism module, and the output of the attention mechanism module is connected to the input of the neck network. The convolution module is used to extract the feature map of the sonar image to be detected, the attention mechanism module calculates the attention value of each block in the feature map based on the relationship between each block and the target object, and generates a first feature map carrying weight information according to the attention value, outputs the first feature map to the neck network, the neck network is used to enhance the first feature map to obtain a second feature map, and the detection network is used to detect the second feature map to obtain a detection result; The convolution module includes a two-dimensional convolution layer, a multi-scale convolution layer, a fusion layer and a dimension reduction layer; the output of the two-dimensional convolution layer is connected to the input of the multi-scale convolution layer, the output of the multi-scale convolution layer is connected to the input of the fusion layer, and the output of the fusion layer is connected to the input of the dimension reduction layer; The two-dimensional convolution layer is used to extract the local features of the sonar image to be detected, the multi-scale convolution layer extracts the features of different scales of the local features through convolution kernels of different sizes to obtain the features of different scales, the fusion layer is used to fuse the features of different scales to obtain fusion features, the dimension reduction layer is used to perform channel dimension reduction on the fusion features to obtain reduced dimension features, and the reduced dimension features are used as the output of the convolution module; The trained target detection model is used to perform target detection on the sonar image to be detected to obtain the target detection result.
2. The method according to claim 1, characterized in that The neck network is used to enhance the first feature map; comprising: Calculate the sampling points of the target object according to the scale information of the first feature map, and form a sampling point collection according to the sampling points; The sampling point collection and the first feature map are processed using a Grid Sample operation to obtain a sampled second feature map.
3. The method according to claim 2, characterized in that The step of calculating the sampling points of the target object according to the scale information of the first feature map comprises: Determining a range factor and an offset, the range factor comprising a static range factor and a dynamic range factor; Calculating a static range factor offset according to the static range factor, and calculating a dynamic range factor offset according to the dynamic range factor; The sampling point is determined according to a sum of the static range factor and the static range factor offset and a sum of the dynamic range factor and the dynamic range factor offset.
4. The method according to claim 1, characterized in that: The step of constructing a training sample set includes: Obtain original seabed side-scan sonar images; Construct background processing dataset and target object processing dataset; Identifying background area pixels in the seabed sidescan sonar raw image; Generate random noise based on the background area pixels, use the random noise to process the target object in the seabed side-scan sonar original image, generate background processing samples, and add them to the background processing data set; Identify the boundary of the target object in the original seabed side-scan sonar image, process the pixels within the boundary, generate target object samples, and add them to the target object processing data set; The background processed data set and the target object processed data set are mixed to obtain a training sample set.
5. The method according to claim 1, characterized in that The constructing of the training sample set also includes: Constructing an initial data set, wherein the initial data set includes a plurality of seabed background sonar images and target object sonar images; Extracting the target object from the target object sonar image, and combining the extracted target object with the seabed background sonar image to form a plurality of first target object sonar images; performing splicing processing on the target object sonar image and the first target object sonar image to form a plurality of second target object sonar images; performing enhancement processing on the target object sonar image, the first target object sonar image, and the second target object sonar image to form a plurality of third target object sonar images; The target object sonar image, the first target object sonar image, the second target object sonar image and the third target object sonar image are updated to the initial data set to obtain a training sample set.
6. The method according to claim 5, characterized in that The forming of a plurality of first target object sonar images comprises: Determining target object pixels according to a binary mask of each pixel in the sonar image of the target object, retaining the target object pixels, and determining the target object according to the target object pixels; The target object is pasted into the seabed background sonar image to obtain the plurality of first target object sonar images.
7. The method according to claim 5, characterized in that The forming of a plurality of sonar images of the second target object comprises: Determine the target object sonar image and the first target object sonar image as a first scale image; Randomly select M images from the first-scale images, and splice the M images to form a second-scale target image; Randomly select N images from the first-scale images, and splice the N images to form a third-scale target image; The first scale image, the second scale target image and the third scale target image are used as the second target object sonar images.
8. The method according to claim 5, characterized in that The forming of a plurality of sonar images of the third target object also includes: Determine the target object in the target object sonar image, the first target object sonar image, and the second target object sonar image; The target object is preprocessed to obtain the sonar image of the third target object.
9. A side-scan sonar target detection device, characterized in that: The device comprises a construction module, a training module and a detection module; wherein, The construction module is used to construct a training sample set, wherein the training sample set includes a plurality of seabed background sonar images and target object sonar images; The training module is used to train the target detection model based on the training sample set; The target detection model is constructed based on the YOLOV8 model. The target detection model includes a backbone network, a neck network and a detection network connected in sequence. The backbone network extracts features from the sonar image to be detected based on the LANConvNeXtv2 module. The LANConvNeXtv2 module includes a convolution module and an attention mechanism module. The output of the convolution module is connected to the input of the attention mechanism module, and the output of the attention mechanism module is connected to the input of the neck network. The convolution module is used to extract the feature map of the sonar image to be detected, the attention mechanism module calculates the attention value of each block in the feature map based on the relationship between each block and the target object, and generates a first feature map carrying weight information according to the attention value, outputs the first feature map to the neck network, the neck network is used to enhance the first feature map to obtain a second feature map, and the detection network is used to detect the second feature map to obtain a detection result; The convolution module includes a two-dimensional convolution layer, a multi-scale convolution layer, a fusion layer and a dimension reduction layer; the output of the two-dimensional convolution layer is connected to the input of the multi-scale convolution layer, the output of the multi-scale convolution layer is connected to the input of the fusion layer, and the output of the fusion layer is connected to the input of the dimension reduction layer; The two-dimensional convolution layer is used to extract the local features of the sonar image to be detected, the multi-scale convolution layer extracts the features of different scales of the local features through convolution kernels of different sizes to obtain the features of different scales, the fusion layer is used to fuse the features of different scales to obtain fusion features, the dimension reduction layer is used to perform channel dimension reduction on the fusion features to obtain reduced dimension features, and the reduced dimension features are used as the output of the convolution module; The detection module is used to perform target detection on the sonar image to be detected using the trained target detection model to obtain a target detection result.
Citation Information
Patent Citations
Seabed target detection method based on BES-YOLOv8 model
CN118736187A
Side-scan sonar shipwreck image target detection method
CN118823328A