A Method for Identifying Suspended Sonar Targets in Water Based on Region Pre-Detection

Through a region pre-detection method, the bright and shadowed areas of the suspended sonar target in water are characterized by stitching and splicing, and the graph attention network is used for identification, which solves the problem of difficulty in the recognition of suspended target in water and improves the recognition accuracy.

CN115810144BActive Publication Date: 2025-07-01NORTHWESTERN POLYTECHNICAL UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211376988.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-04
Publication Date
2025-07-01
Estimated Expiration
2042-11-04

AI Technical Summary

Technical Problem

In the detection and recognition of underwater targets, the bright and shadowed areas of suspended targets in the water are difficult to correlate their characteristics due to the large spacing distance, resulting in poor recognition effect.

Method used

The bright and shadowed areas of the suspended sonar target in water are extracted simultaneously by using a method based on area pre-detection. The bright and shadow areas of the suspended sonar target in water are sheared, matched and spliced ​​through the principle of multi-beam sonar imaging, and converted into graph structure data in combination with SLIC superpixel clustering, and target recognition is used using the graph attention network.

Benefits of technology

The effective combination of the characteristics of bright and shadowed areas is achieved, and the accuracy and efficiency of identification of suspended targets in water are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115810144B_ABST
    Figure CN115810144B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying suspended sonar targets in water based on regional pre-detection. This method simultaneously extracts the target bright area and the shadow area of the suspended target in water through regional pre-detection, and then uses a graph attention network to extract and associate the spatial position features of the detection results of the two areas, thereby effectively improving the recognition effect of sonar targets. First, the YOLOv5s network is used to pre-detect the target bright area and the shadow area of the sonar image, and then regional cropping is performed according to the detection results. After cropping, the two areas belonging to the same target are matched and spliced according to the imaging principle of the multi-beam sonar, so as to achieve the physical connection of the target bright area and the shadow area similar to the bottom target. The regional pre-detection proposed by the present invention can effectively associate the information of the sonar target bright area and the shadow area under the condition that the environmental parameters and target parameters are unknown, thereby effectively improving the recognition performance of sonar targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of underwater target detection and recognition, and particularly relates to a method for identifying underwater suspended sonar targets based on regional pre-detection. Background Art

[0002] Underwater target detection and recognition is one of the most extensive applications of sonar, and can be used for underwater search and rescue, submarine pipeline detection, anti-mine and frogman, etc. Sonar target detection mainly presents in two forms: instantaneous acoustic signal detection and regional imaging detection. Among them, imaging sonar is more widely used because of its more intuitive nature. Therefore, the focus of the present invention is on underwater target detection and recognition under imaging sonar.

[0003] Sonar targets mainly include bottom targets and underwater suspended targets. Their imaging in sonar images includes three main parts: target highlight area, target shadow area and reverberation background. In addition to the target bright area containing some target information, the shadow area also contains target contour and height information, etc. This part of the effective features is often easily ignored. Therefore, combining the features of the bright area and the shadow area for sonar target recognition can achieve better recognition effects. The bright area and the shadow area of the bottom target are closely connected in the sonar image, while the suspended target has a certain height from the bottom, resulting in a distance between the highlight area and the shadow area. Therefore, when extracting the features of the suspended target object from the sonar image, it is difficult to associate the bright area and the shadow features belonging to the same target. If the two areas are physically connected, the features of the two areas still need to be combined during feature extraction. The graph attention network based on non-Euclidean space can well extract spatial position features, and through aggregating neighbor information, realize the learning of neighborhood features and spatial features and the feature combination of the bright area and the shadow area of the sonar target.

[0004] Based on the above considerations, for the detection and recognition of underwater suspended targets, a method for identifying underwater suspended sonar targets based on regional pre-detection is proposed. This method uses regional pre-detection to simultaneously extract the bright area and the shadow area of the suspended target, and then according to the multi-beam sonar imaging principle, the two areas belonging to the same target are sheared, matched and spliced. The splicing result is converted into graph structure data through SLIC superpixel clustering to realize the joint of the two area feature spaces, and finally the graph attention network is used to identify the target of the constructed sonar data. Summary of the Invention

[0005] In view of the above existing technical problems, the present invention discloses a method for identifying underwater suspended sonar targets based on regional pre-detection.

[0006] The object of the present invention is a method for identifying underwater suspended sonar targets based on regional pre-detection, and the steps of this method are as follows:

[0007] S1: According to the multi-beam sonar imaging principle and imaging effect, perform coordinate transformation and image enhancement preprocessing on the original sonar image.

[0008] S2: Based on the YOLOv5s network, pre-detect the bright areas and shadow areas of the suspended targets simultaneously, and obtain the corresponding category information and target box position information.

[0009] S3: Cut the bright areas and acoustic shadow areas respectively according to the category and target box position information.

[0010] S4: Match and splice the bright area and shadow shear results belonging to the same target according to the imaging principle.

[0011] S5: Use the SLIC superpixel algorithm to construct Graph (graph) structure data for the spliced result.

[0012] S6: Construct an underwater suspended target recognition model based on GAT (Graph Attention Network).

[0013] S7: Ablation experiment settings, verify the effectiveness of using regional pre-detection and suspended target acoustic shadow area information, and verify the impact of different SLIC clustering results on the recognition rate.

[0014] Further, step S1 includes the following steps:

[0015] S11: Multi-beam sonar image reconstruction (coordinate transformation):

[0016] Multi-beam forward-looking sonar images are mainly presented in two ways: fan-shaped images and rectangular images. Among them, the fan-shaped sonar image based on the Cartesian coordinate system can well restore the underwater real situation, and the rectangular sonar image based on the polar coordinate system can very intuitively observe the position of the target underwater. And in order to more conveniently use the target box position information to crop the area after marking and pre-detecting the bright area and acoustic shadow area of the target, the following coordinate transformation is used to convert the fan-shaped image into a conventional rectangular image during the preprocessing process:

[0017]

[0018] where φ and R respectively represent the horizontal opening angle and slant range size of the multi-beam sonar, and W and H respectively represent the horizontal and vertical dimensions of the image.

[0019] S12: Multi-beam sonar image enhancement:

[0020] Aiming at the problems of low resolution and serious noise in the sonar image, it is necessary to perform filtering processing and image enhancement on the original image. The preprocessing specifically includes the following steps:

[0021] (1) Pixel enhancement: Multiply all pixel points in each sonar image by a certain factor. Set the pixel value at each pixel position as f(x, y), and the amplification factor is Then the pixel value of each pixel point after pixel enhancement should be:

[0022] (2) Median filtering: Select a square area (neighborhood) centered on a certain pixel, then sort the gray values of each pixel in the neighborhood except the central pixel, and use the sorted median as the new value of the central pixel point. When the neighborhood window slides orderly within the entire image range, the median filtering algorithm can be used to complete the filtering process for the entire image.

[0023] Further, the step S2 includes the following steps:

[0024] S21: Construct a YOLOv5s network model and algorithm for detecting the bright area and shadow area of the suspended target:

[0025] In order to better combine the features of the high-brightness area and shadow area of the same target and achieve fast detection speed, this part uses the YOLOv5s network, which has the fastest recognition speed in the YOLO detection algorithm series, to pre-detect the target high-brightness area and shadow area in the sonar image, and obtain the positions of the target bright area and shadow area in each multi-beam sonar image. In this step, the sonar target is not detected, only the areas with targets and the areas of the target shadow area in the image are detected. The specific steps for using the YOLOv5s network to pre-detect the bright area and shadow area of the suspended target in the multi-beam sonar image are as follows:

[0026] (1) Preprocessing. Use the Mosaic technology at the input end to perform data enhancement on the input multi-beam sonar image, then perform adaptive scaling, and finally send it into the network for training according to the preset batch number.

[0027] (2) Forward propagation. The image is sent into the backbone network for feature extraction, then feature fusion based on the Neck structure is used, and finally predictions are made to obtain the positions and sizes of the bright area and shadow area of the sonar target.

[0028] (3) Error calculation. According to the predicted class results and the predicted target box positions obtained in the forward propagation, use the predefined loss function to calculate the error size between them and the Ground truth.

[0029] (4) Parameter update. According to the calculated error results, use the Adam optimizer to update the network parameters in the forward propagation, and continuously iterate to reduce the error. After the iteration stops, select the network parameters corresponding to the minimum prediction error for detecting the image to be detected.

[0030] (5) Target detection. Replace the finally selected network parameters into the forward propagation, and then detect the bright area and the acoustic shadow area of the sonar target on the image to be detected, and obtain the corresponding target position and the size of the target box.

[0031] S22: Collect sonar image data containing various suspended targets in water, conduct data construction, and use the constructed dataset to train the model:

[0032] Multi-beam sonar image acquisition environment: Mount the multi-beam sonar on a small fishing boat, search for targets in a lake with a water depth of 5 - 10 meters. The sonar is placed underwater 0.5 meters from the water surface, and the inclination angle of the sonar from the water surface is 30°. Five types of targets, namely dummies, tires, spheres, cylinders, and cubes, are respectively placed, and each target is suspended. By sailing the fishing boat along different trajectory lines, and at the same time making the multi-beam sonar irradiate the target objects from different angles, finally form target imaging in multiple directions.

[0033] Dataset construction: After collecting field data, a total of 803 multi-beam sonars containing suspended targets in water are sorted out, among which there are 300 sphere targets, 204 cube targets, 85 cylinder targets, 103 dummy targets, and 111 tire targets. After the dataset is sorted out, it is divided into a training set, a validation set, and test samples according to the ratio of 8:0.5:1.5.

[0034] Model training: Use the Labelme annotation software to annotate all the sample sets with two label types, namely the bright area and the acoustic shadow area of the target. After annotation, modify the category parameters, backbone model, pre-trained weights, etc. in the network, and then send the sample sets into the detection network for model training, and obtain the model parameters under the optimal detection effect.

[0035] S23: Use the trained model to pre-detect the regions of the multi-beam sonar image to be detected, and obtain the target box position information of each region:

[0036] According to the continuous iteration of model training until the training loss function converges, after convergence, select the network parameters with the best recognition for the detection of the bright area and the acoustic shadow area of the sonar target in the image to be detected. Through the detection of the trained network model, category information and target box information are obtained at the same time. The category information includes two situations: the bright area and the dark area; the target box information includes four elements: [x center , y center , w, h], which respectively represent the normalized abscissa of the center position point of the target box, the normalized ordinate of the center position point of the target box, the normalized width of the target box, and the normalized height of the target box. Based on the above information, the region containing the target and its location in the original sonar image can be obtained.

[0037] Further, step S3 includes the following steps:

[0038] S31: Crop the bright area according to the position information of the target box in the bright area:

[0039] In order to separate multiple targets in a sonar image and realize the feature combination of the subsequent target bright area and the sonar shadow area, this part crops the target area of the above detection results, cuts out all areas containing targets in each sonar image, and provides samples for subsequent area matching and combination. For the target highlight area, according to the target box information of the sonar target bright area obtained above, the pixel position coordinates of the upper left corner and the lower right corner of the target box that can represent the original image range of the area are converted according to the following formula:

[0040]

[0041] After conversion, obtain the position area of each bright area target box in the original image, and then according to the obtained [x min , y min , x max , y max coordinates, crop the original image to finally obtain all target bright area images.

[0042] S32: Crop the shadow area according to the target box in the sonar shadow area:

[0043] Similar to the target bright area, for the target shadow area, according to the target box information of the target sonar shadow area obtained above, convert the position of the target box in the YOLO format to the pixel position according to the pixel position conversion formula, and then crop the area according to the pixel points to finally obtain all target shadow area images.

[0044] Further, step S4 includes the following steps:

[0045] S41: Match the bright area and the shadow belonging to the same target according to the imaging principle:

[0046] According to the imaging principle of the multi-beam sonar, the acoustic wave beam is emitted forward, reflected by the underwater medium, and finally the returned acoustic wave signal is received at the receiving transducer. Therefore, according to the acoustic wave propagation principle, during the propagation of the acoustic wave, the acoustic wave is blocked by suspended objects in the water and strongly reflected, thus forming a bright target area in the imaging effect. At the same time, in the area behind the bright target area at a certain length, there is no acoustic wave irradiation, thus forming a shadow area. In the fan-shaped sonar image, with the sonar transmitting transducer as the center point, in the triangular area with the boundary of the target object as the side, both the bright target area and the shadow area are within this range, and the triangular boundary of the target object and the triangular boundary of the shadow area formed by the target object at a certain distance are the same boundary; correspondingly, in the rectangular sonar image, the boundary of the target object and the boundary of the shadow area formed by the target object at the back are also the same boundary. Therefore, according to this imaging principle, the high-brightness area and the shadow area belonging to the same sonar target in a sonar image are matched.

[0047] S42: Stitch the region matching results:

[0048] After completing the above matching work, it is necessary to substantially associate the two parts of the bright area and the shadow area. The most intuitive physical connection is stitching. For underwater targets, the high-brightness area and the shadow area of the target are connected together, so the features of the two areas are directly combined. However, for suspended targets in the water, due to the certain height of the suspended target from the bottom of the water, the shadow imaging will be at a certain distance outside the bright area imaging and will not be directly connected behind the bright area. Therefore, it is necessary to physically connect the two areas of the bright area and the shadow area obtained by pre-detection. After cutting according to the region position information obtained by pre-detection, only the region containing the target information remains, and the redundant background regions have been removed. After matching, the high-brightness area and the shadow area belonging to the same target have been obtained. Inspired by the imaging of underwater targets in the sonar image, in order to be closer to the imaging effect of the multi-beam sonar, the high-brightness area and the shadow area belonging to the same target are vertically stitched, with the bright area below and the shadow above, thus completing the preprocessing of the entire sonar target.

[0049] Further, the step S5 includes the following steps:

[0050] 51: Use the SLIC clustering algorithm to perform superpixel segmentation on the stitching result of the suspended sonar target:

[0051] In order to better associate the bright area features and the shadow area features of the target object, the present invention uses the SLIC clustering algorithm to convert the stitching result of the suspended sonar target into graph structure data, which is conducive to the extraction of the target spatial geometric features. The SLIC clustering algorithm assigns each pixel in the image data to a 5-dimensional vector V[I,a,b,x,y] T , where [l,a,b]T Denote the color feature coordinates, [x, y] T Denote the spatial feature coordinates. The steps of specifically using the SLIC clustering algorithm to perform superpixel segmentation on the stitching result of the suspended sonar target are as follows:

[0052] (1) Initialize the clustering centers.

[0053] (2) Update the positions of the clustering centers.

[0054] (3) Assign initial class labels.

[0055] (4) Measure the distance similarity.

[0056] (5) Iteratively optimize the clustering results.

[0057] S52: Construct Graph (graph) structure data for the suspended sonar target based on superpixel clustering:

[0058] After the SLIC superpixel clustering is completed, each pixel point has a corresponding superpixel center. Therefore, the final feature information (average metric of all pixel points within the region) [l, a, b, x, y] of the superpixel points can be obtained according to the following formula T :

[0059]

[0060] In the formula, N i represents the number of pixels included in the i-th superpixel center. After obtaining the superpixels and their feature information, it is necessary to convert them into graph structure data to realize the association in the feature space of the bright area and the sound shadow area. When converting the data, all the superpixel points obtained from each image are used as the nodes in the graph structure data, the feature information of the superpixel points is used as the node attributes in the graph structure data, and the Euclidean distance between two superpixels is used as the edge and its attribute in the graph structure data.

[0061] Furthermore, the step 52 includes the following steps:

[0062] S521: Graph (graph) structure data representation under the suspended sonar target image:

[0063] Graph (graph) is a non-Euclidean data structure, mainly including two major elements: nodes and edges, denoted as G = {V, E}, where V = {v1,... v M} represents the set of nodes, E = {e1,... e P} represents the set of edges, and both nodes and edges can have their attribute characteristics. In the present invention, the entire graph is discriminated, so it is necessary to define a label category for an entire graph structure, where the category includes 5 types: dummy, tire, cylinder, sphere, and cube. The nodes in each graph structure are defined as a sequence of superpixel centers obtained by SLIC clustering, and the information contained in each node is the feature information of the superpixel points (pixel mean feature sp-intensity and center position feature sp-coord); the edges between nodes are defined as the association between two nodes, and the information of the edge is the distance between two superpixel points.

[0064] S522: Definition of the Graph structure attributes under the hovering sonar target image:

[0065] Both pixel information and spatial position information are reflected in the attributes of the nodes in the graph structure, so it is necessary to define the attributes of the nodes. Among them, the pixel information is the pixel mean of all pixel points in each superpixel region, and the spatial position information is the center position coordinates of each superpixel point. Since the SLIC superpixel clustering result well preserves the target edge, the pixel information value of the superpixel where the target bright area is located is large, the pixel information value of the superpixel where the target sound shadow area is located is small, and the value of the background area is moderate, so the target boundary can be very effectively reflected. The nodes and their attributes are specifically represented as follows:

[0066]

[0067] Among them, (x i , y i ) represents the spatial position information, and f(x i , y i ) represents the pixel information.

[0068] The edge reflects the degree of association between two nodes, and the attribute of the edge can just be defined as the size of the association, that is, the weight of the edge. As mentioned above, in order to eliminate the redundancy of calculation, the edge weight between two nodes at a relatively long distance is assigned a value of 0, and the specific elimination rule is determined by the K-Nearest Neighbor algorithm (KNN). Take the number of neighbors as κ, then the number of edges with edge weight values is κ × M = κM, and the remaining edge weights are 0. When calculating the association of the edges with weights, it is necessary to consider both the pixel distance and the spatial position distance between two nodes at the same time. The specific calculation formula is as follows:

[0069]

[0070] Among them, δ x represents the average value of the spatial position distances between node v i and its κ nearest nodes, and δ f represents the point v iThe average of the pixel distances between the κ nearest nodes, where γ represents the relative ratio between the pixel information and the spatial position information. In summary, the attribute definitions in the graph structure data are summarized as follows:

[0071]

[0072] Furthermore, step S6 includes the following steps:

[0073] S61: Construction of a graph attention network for underwater suspended target recognition:

[0074] After converting the underwater suspended target area image into graph structure data, it is necessary to construct a graph attention network to perform iterative regression on the graph structure. In the graph attention network model, an attention module is used to embed the nodes in the graph. By calculating the attention coefficients between the current node and its neighbor nodes, the neighbor information is aggregated, and the adaptive allocation of different neighbor weights is realized, so as to learn the neighborhood features and spatial features.

[0075] The input and output of the entire graph attention layer can be expressed as follows:

[0076]

[0077] In the formula is the input of the graph attention layer, representing the combination of the pixel features and spatial position features of each node; Q is the weight matrix, representing the linear transformation between the input and the output; σ is the non-linear activation function; α ij represents the attention coefficient between nodes i and j, which can be obtained from the following formula:

[0078]

[0079] In the formula represents the attention mechanism, || represents the concatenation operation, and k represents the number of selected neighborhoods. In order to eliminate redundancy, it is necessary to jointly consider the edge weight matrix after correlation adjustment, so as to obtain the final expression of the attention mechanism coefficient as:

[0080]

[0081] In the formula, W i,j is the updated adjacency matrix. For each node, the weights W i,j of the top k edges with strong connectivity to it are assigned 1, and the weights of the remaining edges are assigned negative infinity.

[0082] S62: Dataset construction:

[0083] The present invention constructs a graph attention network based on the DGL framework. Therefore, it is necessary to convert the graph structure data into the standard DGLGraph data format. First, the number of nodes in the graph structure is used as the node ID in the DGLGraph, and the pixel information and spatial position information of the nodes are used as the node attributes ndata. Secondly, the numbers between the two nodes connected by the edge are used as the edge ID pair, and the weight of the edge is used as the edge attribute edata. Finally, the label of the graph structure is used as the label of the DGLGraph, thus completing the construction of the data in the standard format.

[0084] After completing the conversion of the standard data format, it is necessary to divide all the collected valid data into datasets, which are divided into a training set, a validation set, and a test set according to the ratio of 8:0.5:1.51. The collected data contains a total of 5 types of suspended targets, that is, the graph structure data labels are also divided into 5 types, namely dummy, tire, cylinder, sphere, and cube. A total of 803 graph structures are created. After dividing according to the ratio of 8:0.5:1.51, the training set contains 646 graph structures, the validation set contains 43 graph structures, and the test set contains 114 graph structures.

[0085] S63: Experimental settings and model training:

[0086] Experimental settings: In the hardware device, the CPU model used is silver 4110 CPU@2.10Ghz, and the GPU model is NVIDIA GeForce RTX 3080. In the model framework, the PyTorch deep learning framework is used, and DGL is used as the graph neural network framework. In the model, the parameters of the graph attention layer, the hidden layer, the dimension of the output feature vector, the number of attention mechanisms, and the node aggregation method are set respectively. In the training parameters, the dataset type, the number of target categories, the number of epochs, the batch-size, the initial learning-rate, and the decay coefficient of the learning-rate are set respectively.

[0087] Model training:

[0088] (1) Data reading and preprocessing. After obtaining the original graph structure data, the K-nearest neighbor algorithm is used to eliminate the redundancy of the edge connections, thereby updating the weights of the edges and the adjacency matrix in the graph structure.

[0089] (2) Forward propagation. The graph structure data is sent into the backbone network based on the graph attention mechanism for message passing using GAT convolution.

[0090] (3) Error calculation. The loss function is used to calculate the error between the prediction result obtained in the forward propagation and the true result.

[0091] (4) Iterative optimization and parameter update. By continuously iteratively optimizing the error, the forward propagation parameters of the model are continuously updated until the error converges, the parameters no longer change, and the recognition effect of the model reaches the optimal.

[0092] (5) Model prediction. The multi-beam sonar images in the validation set are sent into the trained model to complete the target recognition work.

[0093] S64: Model testing and result analysis:

[0094] In the testing stage, the training module in the network is turned off, the code for the validation and testing part is enabled, and the category recognition of all the graph structure data in the test set is performed using the model parameters with the best effect during the training process, so as to obtain the recognition accuracy under the test set, and the network model is tested with this index.

[0095] Further, the step S7 includes the following steps:

[0096] S71: Verify the effectiveness of using regional pre-detection and the shadow area information of suspended targets:

[0097] In order to verify the effectiveness of regional pre-detection and the shadow area information of suspended targets, a total of 3 types of graph structure data are made. The first type is the graph structure data made from the most original multi-beam sonar images, and this type of data is directly converted into graph structure data through SLIC clustering without any processing; the second type is after performing regional pre-detection on the original multi-beam sonar images, only extracting the target bright area images, and then converting the images of the highlighted areas into graph structure data; the third type is after performing regional pre-detection on the original images, simultaneously cutting out the target highlighted area and the shadow area, longitudinally splicing the two areas, and finally converting the splicing result into graph structure data. After making the 3 types of graph structure data, each type of data is used for model training, and their respective training losses and the recognition accuracies of the test sets are obtained respectively. Finally, the effectiveness of regional pre-detection and the shadow area information of suspended targets is verified from the perspective of recognition accuracy.

[0098] S72: Verify the influence of different SLIC clustering results on the recognition rate:

[0099] The metric coefficient τ of the relative importance between color similarity and spatial proximity is the most important factor affecting the SLIC clustering effect. The smaller the value of τ, the more the final clustering effect tends to the color feature; the larger the value of τ, the more the final clustering effect favors the spatial location feature. Therefore, the different recognition effects caused by different values of τ can well illustrate the respective importance of the spatial location feature and pixel feature in the multi-beam sonar image. To explore the influence of different clustering effects on the multi-beam sonar target recognition effect, the relative importance metric parameter τ is respectively set to 1, 5, 10, and 20, and the SLIC clustering effects under different parameter values are made into a graph structure dataset, thus obtaining 4 types of graph structure data. Then, each dataset is used to train the model to obtain their respective training losses and the recognition accuracies of the test sets. Finally, the relative importance distribution of the spatial location feature and pixel feature is explained using the recognition accuracy of the model. Brief Description of the Drawings

[0100] Figure 1 is the flowchart of the present invention;

[0101] Figures 2a - 2b is the multi-beam sonar imaging effect diagram of the present invention;

[0102] Figure 3 is the schematic diagram of the multi-beam sonar image preprocessing result of the present invention;

[0103] Figure 4 is the schematic diagram of the detection result of the target bright area and shadow area in the multi-beam sonar image of the present invention;

[0104] Figure 5 is the schematic diagram of the process of shear matching and splicing after pre-detecting the target area of the present invention;

[0105] Figure 6 is the effect diagram after SLIC clustering of the original multi-beam sonar image of the present invention;

[0106] Figure 7 is the SLIC clustering effect diagram after pre-detecting the bright area and shadow area of the present invention;

[0107] Figure 8 is the SLIC clustering effect diagram after pre-detecting the bright area of the present invention;

[0108] Figure 9 is the training loss function and training accuracy result diagram after pre-detecting the area of the present invention;

[0109] Figure 10 is the loss and recognition curve of the validation set of the present invention, as well as the recognition rate curve of the test set;

[0110] Figure 11It is a comparison chart of the training loss function values under the pre-detection of the verification area of the present invention and the effectiveness of the information in the suspended target shadow area;

[0111] Figure 12 It is a comparison chart of the recognition accuracy under the pre-detection of the verification area of the present invention and the effectiveness of the information in the suspended target shadow area;

[0112] Figure 13 It is a comparison chart of the training loss function values under the verification of different SLIC clustering results of the present invention;

[0113] Figure 14 It is a comparison chart of the recognition accuracy under the verification of different SLIC clustering results of the present invention. Detailed implementation manners

[0114] The following will describe the embodiments of the present invention in detail with reference to the accompanying drawings.

[0115] Refer to Figure 1 , Figure 1 It is a flow chart of a method for identifying sonar targets suspended in water based on region pre-detection provided by the present invention, including the following steps:

[0116] S1: According to the multi-beam sonar imaging principle and imaging effect, perform coordinate transformation and image enhancement preprocessing on the original sonar image.

[0117] S2: Based on the YOLOv5s network, pre-detect the bright area and shadow area of the suspended target simultaneously, and obtain the corresponding category information and target box position information.

[0118] S3: Cut the bright area and the shadow area respectively according to the category and target box position information.

[0119] S4: Match and splice the bright area and shadow shear results belonging to the same target according to the imaging principle.

[0120] S5: Use the SLIC superpixel algorithm to construct Graph (graph) structure data for the spliced result.

[0121] S6: Construct a recognition model for suspended targets in water based on GAT (Graph Attention Network).

[0122] S7: Set up ablation experiments to verify the effectiveness of using region pre-detection and the information in the suspended target shadow area, and verify the influence of different SLIC clustering results on the recognition rate.

[0123] Furthermore, the step S1 includes the following steps:

[0124] S11: Multi-beam sonar image reconstruction (coordinate transformation):

[0125] The multi-beam forward-looking sonar is an active sonar, generally installed in front of an underwater mobile platform for observing and detecting the environment in front of the platform. It transmits a single-frequency narrowband acoustic pulse signal through the entire detection area and uses the method of multi-beam electronic scanning to achieve imaging of the target scene. When the system works, the acoustic intensity data collected by the sonar transducer is presented in the form of a fan-shaped sonar image by the system. This form fits the real underwater environment, and its field of view is essentially a three-dimensional space, which is then mapped onto a two-dimensional plane through projective transformation. The three-dimensional field of view range and the two-dimensional imaging plane are as Figure 2a shown. The fan-shaped sonar image based on the Cartesian coordinate system can well restore the real underwater situation, and the rectangular sonar image based on the polar coordinate system can very intuitively observe the position of the target underwater. However, in order to more conveniently use the position information of the target box to crop the area after marking and pre-detecting the bright area and acoustic shadow area of the target, it is necessary to convert the fan-shaped image into a conventional rectangular image using coordinate transformation. The specific conversion is as Figure 2b shown. Let (r, θ) be the coordinate axes of the fan-shaped image and (x, y) be the coordinate axes of the conventional rectangular image. The conversion formula between the two coordinate systems is as follows:

[0126]

[0127] where φ and R represent the horizontal opening angle and slant range of the multi-beam sonar respectively, and W and H represent the horizontal and vertical dimensions of the image respectively. After coordinate transformation, the multi-beam sonar image is in the same form as a conventional optical image. The horizontal axis of the image represents the imaging results at different azimuth angles, and the vertical axis represents the imaging effects at different slant ranges.

[0128] S12: Enhancement of multi-beam sonar image:

[0129] If no gain amplification is performed during the acquisition process of the multi-beam sonar image, the collected acoustic intensity data will be very weak, making it difficult to distinguish the target position and target category with the naked eye. Moreover, due to the low resolution of the sonar, there is a large area of noise in the image. Therefore, in the post-processing, it is necessary to perform filtering processing and image enhancement on the original sonar image. The preprocessing specifically includes the following steps:

[0130] (1) Pixel enhancement: Multiply all pixel points in each sonar image by a factor. Let the pixel value at each pixel position be f(x, y), and the amplification factor be Then the pixel value of each pixel point after pixel enhancement should be: In the present invention

[0131] (2) Median Filtering: The basic idea of median filtering is to replace the gray value of the central pixel with the median gray value in the neighborhood. Median filtering can weaken high-frequency components and achieve good denoising effects when denoising images containing impulse noise. Its main process is as follows: Usually, a square area (neighborhood) centered on a certain pixel is selected, and then the gray values of each pixel in the neighborhood except the central pixel are sorted. The sorted median is used as the new value of the central pixel. When the neighborhood window slides orderly within the entire image range, the entire image can be filtered using the median filtering algorithm. The definition of the selected area median is as follows: Consider the pixels within the neighborhood window to be processed as data x1, x2, x3, … x n , and sort them by size. Then the median is:

[0132]

[0133] Reference Figure 3 shows the multi-beam sonar image after the above preprocessing. Figure 3 The first row is the rectangular multi-beam sonar image after coordinate transformation, and the second row is the result after image preprocessing. From the final processing results, it is very clear to observe the target highlight area of the suspended target and the acoustic shadow area that appears in the background due to the suspension of the target.

[0134] Furthermore, step S2 includes the following steps:

[0135] S21: Construct a YOLOv5s network model and algorithm for detecting the bright area and acoustic shadow area of the suspended target:

[0136] When using a multi-beam sonar to detect suspended targets in deeper waters, the imaging effect of small targets is poor, and the target bright area occupies fewer pixels in the entire sonar image. Therefore, it is very difficult to identify the target category only using the characteristics of the sonar target bright area. Due to the position characteristics of the suspended target, a large area of acoustic shadow area will appear behind the target during multi-beam sonar imaging, and the imaging characteristics of the acoustic shadow area are closely related to the shape of the target object. Therefore, combining the bright area characteristics and acoustic shadow area characteristics of the target can greatly improve the accuracy of target recognition for underwater targets. To better combine the two characteristics, this part uses the YOLOv5s network with the fastest recognition speed to pre-detect the target highlight area and acoustic shadow area in the sonar image to obtain the positions of the target bright area and shadow area in each multi-beam sonar image. In this step, the sonar target is not detected, only the areas with targets and the areas of the acoustic shadow area in the image are detected. Therefore, it is simply called regional pre-detection. The specific steps of the specific regional pre-detection are as follows:

[0137] (1) Preprocessing. At the input end, the Mosaic technique is used to enhance the data of the input multibeam sonar image, followed by adaptive scaling, and finally it is fed into the network for training according to the preset batch number.

[0138] (2) Forward propagation. The image is fed into the backbone network for feature extraction, then feature fusion based on the Neck structure is used, and finally predictions are made to obtain the positions and sizes of the bright areas and shadow areas of the sonar targets.

[0139] (3) Error calculation. According to the predicted class results and the predicted target box positions obtained in the forward propagation, the error size between them and the Ground truth is calculated using a predefined loss function.

[0140] (4) Parameter update. According to the calculated error results, the network parameters in the forward propagation are updated using the Adam optimizer, and the error is reduced through continuous iteration. After the iteration stops, the network parameters corresponding to the minimum prediction error are selected for detecting the image to be detected.

[0141] (5) Object detection. The finally selected network parameters are replaced into the forward propagation, and then the bright areas and shadow areas of the objects in the image to be detected are detected to obtain the corresponding object positions and object box sizes.

[0142] S22: Collect sonar image data containing various underwater suspended targets, perform data construction, and use the constructed dataset to train the model:

[0143] Multibeam sonar image acquisition environment: The multibeam sonar is mounted on a small fishing boat to search for targets. The sonar is placed underwater 0.5 meters from the water surface, and the inclination angle of the sonar from the water surface is 30°. The depth of the entire water area is 5 - 10 meters. Dummies, tires, spheres, cylinders, and cubes are respectively placed as targets, and each target is suspended. By sailing the fishing boat along different trajectory lines and simultaneously irradiating the targets from different angles with the multibeam sonar, finally, target imaging in multiple directions is formed.

[0144] Dataset construction: The purpose of this part is to detect the target areas and shadow areas in the sonar images formed by each ping, without performing class recognition of each target. Therefore, when constructing the dataset, only the target bright areas and target shadow areas in each image are labeled. After the above field data collection, a total of 803 multibeam sonars containing underwater suspended targets are sorted out, including 300 sphere targets, 204 cube targets, 85 cylinder targets, 103 dummy targets, and 111 tire targets. After the dataset is sorted out, it is divided into a training set, a validation set, and test samples according to the ratio of 8:0.5:1.5.

[0145] Model training: After segmentation, all sample sets are labeled with two label types, namely the target highlight area and the acoustic shadow area, using the Labelme annotation software. After annotation, modify the class parameters, backbone model, pre-trained weights, etc. in the network, and then send the sample set into the detection network for model training to obtain the model parameters under the optimal detection effect.

[0146] S23: Use the trained model to perform regional pre-detection on the multi-beam sonar image to be detected, and obtain the target box position information of each region:

[0147] According to the continuous iteration of model training until the training loss function converges, after convergence, select the network parameters with the best recognition for the detection of the sonar target bright area and acoustic shadow area in the image to be detected. Specifically: in the test code, replace the default weight path with the trained weight result and modify the corresponding class parameters, and finally perform the detection recognition and position regression of the bright area and shadow area of the sonar target on the test set data.

[0148] After the detection of the trained network model, both class information and target box information are obtained. The class information includes two cases: bright area and dark area; the target box information includes four elements: [x center ,y center ,w,h], which respectively represent the normalized abscissa of the center position point of the target box, the normalized ordinate of the center position point of the target box, the normalized width of the target box, and the normalized height of the target box. According to the above information, the area containing the target and its location in the original sonar image can be obtained.

[0149] Reference Figure 4 shows the detection results of the target bright area and the shadow area in the sonar image to be detected using YOLOv5s. As shown in the figure, all sonar target areas in the sonar image have been detected. From the detection results, the geometric shapes of the target bright areas are roughly similar, so it is very difficult to classify the target classes only using the information of the target bright areas; due to the floating characteristics of the targets, the area of the shadow area is much larger than that of the bright area, and the target features are more obvious. Therefore, the joint recognition of the information of the two areas can effectively improve the recognition effect of sonar targets.

[0150] Further, the step S3 includes the following steps:

[0151] S31: Crop the bright area according to the target box position information of the bright area:

[0152] When using a graph network for sonar target recognition, if the original image is directly used for recognition, the features of the target bright area and the sonar shadow area cannot be jointly extracted. Instead, specific and independent feature extraction is performed for each area. Moreover, when there are multiple targets in a sonar image, the network will attribute the features of multiple targets to one target, that is, only single-target detection can be performed, resulting in misjudgment of sonar target categories. Considering the above problems, in this part, the above detection results are cropped for the target area, and all areas containing targets in each sonar image are cut out to provide samples for subsequent area matching and combination.

[0153] For the target highlight area, according to the target box information of the sonar target bright area obtained above, the pixel position coordinates of the upper left corner and the lower right corner of the target box that can represent the original image range where the area is located are converted according to the following formula:

[0154]

[0155] After conversion, the position area of each bright area target box in the original image is obtained. Then, according to the obtained [x min , y min , x max , y max coordinates, the original image is cropped, and finally all target bright area images are obtained.

[0156] S32: Crop the shadow area according to the target box of the sonar shadow area:

[0157] Similar to the target bright area, for the target shadow area, according to the target box information of the target sonar shadow area obtained above, the position of the target box in the YOLO format is converted to the pixel position according to the pixel position conversion formula, and then the area is cropped according to the pixel points, and finally all target shadow area images are obtained.

[0158] Further, the step S4 includes the following steps:

[0159] S41: Match the bright area and the shadow belonging to the same target according to the imaging principle:

[0160] After cropping a single original sonar image, all the obtained target bright area and sonar shadow area images are mixed together, and the bright area and the shadow belonging to the same target cannot be combined yet. Therefore, it is necessary to match the highlight area and the shadow area belonging to the same sonar target according to the multi-beam sonar imaging principle.

[0161] According to the imaging principle of the multi-beam sonar, the acoustic wave beam is emitted forward, reflected by the underwater medium, and finally the returned acoustic wave signal is received at the receiving transducer. Therefore, according to the acoustic wave propagation principle, during the propagation of the acoustic wave, the acoustic wave is blocked by the suspended objects in the water and generates strong reflection, thus forming a target bright area in the imaging effect. At the same time, in the area behind the target bright area at a certain length, there is no acoustic wave irradiation, thus forming a shadow area. In the fan-shaped sonar image, in the triangular area with the sonar transmitting transducer as the center point and the boundary of the object as the side, both the target bright area and the shadow area are within this range, and the triangular boundary of the object is the same as the triangular boundary of the shadow area formed by the object at a certain distance; correspondingly, in the rectangular sonar image, the boundary of the object is the same as the boundary of the shadow area formed by the object behind, so according to this imaging principle, the matching of the high-brightness area and the shadow area belonging to the same sonar target in a sonar image is completed.

[0162] Reference Figure 2b It shows the imaging principle of the suspended sphere target in water in the multi-beam sonar image. On the left side of the figure is the imaging of the suspended sphere target under the fan-shaped sonar image. It can be clearly seen from the figure that the bright area and the shadow area belonging to the same target are in the same triangular area, and the acoustic ray with the sonar transducer as the vertex is tangent to both the high-brightness area and the shadow area at the same time. On the right side of the figure is the imaging of the suspended sphere target under the rectangular sonar image. The abscissa in the rectangular coordinate system represents the sonar beam angle, and the ordinate represents the distance. The abscissas of the bright area imaging and the shadow imaging belonging to the same target are the same, thus presenting that the bright area and the shadow are on the same vertical line. According to this principle, the matching of the bright area and the shadow belonging to the same target can be completed.

[0163] S42: Stitch the region matching results:

[0164] After completing the above matching work, it is necessary to substantially associate the bright area and the shadow area. The most intuitive physical connection is splicing. For underwater targets, the high-light area and the shadow area of the target are connected together, so the features of the two areas can be directly combined. However, for the suspended targets in water considered in the present invention, since the suspended target has a certain height from the bottom of the water, the shadow imaging will be at a certain distance after the bright area imaging, and will not be directly connected behind the bright area. Therefore, it is necessary to physically connect the bright area and the shadow area obtained by pre-detection. After cutting according to the region position information obtained by pre-detection, only the region containing the target information remains, and the redundant background regions have been removed. After matching, the high-light area and the shadow area belonging to the same target have been obtained. Inspired by the imaging of underwater targets in sonar images, in order to be closer to the imaging effect of multi-beam sonar, the high-light area and the shadow area belonging to the same target are vertically spliced, with the bright area below and the shadow above, thus completing the preprocessing of the entire sonar target.

[0165] Reference Figure 5 shows the process of shearing, matching and splicing after obtaining the detection results through the YOLO detection network. The first column in the figure shows the high-light area and the shadow area of all targets in each multi-beam sonar image obtained by pre-detection through the YOLOv5s network, as well as the target box (label) information of the corresponding areas. The figure mainly shows the schematic diagrams of the sphere target and the cube target. The second column shows cutting the original image according to the target box information. After converting the target box information obtained during the detection process through coordinate transformation, the pixel coordinate position points to be cut are obtained. Finally, all target areas included in the original sonar image are cut according to the coordinate positions. It is easy to observe from the cut physical diagram that the target bright area and the shadow area each have obvious features. The third column shows the results of matching and splicing the cut areas. The background area between the bright area and the shadow area is removed in the results, and the two areas are fused together, thus effectively ensuring the feature combination of the subsequent high-light area and the shadow area.

[0166] Further, the step S5 includes the following steps:

[0167] 51: Use the SLIC clustering algorithm to perform superpixel segmentation on the splicing result of the suspended sonar target:

[0168] Since graph networks can only recognize graph-structured data, and graph-structured data can contain both pixel features and spatial location features, this part needs to convert the pre-stage stitching result image of the suspended sonar target into graph-structured data. During the data acquisition process, the height information of the mobile platform and the water area height information cannot be accurately known, so the approximate geometric size of the target object cannot be calculated based on the target bright area and shadow area. To better utilize the geometric features of the target object and realize the association between the bright area features and the acoustic shadow area features, the present invention uses the SLIC clustering algorithm to convert the stitching result of the suspended sonar target into graph-structured data, which is conducive to the extraction of the target spatial geometric features. The SLIC clustering algorithm assigns each pixel in the image data to a 5D vector V[I,a,b,x,y] T , where [l,a,b] T represents the color feature coordinates, and [x,y] T represents the spatial feature coordinates. The specific steps of using the SLIC clustering algorithm to segment the sonar target stitching result are as follows:

[0169] (1) Initialize the clustering centers:

[0170] Define the number of clustering center points (superpixels) as M, and make all the initial superpixel clustering centers evenly distributed in the image to be clustered, and the sizes of each superpixel are the same. If the total number of pixels is N, then the area of each superpixel is N / M, and the distance between two adjacent superpixels is In the present invention, the size of the stitching result image of the suspended sonar target is 200×400, N = 80000, and M = 100 is taken.

[0171] (2) Update the positions of the clustering centers:

[0172] Due to the regularity of the initially defined superpixel clustering centers, they are likely to fall on the contour boundaries with large gradients. Therefore, it is necessary to update the positions of the initial clustering centers within the area of the initial clustering centers to achieve a better initialization effect. Usually, it is selected within the 3×3 range of the clustering center points. The specific method is: calculate the gradient values of all pixel points within the 3×3 range of the clustering center points, and find the pixel point with the minimum gradient according to the following formula, which is the updated clustering center position:

[0173]

[0174] (3) Assign the initial class labels:

[0175] After the superpixel clustering center position is updated, it is necessary to assign labels to all pixel points within the area, that is, to define the superpixel numbers to which each pixel point within the defined range belongs. In the SLIC algorithm, the search range is defined within 2S×2S instead of the entire image, which can effectively reduce the workload of distance calculation, and limiting it to a smaller range can better preserve the edges of the objects in the image. The specific method is as follows: In this step, the label array and distance array of each pixel point are mainly defined. The label array is used to store the label sequence number value, and the distance array is used to store the distance value.

[0176] (4) Distance similarity measurement:

[0177] Distance similarity measurement is used to update the belonging domains of all pixel points within the superpixel area, that is, to which superpixel each pixel point should belong. For the measurement of distance similarity, both the pixel feature distance and the spatial position feature distance are considered. Since the orders of magnitude of the two features are different, it is necessary to normalize them before jointly considering the two features, and then perform feature merging. The specific method is as follows: Take the maximum feature distance N of the pixel features c and the maximum feature distance N of the spatial position features s =S. To reduce the computational complexity, define the measure constant τ = N c for the relative importance between the color similarity and the spatial proximity. Then the distance metric D between the pixel point and the clustering center is expressed as follows:

[0178]

[0179] In the formula, d c represents the pixel feature distance, and d s represents the spatial position feature distance. The larger the value of the constant τ, the more important the spatial position feature, and the more compact the formed clustering region; the smaller the value of the constant τ, the more important the pixel color feature, and the better the superpixel region formed preserves the edges of the objects in the image.

[0180] (5) Iteratively optimize the clustering result:

[0181] According to the distance similarity measurement rule, the belonging domains of all pixel points within the superpixel area are iteratively updated. During the iteration process, if the distance from the pixel point to the previous superpixel center is greater than the distance calculated in this iteration, it is determined that the point belongs to the current clustering center, and vice versa. After each iteration, it is necessary to calculate the error between the result and the result of the previous iteration until the error converges. During the experiment, the clustering effect can reach a good effect after 10 iterations. Therefore, in the present invention, the number of iterations is fixed at 10.

[0182] S52: Constructing Graph structure data based on superpixel clustering for suspended sonar targets:

[0183] After SLIC superpixel clustering is completed, each pixel point has a corresponding superpixel center. Therefore, the feature information (average metric of all pixel points within the region) of the final superpixel points can be obtained according to the following formula [l, a, b, x, y]: T :

[0184]

[0185] In the formula, N i represents the number of pixels contained in the i-th superpixel center. After obtaining the superpixels and their feature information, it is necessary to convert them into graph structure data to achieve the association in the feature space of the bright area and the acoustic shadow area. When converting the data, all the superpixel points obtained from each image are used as nodes in the graph structure data, the feature information of the superpixel points is used as the node attributes in the graph structure data, and the Euclidean distance between two superpixel points is used as the edge and its attributes in the graph structure data.

[0186] Reference Figure 6 , 7, and 8 respectively show the effect diagrams after clustering the original multi-beam sonar images, the results after SLIC clustering of the images after pre-detection and shear stitching of the target bright area and the shadow area, and the results after SLIC segmentation of the images after pre-detection and shear of the high-brightness area. From the clustering results, it can be seen that the SLIC algorithm well preserves the edge of the target. Due to the large size of the original sonar image, a single superpixel can cover the target bright area. After the pre-detected image is cropped and matched, the distribution of the superpixel centers is exactly close to the edges of the target bright area and the shadow area, and the superpixel distribution in the background area is relatively uniform. In summary, for the original sonar image, its clustering effect is poor, most of the superpixels contain the background area, and the target information contained is less; for the image after pre-detection and cropping, most of the superpixels contain the information of the target area, and the background reverberation area and the target area in the multi-beam sonar image can be distinguished from the distance between the positions of two superpixel centers.

[0187] Furthermore, step 52 includes the following steps:

[0188] S521: Graph structure data representation for suspended sonar target images:

[0189] Graph is a non-Euclidean data structure, mainly including two major elements: nodes and edges, denoted as G = {V, E}, where V = {v1,... v M} represents the set of nodes, and E = {e1,... e P} represents the set of edges, and both nodes and edges can have their attribute characteristics. In the present invention, the entire graph is discriminated, so a label category needs to be defined for an entire graph structure, where the categories include 5 types: dummy, tire, cylinder, sphere, and cube. The nodes in each graph structure are defined as a sequence of superpixel centers obtained by SLIC clustering, and the information contained in each node is the feature information of the superpixel points (pixel mean feature sp-intensity and center position feature sp-coord):

[0190]

[0191] The edges between nodes are defined as the correlation between pairwise nodes, and the information of the edges is the distance between superpixel points.

[0192] In the graph structure data, in order to eliminate redundancy, the connection between two nodes with very weak connectivity can be eliminated through the adjacency matrix. The adjacency matrix W ∈ R N×N is a two-dimensional array reflecting the correlation between pairwise nodes, and W i,j = w i,j represents the weight of the edge from node v i to v j . If W i,j = 0, it means that there is no edge between nodes v i and v j . In the present invention, for the edges between two nodes that are far apart, the value is assigned as 0, so as to eliminate the connection of the edges with weak correlation. The weights of the remaining edges are defined as the distance values obtained by jointly calculating the spatial position features and pixel features between pairwise superpixel points.

[0193] S522: Definition of the Graph structure attributes under the suspended sonar target image:

[0194] Using traditional deep learning methods to identify sonar targets is based on the Euclidean space, and feature extraction is performed on all pixels to varying degrees. However, due to the characteristics of sonar images such as low resolution, blurred edges, and unclear features, it is very difficult to achieve good recognition effects even with very excellent feature extraction methods. The graph neural network based on the non-Euclidean space adopted in the present invention can effectively utilize the spatial position features of the target bright area and shadow area obtained in the previous preprocessing steps. When constructing graph structure data, in addition to considering pixel value information, more attention is paid to the position relationship between each node, and the spatial position relationship can just reflect the spatial geometric features of the sonar target.

[0195] The pixel information and the spatial position information are both reflected in the attributes of the nodes in the graph structure. Therefore, it is necessary to define the attributes of the nodes. The pixel information is the pixel mean value of all pixel points within each super-pixel region, and the spatial position information is the central position coordinates of each super-pixel point. Since the SLIC super-pixel clustering result well preserves the target edge, the pixel information value of the super-pixel where the target bright area is located is large, the pixel information value of the super-pixel where the target shadow area is located is small, and the value of the background area is moderate, thus the target boundary can be effectively reflected. The nodes and their attributes are specifically represented as follows:

[0196]

[0197] Among them, (x i , y i ) represents the spatial position information, and f(x i , y i ) represents the pixel information.

[0198] In addition to the definition of node attributes, it is also necessary to define the attributes of edges. Edges reflect the degree of association between two nodes, and the attributes of edges can just be defined as the size of the association, that is, the weight of the edge. As mentioned above, in order to eliminate the redundancy of calculation, the edge weight between two nodes at a relatively long distance is assigned 0, and the specific elimination rule is determined by the K-Nearest Neighbor algorithm (KNN). Take the number of neighbors as κ, then the number of edges with edge weight values is κ×M = κM, and the remaining edge weights are 0. When calculating the association for the edges with weights, it is necessary to consider both the pixel distance and the spatial position distance between two nodes. The specific calculation formula is as follows:

[0199]

[0200] Among them, δ x represents the average value of the spatial position distances between node v i and its κ nearest nodes, and δ f represents the average value of the pixel distances between point v i and its κ nearest nodes. γ represents the relative ratio between the pixel information and the spatial position information. In summary, the attribute definitions in the graph structure data are summarized as follows:

[0201]

[0202] Furthermore, the step S6 includes the following steps:

[0203] S61: Construction of a graph attention network for underwater suspended target recognition:

[0204] After converting the image of the suspended target area in water into graph-structured data, it is necessary to construct a graph attention network to perform iterative regression on the graph structure. In the graph attention network model, an attention module is adopted to embed the nodes in the graph. By calculating the attention coefficients between the current node and its neighbor nodes, the neighbor information is aggregated, and the adaptive allocation of different neighbor weights is realized, so as to learn the neighborhood features and spatial features. The most core part of the network structure is the graph attention layer.

[0205] The input and output of the entire graph attention layer can be expressed as follows:

[0206]

[0207] In the formula is the input of the graph attention layer, representing the combination of the pixel features and spatial position features of each node; Q is the weight matrix, representing the linear transformation between the input and the output; σ is the non-linear activation function; α ij represents the attention coefficient between nodes i and j. Among them can be obtained after the construction of the graph-structured data is completed, the Q matrix is obtained through backpropagation, and σ is a pre-defined activation function. Then, among the parameters of the above formula, only the attention coefficient is unknown, and this coefficient needs to be calculated by the following formula:

[0208]

[0209] In the formula represents the attention mechanism, || represents the concatenation operation, and k represents the number of selected neighborhoods. Similarly, when calculating the coefficient, in order to remove the redundancy of the calculation and eliminate the noise generated due to the excessive introduction of redundant relationships between nodes, which reduces the performance of the model, it is necessary to jointly consider the edge weight matrix after the correlation adjustment, so as to obtain the final expression of the attention mechanism coefficient as:

[0210]

[0211] In the formula, W i,j is the updated adjacency matrix. For each node, the weights W i,j of the top k edges with strong connectivity to it are assigned 1, and the weights of the remaining edges are assigned negative infinity.

[0212] After the graph attention layer is defined, select an appropriate number of graph attention layers to build the entire graph attention network. In the network, the graph attention module in the joint attention layer and the adjacency matrix containing the edge weight coefficients are used to generate the attention coefficients between the graph nodes and their neighbor nodes, and the model can well adapt to the specific input of the network by adjusting the coefficient distribution, so as to complete the multi-beam suspended sonar target recognition under various types.

[0213] S62: Dataset construction:

[0214] The present invention constructs a graph attention network based on the DGL framework. Therefore, it is necessary to convert the graph structure data into the standard DGLGraph data format. First, the number of nodes in the graph structure is used as the node ID in the DGLGraph, and the pixel information and spatial position information of the nodes are used as the node attribute ndata. Secondly, the numbers between the two nodes connected by the edge are used as the edge ID pair, and the weight of the edge is used as the edge attribute edata. Finally, the label of the graph structure is used as the label of the DGLGraph, thus completing the construction of the data in the standard format.

[0215] After completing the conversion of the standard data format, it is necessary to divide all the collected valid data into data sets, which are divided into a training set, a validation set, and a test set according to the ratio of 8:0.5:1.51. The collected data contains a total of 5 types of suspended targets, that is, the labels of the graph structure data are also divided into 5 types, namely dummy, tire, cylinder, sphere, and cube. A total of 803 graph structures are created. After dividing according to the ratio of 8:0.5:1.51, the training set contains 646 graph structures, the validation set contains 43 graph structures, and the test set contains 114 graph structures.

[0216] S63: Experimental settings and model training:

[0217] Experimental settings: The present invention uses DGL as the graph neural network framework to complete the model construction under the PyTorch deep learning framework. The CPU model used is silver 4110 CPU@2.10Ghz, and the GPU model is NVIDIA GeForce RTX 3080. In the model, the number of graph attention layers is set to 4, the hidden layer is set to 19, the dimension of the output feature vector is set to 152, the number of attention mechanisms is set to 8, and the node aggregation method is set to average value. In the training parameters, the data set is set to the made suspended sonar target data set, the target category is set to 5, the epochs batch is set to 250, the batch-size is set to 4, the initial learning-rate is set to 0.001, and the decay coefficient of the learning-rate is set to 0.5.

[0218] Model training: After completing the modification of each parameter according to the experimental settings, the training set data made from the suspended target area will be sent into the graph attention network (GAT) for model training, and finally the network parameters with the optimal recognition effect will be obtained. The specific training steps are as follows:

[0219] (1) Data reading and preprocessing. After obtaining the original graph structure data, the K-nearest neighbor algorithm is used to eliminate the redundancy of the edge connections, thereby updating the weights of the edges and the adjacency matrix in the graph structure.

[0220] (2) Forward propagation. The graph-structured data is fed into the backbone network based on the graph attention mechanism to perform message passing using GAT convolution.

[0221] (3) Error calculation. The loss function is used to calculate the error between the predicted result and the true result obtained in the forward propagation.

[0222] (4) Iterative optimization and parameter update. By continuously iteratively optimizing the error, the forward propagation parameters of the model are continuously updated until the error converges and the parameters no longer change, and the recognition effect of the model reaches the optimal.

[0223] (5) Model prediction. The multi-beam sonar images in the validation set are fed into the trained model to complete the target recognition work.

[0224] Reference Figure 9 shows the training loss function and training accuracy under the dataset after pre-detection and shear stitching. From the perspective of the training loss function, the loss value finally approaches stability near 0.02, and the model has already converged near the iteration number 70, and the converged error is also very small, thus indicating the stability and superiority of the model. From the perspective of the training accuracy, the recognition accuracy finally converges to near 0.99, and the model also converges near the iteration number 70, which can also indicate that the model training effect is good.

[0225] S64: Model testing and result analysis:

[0226] In the testing stage, the training module in the network is turned off, the code for the validation and testing part is turned on, and the optimal model parameters during the training process are used to identify the categories of all the graph-structured data in the test set, so as to obtain the recognition accuracy under the test set, and the network model is tested with this indicator.

[0227] Reference Figure 10 shows the loss curve and recognition curve of the validation set and the recognition rate curve of the test set under the dataset after pre-detection and shear stitching. In the validation set, the loss function value converges approximately after 70 iterations, and the recognition accuracy has already converged approximately when the iteration number is 40, and the recognition accuracy of the validation set finally stabilizes near 0.9. In the test set, the recognition curve reaches convergence approximately when the iteration number is 30, and the recognition rate converges to near 0.97. Therefore, from the perspective of the recognition accuracy, the proposed suspension target recognition model for multi-beam sonar has a good convergence effect and recognition effect.

[0228] Furthermore, the step S7 includes the following steps:

[0229] S71: Verify the effectiveness of using regional pre-detection and the information of the shadow area of the suspended target:

[0230] To verify the effectiveness of regional pre-detection and the information of the shadow area of suspended targets, three types of graph structure data were produced. The first type is the graph structure data made from the most original multi-beam sonar images; the second type is the graph structure data obtained by first performing regional pre-detection on the original multi-beam sonar images, then only extracting the target bright area images, and finally converting the images of the highlighted areas into graph structure data; the third type is the graph structure data obtained by first performing regional pre-detection on the original images, then simultaneously cutting out the target highlighted area and the shadow area, longitudinally splicing the two areas, and finally converting the splicing result into graph structure data. After producing the three types of graph structure data, each type of data was used for model training, and the respective training losses and the recognition accuracies of the test sets were obtained. Finally, the effectiveness of regional pre-detection and the information of the shadow area of suspended targets was verified from the perspective of recognition accuracy.

[0231] Reference Figure 11 Shows a comparison graph of the training loss function values of three types of graph structure data under the verification of the effectiveness of regional pre-detection and the information of the shadow area of suspended targets. The figure shows that the dataset under the original sonar image converges the slowest, and the convergence effects of the images that only extract the target bright area images and the images that simultaneously cut out the target highlighted area and the shadow area after pre-detection are roughly the same. The final loss function values of the three types of datasets all converge to around 0.02. From the convergence effect of the loss function values, the suspended target recognition model proposed in the present invention has a good effect on model convergence.

[0232] Reference Figure 12 Shows a comparison graph of the recognition accuracies of three types of graph structure data under the verification of the effectiveness of regional pre-detection and the information of the shadow area of suspended targets. The figure shows the recognition accuracy curves obtained under the three types of datasets. Among them, the recognition effect of the graph structure data based on the original sonar image is the worst, followed by the images that only extract the target bright area images after pre-detection, and the best recognition effect is achieved by the images that simultaneously cut out the target highlighted area and the shadow area. The average recognition accuracy after the curve converges under the original sonar image is 0.913, and the maximum recognition accuracy is 0.921; the average recognition accuracy of the images that only extract the target bright area images after pre-detection is 0.921, and the maximum recognition accuracy is 0.930; the average recognition accuracy of the images that simultaneously cut out the target highlighted area and the shadow area is 0.962, and the maximum recognition accuracy is 0.974. Therefore, from various indicators of recognition accuracy, the effectiveness of the present invention in recognizing suspended targets by combining regional pre-detection and the information of the shadow area of suspended targets can be verified.

[0233] S72: Verify the influence of different SLIC clustering results on the recognition rate:

[0234] The coefficient τ of the relative importance between color similarity and spatial proximity is the most important factor affecting the clustering effect. The smaller the value of τ, the more the clustering effect tends to color features; the larger the value of τ, the more the clustering effect tends to spatial position features. In order to explore the impact of different clustering effects on sonar target recognition, the relative importance measurement parameter τ is set to 1, 5, 10, and 20 respectively. The SLIC clustering effect under different parameter values ​​is made into a graph structure data set, thereby obtaining 4 types of graph structure data. Then, the model is trained using various data sets to obtain the respective training losses and recognition accuracy of the test set. Finally, the recognition accuracy of the model is used to explain the relative importance distribution of spatial position features and pixel features.

[0235] refer to Figure 13 The comparison of the training loss function values ​​of four graph structure data under different SLIC clustering results is shown. The effect is the worst when the relative importance measurement coefficient τ is 1 in the figure, and the loss function converges to around 0.35. Secondly, when τ is 5, the loss function converges slowly. When τ is 10 and 20, the loss function curve converges the fastest, and the loss value converges to around 0.02. Therefore, from the perspective of model training, when τ is more than 10, the model can achieve a good convergence effect, that is, the clustering results under the main proportion of spatial location features can effectively improve the training effect of the model.

[0236] refer to Figure 14 The comparison chart of recognition accuracy of four kinds of graph structure data under different SLIC clustering results is shown. When the relative importance metric parameter τ takes the value of 1, the recognition effect of the test set is the worst, with an average recognition rate of 0.688 and a maximum recognition rate of 0.737; when τ takes the value of 5, the average recognition rate is 0.857 and the maximum recognition rate is 0.886; when τ takes the value of 10, the average recognition rate is 0.962 and the maximum recognition rate is 0.974; when τ takes the value of 20, the average recognition rate is 0.965 and the maximum recognition rate is 0.991; from the recognition rate index, the larger the τ value, the better the recognition effect of the model on the suspended target, which directly shows that when performing SLIC clustering, the spatial position feature accounts for a more important proportion, and when the super-pixel graph structure data is subsequently formed, the network is more inclined to extract and utilize the spatial position feature. This just verifies that the present invention proposes to use the spatial position relationship between the highlight area and the shadow area of ​​the sonar target to effectively improve the recognition rate of the suspended target in the water.

[0237] The preferred embodiments and principles of the present invention are described in detail above. For those skilled in the art, according to the ideas provided by the present invention, there may be changes in the specific implementation methods, and these changes should also be regarded as the protection scope of the present invention.

Claims

1. A method for identifying underwater suspended sonar targets based on regional pre-detection, characterized in that Including the following steps: S1: According to the multi-beam sonar imaging principle and imaging effect, perform coordinate transformation and image enhancement preprocessing on the original sonar image; S2: Based on the YOLOv5s network, pre-detect the bright areas and shadow areas of suspended targets simultaneously, and obtain the corresponding category information and target box position information; S3: Cut the bright areas and shadow areas respectively according to the category and target box position information; S4: Match and splice the bright area and shadow cut results belonging to the same target according to the imaging principle; S5: Use the SLIC superpixel algorithm to construct Graph structure data for the spliced result; S6: Construct a recognition model for suspended targets in water based on GAT (Graph Attention Network); S7: Ablation experiment setting, verify the effectiveness of using regional pre-detection and suspended target shadow area information, and verify the influence of different SLIC clustering results on the recognition rate.

2. The method for identifying underwater suspended sonar targets based on regional pre-detection according to claim 1, wherein The step S1 includes the following steps: S11: Multi-beam sonar image reconstruction (coordinate transformation): The multi-beam forward-looking sonar image is mainly presented in two ways: fan-shaped image and rectangular image; in order to facilitate the subsequent use of the target box position information to crop the area after marking and pre-detecting the bright area and shadow area of the target, the following coordinate transformation is used to convert the fan-shaped image into a conventional rectangular image during the preprocessing process: where φ and R represent the sonar horizontal opening angle and slant range size respectively, and W and H represent the horizontal and vertical dimensions of the image; S12: Multi-beam sonar image enhancement: The image enhancement includes the following steps: (1) Pixel enhancement: Multiply all pixel points in the sonar image by a certain multiple. Set the pixel value at each pixel position as f(x, y), and the magnification factor is Then the pixel value of each pixel point after pixel enhancement should be: (2) Median filtering: Select a square area (neighborhood) centered on a certain pixel, then sort the gray values of each pixel except the central pixel in the neighborhood, and use the sorted median as the new value of the central pixel point. When the neighborhood window slides orderly within the entire image range, the median filtering algorithm can be used to complete the filtering process for the entire image.

3. The method for identifying underwater suspended sonar targets based on regional pre-detection according to claim 1, wherein, The step S2 includes the following steps: S21: Construct a YOLOv5s network model and algorithm for detecting the bright areas and shadow areas of suspended targets: In order to better combine the features of the high-brightness area and shadow area of the same target, this part uses the YOLOv5s network to pre-detect the target high-brightness area and shadow area in the sonar image, and obtain the positions of the target bright area and shadow area in each multi-beam sonar image; in this step, the sonar target is not detected, only the areas with targets and the areas with shadow areas in the image are detected; S22: Collect sonar image data containing various suspended targets in water, perform data construction, and use the constructed dataset to train the model. A multi-beam sonar is mounted on a small fishing boat to search for targets in a lake with a water depth of 5 to 10 meters. The sonar is placed underwater 0.5 meters below the water surface, and the installation inclination angle of the sonar is 30° from the water surface. The targets are all suspended. After collecting field data, a total of 803 images containing suspended targets in water are sorted out, including 300 spheres, 204 cubes, 85 cylinders, 103 dummies, and 111 tire targets. Then, all the sample sets are marked with the target highlight area and the shadow area. After marking, the network parameters are modified, and the sample sets are sent into the network for model training to obtain the model parameters under the optimal detection effect. S23: Use the trained model to perform regional pre-detection on the multi-beam sonar image to be detected and obtain the position information of the target boxes in each region: Use the trained network model to detect the target bright area and shadow area of the image to be detected, and obtain the category information and target box information at the same time. The category information includes two types: bright area and dark area; the target box information includes four elements: [x center , y center , w, h], which respectively represent the normalized abscissa of the center point of the target box, the normalized ordinate of the center point of the target box, the normalized width of the target box, and the normalized height of the target box.

4. A method for identifying sonar targets suspended in water based on regional pre-detection according to claim 1, characterized in that, The step S3 includes the following steps: S31: Crop the bright area according to the position information of the target box in the bright area: For the target highlighted area, according to the target box information of the sonar target bright area obtained above, convert it into the pixel position coordinates of the upper left corner and the lower right corner of the target box that can represent the original image range where the area is located according to the conversion formula. After conversion, obtain the position area of each bright area target box in the original image, and then according to the obtained [x min ,y min ,x max ,y max coordinates, crop the original image to obtain all target bright area images; S32: Crop the shadow area according to the target box in the shadow area: Similar to the target bright area, for the target shadow area, according to the obtained target box information of the target shadow area, the position of the target box in the YOLO format is converted into the pixel position according to the pixel position conversion formula, and then the area is cropped according to the pixel points to finally obtain all the target shadow area images.

5. The method for identifying underwater suspended sonar targets based on regional pre-detection according to claim 1, wherein, The step S4 includes the following steps: S41: Match the bright area and the shadow belonging to the same target according to the imaging principle: According to the imaging principle of the multi-beam sonar, the acoustic wave beam is emitted forward, then reflected by the underwater medium, and finally the returned acoustic wave signal is received at the receiving transducer. In the fan-shaped sonar image, in the triangular area with the sonar transmitting transducer as the center point and the boundary of the target object as the side, both the target bright area and the shadow area are within this range, and the triangular boundary of the target object and the triangular boundary of the shadow area formed by the target object at a certain distance are the same boundary. Correspondingly, in the rectangular sonar image, the boundary of the target object and the boundary of the shadow area formed by the target object at the rear are also the same boundary. Therefore, according to this imaging principle, the high-light area and the shadow area belonging to the same sonar target in a sonar image are matched. S42: Stitch the region matching results: Since the suspended target is at a certain height from the bottom of the water, the shadow imaging will be at a certain distance outside the bright area imaging and will not be directly connected to the rear of the bright area. Therefore, the two regions of the bright area and the shadow area obtained by the pre-detection need to be physically connected. After cutting according to the region position information obtained by the pre-detection, only the regions containing target information remain, and the redundant background regions have been removed. After matching, the high-light area and the shadow area belonging to the same target have been obtained. To be closer to the imaging effect of the multi-beam sonar, inspired by the imaging of the underwater target in the sonar image, the high-light area and the shadow area belonging to the same target are vertically stitched, with the bright area at the bottom and the shadow at the top, thus completing the preprocessing of the entire sonar target.

6. The method for identifying underwater suspended sonar targets based on regional pre-detection according to claim 1, wherein, The step S5 includes the following steps: S51: Use the SLIC clustering algorithm to perform superpixel segmentation on the stitching result of the suspended sonar target: In order to better associate the bright area features and shadow area features of the target, the SLIC clustering algorithm is used to convert the stitching result of the suspended sonar target into graph-structured data, which is beneficial to the extraction of the target's spatial geometric features. First, the clustering centers are initialized, and then the positions of the clustering centers are updated within the area range of the initial clustering centers. After the positions of the superpixel clustering centers are updated, label assignment is performed on all pixel points within the area. Then, according to the distance similarity measurement rule, the attribution domains of all pixel points within the superpixel region are iteratively updated until the error converges, thus completing the superpixel segmentation. S52: Construct graph-structured data for the suspended sonar target based on superpixel clustering: After the SLIC superpixel clustering is completed, the position information and feature information of each superpixel point can be obtained according to the following formula. After obtaining the superpixels and their feature information, it is necessary to convert them into graph-structured data to realize the association of the bright area and shadow area features in space. When converting the data, all the superpixel points obtained from each image are used as nodes in the graph-structured data, the feature information of the superpixel points is used as the node attributes in the graph-structured data, and the Euclidean distance between two superpixels is used as the edge and its attributes in the graph-structured data.

7. A method for identifying sonar targets suspended in water based on regional pre-detection according to claim 1, characterized in that, The step S6 includes the following steps: S61: Construct a graph attention network for underwater suspended target recognition: After converting the underwater suspended target area image into graph-structured data, it is necessary to construct a graph attention network to perform iterative regression on the graph structure. In the network model, the attention mechanism is used to calculate the attention coefficients between the current node and its neighbors, so as to realize the allocation of the importance of different neighbors. When performing aggregation calculation with neighbor nodes, the goal of learning pixel features and neighborhood spatial position features is achieved. In addition, in order to eliminate redundancy, it is necessary to jointly consider the edge weight matrix after the relevance adjustment. S62: Dataset construction: The number of nodes in the graph structure is used as the node ID in DGLGraph, the pixel information and spatial position information of the nodes are used as node attributes. Secondly, the numbers of the two nodes connected by the edge are used as the edge ID pair, and the weight of the edge is used as the edge attribute. Finally, the label of the graph structure is used as the label of DGLGraph, thus completing the construction of the standard format data. Then, the constructed standard data is divided into a dataset, which is divided into a training set, a validation set and a test set according to the ratio of 8:0.5:1.

51. Among them, the graph structure data labels are divided into 5 categories, namely dummy, tire, cylinder, sphere and cube, and a total of 803 graph structures are created. S63: Experiment settings and model training: Experiment settings: Select appropriate GPU and CPU models, and set the parameters of the graph attention layer, hidden layer, output feature vector dimension, number of attention mechanisms and node aggregation method respectively. Set the dataset type, number of target categories, number of epochs, batch-size, initial learning rate and decay coefficient of the learning rate. Model training: First, data is read and preprocessed. Then, the preprocessed graph structure data is fed into the backbone network based on the graph attention mechanism to perform message passing using GAT convolution. The residual between the predicted result and the true result is obtained according to the error calculation formula. After obtaining the error, it is continuously iteratively optimized until the error converges. During the iteration process, the forward propagation parameters of the model are also updated accordingly. When the recognition effect of the model reaches the optimal, the model parameters under this effect are selected and replaced into the forward propagation of the network. Finally, the multi-beam sonar images in the validation set are fed into the model to complete the target recognition work. S64: Model testing and result analysis: In the testing stage, the training module in the network is turned off, and the code for the validation and testing part is enabled. The optimal model parameters during the training process are used to identify the categories of all the graph structure data in the test set, so as to obtain the recognition accuracy under the test set, and the network model is tested with this indicator.

8. A method for identifying underwater suspended sonar targets based on regional pre-detection according to claim 1, characterized in that, The step S7 includes the following steps: S71: Verify the effectiveness of using regional pre-detection and the information of the shadow area of suspended targets: To verify the effectiveness of regional pre-detection and the information of the shadow area of suspended targets, a total of 3 types of graph structure data are made. The first type is the graph structure data made from the most original multi-beam sonar images without any processing. The second type is to perform regional pre-detection on the original image, only extract the target bright area image, and then convert the image of the highlighted area into graph structure data. The third type is to perform regional pre-detection on the original image, and then cut out both the target highlighted area and the shadow area, and splice the two areas vertically. Finally, the splicing result is converted into graph structure data. After making the 3 types of graph structure data, each type of data is used for model training, and their respective training losses and the recognition accuracies of the test set are obtained. Finally, the effectiveness of regional pre-detection and the information of the shadow area of suspended targets is verified from the perspective of recognition accuracy. S72: Verify the influence of different SLIC clustering results on the recognition rate: The measurement coefficient τ of the relative importance between color similarity and spatial proximity is the most important factor affecting the SLIC clustering effect. Different values of τ resulting in different recognition effects can well illustrate the importance of the spatial position features and pixel features in the multi-beam sonar images respectively. To explore the influence of different clustering effects on the multi-beam sonar target recognition effect, the relative importance measurement parameter τ is respectively set to 1, 5, 10, 20, and made into graph structure data sets, so as to obtain 4 types of graph structure data. Then, each type of data set is used to train the model, and their respective training losses and the recognition accuracies of the test set are obtained. Finally, the relative importance distribution of the spatial position features and pixel features is explained using the recognition accuracy of the model.

Citation Information

Patent Citations

  • Hyperspectral image classification method based on graph neural network

    CN111695636A

  • Deep convolutional neural network-based submerged oil sonar detection image recognition method

    WO2021243743A1