Wall recognition method, system and equipment of sonar image and storage medium

The pre-trained anomaly detection model performs pixel-level classification of sonar image data, which solves the problem of difficulty in accurately identifying complex wall shapes in the prior art, and achieves higher wall recognition accuracy.

CN120219939AActive Publication Date: 2025-06-27ZHEJIANG OCEAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510487663.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-06-27
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

The existing wall recognition method based on CNN uses bounding box markings, making it difficult to accurately present the outlines and details of complex walls, resulting in low accuracy of wall recognition.

Method used

The pre-trained anomaly detection model is used to predict the wall classification of each pixel point of the target sonar image data, and train it through an anomaly detection network based on a convolutional neural network, using the sonar image sample data and the target label matrix as the training data.

Benefits of technology

Through pixel-level classification, the wall can be more accurately identified, overcome the limitations of the traditional frame marking method, carefully outline the outline of the wall, accurately display various details of the wall, and improve the accuracy of wall recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219939A_ABST
    Figure CN120219939A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of sonar, and discloses a wall recognition method, system and device for a sonar image and a storage medium, and the method comprises the steps: inputting target sonar image data into a pre-trained anomaly detection model, and carrying out the classification prediction of whether each pixel point is a wall or not, and obtaining a wall recognition result; the step of training the anomaly detection model comprises the steps of obtaining training data, adopting the training data to carry out classification prediction training of whether each pixel point is a wall or not on the initial model to obtain the anomaly detection model, and adopting an anomaly detection network based on a convolutional neural network as the initial model. The training data comprises sonar image sample data and a target label matrix. According to the invention, the anomaly detection model directly identifies whether each pixel point is the classification of the wall or not, the wall can be identified more accurately through pixel-level classification, the contour of the wall can be drawn more meticulously, various details of the wall can be displayed accurately, and the accuracy of wall identification is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of sonar technology and artificial intelligence technology, and particularly to a method, system, device and storage medium for wall recognition in sonar images. Background Art

[0002] Sonar detection technology is widely used in underwater environmental monitoring and obstacle detection. It emits and receives reflected signals, and generates an image reflecting underwater objects based on the reflected signals. The change in gray value in the image can be used to identify various underwater targets. However, due to the complex underwater environment, sonar images are interfered by factors such as noise and blur, and the image quality is unstable during the wall recognition process, which brings difficulties to the positioning and recognition of walls. In recent years, convolutional neural networks (CNNs) have developed significantly in the field of image recognition and have been used for sonar image processing to detect walls. They can automatically extract image features to make the detection more accurate. However, currently, the wall recognition method based on CNN uses bounding boxes to mark the positions of walls. This method can only roughly identify the walls and is difficult to accurately present the contours and details of walls with complex shapes, resulting in low accuracy of the wall recognition results. Summary of the Invention

[0003] Based on this, in view of the technical problem that the existing CNN-based wall recognition method uses bounding boxes to mark the positions of walls, which can only roughly identify the walls and is difficult to accurately present the contours and details of walls with complex shapes, resulting in low accuracy of the wall recognition results, a method, system, device and storage medium for wall recognition in sonar images are proposed.

[0004] In a first aspect, a method for wall recognition in sonar images is provided. The method includes: Obtaining target sonar image data; Inputting the target sonar image data into a pre-trained anomaly detection model for classification prediction of whether each pixel is a wall to obtain a wall recognition result; Wherein, the training steps of the anomaly detection model include: obtaining training data, and using the training data to train an initial model for classification prediction of whether each pixel is a wall to obtain the anomaly detection model. The initial model uses an anomaly detection network based on a convolutional neural network. The training data includes sonar image sample data and a target label matrix.

[0005] Further, the initial model sequentially includes: an input layer, at least two feature extraction units, a feature fusion unit, a decoding unit, and a classification unit. Each of the feature extraction units is arranged in parallel; The feature extraction unit sequentially includes: a first convolutional layer, a first activation layer, a first batch normalization layer, a first max pooling layer, a second convolutional layer, a second activation layer, a second batch normalization layer, and a second max pooling layer; The decoding unit sequentially includes: a first transposed convolutional layer, a third activation layer, a third batch normalization layer, and a second transposed convolutional layer.

[0006] Further, the number of the feature extraction units is four; The sizes of the convolutional kernels of the first convolutional layer and the second convolutional layer of the first feature extraction unit are both 3×3; The size of the convolutional kernel of the first convolutional layer of the second feature extraction unit is 7×7, and the size of the convolutional kernel of the second convolutional layer of the second feature extraction unit is 3×3; The size of the convolutional kernel of the first convolutional layer of the third feature extraction unit is 15×15, and the size of the convolutional kernel of the second convolutional layer of the third feature extraction unit is 5×5; The size of the convolutional kernel of the first convolutional layer of the fourth feature extraction unit is 25×25, and the size of the convolutional kernel of the second convolutional layer of the fourth feature extraction unit is 7×7; The sizes of the convolutional kernels of the first transposed convolutional layer and the second transposed convolutional layer are both 2×2.

[0007] Further, before the step of obtaining the training data, the following steps are further included: Obtain sonar echo signal sample data; Perform two-dimensional image matrix reshaping, normalization processing, jet color map processing, visualization, and obtain anomaly point markers based on the visualization on the sonar echo signal sample data to obtain a wall coordinate set corresponding to each wall; Generate a polynomial equation of multiple degrees according to each wall coordinate set; Generate a two-dimensional label matrix according to the sonar echo signal sample data and each polynomial equation of multiple degrees; Perform label expansion on the two-dimensional label matrix to obtain the target label matrix.

[0008] Further, the polynomial equation of multiple degrees adopts a quartic polynomial equation.

[0009] Further, the step of generating a two-dimensional label matrix according to the sonar echo signal sample data and each polynomial equation of multiple degrees includes: Traverse all data points in the sonar echo signal sample data by using a double-loop traversal method, and calculate the absolute value of the difference between the y-axis coordinate corresponding to each data point and each equation value corresponding to the data point to obtain the target distance, where the equation value is the value calculated by substituting the x-axis coordinate corresponding to the data point into the multiple polynomial equation; If the target distance corresponding to the data point is less than a preset threshold, then in the two-dimensional label matrix, set the label value corresponding to the data point to a first value; If the target distance corresponding to the data point is not less than the preset threshold, then in the two-dimensional label matrix, set the label value corresponding to the data point to a second value.

[0010] Further, after the step of obtaining the sonar echo signal sample data, the following steps are further included: Perform three-dimensional image reshaping on the sonar echo signal sample data to obtain sample data to be processed; Perform geometric transformation and Gamma correction on the sample data to be processed in sequence to obtain the sonar image sample data.

[0011] In a second aspect, a wall recognition system for sonar images is provided. The system includes: an electronic device configured to implement the steps of the wall recognition method for sonar images described in any one of the above.

[0012] In a third aspect, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the wall recognition method for sonar images described in any one of the above are implemented.

[0013] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the wall recognition method for sonar images described in any one of the above are implemented.

[0014] The wall recognition method, system, device and storage medium for sonar images of the present application obtain target sonar image data and input it into a pre-trained anomaly detection model to perform wall classification prediction on pixel points to obtain a wall recognition result. The anomaly detection model uses an anomaly detection network based on a convolutional neural network and is trained with training data including sonar image sample data and a target label matrix. Traditional box marking methods can only roughly identify walls and are difficult to accurately present the contours and details of walls with complex shapes. The anomaly detection model of the present application directly classifies each pixel point as to whether it is a wall. Through pixel-level classification, walls can be more accurately identified. This means that when facing complex underwater wall structures, the present application can overcome the limitations of traditional box marking methods, more precisely outline the contours of walls, accurately present various details of walls, and improve the accuracy of wall recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0016] Among them: Figure 1 It is an application environment diagram of the wall recognition method for sonar images in an embodiment; Figure 2 It is a flowchart of the wall recognition method for sonar images in an embodiment; Figure 3 It is a flowchart of the wall recognition method for sonar images in an embodiment; Figure 4 It is a flowchart of the wall recognition method for sonar images in an embodiment; Figure 5 It is a schematic diagram of the network structure of the initial model of the present application; Figure 6 It is a corresponding diagram of the sonar image sample data and the target label matrix of the present application; Figure 7 It is a block diagram of the structure of an electronic device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0018] The wall recognition method for sonar images provided by the embodiments of the present application can be applied in an application environment such as Figure 1 where the client 110 communicates with the server 120 through a network.

[0019] The server 120 is configured to implement the steps of the wall recognition method for sonar images of the present application: obtain target sonar image data through the client 110; input the target sonar image data into a pre-trained anomaly detection model to perform classification prediction on whether each pixel point is a wall to obtain a wall recognition result; wherein, the training steps of the anomaly detection model include: obtaining training data, and using the training data to train an initial model for classification prediction on whether each pixel point is a wall to obtain the anomaly detection model. The initial model uses an anomaly detection network based on a convolutional neural network, and the training data includes: sonar image sample data and a target label matrix. By obtaining target sonar image data and inputting it into a pre-trained anomaly detection model for pixel-level wall classification prediction to obtain a wall recognition result, where the anomaly detection model uses an anomaly detection network based on a convolutional neural network and is trained with training data including sonar image sample data and a target label matrix. The traditional box marking method can only roughly identify the wall, and it is difficult to accurately present the outline and details of complex-shaped walls. However, the anomaly detection model of the present application directly performs classification recognition on whether each pixel point is a wall. Through pixel-level classification, the wall can be more accurately recognized. This means that when facing complex underwater wall structures, the present application can overcome the limitations of the traditional box marking method, more carefully outline the wall contour, accurately display various details of the wall, and improve the accuracy of wall recognition.

[0020] Among them, the client 110 (that is, an electronic device) can be, but is not limited to, various personal computers, laptop computers, smartphones, tablet computers, and portable wearable devices. The server 120 (that is, an electronic device) can be implemented by an independent server or a server cluster composed of multiple servers. That is to say, both the client 110 and the server 120 can adopt computer devices.

[0021] The following describes the present application in detail through specific embodiments.

[0022] Please refer to Figure 2 as shown Figure 2A flowchart of a method for identifying walls in sonar images provided by an embodiment of the present application includes the following steps: S1: Obtain target sonar image data; First, use a sonar detection device to perform sonar scanning on the target underwater environment, and then reshape the sonar echo signal data obtained from the scanning into a three-dimensional image, and use the reshaped image as the target sonar image data.

[0023] The target underwater environment is the underwater environment where wall identification is desired.

[0024] The sonar echo signal data is a.dat file, and the sonar echo signal data includes the echo intensity of each position point (also called a data point) in the target underwater environment. That is to say, the sonar echo signal data is one-dimensional data.

[0025] It can be understood that the pixel points in the target sonar image data correspond one-to-one with the data points in the sonar echo signal data.

[0026] It can be understood that when reshaping the three-dimensional image, the content of the data is essentially not changed, but only the dimension of the data is increased. The specific implementation method of three-dimensional image reshaping can be selected from the prior art and will not be elaborated here.

[0027] Specifically, the target sonar image data can be obtained through a client according to user input, or the target sonar image data can be obtained from a preset storage space, or the target sonar image data sent by a third-party application can also be obtained.

[0028] S2: Input the target sonar image data into a pre-trained anomaly detection model to perform classification prediction on whether each pixel point is a wall, and obtain a wall identification result; Specifically, input the target sonar image data into a pre-trained anomaly detection model to perform classification prediction on whether each pixel point is a wall, and use the vector obtained from the classification prediction as the wall identification result.

[0029] Each vector value in the wall identification result is the probability value of being a wall.

[0030] It can be understood that the vector value at the i-th row and j-th column in the wall identification result corresponds to the same position point as the pixel value at the i-th row and j-th column in the target sonar image data, and both i and j are positive integers greater than 0.

[0031] Among them, the training steps of the anomaly detection model include: obtaining training data, and using the training data to train the initial model for classifying and predicting whether each pixel point is a wall, so as to obtain the anomaly detection model. The initial model uses an anomaly detection network based on a convolutional neural network. The training data includes: sonar image sample data and a target label matrix.

[0032] The anomaly detection model is a model for performing anomaly detection tasks. The anomaly detection task is a supervised learning task (that is, using the sonar image sample data as the input of the model and the target label matrix as the supervised target, and training the output of the initial model to get closer and closer to the target label matrix). It involves identifying data points that do not conform to the norm in the data (that is, the data points corresponding to the wall). The anomaly detection task can be divided into binary classification and multi-classification problems. Among them, the binary classification problem usually divides the data into an anomaly class (wall) and a normal class (the environment other than the wall). That is to say, using the training data to train the initial model for classifying and predicting whether each pixel point is a wall is the training of a classification model. The specific training method can be selected from the prior art and will not be elaborated here.

[0033] The target label matrix is a three-dimensional matrix. When the element value (also called the label value) in the target label matrix is the first value, the data point corresponding to the element value is a wall. When the element value in the target label matrix is the second value, the data point corresponding to the element value is not a wall.

[0034] The element value at the m-th row and k-th column in the target label matrix corresponds to the same position point as the pixel value at the m-th row and k-th column in the sonar image sample data, where both m and k are positive integers greater than 0.

[0035] It can be understood that using the training data to train the initial model for classifying and predicting whether each pixel point is a wall, when the training is completed, the initial model that has completed the training is used as the anomaly detection model.

[0036] In this embodiment, by obtaining the target sonar image data and inputting it into the pre-trained anomaly detection model for wall classification prediction of pixel points, the wall recognition result is obtained. The anomaly detection model uses an anomaly detection network based on a convolutional neural network and is trained by training data including sonar image sample data and a target label matrix. The traditional box marking method can only roughly identify the wall and it is difficult to accurately present the contour and details of the wall with complex shapes. However, the anomaly detection model of the present application directly performs classification recognition on whether each pixel point is a wall. Through pixel-level classification, the wall can be recognized more accurately. This means that when facing complex underwater wall structures, the present application can overcome the limitations of the traditional box marking method, more carefully outline the contour of the wall, accurately present various details of the wall, and improve the accuracy of wall recognition.

[0037] Please participate Figure 5 , in one embodiment, the initial model sequentially includes: an input layer, at least two feature extraction units, a feature fusion unit, a decoding unit, and a classification unit, and the feature extraction units are arranged in parallel; The feature extraction unit sequentially includes: a first convolutional layer, a first activation layer, a first batch normalization layer, a first max pooling layer, a second convolutional layer, a second activation layer, a second batch normalization layer, and a second max pooling layer; The decoding unit sequentially includes: a first transposed convolutional layer, a third activation layer, a third batch normalization layer, and a second transposed convolutional layer.

[0038] The input layer is used to receive the data input into the initial model, and then input the received data into each feature extraction unit.

[0039] Each of the feature extraction units constitutes an encoder. Each of the feature extraction units performs feature extraction at different scales.

[0040] The feature fusion unit is used to splice the features extracted by each of the feature extraction units in the channel dimension to form a feature map that fuses multi-scale features, and input the feature map into the decoding unit.

[0041] For example, the number of the feature extraction units is 4, and the feature fusion unit splices these 4 feature maps (each feature extraction unit outputs one feature map) in the dimension. The specific principle is: splicing in the channel dimension: in the NCHW (N: batch size; C: number of channels; H: height of the feature map; W: width of the feature map) format of PyTorch (PyTorch is an open-source deep learning framework for machine learning and deep learning): when torch.cat (a function for splicing tensors) splices in the channel dimension (dim = 1), it is necessary to ensure that the dimensions other than the number of channels are the same, otherwise they cannot be directly spliced. Only when their N, H / 4, and W / 4 dimensions are exactly the same can they be successfully cat (spliced) in the first dimension (channel dimension). The first feature extraction unit and the second feature extraction unit have 128 dimensions, and the third feature extraction unit and the fourth feature extraction unit have 64 dimensions. So after splicing, there are 384 channels (that is, the feature fusion unit inputs data with 384 channels). That is to say, the feature fusion unit realizes the fusion of parallel features.

[0042] In the fields of machine learning and deep learning, batch size refers to the number of data samples used to train a model in one iteration.

[0043] For the aggregation of parallel features, from the perspective of data structure, the concatenation operation will perform "horizontal expansion" in the channel direction, enabling each convolutional kernel in the subsequent network to see features from different branches. It can be understood that the network splits the image into multiple branches with different receptive fields at the front (each feature extraction unit serves as a branch), each branch goes through its own convolutional and pooling paths, and then "converges" at the same depth level before entering the decoding unit.

[0044] The decoding unit is used to decode according to the feature map input by the feature fusion unit and finally outputs a single-channel image.

[0045] The classification unit is a fully connected layer using the Sigmoid activation function. The classification unit is used to perform classification according to the input of the decoding unit.

[0046] The Sigmoid activation function is an S-shaped function commonly found in biology, also known as the S-shaped growth curve. In information science, due to its properties such as being monotonically increasing and having a monotonically increasing inverse function, the Sigmoid activation function is often used as the activation function of neural networks to map variables between 0 and 1.

[0047] Both the first convolutional layer and the second convolutional layer adopt convolutional layers.

[0048] Both the first activation layer, the second activation layer, and the third activation layer adopt the ReLU activation function.

[0049] The ReLU (Rectified Linear Unit) activation function is one of the most commonly used activation functions in convolutional neural networks (CNNs) and many deep learning models. Its main role is to introduce non-linearity, enabling the model to learn and express more complex features.

[0050] The features extracted by the feature extraction unit, after being processed by the ReLU activation function, can map linearly inseparable data to a new space, enabling the initial model to learn more complex patterns and relationships, thereby enhancing the non-linear expression ability of the model.

[0051] Both the first batch normalization layer, the second batch normalization layer, and the third batch normalization layer adopt batch normalization layers.

[0052] Both the first max pooling layer and the second max pooling layer adopt max pooling layers. During the operation of the initial model, the max pooling layer gradually reduces the spatial dimension of the feature map. This operation directly reduces the computational complexity first because the smaller size of the feature map means a smaller input scale for the subsequent layers, thus reducing the computational complexity of the connection weights between neurons. At the same time, this simplification operation on the feature map helps prevent overfitting. It removes some local details and noises in the data, avoiding the model from over-relying on special cases in the training data, and thus improving the generalization ability of the model. More importantly, the max pooling layer retains the most significant features when reducing the spatial dimension of the feature map. By selecting the maximum value in the local area, it retains the most representative feature information in each local area. These retained significant features are crucial for the subsequent processing layers to accurately analyze, recognize, or classify tasks, laying a foundation for the efficient and accurate operation of the entire model.

[0053] Both the first transposed convolution layer and the second transposed convolution layer adopt transposed convolution layers, which are also known as convolutional transpose layers. The spatial resolution of the image is gradually restored through the transposed convolution layer to ensure the restoration of image details.

[0054] Through this multi-branch (taking each feature extraction unit as a branch) design for multi-scale feature extraction, the multi-branch design can effectively fuse image features at different scales, enabling the initial model to learn from complex images and enhance detailed information, and finally output an enhanced image, making the wall more prominent and enhancing the detectability of the wall. The advantage of this design is that it can effectively cope with the challenges of complex backgrounds and multi-scale features, especially suitable for tasks such as sonar image sample data with a large amount of background noise and target variations.

[0055] Optionally, the sizes of the convolution kernels of each first convolution layer are different.

[0056] It can be understood that a first activation layer, a first batch normalization layer, and a first max pooling layer are connected after the first convolution layer of the feature extraction unit to gradually extract image features and reduce the spatial resolution.

[0057] Through the improvement of the model in this embodiment, the obtained initial model outlines the contour of the wall more meticulously and precisely shows various details of the wall. The initial model can maintain a high recognition accuracy in the application scenario of the underwater environment with more noise and poor image quality, and has stronger robustness than traditional methods, capable of coping with the challenges in complex underwater conditions.

[0058] In one embodiment, the number of the feature extraction units is four; The sizes of the convolution kernels of the first convolution layer and the second convolution layer of the first described feature extraction unit are both 3×3; The size of the convolution kernel of the first convolutional layer of the second feature extraction unit is 7×7, and the size of the convolution kernel of the second convolutional layer of the second feature extraction unit is 3×3; The size of the convolution kernel of the first convolutional layer of the third feature extraction unit is 15×15, and the size of the convolution kernel of the second convolutional layer of the third feature extraction unit is 5×5; The size of the convolution kernel of the first convolutional layer of the fourth feature extraction unit is 25×25, and the size of the convolution kernel of the second convolutional layer of the fourth feature extraction unit is 7×7; The size of the convolution kernels of the first deconvolutional layer and the second deconvolutional layer are both 2×2.

[0059] The first convolutional layer of the first feature extraction unit is used to extract the edges of the wall; the first convolutional layer of the second feature extraction unit is used to capture the local shape; the first convolutional layer of the third feature extraction unit is used to extract the contours and complex structures of the corresponding area of the wall; the first convolutional layer of the fourth feature extraction unit is used to obtain the overall shape and background information of the wall. That is to say, small-sized convolution kernels (3×3) extract low-level detailed features of the image, such as the edges of the wall; medium-sized convolution kernels (7×7) capture the local shape of the wall; medium-large-sized convolution kernels (15×15) extract the contours and object structures of the corresponding area of the wall; large-sized convolution kernels (25×25) extract global features, such as the general shape and background information of the wall. By integrating these features through the feature fusion unit, the initial model can comprehensively analyze the multi-scale information of the image, from the local details of the wall to the overall structure of the wall (such as the contour of the wall), to achieve accurate detection and recognition of the wall.

[0060] Please refer to Figure 3 As shown, in one embodiment, before the step of obtaining the training data, the following steps are further included: S31: Obtain sonar echo signal sample data; Use a sonar detection device to perform sonar scanning on the underwater environment, and use the obtained sonar echo signal data as sonar echo signal sample data.

[0061] Specifically, the sonar echo signal sample data input by the user can be obtained through the client, or the sonar echo signal sample data can be obtained from a preset storage space, or the sonar echo signal sample data sent by a third-party application can be obtained.

[0062] S32: Perform two-dimensional image matrix reshaping, normalization processing, jet color map processing, visualization, and obtain anomaly point markings based on the visualization on the sonar echo signal sample data to obtain a wall coordinate set corresponding to each wall; The jet color map is a commonly used color mapping in fields such as data visualization. The jet color map contains a continuous color change from blue to green, then to red, and finally to yellow. It usually maps the range of data values to this color range, with lower values corresponding to the blue end, higher values corresponding to the yellow end, and intermediate values distributed in the green and red regions. The jet color map is designed to provide an intuitive way to represent the range of data changes. By using different colors, the human eye can more easily distinguish different numerical ranges in the data, thus helping people quickly understand the data distribution and change trends.

[0063] Specifically, perform two-dimensional image matrix reshaping on the sonar echo signal sample data, normalize the reshaped data, and use the jet color map to enhance the visibility of the reflection intensity. Two-dimensional image matrix reshaping and normalization ensure that the data is adapted to subsequent processing while ensuring that the initial model can learn stably during training, avoiding the impact of size differences and range differences of the data input into the initial model on the performance of the initial model.

[0064] Perform two-dimensional image matrix reshaping and normalization on the sonar echo signal sample data to ensure stable operation of the data in subsequent steps. Through the matrix reshaping operation, the sonar echo signal sample data is transformed from a one-dimensional vector into a two-dimensional image matrix (for example, if the original data is 100 frames of data, then it has a total of 30,720,000 data points, and we reshape it into 100 * 600 * 512, that is, 100 two-dimensional image matrices) for subsequent image processing and feature extraction. Normalization ensures the scale consistency of the data and avoids the negative impact of data with different scales on the training of the initial model.

[0065] Display the data processed by the jet color map on the interface to achieve visualization.

[0066] For example, as Figure 6 shown, Figure 6 in the sub- Figure 6 a is the visualization effect diagram of the data input into the initial model, Figure 6 in the sub- Figure 6 b is the visualization effect diagram of the data processed by the jet color map on the interface to achieve visualization, Figure 6 Both a and 6b show the wall q1 and the wall q2.

[0067] The specific steps for obtaining anomaly point markings based on visualization for each wall include: According to the content displayed on the interface, when we think there is an anomaly in a certain area (i.e., the wall), we use the mouse for interactive marking. When the mouse is pressed, the starting position is recorded and the coordinate recording is initialized; when the mouse is dragged, the current mouse position is recorded and the coordinate list is updated; when the mouse is released, all the recorded coordinate points are saved. It should be noted that if the mouse is pressed again, the previously recorded coordinate list will be cleared. Finally, all the coordinate points will be saved to a text file, and the file format is simple, with each line recording a coordinate point (x, y). This text file is used as the wall coordinate set corresponding to this wall.

[0068] For example, if we think there are two areas with anomalies (i.e., there are walls) in the read image, then we need to first drag the mouse in one of the areas (without releasing during the process), and record as many coordinates as possible while ensuring accuracy. After completion, we need to modify the path of the saved file, use this file as a wall coordinate set, then restart the program, drag the mouse in the other area, and record as many coordinates as possible while ensuring accuracy. After completion, we need to modify the path of the saved file and use this file as a wall coordinate set.

[0069] S33: Generate a multiple polynomial equation according to each of the wall coordinate sets; Specifically, fit a multiple polynomial equation for each of the wall coordinate sets.

[0070] The multiple polynomial equation can be a quadratic polynomial equation, a cubic polynomial equation, or a quartic polynomial equation. That is to say, the multiple polynomial equation of this application is a polynomial equation of at least degree two.

[0071] Since the wall is relatively regular and its position is relatively fixed, the wall is represented by the multiple polynomial equation obtained by fitting.

[0072] S34: Generate a two-dimensional label matrix according to the sonar echo signal sample data and each of the multiple polynomial equations; Specifically, traverse all data points through the multiple polynomial equations of all walls to generate a two-dimensional label matrix, where the element value in the two-dimensional label matrix is the first value or the second value. The first value is an anomaly point (the data point corresponding to the wall), and the second value is a normal point (the data point corresponding to the background outside the wall).

[0073] S35: Perform label expansion on the two-dimensional label matrix to obtain the target label matrix.

[0074] Specifically, label expansion is used to adjust the dimensions of the tensor of the two-dimensional label matrix so that the resulting target label matrix after label expansion matches the shape of the data of the input initial pattern, which enables the initial model to make comparisons based on the images and label values (the element values in the two-dimensional label matrix) during the training process.

[0075] In this embodiment, the shape and edges of the wall are accurately displayed through color marking (i.e., jet color map processing), avoiding the defect that the traditional box marking method cannot depict the details of the wall in detail. This not only improves the recognition accuracy but also enhances the visualization effect of the results.

[0076] In one embodiment, the multiple polynomial equation uses a quartic polynomial equation.

[0077] For example, for Figure 6 the wall q1 in, the fitted quartic polynomial equation is: , For Figure 6 the wall q2 in, the fitted quartic polynomial equation is: , where x is the x-axis coordinate of the data point.

[0078] It can be understood that the degree in the quartic polynomial equation is the largest power of x.

[0079] Please refer to Figure 4 As shown, in one embodiment, the step of generating a two-dimensional label matrix according to the sonar echo signal sample data and each of the multiple polynomial equations includes: S341: Traverse all data points in the sonar echo signal sample data in a double-loop traversal manner, and calculate the absolute value of the difference between the y-axis coordinate of the corresponding position of each data point and the value of each equation corresponding to the data point to obtain the target distance, where the equation value is the value obtained by substituting the x-axis coordinate of the corresponding position of the data point into the multiple polynomial equation; Specifically, traverse all data points in the sonar echo signal sample data in a double-loop traversal manner, subtract the equation value obtained by substituting the x-axis coordinate of the corresponding position of the data point into the multiple polynomial equation from the y-axis coordinate of the corresponding position of the data point, and take the absolute value of the subtracted data as a target distance.

[0080] That is to say, each data point will calculate a target distance for each of the multiple polynomial equations.

[0081] S342: If the target distance of the data point is less than a preset threshold, then in the two-dimensional label matrix, set the label value corresponding to the data point to a first value; Specifically, if the target distance of the data point is less than the preset threshold, it means that the data point is a wall at this time. Therefore, in the two-dimensional label matrix, the label value corresponding to the data point is set to the first value.

[0082] S343: If the target distance of the data point is not less than the preset threshold, then in the two-dimensional label matrix, the label value corresponding to the data point is set to the second value.

[0083] Specifically, if the target distance of the data point is not less than the preset threshold, it means that the data point is not a wall at this time. Therefore, in the two-dimensional label matrix, the label value corresponding to the data point is set to the second value.

[0084] Optionally, the first value is set to 1 and the second value is set to 0.

[0085] In this embodiment, the target distance is determined by calculating the absolute value of the difference between the y-axis coordinate of the position corresponding to each data point and the equation value obtained by substituting into the polynomial equation multiple times, which can effectively screen and classify data points. When there is a data point with a target distance less than the preset threshold, its label value is set to the first value in the two-dimensional label matrix, which helps to accurately mark the data points that fit well with the polynomial equation, and may indicate that these data points conform to a certain expected characteristic or pattern. For the data points where the target distance is not less than the preset threshold, they are set to the second value, which can distinguish the data points that do not fit well with the polynomial equation, which helps to distinguish different types of data points in the analysis of sonar echo signal data, such as distinguishing the signals corresponding to the wall from the signals corresponding to the points outside the wall, thereby improving the accuracy of understanding, analyzing the sonar echo signal data and subsequent processing that may be based on this data.

[0086] Please refer to Figure 3 As shown, in one embodiment, after the step of obtaining the sonar echo signal sample data, the following steps are further included: S36: Perform three-dimensional image reshaping on the sonar echo signal sample data to obtain the sample data to be processed; Specifically, first, it is necessary to determine the structure and relevant parameters of the sonar echo signal sample data, such as the number of data points, the distribution in different dimensions, etc. Then, according to the preset rules, plan the spatial layout of these discrete data points. For the three dimensions of x, y, and z, determine the coordinate positions of each data point in the new three-dimensional space respectively. Mathematical methods such as interpolation can be used to fill the gaps between data points to ensure the integrity and coherence of the three-dimensional image. Next, according to the characteristics of the sonar echo signal sample data, assign corresponding color or grayscale values to each data point to represent the attribute of the echo intensity. Finally, through specialized graphics rendering techniques and tools, combine these processed data points to form a complete and intuitive three-dimensional image, which can visually present the important features of the sonar echo signal sample data (such as distribution, echo intensity).

[0087] S37: Perform geometric transformation and Gamma correction on the to-be-processed sample data in sequence to obtain the sonar image sample data.

[0088] Geometric transformation includes rotation, translation, scaling, and flipping. For example, rotate a three-dimensional image through a rotation matrix Rotate the wall edge by 30° and move it to the edge of the three-dimensional image to simulate different perspectives. Here, cos is the cosine function and sin is the sine function.

[0089] Gamma correction (the full English name is Gamma Correction) is an image processing technology that performs a non-linear transformation on the pixel values of an image to improve the display effect and visual quality of the image. Gamma correction further highlights the brightness difference between the wall and the background by adjusting the brightness curve, thereby enhancing the robustness and recognition ability of the initial model for complex scenes.

[0090] Specifically, first perform geometric transformation on the to-be-processed sample data, and then perform Gamma correction on the data obtained from the geometric transformation. Take the data obtained from Gamma correction as the sonar image sample data.

[0091] In this embodiment, by performing geometric transformation and Gamma correction on the to-be-processed sample data in sequence, the originally blurred wall becomes clear and bright after processing, and the originally strong noise is also suppressed after processing.

[0092] In one embodiment, a wall recognition system for a sonar image is provided. The system includes: an electronic device, and the electronic device is configured to implement the steps of the wall recognition method for a sonar image described in any one of the above.

[0093] In this embodiment, the target sonar image data is obtained and input into a pre-trained anomaly detection model for wall classification prediction of pixel points to obtain a wall recognition result. The anomaly detection model uses an anomaly detection network based on a convolutional neural network and is trained with training data including sonar image sample data and a target label matrix. The traditional box marking method can only roughly identify the wall and is difficult to accurately present the contour and details of the wall with complex shapes. However, the anomaly detection model of this application directly classifies each pixel point to identify whether it is a wall. Through pixel-level classification, the wall can be more accurately identified. This means that when facing a complex underwater wall structure, this application can overcome the limitations of the traditional box marking method, more precisely outline the contour of the wall, accurately present various details of the wall, and improve the accuracy of wall recognition.

[0094] In one embodiment, a computer device is provided. The computer device may be a server 120, and its internal structure diagram may be as Figure 7 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client 110 through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server 120 side of a method for wall recognition of sonar images.

[0095] In one embodiment, an electronic device is proposed, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are realized: Obtain target sonar image data; Input the target sonar image data into a pre-trained anomaly detection model for classification prediction of whether each pixel point is a wall to obtain a wall recognition result; Among them, the training steps of the anomaly detection model include: obtaining training data, and using the training data to train an initial model for classification prediction of whether each pixel point is a wall to obtain the anomaly detection model. The initial model uses an anomaly detection network based on a convolutional neural network. The training data includes sonar image sample data and a target label matrix.

[0096] In this embodiment, the target sonar image data is obtained and input into a pre-trained anomaly detection model for wall classification prediction of pixel points to obtain a wall recognition result. The anomaly detection model uses an anomaly detection network based on a convolutional neural network and is trained with training data including sonar image sample data and a target label matrix. The traditional bounding box marking method can only roughly identify the wall and it is difficult to accurately present the contour and details of the wall with a complex shape. However, the anomaly detection model of this application directly performs classification recognition on each pixel point to determine whether it is a wall. Through pixel-level classification, the wall can be more accurately recognized. This means that when facing a complex underwater wall structure, this application can overcome the limitations of the traditional bounding box marking method, more precisely outline the contour of the wall, accurately present various details of the wall, and improve the accuracy of wall recognition.

[0097] In one embodiment, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the following steps are implemented: Obtain target sonar image data; Input the target sonar image data into a pre-trained anomaly detection model for classification prediction of whether each pixel point is a wall to obtain a wall recognition result; Among them, the training steps of the anomaly detection model include: obtaining training data, and using the training data to train an initial model for classification prediction of whether each pixel point is a wall to obtain the anomaly detection model. The initial model uses an anomaly detection network based on a convolutional neural network, and the training data includes sonar image sample data and a target label matrix.

[0098] In this embodiment, the target sonar image data is obtained and input into a pre-trained anomaly detection model for wall classification prediction of pixel points to obtain a wall recognition result. The anomaly detection model uses an anomaly detection network based on a convolutional neural network and is trained with training data including sonar image sample data and a target label matrix. The traditional bounding box marking method can only roughly identify the wall and it is difficult to accurately present the contour and details of the wall with a complex shape. However, the anomaly detection model of this application directly performs classification recognition on each pixel point to determine whether it is a wall. Through pixel-level classification, the wall can be more accurately recognized. This means that when facing a complex underwater wall structure, this application can overcome the limitations of the traditional bounding box marking method, more precisely outline the contour of the wall, accurately present various details of the wall, and improve the accuracy of wall recognition.

[0099] It should be noted that for the functions or steps that the above computer-readable storage medium or electronic device can achieve, reference can be made to the relevant descriptions on the server 120 side and the client 110 side in the foregoing method embodiments. To avoid repetition, they will not be described in detail here.

[0100] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0101] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.

[0102] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A wall recognition method using sonar images, characterized in that: The method comprises: Acquire target sonar image data; Input the target sonar image data into a pre-trained anomaly detection model to classify and predict whether each pixel is a wall, and obtain a wall recognition result; Among them, the training step of the anomaly detection model includes: obtaining training data, using the training data to train the initial model for classification prediction of whether each pixel point is a wall, and obtaining the anomaly detection model, the initial model uses an anomaly detection network based on a convolutional neural network, and the training data includes: sonar image sample data and a target label matrix.

2. The wall recognition method of sonar images according to claim 1, characterized in that: The initial model includes in sequence: an input layer, at least two feature extraction units, a feature fusion unit, a decoding unit and a classification unit, and each of the feature extraction units is arranged in parallel; The feature extraction unit includes, in sequence: a first convolutional layer, a first activation layer, a first batch normalization layer, a first maximum pooling layer, a second convolutional layer, a second activation layer, a second batch normalization layer and a second maximum pooling layer; The decoding unit sequentially comprises: a first deconvolution layer, a third activation layer, a third batch normalization layer and a second deconvolution layer.

3. The wall recognition method of sonar images according to claim 2, characterized in that: The number of the feature extraction units is four; The sizes of the convolution kernels of the first convolution layer and the second convolution layer of the first feature extraction unit are both 3×3; The size of the convolution kernel of the first convolution layer of the second feature extraction unit is 7×7, and the size of the convolution kernel of the second convolution layer of the second feature extraction unit is 3×3; The size of the convolution kernel of the first convolution layer of the third feature extraction unit is 15×15, and the size of the convolution kernel of the second convolution layer of the third feature extraction unit is 5×5; The size of the convolution kernel of the first convolution layer of the fourth feature extraction unit is 25×25, and the size of the convolution kernel of the second convolution layer of the fourth feature extraction unit is 7×7; The sizes of the convolution kernels of the first deconvolution layer and the second deconvolution layer are both 2×2.

4. The wall recognition method of sonar images according to claim 1, characterized in that: Before the step of obtaining training data, the method further includes: Get sonar echo signal sample data; According to the sonar echo signal sample data, two-dimensional image matrix reshaping, normalization processing, jet color map processing, visualization and abnormal point marking based on visualization are performed to obtain a wall coordinate set corresponding to each wall; Generating a multi-order polynomial equation according to each of the wall coordinate sets; Generate a two-dimensional label matrix according to the sonar echo signal sample data and each of the multi-order polynomial equations; Perform label expansion on the two-dimensional label matrix to obtain the target label matrix.

5. The wall recognition method of sonar images according to claim 4, characterized in that: The multi-order polynomial equation adopts a quartic polynomial equation.

6. The wall recognition method of sonar images according to claim 4, characterized in that: The step of generating a two-dimensional label matrix according to the sonar echo signal sample data and each of the multi-order polynomial equations comprises: All data points in the sonar echo signal sample data are traversed in a double-layer loop traversal manner, and the absolute value of the difference between the y-axis coordinate of the position corresponding to each data point and each equation value corresponding to the data point is calculated to obtain the target distance, wherein the equation value is the value obtained by substituting the x-axis coordinate of the position corresponding to the data point into the multi-order polynomial equation; If the target distance between the data point and the target is less than a preset threshold, then in the two-dimensional label matrix, the label value corresponding to the data point is set to a first value; If the data point does not exist and the target distance is less than a preset threshold, then in the two-dimensional label matrix, the label value corresponding to the data point is set to a second value.

7. The wall recognition method of sonar images according to claim 4, characterized in that: After the step of obtaining the sonar echo signal sample data, the method further includes: Reconstructing the sonar echo signal sample data into a three-dimensional image to obtain sample data to be processed; The sample data to be processed are subjected to geometric transformation and Gamma correction in sequence to obtain the sonar image sample data.

8. A sonar image wall recognition system, characterized in that: The system comprises: an electronic device configured to implement the steps of the wall recognition method of sonar images as claimed in any one of claims 1 to 7.

9. An electronic device, characterized in that: The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the wall recognition method of the sonar image according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the wall recognition method of sonar images as claimed in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • A method for seabed bottom sonar image classification based on convolution neural network

    CN109086824A

  • Method for detecting cigarette packet seal defects based on RANSAC and CNN algorithms

    CN114140400A

  • Open set identification method and device for long-tail sonar image, medium and product

    CN118644769A

  • Target identification method based on sonar image segmentation and computing device

    CN119131567A

  • Image segmentation method and apparatus, and server

    WO2021115061A1