A target recognition method and computing device based on sonar image segmentation
By building a target recognition model for encoder and decoder, combined with initial training, fine-tuning training, sample augmentation and cost-sensitive training, the problem of low sonar image recognition accuracy in the existing technology is solved, and more accurate and efficient target recognition is achieved.
Patent Information
- Application Number
- CN202411121879.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-15
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-08-15
AI Technical Summary
The existing target recognition method based on sonar images has problems such as high false alarm rate and false detection rate and reduced model accuracy, especially when the amount of image data is small, the target category is unbalanced and the imaging quality is poor.
The target recognition method based on sonar image segmentation is adopted, and the target recognition model composed of encoder and decoder is built, and initial training and fine-tuning training are carried out, combining sample augmentation and cost-sensitive training to improve the recognition accuracy.
Accurate identification of targets is achieved, recognition accuracy and generalization capabilities are improved, and false alarm rates and false detection rates are reduced.
Smart Images

Figure CN119131567B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of sonar image processing, and relates to a target recognition method and a computing device, and in particular to a target recognition method and a computing device based on sonar image segmentation. Background Art
[0002] Existing target recognition methods based on sonar images mainly include two types: target recognition methods based on traditional image algorithms and target recognition methods based on deep learning. Among them, the target recognition method based on traditional image algorithms detects the target through image enhancement, dynamic threshold and other methods, and then identifies the target category through geometric features such as area, shape aspect ratio, target intensity, etc., and its false alarm rate and false detection rate are relatively high. The target recognition method based on deep learning is mainly implemented based on mainstream target detection frameworks such as RCNN and YOLO. When sonar images generally have problems such as small image data volume, unbalanced target categories, and large imaging quality differences caused by different sonar equipment, the accuracy of the trained model is somewhat lower than that of general visual target detection tasks, resulting in lower recognition accuracy.
[0003] Therefore, in view of the defects existing in the above-mentioned prior art, it is necessary to develop a new target recognition method based on sonar image segmentation. Summary of the invention
[0004] In order to overcome the defects of the prior art, the present invention proposes a target recognition method based on sonar image segmentation and a computing device, wherein the target recognition method based on sonar image segmentation can realize accurate recognition of the target.
[0005] In order to achieve the above object, the present invention provides a target recognition method based on sonar image segmentation, characterized in that it comprises the following steps:
[0006] 1) Obtain a dataset of existing sonar images;
[0007] 2) Build a target recognition model consisting of an encoder and a decoder;
[0008] 3) inputting the data set from the encoder into the target recognition model to perform initial training on the target recognition model;
[0009] 4) performing sample augmentation on the categories with fewer samples in the data set to obtain an augmented feature domain data set;
[0010] 5) using the augmented feature domain dataset to fine-tune the target recognition model;
[0011] 6) After fine-tuning and training the target recognition model, the newly acquired sonar image is input into the target recognition model to obtain a segmentation map of the target.
[0012] Preferably, in the step 2), the detection module RCN of the encoder of the target recognition model is jump-connected to the detection module RCN of the decoder, and the high-resolution features of the encoder are added to the corresponding resolution features of the decoder to help the decoder restore the feature resolution and output the segmentation map.
[0013] Preferably, in step 2), the image input to the encoder is processed by the first RC1 of the encoder and then input to the first RC2 of the encoder, and the first RC2 outputs the feature Figure 1 ; The image processed by the first RC2 is input into the first RC3 of the encoder, and the first RC3 outputs the feature Figure 2 The image processed by the first RC3 is input into the first RC5 of the encoder, and the feature map processed by the first RC5 is recorded as the feature Figure 3 ; The characteristics Figure 3 The second RC5 of the decoder is input, the first RC3 of the encoder is jump-connected to the second RC3 of the decoder, and the image processed by the second RC5 is the same as the feature output of the first RC3. Figure 2 Channel stitching is performed, and the stitched image is passed through the second RC3; the image processed by the second RC3 is combined with the feature Figure 1 Channel splicing is performed, and the spliced image passes through the second RC2 of the decoder; the image processed by the second RC2 passes through the second RC1 of the decoder again, and then the segmentation map is output.
[0014] Preferably, in step 3), the loss function used when initially training the target recognition model is:
[0015]
[0016] In the formula, N represents the number of categories in the data set; i represents the current category number; w i represents the loss weight of this category when the target recognition model is initially trained; y i represents the true probability of this category when the target recognition model is initially trained, and its value is 0 or 1; p i It represents the predicted probability of the pixel belonging to this category by the target recognition model when the target recognition model is initially trained.
[0017] Preferably, in the step 3), the degree of attention paid to the target category in the data set is greater than the degree of attention paid to the background category in the data set, and the number of pixels of different target categories in the data set also differs. The loss weight is designed based on the degree of attention paid to each category in the data set and the difference in the number of pixels of the target recognition model in the initial training, wherein the loss weight of the background category ≤ the loss weight of the target category with a large number of pixels ≤ the loss weight of the target category with a small number of pixels.
[0018] Preferably, in step 4), the following steps are further included:
[0019] 41) Organizing a feature domain data set: selecting a feature map output by the encoder as a feature domain data set;
[0020] 42) Random adjacent sample selection: Expand the feature graph of the category to be augmented into a one-dimensional feature vector, traverse each one-dimensional feature vector of the same category sample and calculate the distance between it and other one-dimensional feature vectors of the same category sample, take the five closest samples of the same category, and randomly select a one-dimensional feature vector from them as a random adjacent sample;
[0021] 43) Random interpolation resampling: randomly select a proportional coefficient in the range of [0,1], interpolate between the one-dimensional feature vector corresponding to the feature domain data set and the selected random adjacent samples to generate new samples, add the new samples to the feature domain data set, and obtain the augmented feature domain data set.
[0022] Preferably, in the step 5), the decoder is first separated from the encoder, the augmented feature domain data set is input into the decoder, and the decoder of the target recognition model is fine-tuned and trained. After the decoder completes the fine-tuning training, the fine-tuned decoder is communicatively connected to the encoder, the encoder loads the parameters during the initial training, and the decoder loads the parameters during the fine-tuning training, and the data set obtained in step 1) is used as input to fine-tune the target recognition model.
[0023] Preferably, in step 5), the loss function used when fine-tuning the target recognition model is:
[0024]
[0025] In the formula, N represents the number of categories in the data set; a represents the current category number; y a It indicates the true probability of this category when fine-tuning the target recognition model, and its value is 0 or 1; p a It represents the predicted probability of the pixel belonging to this category by the target recognition model when fine-tuning the target recognition model.
[0026] Preferably, in the step 1), each of the sonar images is subjected to Gaussian filtering to obtain an enhanced image; the enhanced image is annotated to generate a label map by performing image segmentation and annotation on the enhanced image; the label map is subjected to one-hot encoding processing, and after processing, the label is a 0-1 binary image with the same number of channels as the number of target categories, wherein each pixel with a channel value of 1 indicates that the true probability that the pixel position belongs to the category corresponding to the channel is 1, and each pixel with a channel value of 0 indicates that the pixel position does not belong to the channel, that is, the true probability of the pixel corresponding to the category is 0.
[0027] According to another aspect of the present invention, the present invention provides a computing device, the computing device comprising:
[0028] Processor; and
[0029] A memory, wherein computer program instructions are stored in the memory, and when the computer program instructions are executed by the processor, the processor executes the target recognition method based on sonar image segmentation.
[0030] Compared with the prior art, the target recognition method and computing device based on sonar image segmentation of the present invention have one or more of the following beneficial technical effects:
[0031] 1. The present invention uses an image segmentation model to realize target recognition, which can accurately detect pixel by pixel.
[0032] 2. The present invention uses cost-sensitive training, which has a better detection effect on targets with fewer pixels.
[0033] 3. The present invention resamples and augments the unbalanced training data set to achieve better target recognition results for categories with few samples.
[0034] 4. The present invention adds the high-resolution features of the encoder to the corresponding resolution features of the decoder through jump connections to help the decoder restore the feature resolution and output a segmentation map, which is beneficial to improving the target recognition effect.
[0035] 5. The present invention performs augmentation in the feature domain rather than on the image, so that the obtained augmented samples have a higher degree of abstraction and a wider feature dimension. In addition, based on the feature map form obtained by the encoder of the initially trained target recognition model, the augmented feature domain data set is more easily accepted by the decoder.
[0036] 6. In the process of fine-tuning and training the target recognition model of the present invention, the target recognition model is trained uniformly without emphasis, which is conducive to optimizing the final output result of the model and improving the accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a flow chart of the target recognition method based on sonar image segmentation of the present invention.
[0038] Figure 2 It is a schematic diagram of the structure of the residual structure detection module RCN of the present invention.
[0039] Figure 3 It is a schematic diagram of the composition of the target recognition model of the present invention. DETAILED DESCRIPTION
[0040] Before describing in detail any embodiment of the present invention, it should be understood that the present invention is not limited in its application to the construction and arrangement details of the components described below or illustrated in the following figures. The present invention can have other embodiments and can be practiced or carried out in various ways. In addition, it should be understood that the words and terms used here are for descriptive purposes and should not be considered restrictive. The use of "including" or "having" and its variations herein is intended to cover the items and their equivalents and additional items displayed below. Unless otherwise specified or limited, the terms "install", "connect", "support" and "couple" and their variations are widely used and cover direct installation and indirect installation, connection, support and connection. In addition, "connect" and "couple" are not limited to physical or mechanical connections or connections.
[0041] Furthermore, on the first aspect, in the disclosure of the present invention, the orientation or positional relationship indicated by terms such as "longitudinal", "transverse", "upper", "lower", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside" and "outside" are based on the orientation or positional relationship shown in the accompanying drawings, which are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore the above terms cannot be understood as limitations on the present invention; on the second aspect, the term "one" should be understood as "at least one" or "one or more", that is, in one embodiment, the number of an element may be one, while in another embodiment, the number of the element may be multiple, and the term "one" cannot be understood as a limitation on the quantity.
[0042] Figure 1 FIG. 2 shows a flow chart of a target recognition method based on sonar image segmentation of the present invention. Figure 1 As shown, the target recognition method based on sonar image segmentation of the present invention comprises the following steps:
[0043] 1. Obtain a dataset of existing sonar images.
[0044] In the present invention, the existing sonar images are first sorted out, and these sonar images are uniformly enhanced and numerically normalized to reduce the numerical differences caused by factors such as different sonar devices for collecting sonar images and different collection environments, and then the targets are image segmented and labeled.
[0045] In a specific example of the present invention, a total of 5,500 sonar images containing two types of targets, reefs and shipwrecks, are selected, and Gaussian filtering is performed on each of the sonar images to obtain an enhanced image, thereby suppressing noise signals.
[0046] Next, the enhanced image is normalized to control its numerical domain. The enhanced image after the above processing is annotated, and the enhanced image is segmented and annotated to generate a label map. The pixel values {0, 1, 2} in the label map represent the pixel as background, reef and shipwreck respectively.
[0047] Subsequently, in order to facilitate the output of the subsequent target recognition model and the calculation of the loss function, the label map is one-hot encoded, and the processed label is a 0-1 binary image with the same number of channels as the number of categories. For example, in a specific example of the present invention, the number of categories is 3 (which includes the background and two targets, i.e., reefs and shipwrecks), so the processed label is a 0-1 binary image with 3 channels. Furthermore, each pixel with a channel value of 1 indicates that the pixel is the category corresponding to its channel number, that is, the true probability that the pixel position belongs to the category corresponding to the channel is 1 (100%), and each pixel with a channel value of 0 indicates that the pixel position does not belong to the channel, that is, the true probability of the pixel corresponding to the category is 0.
[0048] Since the acquisition difficulty of reefs and shipwrecks is different, and the proportion of the two types of targets in the data set is different, in this specific example of the present invention, the selected data set contains 3624 sonar images of reefs, 1388 sonar images of shipwrecks, and the rest are sonar images without targets. It should be understood by those skilled in the art that the total number of images in the data set, the number of target images, and the proportion are only examples, and cannot be a limitation on the content and scope of the target recognition method described in the present invention. More or fewer images can also be collected to form a data set.
[0049] 2. Build a target recognition model consisting of an encoder and a decoder.
[0050] The target recognition model constructed by the present invention is a target recognition model of an encoder-decoder system, which includes an encoder and a decoder.
[0051] Wherein, the encoder and the decoder both include a plurality of residual structure detection modules RCN. Figure 2As shown, each detection module RCN of the residual structure is composed of two dilation rate N hollow convolution layers, a regularization layer and an activation function layer. N is 1, 2, 3 or 5. Preferably, the activation function layer is a ReLU activation function layer. In addition, the high-resolution features of the encoder are added to the corresponding resolution features of the decoder through a jump connection to help the decoder restore the feature resolution and output a segmentation map.
[0052] Specifically, Figure 3 As shown, the encoder includes a first RC1, a first RC2, a first RC3 and a first RC5. The decoder includes a second RC1, a second RC2, a second RC3 and a second RC5. The first RC1 and the second RC1 are both composed of two dilated convolution layers with a dilation rate of 1, a regularization layer and an activation function layer. The first RC2 and the second RC2 are both composed of two dilated convolution layers with a dilation rate of 2, a regularization layer and an activation function layer. The first RC3 and the second RC3 are both composed of two dilated convolution layers with a dilation rate of 3, a regularization layer and an activation function layer. The first RC5 and the second RC5 are both composed of two dilated convolution layers with a dilation rate of 5, a regularization layer and an activation function layer. That is, the encoder and the decoder have the same number of layers, the corresponding number of layers have the same detection module RCN dilation rate, and the dilation rate of the detection module RCN is the order of the detection module RCN.
[0053] The first RC1 is communicatively connected to the first RC2, the image input to the encoder is processed by the first RC1 and then input to the first RC2, and the first RC2 can output a feature Figure 1 The first RC2 is communicatively connected to the first RC3, and the image processed by the first RC2 is input to the first RC3, and the first RC3 can output feature Figure 2 The first RC3 is communicatively connected to the first RC5, the image processed by the first RC3 is input to the first RC5, and the feature map processed by the first RC5 is recorded as a feature map. Figure 3 .
[0054] The first RC5 of the encoder is communicatively connected to the second RC5 of the decoder, and the image (i.e., the feature Figure 3 ) inputs the second RC5 of the decoder. The second RC3 of the decoder is communicatively connected to the second RC5 of the decoder, the first RC3 of the encoder is jump-connected to the second RC3, and the image processed by the second RC5 is the same as the feature output of the first RC3 Figure 2Channel splicing is performed, and the spliced image passes through the second RC3. The second RC2 of the decoder is communicatively connected to the second RC3 of the decoder, and the first RC2 of the encoder is jump-connected to the second RC2 of the decoder. The image processed by the second RC3 is connected to the feature Figure 1 Channel splicing is performed, and the spliced image passes through the second RC2 of the decoder. The second RC2 is communicatively connected to the second RC1, and the image processed by the second RC2 module passes through the second RC1 again, and then the segmentation map is output. In this way, the high-resolution features of the encoder are added to the corresponding resolution features of the decoder through the skip connection to help the decoder restore the feature resolution and output the segmentation map.
[0055] In a specific example of the present invention, the input of the encoder is a single-channel image of [512×512×1], the single-channel image of [512×512×1] passes through the first RC1 of the encoder and increases the feature map channel to [512×512×8], the feature map of [512×512×8] passes through the first RC2 and increases the feature map channel to [512×512×16], and the output of the first RC2 is retained as the feature Figure 1 Then the [512×512×16] feature map continues to pass through the first RC3 and increases the feature map channel to [512×512×32], retaining the output of the first RC3 as the feature Figure 2 Then the [512×512×32] feature map passes through the first RC5 and increases the feature map channel to [512×512×64], and uses the [512×512×64] feature map as the feature Figure 3 . Next feature Figure 3 Through the decoder and reduce the feature map channel to [512×512×32], then the [512×512×32] feature map is concatenated Figure 2 , get the feature map of [512×512×64], the feature map of [512×512×64] passes through the second RC3 of the decoder, and reduces the feature map channel to [512×512×16]. Then, concatenate the feature map of [512×512×16] Figure 1 The feature map of [512×512×32] is obtained, and the feature map of [512×512×32] is passed through the second RC2 of the decoder and the feature map channel is reduced to [512×512×8], and the feature map of [512×512×8] is passed through the second RC1 and the feature channel is reduced to [512×512×N], and the segmentation map is output. At this time, N represents the number of categories detected by the target recognition model. In the present invention, N=3 is taken, which includes two targets (reefs and shipwrecks) and background.
[0056] 3. Input the data set from the encoder into the target recognition model to perform initial training on the target recognition model.
[0057] After the data set is obtained and the target recognition model is built, the data set can be input into the target recognition model to perform initial training on the target recognition model.
[0058] In the present invention, when the target recognition model is initially trained, the model parameters of the target recognition model are adjusted by gradient descent, and cost-sensitive training is used to make the loss function focus on the segmentation accuracy of non-background pixels, so that the target recognition model can effectively perform target segmentation on sonar images.
[0059] Specifically, the loss function used when initially training the target recognition model is:
[0060]
[0061] Wherein, N represents the number of categories in the data set, and in an example of the present invention, N is 3; i represents the current category number; w i represents the loss weight of this category when the target recognition model is initially trained; y i represents the true probability of this category when the target recognition model is initially trained, and its value is 0 or 1. The true probability of this category is the pixel value at the corresponding position of the label map generated in step 1; p i It indicates that when the target recognition model is initially trained, the target recognition model predicts the probability that the pixel belongs to the category, and the target recognition model predicts the probability that the pixel belongs to the category is the output value of the target recognition model.
[0062] Further, in step three of the present invention, the degree of attention paid to the target category in the data set is greater than the degree of attention paid to the background category in the data set, and there are also differences in the number of pixels of different target categories in the data set. The category with fewer pixels has a lower impact on the loss function value in the training process of the target recognition model, and its learning difficulty is greater, so it is necessary to set a higher loss weight for it to improve the sensitivity of the target recognition model to this type. Therefore, the loss weight is designed according to the degree of attention paid to each category in the data set and the difference in the number of pixels of the target recognition model in the initial training. Preferably, the loss weight of the background category ≤ the loss weight of the target category with a large number of pixels ≤ the loss weight of the target category with a small number of pixels.
[0063] In a specific example of the present invention, the categories in the sonar image include background, reefs and shipwrecks. The data set contains 3624 sonar images of reefs, 1388 sonar images of shipwrecks, and the rest are sonar images without targets. Among them, the average pixel area of the reef is about 12 pixels, and the average pixel area of the shipwreck is about 10,000 pixels, that is, the number of pixels of the reef is 43,488, and the number of pixels of the shipwreck is 138,800,000. When the target recognition model is initially trained, the loss weight w1 of the background is set ≤ the loss weight w3 of the shipwreck ≤ the loss weight w2 of the reef. For example, but not limited to, the loss weights {w1, w2, w3} of the background, reef and shipwreck are set to {0.2, 0.5, 0.3} respectively. At this time, through training, the target recognition model is more sensitive to the accuracy of reef and shipwreck resolution, and is slightly inclined to the reef. It is worth mentioning that the numerical values such as the number of sonar images, pixels and specific loss weights are only used as examples and do not limit the content and scope of the target recognition method of the present invention.
[0064] Fourth, perform sample augmentation on categories with fewer samples in the data set to obtain an augmented feature domain data set.
[0065] Since there are fewer samples in the data set, the target recognition model is prone to overfitting on this category, resulting in insufficient learning of the target recognition model for this category, and the detection accuracy is reduced, which is lower than that of other categories. Therefore, it is necessary to perform sample augmentation on this category to improve the overall accuracy and generalization ability of the target recognition model. In an exemplary example of the present invention, the number of sonar images containing the shipwreck category is small, so it is necessary to perform sample augmentation on the sonar images of the shipwreck category.
[0066] Specifically, the category distribution of the data set is analyzed, the feature maps at each level output by the encoder are used as feature domains, and the feature maps at each level are organized into feature domain data sets, the feature domains of the few-sample categories are resampled, the number of samples of the few-sample categories is increased proportionally, and the new samples obtained after resampling are added to the feature domain data set.
[0067] Among them, resampling the feature domain of the few-sample category specifically includes:
[0068] 1. Organize feature domain datasets.
[0069] The output of the encoder is selected as the feature domain data set, that is, the feature maps corresponding to the existing data set samples are organized into the feature domain data set. That is, in this specific embodiment of the present invention, each sample in the feature domain data set contains the feature Figure 1 ,feature Figure 2 and Features Figure 3 .
[0070] 2. Random adjacent sample selection.
[0071] Each sample in the feature domain dataset is a multidimensional vector. First, the feature map of the category that needs to be augmented, such as the feature map of a shipwreck, is Figure 1 ,feature Figure 2 and Features Figure 3 Each of them is expanded into a one-dimensional feature vector, and each one-dimensional feature vector of the same category samples is traversed and the distance between it and other one-dimensional feature vectors of the same category samples is calculated. The five closest samples of the same category are taken, and a one-dimensional feature vector is randomly selected from them as a random adjacent sample.
[0072] 3. Random interpolation resampling.
[0073] An interpolation coefficient is randomly selected in the range of [0,1], and a new sample is generated by interpolating between the one-dimensional feature vector corresponding to the feature domain data set and the selected random adjacent sample. The new sample is added to the feature domain data set to obtain the augmented feature domain data set.
[0074] Preferably, the categories with fewer samples in the data set are augmented until the number of samples in each category is close or consistent, so that the number of samples in each category is balanced. This helps to avoid the overfitting phenomenon of the target recognition model on the category with fewer samples, resulting in insufficient learning of the target recognition model for the category, thereby improving the overall accuracy and generalization ability of the target recognition model.
[0075] The present invention performs augmentation in the feature domain rather than on the image, and the obtained augmented samples have a higher degree of abstraction and a wider feature dimension. In addition, based on the feature map form obtained by the encoder of the initially trained target recognition model, the augmented feature domain data set is more easily accepted by the decoder.
[0076] 5. Use the augmented feature domain dataset to fine-tune the target recognition model.
[0077] First, the decoder is separated from the encoder, the augmented feature domain data set is input into the decoder, and the decoder of the target recognition model is fine-tuned and trained.
[0078] In the present invention, when fine-tuning the decoder of the target recognition model, cost-sensitive training is also used to make the loss function focus on the segmentation accuracy of non-background pixels, and the loss function used in fine-tuning training is:
[0079]
[0080] In the formula, N represents the number of categories in the data set; a represents the current category number; y aIt indicates the true probability of this category when fine-tuning the target recognition model, and its value is 0 or 1; p a It represents the predicted probability of the pixel belonging to this category by the target recognition model when fine-tuning the target recognition model.
[0081] In the process of fine-tuning and training the target recognition model, the target recognition model is trained evenly without emphasis, which is conducive to optimizing the final output result of the model and improving the accuracy of the model.
[0082] Furthermore, the decoder and the encoder after fine-tuning are communicatively connected, the encoder loads the parameters of the initial training, and the decoder loads the parameters of the fine-tuning training, and the data set obtained in step 1 is used as input to fine-tune the entire target recognition model. After sample augmentation and fine-tuning training, the target recognition model has higher detection accuracy on the few sample categories and stronger model generalization.
[0083] 6. Target identification.
[0084] The parameters of the target recognition model after fine-tuning the training in step five are loaded, and the newly obtained sonar image is input into the target recognition model to obtain a segmentation map of the target.
[0085] Before the newly acquired sonar image is input into the target recognition model, the sonar image is firstly subjected to image enhancement (eg, Gaussian filtering) and numerical normalization operations so as to facilitate better recognition by the target recognition model.
[0086] The target recognition method based on sonar image segmentation of the present invention is implemented based on the target recognition model, and can accurately detect pixel by pixel. Moreover, the present invention uses cost-sensitive training, which has a better detection effect on targets with fewer pixels. In addition, the present invention performs resampling data enhancement on the unbalanced training data set, so that the target recognition effect of the small sample category is better.
[0087] According to another aspect of the present invention, the present invention further provides a computing device, wherein the computing device includes a processor and a memory, wherein computer program instructions are stored in the memory, and when the computer program instructions are executed by the processor, the processor executes the target recognition method based on sonar image segmentation described in the present invention.
[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention, rather than to limit the scope of protection of the present invention. Those skilled in the art can modify or replace the technical solution of the present invention according to the idea of the present invention without departing from the essence and scope of the technical solution of the present invention.
Claims
1. A target recognition method based on sonar image segmentation, characterized in that: The following steps are involved: 1) Obtain a dataset of existing sonar images; 2) Building a target recognition model composed of an encoder and a decoder, wherein the encoder and the decoder both include a plurality of residual structure detection modules RCN, each of which is composed of two hole convolution layers with an expansion rate of N, a regularization layer, and an activation function layer, where N is 1, 2, 3, or 5, wherein an image input to the encoder is processed by the first RC1 of the encoder and then input to the first RC2 of the encoder, and the first RC2 outputs a feature map 1; the image processed by the first RC2 is input to the first RC3 of the encoder, and the first RC3 outputs a feature map 2; the image processed by the first RC3 is input to the The first RC5 of the encoder, and the feature map processed by the first RC5 is recorded as feature map 3; the feature map 3 is input into the second RC5 of the decoder, the first RC3 of the encoder is jump-connected to the second RC3 of the decoder, the image processed by the second RC5 is channel-joined with the feature map 2 output by the first RC3, and the joined image passes through the second RC3; the image processed by the second RC3 is channel-joined with the feature map 1, and the joined image passes through the second RC2 of the decoder; the image processed by the second RC2 passes through the second RC1 of the decoder, and then the segmentation map is output; 3) inputting the data set from the encoder into the target recognition model to perform initial training on the target recognition model; 4) performing sample augmentation on the categories with fewer samples in the data set to obtain an augmented feature domain data set; 5) Using the augmented feature domain data set, fine-tuning the target recognition model, wherein the decoder is first separated from the encoder, the augmented feature domain data set is input into the decoder, and the decoder of the target recognition model is fine-tuned. After the decoder completes the fine-tuning training, the fine-tuned decoder is communicatively connected to the encoder, the encoder loads the parameters of the initial training, and the decoder loads the parameters of the fine-tuning training, and the data set obtained in step 1) is used as input to fine-tune the target recognition model; 6) After fine-tuning and training the target recognition model, the newly acquired sonar image is input into the target recognition model to obtain a segmentation map of the target.
2. The target recognition method based on sonar image segmentation according to claim 1, characterized in that: In step 3), the loss function used when initially training the target recognition model is: In the formula, N represents the number of categories in the data set; i represents the current category number; w i represents the loss weight of category i when the target recognition model is initially trained; y i represents the true probability of category i when the target recognition model is initially trained, and its value is 0 or 1; p i It represents the predicted probability of the pixel belonging to category i by the target recognition model when the target recognition model is initially trained.
3. The target recognition method based on sonar image segmentation according to claim 2 is characterized in that: In the step 3), the degree of attention paid to the target category in the data set is greater than the degree of attention paid to the background category in the data set, and the number of pixels of different target categories in the data set also varies. The loss weight is designed based on the degree of attention paid to each category in the data set and the difference in the number of pixels of the target recognition model in the initial training, wherein the loss weight of the background category is ≤ the loss weight of the target category with a large number of pixels ≤ the loss weight of the target category with a small number of pixels.
4. The target recognition method based on sonar image segmentation according to claim 1, characterized in that: In the step 4), the following steps are further included: 41) Organizing a feature domain data set: selecting a feature map output by the encoder as a feature domain data set; 42) Random adjacent sample selection: Expand the feature graph of the category to be augmented into a one-dimensional feature vector, traverse each one-dimensional feature vector of the same category sample and calculate the distance between it and other one-dimensional feature vectors of the same category sample, take the five closest samples of the same category, and randomly select a one-dimensional feature vector from them as a random adjacent sample; 43) Random interpolation resampling: randomly select a proportional coefficient in the range of [0,1], interpolate between the one-dimensional feature vector corresponding to the feature domain data set and the selected random adjacent samples to generate new samples, add the new samples to the feature domain data set, and obtain the augmented feature domain data set.
5. The target recognition method based on sonar image segmentation according to claim 4 is characterized in that: In step 5), the loss function used when fine-tuning the target recognition model is: In the formula, N represents the number of categories in the data set; a represents the current category number; y a represents the true probability of category a when fine-tuning the target recognition model, and its value is 0 or 1; p a It represents the predicted probability of the pixel belonging to category a by the target recognition model when fine-tuning the target recognition model.
6. The target recognition method based on sonar image segmentation according to any one of claims 1 to 5, characterized in that: In the step 1), Gaussian filtering is performed on each of the sonar images to obtain an enhanced image; The enhanced image is labeled to generate a label map by image segmentation and labeling of the enhanced image; the label map is one-hot encoded, and the label after processing is a 0-1 binary image with the same number of channels as the number of target categories, wherein each pixel with a channel value of 1 indicates that the true probability of the pixel position belonging to the category corresponding to the channel is 1, and each pixel with a channel value of 0 indicates that the pixel position does not belong to the channel, that is, the true probability of the pixel corresponding to the category is 0.
7. A computing device, characterized in that: The computing device comprises: Processor; and A memory, wherein computer program instructions are stored in the memory, and when the computer program instructions are executed by the processor, the processor executes the target recognition method based on sonar image segmentation as claimed in any one of claims 1 to 6.
Citation Information
Patent Citations
Side-scan sonar image domain adaptive learning real-time segmentation method based on AUV
CN112613518A
Multi-scene iris recognition method based on deep learning
CN113591747A
Deep learning-based qualification image classification method and system
CN117788957A