UUV sonar image segmentation method, apparatus and device, and storage medium

Through the method based on convolutional neural network, preprocessing, feature extraction and multi-layer feature fusion of sonar images is solved, and the problem of low segmentation accuracy and insufficient real-time performance of sonar images is achieved, achieving more efficient and accurate segmentation effects.

CN120047689AInactive Publication Date: 2025-05-27TIANJIN QINGRUNBO INTELLIGENT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510526304.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing sonar image segmentation methods have problems such as low segmentation accuracy and difficulty in meeting real-time requirements, especially in underwater detection tasks.

Method used

The method based on convolutional neural network is used to preprocess, feature extraction and multi-layer feature fusion on the sonar image, and finally image segmentation is performed by calculating the probability on the feature map.

Benefits of technology

It improves the accuracy and efficiency of sonar image segmentation, and can better adapt to complex underwater environments and real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047689A_ABST
    Figure CN120047689A_ABST
Patent Text Reader

Abstract

The invention discloses a UUV sonar image segmentation method, device and equipment and a storage medium, and relates to the technical field of image processing, and the method comprises the steps: carrying out the feature extraction of a preprocessed sonar image, and obtaining an initial target feature map; inputting the initial target feature map into a second network model to generate an intermediate target feature map; inputting the initial target feature map and the intermediate target feature map into a first network model to obtain an output target feature map; and calculating the probability that each position point of the output target feature map belongs to the target category, and performing image segmentation on the sonar image according to the probability that each position point belongs to the target category. The method can improve the accuracy of sonar image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and particularly relates to a method, device, equipment and storage medium for UUV sonar image segmentation. Background Art

[0002] As an important underwater detection means, sonar technology plays an irreplaceable role in underwater topographic mapping, underwater archaeology and other aspects. As the direct product of sonar technology, sonar images can reflect the structural characteristics and object shapes of the underwater environment. However, due to the complexity of the underwater environment and the limitations of sonar equipment itself, sonar images often have problems such as noise interference, low resolution and poor contrast, which pose great challenges to subsequent image analysis and understanding.

[0003] Regarding the segmentation problem of sonar images, traditional solutions mainly rely on image processing algorithms and machine learning technologies. For example, the threshold segmentation method distinguishes the target from the background by setting an appropriate gray threshold; the edge detection method uses the edge information in the image to outline the contour of the target. In addition, there are clustering-based methods, such as K-means clustering, fuzzy C-means clustering, etc., which can divide the image into different categories according to the gray value or feature vector of the pixels. These methods have improved the accuracy of sonar image segmentation to a certain extent, but often need to make a trade-off between computational efficiency and segmentation effect.

[0004] Although traditional sonar image segmentation schemes have achieved certain results in specific application scenarios, they generally have the problem of low segmentation accuracy. In addition, with the increasing number of underwater detection tasks and the continuous improvement of real-time requirements, traditional segmentation algorithms have been difficult to meet the needs of actual applications. Therefore, how to improve the accuracy of sonar image segmentation has become an urgent problem to be solved. Summary of the Invention

[0005] The present application provides a method, device, equipment and storage medium for UUV sonar image segmentation, which can improve the accuracy of sonar image segmentation.

[0006] To achieve the above object, the present application adopts the following technical solutions: In a first aspect, the present application provides a method for UUV sonar image segmentation, the method comprising: Obtaining a first sonar image collected by a UUV; Performing preprocessing on the first sonar image to obtain a second sonar image; Performing feature extraction on the second sonar image to obtain an initial target feature map of the second sonar image; Inputting the initial target feature map into a second network model to generate an intermediate target feature map; Input the initial target feature map and the intermediate target feature map into a first network model to obtain an output target feature map; the scale of the output target feature map is half of the size of the initial target feature map. Calculate the probability that each position point of the output target feature map belongs to the target category, and perform image segmentation on the second sonar image according to the probability that each position point belongs to the target category.

[0007] In some possible implementation manners, the first network model includes a first convolution module, a second convolution module, a third convolution module, a fourth convolution module, a first target information fusion module, a second target information fusion module, and a third target information fusion module. The intermediate target feature map includes a fourth target feature map, a fifth target feature map, and a sixth target feature map. Inputting the initial target feature map and the intermediate target feature map into the first network model to obtain an output target feature map includes: Input the initial target feature map into the first convolution module to obtain a first target feature map, splice the first target feature map and the fourth target feature map in the first target information fusion module to obtain a seventh target feature map, input the seventh target feature map into the second convolution module to obtain a second target feature map, splice the second target feature map and the fifth target feature map in the second target information fusion module to obtain an eighth target feature map, input the eighth target feature map into the third convolution module to obtain a third target feature map, splice the third target feature map and the sixth target feature map in the third target information fusion module to obtain a ninth target feature map, and input the ninth target feature map into the fourth convolution module to obtain an output target feature map. Among them, the scale of the first target feature map is the same as that of the fourth target feature map, the scale of the second target feature map is the same as that of the fifth target feature map, and the scale of the third target feature map is the same as that of the seventh target feature map.

[0008] In some possible implementation manners, calculating the probability that each position point of the output target feature map belongs to the target category includes:

[0009] where is the probability that x is the target category, and x is a position point in the output target feature map.

[0010] In some possible implementation manners, the method further includes: Before splicing the first target feature map, the second target feature map, and the third target feature map with the fourth target feature map, the fifth target feature map, and the sixth target feature map respectively, align the sizes of each pair of feature maps to ensure that the spatial position information of the feature maps is consistent during splicing.

[0011] In some possible implementations, preprocessing the first sonar image to obtain a second sonar image includes: Normalize the pixel values of the first sonar image, and calculate through the following formula:

[0012] where a is the pixel value of a pixel point in the first sonar image, and b is the pixel value of a pixel point in the second sonar image.

[0013] In a second aspect, the present application provides a UUV sonar image segmentation device, and the device includes: An acquisition module, configured to acquire a first sonar image collected by a UUV; preprocess the first sonar image to obtain a second sonar image; extract features from the second sonar image to obtain an initial target feature map of the second sonar image; A processing module, configured to input the initial target feature map into a second network model to generate an intermediate target feature map; input the initial target feature map and the intermediate target feature map into a first network model to obtain an output target feature map; A segmentation module, configured to calculate the probability that each position point of the output target feature map belongs to a target category, and perform image segmentation on the second sonar image according to the probability that each position point belongs to the target category.

[0014] In a third aspect, the present application provides a computing device, including a memory and a processor; wherein, one or more computer programs are stored in the memory, and the one or more computer programs include instructions; when the instructions are executed by the processor, the computing device is enabled to execute the method described in any item of the first aspect.

[0015] In a fourth aspect, the present application provides a computer-readable storage medium, and the computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method described in any item of the first aspect.

[0016] In a fifth aspect, the present application provides a computer program product, and the computer program product includes one or more computer instructions. When the computer instructions are executed by a computer, the computer executes the method described in any item of the first aspect.

[0017] It can be seen from the above technical solutions that the present application has at least the following beneficial effects: In this application, a first sonar image collected by a UUV is obtained; the first sonar image is preprocessed to obtain a second sonar image; feature extraction is performed on the second sonar image to obtain an initial target feature map of the second sonar image; the initial target feature map is input into a second network model to generate an intermediate target feature map; the initial target feature map and the intermediate target feature map are input into a first network model to obtain an output target feature map; the probability that each position point of the output target feature map belongs to a target category is calculated, and the second sonar image is segmented according to the probability that each position point belongs to the target category. It can be seen that based on two convolutional neural network models, this application processes the initial target feature map to obtain an output target feature map, and by calculating the probability that each position point on the output target feature map belongs to the target category, the segmentation process of the sonar image is realized, and the accuracy of sonar image segmentation is improved.

[0018] It should be understood that the description of technical features, technical solutions, beneficial effects or similar languages in this application does not imply that all features and advantages can be achieved in any single embodiment. On the contrary, it can be understood that the description of features or beneficial effects means that at least one embodiment includes specific technical features, technical solutions or beneficial effects. Therefore, the description of technical features, technical solutions or beneficial effects in this specification does not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions and beneficial effects described in this embodiment can be combined in any appropriate manner. Those skilled in the art will understand that an embodiment can be implemented without one or more specific technical features, technical solutions or beneficial effects of a specific embodiment. In other embodiments, additional technical features and beneficial effects can also be identified in specific embodiments that do not embody all embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flowchart of a method for segmenting UUV sonar images provided by an embodiment of this application; Figure 2 It is a schematic diagram of information fusion of target feature maps provided by an embodiment of this application; Figure 3 It is a schematic diagram of information splicing of target feature maps provided by an embodiment of this application; Figure 4 It is a schematic structural diagram of a method for segmenting UUV sonar images provided by an embodiment of this application; Figure 5 It is a flowchart of network model training provided by an embodiment of this application; Figure 6 It is a schematic diagram of a device for segmenting UUV sonar images provided by an embodiment of this application; Figure 7Schematic diagram of a computing device provided by an embodiment of the present application. Detailed implementation manners

[0020] Terms such as "first", "second", and "third" in the description and drawings of the present application are used to distinguish different objects, rather than to limit a specific order.

[0021] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0022] For the sake of clear and concise description of the following embodiments, a brief introduction to the related technologies is given first: YOLOv5, namely You Only Look Once v5, is a fast and efficient single-stage object detection algorithm developed by the Ultralytics team. YOLOv5 regards the object detection task as a regression problem and directly predicts the position and class probability of the bounding box through a neural network. The model adopts the Feature Pyramid Network (FPN) structure to fuse feature maps of different scales and enhance the detection ability for objects of different sizes.

[0023] U-Net is a convolutional neural network for image segmentation proposed by Olaf Ronneberger et al. in 2015. U-Net consists of a contracting path (downsampling) and an expanding path (upsampling). The contracting path gradually reduces the image resolution through consecutive convolution and pooling operations to extract high-level semantic features of the image; the expanding path fuses low-level features with high-level features through transposed convolution (deconvolution) and skip connections, gradually restores the image resolution, and achieves precise segmentation of different regions in the image.

[0024] CSP-Darknet53 is a neural network structure for object detection tasks, which is combined by CSPNet (CrossStage Partial Network) and Darknet53 and is widely used in models such as YOLOv4. CSP-Darknet53 is mainly composed of convolutional layers, CSP modules and residual structures. The convolutional layers are used to extract the basic features of the image; the CSP module is its core structure. By dividing the feature map of the basic layer into two parts, one part is directly passed to the next stage, and the other part is merged with the former after a series of convolutions and residual processes, so as to reduce the computational amount and enhance the reusability of features; the residual structure solves the problem of gradient disappearance in the training process of deep neural networks, enabling the network to learn richer features. The design of this structure aims to improve the computational efficiency and detection accuracy of the model. The CSPNet part reduces redundant calculations, lowers the memory cost, and enhances the propagation of the gradient flow through cross-stage feature fusion, making the model more stable during training. While Darknet53 provides powerful feature extraction capabilities. It draws on the residual structure of ResNet and can build deeper networks to extract more advanced semantic features.

[0025] FPN (Feature Pyramid Network) and PAN (Path Aggregation Network) are important network structures for processing multi-scale objects in the field of object detection. The combination of the two can further improve the model's detection ability for different-scale objects.

[0026] FPN aims to improve the detection performance for different-sized objects by constructing a feature pyramid and fusing feature maps of different scales. In a top-down manner, it upsamples the high-level semantic feature maps and then fuses them with the low-level feature maps of the corresponding scales. The high-level feature maps are rich in semantic information and are suitable for detecting large objects; the low-level feature maps have more detailed information and are beneficial for detecting small objects. Through this fusion method, each scale of feature map has rich semantic and detailed information.

[0027] PAN is proposed based on FPN. It adds a bottom-up path to supplement more low-level information. In the bottom-up path, the low-level feature maps are fused with the high-level feature maps through downsampling. This two-way feature fusion path enables the model to better utilize features at different levels and enhances the efficiency of feature propagation.

[0028] After combining FPN and PAN, the model has a more powerful multi-scale feature fusion ability. It can detect large targets using high-level semantic information like FPN and supplement low-level detailed information through a bottom-up path like PAN to more accurately detect small targets. This two-way feature fusion mechanism greatly improves the model's detection performance for targets of different scales in complex scenarios.

[0029] Most current sonar image segmentation methods are based on traditional image processing algorithms, and it often takes more than one second to process an image. Traditional sonar image segmentation methods usually rely on manually designed features, have poor robustness, are sensitive to noise and interference, and are difficult to cope with the complex UUV operating environment.

[0030] In view of this, an embodiment of the present application provides a UUV sonar image segmentation method, which is applied to a processing device. In this method, a first sonar image collected by the UUV is obtained; the first sonar image is preprocessed to obtain a second sonar image; feature extraction is performed on the second sonar image to obtain an initial target feature map of the second sonar image; the initial target feature map is input into a second network model to generate an intermediate target feature map; the initial target feature map and the intermediate target feature map are input into a first network model to obtain an output target feature map; the probability that each position point of the output target feature map belongs to the target category is calculated, and the second sonar image is segmented according to the probability that each position point belongs to the target category. It can be seen that the present application processes the initial target feature map based on two convolutional neural network models to obtain an output target feature map, and realizes the segmentation processing of the sonar image by calculating the probability that each position point on the output target feature map belongs to the target category, which can improve the accuracy of sonar image segmentation on the basis of the traditional solution.

[0031] To make the technical solution of the present application clearer and easier to understand, the following introduces a UUV sonar image segmentation method provided by an embodiment of the present application in conjunction with the accompanying drawings. As Figure 1 shown, this figure is a flowchart of a UUV sonar image segmentation method provided by an embodiment of the present application. The UUV sonar image segmentation method includes: S101. Construct a sonar image data set.

[0032] The processing device first determines the data collection scenario and type: According to the actual application conditions of the sonar image, it is necessary to determine the environmental conditions and target types for sonar image acquisition, such as different water depths and water qualities, so that the data set samples have pertinence and practicability.

[0033] Then, data acquisition is carried out: Sonar devices of different types vary in terms of resolution, operating frequency, detection range, etc., resulting in differences in the acquired image features and quality. Therefore, it is necessary to use appropriate sonar devices for image acquisition according to the requirements of the dataset samples; set specific acquisition schemes, such as planning the acquisition area, depth range, target type, etc., and record relevant parameters during acquisition, such as the configuration of the sonar device, acquisition time, location, environment, etc.; in addition, attention should also be paid to the standardization of operations during the acquisition process to avoid interference caused by human factors.

[0034] Next, the data is processed: Format conversion, converting the acquired images into a unified format, such as JPEG, PNG, etc.; data cleaning, removing invalid acquired images due to various reasons to ensure data quality; image annotation (using LabelMe, LabelImg, etc.), for example, target detection and target segmentation require annotating the position and category, while classification only requires the image category; data augmentation, it is difficult to obtain a large number of samples for sonar images, so it is necessary to augment the data, such as rotation, cropping, adding noise, etc.; dataset division, that is, dividing the model training dataset and the validation dataset, taking 80% of the data as the training set and 20% as the validation set.

[0035] Among them, for image format conversion, it can be based on the conversion software provided by the sonar device manufacturer or existing open-source conversion tools; for data cleaning, invalid acquired images can be screened manually, or scripts can be written to clean the images based on certain rules; for image annotation, images can be annotated based on LabelMe, LabelImg, etc., and the labels are the position, category of the target, and the target segmentation map. The format of the position label is [x, y, w, h], where x is the x coordinate of the center point of the target, y is the y coordinate of the center point of the target, w is the width of the target, and h is the height of the target; the category label is in the format of a vector, such as [0, 1, 0, 0], and this vector corresponds one-to-one with the actual category values, such as [car, bus, people, motorbike], and [0, 1, 0, 0] represents the category of bus; the segmented label map is a binary map, with pixel value 1 representing the target and 0 representing the background; for data augmentation, data augmentation can be performed based on the library interfaces in OpenCV, such as rotation, cropping, adding noise, etc.

[0036] The dataset is obtained through the above steps for training the model.

[0037] S102. The processing device obtains the first sonar image.

[0038] The processing device obtains the sonar image from the sensor and determines this sonar image as the first sonar image.

[0039] S103. The processing device preprocesses the first sonar image to obtain the second sonar image.

[0040] Convolutional neural networks usually require the input images to have a fixed size or meet certain conditions (such as being a multiple of 32, etc.). In addition, normalizing the pixel values of the input images also helps the network converge quickly.

[0041] First, perform image size normalization on the first sonar image: When the actual image size is inconsistent with the network requirements, image scaling and padding are required. Scaling can be divided into proportional scaling and non-proportional scaling. Non-proportional scaling will cause image distortion and is generally not used. When scaling, interpolation algorithms are used to calculate new pixel values, such as nearest neighbor interpolation, bilinear interpolation, etc. The nearest neighbor interpolation algorithm has a fast calculation speed but will lose image details; bilinear interpolation has a better scaling effect but more computational effort. In order to retain more details for subsequent processing and analysis, this application uses the bilinear interpolation algorithm for image scaling.

[0042] Next, perform image pixel value normalization on the size-normalized first sonar image: In the RGB color image model, each pixel is composed of three color channels: red, green, and blue, and the pixel range of each channel is from 0 to 255. Normalization is to normalize all pixel values to a fixed range. For image segmentation tasks, linear normalization is usually used to map the pixel values to the range between -1 and 1 to obtain the second sonar image, which is calculated by the following formula:

[0043] where a is the pixel value of the first sonar image, and b is the pixel value of the second sonar image.

[0044] By adjusting the first sonar image to a unified size, it can be ensured that all images have the same size in the subsequent processing process, thus avoiding the additional calculations and adjustments required when processing images of different sizes. Moreover, the unified size can simplify the image processing process and improve the running efficiency of the algorithm. The size-normalized images usually have a smaller resolution, thus reducing the storage requirements of the image data.

[0045] Pixel value normalization can convert the pixel values in the image to a smaller range (such as between 0 and 1), making the pixel values between different images comparable. This helps the algorithm extract the texture, edges, and other features of the image more accurately, thereby improving the effect of image processing and computer vision tasks. At the same time, the normalized pixel values have a smaller change range, which can improve the numerical stability and convergence speed of the algorithm. By normalizing the image pixel values to a unified range, the changes in brightness and contrast between different images can be reduced. This helps reduce the sensitivity of the network model to these changes and makes it perform more stably on different image data.

[0046] S104. The processing device extracts features from the second sonar image to obtain the initial target feature map of the second sonar image.

[0047] The processing device uses a feature extraction network to extract features from the second sonar image. The feature extraction network is mainly constructed based on the CSP-Darknet53 structure and consists of a Slice layer, a Cross Stage Partial (CSP) layer, a Spatial Pyramid Pooling (SPP) layer, and a Convolutional Block (CBL).

[0048] Among them, the Slice layer is used to slice, splice, and perform convolution on the input feature map. Since the computing power on the UUV is limited, the Slice layer can reduce the image size and computational amount. Compared with pooling, slice splicing can retain more detailed information, enabling the network to learn richer spatial features and facilitating the improvement of model performance.

[0049] The CSP (Cross Stage Partial) layer is a cross-stage local fusion strategy that divides the input feature map into two parts. One part undergoes convolution operations, and the other part is directly passed to the next stage. Finally, the two are merged. This strategy can reduce the number of parameters without affecting the network accuracy, alleviate the gradient disappearance problem, and accelerate network convergence.

[0050] The SPP (Spatial Pyramid Pooling) layer is spatial pyramid pooling. The receptive field of a conventional convolution is fixed, and it cannot capture multi-scale information well for particularly large or small targets. The SPP layer performs pooling using pooling windows of different sizes and stitches together the results of each pooling to form a feature map with multi-scale information, which can reduce the size of the second sonar image and help improve the model's perception ability for targets of different sizes.

[0051] The second sonar image first undergoes size adjustment and channel stitching through the Slice layer, then feature extraction through several CSP layers and ordinary convolutional layers, and passes through an SPP layer in the middle for target multi-scale information extraction. Thus, the initial target feature map of the second sonar image is obtained.

[0052] S105. Input the initial target feature map into the second network model to generate an intermediate target feature map.

[0053] The intermediate target feature map includes a fourth target feature map, a fifth target feature map, and a sixth target feature map. In order to perform a splicing operation with the first target feature map, the second target feature map, and the third target feature map of the first network model, the fourth target feature map, the fifth target feature map, and the sixth target feature map of the second network model are respectively subjected to adaptive pooling or an interpolation algorithm to unify the scales with the first target feature map, the second target feature map, and the third target feature map.

[0054] In the embodiment of the present application, the second network model adopts the detection architecture of YOLOv5, which is obtained by combining FPN and PAN. High-level features usually contain stronger semantic information, while low-level features contain richer localization information. FPN is a top-down process that increases the size of the feature map through upsampling, fusing deep semantic information with low-level detailed information to form a feature pyramid; based on FPN, PAN adds a bottom-up pyramid to transfer strong localization features at the low level and enhance the model's ability to perceive small targets. This "twin-tower strategy" not only retains the semantic enhancement ability of FPN but also supplements the transmission of localization information, making the network more comprehensive and accurate in processing multi-scale targets. In PAN, low-level features are transferred to the high level through downsampling and horizontal connections, complementing the top-down transmission of FPN. In this way, the network can utilize both high-level semantic information and low-level localization information, improving the accuracy of object detection.

[0055] In this step, the processing device generates target boxes, class probabilities, and center points to predict information such as the target position, class, and confidence in the second sonar image.

[0056] Therefore, the processing device inputs the initial target feature map into the second network model to generate a fourth target feature map, a fifth target feature map, and a sixth target feature map. To facilitate splicing with the corresponding target feature maps in S105, the fourth target feature map has the same scale as the first target feature map, the fifth target feature map has the same scale as the second target feature map, and the sixth target feature map has the same scale as the third target feature map.

[0057] S106. Input the initial target feature map and the intermediate target feature map into the first network model to obtain an output target feature map.

[0058] The first network model is the decoder in the image segmentation network. Conventional image segmentation networks generally based on encoder-decoder structures, where the encoder extracts image features and the decoder restores the extracted feature map to nearly the size of the original input image to achieve pixel-level segmentation. The feature extraction network in the present application is the encoder in this image segmentation network.

[0059] In this application, the decoder enlarges the size of the initial target feature map by using skip connections and upsampling to generate target feature maps of different scales. Although this method combines low-level spatial information and high-level semantic information, since the main function of the image segmentation network is target segmentation and no special design and training are carried out for target localization, the decoder of the image segmentation network does not really utilize effective target position information.

[0060] Therefore, in this application, by adding the method in S104, the category and position information of the target are fused into the decoder of the image segmentation network to improve the accuracy of sonar image segmentation.

[0061] As Figure 2 shown, this figure is a schematic diagram of the fusion of target feature map information provided by an embodiment of this application. The first network model includes a first convolutional module, a second convolutional module, a third convolutional module, a fourth convolutional module, a first target information fusion module, a second target information fusion module, and a third target information fusion module. The intermediate target feature maps include a fourth target feature map, a fifth target feature map, and a sixth target feature map. Inputting the initial target feature map and the intermediate target feature maps into the first network model to obtain an output target feature map, and the scale of the output target feature map is half of the size of the initial target feature map, including: Input the initial target feature map into the first convolutional module to obtain a first target feature map. Input the first target feature map and the fourth target feature map into the first target information fusion module for splicing to obtain a seventh target feature map. Input the seventh target feature map into the second convolutional module to obtain a second target feature map. Input the second target feature map and the fifth target feature map into the second target information fusion module for splicing to obtain an eighth target feature map. Input the eighth target feature map into the third convolutional module to obtain a third target feature map. Input the third target feature map and the sixth target feature map into the third target information fusion module for splicing to obtain a ninth target feature map. Input the ninth target feature map into the fourth convolutional module to obtain the output target feature map. Among them, the scale of the first target feature map is the same as that of the fourth target feature map, the scale of the second target feature map is the same as that of the fifth target feature map, and the scale of the third target feature map is the same as that of the seventh target feature map.

[0062] In this process, the target feature maps generated in different stages are concatenated and fused. Feature maps in different stages often contain information at different scales. Shallow feature maps may retain more detailed information, while deep feature maps contain more abstract and high-level semantic information. Through this concatenation and fusion operation, feature information at different scales can be integrated together, enabling the final output target feature map to take into account both detailed and semantic information, thus enhancing the expressive power of the features. Each convolution operation further extracts and transforms the features of the feature map, mining out more representative features. The concatenation and fusion of feature maps establish connections between different features, making the feature information more rich and diverse. This rich feature representation helps the model make more accurate decisions and judgments in subsequent tasks. Through multiple feature fusions and convolution operations, the effective propagation of features in the model can be promoted. Moreover, each convolution operation and feature fusion are performed on relatively small feature maps, reducing the computational amount and memory consumption and improving the running efficiency of the model.

[0063] During the concatenation process, the feature information of the previous stage can be passed to the subsequent stage, avoiding information loss and the vanishing gradient problem. This enables the model to better learn global feature information, improving the learning ability and generalization ability of the model. The entire process adopts a modular design, including a convolution module and a target information fusion module. This modular structure makes the construction and adjustment of the model more flexible, easy to expand and improve.

[0064] The embodiment of the present application also provides a schematic diagram of the concatenation of target feature map information, as Figure 3 shown. This figure is a schematic diagram of the concatenation of target feature map information provided by the embodiment of the present application. Before the first target feature map, the second target feature map, and the third target feature map are respectively concatenated with the fourth target feature map, the fifth target feature map, and the sixth target feature map, the sizes of each pair of feature maps are aligned to ensure that the spatial position information of the feature maps is consistent during concatenation.

[0065] S107. Calculate the probability that each position point of the output target feature map belongs to the target category, and perform image segmentation on the second sonar image according to the probability that each position point belongs to the target category.

[0066] The output target feature map is a feature map with a width of w, a height of h, and a channel number of 1. That is to say, there are w times h position points on this feature map. Calculating the probability that each position point of the output target feature map belongs to the target category includes:

[0067] Among them, is the probability that x is the target category, and x is the position point in the output target feature map.

[0068] When the value is greater than 0.5, the position point x is considered as the target class. When the value is less than 0.5, the position point x is considered as the background. Calculate and judge each position point until all position points are judged, and one or more groups of clusters with values greater than 0.5 are obtained to realize the segmentation of the sonar image.

[0069] To make the technical solution of the present application clearer and easier to understand, the following introduces a UUV sonar image segmentation method provided by an embodiment of the present application in conjunction with the accompanying drawings. As Figure 4 shown, this figure is a schematic structural diagram of a UUV sonar image segmentation method provided by an embodiment of the present application.

[0070] It can be seen from the figure that the feature extraction module is composed of a slice layer Slice, a layer CBL composed of a combination of convolution, normalization and activation functions, a cross-stage local fusion strategy CSP, and a spatial pyramid pooling SPP.

[0071] The target detection module is obtained by combining FPN and PAN; the Decoder is a decoder.

[0072] The target information fusion module transmits the fourth target feature map, the fifth target feature map and the sixth target feature map generated in the target detection module to the decoder, and splices them with the first target feature map, the second target feature map and the third target feature map respectively, and then performs a convolution operation to realize the segmentation of the sonar image.

[0073] Based on the above content, the present application processes the initial target feature map based on two convolutional neural network models to obtain an output target feature map, and realizes the segmentation processing of the sonar image by calculating the probability that each position point on the output target feature map belongs to the target class, which can improve the accuracy of sonar image segmentation on the basis of the traditional scheme.

[0074] In order to enable the network model to learn effective target information, it is necessary to train the target detection module of the network separately. The overall training strategy is divided into two parts, namely target detection module training and target segmentation module training. The target segmentation training is based on the completion of the target detection training. Since it is difficult to collect a large amount of sonar image data, when training the target detection module, first pre-train based on the open source dataset, and then fine-tune the model on the sonar image dataset. As Figure 5 shown, this figure is a flowchart of a network model training provided by an embodiment of the present application.

[0075] The embodiment of the present application also provides a UUV sonar image segmentation device, as Figure 6As shown in the figure, this figure is a schematic diagram of a UUV sonar image segmentation device provided by an embodiment of the present application. The device includes: an acquisition module 601, a processing module 602, and a segmentation module 603; The acquisition module 601 is configured to acquire a first sonar image; perform preprocessing on the first sonar image to obtain a second sonar image; and perform feature extraction on the second sonar image to obtain an initial target feature map of the second sonar image; The processing module 602 is configured to input the initial target feature map into a second network model to generate an intermediate target feature map; and input the initial target feature map and the intermediate target feature map into a first network model to obtain an output target feature map; The segmentation module 603 is configured to calculate the probability that each position point of the output target feature map belongs to the target category, and perform image segmentation on the second sonar image according to the probability that each position point belongs to the target category.

[0076] In some possible implementation manners, the first network model includes a first convolution module, a second convolution module, a third convolution module, a fourth convolution module, a first target information fusion module, a second target information fusion module, and a third target information fusion module. The intermediate target feature map includes a fourth target feature map, a fifth target feature map, and a sixth target feature map. The processing module 602 is specifically configured to input the initial target feature map into the first convolution module to obtain a first target feature map, splice the first target feature map and the fourth target feature map in the first target information fusion module to obtain a seventh target feature map, input the seventh target feature map into the second convolution module to obtain a second target feature map, splice the second target feature map and the fifth target feature map in the second target information fusion module to obtain an eighth target feature map, input the eighth target feature map into the third convolution module to obtain a third target feature map, splice the third target feature map and the sixth target feature map in the third target information fusion module to obtain a ninth target feature map, and input the ninth target feature map into the fourth convolution module to obtain an output target feature map. Wherein, the scale of the first target feature map is the same as the scale of the fourth target feature map, the scale of the second target feature map is the same as the scale of the fifth target feature map, and the scale of the third target feature map is the same as the scale of the seventh target feature map.

[0077] In some possible implementation manners, the segmentation module 603 is specifically configured to calculate the probability that each position point of the output target feature map belongs to the target category, including:

[0078] Wherein, is the probability that x is the target category, and x is the position point in the output target feature map.

[0079] In some possible implementations, the processing module 602 is further configured to align the sizes of each pair of feature maps before splicing the first target feature map, the second target feature map, and the third target feature map with the fourth target feature map, the fifth target feature map, and the sixth target feature map respectively, so as to ensure that the spatial position information of the feature maps is consistent during splicing.

[0080] In some possible implementations, the obtaining module 601 is specifically configured to preprocess the first sonar image to obtain a second sonar image, including: Normalize the pixel values of the first sonar image, and calculate through the following formula:

[0081] where a is the pixel value of a pixel point in the first sonar image, and b is the pixel value of a pixel point in the second sonar image.

[0082] The embodiment of the present application further provides a computing device. As Figure 7 shown, this figure is a schematic diagram of a computing device provided by the embodiment of the present application. The computing device 400 includes a bus 401, a processor 402, a communication interface 403, and a memory 404. The processor 402, the memory 404, and the communication interface 403 communicate with each other through the bus 401.

[0083] The bus 401 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 7 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0084] The processor 402 may be any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), etc.

[0085] The communication interface 403 is used for external communication.

[0086] The memory 404 may include volatile memory, such as random access memory (RAM). The memory 404 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0087] Executable code is stored in the memory 404, and the processor 402 executes the executable code to perform the foregoing UUV sonar image segmentation method.

[0088] An embodiment of this application also provides a computer-readable storage medium. The computer-readable storage medium may be any available medium that a computing device can store or a data storage device such as a data center that includes one or more available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to perform the foregoing UUV sonar image segmentation method.

[0089] An embodiment of this application also provides a computer program product that includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the processes or functions described in the embodiments of this application are fully or partially generated.

[0090] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center in a wired manner (such as coaxial cable, optical fiber) or a wireless manner (such as infrared, wireless, microwave, etc.).

[0091] When the computer program product is executed by a computer, the computer performs any of the foregoing UUV sonar image segmentation methods. The computer program product may be a software installation package. In the case where any of the foregoing UUV sonar image segmentation methods is needed, the computer program product may be downloaded and executed on the computer.

[0092] The descriptions of the processes or structures corresponding to the foregoing various drawings each have their own emphases. For parts not detailed in a certain process or structure, reference may be made to the relevant descriptions of other processes or structures.

[0093] As described above, it is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application shall be covered by the protection scope of this application.

Claims

1. A UUV sonar image segmentation method, characterized in that: The method comprises: Acquire the first sonar image collected by the UUV; Preprocessing the first sonar image to obtain a second sonar image; Performing feature extraction on the second sonar image to obtain an initial target feature map of the second sonar image; Inputting the initial target feature map into a second network model to generate an intermediate target feature map; Inputting the initial target feature map and the intermediate target feature map into a first network model to obtain an output target feature map; the scale of the output target feature map is half the size of the initial target feature map; The probability that each position point of the output target feature map belongs to the target category is calculated, and the second sonar image is segmented according to the probability that each position point belongs to the target category.

2. The method according to claim 1, characterized in that The first network model includes a first convolution module, a second convolution module, a third convolution module, a fourth convolution module, a first target information fusion module, a second target information fusion module and a third target information fusion module, the intermediate target feature map includes a fourth target feature map, a fifth target feature map and a sixth target feature map, the initial target feature map and the intermediate target feature map are input into the first network model to obtain an output target feature map, including: The initial target feature map is input into the first convolution module to obtain the first target feature map, the first target feature map and the fourth target feature map are input into the first target information fusion module for splicing to obtain the seventh target feature map, the seventh target feature map is input into the second convolution module to obtain the second target feature map, the second target feature map and the fifth target feature map are input into the second target information fusion module for splicing to obtain the eighth target feature map, the eighth target feature map is input into the third convolution module to obtain the third target feature map, the third target feature map and the sixth target feature map are input into the third target information fusion module for splicing to obtain the ninth target feature map, and the ninth target feature map is input into the fourth convolution module to obtain the output target feature map, wherein the scale of the first target feature map is the same as that of the fourth target feature map, the scale of the second target feature map is the same as that of the fifth target feature map, and the scale of the third target feature map is the same as that of the seventh target feature map.

3. The method according to claim 1, characterized in that The calculating and outputting the probability that each position point of the target feature map belongs to the target category includes: in, x is the probability of the target category, and x is the position point in the output target feature map.

4. The method according to claim 1, characterized in that: The method further comprises: Before splicing the first target feature map, the second target feature map and the third target feature map with the fourth target feature map, the fifth target feature map and the sixth target feature map respectively, the size of each pair of feature maps is aligned to ensure that the spatial position information of the feature maps is consistent when splicing.

5. The method according to claim 1, characterized in that The preprocessing of the first sonar image to obtain the second sonar image includes: The pixel values ​​of the first sonar image are normalized and calculated using the following formula: Wherein, a is the pixel value of the first sonar image, and b is the pixel value of the second sonar image.

6. A UUV sonar image segmentation device, characterized in that: The device comprises: The acquisition module is used to acquire a first sonar image collected by the UUV; preprocess the first sonar image to obtain a second sonar image; perform feature extraction on the second sonar image to obtain an initial target feature map of the second sonar image; A processing module, configured to input the initial target feature map into the second network model to generate an intermediate target feature map; input the initial target feature map and the intermediate target feature map into the first network model to obtain an output target feature map, wherein the scale of the output target feature map is half the size of the initial target feature map; The segmentation module is used to calculate the probability that each position point of the output target feature map belongs to the target category, and perform image segmentation on the second sonar image according to the probability that each position point belongs to the target category.

7. The device according to claim 6, characterized in that The first network model includes a first convolution module, a second convolution module, a third convolution module, a fourth convolution module, a first target information fusion module, a second target information fusion module and a third target information fusion module, the intermediate target feature map includes a fourth target feature map, a fifth target feature map and a sixth target feature map, the processing module is specifically used to input the initial target feature map into the first convolution module to obtain the first target feature map, input the first target feature map and the fourth target feature map into the first target information fusion module for splicing to obtain the seventh target feature map, input the seventh target feature map into the second convolution module to obtain the second target feature map, and input the second target feature map into the second convolution module to obtain the second target feature map. The first target feature map and the fifth target feature map are input into the second target information fusion module for splicing to obtain an eighth target feature map, the eighth target feature map is input into the third convolution module to obtain a third target feature map, the third target feature map and the sixth target feature map are input into the third target information fusion module for splicing to obtain a ninth target feature map, and the ninth target feature map is input into the fourth convolution module to obtain an output target feature map, wherein the scale of the first target feature map is the same as that of the fourth target feature map, the scale of the second target feature map is the same as that of the fifth target feature map, and the scale of the third target feature map is the same as that of the seventh target feature map.

8. The device according to claim 6, characterized in that The processing module is also used to align the size of each pair of feature maps before splicing the first target feature map, the second target feature map and the third target feature map with the fourth target feature map, the fifth target feature map and the sixth target feature map respectively, to ensure that the spatial position information of the feature maps is consistent during splicing.

9. A computing device, characterized in that including memory and processor; One or more computer programs are stored in the memory, and the one or more computer programs include instructions; when the instructions are executed by the processor, the computing device executes the method as claimed in any one of claims 1 to 5.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Sonar target detection method and device based on improved YOLOv4

    CN116152649A

  • Target real-time detection and classification method based on sonar image, medium and system

    CN118865003A

  • Multi-task panoptic driving perception method and system based on improved yolov5

    WO2024060605A1