Semantic segmentation-based submarine sonar image processing method

By improving the U-Net dual-branch semantic segmentation network combined with the submarine landform boundary prior map and the landform mutual feed gating module, the problems of blurred boundaries and poor cross-region adaptability in submarine sonar image segmentation are solved, and high-precision and real-time submarine landform segmentation are achieved, which is suitable for marine geological surveys and underwater operations.

CN120355928AActive Publication Date: 2025-07-22SHENYANG LIAOHAI EQUIP

Patent Information

Application Number
CN202510837340.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-07-22
Estimated Expiration
2045-06-23

AI Technical Summary

Technical Problem

When facing the influence of high noise, speckle and shadow, the existing subsea sonar image segmentation method is difficult to accurately segment the geomorphological boundaries, and lacks the integration of prior knowledge of physical geomorphology, resulting in blurred or misaligned segmentation results, which is difficult to meet the requirements of geomorphological vectorization and GIS entry. At the same time, the cross-region adaptability is poor, and a large amount of labeled data is required. The existing network structure cannot take into account global category discrimination and submeter-level boundary positioning.

Method used

A improved U-Net dual-branch semantic segmentation network is constructed, combined with semantic branches and boundary branches, and introduced a priori map of the seabed landform boundary, and feature fusion is performed through the geomorphological mutual feeding gating module. Depth separable convolution and frequency band maintenance convolution are adopted, and supervision and training is performed based on category cross entropy, boundary Dice loss and topological consistency loss, and deployed to an autonomous underwater vehicle for real-time segmentation.

Benefits of technology

It improves the segmentation accuracy and stability of the seabed landform boundaries, and can improve the average mIoU indicator in the mixed scenarios of small goals and multiple categories, reduce labeling costs, and achieve real-time high-quality segmentation. It is suitable for marine geological surveys and underwater operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355928A_ABST
    Figure CN120355928A_ABST
Patent Text Reader

Abstract

The invention discloses a seabed sonar image processing method based on semantic segmentation. The method comprises the following steps: S1, obtaining preprocessed sonar image blocks; s2, generating a submarine landform boundary priori map corresponding to the preprocessed sonar image blocks pixel by pixel; s3, an improved U-Net double-branch semantic segmentation network is constructed, the improved U-Net double-branch semantic segmentation network comprises a semantic branch and a boundary branch, and a landform mutual feedback gating module is configured in each scale jumper connection layer; s4, obtaining a fusion feature map; s5, deploying to the autonomous underwater vehicle to form an onboard reasoning model; and S6, performing online semantic segmentation reasoning on the preprocessed sonar image block acquired in real time by using the onboard reasoning model, and outputting a semantic category diagram and a boundary vector diagram corresponding to the preprocessed sonar image block. According to the method, the average mIoU index of the model is improved in a small-target and multi-category hybrid scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sonar, and in particular to a method for processing submarine sonar images based on semantic segmentation. Background Art

[0002] With the continuous expansion of ocean resource development, submarine geological exploration and underwater engineering application scenarios, the technology of submarine geomorphology recognition and mapping based on sonar imaging has gradually become an important means for modern ocean information acquisition and environmental assessment. Currently, the mainstream semantic segmentation methods for submarine sonar images mostly use deep convolutional neural networks to automatically process sonar images, dividing the original sonar images into different category regions such as sediments, bedrock, and artificial structures, providing a data basis for geological analysis, resource exploration, and underwater operations.

[0003] However, the existing technologies have exposed many deficiencies in practical applications. First of all, sonar images naturally have high noise, speckle, and shadow effects, making the geomorphological boundaries appear blurred or broken in the imaging results. The traditional U-Net segmentation method often renders the real hard boundaries as wide transition zones when processing such data, which is difficult to meet the requirements of subsequent geomorphological vectorization and accurate GIS entry. In addition, the existing network structures generally lack the integration of prior knowledge of physical geomorphology and often only rely on pixel-level gradients or simple attention mechanisms. As a result, when the model faces terrain breaks, complex shadows, and multi-scale heterogeneous targets, the segmentation results are prone to pseudo-edges, contour gaps, and category misclassification, seriously affecting the engineering value of geomorphological mapping and feature extraction.

[0004] Secondly, the acquisition of submarine sonar data is limited by the high cost of ship time and the difficulty of on-site collection. The sample size of the public dataset is limited, and there are significant differences in different survey equipment, regions, and annotation standards, making it difficult for the existing methods to generalize in cross-regional and cross-sensor scenarios or to rely on a large amount of pixel-level fully annotated data, reducing the actual deployment efficiency and adaptability of the algorithm. At the same time, the existing networks based on single-branch or simple multi-branch structures usually cannot balance the dual requirements of global category discrimination and sub-meter boundary localization, and lack the ability to retain the high-frequency details of sonar echoes, resulting in the output geomorphological segmentation maps being difficult to provide high-quality decision support for subsequent underwater navigation, resource screening, and environmental monitoring.

[0005] Therefore, there is an urgent need for an innovative method that can improve the segmentation accuracy, enhance the boundary expression, reduce the annotation dependence, and have real-time inference ability to better serve the actual needs of marine geological surveys and intelligent underwater operations. Summary of the Invention

[0006] An object of the present invention is to propose a method for processing submarine sonar images based on semantic segmentation. In the scenario of small targets and multi-category mixing, the average mIoU index of the model has been improved.

[0007] A method for processing underwater sonar images based on semantic segmentation according to an embodiment of the present invention includes the following steps: S1. Collect original underwater sonar image data for standardization processing, preprocess the standardized sonar image blocks to obtain preprocessed sonar image blocks; S2. Calculate the terrain slope feature based on the preprocessed sonar image blocks, fuse the features, and generate a prior map of the underwater geomorphic boundary corresponding to each pixel of the preprocessed sonar image blocks; S3. Construct an improved U-Net dual-branch semantic segmentation network, and the improved U-Net dual-branch semantic segmentation network includes a semantic branch and a boundary branch; S4. Configure a geomorphic mutual feedback gating module in each scale skip connection layer, embed the prior map of the underwater geomorphic boundary into the multi-scale skip connection structure of the improved U-Net dual-branch semantic segmentation network through the geomorphic mutual feedback gating module, and perform weighted fusion on the semantic feature map output by the semantic branch to obtain a fused feature map; S5. Use the preprocessed sonar image blocks, the prior map of the underwater geomorphic boundary, and the corresponding pixel-level class labels as training data, use a joint loss function to supervise and train the improved U-Net dual-branch semantic segmentation network to obtain a trained segmentation model, and deploy it to an autonomous underwater vehicle to form an on-board inference model; S6. Use the on-board inference model to perform online semantic segmentation inference on the preprocessed sonar image blocks obtained in real time, and output a semantic class map and a boundary vector map corresponding to the preprocessed sonar image blocks.

[0008] Optionally, the S1 includes the following steps: S11. Use an underwater detection device to collect original underwater sonar image data to form an original image set , where represents the th original underwater sonar image, is the total number of collected images; S12. Perform size regularization and image block cutting on the original image set , and divide each original image into a number of non-overlapping or partially overlapping sub-blocks according to a preset block size to generate a set of standardized sonar image blocks; S13. Perform an adaptive spatial domain filtering operation on the set of standardized sonar image blocks, construct a local response window, and apply a weighted denoising model in each window to obtain denoised image blocks; S14. Perform texture-preserving enhancement processing on each denoised image block, and construct a set of structural response operators in multiple directions , where the set of structural response operators includes horizontal, vertical, and two diagonal directions, corresponding to the direction angles respectively. Gradient response features of image patches are extracted in each direction. By comparing the response intensities in different directions and selecting the direction with the maximum response for enhancement, image patches with enhanced texture structures are generated; S15. Perform gray dynamic range equalization processing on the image patches with enhanced texture structures, map the original pixel gray values to the equalized gray values, form preprocessed sonar image patches, and combine them to form a set of preprocessed sonar image patches .

[0009] Optionally, the S2 includes the following steps: S21. Calculate the local gray gradient amplitude map of each pixel point in the preprocessed sonar image patch , and the local gray gradient amplitude map is used to characterize the gray change degree of the preprocessed sonar image patch in the horizontal and vertical directions; S22. Calculate the shadow direction response map of each pixel point in the preprocessed sonar image patch , and the shadow direction response map is used to characterize the sonar shadow significance of the pixel point in multiple directions; S23. Perform slope calculation on the historical digital elevation model data corresponding to the spatial position of the preprocessed sonar image patch , define the slope gradient components of each terrain grid unit as and , calculate the slope amplitude map , and the slope amplitude map is used to measure the landform slope size of the terrain at each pixel point; S24. Perform normalization processing on the local gray gradient amplitude map, shadow direction response map, and slope amplitude map. After normalization, the normalized gray gradient map , normalized shadow response map and normalized slope amplitude map are obtained respectively; S25. Construct a pixel-level fusion function, and perform weighted fusion on the normalized gray gradient map, normalized shadow response map, and normalized slope amplitude map according to preset weight coefficients to generate a prior map of the seabed geomorphic boundary ; S26. Combine the prior maps of the seabed geomorphic boundaries of all preprocessed sonar image patches to form a set of prior maps of the seabed geomorphic boundaries .

[0010] Optionally, the S3 includes the following steps: S31. Construct a multi-scale sonar feature encoder, with the input being the set of preprocessed sonar image patches and the set of prior maps of the seabed geomorphic boundaries , at each scale, the multi-scale sonar feature encoder uses a parallel stack of depthwise separable convolutions and band-preserving convolutions to generate an encoded feature map ; S32. At each scale, after the encoded feature map , a geomorphic attention unit is concatenated. The geomorphic attention unit constructs a geomorphic feature vector from the global average mapping of the normalized gray gradient map and the normalized slope amplitude map, performs a linear transformation on the geomorphic feature vector through a learnable mapping matrix, and normalizes and constrains the probability distribution to obtain a geomorphic attention weight map ; S33. Construct a semantic branch, perform geomorphic attention weight modulation on the encoded feature map at each scale, and obtain a modulated semantic feature map by performing element-wise multiplication of each pixel in the encoded feature map with the geomorphic attention weight at the corresponding position , which is used to highlight key regions containing sediment boundaries, fracture lines, or artificial structures; S34. Construct a boundary branch. The boundary branch takes the corresponding prior map of the seabed geomorphic boundary as input, performs a multi-scale dilated convolution operation to generate a preliminary boundary initial feature map, and performs a local contrast enhancement operation on the preliminary boundary initial feature map. The local contrast enhancement operation obtains a locally enhanced boundary initial feature map of the boundary response by applying large-scale average pooling to the prior map of the boundary and subtracting the result of the dilated convolution .

[0011] Optionally, the S4 includes the following steps: S41. Construct a geomorphic mutual feedback gating module in the skip connection layers at each scale. The geomorphic mutual feedback gating module jointly analyzes the modulated semantic feature map and the boundary initial feature map, respectively extracts features through independent learnable convolution kernels and then sums them up, and generates a pixel-level mutual feedback coefficient after being processed by the Sigmoid normalization function ; S42. Perform weighted fusion on the modulated semantic feature map and the boundary initial feature map at each scale with the pixel-level mutual feedback coefficient as the weight to generate a fused feature map .

[0012] Optionally, the S5 includes the following steps: S51. Construct a training dataset, which is jointly composed of a preprocessed sonar image patch set, a corresponding prior map set of the seabed geomorphic boundary, and a per-pixel class label map set , where represents the semantic class label to which each pixel in the preprocessed sonar image patch belongs, and the class labels cover semantic items of seabed sediments, bedrock structures, artificial structures, and background regions; S52. Define a multi-component joint loss function As the objective function for supervised training, the multi-component joint loss function is calculated by weighted summation of the categorical cross-entropy loss , the boundary region Dice loss and the topological consistency preservation loss ; S53. Use the multi-component joint loss function to perform end-to-end supervised training on the constructed improved U-Net dual-branch semantic segmentation network. During the training process, simultaneously optimize the parameter weights of the semantic branch and the boundary branch until the improved U-Net dual-branch semantic segmentation network converges on the validation set to obtain a trained segmentation model; S54. Optimize the structure and compress the model of the trained segmentation model. Replace the redundant convolution modules with depthwise separable convolutions and fuse the boundary guidance weight maps to form a lightweight network structure. Deploy the segmentation model to the on-board graphics processing unit of an autonomous underwater vehicle or a remotely operated underwater platform to form an on-board inference model.

[0013] Optionally, the categorical cross-entropy loss is used to measure the classification consistency across the entire image between the final predicted per-pixel class label map output by the semantic branch and the actual per-pixel class label map : ; where represents the height of each preprocessed sonar image patch, represents the width of each preprocessed sonar image patch, represents the total number of classes in the segmentation task, including seabed sediment, bedrock structure, artificial structure, and background classes, represents the horizontal coordinate index of the current pixel, represents the vertical coordinate index of the current pixel, represents the semantic class index, represents that the pixel at position belongs to the ground truth label of the -th class, represents the predicted probability value of the improved U-Net dual-branch semantic segmentation network for the -th class at position .

[0014] Optionally, the boundary region Dice loss is used to supervise the accuracy of the boundary branch in the geomorphic boundary region. The boundary region is extracted from the seabed geomorphic boundary prior map using a binary boundary mask map: ; where represents the boundary prediction map output by the improved U-Net dual-branch semantic segmentation network, represents the binary boundary mask map, which is obtained by extracting the prior map of the seabed geomorphic boundary according to a preset threshold, represents the sum of the pixel-level products of the binary boundary mask map and the predicted boundary map, represents the number of all pixels belonging to the boundary region in the binary boundary mask map, represents the sum of the confidence levels of all pixels in the predicted boundary map.

[0015] Optionally, the topological consistency preservation loss is used to constrain the consistency of the spatial topological relationship between the output contour structure of the improved U-Net dual-branch semantic segmentation network and the real contour. The topological consistency preservation loss is obtained by separately calculating the Euler number of each category region output by the improved U-Net dual-branch semantic segmentation network and the Euler number of the corresponding category region in the real label map, and taking the average value after calculating the absolute value of the difference between the two Euler numbers.

[0016] Optionally, S6 includes the following steps: S61. The on-board inference model is docked with the sonar data acquisition system of the seabed detection device in real time; S62. Use the on-board inference model to perform batch or streaming input on the real-time collected and pre-normalized and pre-processed sonar image block sets, and execute the online semantic segmentation inference process. For each pre-processed sonar image block, parallel inference outputs the corresponding semantic category map and boundary vector map. Among them, the semantic category map is the pixel-level category distribution result with the same spatial size as the input sonar image block, and the boundary vector map is a multi-channel feature map encoding the boundary response intensity and direction of each pixel.

[0017] The beneficial effects of the present invention are: (1) The present invention explicitly introduces the prior map of the seabed geomorphic boundary into the network structure, utilizes multi-source physical information such as normalized gray gradient, shadow direction and slope amplitude, and realizes the joint modeling of physical geomorphic features and depth semantic features through the fusion of the geomorphic attention unit and the multi-scale feature encoder. It can not only actively focus on the key slope breaks, faults, and complex regions of fold-back shadows in the sonar image, but also accurately suppress the false contours caused by noise and speckle interference, enabling the model to obtain sub-meter-level segmentation resolution at the geomorphic boundary and significantly improving the spatial accuracy of geomorphic elements.

[0018] (2) The present invention designs a geomorphic feedback gating module, which realizes the pixel-level weight dynamic fusion of the modulation semantic feature map and the boundary initial feature map. Through the mutual feedback learning of multi-scale geomorphic responses, the network can adaptively adjust the attention to the semantic category region and the boundary region, taking into account both global category discrimination and fine-grained boundary recognition. It is applicable to the accurate extraction of thin layers of submarine sediments, debris flow fans, and weak boundary targets of artificial structures. In the scenario of small targets and mixed multi-categories, the average mIoU index of the model has been improved.

[0019] (3) The present invention adopts depthwise separable convolution, band-preserving convolution, and lightweight design of the model structure compression, enabling the trained segmentation model to run in real time on the embedded platform of an autonomous underwater vehicle at an inference speed of 8 - 12 fps. At the same time, combined with the boundary prior map and the joint loss optimization strategy, high-quality training can be completed using a very small amount of pixel-level annotations or only rough-annotated boundaries, significantly reducing the annotation cost and adaptation threshold of submarine sonar data. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings: Figure 1 is a flowchart of a method for processing submarine sonar images based on semantic segmentation proposed by the present invention; Figure 2 is a block diagram of an improved U-Net dual-branch semantic segmentation network structure in a method for processing submarine sonar images based on semantic segmentation proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0022] Refer to Figure 1 - Figure 2 , a method for processing submarine sonar images based on semantic segmentation, includes the following steps: S1. Collect the original submarine sonar image data, cut and standardize it according to the unified pixel resolution and file format to obtain the standardized sonar image blocks. Preprocess the standardized sonar image blocks, use adaptive spatial domain filtering to remove speckle noise, perform texture-preserving enhancement, and gray-level dynamic range equalization processing to obtain the preprocessed sonar image blocks; S2. Calculate the terrain slope features based on the local gray-level gradient, sonar shadow direction information, and historical digital elevation model data of the preprocessed sonar image blocks, fuse the features, and generate a submarine geomorphic boundary prior map corresponding to each pixel of the preprocessed sonar image block; S3. Construct an improved U-Net dual-branch semantic segmentation network. The improved U-Net dual-branch semantic segmentation network includes a semantic branch and a boundary branch. The semantic branch takes the preprocessed sonar image patches as input to extract the features of seabed sediments, bedrock structures, and artificial structures. The boundary branch takes the seabed geomorphic boundary prior map as input to extract boundary saliency features; S4. Configure a geomorphic feedback gating module in each scale skip connection layer. Embed the seabed geomorphic boundary prior map into the multi-scale skip connection structure of the improved U-Net dual-branch semantic segmentation network through the geomorphic feedback gating module, and perform weighted fusion on the semantic feature maps output by the semantic branch to obtain fused feature maps; S5. Use the preprocessed sonar image patches, the seabed geomorphic boundary prior map, and the corresponding pixel-level class labels as training data. Adopt a joint loss function that includes categorical cross-entropy loss, boundary region Dice loss, and topological consistency loss to perform supervised training on the improved U-Net dual-branch semantic segmentation network, obtain a trained segmentation model, and deploy it to an autonomous underwater vehicle to form an on-board inference model; S6. Use the on-board inference model to perform online semantic segmentation inference on the preprocessed sonar image patches obtained in real time, and output a semantic class map and a boundary vector map corresponding to the preprocessed sonar image patches.

[0023] In this embodiment, S1 includes the following steps: S11. Use a seabed detection device to collect original seabed sonar image data. The original seabed sonar image data is a two-dimensional grayscale image. During the collection process, unified image resolution and frequency band parameters are adopted to form an original image set , where represents the th original seabed sonar image, is the total number of collected images; S12. Perform size regularization and image chunking on the original image set . According to the preset chunk size , divide each original image into several non-overlapping or partially overlapping sub-blocks to generate a standardized sonar image patch set , where represents the rd image patch of the th image, and are the width and height of the image patch respectively; S13. Perform an adaptive spatial domain filtering operation on the standardized sonar image patch set . Construct a local response window , and apply a weighted denoising model in each window to obtain the denoised image patch: ; Among them, represents the gray value of the pixel points within the window, represents the weight coefficient based on local variance and structure preservation, is the central pixel value after denoising; S14. Perform texture-preserving enhancement processing on each denoised image block to construct a set of structure response operators in multiple directions , where the set of structure response operators includes horizontal, vertical, and two diagonal directions, corresponding to direction angles respectively. Extract the gradient response features of the image block in each direction, generate an image block with enhanced texture structure by comparing the response intensities in different directions and selecting the maximum response direction for enhancement; S15. Perform gray-level dynamic range equalization processing on the image block with enhanced texture structure, adopt a gray-level transformation method based on local histogram, construct a gray-level mapping function of pixel values by calculating the frequencies of each gray value appearing in the entire image block, map the original pixel gray value to the equalized gray value, form a preprocessed sonar image block, and combine them to form a set of preprocessed sonar image blocks .

[0024] The gray-level mapping function is constructed as follows: for each normalized sonar image block, count the frequencies of pixels at each gray level within the image block to form a local gray-level histogram, remap each gray level according to the cumulative distribution function. The expression of the mapping function is: linearly map the original pixel gray value according to its cumulative probability in this image block to the target gray-level interval to achieve the redistribution of gray levels, enhance the local contrast and compress the noise interval. The output gray value of each pixel point is equal to its cumulative probability multiplied by the upper limit value of the gray level, and then rounded down to obtain an integer gray level.

[0025] In this embodiment, S2 includes the following steps: S21. Calculate the local gray-level gradient magnitude map of each pixel point in the preprocessed sonar image block , where the local gray-level gradient magnitude map is used to characterize the gray-level change degree of the preprocessed sonar image block in the horizontal and vertical directions. The local gray-level gradient magnitude map is obtained by separately calculating the gradient components of the image in the horizontal direction and the vertical direction, and then taking the square root of the sum of the squares of the two gradient components; S22. Calculate the shadow direction response map of each pixel point in the preprocessed sonar image block , the shadow direction response map is used to characterize the sonar shadow significance of pixel points in multiple directions. The shadow direction response map applies linear direction enhancement filters in multiple preset directions to the preprocessed sonar image block, calculates the direction response value of pixel points for each direction, and selects the maximum value from all direction response values as the response intensity of the corresponding pixel point in the shadow direction response map; S23. For the historical digital elevation model data corresponding to the spatial position of the preprocessed sonar image block perform slope calculation, and define the slope gradient component of each terrain grid unit as and , calculate the slope magnitude map , the slope magnitude map is used to measure the topographic slope size at each pixel point: ; S24. Normalize the local gray gradient magnitude map, the shadow direction response map, and the slope magnitude map. After normalization, the normalized gray gradient map , the normalized shadow response map and the normalized slope magnitude map are obtained; S25. Construct a pixel-level fusion function, and perform weighted fusion on the normalized gray gradient map, the normalized shadow response map, and the normalized slope magnitude map according to preset weight coefficients to generate a prior map of the seabed geomorphic boundary , which is used to represent the significance of each pixel point belonging to the geomorphic boundary. During the fusion process, the boundary prior value of each pixel point is obtained by summing the product of the normalized gray gradient value, the normalized shadow response value, and the normalized slope magnitude multiplied by the corresponding fusion weight coefficient respectively; The pixel-level fusion function normalizes the three types of feature maps of gray gradient, shadow response, and slope magnitude to the [0,1] interval for each preprocessed sonar image block, assigns independent fusion weights to each feature map at the pixel level, and the fusion weights can be obtained according to scene requirements, empirical knowledge, or through prior data statistics. For each pixel point in the image, the fusion function calculates the product of the three normalized feature values and the corresponding weights respectively, and sums the three to obtain the final pixel-level prior value.

[0026] S26. Combine the prior maps of the seabed geomorphic boundary of all preprocessed sonar image blocks to form a prior map set of the seabed geomorphic boundary , and each prior map of the seabed geomorphic boundary corresponds one-to-one with the corresponding preprocessed sonar image block in the spatial dimension.

[0027] In this embodiment, S3 includes the following steps: S31. Construct a multi-scale sonar feature encoder, and the input is the set of preprocessed sonar image blocks And the prior atlas of submarine geomorphic boundaries At each scale, the multi-scale sonar feature encoder uses a parallel stack of depthwise separable convolutions and band-preserving convolutions to generate an encoded feature map Among them, the sampling rate of the band-preserving convolution is consistent with the transmission frequency of the sonar signal, which is used to prevent the high-frequency detail features in the sonar image from being smoothed or misclipped during the convolution process, so as to retain the structural integrity of the sonar echo texture; S32. At each scale, after the encoded feature map A geomorphic attention unit is concatenated. The geomorphic attention unit constructs a geomorphic feature vector from the global average mapping of the normalized gray gradient map and the normalized slope amplitude map, and performs a linear transformation on the geomorphic feature vector through a learnable mapping matrix and normalizes and constrains the probability distribution to obtain a geomorphic attention weight map The geomorphic attention weight map is used to indicate which regions are more likely to be geomorphic boundary or edge regions at the current scale; S33. Construct a semantic branch, modulate the encoded feature map with the geomorphic attention weight at each scale, and obtain a modulated semantic feature map by element-wise multiplying each pixel in the encoded feature map with the geomorphic attention weight at the corresponding position It is used to highlight key regions containing sediment boundaries, fracture lines or artificial structures, and suppress noise regions or low-response regions, so that the regions concerned by the semantic branch are more concentrated on the pixel distribution with geomorphic structural significance; S34. Construct a boundary branch. The boundary branch takes the corresponding prior map of submarine geomorphic boundaries as input, performs a multi-scale dilated convolution operation to generate a preliminary boundary initial feature map, and performs a local contrast enhancement operation on the preliminary boundary initial feature map. The local contrast enhancement operation is obtained by applying large-scale average pooling to the boundary prior map and subtracting the result of the dilated convolution, and obtains a boundary initial feature map with local enhancement of the boundary response The local contrast enhancement operation is used to improve the gray difference between the weak boundary region in the image block and its neighborhood, so as to enhance the boundary information. The boundary initial feature map is used to describe the texture and response features of the suspected boundary regions in the image; ; Among them, is the dilated convolution operator, and the dilation rate corresponds to the scale , is the window size of the average pooling.

[0028] In this embodiment, S4 includes the following steps: S41. Construct a geomorphic interaction gating module in each scale skip connection layer. The geomorphic interaction gating module jointly analyzes the modulated semantic feature map and the boundary initial feature map, extracts features through independent learnable convolutional kernels respectively and then sums them up, and generates pixel-level interaction coefficients after being processed by the Sigmoid normalization function. , and the pixel-level interaction coefficients are used to measure the contribution degrees of semantic features and boundary features to the fusion result at each spatial position; The geomorphic interaction gating module is a pixel-level and dynamic fusion mechanism proposed based on the improved U-Net dual-branch semantic segmentation network to meet the requirements of complex expression of multi-source heterogeneous geomorphic features and high-precision boundary discrimination in submarine sonar images. The geomorphic interaction gating module mainly acts on the multi-scale skip connection layer of the network, constructs a cross-branch information feedback and weighted modulation path between the semantic branch and the boundary branch, and realizes the deep coupling and dynamic complementarity of semantic features and boundary features.

[0029] The fusion process of the geomorphic interaction gating module is not just a simple feature splicing or linear superposition, but through convolutional extraction, non-linear normalization and spatial adaptive weighting, fully utilizes the physical saliency and semantic discrimination information in the geomorphic boundary area of the sonar image, realizes dynamic adjustment and enhancement of key pixel responses, significantly improves the segmentation effect in weak boundary, small target and complex terrain areas, and effectively alleviates the problems of category misclassification and contour fracture caused by boundary blur and noise interference in traditional single-branch or static fusion methods.

[0030] S42. Weightedly fuse the modulated semantic feature map and the boundary initial feature map at each scale with the pixel-level interaction coefficient as the weight to generate a fusion feature map . The final value of each pixel in the fusion feature map is the sum of the modulated semantic feature map and the boundary initial feature map multiplied by the pixel-level interaction coefficient and its complementary value respectively; ; In this embodiment, S5 includes the following steps: S51. Construct a training data set, which is jointly composed of a preprocessed sonar image block set, a corresponding submarine geomorphic boundary prior map set, and a per-pixel class label map set , where represents the semantic class label of each pixel in the preprocessed sonar image block, and the class labels cover semantic items of submarine sediments, bedrock structures, artificial structures, and background areas; S52. Define a multi-component joint loss function as the objective function for supervised training. The multi-component joint loss function consists of a categorical cross-entropy loss , a boundary region Dice loss Topological Consistency Preservation Loss Calculated by weighted computation; S53. Use the multi-component joint loss function Perform end-to-end supervised training on the constructed improved U-Net double-branch semantic segmentation network. During the training process, simultaneously optimize the parameter weights of the semantic branch and the boundary branch until the improved U-Net double-branch semantic segmentation network converges on the validation set to obtain the trained segmentation model; S54. Optimize the structure and compress the model of the trained segmentation model. Replace the redundant convolution module with depthwise separable convolution and fuse the boundary guidance weight map to form a lightweight network structure. Deploy the segmentation model to the on-board graphics processing unit of the autonomous underwater vehicle or the remotely operated underwater platform to form an on-board inference model.

[0031] In this embodiment, the categorical cross-entropy loss is used to measure the classification consistency of the final predicted per-pixel categorical label map output by the semantic branch and the actual per-pixel categorical label map over the entire image: ; wherein, represents the height of each preprocessed sonar image patch, represents the width of each preprocessed sonar image patch, represents the total number of categories in the segmentation task, including seabed sediment, bedrock structure, artificial structure, and background categories, represents the horizontal coordinate index of the current pixel, and the value range is , represents the vertical coordinate index of the current pixel, and the value range is , represents the semantic category index, and the value range is , represents the position where the pixel at the position belongs to the true label of the th category. If it belongs to this category, the value is 1; otherwise, it is 0, represents the predicted probability value of the improved U-Net double-branch semantic segmentation network for the th category at the position , and the value range is from 0 to 1.

[0032] In this embodiment, the boundary region Dice loss is used to supervise the accuracy of the boundary branch in the geomorphic boundary region. The boundary region is extracted from the seabed geomorphic boundary prior map through the binary boundary mask map: ; Among them, represents the boundary prediction map output by the improved U-Net dual-branch semantic segmentation network, represents the binary boundary mask map, which is obtained by extracting the seafloor geomorphic boundary prior map according to a preset threshold. Each pixel value in is 1 indicating belonging to the boundary area and 0 indicating non-boundary area. represents the sum of the pixel-level products of the binary boundary mask map and the predicted boundary map, indicating the total overlapping area between the predicted boundary of the model and the true boundary. represents the number of all pixels belonging to the boundary area in the binary boundary mask map. represents the sum of the confidence levels of all pixels in the predicted boundary map.

[0033] In this embodiment, the topological consistency preservation loss is used to constrain the consistency of the spatial topological relationship between the output contour structure of the improved U-Net dual-branch semantic segmentation network and the true contour. The topological consistency preservation loss is obtained by separately calculating the Euler number of each category area output by the improved U-Net dual-branch semantic segmentation network and the Euler number of the corresponding category area in the true label map, and taking the average value after calculating the absolute value of the difference between the two Euler numbers. The Euler number is used to measure the number of connected domains, the number of holes, and the integrity of the topological structure of each category area. The topological consistency preservation loss takes the average value of the Euler number differences of all categories as the final loss value, and the Euler number calculation method remains consistent within each category area.

[0034] In this embodiment, S6 includes the following steps: S61. The on-board inference model is docked with the sonar data acquisition system of the seafloor detection device in real time; S62. Use the on-board inference model to perform batch or streaming input on the real-time collected and standardized and preprocessed sonar image block sets, and execute the online semantic segmentation inference process. For each preprocessed sonar image block, parallelly infer and output the corresponding semantic category map and boundary vector map. Among them, the semantic category map is the pixel-level category distribution result with the same spatial size as the input sonar image block, and the boundary vector map is a multi-channel feature map encoding the boundary response intensity and direction of each pixel. S63. Perform real-time post-processing on the online semantic segmentation inference results. The semantic category map obtains the final category label through pixel-level probability thresholding. The category labels include seafloor sediment area, bedrock structure area, artificial structure area, and background area: The category label of the seafloor sediment area corresponds to the pixel points where the maximum category probability belongs to the sediment category and the response intensity of the pixel in the seafloor geomorphic boundary prior map is low and the gray level is uniform. The bedrock structure area category label corresponds to the pixel point with the maximum category probability belonging to the bedrock structure category and the pixel is in the local extreme value area of the slope amplitude. The artificial structure region category label corresponds to the pixel point with the maximum category probability belonging to the artificial structure category and the pixel belongs to the regular texture region or has a significant boundary enhancement response. The background area category label corresponds to the pixel points whose maximum probability of all categories is lower than the preset confidence threshold or does not conform to other category structures or landform features; The boundary vector map extracts the boundary pixel set through non-maximum suppression and threshold discrimination to obtain the spatial distribution information of the landform boundary. The boundary type in the boundary vector map maintains a corresponding relationship with the structural consistency of the semantic category map: The sediment boundary type corresponds to the semantic category map where the boundary pixel has a seabed sediment area category label on one side and a background area category label or other category label on the other side, and the response value of the boundary in the seabed landform boundary prior map is in the medium-low range; The structural boundary type corresponds to the boundary pixel connection area being in the high slope response area or the gradient extreme value area, and the border position of the bedrock structure area category label and the seabed sediment area category label or the artificial structure area category label is distinguished in the semantic category map; The artificial boundary type corresponds to the boundary presenting a long straight line or a regular closed shape, and overlaps with the high confidence area of the artificial structure area category label, and the boundary response is continuous and the boundary closure degree is high in the local structure contrast enhancement result.

[0035] Example 1: A marine survey team carried out a mission from April to June 2024. The water depth in the operating area was 400 to 1100 meters, and the sea area was about 130 square kilometers, covering mud sedimentary fans, fault zones and artificial pipeline areas. The survey team used a certain type of multi-beam sonar (operating frequency 300kHz), synchronous side-scan sonar equipment (resolution 0.25 meters / pixel) and an underwater high-performance GPU (NVIDIA Jetson AGX Orin, 32GB RAM) on the AUV platform to acquire and process large-area underwater landform data in real time.

[0036] During the data collection phase, the AUVs were deployed at equidistant routes at a speed of 3 knots, and the operation time of a single voyage was about 12 hours. About 11,000 original sonar images were collected each time, and the size of a single original image was 2048×4096 pixels. The original data had serious sonar shadows, speckle noise, and echo bands. The traditional U-Net, SEAUNet, FPUA-UNet, and the segmentation method based on the improved U-Net and seabed landform boundary enhancement proposed in this invention were used as comparison objects, and the sonar images collected in the same batch were used for model training and testing.

[0037] In the data preprocessing stage, the resolution of all original sonar images is unified to 0.25 meters per pixel. Adaptive spatial filtering is used to remove speckle noise, and the original images are cut into small blocks of 512×512 pixels. A total of 4,800 standardized training samples are collected and constructed for the muddy sediment area and the fault zone area respectively. Information on local slope, gray gradient, and main shadow direction is automatically extracted from historical survey data to generate a prior map of the geomorphic boundary, and the total number of prior map samples is the same as that of the image blocks.

[0038] In the model training stage, comparative experiments are carried out using the method proposed in the present invention and three traditional segmentation methods respectively. The number of samples in each model training set is 4,000 blocks, and the number of samples in the test set is 800 blocks. The pixel-level categories and boundaries of each image block are manually labeled. All models are trained and tested on the same hardware platform (NVIDIA Jetson AGX Orin). The Adam optimizer is used in the training process, with an initial learning rate of 0.001 and a maximum of 50 iterations. The average time per iteration is 40 minutes.

[0039] Applying the method of the present invention, the input of the segmentation network is the preprocessed sonar image block, the prior map of the geomorphic boundary, and the pixel-level annotation. The backbone adopts a multi-scale feature extraction structure of depthwise separable convolution and band-preserving convolution, and shows the fusion of slope, gradient, and shadow priors. The semantic branch focuses on the extraction of sediment, bedrock, and artificial target areas. The boundary branch uses the geomorphic prior to dynamically guide the segmentation contour. The multi-scale geomorphic mutual feedback gating module performs pixel-level weighted fusion on the semantic and boundary responses to achieve refined contour expression. The model training uses the joint optimization of categorical cross-entropy, boundary Dice, and topological consistency loss. Finally, the inference model is compressed and deployed on the AUV on-board GPU to achieve full-process automatic inference and real-time output.

[0040] In actual tests, among the 800 test image blocks collected, typical segmentation scenarios include sediment fan transition zones, fault boundary areas, artificially laid pipeline areas, and debris flow fan areas. Based on the manual annotation, the performance of each method in terms of mIoU (mean intersection over union), boundary pixel accuracy (Boundary F1), boundary Hausdorff distance (HD, unit: pixel), and inference speed (fps) is shown in Table 1 below: Table 1 Data comparison between the present invention and three traditional segmentation methods

[0041] In addition, in terms of feature retention and small target detection, the method of the present invention has significantly improved the detection rate of fracture zones and slender pipeline structures. Compared with manual annotation, the average small target recall rate has increased from 62.8% of the traditional U-Net to 80.4% of the present invention, and the segmented fault line boundaries are continuous without obvious transition zones or fracture artifacts. In the area of artificial structures (such as laid pipelines), the end-to-end connectivity rate (the coincidence degree of the connected length and the real structure) detected by the method of the present invention has increased to 93.6%, which is more than 10% higher than that of SEAUNet.

[0042] Regarding the expression of geomorphic boundaries in a high-noise environment, the method of the present invention has achieved a pixel-level accuracy of 89.3% in the fault and slope-break shadow areas, which is 13.4% higher than that of FPUA-UNet. In specific shadow-overlapping samples, the number of boundary artifacts has been reduced by more than 40%, and the average number of pixels that need to be manually post-processed and repaired in all test samples has been reduced to 110 pixels / block (the traditional U-Net needs to repair 275 pixels / block).

[0043] In terms of inference speed, the real-time inference frame rate of the method of the present invention on the board-mounted GPU is stable at 8-12fps, which can meet the requirements of online autonomous recognition, real-time path planning, and multi-target geomorphic tracking of AUV or ROV, and significantly shorten the processing cycle of subsea mapping data.

[0044] Training sample display: Samples in the sediment fan area: The original image has strong noise, and the slope prior shows a zonal transition. The boundary of the sediment area segmented by the present invention fits the manual annotation, and the fault boundary is clear; Samples in the pipeline area: The artificial structure is long and strip-shaped, and there are many discontinuous pseudo-edges in the traditional method. The present invention accurately detects all line segments at the boundary branches and the endpoints are connected; Samples with strong shadow interference: The segmentation result of the traditional method has blurred edges. The model of the present invention significantly suppresses shadow pseudo-edges through geomorphic priors, and the contour is continuous.

[0045] The present invention explicitly introduces the subsea geomorphic boundary prior map into the network structure, utilizes multi-source physical information such as normalized gray gradient, shadow direction, and slope amplitude, and realizes the joint modeling of physical geomorphic features and depth semantic features through the fusion of the geomorphic attention unit and the multi-scale feature encoder. It can not only actively focus on key slope breaks, faults, and complex areas of folding-back shadows in sonar images, but also accurately suppress false contours caused by noise and speckle interference, enabling the model to obtain sub-meter-level segmentation resolution at geomorphic boundaries and significantly improving the spatial accuracy of geomorphic elements.

[0046] The present invention designs a geomorphic interaction gating module, which realizes the pixel-level weight dynamic fusion of the modulated semantic feature map and the boundary initial feature map. Through the interaction learning of multi-scale geomorphic responses, the network can adaptively adjust the attention to the semantic category region and the boundary region, taking into account both global category discrimination and fine-grained boundary recognition. It is applicable to the accurate extraction of thin layers of submarine sediments, debris flow fans, and weak boundary targets of artificial structures. In the scenario of small targets and mixed multi-categories, the average mIoU index of the model has been improved.

[0047] The present invention adopts depthwise separable convolution, band-preserving convolution, and lightweight design of the model structure, enabling the trained segmentation model to run in real time at an inference speed of 8–12 fps on an autonomous underwater vehicle embedded platform. At the same time, combined with the boundary prior map and the joint loss optimization strategy, high-quality training can be completed with a very small amount of pixel-level annotations or only rough-annotated boundaries, significantly reducing the annotation cost and adaptation threshold of submarine sonar data.

[0048] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. An underwater sonar image processing method based on semantic segmentation, characterized in that The steps include: S1. Collecting original seabed sonar image data for standardization, preprocessing the standardized sonar image blocks to obtain preprocessed sonar image blocks; S2. Calculate the terrain slope features based on the preprocessed sonar image blocks, fuse the features, and generate a priori map of the seabed landform boundary corresponding to the preprocessed sonar image blocks pixel by pixel; S3. Construct an improved U-Net dual-branch semantic segmentation network, which includes a semantic branch and a boundary branch. S4. Configure a geomorphic mutual feedback gating module at each scale skip layer, embed the seabed geomorphic boundary prior map into the multi-scale skip structure of the improved U-Net dual-branch semantic segmentation network through the geomorphic mutual feedback gating module, and perform weighted fusion on the semantic feature map output by the semantic branch to obtain a fused feature map; S5. Using the pre-processed sonar image blocks, the seafloor topography boundary prior map and the corresponding pixel-level category labels as training data, the improved U-Net dual-branch semantic segmentation network is supervised trained using a joint loss function to obtain a trained segmentation model, which is then deployed to the autonomous underwater vehicle to form an onboard inference model; S6. Perform online semantic segmentation inference on the pre-processed sonar image blocks acquired in real time using the onboard inference model, and output a semantic category map and a boundary vector map corresponding to the pre-processed sonar image blocks.

2. The method for processing underwater sonar images based on semantic segmentation according to claim 1, wherein The S1 comprises the following steps: S11. Collect the original underwater sonar image data using underwater detection equipment to form an original image set , where represents the th original underwater sonar image, is the total number of collected images; S12. For the set of original images perform size regularization and image chunking processing, and according to the preset chunk size divide each original image into a number of non-overlapping or partially overlapping sub-blocks to generate a set of standardized sonar image blocks; S13. performing an adaptive spatial filtering operation on the standardized sonar image block set, constructing a local response window, applying a weighted denoising model in each window, and obtaining a denoised image block; S14. Perform texture-preserving enhancement processing on each denoised image patch to construct a set of structure response operators in multiple directions , where the set of structure response operators includes horizontal, vertical, and two diagonal directions, corresponding to direction angles respectively. Extract the gradient response features of the image patch in each direction, enhance by comparing the response intensities in different directions and selecting the maximum response direction, and generate an image patch with enhanced texture structure; S15. Perform gray - scale dynamic range equalization processing on the image block with enhanced texture structure, map the original pixel gray - scale value to the equalized gray - scale value, form a pre - processed sonar image block, and combine them to form a set of pre - processed sonar image blocks .

3. The method for processing submarine sonar images based on semantic segmentation according to claim 2, wherein, The S2 comprises the following steps: S21. Calculate the local gray-scale gradient magnitude map of each pixel in the preprocessed sonar image block , where the local gray-scale gradient magnitude map is used to characterize the gray-scale change degree of the preprocessed sonar image block in the horizontal and vertical directions; S22. Calculate the shadow direction response map for each pixel in the preprocessed sonar image block , where the shadow direction response map is used to characterize the sonar shadow saliency of the pixel in multiple directions; S23. Calculate the slope of the historical digital elevation model data corresponding to the spatial position of the preprocessed sonar image block Define the slope gradient components of each topographic grid cell as and Calculate the slope amplitude map The slope amplitude map is used to measure the topographic slope at each pixel point; S24. Normalize the local gray-scale gradient magnitude map, the shadow direction response map, and the slope magnitude map, and respectively obtain a normalized gray-scale gradient map , a normalized shadow response map , and a normalized slope magnitude map ; S25. Construct a pixel-level fusion function to perform weighted fusion on the normalized gray gradient map, the normalized shadow effect map, and the normalized slope amplitude map according to preset weight coefficients to generate a prior map of the seabed geomorphic boundary ; S26. Combine the prior maps of the seabed geomorphic boundaries of all preprocessed sonar image patches to form a prior map set of the seabed geomorphic boundaries .

4. A method for processing submarine sonar images based on semantic segmentation according to claim 3, characterized in that, The S3 comprises the following steps: S31. Construct a multi-scale sonar feature encoder, with the input being the preprocessed sonar image block set and the prior atlas of submarine geomorphic boundaries , at each scale, the multi-scale sonar feature encoder uses parallel stacking of depthwise separable convolution and band-preserving convolution to generate encoded feature maps ; S32. At each scale, after encoding the feature map a geomorphic attention unit is concatenated. The geomorphic attention unit constructs a geomorphic feature vector from the global average mapping of the normalized gray gradient map and the normalized slope amplitude map, linearly transforms the geomorphic feature vector through a learnable mapping matrix, normalizes it, and applies a probability distribution constraint to obtain a geomorphic attention weight map ; S33. Construct a semantic branch, perform geomorphic attention weight modulation on the encoded feature map at each scale, and obtain the modulated semantic feature map by element-wise multiplying each pixel point in the encoded feature map with the geomorphic attention weight at the corresponding position , which is used to highlight key regions containing sediment boundaries, fracture lines or artificial structures; S34. Construct a boundary branch. The boundary branch takes the corresponding prior map of the seabed geomorphic boundary as input, performs multi-scale dilated convolution operations to generate a preliminary boundary initial feature map, and performs a local contrast enhancement operation on the preliminary boundary initial feature map. The local contrast enhancement operation obtains a locally enhanced boundary initial feature map of the boundary response by applying large-scale average pooling to the boundary prior map and subtracting the result of the dilated convolution .

5. A method for processing underwater sonar images based on semantic segmentation according to claim 4, characterized in that, The S4 comprises the following steps: S41. Build a geomorphic interaction gating module in each scale skip connection layer. The geomorphic interaction gating module jointly analyzes the modulated semantic feature map and the boundary initial feature map, extracts features through independent learnable convolutional kernels respectively, sums them up, and generates pixel-level interaction coefficients after being processed by the Sigmoid normalization function. ; S42. For the modulated semantic feature map and the initial boundary feature map at each scale, generate a fused feature map by weighted fusion with the pixel-level mutual feedback coefficient as the weight .

6. A method for processing underwater sonar images based on semantic segmentation according to claim 5, characterized in that, The S5 comprises the following steps: S51. Construct a training data set, which is composed of a preprocessed sonar image patch set, a corresponding prior atlas of seabed geomorphic boundaries, and a pixel-by-pixel class label atlas collectively, where represents the semantic class label to which each pixel in the preprocessed sonar image patch belongs, and the class labels cover semantic items of seabed sediments, bedrock structures, artificial structures, and background areas; S52. Define the multi-component joint loss function As the objective function for supervised training, the multi-component joint loss function is calculated by weighted combination of the categorical cross-entropy loss , the boundary region Dice loss and the topological consistency preservation loss ; S53. Using a multi-component joint loss function Perform end-to-end supervised training on the constructed improved U-Net double-branch semantic segmentation network. During the training process, simultaneously optimize the parameter weights of the semantic branch and the boundary branch until the improved U-Net double-branch semantic segmentation network converges on the validation set to obtain a trained segmentation model; S54. Perform structural optimization and model compression on the trained segmentation model, use depthwise separable convolution to replace redundant convolution modules, and fuse the boundary-guided weight map to form a lightweight network structure. Deploy the segmentation model to the onboard graphics processing unit of the autonomous underwater vehicle or remote-controlled underwater platform to form an onboard inference model.

7. A method for processing submarine sonar images based on semantic segmentation according to claim 6, characterized in that The cross-entropy loss of the category is used to measure the per-pixel category label map of the final prediction output by the semantic branch and the actual per-pixel category label map for the classification consistency across the entire image: ; Among them, represents the height of each preprocessed sonar image block, represents the width of each preprocessed sonar image block, represents the total number of categories in the segmentation task, including seafloor sediments, bedrock structures, artificial structures, and background categories, represents the horizontal coordinate index of the current pixel, represents the vertical coordinate index of the current pixel, represents the semantic category index, represents the position where the pixel at this position belongs to the ground truth label of the th class, represents the predicted probability value of the improved U-Net double-branch semantic segmentation network for the th class at the position ​ 8. A method for processing underwater sonar images based on semantic segmentation according to claim 6, characterized in that, The boundary region Dice loss used to supervise the accuracy of the boundary branch in the geomorphic boundary region, and the boundary region is extracted from the prior map of the seabed geomorphic boundary through a binary boundary mask map as follows: ; Among them, represents the boundary prediction map output by the improved U-Net double-branch semantic segmentation network, represents the binary boundary mask map, which is obtained by extracting the prior map of the seabed geomorphic boundary according to a preset threshold, represents the sum of the pixel-level products of the binary boundary mask map and the predicted boundary map, represents the number of all pixels belonging to the boundary region in the binary boundary mask map, represents the sum of the confidence levels of all pixels in the predicted boundary map.

9. A method for processing underwater sonar images based on semantic segmentation according to claim 6, characterized in that, The topological consistency preservation loss is used to constrain the consistency of the spatial topological relationship between the output contour structure of the improved U-Net double-branch semantic segmentation network and the real contour. The topological consistency preservation loss is obtained by separately calculating the Euler number of each category region output by the improved U-Net double-branch semantic segmentation network and the Euler number of the corresponding category region in the real label map, and taking the average value after calculating the absolute value of the difference between the two Euler numbers.

10. A method for processing underwater sonar images based on semantic segmentation according to claim 6, characterized in that, The S6 comprises the following steps: S61. Real-time connection between the onboard reasoning model and the sonar data acquisition system of the seabed detection equipment; S62. Use the onboard inference model to batch or stream input the sonar image block sets collected in real time and after standardization and preprocessing, and perform the online semantic segmentation inference process. For each preprocessed sonar image block, parallel inference outputs the corresponding semantic category map and boundary vector map, where the semantic category map is the pixel-level category distribution result consistent with the spatial size of the input sonar image block, and the boundary vector map is a multi-channel feature map that encodes the strength and direction of each pixel boundary response.

Citation Information

Patent Citations

  • Sonar-based time-space association map real-time construction method

    CN113052940A

  • Submarine cable detection tracking method based on sonar image semantic segmentation

    CN116597141A

  • Drainage basin division method based on deep learning semantic segmentation

    CN116883651A

  • Temporal semantic boundary loss for video semantic segmentation networks

    US20240177318A1

Cited By

  • Deep learning-based war wound ultrasonic image diagnosis system and method

    CN121213554A

  • Miniaturized unmanned aerial vehicle mounting type high branch and leaf intelligent sampling method and system, and storage medium

    CN121353943A

  • Object segmentation method and system for sonar image attribute modeling and closed-loop enhancement

    CN121883516A

  • Submarine pipeline side-scan sonar detection method and device and electronic equipment

    CN122072338A

  • Submarine line detection method based on deep neural network and multi-beam sonar data

    CN122473626A