Sonar image segmentation method, system and device, medium and product

By combining the multi-view feature module and the gradient enhancement transformation module, the problem of noise interference and edge feature extraction difficulties in underwater sonar image segmentation is solved, and higher segmentation accuracy and target segmentation effect in complex environments are achieved.

CN120339609APending Publication Date: 2025-07-18SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510384028.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art has problems such as noise interference, image blur, and difficulty in extracting edge features in underwater sonar image segmentation. Especially in complex underwater environments, it is difficult to effectively extract high-frequency gradient information, resulting in poor segmentation effect.

Method used

The multi-view feature module (MPFM) and the gradient enhancement transformation module (GETM) are combined to reduce the computational complexity through parallel attention mechanism and channel enhancement convolution, capture local details and global features, and extract edge features using gradient operations, and combine SAM encoder and decoder for feature fusion to enhance the model's sensitivity to boundaries.

Benefits of technology

It improves the accuracy and segmentation effect of sonar image segmentation, which is significantly better than the existing methods, improves Dice score, and can effectively segment goals in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339609A_ABST
    Figure CN120339609A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and provides a sonar image segmentation method, system and device, a medium and a product. The method comprises the following steps: based on an input feature map, extracting a global feature map and a local region feature map by adopting an MPFM module, and fusing the local region feature map with the global feature map to obtain a first fused feature map; fusing the feature map output by the window attention module with the first fusion feature map to obtain a second fusion feature map; based on the input feature map, an SAM encoder is adopted, and a multi-layer feature map is obtained through an MLP layer; fusing the multi-layer feature map with the second fusion feature map to obtain a third fusion feature map; enhancing vertical gradient information and horizontal gradient information of each channel by adopting a gradient enhancement transformation module based on the input feature map to obtain an edge feature map; fusing the edge feature map and the edge feature map to obtain a fourth fused feature map; and inputting the fourth fusion feature map into a decoder to obtain a target segmented image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning and image processing technology, and in particular to a sonar image segmentation method, system, device, medium and product. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Sonar imaging has become a key technology in the marine sector due to its ability to penetrate water and certain obstacles.

[0004] Existing technologies face many challenges in underwater sonar image segmentation, mainly in terms of noise interference, image blur, and difficulty in edge feature extraction. Due to the complex underwater environment, insufficient lighting, low visibility, and interference from marine organisms, traditional visual methods are difficult to apply. Although sonar images can penetrate water bodies, their inherent high noise, low resolution, and blurred boundaries seriously affect the segmentation effect. Mainstream segmentation networks are difficult to directly apply to sonar images, especially in terms of maintaining high-frequency information. In addition, although large models (such as SAM) have wide applicability, their performance in complex underwater environments is limited due to the lack of prior knowledge in specific fields, and their large number of parameters leads to huge computational overhead for complete fine-tuning and may cause catastrophic forgetting problems.

[0005] The imaging principle of sonar images is different from that of optical systems. The target boundaries are usually blurred and have more noise interference. Existing segmentation models (such as SAM) often generalize features and have difficulty accurately capturing subtle edge details in complex backgrounds. Although some studies have attempted to improve segmentation accuracy through user interaction, iterative optimization and other methods, these methods rely on precise user prompts or high-quality edge feature extraction, and are still difficult to generalize in high-noise, low-contrast scenes. Especially in low-quality images or complex environments, how to effectively extract high-frequency gradient information to enhance boundary clarity is still a key shortcoming of existing methods. Summary of the invention

[0006] To solve the technical problems existing in the above-mentioned background art, the present invention provides a sonar image segmentation method, system, device, medium and product. First, the Multi-Perspective Feature Module (MPFM) uses a parallel attention mechanism and channel-enhanced convolution to reduce the computational complexity, suppress noise, and simultaneously capture local details and global features. Next, the Gradient-Enhanced Transformation Module (GETM) uses gradient-based operations to extract edge features, thereby enhancing the model's sensitivity to boundaries. Finally, the sonar adapter module integrates the specific task knowledge in the MPFM with the general characteristics in the SAM. Through the decoder, the target segmentation image is output. The present invention can improve the accuracy of target segmentation.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] The first aspect of the present invention provides a sonar image segmentation method.

[0009] A sonar image segmentation method includes:

[0010] Based on the acquired underwater sonar image, an input feature map is extracted;

[0011] Based on the input feature map, the MPFM module is used to extract a global feature map and a local region feature map, and the local region feature map is fused with the global feature map to obtain a first fused feature map; the feature map output by the window attention module is fused with the first fused feature map to obtain a second fused feature map;

[0012] Based on the input feature map, the SAM encoder is used, and through the MLP layer, a multi-layer feature map is obtained; the multi-layer feature map is fused with the second fused feature map to obtain a third fused feature map;

[0013] Based on the input feature map, the gradient enhancement transformation module is used to enhance the vertical gradient information and horizontal gradient information of each channel to obtain an edge feature map; the edge feature map is fused with the edge feature map to obtain a fourth fused feature map;

[0014] The fourth fused feature map is input into the decoder to obtain the target segmentation image;

[0015] Among them, the MPFM module, the SAM encoder, the gradient enhancement transformation module and the decoder constitute an image segmentation model.

[0016] Further, based on the input feature map, the MPFM module is used to extract the global feature map. The method includes: based on the input feature map, a feature extraction layer is used to obtain the basic feature information map; based on the basic feature information map, the CoreFormer module is used to generate query vectors, key vectors, and value vectors; based on the query vectors and key vectors, the information in the basic feature information map is compressed into a low-dimensional space by base point projection to obtain base points; based on the base points and query vectors, a first attention map is obtained; based on the base points and key vectors, a second attention map is obtained; based on the first attention map, the second attention map, and the value vectors, the global feature map is obtained.

[0017] Further, based on the input feature map, the MPFM module is used to extract the local region feature map. The method includes: dividing the input feature map into multiple non-overlapping windows, performing attention calculation on each window, extracting the key information within the window, and generating the local region feature map.

[0018] Further, the feature map output by the window attention module is fused with the first fused feature map to obtain the second fused feature map. The method includes: using the window attention module to splice the local region feature maps of all windows along the channel dimension and then fuse them with the first fused feature map to obtain the second fused feature map.

[0019] Further, the multi-layer feature map is fused with the second fused feature map to obtain the third fused feature map. The method includes: integrating a Sonar-Fusion adapter into each Transformer block in the SAM encoder, and in the Sonar-Fusion adapter, fusing the multi-layer feature map with the second fused feature map to obtain the third fused feature map.

[0020] Further, based on the input feature map, the gradient enhancement transformation module is used to enhance the vertical gradient information and horizontal gradient information of each channel to obtain the edge feature map. The method includes: based on the input image features, using the gradient enhancement transformation module to extract the vertical gradient information and horizontal gradient information of each channel of the input feature map respectively, integrating the vertical gradient information and horizontal gradient information into the corresponding channels to obtain the gradient enhancement feature map; encoding the gradient enhancement feature map to obtain the edge feature map.

[0021] The second aspect of the present invention provides a sonar image segmentation system.

[0022] A sonar image segmentation system includes:

[0023] A feature extraction module, which is configured to: based on the acquired underwater sonar image, extract the input feature map;

[0024] The MPFM module is configured to: based on the input feature map, extract the global feature map and the local region feature map, and fuse the local region feature map with the global feature map to obtain the first fused feature map; fuse the feature map output by the window attention module with the first fused feature map to obtain the second fused feature map;

[0025] The SAM encoder module is configured to: based on the input feature map, obtain the multi-layer feature map through the MLP layer; fuse the multi-layer feature map with the second fused feature map to obtain the third fused feature map;

[0026] The gradient enhancement transformation module is configured to: based on the input feature map, enhance the vertical gradient information and the horizontal gradient information of each channel to obtain the edge feature map; fuse the edge feature map with the edge feature map to obtain the fourth fused feature map;

[0027] The decoder module is configured to: based on the fourth fused feature map, obtain the target segmentation image;

[0028] Among them, the MPFM module, the SAM encoder, the gradient enhancement transformation module and the decoder constitute an image segmentation model.

[0029] The third aspect of the present invention provides a computer device, which includes:

[0030] A processor, adapted to execute a computer program;

[0031] A computer-readable storage medium, in which a computer program is stored. When the computer program is executed by the processor, the steps in the sonar image segmentation method described in the first aspect above are implemented.

[0032] The fourth aspect of the present invention provides a computer-readable storage medium, which stores a computer program, and the computer program is adapted to be loaded and executed by the processor to implement the steps in the sonar image segmentation method described in the first aspect above.

[0033] The fifth aspect of the present invention provides a computer program product or a computer program.

[0034] The present invention provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the sonar image segmentation method described in the first aspect above.

[0035] Compared with the prior art, the beneficial effects of the present invention are:

[0036] The present invention provides a sonar image segmentation method, system, device, medium and product, which integrates denoising and multi-scale boundary feature extraction and is dedicated to underwater sonar image segmentation; a new Transformer-based module, the multi-perspective feature module (MPFM), is introduced, which improves the image denoising quality by effectively compressing and summarizing global, regional and local information from a unified perspective; a gradient enhancement transformation module (GETM) is proposed, which significantly enhances the model's ability to extract and utilize sonar edge information; the Sonar-Fusion adapter is incorporated into the Transformer block branch of each SAM layer to promote the proportional fusion of task-specific knowledge from MPFM and the general knowledge of SAM. This strategy enhances the adaptability and segmentation performance of the model on sonar datasets while significantly reducing the number of trainable parameters.

[0037] In view of the complexity of the underwater environment, combined with the problems of the inherently low resolution and noise interference of sonar images, the present invention proposes a new segmentation model GM-SAM that integrates denoising and multi-scale boundary feature extraction. First, the multi-perspective feature module (MPFM) adopts a parallel attention mechanism and channel-enhanced convolution to reduce the computational complexity, suppress noise, and at the same time capture local details and global features. Then, the gradient enhancement transformation module (GETM) uses gradient-based operations to extract edge features, thereby enhancing the model's sensitivity to boundaries. Finally, the sonar fusion adapter module integrates the task-specific knowledge in MPFM with the general characteristics in SAM. Experimental results show that the segmentation effect of GM-SAM is significantly better than existing methods, can obtain superior Dice scores, and improves the accuracy of target segmentation. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0039] Figure 1 is a flowchart of the sonar image segmentation method shown in the embodiments of the present invention;

[0040] Figure 2 is an architecture diagram of the GM-SAM network shown in the embodiments of the present invention;

[0041] Figure 3 is a framework diagram of the multi-perspective feature module shown in the embodiments of the present invention;

[0042] Figure 4 is a framework diagram of the Sonar-Fusion adapter shown in the embodiments of the present invention;

[0043] Figure 5It is a diagram of the gradient enhancement transformation module shown in the embodiments of the present invention;

[0044] Figure 6 It is a comparison diagram of the visualization results of the present invention and other segmentation models on the marine debris dataset shown in the embodiments of the present invention;

[0045] Figure 7 It is a structural diagram of a computer device shown in the embodiments of the present invention. Detailed implementation manners

[0046] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0047] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0048] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless otherwise clearly specified in the context, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0049] As Figure 1 shown, this embodiment provides a sonar image segmentation method, including:

[0050] Based on the acquired underwater sonar image, an input feature map is extracted;

[0051] Based on the input feature map, the MPFM module is used to extract a global feature map and a local region feature map, and the local region feature map is fused with the global feature map to obtain a first fused feature map; the feature map output by the window attention module is fused with the first fused feature map to obtain a second fused feature map;

[0052] Based on the input feature map, the SAM encoder is used, and through the MLP layer, a multi-layer feature map is obtained; the multi-layer feature map is fused with the second fused feature map to obtain a third fused feature map;

[0053] Based on the input feature map, the gradient enhancement transformation module is used to enhance the vertical gradient information and horizontal gradient information of each channel to obtain an edge feature map; the edge feature map is fused with the edge feature map to obtain a fourth fused feature map;

[0054] The fourth fused feature map is input into the decoder to obtain the target segmentation image;

[0055] Among them, the MPFM module, the SAM encoder, the gradient enhancement transformation module, and the decoder constitute the image segmentation model.

[0056] First, the multi-view feature module (MPFM) uses a parallel attention mechanism and channel-enhanced convolution to reduce computational complexity, suppress noise, and capture local details and global features simultaneously. Next, the gradient enhancement transformation module (GETM) uses gradient-based operations to extract edge features, thereby enhancing the model's sensitivity to boundaries. Finally, the sonar fusion adapter module integrates the task-specific knowledge in MPFM with the general characteristics in SAM. The experimental results show that the segmentation effect of GM-SAM is significantly better than existing methods, can obtain a superior Dice score, and effectively segment valid targets.

[0057] As Figure 2 shown, the architecture of the GM-SAM network is designed to enhance the segmentation performance in underwater sonar scenes and support joint training with the sonar fusion adapter. While retaining the advantages of SAM, it extracts edge, gradient, and semantic information, aligns with the original features of SAM to achieve multi-level fusion, and enhances the ability of GM-SAM to segment target objects in complex scenes, noisy, and low-quality images. To this end, three key modules are designed: the multi-view feature module (MPFM), the adapter module, and the gradient enhancement transformation module (GETF).

[0058] Based on the obtained underwater sonar image, an input feature map is extracted;

[0059] In some embodiments, based on the input feature map, a feature extraction layer is used to obtain a basic feature information map; based on the basic feature information map, a CoreFormer module is used to generate query vectors, key vectors, and value vectors; based on the query vectors and key vectors, a base point projection is used to compress the information in the basic feature information map into a low-dimensional space to obtain a base point; based on the base point and the query vectors, a first attention map is obtained; based on the base point and the key vectors, a second attention map is obtained; based on the first attention map, the second attention map, and the value vectors, a global feature map is obtained. Specifically:

[0060] Module 1: The structure of MPFM is as Figure 3 shown. This module first captures the preliminary feature representation of the input image through a feature extraction layer to construct basic feature information. Then, this information will be input into the CoreFormer core module, which is a key component in the Transformer stage. CoreFormer uses a three-layer hierarchical structure to capture unified features, which are composed of global, regional, and local features respectively. This architecture is beneficial for effective denoising while retaining basic semantic details.

[0061] 1) Global feature extraction: Specifically, CoreFormer captures the global features of an image through a global perspective attention mechanism. The global feature extraction equation is as follows:

[0062]

[0063] Among them, q represents Query; d represents the dimension of the Query and Key vectors; F represents the attention map between the Query and Key in Self-Attention; h represents an attention mapping that describes the relationship between the base points; F << N, A ∈ R F×d represents the base point, F w ∈ R F×N and F h ∈ R F×F represent the attention mappings between the query-base point pair and the base point-key pair respectively. N represents the number of pixels in the image feature map. This module first compresses the information in the feature map into a low-dimensional space using the base point projection to obtain the base point A. This base point acts as an intermediary to decompose the original attention map into two compact attention maps F w and F h . By transmitting and fusing between different stripes, the processed feature map G is obtained.

[0064] In some embodiments, the input feature map is divided into multiple non-overlapping windows, and attention calculation is performed on each window to extract the key information within the window and generate a local region feature map. Specifically:

[0065] 2) Region feature extraction: To capture the region features of an image, the window attention mechanism divides the feature map into multiple non-overlapping local windows. The specific details of the division formula are as follows:

[0066]

[0067] Among them, H represents the height of the feature map. Each window independently performs attention calculation to extract the key information from the local region. The attention calculation formula is as follows:

[0068]

[0069] Among them, W Q , W K , W V are learnable projection matrices. The calculated feature information is concatenated with the feature map G along the channel dimension.

[0070] In some embodiments, a window attention module is adopted, and the local region feature maps of all windows are concatenated along the channel dimension and then fused with the first fused feature map to obtain a second fused feature map; specifically:

[0071] 3) Local feature refinement: In addition, to obtain a more refined feature representation, MPFM uses a channel attention enhanced convolution module. This module integrates convolution operations and a channel attention mechanism to improve the feature expression ability. The channel attention mechanism equation is described as follows:

[0072]

[0073] where g(X) represents the global average pooling operation applied to the input feature map X, which compresses it along the channel dimension. W1 and W2 are learnable weight matrices, σ represents the activation function, represents the convolution operation. Subsequently, the vector g(X) is passed through a fully connected layer to calculate the attention weight vector W c , which is used to adjust the weights of the feature map channel by channel to obtain a weighted feature map, and then a convolution operation is performed to obtain the final output feature map Y. Finally, the feature map Y is concatenated along the channel dimension with the feature map generated by the parallel structure to produce the final feature representation. This fusion process enhances the model's ability to suppress noise interference while emphasizing key features.

[0074] In some embodiments, multiple-layer feature maps are fused with the second fused feature map to obtain a third fused feature map; the method includes: integrating a Sonar-Fusion adapter into each Transformer block in the SAM encoder, and in the Sonar-Fusion adapter, fusing the multiple-layer feature maps with the second fused feature map to obtain a third fused feature map; specifically:

[0075] Module 2: Sonar-Fusion adapter is as Figure 4 shown. The present invention integrates a Sonar-Fusion adapter into each Transformer block in the SAM image encoder. This module aims to combine the task-specific knowledge from MPFM with the output of the MLP in a simple and effective way, thereby improving the model performance for the dedicated sonar segmentation task. The Sonar-Fusion adapter is lightweight, with only two linear layers and one non-linear activation function. The process equation is as follows:

[0076] P s = Up(ReLU(Down(F i ))) (5)

[0077] where Up and Down respectively represent upward and downward linear projections. ReLU represents the activation operation, F iDenotes the initial features extracted by the SAM image encoder, P s Denotes the learned features.

[0078] Subsequently, the output of the MPFM module is fused with P in the Sonar-Fusion adapter s This method fuses the global features of SAM with the task-specific knowledge of MPFM and learns them in a certain proportion to form a backbone feature representation that better suits the specific task requirements. The fusion equation is as follows:

[0079] F fusion = σ(α) ⊙ F MLP + σ(β ⊙ F MPFM (6)

[0080] Where σ(·) represents the Sigmoid activation function, which is used to constrain the weights α and β within the interval (0, 1), and ⊙ represents element-wise multiplication. Finally, the fused feature information is returned to the SAM image encoder for the next layer. This iterative refinement strategy enables the Sonar-Fusion Adapter to dynamically inject task-specific knowledge into SAM while maintaining computational efficiency.

[0081] In some embodiments, based on the input image features, a gradient enhancement transformation module is adopted to extract the vertical gradient information and horizontal gradient information of each channel of the input feature map respectively, integrate the vertical gradient information and horizontal gradient information into the corresponding channels to obtain a gradient-enhanced feature map; encode the gradient-enhanced feature map to obtain an edge feature map; specifically:

[0082] Module three: The gradient enhancement transformation module is as Figure 5 shown. Aiming at the problem of blurred boundary information in underwater sonar images, an edge extraction method for underwater sonar images based on edge extraction is proposed. GETM mainly extracts gradient-based edge features from sonar images and interacts with the information extracted by the SAM image encoder to complete depth feature enhancement. The core innovation of GETM lies in its explicit gradient operation design, which directly extracts the gradients in the horizontal and vertical directions and calculates the gradient feature X' α . The formalization of this process is as follows:

[0083]

[0084] Where represent the vertical gradient and horizontal gradient respectively, x i represents the i-th channel of the input feature map, weight v 、weight h represent the vertical convolution kernel and horizontal convolution kernel respectively. ∈ is a very small constant to prevent division by zero errors. X' αIt is passed to a series of ViT encoding blocks for further processing to generate edge features X α Padding indicates the padding method of the convolution operation, which is used to control the distance between the convolution kernel and the edge of the input feature map. Dim represents the dimension used in the concatenation operation. Finally, the result is passed to the SAM image encoder for fusion. The integration of gradient features and semantic information enables the model to achieve superior segmentation performance, especially in cases where edge sharpness is crucial.

[0085] The present invention uses the Dice coefficient as the main evaluation metric. In this embodiment, the GM-SAM model and other mainstream segmentation models are compared multiple times on the marine debris dataset, including SAM, MedSAM, SAMed, SonarSAM, SAM-Adapter, and Medical-Adapter. Figure 6 Shows a visual comparison of the model described in the present invention with mainstream models. It can be seen that the method described in this embodiment has clearer edges than other methods in complex scenes and low-quality images, and maintains high performance even in the presence of blurred edges and noise interference. Table 1 proves that the method described in this embodiment shows significantly superior performance on various segmentation metrics in the test set compared to other models. It is worth noting that GM-SAM achieves a Dice coefficient of 96.6%, which is significantly better than other methods.

[0086] Table 1 Quantitative performance comparison between the present invention and other segmentation models

[0087]

[0088] The above combination Figure 1 The sonar image segmentation method provided by the embodiment of the present invention has been introduced in detail. Next, the sonar image segmentation system provided by the embodiment of the present invention will be introduced in combination with the accompanying drawings.

[0089] In some possible implementation manners, the present invention provides a sonar image segmentation system, including:

[0090] A feature extraction module, which is configured to: based on the acquired underwater sonar image, extract an input feature map;

[0091] An MPFM module, which is configured to: based on the input feature map, extract a global feature map and a local region feature map, and fuse the local region feature map with the global feature map to obtain a first fused feature map; fuse the feature map output by the window attention module with the first fused feature map to obtain a second fused feature map;

[0092] The SAM encoder module is configured to: based on the input feature map, obtain a multi-layer feature map through an MLP layer; fuse the multi-layer feature map with the second fused feature map to obtain a third fused feature map;

[0093] The gradient enhancement transformation module is configured to: based on the input feature map, enhance the vertical gradient information and horizontal gradient information of each channel to obtain an edge feature map; fuse the edge feature map with the edge feature map to obtain a fourth fused feature map;

[0094] The decoder module is configured to: based on the fourth fused feature map, obtain the target segmentation image;

[0095] Wherein, the MPFM module, the SAM encoder, the gradient enhancement transformation module and the decoder constitute an image segmentation model.

[0096] In some embodiments, the MPFM module is specifically configured to: based on the input feature map, use a feature extraction layer to obtain a basic feature information map; based on the basic feature information map, use a CoreFormer module to generate query vectors, key vectors and value vectors; based on the query vectors and key vectors, use a base point projection to compress the information in the basic feature information map into a low-dimensional space to obtain a base point; based on the base point and the query vectors, obtain a first attention map; based on the base point and the key vectors, obtain a second attention map; based on the first attention map, the second attention map and the value vectors, obtain a global feature map.

[0097] In some embodiments, the MPFM module is specifically further configured to: divide the input feature map into multiple non-overlapping windows, perform attention calculation on each window, extract the key information within the window, and generate a local region feature map.

[0098] In some embodiments, the MPFM module is specifically further configured to: use a window attention module to splice the local region feature maps of all windows along the channel dimension and fuse them with the first fused feature map to obtain a second fused feature map.

[0099] In some embodiments, the SAM encoder module is specifically configured to: integrate a Sonar-Fusion adapter in each Transformer block in the SAM encoder, and fuse the multi-layer feature map with the second fused feature map in the Sonar-Fusion adapter to obtain a third fused feature map.

[0100] In some embodiments, the gradient enhancement transformation module is specifically configured to: based on the input image features, use the gradient enhancement transformation module to separately extract the vertical gradient information and horizontal gradient information of each channel of the input feature map, integrate the vertical gradient information and horizontal gradient information into the corresponding channels to obtain a gradient-enhanced feature map; perform encoding processing on the gradient-enhanced feature map to obtain an edge feature map.

[0101] According to an embodiment of the present invention, the sonar image segmentation system may correspond to executing the methods described in the embodiments of the present invention, and the above and other operations and / or functions of each module of the sonar image segmentation system are respectively for implementing Figure 1 the corresponding processes of the respective methods in, and for the sake of brevity, will not be described herein again.

[0102] Referring to Figure 7 the structural diagram of the computer device shown, the computer device includes a processor, a communication interface, and a computer-readable storage medium. Among them, the processor, the communication interface, and the computer-readable storage medium can be connected by a bus or other means. Among them, the communication interface is used to receive and send data. The computer-readable storage medium can be stored in the memory of the computer device. The computer-readable storage medium is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer-readable storage medium. The processor (or CPU (Central Processing Unit, central processor)) is the computing core and control core of the computer device, and is adapted to implement one or more instructions, and is specifically adapted to load and execute one or more instructions to implement the corresponding steps in the embodiment of the sonar image segmentation method.

[0103] This embodiment provides a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in the computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and this storage space stores the processing system of the computer device.

[0104] And, one or more instructions adapted to be loaded and executed by the processor are also stored in this storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory, or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.

[0105] In one embodiment, one or more instructions are stored in the computer-readable storage medium; the processor loads and executes one or more instructions stored in the computer-readable storage medium to implement the corresponding steps in the embodiment of the above sonar image segmentation method.

[0106] This embodiment provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the corresponding steps in the above-described embodiment of the sonar image segmentation method.

[0107] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of an embodiment implemented in hardware, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.

[0108] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the specified functions in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0109] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0110] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.

[0111] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0112] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A sonar image segmentation method, characterized in that, Including: Based on the obtained underwater sonar image, an input feature map is extracted; Based on the input feature map, using the MPFM module, a global feature map and a local region feature map are extracted, and the local region feature map is fused with the global feature map to obtain a first fused feature map; The feature map output by the window attention module is fused with the first fused feature map to obtain a second fused feature map; Based on the input feature map, using the SAM encoder, through the MLP layer, a multi-layer feature map is obtained; the multi-layer feature map is fused with the second fused feature map to obtain a third fused feature map; Based on the input feature map, using the gradient enhancement transformation module, the vertical gradient information and horizontal gradient information of each channel are enhanced to obtain an edge feature map; The edge feature map is fused with the edge feature map to obtain a fourth fused feature map; The fourth fused feature map is input into the decoder to obtain a target segmentation image; Among them, the MPFM module, the SAM encoder, the gradient enhancement transformation module and the decoder constitute an image segmentation model.

2. The sonar image segmentation method according to claim 1, wherein Based on the input feature map, using the MPFM module, a global feature map is extracted; the method includes: based on the input feature map, using a feature extraction layer to obtain a basic feature information map; based on the basic feature information map, using the CoreFormer module to generate query vectors, key vectors and value vectors; based on the query vectors and key vectors, using base point projection to compress the information in the basic feature information map into a low-dimensional space to obtain base points; based on the base points and query vectors, obtaining a first attention map; based on the base points and key vectors, obtaining a second attention map; based on the first attention map, the second attention map and the value vectors, obtaining a global feature map.

3. The sonar image segmentation method according to claim 1, wherein Based on the input feature map, using the MPFM module, a local region feature map is extracted; the method includes: dividing the input feature map into multiple non-overlapping windows, performing attention calculation on each window, extracting the key information within the window, and generating a local region feature map.

4. The sonar image segmentation method according to claim 3, characterized in that The feature map output by the window attention module is fused with the first fused feature map to obtain a second fused feature map; the method includes: using the window attention module to splice the local region feature maps of all windows along the channel dimension and then fuse them with the first fused feature map to obtain a second fused feature map.

5. The sonar image segmentation method according to claim 1, wherein The multi-layer feature map is fused with the second fused feature map to obtain a third fused feature map; the method includes: integrating a Sonar-Fusion adapter into each Transformer block in the SAM encoder, and fusing the multi-layer feature map with the second fused feature map in the Sonar-Fusion adapter to obtain a third fused feature map.

6. The sonar image segmentation method according to claim 1, characterized in that Based on the input feature map, using the gradient enhancement transformation module, the vertical gradient information and horizontal gradient information of each channel are enhanced to obtain an edge feature map; the method includes: based on the input image features, using the gradient enhancement transformation module to separately extract the vertical gradient information and horizontal gradient information of each channel of the input feature map, integrating the vertical gradient information and horizontal gradient information into the corresponding channels to obtain a gradient enhancement feature map; encoding the gradient enhancement feature map to obtain an edge feature map.

7. A sonar image segmentation system, characterized in that Including: A feature extraction module, which is configured to: extract an input feature map based on the obtained underwater sonar image; An MPFM module, which is configured to: extract a global feature map and a local region feature map based on the input feature map, fuse the local region feature map with the global feature map to obtain a first fused feature map; fuse the feature map output by the window attention module with the first fused feature map to obtain a second fused feature map; A SAM encoder module, which is configured to: obtain a multi-layer feature map based on the input feature map through an MLP layer; fuse the multi-layer feature map with the second fused feature map to obtain a third fused feature map; A gradient enhancement transformation module, which is configured to: enhance the vertical gradient information and horizontal gradient information of each channel based on the input feature map to obtain an edge feature map; fuse the edge feature map with the edge feature map to obtain a fourth fused feature map; A decoder module, which is configured to: obtain a target segmentation image based on the fourth fused feature map; Wherein, the MPFM module, the SAM encoder, the gradient enhancement transformation module and the decoder constitute an image segmentation model.

8. A computer device, characterized in that A processor, adapted to execute a computer program; A computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by the processor, the steps in the sonar image segmentation method according to any one of claims 1-6 are implemented.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps in the sonar image segmentation method according to any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by the processor, the steps in the sonar image segmentation method according to any one of claims 1-6 are implemented.