Underwater forward-looking sonar image segmentation method, system and device based on CNN and Transform
By combining CNN and Transformer methods, low resolution, noise interference and complex background problems in underwater forward sonar image segmentation are solved, and more accurate target area segmentation is achieved, improving the effect of underwater sonar image processing.
Patent Information
- Application Number
- CN202510463417.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-08-08
AI Technical Summary
Underwater forward-view sonar image segmentation has problems such as low resolution, noise interference, target blur and occlusion, and complex background interference, resulting in inaccurate target area segmentation.
Using CNN and Transformer methods, feature extraction is performed through cascaded convolution module and hybrid Transformer model, combined with edge feature modules to learn edge information of the target area, data augmentation technology is used to generate high-quality data, and image features are restored through dense convolution and decoder, and segmentation model is optimized using joint loss function.
It improves the accuracy and accuracy of underwater forward-view sonar image segmentation, and can effectively deal with target area segmentation in complex underwater environments.
Smart Images

Figure CN120451191A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital image processing and pattern recognition, and in particular to a method, system and device for underwater forward-looking sonar image segmentation based on CNN and Transformer.
[0002] Acquiring underwater information has long been a challenge for ocean development planning worldwide. With the advent of modern sonar, it has gradually become an effective tool for countries worldwide to obtain underwater information. Underwater sonar images are a key channel for humans to obtain ocean information, and they can help us extract accurate and useful underwater information from images. Currently, sonar is mainly divided into forward-looking sonar, side-scan sonar, multi-beam sonar, and synthetic aperture sonar based on their operating principles. Forward-looking sonar (FLS) actively transmits sound waves and receives echoes, generating real-time two-dimensional or three-dimensional images of underwater scenes. It is widely used in underwater robotics (ROVs / AUVs), seabed mapping, shipwreck salvage, pipeline inspection, and military target identification.
[0003] The processing of underwater sonar images is generally divided into four steps: preprocessing, segmentation, feature extraction, and classification. Accurately segmenting the target area is a key step in image processing and is also a necessary process for subsequent identification. However, the segmentation of effective areas in underwater sonar images still faces huge challenges: 1) Low resolution and noise interference: The long wavelength of sound waves leads to low image resolution and is susceptible to reverberation, multipath effects, and random noise; 2) Target blur and occlusion: The shadow effect and edge diffusion phenomenon of acoustic imaging lead to unclear target outlines; 3) Complex background interference: Dynamic interference objects such as mud, bubbles, and organisms in underwater scenes can introduce false signals. Therefore, there is an urgent need to develop accurate, reliable, and automated technologies to accurately segment target areas using images collected by underwater forward-looking sonar, which will facilitate subsequent practical applications.
[0004] In light of this, the present invention aims to provide a method, system, and device for underwater forward-looking sonar image segmentation based on CNN and Transformer. By using a convolutional architecture and Transformer to focus on local and spatial information in an image, this approach provides more comprehensive and specific image features. Furthermore, a multi-task edge detection module is employed to learn edge information of target areas in sonar images, explicitly correcting edge features and thereby improving overall segmentation accuracy. Summary of the Invention
[0005] The purpose of the present invention is to provide a method, system and device for underwater forward-looking sonar image segmentation based on CNN and Transformer, which solves the problems existing in the prior art.
[0006] To solve the above technical problems, the present invention is achieved through the following technical solutions:
[0007] The present invention is a method for underwater forward-looking sonar image segmentation based on CNN and Transformer, comprising the following steps: S1: collecting image data collected by forward-looking sonar and performing corresponding annotations, and obtaining a training set and a validation set after data enhancement through data processing; S2: extracting features from the image through a cascaded convolution module and a hybrid Transformer model, the convolution module focuses on the local semantic information of the image, the Transformer encoder is used to capture global context information, the cross-fusion module realizes the interaction and fusion of dual-encoded information in the same stage, and uses parallel paths to deeply fuse the local features captured by the CNN branch and the global features captured by the Transformer encoder; S3: the features in the encoding stage are used as the source of extracting edge information, and the edge feature extraction module is used to fully learn them, and finally the edge information of the target area is obtained; S4: constructing a decoder through a dense convolution method, restoring the features of the encoder layer by layer, and splicing the edge detection results with the results of the image features restored in the decoder to obtain the segmentation results, and saving them in the model with the best effect in the validation set; S5: inputting the sonar image collected in real time into the trained segmentation model to obtain the segmented result map.
[0008] The present invention is further configured such that S1 includes the following steps: S11: collecting underwater sonar image data collected by sonar equipment and corresponding pixel-level annotation data; S12: enhancing all collected data to obtain enhanced data. The present invention adopts a comprehensive data enhancement method, integrating Contourlet transform and multi-scale Reinex into a generative adversarial network (GAN) to address problems such as brightness, contrast, and lack of details in low-light images, and generate high-quality data; S13: normalizing the entire sonar image so that the mean of the entire image is 0 and the standard deviation is 1, that is, the grayscale distribution of the image obeys a normal distribution.
[0009] The present invention is further configured such that S2 includes the following steps: S21: using a CNN encoding block to extract local features of an underwater sonar image, the size of the input image is H×W, and an initial convolution block is used to perform initial convolution before the first layer of the encoder. Each small convolution module includes: convolution Conv, batch normalization BN, and a linear layer ReLU to obtain an initial feature map of size H×W×C, where the number of channels is C; then a cascaded convolution module fused with a depthwise separable convolution is used to enhance feature learning; then a 2×2 convolution with a step size of 2 is used to perform downsampling, and the size of the compressed feature map is By repeatedly combining cascaded convolution modules and downsampling, feature maps containing more image information are obtained, with sizes of Among them, the formulas of each convolution module CBR(*) and cascade convolution DC(*) are as follows:
[0010] CBR(X)=ReLU(BN(Conv(X)))
[0011]
[0012] S22: Use the Transformer encoding block to extract global features of the image. First, the feature map obtained from the data map is serialized and the sequence is learned using the Enhanced Transformer (ET). This module consists of the Swin Transformer (ST) and the Axial Transformer (AT). The input image is divided into N image blocks through convolution operations and flattened into a sequence to obtain a size of The initial sequence is then transformed linearly in the linear embedding layer through convolution operations to map it to a space with a dimension of 4C. The sequence size is Then, after downsampling and ET encoding module, we get The feature sequence obtained by repeated operations is The formulas for ST, AT, and ET code blocks are as follows:
[0013] ST(x l-1 )=MLP(LN(W_MSA(LN(x l-1 ))+x l-1 ))+W_MSA(LN(x l-1 ))+x l-1 ,
[0014] z l+1 =MLP(LN(SW_MSA(LN(z l ))+z l ))+SW_MSA(LN(z l ))+z l
[0015] AT(x l-1 )=MLP(LN(reshape(WSA(HSA(CBR(reshape(x l-1 )))))+x l-1 ))
[0016] +reshape(WSA(HSA(CBR(reshape(x l-1 ))))),
[0017] ET(x l-1)=ST(x l-1 )+AT(x l-1 )
[0018] In the above formula, MLP(*) is a multi-layer perceptron layer; LN(*) represents the normalization layer; W_MSA(*) and SW_MSA(*) represent the window-based multi-head self-attention mechanism and the window-shift-based multi-head self-attention mechanism, respectively; CBR(*) represents the convolution module, and reshape(*) represents the feature mapping module, which maps features between the H×W×C and (H×W)×C domains to meet different computing requirements; HSA(*) and WSA(*) represent the self-attention mechanisms based on the height axis and width axis, respectively; S23: The feature fusion module is used to aggregate the feature maps of the two paths S21 and S22, and the complementary features are fully learned to improve the feature expression ability. The features of the S21 and S22 paths at the same stage are respectively represented by CT. i and FT i Indicates that i represents each layer of S21 and S22, with sizes H i ×W i ×C i and pair (H i ×W i )×C i CT i Feature reshape to FT i Under the same size, global average pooling and global maximum pooling are performed respectively to obtain two sets of 1×1×C global feature vectors, and CT i The two sets of global feature vectors obtained are fused into FT i In the i The two sets of global vectors obtained are fused into CT i Then the two sets of fused features are input into Swin Transformer for interaction and fusion of global and local information to improve the feature representation ability, and reshape operation is performed to obtain two sets of H i ×W i ×C i Feature representation; finally, the capsule network is used to fuse the two sets of features to obtain H i ×W i ×C i The specific process is shown in the formula:
[0019]
[0020]
[0021]
[0022] Here, GAP(*) represents global average pooling, CMP(*) represents global maximum pooling, and Caps(*) represents the capsule network operation. The capsule network uses vectors to represent a set of neurons, which can effectively characterize the spatial relationship between features and the probability of feature existence. Its core lies in the dynamic routing algorithm, which uses iterative calculations to better model the relationship between local features and the overall structure. The mathematical expression of the capsule network is as follows:
[0023]
[0024] In the above formula, take is the capsule block set of the first layer, is the set of capsule blocks in the l+1th layer, W ij For the learned weight matrix, the capsule block of the first layer is Linearly mapped to the capsule block in the l+1th layer r ij is the coupling coefficient between capsules i and j assigned according to the dynamic routing assignment algorithm.
[0025] The present invention is further configured such that the S3 specifically includes the following steps: S31: using the second and fourth layer features of the encoding stage as the edge information source, wherein the high-level features suppress the noise of the middle and low-level features, and the middle and low-level features supplement the edge details of the high-level features, S32: since the shallow features have more sensitive texture and edge information, the deep features have semantic information such as shape, structure or contextual relationship of higher layers, the second layer E2 in the CNN encoding module is used, with a size of and the last layer E4, with a size of The feature is used as an edge-aware structure to process edge information input. First, the two inputs are mapped to the same size channel through the convolution module CBR, and E2 is passed through the perception attention module based on the parallel architecture to emphasize the query area, effectively integrating the edge detail information retained in the middle and low layers; then, upsampling and downsampling are used to fuse the important features of the two layers; then the fused feature E 24 Continue to operate, 24 Global maximum pooling is used to compress global information and to calibrate the feature E 24 , and compare the obtained features with E 24 Add and pass Sigmoid to obtain the initial features Finally, after convolution and regularization, the initial edge contour f is output 24 .
[0026] The present invention is further configured such that S4 specifically includes the following steps: S41: constructing a decoder through dense convolution and deconvolution, amplifying the feature map in the encoding stage, continuously halving the number of channels and doubling the size, and using jump connections to splice it with the corresponding underlying feature map in the channel dimension until the feature map is restored to the same size as the input feature; S42: the entire network is trained using a joint loss function, and the best performing model on the validation set is saved, using the loss function shown below:
[0027] L(k)=L BCE (k)+L SSIM (k)+L IoU (k)
[0028] L BCE (k), L SSIM (k) and L IoU (k) represents the application of binary cross entropy loss (BCE), structural similarity loss (SSIM) and intersection over union loss (IoU) to k. BCE loss is widely used in classification tasks and can effectively handle unbalanced data:
[0029]
[0030] Where N is the number of samples, y i represents the true label of the i-th sample (0 or 1), p i It represents the probability that the i-th sample predicted by the model belongs to the positive class. SSIM is a metric used to measure the similarity between two images. It not only considers the similarity at the pixel level, but also factors such as brightness, contrast, and structural information:
[0031]
[0032] Among them, μ x and μ y represent the means of x and y respectively, and Denote the variance of x and y, σ xy represents the covariance between x and y. C1 and C2 are small constants to avoid division by zero. Typically, these constants are set to C1 = 0.012 and C2 = 0.032. IOU mainly measures the degree of overlap between the predicted bounding box and the true bounding box. It is often used to evaluate the model's ability to accurately locate objects in an image. The calculation formula for IOU is as follows:
[0033]
[0034] Among them, S(x,y)∈{0,1} represents the predicted saliency probability of pixel (x,y), and G(x,y)∈{0,1} represents the true label of pixel (x,y).
[0035] A CNN and Transformer-based underwater forward-looking sonar image segmentation system includes: a data enhancement module, a feature extraction module, an edge feature extraction module, and a feature decoding module. The feature extraction module includes a cascaded convolution feature extraction module, a hybrid Transformer feature extraction module, and a cross-feature fusion module.
[0036] A CNN- and Transformer-based underwater forward-looking sonar image segmentation device comprises a memory, a processor, and a computer program and related image data stored in the memory and executable on the processor. When the processor executes the program, the steps of an underwater forward-looking sonar image segmentation method are implemented.
[0037] The present invention has the following beneficial effects:
[0038] This paper processes the initial sonar data by integrating Contourlet transform, multi-scale Retinex, and GAN, achieving data enhancement. The network's encoder then leverages the superior feature learning capabilities of convolution and the long-range modeling capabilities of the hybrid Transformer to effectively learn the local and spatial information of the image data. Finally, a cross-fusion network based on the Swin Transformer and capsule network is used to deeply fuse the acquired local and spatial information. Furthermore, the invention's edge feature module effectively learns edge information of target areas in sonar images, thereby explicitly correcting segmentation results and improving overall network segmentation performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments.
[0040] Figure 1 This is a flowchart of a method for underwater forward-looking sonar image segmentation based on CNN and Transformer disclosed in the implementation of the present invention;
[0041] Figure 2 This is a diagram showing the overall architecture of a CNN and Transformer-based underwater forward-looking sonar image segmentation method disclosed in the implementation of the present invention;
[0042] Figure 3 Schematic diagram of the Swin Transformer module structure of a CNN and Transformer-based underwater forward-looking sonar image segmentation method disclosed in the embodiment of the present invention;
[0043] Figure 4This is a schematic diagram of the axial Transformer module structure of a CNN and Transformer-based underwater forward-looking sonar image segmentation method disclosed in the embodiment of the present invention;
[0044] Figure 5 Schematic diagram of the multi-head self-attention mechanism structure of a CNN and Transformer-based underwater sonar image segmentation method disclosed in the embodiment of the present invention;
[0045] Figure 6 Schematic diagram of the cross-feature fusion module structure of a CNN and Transformer-based underwater forward-looking sonar image segmentation method disclosed in the embodiment of the present invention;
[0046] Figure 7 This is a schematic diagram of the structure of an edge screening perception module of a CNN and Transformer-based underwater forward-looking sonar image segmentation method disclosed in the embodiment of the present invention;
[0047] Figure 8 This is a structural diagram of an underwater forward-looking sonar image segmentation system based on CNN and Transformer disclosed in the implementation of the present invention. DETAILED DESCRIPTION
[0048] The technical solutions in the embodiments of the present invention will be described below in conjunction with the drawings in the embodiments of the present invention. The described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0049] Example 1
[0050] Please refer to Figure 1-8, a method for underwater forward-looking sonar image segmentation based on CNN and Transformer, the method comprising the following steps: S1: collecting image data collected by the forward-looking sonar and performing corresponding annotations, and obtaining a training set and a validation set after data enhancement through data processing, S1 comprises the following steps: S11: collecting underwater sonar image data collected by the sonar equipment and the corresponding pixel-level annotation data; S12: enhancing all collected data to obtain enhanced data, the present invention adopts a comprehensive data enhancement method, integrating Contourlet transform and multi-scale Reinex into the generative adversarial network GAN, which is used to deal with the problems of brightness, contrast and lack of details in low-light images and generate high-quality data; S13: normalizing the entire sonar image so that the mean of the entire image is 0 and the standard deviation is 1, even if the image grayscale distribution obeys the normal distribution; S2: through cascaded convolution modules and hybrid Transforme The r model extracts features from the image. The convolution module focuses on the local semantic information of the image. The Transformer encoder is used to capture global contextual information. The cross-fusion module realizes the interaction and fusion of the dual-encoded information in the same stage. The parallel path is used to deeply fuse the local features captured by the CNN branch and the global features captured by the Transformer encoder. S2 includes the following steps: S21: Use the CNN encoding block to extract local features of the underwater sonar image. The size of the input image is H×W. The initial convolution block is used for initial convolution before the first layer of the encoder. Each small convolution module includes: convolution Conv, batch normalization BN and linear layer ReLU to obtain an initial feature map of size H×W×C, where the number of channels is C; then use the cascaded convolution module that integrates the depth-separable convolution to enhance the learning of the features; then use the 2×2 convolution with a step size of 2 to downsample, and the size of the compressed feature map is By repeatedly combining cascaded convolution modules and downsampling, feature maps containing more image information are obtained, with sizes of Among them, the formulas of each convolution module CBR(*) and cascade convolution DC(*) are as follows:
[0051] CBR(X)=ReLU(BN(Conv(X)))
[0052]
[0053] S22: Use the Transformer encoding block to extract global features of the image. First, the feature map obtained from the data map is serialized and the sequence is learned using the Enhanced Transformer (ET). This module consists of the Swin Transformer (ST) and the Axial Transformer (AT). The input image is divided into N image blocks through convolution operations and flattened into a sequence to obtain a size of The initial sequence is then transformed linearly in the linear embedding layer through convolution operations to map it to a space with a dimension of 4C. The sequence size is Then, after downsampling and ET encoding module, we get The feature sequence obtained by repeated operations is The formulas for ST, AT, and ET code blocks are as follows:
[0054] ST(x l-1 )=MLP(LN(W_MSA(LN(x l-1 ))+x l-1 ))+W_MSA(LN(x l-1 ))+x l-1 ,
[0055] z l+1 =MLP(LN(SW_MSA(LN(z l ))+z l ))+SW_MSA(LN(z l ))+z l
[0056] AT(x l-1 )=MLP(LN(reshape(WSA(HSA(CBR(reshape(x l-1 )))))+x l-1 ))
[0057] +reshape(WSA(HSA(CBR(reshape(x l-1 ))))),
[0058] ET(x l-1 )=ST(x l-1 )+AT(x l-1 )
[0059] In the above formula, MLP(*) is a multi-layer perceptron layer; LN(*) represents the normalization layer; W_MSA(*) and SW_MSA(*) represent the window-based multi-head self-attention mechanism and the window-shift-based multi-head self-attention mechanism, respectively; CBR(*) represents the convolution module, and reshape(*) represents the feature mapping module, which maps features between the H×W×C and (H×W)×C domains to meet different computing requirements; HSA(*) and WSA(*) represent the self-attention mechanisms based on the height axis and width axis, respectively; S23: The feature fusion module is used to aggregate the feature maps of the two paths S21 and S22, and the complementary features are fully learned to improve the feature expression ability. The features of the S21 and S22 paths at the same stage are respectively represented by CT. i and FT i Indicates that i represents each layer of S21 and S22, with sizes H i ×W i ×C i and pair (H i ×W i )×C i CT i Feature reshape to FT i Under the same size, global average pooling and global maximum pooling are performed respectively to obtain two sets of 1×1×C global feature vectors, and CT i The two sets of global feature vectors obtained are fused into FT i In the i The two sets of global vectors obtained are fused into CT i Then the two sets of fused features are input into Swin Transformer for interaction and fusion of global and local information to improve the feature representation ability, and reshape operation is performed to obtain two sets of H i ×W i ×C i Feature representation; finally, the capsule network is used to fuse the two sets of features to obtain H i ×W i ×C i The specific process is shown in the formula:
[0060]
[0061]
[0062]
[0063] Here, GAP(*) represents global average pooling, CMP(*) represents global maximum pooling, and Caps(*) represents the capsule network operation. The capsule network uses vectors to represent a set of neurons, which can effectively characterize the spatial relationship between features and the probability of feature existence. Its core lies in the dynamic routing algorithm, which uses iterative calculations to better model the relationship between local features and the overall structure. The mathematical expression of the capsule network is as follows:
[0064]
[0065] In the above formula, take is the capsule block set of the first layer, is the set of capsule blocks in the l+1th layer, W ij For the learned weight matrix, the capsule block of the first layer is Linearly mapped to the capsule block in the l+1th layer r ij It is the coupling coefficient between capsules i and j assigned according to the dynamic routing allocation algorithm; S3: The features of the encoding stage are used as the source of edge information extraction, and the edge feature extraction module is used to fully learn it, and finally the edge information of the target area is obtained. S3 specifically includes the following steps: S31: Use the second and fourth layer features of the encoding stage as the source of edge information, where high-level features suppress the noise of low-level features, and low-level features supplement the edge details of high-level features. S32: Since shallow features have more sensitive texture and edge information, deep features have higher-level semantic information such as shape, structure or contextual relationship, and the second layer E2 in the CNN encoding module is used, with a size of and the last layer E4, with a size of The feature is used as an edge-aware structure to process edge information input. First, the two inputs are mapped to the same size channel through the convolution module CBR, and E2 is passed through the perception attention module based on the parallel architecture to emphasize the query area, effectively integrating the edge detail information retained in the middle and low layers; then, upsampling and downsampling are used to fuse the important features of the two layers; then the fused feature E 24 Continue to operate, 24 Global maximum pooling is used to compress global information and to calibrate the feature E 24 , and compare the obtained features with E 24 Add and pass Sigmoid to obtain the initial features Finally, after convolution and regularization, the initial edge contour f is output 24; S4: Construct a decoder through dense convolution method, restore the features of the encoder layer by layer, and splice the edge detection results with the results of the image features restored in the decoder to obtain the segmentation results, and save the model with the best effect in the verification set. S4 specifically includes the following steps: S41: Construct a decoder through dense convolution and deconvolution, amplify the feature map of the encoding stage, halve the number of channels and double the size, and use jump connections to splice it with the corresponding underlying feature map in the channel dimension until it is restored to the feature map with the same size as the input feature; S42: The entire network is trained using a joint loss function, and the model with the best performance on the verification set is saved, using the loss function shown below:
[0066] L(k)=L BCE (k)+L SSIM (k)+L IoU (k)
[0067] L BCE (k), L SSIM (k) and L IoU (k) represents the application of binary cross entropy loss (BCE), structural similarity loss (SSIM) and intersection over union loss (IoU) to k. BCE loss is widely used in classification tasks and can effectively handle unbalanced data:
[0068]
[0069] Where N is the number of samples, y i represents the true label of the i-th sample (0 or 1), p i It represents the probability that the i-th sample predicted by the model belongs to the positive class. SSIM is a metric used to measure the similarity between two images. It not only considers the similarity at the pixel level, but also factors such as brightness, contrast, and structural information:
[0070]
[0071] Among them, μ x and μ y represent the means of x and y respectively, and Denote the variance of x and y, σ xy represents the covariance between x and y. C1 and C2 are small constants to avoid division by zero. Typically, these constants are set to C1 = 0.012 and C2 = 0.032. IOU mainly measures the degree of overlap between the predicted bounding box and the true bounding box. It is often used to evaluate the model's ability to accurately locate objects in an image. The calculation formula for IOU is as follows:
[0072]
[0073] Among them, S(x,y)∈{0,1} represents the predicted saliency probability of pixel (x,y), and G(x,y)∈{0,1} represents the true label of pixel (x,y); S5: Input the real-time sonar image into the trained segmentation model to obtain the segmented result image.
[0074] Example 2
[0075] Please refer to Figure 8 , an underwater forward-looking sonar image segmentation system based on CNN and Transformer, including: data enhancement module (random cropping, Contourlet transform, multi-scale Retinex, GAN), feature extraction module (convolutional feature extraction module, Swin Transformer / axial Transformer feature extraction module and cross feature fusion module), edge feature extraction module and feature decoding module.
[0076] Among them, the data enhancement module mainly uses GAN network, random cropping, Contourlet transform and multi-scale Retinex to enhance the data of the segmented sonar images, and normalizes the pixel distribution of the image.
[0077] The feature extraction module consists of three parts: the convolutional feature extraction module, the Swin Transformer and axial Transformer feature extraction modules, and the cross-feature fusion module. The convolutional feature extraction module uses a convolutional network to perform local feature extraction on the input image and downsamples it using a 2×2 convolution with a stride of 2 to obtain the convolutional features. The Swin Transformer and axial Transformer feature extraction modules perform layer-by-layer feature extraction on the serialized feature maps of the input image, downsample using a convolution with a stride of 2, and finally obtain the Transformer global features. The cross-feature fusion module deeply fuses the output convolutional features with the Transformer features to obtain the final output of the feature extraction module.
[0078] The edge feature extraction module primarily uses the output features of the intermediate layers of the convolutional feature extraction module to learn edge features of image data. Low-level features are designed to provide local edge information, while high-level features have a larger receptive field and are more spatially invariant. Edge information is learned using methods such as perceptual attention.
[0079] The feature decoding module receives the fused features output by the feature extraction module and restores the size of the feature map through dense convolution and upsampling layer by layer. Skip connections can transmit semantic features with higher resolution at different layers to the upsampling network. The output of the decoding module is fused with the output of the edge feature extraction module to obtain the final segmentation result.
[0080] Example 3
[0081] A CNN- and Transformer-based underwater forward-looking sonar image segmentation device includes a memory, a processor, and a computer program and related image data stored in the memory and executable on the processor. The device is characterized in that when the processor executes the program, the steps of the underwater forward-looking sonar image segmentation method are implemented.
[0082] In this embodiment, the memory is used to store computer programs and related data, and may include various types of storage media, including random access memory (RAM), including static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), and double data rate random access memory (DDR RAM);
[0083] Read-only memory (ROM), including programmable read-only memory (PROM) and electrically erasable programmable read-only memory (EEPROM); magnetic storage, including hard disk drives (HDDs) and solid-state drives (SSDs); optical disk storage and removable storage, etc.
[0084] In this example, the processor may be a central processing unit (CPU), a microprocessor, or other data processing chip, which is primarily used to execute computer programs or process data in the memory to implement the processing of the underwater forward-looking sonar image segmentation method based on CNN and Transformer in the aforementioned embodiment, thereby obtaining a segmentation result of feature information such as the target area in the image based on the image collected by the sonar device.
[0085] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all details in detail, nor do they limit the invention to only the specific implementation methods described. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can better understand and utilize the present invention.
Claims
1. A method for underwater forward-looking sonar image segmentation based on CNN and Transformer, characterized in that: The method comprises the following steps: S1: Collect image data collected by forward-looking sonar, perform corresponding annotations, and obtain data-enhanced training and validation sets through data processing; S2: Image feature extraction is performed through cascaded convolutional modules and a hybrid Transformer model. The convolutional module focuses on local semantic information of the image, while the Transformer encoder is used to capture global contextual information. The cross-fusion module enables interaction and fusion of dual-encoded information at the same stage, and uses parallel paths to deeply fuse local features captured by CNN branches and global features captured by the Transformer encoder. S3: The features in the encoding stage are used as the source of edge information extraction. The edge feature extraction module fully learns them and finally obtains the edge information of the target area. S4: Build a decoder using dense convolution to recover the encoder features layer by layer. Combine the edge detection results with the image features recovered in the decoder to obtain the segmentation results. The model with the best performance is saved in the validation set. S5: Input the real-time collected sonar image into the trained segmentation model to obtain the segmented result image.
2. The underwater forward-looking sonar image segmentation method based on CNN and Transformer according to claim 1, characterized in that: Said S1 comprises the following steps: S11: Collect underwater sonar image data collected by sonar equipment and corresponding pixel-level annotation data; S12: Enhance all collected data to obtain enhanced data. The present invention adopts a comprehensive data enhancement method, integrating Contourlet transform and multi-scale Reinex into a generative adversarial network (GAN) to address issues such as brightness, contrast, and lack of details in low-light images, and generate high-quality data. S13: Normalize the entire sonar image so that the mean of the entire image is 0 and the standard deviation is 1, that is, the grayscale distribution of the image obeys the normal distribution.
3. The underwater forward-looking sonar image segmentation method based on CNN and Transformer according to claim 1, characterized in that: The S2 comprises the following steps: S21: Use CNN encoding blocks to extract local features of underwater sonar images. The size of the input image is H×W. The initial convolution block is used before the first layer of encoder for initial convolution. Each small convolution module includes: convolution Conv, batch normalization BN and linear layer ReLU, and obtains an initial feature map of size H×W×C, where the number of channels is C; then use cascaded convolution modules that fuse depthwise separable convolution to enhance feature learning; then use 2×2 convolution with a stride of 2 to downsample, and the size of the compressed feature map is By repeatedly combining cascaded convolution modules and downsampling, feature maps containing more image information are obtained, with sizes of Among them, the formulas of each convolution module CBR(*) and cascade convolution DC(*) are as follows: CBR(X)=ReLU(BN(Conv(X))) S22: Use the Transformer encoding block to extract global features of the image. First, the feature map obtained from the data map is serialized and the sequence is learned using the Enhanced Transformer (ET). This module consists of the Swin Transformer (ST) and the Axial Transformer (AT). The input image is divided into N image blocks through convolution operations and flattened into a sequence to obtain a size of The initial sequence is then transformed linearly in the linear embedding layer through convolution operations to map it to a space with a dimension of 4C. The sequence size is Then, after downsampling and ET encoding module, we get The feature sequence obtained by repeated operations is The formulas for ST, AT, and ET code blocks are as follows: ST(x l-1 )=MLP(LN(W_MSA(LN(x l-1 ))+x l-1 ))+W_MSA(LN(x l-1 ))+x l-1 , z l+1 =MLP(LN(SW_MSA(LN(z l ))+z l ))+SW_MSA(LN(z l ))+z l AT(x l-1 )=MLP(LN(reshape(WSA(HSA(CBR(reshape(x l-1 )))))+x l-1 ))+reshape(WSA(HSA(CBR(reshape(x l-1 ))))), ET(x l-1 )=ST(x l-1 )+AT(x l-1 ) In the above formula, MLP(*) is a multi-layer perceptron layer; LN(*) represents a normalization layer; W_MSA(*) and SW_MSA(*) represent window-based multi-head self-attention mechanisms and window-shift-based multi-head self-attention mechanisms, respectively; CBR(*) represents a convolution module, and reshape(*) represents a feature mapping module, which maps features between the H×W×C and (H×W)×C domains to meet different computational requirements; HSA(*) and WSA(*) represent self-attention mechanisms based on the height axis and width axis, respectively. S23: The feature fusion module is used to aggregate the feature maps of the two paths S21 and S22, and fully learn the complementary features to improve the feature expression ability. The features of the S21 and S22 paths in the same stage are respectively CT i and FT i Indicates that i represents each layer of S21 and S22, with sizes H i ×W i ×C i and pair (H i ×W i )×C i CT i Feature reshape to FT i Under the same size, global average pooling and global maximum pooling are performed respectively to obtain two sets of 1×1×C global feature vectors, and CT i The two sets of global feature vectors obtained are fused into FT i In the i The two sets of global vectors obtained are fused into CT i Then the two sets of fused features are input into Swin Transformer for interaction and fusion of global and local information to improve the feature representation ability, and reshape operation is performed to obtain two sets of H i ×W i ×C i Feature representation; finally, the capsule network is used to fuse the two sets of features to obtain H i ×W i ×C i The specific process is shown in the formula: Among them, GAP(*) represents global average pooling, CMP(*) represents global maximum pooling, and Caps(*) represents capsule network operation; The capsule network uses vectors to represent a set of neurons, effectively characterizing the spatial relationship between features and the probability of their existence. Its core lies in the dynamic routing algorithm, which uses iterative calculations to better model the relationship between local features and the overall structure. The mathematical expression of the capsule network is as follows: In the above formula, take is the capsule block set of the first layer, is the set of capsule blocks in the l+1th layer, W ij For the learned weight matrix, the capsule block of the first layer is Linearly mapped to the capsule block in the l+1th layer r ij is the coupling coefficient between capsules i and j assigned according to the dynamic routing assignment algorithm.
4. The underwater forward-looking sonar image segmentation method based on CNN and Transformer according to claim 1, characterized in that: The S3 specifically includes the following steps: S31: Use the second and fourth layer features of the encoding stage as the source of edge information, where high-level features suppress the noise of low- and medium-level features, and low- and medium-level features supplement the edge details of high-level features; S32: Since shallow features have more sensitive texture and edge information, deep features have higher-level semantic information such as shape, structure or context, the second layer E2 in the CNN encoding module is used, with a size of and the last layer E4, with a size of The feature is used as an edge-aware structure to process edge information input. First, the two inputs are mapped to the same size channel through the convolution module CBR, and E2 is passed through the perception attention module based on the parallel architecture to emphasize the query area, effectively integrating the edge detail information retained in the middle and low layers; then, upsampling and downsampling are used to fuse the important features of the two layers; then the fused feature E 24 Continue to operate, 24 Global maximum pooling is used to compress global information and to calibrate the feature E 24 , and compare the obtained features with E 24 Add and pass Sigmoid to obtain the initial features Finally, after convolution and regularization, the initial edge contour f is output 24 .
5. The underwater forward-looking sonar image segmentation method based on CNN and Transformer according to claim 1, characterized in that: The S4 specifically includes the following steps: S41: The decoder is constructed by dense convolution and deconvolution. The feature map in the encoding stage is enlarged, the number of channels is continuously halved and the size is doubled. Then, skip connections are used to splice it with the corresponding underlying feature map in the channel dimension until the feature map is restored to the same size as the input feature map. S42: The entire network is trained using a joint loss function, and the best performing model on the validation set is saved. The loss function is as follows: L(k)=L BCE (k)+L SSIM (k)+L IoU (k) L BCE (k), L SSIM (k) and L IoU (k) represents the application of binary cross entropy loss (BCE), structural similarity loss (SSIM) and intersection over union loss (IoU) to k. BCE loss is widely used in classification tasks and can effectively handle unbalanced data: Where N is the number of samples, y i represents the true label of the i-th sample (0 or 1), p i It represents the probability that the i-th sample predicted by the model belongs to the positive class. SSIM is a metric used to measure the similarity between two images. It not only considers the similarity at the pixel level, but also factors such as brightness, contrast, and structural information: Among them, μ x and μ y represent the means of x and y respectively, and represent the variance of x and y respectively, σ xy represents the covariance between x and y. C1 and C2 are small constants to avoid division by zero. Typically, these constants are set to C1 = 0.012 and C2 = 0.
032. IOU mainly measures the degree of overlap between the predicted bounding box and the true bounding box. It is often used to evaluate the model's ability to accurately locate objects in an image. The calculation formula for IOU is as follows: Among them, S(x,y)∈{0,1} represents the predicted saliency probability of pixel (x,y), and G(x,y)∈{0,1} represents the true label of pixel (x,y).
6. The underwater forward-looking sonar image segmentation system based on CNN and Transformer according to any one of claims 1 to 5, comprising: The data enhancement module, the feature extraction module, the edge feature extraction module and the feature decoding module are characterized in that the feature extraction module includes a cascade convolution feature extraction module, a hybrid Transformer feature extraction module and a cross feature fusion module.
7. The underwater forward-looking sonar image segmentation device based on CNN and Transformer according to any one of claims 1, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and related image data, characterized in that: When the processor executes the program, the steps of the underwater forward-looking sonar image segmentation method are implemented.
Citation Information
Cited By
Deep sea image instance segmentation system based on lightweight mixed visual structure
CN120807939A
Sonar target detection method based on neural architecture search technology
CN121213947A