Blood vessel image segmentation method based on vision Mangbar context perception semantics
By adopting a visual Mamba-based encoder-decoder network model in vascular image segmentation, using the improved CSVM module and multi-scale edge guidance module, the problems of diverse morphological changes and low contrast in vascular image segmentation are solved, achieving higher segmentation accuracy and effect.
Patent Information
- Application Number
- CN202510090495.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The prior art has difficulties in vascular image segmentation, mainly because the vascular structures vary in various shapes, low contrast between small blood vessels, diverse shapes, non-vascular structures or lesions affect segmentation performance, and artificial products due to improper data collection are difficult to accurately reduce vascular images.
The vascular image segmentation method based on visual Mamba context-aware semantics is adopted, and the U-shaped structure network model of the encoder-decoder is segmented. The encoder takes the improved CSVM module as the core. A multi-scale edge guide module and a dynamic boundary perception module are set up between the decoder and the encoder to capture local and global features and aggregate boundary features and semantic information.
It significantly improves the accuracy and effect of vascular image segmentation, surpasses the existing SOTA method, can position blood vessel boundaries more accurately, reduce the problems of oversegmentation and insufficient segmentation, and is suitable for medical image segmentation tasks.
Smart Images

Figure CN120107578A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vascular image processing, and in particular to a vascular image segmentation method based on visual Mamba context-aware semantics. Background Art
[0002] Early diagnosis and medical intervention are crucial to improving the prognosis of patients with vascular diseases. Clinicians can use vascular imaging technology to observe abnormalities in blood vessels and intervene in the early stages of the disease. However, clinicians need to have sufficient clinical experience to make accurate diagnoses. Deep learning (DL) technology that effectively utilizes data has opened up important new research avenues for medical image analysis.
[0003] In the past, deep learning models used for image segmentation tasks were generally based on a U-shaped structure network of encoder-decoder, in which CNN or Vision Transformer was used in the encoder part. "U-Net: Convolutional Networks for Biomedical Image Segmentation" points out that UNet is a convolutional neural network (CNN) architecture for medical image segmentation. It has become a classic model in image segmentation tasks with its excellent segmentation performance and simple structure. UNet is mainly used for pixel-level classification tasks, especially for scenarios that require precise boundary segmentation, such as medical image analysis. UNet consists of two main parts: the encoder (downsampling path) and the decoder (upsampling path), as well as skip connections between them.
[0004] Each convolution layer in the convolutional neural network operates on the vascular image through a small convolution kernel (such as 3x3 or 5x5). Although the small convolution kernel can effectively capture local features, the small convolution kernel will lead to a limited receptive field, that is, each convolution kernel can only focus on the local area of the image, and cannot directly capture a larger range of global information.
[0005] Although Vision Transformer can capture global image information, it is limited by the attention mechanism, especially when dealing with long sequence modeling, it has high quadratic complexity, which leads to expensive computational overhead when dealing with downstream dense prediction tasks (such as target detection, semantic segmentation, etc.). In addition, in medical imaging, high-resolution images such as whole-slice pathological images are common, where transformer-based models may have overfitting problems and require a large amount of data for effective training.
[0006] Current networks based on Visual Mamba (VMamba) are able to achieve a global receptive field similar to that of visual transformers while maintaining linear complexity in the number of tokens. However, existing VMamba models still have difficulty maintaining spatial local and global dependencies of tokens in high-dimensional arrays due to their sequential nature.
[0007] In summary, although the current encoder-decoder based U-shaped structure network performs well in most medical image segmentation tasks, it has certain limitations in vascular image segmentation. The main reasons are: the vascular structure shows obvious morphological changes, including thick blood vessels, thin blood vessels and filamentous retinal vessels; small blood vessels have low contrast, diverse shapes, and lengths of generally less than 10 pixels, and some are even less than 1 pixel; non-vascular structures or lesions (such as the optic disc) may affect the segmentation performance; artifacts caused by improper data acquisition cannot restore vascular images well. Summary of the invention
[0008] The present invention aims to solve the above-mentioned technical problems existing in the prior art and provides a vascular image segmentation method based on visual Mamba context-aware semantics.
[0009] The technical solution of the present invention is: a vascular image segmentation method based on visual Mamba context-aware semantics, which performs vascular image segmentation based on a U-shaped structure network model of an encoder-decoder, wherein the encoder is an encoder constructed with a CSVM module as the core, and the CSVM module is an improved VSS Block structure, specifically, the DW convolution layer in front of SS2D in the VSS Block is replaced by a CCSA module, so as to display the local and global dependencies in the reserved space in a compressed form; a multi-scale edge guidance module is set between the encoder and the decoder, and the Laplace operator is used to emphasize the boundary features of the underlying features; at the same time, a dynamic boundary perception module is set between the encoder, the multi-scale edge guidance module and the decoder to aggregate the boundary features of the underlying features and the semantic information of the high-level features.
[0010] The encoder is provided with a residual block ResBlock1-1, a downsampling module Down sampling1-1, a CSVM module 1-1, a downsampling module Down sampling1-2, a CSVM module 1-2, a downsampling module Downsampling1-3 and a CSVM module 1-3 from top to bottom; the CCSA is provided with a channel attention module ChannelAttention, a channel prior module Channel Prior, two average pooling layers XAvgPool and YAvgPool, a deep one-dimensional convolutional layer MS-DW Conv with kernel sizes of 3, 5, 7 and 9 respectively, two concat function layers, a group normalization layer Group Norn and a ReLU activation function layer;
[0011] The decoder is provided with a residual block ResBlock2-1, an upsampling module Up sampling2-1, a residual block ResBlock2-2, an upsampling module Up sampling2-2, a residual block ResBlock2-3, an upsampling module Upsampling2-3, a residual block ResBlock2-4 and a segmentation head Seg Head in sequence from bottom to top;
[0012] The residual block ResBlock1-1 is connected to the residual block ResBlock2-4 through a multi-scale edge guidance module EGAA1-1; the CSVM module 1-1 is connected to the up-sampling module Up sampling2-3 through a multi-scale edge guidance module EGAA1-2, and the CSVM module 1-2 is connected to the up-sampling module Up sampling2-2 through a multi-scale edge guidance module EGAA1-3;
[0013] The CSVM module 1-3 is connected to the residual block ResBlock2-1 through a jump connection and another connection through a dynamic boundary perception module DBA.
[0014] The multi-scale edge guiding modules EGAA1-1, EGAA1-2, and EGAA1-3 have the same structure and are provided with a parallel reverse operation module Reverse1, a Gaussian filter GF, and a deep convolution layer DWConv. The outputs of the reverse operation module Reverse1, the Gaussian filter GF, and the deep convolution layer DWConv pass through the convolution layer Conv1 and are then connected to the parallel convolution layers Conv2, Convolution layer Conv3, and Convolution layer Conv4.
[0015] The dynamic boundary perception module DBA is provided with a dynamic filter Dynamic Filter, and the output of the dynamic filter DynamicFilter is divided into two paths: one path is provided with a linear layer Linear layer1-1, a normalization layer LayerNorm1-1, a convolution layer Conv5, an activation function Sigmoid1 and a reverse operation module Reverse2 in sequence, and the other path is provided with a linear layer Linear layer2-1, a normalization layer LayerNorm2-1, a convolution layer Conv6, and an activation function Sigmoid2 in sequence; the output features are then transmitted through the two paths and connected to the Softmax activation function layer, one path is provided with a linear layer Linear layer3-1, a normalization layer LayerNorm3-1 in sequence, and the other path is provided with a linear layer Linearlayer4-1, a normalization layer LayerNorm4-1 in sequence.
[0016] Input the blood vessel image into the network model and train it according to the following steps:
[0017] Step 1. Use the residual block ResBlock1-1 to process the blood vessel image I IN Processing to obtain feature map
[0019] Step 2. Use the down sampling module Down sampling1-1 and CSVM module 1-1 to Processing to obtain feature map Using the multi-scale edge guidance module EGAA1-1 Processing to obtain feature map
[0020] Step 3. Use the down sampling module Down sampling1-2 and CSVM module 1-2 to Processing to obtain feature map Using the multi-scale edge guidance module EGAA1-2 Processing to obtain feature map
[0021] Step 4. Use the down sampling module Down sampling1-3 and CSVM module 1-3 to Processing to obtain feature map Using the multi-scale edge guidance module EGAA1-3 Processing to obtain feature map
[0022] Step 5. After connecting with Input to the dynamic boundary perception module DBA for processing to obtain the feature map IF DBA ;
[0023] Step 6. and IF DBA The input is processed by the residual block ResBlock2-1 and then goes through the upsampling module Upsampling2-1 to obtain the feature map.
[0024] Step 7. and The input is processed by the residual block ResBlock2-2 and then goes through the upsampling module Upsampling2-2 to obtain the feature map.
[0025] Step 8. and The input is processed by the residual block ResBlock2-3 and then goes through the upsampling module Upsampling2-3 to obtain the feature map.
[0026] Step 9. and Input to the residual block ResBlock2-4 for processing to obtain the feature map
[0028] Step 10. Use the segmentation head Seg Head to segment the feature map Processing to obtain segmentation results
[0029] The training process of the CCSA modules in the CSVM modules 1-1, CSVM modules 1-2, and CSVM modules 1-3 is as follows: Input feature F∈R H×W×C Through the channel attention module ChannelAttention, a 1D channel attention map Mc∈R is inferred C×1×1 , then multiply Mc by the input feature F, and obtain the refined feature Fc∈R with channel attention through the channel prior module Channel Prior C×H×W ; The refined features Fc are processed by the average pooling layers XAvgPool and YAvgPool respectively to obtain two one-dimensional sequence structures Fc 1 ∈R C×H , Fc 2 ∈R C×W; The two one-dimensional sequence structures are processed by a deep one-dimensional convolutional layer MS-DW Conv with kernel sizes of 3, 5, 7, and 9 and two concat function layers, and then the two features are element-wise multiplied and passed through a group normalization layer Group Norn and a ReLU activation function layer.
[0030] The training process of the multi-scale edge guidance modules EGAA1-1, EGAA1-2, and EGAA1-3 is as follows: first, the input feature map is processed by the reverse operation module Reverse1, the Gaussian filter GF, and the deep convolution layer DWConv, and then the output result is processed by the convolution layer Conv1, and then by the parallel convolution layer Conv2, the convolution layer Conv3, and the convolution layer Conv4, respectively. The features extracted by the convolution layer Conv2 and the convolution layer Conv3 are multiplied and then multiplied with the features extracted by the convolution layer Conv4, and then residually connected with the input image features.
[0031] The training process of the dynamic boundary perception module DBA is as follows: first, the input features are processed by the dynamic filter Dynamic Filter, and then through two paths: one path is provided with a linear layer Linear layer1-1, a normalization layer LayerNorm1-1, a convolution layer Conv5, an activation function Sigmoid1 and a reverse operation module Reverse2 in sequence, and the other path is provided with a linear layer Linear layer2-1, a normalization layer LayerNorm2-1, a convolution layer Conv6, and an activation function Sigmoid2 in sequence; the results obtained by the two paths are multiplied with the input features respectively, and then the two multiplication results are added to obtain a feature map Then Then two feature maps are obtained through two paths, namely, feature maps and feature map The feature map and feature map After the connection, it passes through the Softmax activation function layer and the feature map After multiplication, we get the feature map IF DBA .
[0032] The present invention is based on the U-shaped structure network model of encoder-decoder for vascular image segmentation. The encoder is an encoder built with CSVM module as the core, and the CSVM module is an improved VSS Block structure. Specifically, the DW convolution layer in VSS Block is replaced by CCSA module, so as to display the local and global dependencies in the reserved space in a compressed form, allowing VMamba to access the local and global context before reaching the last token; a multi-scale edge guidance module EGAA is set between the encoder and the decoder, and the Laplacian operator is used to emphasize the boundary features of the underlying features to achieve more accurate boundary positioning; in addition, a dynamic boundary perception module is set between the encoder, the multi-scale edge guidance module and the decoder to aggregate the boundary features of the underlying features and the semantic information of the high-level features, revealing the visual details that are no longer obvious in the image segmentation strategy. Experimental results show that the present invention surpasses the existing SOTA method and demonstrates its effectiveness and practicality in processing medical images. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a schematic diagram of the framework of an embodiment of the present invention.
[0034] Figure 2 It is a schematic diagram of the improved VSS Block structure according to an embodiment of the present invention.
[0035] Figure 3 Schematic diagram of the structure of the CCSA module in an embodiment of the present invention.
[0036] Figure 4 Schematic diagram of the structure of the multi-scale edge guidance module EGAA according to an embodiment of the present invention.
[0037] Figure 5 Schematic diagram of the structure of the multi-scale edge guidance module EGAA according to an embodiment of the present invention. DETAILED DESCRIPTION
[0038] The present invention provides a vascular image segmentation method based on visual Mamba context-aware semantics, which performs vascular image segmentation based on a U-shaped structure network model of an encoder-decoder, wherein the encoder is an encoder constructed with a CSVM module as the core, and the CSVM module is as follows Figure 2The improved VSS Block structure shown is specifically that the DW convolution layer in front of SS2D in VMamba's VSS Block is replaced by a CCSA module, so that the local and global dependencies in the reserved space are displayed in a compressed form; a multi-scale edge guidance module is set between the encoder and the decoder, and the Laplace operator is used to emphasize the boundary features of the underlying features; at the same time, a dynamic boundary perception module is set between the encoder, the multi-scale edge guidance module and the decoder to aggregate the boundary features of the underlying features and the semantic information of the high-level features.
[0039] The specific framework (CSEM-Net) is as follows Figure 1 As shown:
[0040] The encoder is provided with a residual block ResBlock1-1 (ResBlock×2), a downsampling module Down sampling1-1, a CSVM module 1-1 (CSVM×2), a downsampling module Down sampling1-2, a CSVM module 1-2 (CSVM×2), a downsampling module Down sampling1-3 and a CSVM module 1-3 (CSVM×2) from top to bottom; the CCSA module structure is as follows Figure 3 As shown in the figure, there are a channel attention module ChannelAttention, a channel prior module Channel Prior, two average pooling layers XAvgPool and YAvgPool, a deep one-dimensional convolution layer MS-DW Conv with kernel sizes of 3, 5, 7, and 9 respectively, two concat function layers, a group normalization layer Group Norn, and a ReLU activation function layer;
[0041] The decoder is provided with a residual block ResBlock2-1 (ResBlock×2), an upsampling module Upsampling2-1, a residual block ResBlock2-2 (ResBlock×2), an upsampling module Up sampling2-2, a residual block ResBlock2-3 (ResBlock×2), an upsampling module Up sampling2-3, a residual block ResBlock2-4 (ResBlock×2) and a segmentation head Seg Head in sequence from bottom to top;
[0042] The residual block ResBlock1-1 is connected to the residual block ResBlock2-4 through a multi-scale edge guidance module EGAA1-1; the CSVM module 1-1 is connected to the up-sampling module Up sampling2-3 through a multi-scale edge guidance module EGAA1-2, and the CSVM module 1-2 is connected to the up-sampling module Up sampling2-2 through a multi-scale edge guidance module EGAA1-3;
[0043] The CSVM module 1-3 is connected to the residual block ResBlock2-1 through a jump connection and another connection through a dynamic boundary perception module DBA.
[0044] The multi-scale edge guidance modules EGAA1-1, EGAA1-2, and EGAA1-3 have the same structure. Figure 4 As shown: there are parallel reverse operation modules Reverse1, Gaussian filters GF, and deep convolution layers DWConv (3×3). The outputs of the reverse operation modules Reverse1, Gaussian filters GF, and deep convolution layers DWConv pass through the convolution layer Conv1 (1×1) and are then connected to the parallel convolution layers Conv2 (1×3), Conv3 (3×1), and Conv4 (3×3).
[0045] The dynamic boundary perception module DBA is as follows Figure 5 As shown, a dynamic filter Dynamic Filter is provided, and the output of the dynamic filter Dynamic Filter is divided into two paths: one path is provided with a linear layer Linear layer1-1, a normalization layer LayerNorm1-1, a convolution layer Conv5 (1×1), an activation function Sigmoid1 and a reverse operation module Reverse2 in sequence, and the other path is provided with a linear layer Linear layer2-1, a normalization layer LayerNorm2-1, a convolution layer Conv6 (1×1), and an activation function Sigmoid2 in sequence; the output features are then transmitted through two paths and connected to the Softmax activation function layer, one path is provided with a linear layer Linear layer3-1, a normalization layer LayerNorm3-1 in sequence, and the other path is provided with a linear layer Linear layer4-1, and a normalization layer LayerNorm4-1 in sequence.
[0046] Input the blood vessel image into the network model and train it according to the following steps:
[0047] Step 1. Use the residual block ResBlock1-1 to process the blood vessel image I IN Processing to obtain feature map
[0049] Step 2. Use the down sampling module Down sampling1-1 and CSVM module 1-1 to Processing to obtain feature map Using the multi-scale edge guidance module EGAA1-1 Processing to obtain feature map
[0050] Step 3. Use the down sampling module Down sampling1-2 and CSVM module 1-2 to Processing to obtain feature map Using the multi-scale edge guidance module EGAA1-2 Processing to obtain feature map
[0051] Step 4. Use the down sampling module Down sampling1-3 and CSVM module 1-3 to Processing to obtain feature map Using the multi-scale edge guidance module EGAA1-3 Processing to obtain feature map
[0052] Step 5. After connecting with Input to the dynamic boundary perception module DBA for processing to obtain the feature map IF DBA ;
[0053] Step 6. and IF DBA The input is processed by the residual block ResBlock2-1 and then goes through the upsampling module Upsampling2-1 to obtain the feature map.
[0054] Step 7. and The input is processed by the residual block ResBlock2-2 and then goes through the upsampling module Upsampling2-2 to obtain the feature map.
[0055] Step 8. and The input is processed by the residual block ResBlock2-3 and then goes through the upsampling module Upsampling2-3 to obtain the feature map.
[0056] Step 9. and Input to the residual block ResBlock2-4 for processing to obtain the feature map
[0058] Step 10. Use the segmentation head Seg Head to segment the feature map Processing, get the segmentation result I Output .
[0059] The training process of the CCSA modules in the CSVM modules 1-1, CSVM modules 1-2, and CSVM modules 1-3 is as follows: Input feature F∈R H×W×C Through the channel attention module ChannelAttention, a 1D channel attention map Mc∈R is inferred C×1×1 , then multiply Mc by the input feature F, and obtain the refined feature Fc∈R with channel attention through the channel prior module Channel Prior C×H×W ; The refined features Fc are processed by the average pooling layers XAvgPool and YAvgPool respectively to obtain two one-dimensional sequence structures Fc 1 ∈R C×H , Fc 2 ∈R C×W ; The two one-dimensional sequence structures are processed by a deep one-dimensional convolutional layer MS-DW Conv with kernel sizes of 3, 5, 7, and 9 and two concat function layers, and then the two features are element-wise multiplied and passed through a group normalization layer Group Norn and a ReLU activation function layer.
[0060] The training process of the multi-scale edge guidance modules EGAA1-1, EGAA1-2, and EGAA1-3 is as follows: first, the input feature map is processed by the reverse operation module Reverse1, the Gaussian filter GF, and the deep convolution layer DWConv, and then the output result is processed by the convolution layer Conv1, and then by the parallel convolution layer Conv2, the convolution layer Conv3, and the convolution layer Conv4, respectively. The features extracted by the convolution layer Conv2 and the convolution layer Conv3 are multiplied and then multiplied with the features extracted by the convolution layer Conv4, and then residually connected with the input image features.
[0061] The training process of the dynamic boundary perception module DBA is as follows: first, the input features are processed by the dynamic filter Dynamic Filter, and then through two paths: one path is provided with a linear layer Linear layer1-1, a normalization layer LayerNorm1-1, a convolution layer Conv5, an activation function Sigmoid1 and a reverse operation module Reverse2 in sequence, and the other path is provided with a linear layer Linear layer2-1, a normalization layer LayerNorm2-1, a convolution layer Conv6, and an activation function Sigmoid2 in sequence; the results obtained by the two paths are multiplied with the input features respectively, and then the two multiplication results are added to obtain a feature map Then Then two feature maps are obtained through two paths, namely, feature maps and feature map The feature map and feature map After the connection, it passes through the Softmax activation function layer and the feature map After multiplication, we get the feature map IF DBA .
[0062] experiment:
[0063] 1. Algorithm comparison experiments on public datasets
[0064] The performance results of the present invention (CSEM-Net) on the FIVES dataset are shown in Table 1 compared with some of the most relevant models including U-Net, R2U-Net, Attention-UNet, Global Convolutional Network (GCN), Deeplab V3+, Selective Kernel (SK), CBAM, PSPNet, ENet, SegNet, SwinUNet, TransUNet and sgat-net. The results show that CSEM-Net can achieve the best segmentation performance in most cases. First, CSEM-Net outperforms the most advanced methods, such as SGAT-Net and Skelcon, in the IOU metric, with an improvement of 0.20% and 0.35% respectively. IOU measures the overlap between the predicted segmented area and the true annotated area. The IOU of the present invention well proves that the method of the present invention effectively avoids the problem of over-segmentation and under-segmentation. Secondly, the SE index of the present invention has achieved a competitive level (91.87%) with methods other than CBAM (93.30%). SE can reflect the under-segmentation of foreground targets, indicating that CSEM-Net can successfully segment more retinal vessels. Again, the SP index of CSEM-Net achieved suboptimal performance, which was 0.10% lower than that of Skelcon. SP is used to evaluate the over-segmentation of the background. The performance comparison with the existing methods shows that CSEM-Net can meet the classification requirements of background pixels. For the two comprehensive evaluation indicators, ACC and F1, CSEM-Net outperforms the most advanced methods, including Genetic U-Net, Skelcon, and SGAT-Net. The ACC metric considers both foreground and background segmentation results, and the performance of CSEM-Net on FIVES reaches 99.07%, which is 0.10%, 0.31%, and 0.21% higher than the state-of-the-art performance, respectively. The F1 metric can be regarded as a compromise between Sensitivity and Precision, which can more balancedly reflect the classification accuracy of vascular pixels. The F1 index of CSEM-Net reached 91.11%, which is 0.40%, 0.47%, and 0.60% higher than the above three methods. The improvement of the comprehensive indicators shows that CSEM-Net can not only accurately segment retinal blood vessels, but also effectively suppress background noise. In order to evaluate the changes in segmentation results under different probability thresholds, the area under the receiver operator characteristic (ROC) curve, i.e., AUC, was also calculated. The CSEM-Net of the present invention achieved indicators that are competitive with Genetic U-Net and Skelcon, but slightly lower than SGAT-Net. The competitive AUC indicator shows that CSEM-Net can more confidently distinguish retinal vessels from the background.
[0065] Table 1
[0066]
[0067] 2. Algorithm Comparison Experiments on Private Datasets
[0068] The present invention selects the performance of U-Net, UNet++, TransUnet, Swin-UNet, VM-UNet and CSEM-Net of the present invention on coronary artery images. The training strategy adopted is to randomly select 40 images for training and the remaining 10 images for testing. The comparative experimental results are shown in Table 2. From the results in Table 2, it can be seen that the CSEM-Net of the present invention has achieved significant results in the coronary artery segmentation task, and has achieved the best results in sensitivity, specificity, accuracy, F1, and AUC, and has achieved suboptimal results in IOU, which is about 0.25% lower than the optimal SwinUNet.
[0069] The experimental results finally show that the CSEM-UNet model of the present invention significantly improves the accuracy of medical image segmentation, especially in vascular imaging, through the innovative CSVM module, dynamic edge aggregation module DBA and multi-scale edge guidance module (EGAA). The CSVM module combines the linear time complexity advantage of Mamba and the global feature selection ability of CCSA, while the DBA module strengthens the edge information in the feature map by simulating the biological visual perception process, effectively improving the reconstruction of details. In addition, the introduction of the EGAA module combines traditional edge detection and deep learning methods, improving the model's ability to detect weak boundaries. It has surpassed the existing SOTA methods and demonstrated its effectiveness and practicality in processing medical images.
Claims
1. A vascular image segmentation method based on visual Mamba context-aware semantics is based on an encoder-decoder U-shaped structure network model for vascular image segmentation, characterized by: The encoder is an encoder built with the CSVM module as the core. The CSVM module is an improved VSS module structure. Specifically, the DW convolution layer in front of SS2D in the VSS module is replaced by a CCSA module, so that the local and global dependencies in the reserved space are displayed in a compressed form; a multi-scale edge guidance module is set between the encoder and the decoder, and the Laplace operator is used to emphasize the boundary features of the underlying features; at the same time, a dynamic boundary perception module is set between the encoder, the multi-scale edge guidance module and the decoder to aggregate the boundary features of the underlying features and the semantic information of the high-level features.
2. The vascular image segmentation method based on Visual Mamba context-aware semantics according to claim 1, characterized in that: The encoder is provided with a residual block ResBlock1-1, a downsampling module Down sampling1-1, a CSVM module 1-1, a downsampling module Down sampling1-2, a CSVM module 1-2, a downsampling module Down sampling1-3 and a CSVM module 1-3 from top to bottom; the CCSA is provided with a channel attention module ChannelAttention, a channel prior module Channel Prior, two average pooling layers XAvgPool and YAvgPool, a deep one-dimensional convolutional layer MS-DW Conv with kernel sizes of 3, 5, 7 and 9 respectively, two concat function layers, a group normalization layer Group Norn and a ReLU activation function layer; The decoder is provided with a residual block ResBlock2-1, an upsampling module Up sampling2-1, a residual block ResBlock2-2, an upsampling module Up sampling2-2, a residual block ResBlock2-3, an upsampling module Upsampling2-3, a residual block ResBlock2-4 and a segmentation head Seg Head in sequence from bottom to top; The residual block ResBlock1-1 is connected to the residual block ResBlock2-4 through a multi-scale edge guidance module EGAA1-1; the CSVM module 1-1 is connected to the up-sampling module Up sampling2-3 through a multi-scale edge guidance module EGAA1-2, and the CSVM module 1-2 is connected to the up-sampling module Up sampling2-2 through a multi-scale edge guidance module EGAA1-3; The CSVM module 1-3 is connected to the residual block ResBlock2-1 in one jump connection and in another connection through a dynamic boundary perception module DBA.
3. The vascular image segmentation method based on Visual Mamba context-aware semantics according to claim 2, characterized in that: The multi-scale edge guiding modules EGAA1-1, EGAA1-2, and EGAA1-3 have the same structure and are provided with a parallel reverse operation module Reverse1, a Gaussian filter GF, and a deep convolution layer DWConv. The outputs of the reverse operation module Reverse1, the Gaussian filter GF, and the deep convolution layer DWConv pass through the convolution layer Conv1 and are then connected to the parallel convolution layers Conv2, Convolution layer Conv3, and Convolution layer Conv4.
4. The method for vascular image segmentation based on Visual Mamba context-aware semantics according to claim 3, characterized in that: The dynamic boundary perception module DBA is provided with a dynamic filter Dynamic Filter, and the output of the dynamic filter DynamicFilter is divided into two paths: one path is provided with a linear layer Linear layer1-1, a normalization layer LayerNorm1-1, a convolution layer Conv5, an activation function Sigmoid1 and a reverse operation module Reverse2 in sequence, and the other path is provided with a linear layer Linear layer2-1, a normalization layer LayerNorm2-1, a convolution layer Conv6, and an activation function Sigmoid2 in sequence; The output features are then passed through two paths and connected to the Softmax activation function layer. One path is equipped with a linear layer Linearlayer3-1 and a normalization layer LayerNorm3-1, and the other path is equipped with a linear layer Linear layer4-1 and a normalization layer LayerNorm4-1.
5. The vascular image segmentation method based on Visual Mamba context-aware semantics according to claim 4 is characterized in that Input the blood vessel image into the network model and train it according to the following steps: Step 1. Use the residual block ResBlock1-1 to process the blood vessel image I IN Processing to obtain feature map Step 2. Use the down sampling module Down sampling1-1 and CSVM module 1-1 to Processing to obtain feature map Using the multi-scale edge guidance module EGAA1-1 Processing to obtain feature map Step 3. Use the down sampling module Down sampling1-2 and CSVM module 1-2 to Processing to obtain feature map Using the multi-scale edge guidance module EGAA1-2 Processing to obtain feature map Step 4. Use the down sampling module Down sampling1-3 and CSVM module 1-3 to Processing to obtain feature map Using the multi-scale edge guidance module EGAA1-3 Processing to obtain feature map Step 5. After connecting with Input to the dynamic boundary perception module DBA for processing to obtain the feature map IF DBA ; Step 6. and IF DBA The input is processed by the residual block ResBlock2-1 and then goes through the upsampling module Upsampling2-1 to obtain the feature map. Step 7. and The input is processed by the residual block ResBlock2-2 and then goes through the upsampling module Upsampling2-2 to obtain the feature map. Step 8. and The input is processed by the residual block ResBlock2-3 and then goes through the upsampling module Upsampling2-3 to obtain the feature map. Step 9. and Input to the residual block ResBlock2-4 for processing to obtain the feature map Step 10. Use the segmentation head Seg Head to segment the feature map Processing, get the segmentation result I Output .
6. The vascular image segmentation method based on Visual Mamba context-aware semantics according to claim 5, characterized in that The training process of the CCSA modules in the CSVM modules 1-1, CSVM modules 1-2, and CSVM modules 1-3 is as follows: Input feature F∈R H×W×C Through the channel attention module ChannelAttention, a 1D channel attention map Mc∈R is inferred C×1×1 , then multiply Mc by the input feature F, and obtain the refined feature Fc∈R with channel attention through the channel prior module Channel Prior C×H×W ; The refined features Fc are processed by the average pooling layers XAvgPool and YAvgPool respectively to obtain two one-dimensional sequence structures Fc1∈R C×H , Fc2∈R C×W ; The two one-dimensional sequence structures are processed by a deep one-dimensional convolutional layer MS-DW Conv with kernel sizes of 3, 5, 7, and 9 and two concat function layers, and then the two features are element-wise multiplied and passed through a group normalization layer Group Norn and a ReLU activation function layer.
7. The vascular image segmentation method based on Visual Mamba context-aware semantics according to claim 6, characterized in that The training process of the multi-scale edge guidance modules EGAA1-1, EGAA1-2, and EGAA1-3 is as follows: first, the input feature map is processed by the reverse operation module Reverse1, the Gaussian filter GF, and the deep convolution layer DWConv, and then the output result is processed by the convolution layer Conv1, and then by the parallel convolution layer Conv2, the convolution layer Conv3, and the convolution layer Conv4, respectively. The features extracted by the convolution layer Conv2 and the convolution layer Conv3 are multiplied and then multiplied with the features extracted by the convolution layer Conv4, and then residually connected with the input image features.
8. The vascular image segmentation method based on Visual Mamba context-aware semantics according to claim 6, characterized in that The training process of the dynamic boundary perception module DBA is as follows: first, the input features are processed by the dynamic filter Dynamic Filter, and then through two paths: one path is provided with a linear layer Linear layer1-1, a normalization layer LayerNorm1-1, a convolution layer Conv5, an activation function Sigmoid1 and a reverse operation module Reverse2 in sequence, and the other path is provided with a linear layer Linear layer2-1, a normalization layer LayerNorm2-1, a convolution layer Conv6, and an activation function Sigmoid2 in sequence; the results obtained by the two paths are multiplied with the input features respectively, and then the two multiplication results are added to obtain a feature map Then Then two feature maps are obtained through two paths, namely, feature maps and feature map The feature map and feature map After the connection, it passes through the Softmax activation function layer and the feature map After multiplication, we get the feature map IF DBA .
Citation Information
Patent Citations
Lightweight semantic segmentation method for high-resolution remote sensing image
CN112183360A
Brain tumor image segmentation method based on multi-scale convolution and Mama structure
CN118447244A
Vehicle image segmentation method based on edge guidance and dynamic pruning
CN118823343A
Multi-task blood vessel segmentation model construction method based on hybrid encoder
CN119251246A
An edge-guided RGBD underwater salient object detection method with multi-attention
JP7605548B1