Image segmentation method and device based on SM-UNet image segmentation model

By introducing sliding window mechanism and Mamba architecture SM-UNet model in medical image segmentation, the problems of high computational complexity and insufficient feature capture in the prior art are solved, and more efficient and accurate image segmentation effect is achieved.

CN120088269AActive Publication Date: 2025-06-03SICHUAN UNIV

Patent Information

Application Number
CN202510254398.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-03
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

The prior art has problems in medical image segmentation with high computational complexity, inability to effectively capture local features and insufficient grasp of multi-size features.

Method used

A method based on SM-UNet image segmentation model is proposed. Through the sliding window mechanism and Mamba architecture, combined with convolution module, encoder, jump connection module and decoder, the image segmentation process is optimized, the computational complexity is reduced and the feature extraction ability is improved.

Benefits of technology

It effectively reduces the computational complexity, improves the ability to capture local and multi-size features, and significantly improves the efficiency and accuracy of medical image segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088269A_ABST
    Figure CN120088269A_ABST
Patent Text Reader

Abstract

The invention relates to the field of image segmentation, in particular to an image segmentation method and device based on an SM-UNet image segmentation model. The invention provides a novel encoder and decoder Mama framework, a sliding window is applied to a UNet segmentation model based on Mama to establish an SM-UNet image segmentation model, Swin Mama UNet (SMUNet) is used for automatic medical image segmentation, and the advantage of Mama linear complexity and the cross-window interaction capability brought by the sliding window are mainly combined to effectively optimize the structure of a standard U-shaped framework.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image segmentation, and particularly to an image segmentation method and device based on an SM-UNet image segmentation model. Background Art

[0002] Existing solutions in the field of medical segmentation are mostly based on TransUNet. TransUNet is a deep learning model that combines the Transformer and UNet architectures and is specifically used for medical image segmentation. The Transformer has achieved great success in the field of natural language processing (NLP). With the development of technology, it has also been tried in the field of computer vision. However, when directly applying the Transformer to computer vision tasks, some defects are exposed:

[0003] (I) High computational complexity

[0004] 1. Slow calculation due to too long sequence

[0005] When applying the Transformer to computer vision tasks, the image is usually segmented into small patches. For example, for high-resolution images, in order to obtain more features by subdividing the patches, the generated sequence will be very long. Taking the processing of high-resolution images (such as 1000×1000 or 800×800) as an example, when the patch size is 16×16, a large number of patches will be generated, resulting in a sequence length of up to thousands.

[0006] 2. Computational complexity problem of the self-attention mechanism

[0007] The computational complexity of the self-attention mechanism is O(n 2 ), where n is the length of the sequence. When processing images, due to the long sequence length, as the image size increases, the computational complexity of the self-attention mechanism grows rapidly in a quadratic manner. For example, when the image size doubles, the sequence length may double, and the amount of calculation will become four times the original.

[0008] (II) Unable to effectively capture local features

[0009] Global relationship processing and the locality of visual information. A lot of visual information in the image depends on local relationships, such as features like textures and edges inside objects. Traditional Transformer processes global relationships. It performs self-attention calculations on the entire input sequence, ignoring the characteristics of this local information. In a convolutional neural network (CNN), the convolutional kernel slides within a local area for convolutional operations, which can effectively capture local features. However, the Transformer lacks such locality operations.

[0010] (3) Insufficient grasp of multi - scale features

[0011] Single - scale feature extraction, such as in early Transformer - based vision models like ViT (Vision Transformer), extracts features at a single scale. For example, it always processes with a patch size of 16×16, and the feature dimension remains unchanged throughout the network. In many downstream visual tasks (such as object detection and image segmentation), multi - scale features are very important. In object detection, objects of different sizes require features of different scales for accurate recognition; in image segmentation, different regions may require features of different resolutions to accurately divide boundaries. Single - scale feature extraction performs poorly in these tasks because it cannot adapt to the multi - scale feature requirements in different objects and scenarios.

[0012] Therefore, there is a need for image segmentation methods and devices with lower complexity that can extract features in a multi - scale and comprehensive manner. Summary of the Invention

[0013] The object of the present invention is to overcome the above - mentioned deficiencies in the prior art and provide an image segmentation method and device based on the SM - UNet image segmentation model.

[0014] To achieve the above - mentioned object of the invention, the present invention provides the following technical solutions:

[0015] An image segmentation method based on the SM - UNet image segmentation model, comprising the following steps:

[0016] Obtain the image to be segmented;

[0017] Input the image to be segmented into a pre - constructed SM - UNet image segmentation model;

[0018] The SM - UNet image segmentation model outputs an image segmentation result;

[0019] Among them, the SM - UNet image segmentation model includes a convolution module, an encoder, a skip - connection module, and a decoder;

[0020] The convolution module is used to perform convolution segmentation processing on the image to be segmented and output a number of feature map patches;

[0021] The encoder includes two branches for sliding window analysis and is used to gradually extract and compress features from a number of the feature map patches;

[0022] The skip - connection module is used to obtain the lost spatial details during the process;

[0023] The decoder includes two branches for performing sliding window analysis to gradually restore the resolution of the feature map, perform detail reconstruction, and output the image segmentation result.

[0024] As a preferred solution of the present invention, the convolution module includes a convolution layer and a Patch Embedding segmentation layer;

[0025] The convolution layer is used to perform convolution processing on the image to be segmented into a feature map;

[0026] The Patch Embedding segmentation layer is used to segment the feature map into a number of feature map patches of the same size.

[0027] As a preferred solution of the present invention, the encoder includes a number of SwinVSS Block blocks and a Patch Merging patch merging layer;

[0028] The SwinVSS Block block is used to capture global and local features;

[0029] The Patch Merging patch merging layer is used to merge adjacent feature map patches.

[0030] As a preferred solution of the present invention, the SwinVSS Block block includes a first VSS block and a second VSS block connected in sequence;

[0031] The first VSS block includes a normalization layer and a fully connected layer connected in sequence; among them, a first branch and a second branch are provided between the normalization layer and the fully connected layer; the first branch includes a fully connected layer, a depthwise separable convolution layer, a W-2D selective scanning layer, and a normalization layer connected in sequence; the second branch includes a fully connected layer.

[0032] The second VSS block includes a normalization layer and a fully connected layer connected in sequence; among them, a third branch and a fourth branch are provided between the normalization layer and the fully connected layer; the third branch includes a fully connected layer, a depthwise separable convolution layer, a SW-2D selective scanning layer, and a normalization layer connected in sequence; the fourth branch includes a fully connected layer.

[0033] As a preferred solution of the present invention, the W-2D selective scanning layer is used to divide the feature map patches that make up the image to be segmented into M*M calculation windows;

[0034] The SW-2D selective scanning layer is used to slide the calculation window of the W-2D selective scanning layer by 1 / 2 window unit length in both the height and width dimensions.

[0035] As a preferred solution of the present invention, the output expression of the W-2D selective scanning layer in the first VSS block is:

[0036] z l = W - SS2D(LN(z l-1 )) + z l-1 ,

[0037] where z l is the output of the W - 2D selective scanning layer, and z l-1 is the input of the W - 2D selective scanning layer, W is the initial window selection, SS2D() is 2D selective scanning, and LN() is normalization processing;

[0038] The output expression of the SW - 2D selective scanning layer in the second VSS block is:

[0039] Z l+1 = SW - SS2D(LN(z l )) + z l ,

[0040] where z l+1 is the output of the SW - 2D selective scanning layer, and SW is sliding window processing.

[0041] As a preferred embodiment of the present invention, the depth - separable convolutional layer includes 8 convolutional layers, and the number of channels in each layer is [16, 32, 64, 128, 128, 64, 32, 16] in sequence.

[0042] As a preferred embodiment of the present invention, the decoder includes several Patch Expanding layers and SwinVSS Block blocks;

[0043] The Patch Expanding layer is used to restore the resolution of the feature map and increase the number of feature map patches.

[0044] As a preferred embodiment of the present invention, the skip connection module includes two SwinVSS Block blocks.

[0045] An image segmentation device based on the SM - UNet image segmentation model includes at least one processor and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described in any one of the above.

[0046] Compared with the prior art, the beneficial effects of the present invention:

[0047] The present invention proposes a novel encoder-decoder Mamba framework, applies a sliding window in a UNet segmentation model based on Mamba to establish an SM-UNet image segmentation model, and realizes Swin Mamba UNet (SMUNet) for automatic medical image segmentation. It mainly combines the advantages of the linear complexity of Mamba and the cross-window interaction ability brought by the sliding window to effectively optimize the structure of the standard U-shaped architecture. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a schematic flowchart of an image segmentation method based on the SM-UNet image segmentation model described in Embodiment 1 of the present invention;

[0049] Figure 2 It is a schematic structural diagram of the SM-UNet image segmentation model in an image segmentation method based on the SM-UNet image segmentation model described in Embodiment 2 of the present invention;

[0050] Figure 3 It is a schematic structural diagram of the SwinVSS Block in an image segmentation method based on the SM-UNet image segmentation model described in Embodiment 2 of the present invention;

[0051] Figure 4 It is a schematic diagram of the effect of the shifted window method in an image segmentation method based on the SM-UNet image segmentation model described in Embodiment 2 of the present invention;

[0052] Figure 5 It is a schematic flowchart of the working process of the SS2D module in an image segmentation method based on the SM-UNet image segmentation model described in Embodiment 2 of the present invention;

[0053] Figure 6 It is a schematic structural diagram of an image segmentation device based on the SM-UNet image segmentation model using the image segmentation method based on the SM-UNet image segmentation model described in Embodiment 1 of the present invention in Embodiment 4 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The present invention will be further described in detail below in combination with test examples and specific embodiments. However, this should not be construed as limiting the scope of the above-mentioned subject matter of the present invention to the following embodiments. All technologies implemented based on the content of the present invention belong to the scope of the present invention.

[0055] TransUNet is a deep learning model that combines the Transformer and UNet architectures and is specifically used for medical image segmentation. The following is a detailed explanation of how TransUNet realizes medical image segmentation:

[0056] I. Component Units and Constituent Parts

[0057] Component Units:

[0058] Transformer Layer: Used to process sequential data and capture long - range dependencies.

[0059] CNN (Convolutional Neural Network): Used to extract local features.

[0060] MLP (Multi - Layer Perceptron): Used for non - linear transformation.

[0061] LayerNorm (Layer Normalization): Used to stabilize the training process.

[0062] MSA (Multi - Head Self - Attention Mechanism): Used to capture relationships between different parts.

[0063] Constituent Parts:

[0064] Embedded Sequence: The sequence obtained after pre - processing the input medical image.

[0065] Transformer Block: Contains multiple Transformer Layers.

[0066] CNN Block: Contains multiple convolutional layers.

[0067] Upsampling and Downsampling Module: Used to adjust the resolution of the feature map.

[0068] Segmentation Head: Used to generate the final segmentation result.

[0069] II. Connection Relationships between Components

[0070] a) The embedded sequence first undergoes normalization processing by LayerNorm.

[0071] b) The normalized sequence enters the MSA to capture global information through the multi - head self - attention mechanism.

[0072] c) The features after passing through the MSA are further processed by LayerNorm and MLP.

[0073] d) The processed features are converted into a shape suitable for CNN processing through a reshape operation.

[0074] e) The CNN block performs convolutional operations on the features to extract local features.

[0075] f) The resolution of the feature map is gradually reduced through the downsampling module to increase the receptive field.

[0076] g) During the downsampling process, the feature maps fuse features of different scales through Feature Concatenation.

[0077] h) The resolution of the feature maps is gradually restored through the upsampling module.

[0078] i) Finally, the feature maps generate segmentation results through the segmentation head.

[0079] III. Implementation and Functions of Each Unit

[0080] Transformer Layer: Captures global dependencies through the self-attention mechanism, enhancing the model's ability to process long-range information.

[0081] CNN: Extracts local features through convolutional operations, capturing the detailed information of the image.

[0082] MLP: Introduces non-linear transformations to enhance the model's expressive power.

[0083] LayerNorm: Stabilizes the training process and accelerates convergence.

[0084] MSA: Captures the relationships between different parts through the multi-head self-attention mechanism, enhancing the model's global perception ability.

[0085] Downsampling: Reduces the resolution of the feature maps through pooling or convolutional operations, increasing the receptive field.

[0086] Upsampling: Restores the resolution of the feature maps through deconvolution or interpolation operations, generating high-resolution segmentation results.

[0087] Segmentation Head: Usually composed of convolutional layers and activation functions, generating the final segmentation mask.

[0088] IV. Implementation Process of the Program Algorithm

[0089] 1. Data Preprocessing: Perform preprocessing operations such as normalizing, cropping, and scaling on medical images.

[0090] 2. Embedded Sequence Generation: Convert the preprocessed image into a sequence form as the input of the model.

[0091] 3. Transformer Block Processing:

[0092] i. Perform layer normalization on the embedded sequence.

[0093] ii. Capture global information through MSA.

[0094] iii. Perform layer normalization and MLP processing again.

[0095] 4. Feature reshape: Convert the features output by the Transformer block into a shape suitable for CNN processing.

[0096] 5. CNN block processing:

[0097] iv. Extract local features through the convolutional layer.

[0098] v. Reduce the resolution of the feature map through the downsampling module.

[0099] vi. Fuse features of different scales through feature concatenation.

[0100] 6. Upsampling module processing:

[0101] vii. Gradually restore the resolution of the feature map through the upsampling operation.

[0102] 7. The segmentation head generates the segmentation result:

[0103] viii. Generate the final segmentation mask through the convolutional layer and the activation function.

[0104] 8. Loss calculation and optimization: Calculate the loss between the segmentation result and the ground truth label, and optimize the model parameters through backpropagation.

[0105] Through the above steps, TransUNet can effectively combine the advantages of Transformer and CNN to achieve high-quality medical image segmentation.

[0106] Example 1

[0107] As Figure 1 shown, an image segmentation method based on the SM-UNet image segmentation model includes the following steps:

[0108] Obtain the image to be segmented;

[0109] Input the image to be segmented into the pre-constructed SM-UNet image segmentation model;

[0110] The SM-UNet image segmentation model outputs the image segmentation result;

[0111] Among them, the SM-UNet image segmentation model includes a convolutional module, an encoder, a skip connection module, and a decoder;

[0112] The convolutional module is used to perform convolutional segmentation processing on the image to be segmented and output several feature map patches;

[0113] The encoder includes two branches for sliding window analysis, and is used to gradually extract and compress features from several of the feature map patches;

[0114] The skip connection module is used to obtain the lost spatial details during the process;

[0115] The decoder includes two branches for performing sliding window analysis, which is used to gradually restore the resolution of the feature map, perform detail reconstruction, and output the image segmentation result.

[0116] Furthermore, the SM-UNet image segmentation model is established based on the VMamba architecture, and the VMamba architecture has the following characteristics:

[0117] - Built based on the state space model (SSM), and sequence data is processed through the Scan mechanism;

[0118] - Utilize the convolution property to achieve a global receptive field and are good at capturing long-range dependencies;

[0119] - Adopt the Cross-Scan strategy to process two-dimensional spatial relationships;

[0120] - Have the advantage of linear computational complexity.

[0121] However, SSM still has deficiencies in fine positioning and the disadvantage that information between windows is not interconnected. Therefore, the present invention introduces a sliding window mechanism into the VMamba architecture to form the SM-UNet image segmentation model of the present invention. After combining the two, the SM-UNet image segmentation model of the present invention has the following mechanisms:

[0122] 1. Local enhancement mechanism of the sliding window for SSM:

[0123] The traditional state space model (SSM) constructs global sequence dependencies through the Scan mechanism, but its continuous scanning characteristics face two inherent challenges in visual tasks:

[0124] (1) Weakening of local details: The global scanning strategy of SSM (such as cross-scan) unfolds the two-dimensional space into a one-dimensional sequence, which may blur the local spatial relationship between pixels. For example, in an edge detection task, the mutation features of adjacent pixels may be smoothed and diluted by the continuity of the scanning process.

[0125] (2) Information islands between windows: Although the global parameter sharing mechanism of the native SSM can capture long-range dependencies, it lacks an explicit local inductive bias. When dealing with complex textures, similar patterns in different windows may not be able to interact effectively, resulting in repeated calculations.

[0126] The sliding window mechanism of the present invention makes targeted improvements to the above defects in the following ways:

[0127] (1) Local feature focusing: Force the calculation to be constrained within a fixed window, and preserve the integrity of the local structure through spatial hard partitioning. Experiments show that in the key point detection task, introducing a sliding window can reduce the positioning error because it strengthens the local correlation modeling.

[0128] (2) Controllable cross-window interaction: Achieve gradient penetration between adjacent windows through the shifted window strategy. The shift operation makes the neurons at the window boundary belong to different windows in two calculations, forming an implicit information bridge. Mathematically, this is equivalent to adding a cross-window correlation constraint term to the loss function.

[0129] (3) Hierarchical feature distillation: In cooperation with hierarchical downsampling, the sliding window forms a pyramidal feature abstraction. Shallow small windows (4×4) capture details such as lines / corners, and deep large windows (14×14) integrate part-level semantics, which complements the single-scale scanning of SSM.

[0130] 2. Global compensation mechanism of SSM for the sliding window:

[0131] The pure sliding window architecture has essential limitations:

[0132] (1) Limited field of view: A single window can only capture local context. For tasks that require global reasoning, samples that are locally similar but globally contradictory may be misjudged.

[0133] (2) Inefficient long-range modeling: Rely on layer-by-layer window shifting to transmit information, and the theoretical receptive field expands slowly. While SSM can achieve fast global coverage through state transfer.

[0134] The compensation of SSM in the present invention is reflected in global state memory and dynamic receptive field adjustment. Through cross-modal feature fusion and designing a direction-aware scanning strategy, SSM can establish associations between different scanning paths. For example, in medical image analysis, the scanning path along the blood vessel direction can strengthen the global connectivity modeling of the lesion area, which forms a dynamic complement to the rigid partitioning of the sliding window.

[0135] Embodiment 2

[0136] This embodiment is a specific implementation manner of the SM-UNet image segmentation model in the image segmentation method based on the SM-UNet image segmentation model described in Embodiment 1.

[0137] The SM-UNet image segmentation model includes a convolution module, an encoder, a skip connection module, and a decoder. As Figure 2As shown, the input 2D grayscale image has a size of H×W×1 and is first segmented into patches similar to those of ViT and VMamba by a convolution module, and then converted into a 1D sequence. The initial Patch Embedding segmentation layer adjusts the feature dimension to an arbitrary size, denoted as C. These patch tokens are then processed through multiple SwinVSS blocks and Patch Merging layers to create hierarchical features. The Patch Merging layer handles downsampling and dimension increase, while the SwinVSS block focuses on learning feature representations. The output resolutions at each stage of the encoder are H / 4×W / 4×C, H / 8×W / 8×2C, H / 16×W / 16×4C, and H / 32×W / 32×8C respectively. The decoder consists of SwinVSS blocks and patch expansion layers following the encoder style, capable of outputting exactly the same feature size, thus enhancing the spatial details lost during the downsampling process through skip connections. In the encoder and decoder, 2 SwinVSS blocks are used respectively.

[0138] Furthermore, the convolution module is used to perform convolution segmentation on the image to be segmented and output a number of feature map patches; the convolution module includes a convolutional layer and a Patch Embedding segmentation layer; the convolutional layer is used to convolve the image to be segmented into a feature map; the Patch Embedding segmentation layer is used to segment the feature map and embed these small blocks into a high-dimensional space, and output a number of feature map patches of the same size.

[0139] Convolutional layer: Use a convolution kernel to perform a convolution operation on the input image. The calculation formula is:

[0140]

[0141] where I is the input image, W is the convolution kernel, Y is the output feature map, F is the convolution kernel size, and C in is the number of input channels. The convolution kernel slides on the input image, the stride S controls the sliding speed, and the padding P is used to maintain the output size. After the convolution operation, an activation function (such as ReLU) is usually applied to increase non-linearity, and the final output feature map is used as the input for the next layer to continue with feature extraction or classification tasks.

[0142] The encoder includes two branches for sliding window analysis to gradually extract and compress features of a number of the feature map patches; the encoder includes a number of SwinVSS Block blocks and a Patch Merging layer; the SwinVSS Block block has a self-attention mechanism to capture global and local features; the Patch Merging layer is used to merge adjacent feature map patches, thereby reducing the resolution of the feature map while increasing the feature dimension to capture higher-level features.

[0143] As Figure 3 shown, the SwinVSS Block block includes a first VSS block and a second VSS block connected in sequence.

[0144] The first VSS block includes a normalization (Layer Norm) layer and a fully connected layer (Linear) connected in sequence; among them, a first branch and a second branch are provided between the normalization layer and the fully connected layer; the first branch includes a fully connected layer, a depthwise separable convolution layer (Depth-wise Conv), a W-2D selective scanning layer, and a normalization (LayerNorm) layer connected in sequence; the second branch includes a fully connected layer.

[0145] The second VSS block includes a normalization (Layer Norm) layer and a fully connected layer (Linear) connected in sequence; among them, a third branch and a fourth branch are provided between the normalization layer and the fully connected layer; the third branch includes a fully connected layer, a depthwise separable convolution layer (Depth-wise Conv), a SW-2D selective scanning layer, and a normalization (LayerNorm) layer connected in sequence; the fourth branch includes a fully connected layer.

[0146] The W-2D selective scanning layer is used to divide the feature map patches that make up the image to be segmented into M*M calculation windows.

[0147] The SW-2D selective scanning layer is used to slide the calculation windows of the W-2D selective scanning layer by 1 / 2 window unit length in both the height and width dimensions.

[0148] The output expression of the W-2D selective scanning layer in the first VSS block is:

[0149] z l = W-SS2D(LN(z l-1 )) + z l-1 ,

[0150] where z l is the output of the W-2D selective scanning layer, z l-1Input for the W-2D selective scanning layer, where W is the initial window selection, SS2D() is the 2D selective scanning, and LN() is the normalization process;

[0151] The output expression of the SW-2D selective scanning layer in the second VSS block is:

[0152] Z l+1 = SW-SS2D(LN(z l )) + z l ,

[0153] where z l+1 is the output of the SW-2D selective scanning layer, and SW is the sliding window process.

[0154] Specifically, the traditional state space models (SSMs), as a linear time-invariant system function, project onto through a hidden state given as the evolution parameter, B, as the projection parameter of the state size, and there is a skip connection This model can be expressed as linear ordinary differential equations (ODEs), as shown in the formula:

[0155] h′(t) = Ah(t) + Bx(t),

[0156] y(t) = Ch(t) + Dx(t).

[0157] The discrete version of this linear model can be obtained through the zero-order hold transformation, given a time scale parameter

[0158] h t = Ah k-2 + Bx k

[0159] y t = Ch k + Dx k

[0160] A = e ΔA

[0161] B = (e ΔA - I)A -1 B

[0162] C = C

[0163] where, Using the first-order Taylor series approximation for B, we get B = (e ΔA - I)A -1 B ≈ (ΔA)(ΔA)-1 ΔB = ΔB.

[0164] The depthwise separable convolutional layer includes 8 convolutional layers, and the number of channels in each layer is [16, 32, 64, 128, 128, 64, 32, 16] in sequence.

[0165] Vision Mamba further introduces a Cross-Scan Module (CSM), and then integrates the convolutional operation into this module. In the VSS block, the input feature first encounters a linear embedding layer and then bifurcates into a dual path. One of the branches goes through depthwise separable convolution and SiLU activation, then enters the WSS2D module, and after layer normalization, merges with the other branch after SiLU activation. Different from typical vision transformers, this SwinVSS block does not use positional embeddings, but adopts a streamlined structure that removes the MLP stage, enabling denser block stacking within the same depth budget.

[0166] In the standard Transformer architecture, each token needs to calculate its relationship with all other tokens, where the computational complexity has a quadratic relationship with the number of tokens, making it unacceptable in many dense prediction and high-resolution image tasks.

[0167] For efficient modeling, the SMUNet described in this embodiment uses window-based SS2D (WSS2D) and shifted-window-based SS2D (SWSS2D). As Figure 4 shown, it is a schematic diagram of the shifted-window method for calculating self-attention in the proposed Swin Transformer architecture. In the first layer (as Figure 4 shown in a), a conventional window partitioning scheme is adopted, and Vision Mamba calculations are performed within each window. In the next layer (as Figure 4 shown in b), the window partitioning shifts, resulting in new windows. The Vision Mamba calculations in the new windows span the boundaries of the previous windows in the first layer, establishing connections between them. Among them, the working process of SS2D is as Figure 5 shown, the input patches are traversed along four different scan paths (cross-scan), and each sequence is independently processed by a separate S6 module. Then the results are merged to construct the final two-dimensional feature map (cross-merge).

[0168] In WSS2D, the input feature will be divided into non-overlapping windows, and each window contains M×M patches. WSS2D only performs Vision Mamba calculations in two directions within the local window. Suppose z l represents the output of the l-th layer of WSS2D, and the calculation is as follows:

[0169] z l = W-SS2D(LN(zl-1 )) + z l-1 。

[0170] The problem of WMSA is the lack of effective information interaction between windows. To introduce cross-window interaction without additional computation, SWMSA is used. The window configuration of SWMSA is different from the previous WMSA layer. It uses an efficient batch processing method by circularly shifting to the upper left corner. After this shift, the batch window may consist of multiple non-adjacent sub-windows in the feature map and maintains the same number of batch windows as the regular partition. When performing visual Mamba calculations within the local window in WMSA and SWMSA, the relative position bias is included in the calculation of similarity. Through this shifted window partitioning mechanism, the output of the SWSS2D module can be written as:

[0171] z l+1 = SW-SS2D(LN(z l )) + z l 。

[0172] In the encoder, the tokenized input of C dimensions undergoes feature learning through two consecutive SwinVSS blocks at a reduced resolution, maintaining the dimension and resolution. Patch merging is used three times in the encoder as the downsampling process. By splitting the input into quarter regions, connecting them, and then normalizing the dimension by layernorm each time, the number of tokens is reduced by half while the feature dimension is doubled.

[0173] The skip connection module is used to obtain the lost spatial details during the process; the skip connection module includes two SwinVSS Block blocks.

[0174] Two SwinVSS blocks are used in the bottleneck part. Skip connections are adopted at each level of the encoder and decoder to fuse multi-scale features with the upsampled output, enhancing spatial details by combining shallow and deep layers. The subsequent linear layer maintains the dimension of this integrated feature set, ensuring consistency with the upsampled resolution.

[0175] The decoder includes two branches for sliding window analysis, used to gradually restore the resolution of the feature map, perform detail reconstruction, and output the image segmentation result.

[0176] The decoder includes several Patch Expanding layers and SwinVSS Block blocks;

[0177] The Patch Expanding layer is used to restore the resolution of the feature map, increasing the number of patches in the feature map for more detailed feature extraction.

[0178] Similar to the encoder, the decoder uses two consecutive SwinVSS blocks for feature reconstruction, and uses a patch expansion layer instead of a merging layer to amplify the depth features. These layers halve the feature dimension through the initial layer, and then double the feature dimension before reorganizing and reducing them to improve the resolution, thereby enhancing the resolution.

[0179] In summary, the present invention applies a sliding window to a Mamba-based UNet segmentation model. By dividing the image into multiple local windows and performing calculations within these windows, the computational amount is greatly reduced. The cross-window interaction ability brought by the sliding window is fully utilized, thereby achieving global modeling ability to a certain extent while maintaining the efficiency of local calculations.

[0180] Example 3

[0181] This example is an actual application example of an image segmentation method using the image segmentation model based on SM-UNet described in Example 2. The specific process is as follows:

[0182] 1. Dataset

[0183] In this example, experiments are carried out using the publicly available ACDC magnetic resonance imaging (MRI) cardiac segmentation dataset from the MICCAI 2017 challenge. This dataset contains magnetic resonance imaging scans of 100 patients, and various cardiac structures are annotated, such as the right ventricle and the endocardial and epicardial walls of the left ventricle. It covers a variety of pathological conditions and is classified into five subgroups: normal, myocardial infarction, dilated cardiomyopathy, hypertrophic cardiomyopathy, and abnormal right ventricle, thus ensuring a wide distribution of feature information. Four regions of interest (ROIs) are verified in ACDC.

[0184] 2. Implementation details

[0185] This example is carried out on the Ubuntu 20.04 system, using Python 3.8.8, PyTorch 1.10, and CUDA 11.3. The hardware settings include an Nvidia GeForce RTX 4080 GPU and an Intel Core i9-10900K CPU. For the ACDC dataset, the average running time is about 5 hours, which includes data transfer, model training, and inference processes. The dataset is specifically processed for two-dimensional image segmentation. The model is trained for 10,000 iterations with a batch size of 24. The Stochastic Gradient Descent (SGD) optimizer is used with a learning rate of 0.01, a momentum of 0.9, and a weight decay set to 0.0001. The network performance is evaluated on the validation set every 200 iterations, and the model weights are saved only when new best performance is achieved on the validation set.

[0186] 3. Evaluation Metrics

[0187] The evaluation of SWin - Mamba - UNet and the baseline methods used a wide range of evaluation metrics. Similarity metrics, where higher values are preferred, include: Dice, Intersection over Union (IoU), accuracy, precision, sensitivity, and specificity, denoted by an upward arrow (↑), indicating that higher values represent better performance. Conversely, dissimilarity metrics such as the 95% Hausdorff distance (HD) and average surface distance (ASD), denoted by a downward arrow (↓), are better when lower, indicating a closer similarity between the prediction and the ground truth segmentation.

[0188]

[0189] Among them, TP represents the number of true positives, TN represents the number of true negatives, FP represents the number of false positives, and FN represents the number of false negatives.

[0190]

[0191] Among them, a and b represent the sets of points on the predicted and ground truth surfaces respectively. d(a, b) represents the Euclidean distance between two points.

[0192] 4. Results

[0193]

[0194] The experimental results are shown in the above table. According to the quantitative results, it shows that the SMUNet image segmentation model described in the present invention is more likely to predict accurate segmentation masks.

[0195] Example 4

[0196] As Figure 6 shown, an image segmentation device based on the SM - UNet image segmentation model includes at least one processor, a memory communicatively connected to the at least one processor, and at least one input - output interface communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the image segmentation method based on the SM - UNet image segmentation model described in the foregoing embodiments. The input - output interface may include a display, a keyboard, a mouse, and a USB interface for inputting and outputting data.

[0197] Those skilled in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments. The aforementioned storage medium includes various media that can store program codes, such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs.

[0198] When the above integrated unit of the present invention is implemented in the form of a software functional unit and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media that can store program codes, such as removable storage devices, ROM, magnetic disks, or optical discs.

[0199] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An image segmentation method based on the SM-UNet image segmentation model, characterized in that: The following steps are involved: Obtain the image to be segmented; Input the image to be segmented into the pre-built SM-UNet image segmentation model; The SM-UNet image segmentation model outputs an image segmentation result; The SM-UNet image segmentation model includes a convolution module, an encoder, a jump connection module and a decoder; The convolution module is used to perform convolution segmentation processing on the image to be segmented, and output a plurality of feature map patches; The encoder includes two branches for sliding window analysis, for gradually extracting and compressing features of a plurality of feature map patches; The skip connection module is used to obtain the spatial details lost in the process; The decoder includes two branches for sliding window analysis, which are used to gradually restore the resolution of the feature map, reconstruct details, and output image segmentation results.

2. According to claim 1, the image segmentation method based on the SM-UNet image segmentation model is characterized in that: The convolution module includes a convolution layer and a Patch Embedding segmentation layer; The convolution layer is used to convolute the image to be segmented into a feature map; The Patch Embedding segmentation layer is used to segment the feature map into a number of feature map patches of the same size.

3. The image segmentation method based on the SM-UNet image segmentation model according to claim 1, characterized in that: The encoder includes several SwinVSS Block blocks and PatchMerging patch merging layers; The SwinVSS Block is used to capture global and local features; The Patch Merging layer is used to merge adjacent feature map patches.

4. The image segmentation method based on the SM-UNet image segmentation model according to claim 3, characterized in that: The SwinVSS Block includes a first VSS block and a second VSS block connected in sequence; The first VSS block includes a normalization layer and a fully connected layer connected in sequence; wherein a first branch and a second branch are provided between the normalization layer and the fully connected layer; the first branch includes a fully connected layer, a depthwise separable convolutional layer, a W-2D selective scanning layer, and a normalization layer connected in sequence; and the second branch includes a fully connected layer; The second VSS block includes a normalization layer and a fully connected layer connected in sequence; wherein, a third branch and a fourth branch are provided between the normalization layer and the fully connected layer; the third branch includes a fully connected layer, a depthwise separable convolutional layer, a SW-2D selective scanning layer, and a normalization layer connected in sequence; and the fourth branch includes a fully connected layer.

5. The image segmentation method based on the SM-UNet image segmentation model according to claim 4, characterized in that: The W-2D selective scanning layer is used to divide the feature map patches constituting the image to be segmented into M*M calculation windows; The SW-2D selective scanning layer is used to slide the calculation window of the W-2D selective scanning layer by 1 / 2 of the window unit length in the height and width dimensions respectively.

6. The image segmentation method based on the SM-UNet image segmentation model according to claim 5, characterized in that: The output expression of the W-2D selective scanning layer in the first VSS block is: z l =W-SS2D(LN(z l-1 ))+z l-1 , Among them, z l is the output of the W-2D selective scanning layer, z l-1 is the input of the W-2D selective scanning layer, W is the initial window selection, SS2D() is the 2D selective scanning, and LN() is the normalization process; The output expression of the SW-2D selective scanning layer in the second VSS block is: WITH l+1 =SW-SS2D(LN(z l ))+z l , Among them, z l+1 is the output of the SW-2D selective scanning layer, and SW is the sliding window processing.

7. The image segmentation method based on the SM-UNet image segmentation model according to claim 6, characterized in that: The depthwise separable convolutional layer includes 8 convolutional layers, and the number of channels in each layer is [16, 32, 64, 128, 128, 64, 32, 16] respectively.

8. The image segmentation method based on the SM-UNet image segmentation model according to claim 6, characterized in that: The decoder includes several Patch Expanding layers and SwinVSS Blocks; The Patch Expanding layer is used to restore the resolution of the feature map and increase the number of feature map patches.

9. The image segmentation method based on the SM-UNet image segmentation model according to claim 6, characterized in that: The jump connection module includes two SwinVSS Blocks.

10. An image segmentation device based on the SM-UNet image segmentation model, characterized in that: It includes at least one processor and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Jump connection method based on combination of progressive contraction strategy and shift window

    CN117670904A

  • Road surface crack image segmentation method based on improved encoder-decoder structure

    CN118736215A

  • CT image segmentation method and system based on improved Swinin-Unet

    CN119151963A

  • Systems and methods for multimodal fusion of missing and unpaired image and tabular data

    KR1020240141126A

  • Multi-tissue segmentation and deformation for breast cancer surgery

    WO2024123484A1

Cited By

  • Double-domain Swin Mama-based generative adversarial network MRI reconstruction method

    CN120635248A