IG-MambaUNet image segmentation model, model training method and application method thereof

By proposing the IG-MambaUNet model in medical image segmentation, combining the IG-Mamba module and the Patch Merging/Expanding module, the existing methods are solved, and the existing methods consume large computing resources and poor feature fusion effects are achieved when processing complex ganglion images, and efficient image segmentation and refined segmentation effects are achieved.

CN120088787APending Publication Date: 2025-06-03HUST SUZHOU INST FOR BRAINMATICS

Patent Information

Application Number
CN202510101512.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing medical image segmentation methods have problems such as high computing resource consumption, lack of global feature capture capabilities and poor feature range fusion effect when processing ganglion images with complex textures and fine structures.

Method used

The IG-MambaUNet image segmentation model is proposed. This model is based on the UNet network architecture, combined with the IG-Mamba module and the Patch Merging/Expanding module, and the fusion of local feature extraction, global feature extraction and gated attention modules can realize the refined segmentation of ganglion images.

Benefits of technology

The performance of ganglion image segmentation is improved, effective fusion of local and global features is achieved, the consumption of computing resources is reduced, and the precision and accuracy of segmentation results are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088787A_ABST
    Figure CN120088787A_ABST
Patent Text Reader

Abstract

The invention discloses an IG-MambaUNet image segmentation model, a model training method and an application method thereof, the IG-MambaUNet image segmentation model comprises an encoder and a decoder which are symmetrically arranged in parallel, the encoder comprises three first encoding modules and a second encoding module, the three first encoding modules are sequentially connected from top to bottom, the second encoding module is connected with the first encoding module on the bottommost layer, and the third encoding module is connected with the second encoding module on the bottommost layer. The input end of the first coding module on the topmost layer is connected with a first convolution and maximum pooling operation module, the first coding module is formed by coupling an IG-Mama module and a Patch Merging module, the second coding module comprises an IG-Mama module, the IG-Mama module comprises a local feature extraction module, a global feature extraction module and a gating attention module, and the global feature extraction module comprises a local feature extraction module, a global feature extraction module and a gating attention module. The decoder comprises a second coding module and a first coding module which are sequentially connected from bottom to top, and the output end of the first decoding module on the topmost layer is connected with a second convolution and maximum pooling operation module. According to the scheme, fine segmentation is realized, and the method can be suitable for automatic cell body segmentation of histological dyed ganglion images with complex textures and fine structures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image segmentation, and particularly to an IG-MambaUNet image segmentation model, a model training method and an application method thereof for automatically segmenting cell bodies of histological stained ganglion images. Background Art

[0002] Image segmentation methods are used to quickly identify and divide regions of interest in images, thereby accelerating the formulation of relevant treatment plans or research plans. The traditional manual segmentation technology is cumbersome and time-consuming, and must be completed by professionals with professional knowledge. Therefore, this method cannot be adapted to the situation where a large number of target images need to be processed, especially the segmentation of medical images with a large amount of data and high-precision requirements, which needs to be realized by automated segmentation technology.

[0003] Existing image segmentation methods usually target specific types of objects, such as neuron axons, vascular networks, etc., with low generalization ability. Moreover, these methods rely on developers' profound professional knowledge in the field of segmentation objects, with a long development cycle, high cost, low input-output ratio, and are difficult to be efficiently put into use.

[0004] Models based on Convolutional Neural Network (CNN) and Transformer have performed excellently in visual tasks and achieved remarkable results in medical image segmentation. For example: The U-Net network model based on CNN, known for its simple structure and strong scalability, performs excellently in extracting local features of images and has become the infrastructure of many medical image segmentation models; Another example: Vision Transformer (ViT) is a pioneering model that first applied the large language model Transformer to image processing. By dividing the image into a sequence of patches for processing, it shows advantages in global information extraction. And by integrating Swin Transformer with the U-shaped architecture, the Swin-UNet model is derived, which shows excellent performance in medical image segmentation; In addition, in order to better combine the advantages of CNN and Transformer, there are currently many studies on fusing Transformer and CNN, such as: TransUNet model, UNETR model, nnFormer model, SwinUNETR model, etc.; In existing medical image segmentation methods, there are also related cases of fusing and applying Transformer and CNN, such as a colon polyp image segmentation method based on the fusion of CNN and Transformer disclosed in the invention patent with the authorization announcement number CN115018824B; And a medical image segmentation method based on a CNN-Transformer parallel encoder disclosed in the invention patent with the authorization announcement number CN118297961B, etc.

[0005] Although many deep learning models have been proposed currently and show strong feature extraction performance in medical image segmentation tasks, there are still some problems when dealing with ganglion images with complex textures and fine structures:

[0006] (1) Limited by its position encoding and attention mechanism, in Transformer, as the number of sequence patches increases during calculation, the self-attention mechanism will cause the consumption of spatial and temporal resources to increase quadratically, requiring a large amount of computing resources during the training and inference processes, which greatly limits its application in resource-constrained situations.

[0007] (2) Although CNN performs excellently in extracting local features, it lacks the ability to capture global features and has the problem of losing some feature information during segmentation.

[0008] (3) Insufficient fusion of feature ranges: The mutual fusion effect of feature ranges based on different network models has not been fully explored currently, resulting in poor fusion effects for complex features.

[0009] In recent years, State Space Models (SSMs) have also been widely studied. They can linearly model one-dimensional sequences and learn long-range dependencies. SSMs have achieved good performance in data analysis tasks of continuous long sequences such as Natural Language Processing (NLP). The existing Mamba model further optimizes the ability of SSMs in discrete data modeling, effectively modeling long-range dependencies between complex features through a selection mechanism and a hardware-aware algorithm, providing a new alternative to Transformers. However, the mutual fusion effect of the usage feature range based on the Mamba model has not been fully explored, resulting in insufficient fusion effect of complex features. Summary of the Invention

[0010] Therefore, to solve the above problems and achieve the following, the present invention provides an IG-MambaUNet image segmentation model, a model training method, and an application method thereof.

[0011] The present invention is implemented through the following technical solutions:

[0012] The IG-MambaUNet image segmentation model is based on the UNet network architecture and includes an encoder and a decoder arranged symmetrically side by side. The encoder includes three first encoding modules connected in sequence from top to bottom and a second encoding module connected to the bottommost first encoding module. The topmost first encoding module serves as the input, and the input end of the topmost first encoding module is connected to a first convolution and max pooling operation module. The first encoding module is composed of a coupling of an IG-Mamba module and a PatchMerging module. The second encoding module includes an IG-Mamba module. The IG-Mamba module includes a local feature extraction module, a global feature extraction module, and a gated attention module. The decoder includes three first decoding modules connected in sequence from bottom to top and a second decoding module connected to the bottommost first decoding module. The topmost first decoding module serves as the output, and the output end of the topmost first decoding module is connected to a second convolution and max pooling operation module. The first decoding module is composed of a coupling of a global feature extraction module and a Patch Expanding module. The second decoding module includes a global feature extraction module. Each first encoding module is skip-connected to its symmetrically arranged first decoding module. The output end of the second encoding module is connected to the input end of the second decoding module.

[0013] The training method of the IG-MambaUNet image segmentation model includes the following steps:

[0014] Obtain a histological stained ganglion image dataset, uniformly preprocess all ganglion images in the dataset, adjust the size of the ganglion images, and perform data augmentation on the ganglion images;

[0015] Input the dataset into the IG-MambaUNet image segmentation model as described above, and train the IG-MambaUNet image segmentation model;

[0016] Further train the IG-MambaUNet image segmentation model through the loss function.

[0017] Preferably, the data augmentation process includes vertical flipping, horizontal flipping, and random rotation.

[0018] Preferably, the loss function includes the Dice loss function and the cross-entropy loss function, and weights of 1 and 1 are respectively assigned to the two loss functions. Set the initial learning rate to 0.001, the minimum learning rate to 0.00001, and use the cosine annealing learning rate scheduler to adaptively adjust the learning rate.

[0019] The application method of the IG-MambaUNet image segmentation model, which is applicable to the automatic segmentation of cell bodies of histological stained ganglion images, includes the following steps:

[0020] S1: Preprocess the ganglion image to be segmented, adjust the size of the ganglion image, and perform data augmentation on the ganglion image;

[0021] S2: Input the data-augmented ganglion image into the IG-MambaUNet image segmentation model obtained by training with the training method of the IG-MambaUNet image segmentation model as described above, and perform automatic segmentation of the cell body of the ganglion image;

[0022] Among them, step S2 includes:

[0023] S21: Input the data-augmented ganglion image into the first convolution and max-pooling operation module for preliminary feature extraction;

[0024] S22: The ganglion image after extracting the preliminary features is output after traversing the first encoding module and the second encoding module from top to bottom starting from the topmost first encoding module;

[0025] S23: The ganglion image output from the bottommost first encoding module is output after traversing the second decoding module and the first decoding module from bottom to top starting from the bottommost first decoding module;

[0026] S24: The ganglion image output from the topmost first decoding module is input into the second convolution and max-pooling operation module and the final result is output.

[0027] Preferably, in step S22, the working process of each of the first encoding modules includes the following steps:

[0028] Input the ganglion images in parallel into the local feature extraction module and the global feature extraction module of the IG-Mamba module to extract features, and fuse the ganglion images output from the two paths;

[0029] Input the ganglion image after fusing features into the gated attention module, enhance the fused features by suppressing invalid information, and output the ganglion image after enhancing the features;

[0030] Input the ganglion image after enhancing the features into the Patch Merging module, increase the number of feature channels of the image and reduce the image feature size, and then output.

[0031] Preferably, in step S23, the working process of each of the first decoding modules includes the following steps:

[0032] Input the ganglion image into the Patch Expanding module to reduce the number of feature channels of the ganglion image and increase the image feature size;

[0033] Input the ganglion image into the global feature extraction module, perform a linear calculation on the feature map of the ganglion image, and output an image close to semantic segmentation.

[0034] Preferably, the local feature extraction module is implemented based on a convolutional neural network, and its working process includes the following steps:

[0035] Perform convolution and rearrangement on the input ganglion image, adjust its size to (C, HW), and obtain a preliminary feature image;

[0036] Expand the number of channels of the preliminary feature image by four times, so that its size becomes (4C, HW);

[0037] Use shared-weight convolution to restore the dimension of the preliminary feature image, restore the preliminary feature image to the original size through rearrangement operation, and output the preliminary feature image again through convolution operation, and add the preliminary feature image to the residual information;

[0038] Divide the preliminary feature image into four parts in the channel dimension, and obtain multi-scale information of the image through dilated convolution;

[0039] Perform a splicing operation in the channel dimension to restore the preliminary feature image to the original size, and perform an interaction on the multi-scale feature information of the image through convolution operation.

[0040] Preferably, the working process of the global feature extraction module includes the following steps:

[0041] The input ganglion image is processed by layer normalization and then input in parallel to two branches;

[0042] In the first branch, the input ganglion image passes through a linear layer and an activation function and then outputs a ganglion image;

[0043] In the second branch, the input ganglion image successively passes through a linear layer, a depthwise separable convolution, and an activation function, and then enters the SS2D module to be unfolded into a sequence in four different directions. The S6 module performs dynamic selective feature extraction on the sequence, scans the information in each direction, and then restores the sequence into a ganglion image with the same size as the input ganglion image through a scan merging operation;

[0044] The ganglion image output by the second branch is processed by layer normalization again and multiplied element-wise with the ganglion image output by the first branch to fuse the ganglion images output by the two branches;

[0045] The fused ganglion image is mixed with a linear layer and output through a residual connection to obtain a ganglion image with global features extracted.

[0046] Preferably, after the ganglion image is input into the gated attention module, the gated attention module generates an attention map with the same shape as the input ganglion image through a depthwise separable convolution and suppresses the feature information of unimportant regions transmitted in the fused features of the local feature extraction module and the global feature extraction module.

[0047] The beneficial effects of the technical solution of the present invention are mainly reflected in:

[0048] 1. In the IG-MambaUNet image segmentation model, in the encoding part, the first convolution and max pooling operation module is used to initially identify the ganglion image features, and then it is iterated four times in the IG-Mamba network structure to learn the deep feature information of the image; in the decoding part, the Vision Mamba network is used to gradually restore the deep feature information, and the second convolution and max pooling operation module is further used to refine the segmentation result to ensure that the segmentation output maintains the same fineness as the true binary label image.

[0049] 2. In the IG-Mamba module, the image is input in parallel to the local feature extraction module and the global feature extraction module. The reverse surface attention network and dilated convolution built by CNN are used, and the local and global feature information in the histological staining image is obtained and fused through the Vision Mamba module. Moreover, the unimportant regions in the fused feature information are suppressed by the gated attention module, enabling the model to focus more on important information. The characteristics of the global feature network and the local feature network are fully integrated, achieving refined segmentation, improving the segmentation performance, enabling targeted segmentation and recognition of histological staining ganglion images, and ensuring the fusion effect of local and global features in the ganglion images. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 is a schematic diagram of the architecture of the IG-MambaUNet image segmentation model;

[0051] Figure 2 is a schematic diagram of the structure of the IG-Mamba module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] To clearly and detailedly demonstrate the objectives, advantages, and features of the present invention, it will be illustrated and explained through the non-limiting description of the following preferred embodiments. This embodiment is only a typical example of applying the technical solution of the present invention, and any technical solutions formed by equivalent replacement or equivalent transformation fall within the scope of protection required by the present invention.

[0053] At the same time, it is stated that in the description of the solution, it should be noted that the orientation or positional relationship indicated by terms such as "top layer" and "bottom layer" is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of description and simplification of the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention.

[0054] In addition, the terms "first" and "second" in this solution are only used for descriptive purposes and cannot be understood as indicating or implying the ranking of importance or implicitly indicating the quantity of the indicated technical features. Therefore, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the present invention, the meaning of "a plurality" is two or more, unless otherwise specifically defined.

[0055] The present invention discloses an IG-MambaUNet image segmentation model, as Figure 1As shown, the model is based on the UNet network architecture and includes an encoder and a decoder arranged symmetrically side by side. The encoder includes three first encoding modules connected in sequence from top to bottom, and a second encoding module connected to the bottommost first encoding module. The topmost first encoding module serves as the input, and the input end of the topmost first encoding module is connected to the first convolution and max pooling operation module. The first encoding module is composed of a coupling of an IG-Mamba module and a Patch Merging module. The second encoding module includes an IG-Mamba module, as Figure 2 shown. The IG-Mamba module includes a local feature extraction module, a global feature extraction module, and a gated attention module. The decoder includes three first encoding modules connected in sequence from bottom to top, and a second encoding module connected to the bottommost first encoding module. The topmost first decoding module serves as the output, and the output end of the topmost first decoding module is connected to the second convolution and max pooling operation module. The first decoding module is composed of a coupling of a global feature extraction module and a PatchExpanding module. The second decoding module includes a global feature extraction module. Each layer of the first encoding module is skip-connected to the symmetrically arranged first decoding module. The output end of the second encoding module is connected to the input end of the second decoding module.

[0056] As Figure 1 shown, in the encoder part, the features of the ganglion image are initially extracted by the first convolution and max pooling operation module, and the features of the ganglion image are iteratively processed four times through four layers of IG-Mamba modules to learn the deep feature information of the image. In the decoder part, the deep features extracted in the previous step are restored layer by layer through four layers of global feature extraction modules, and the second convolution and max pooling operation module is used to further refine the segmentation result to ensure that the final output result after segmentation maintains the same fineness as the binary label image. Among them, both the first convolution and max pooling operation module and the second convolution and max pooling operation module are two-layer convolution architectures. The local feature extraction module is implemented based on a convolutional neural network (CNN), abbreviated as the IEAD module. The global feature extraction module uses the Vision Mamba model, and the gated attention module uses the Gated Attention model.

[0057] The present invention also discloses a training method for the IG-MambaUNet image segmentation model, including the following steps:

[0058] Obtain a histological stained ganglion image dataset, uniformly preprocess all ganglion images in the dataset, adjust the size of the ganglion images, and perform data augmentation on the ganglion images. In a preferred embodiment, the size of all ganglion images in the dataset is adjusted to 512×512, and the data augmentation processing includes vertical flipping, horizontal flipping, and random rotation;

[0059] Input the preprocessed dataset ganglion images into the IG-MambaUNet image segmentation model as described above, and train the IG-MambaUNet image segmentation model;

[0060] Further train the IG-MambaUNet image segmentation model through the loss function, calculate the loss value using the loss function, and update the parameters of the model according to the loss value, so as to further reduce the prediction error. In a preferred embodiment, the loss function includes the Dice loss function and the cross-entropy loss function, and weights of 1 and 1 are respectively assigned to the two combined loss functions. Set the initial learning rate to 0.001, the minimum learning rate to 0.00001, and use the cosine annealing learning rate scheduler to adaptively adjust the learning rate. Among them, the cross-entropy loss function performs well in dealing with classification problems, while the Dice loss function is more effective in dealing with image segmentation problems. By combining these two loss functions, the robustness of the model can be enhanced.

[0061] The present invention also discloses an application method of the IG-MambaUNet image segmentation model, which is applicable to the automatic segmentation of cell bodies of histological stained ganglion images. Among them, the IG-MambaUNet image segmentation model obtained by training using the training method of an IG-MambaUNet image segmentation model disclosed above is adopted in this application method.

[0062] The application method of the IG-MambaUNet image segmentation model includes the following steps:

[0063] S1: Preprocess the ganglion image to be segmented, adjust the size of the ganglion image, and perform data augmentation on the ganglion image. In an embodiment, the size of all ganglion images in the dataset is adjusted to 512×512, and the data augmentation processing includes vertical flipping, horizontal flipping, and random rotation;

[0064] S2: Input the data-augmented ganglion image into the IG-MambaUNet image segmentation model obtained by training using the training method of the IG-MambaUNet image segmentation model as described above, and perform automatic segmentation of the cell bodies of the ganglion image;

[0065] Among them, step S2 includes:

[0066] S21: Input the ganglion image after data augmentation into the first convolutional and max pooling operation module, and perform preliminary feature extraction of the ganglion image after passing through two convolutional layers;

[0067] S22: The ganglion image after extracting preliminary features is output after traversing the first encoding module and the second encoding module from top to bottom starting from the topmost first encoding module, so that the ganglion image features pass through four IG-Mamba modules in sequence to complete four iterations and learn the deep feature information of the image;

[0068] S23: The ganglion image output from the bottommost first encoding module is output after traversing the second decoding module and the first decoding module from bottom to top starting from the bottommost first decoding module. Among them, the ganglion image restores the deep features extracted in the previous step layer by layer through four global feature extraction modules.

[0069] S24: The ganglion image output from the topmost first decoding module is input into the second convolutional and max pooling operation module and the final result is output.

[0070] In some embodiments, the working process of each of the first encoding modules in step S22 includes the following steps:

[0071] Input the ganglion image in parallel into the local feature extraction module and the global feature extraction module of the IG-Mamba module to extract features, and fuse the ganglion images output from the two paths;

[0072] Input the ganglion image after fusing features into the gated attention module, enhance the fused features by suppressing invalid information, and output the ganglion image after enhancing the features;

[0073] Input the ganglion image after enhancing the features into the Patch Merging module, increase the number of feature channels of the image and reduce the image feature size, and then output.

[0074] In some embodiments, in step S23, the working process of each of the first decoding modules includes the following steps:

[0075] Input the ganglion image into the Patch Expanding module to reduce the number of feature channels of the ganglion image and increase the image feature size;

[0076] Input the ganglion image into the global feature extraction module, perform linear calculation on the feature map of the ganglion image, and output an image close to semantic segmentation.

[0077] Such as Figure 2As shown, in a preferred embodiment, the local feature extraction module (IEAD module) in the IG-Mamba module is implemented based on a convolutional neural network, effectively enhancing the model's ability to express local features and capture details when processing image data. Its working process includes the following steps:

[0078] Perform convolution and rearrangement on the input ganglion image, adjust its size to (C, HW), and obtain a preliminary feature image;

[0079] Expand the number of channels of the preliminary feature image by four times, making its size become (4C, HW);

[0080] Use shared-weight convolution to restore the dimension of the preliminary feature image, restore the preliminary feature image to its original size through rearrangement operation, and output the preliminary feature image again through convolution operation, and add the preliminary feature image to the residual information;

[0081] Divide the preliminary feature image into four parts in the channel dimension, and obtain multi-scale information of the image through dilated convolution;

[0082] Perform a splicing operation in the channel dimension to restore the preliminary feature image to its original size, and make the multi-scale feature information of the image interact through convolution operation.

[0083] As Figure 2 As shown, in a preferred embodiment, the global feature extraction module uses the Vision Mamba model. The working process of the global feature extraction module includes the following steps:

[0084] After layer normalization processing, the input ganglion image is connected in parallel and input into two branches;

[0085] In the first branch, the input ganglion image passes through a linear layer and an activation function and then outputs a ganglion image;

[0086] In the second branch, the input ganglion image passes through a linear layer, a depthwise separable convolution, and an activation function in sequence, and then enters the SS2D module and unfolds into a sequence along four different directions. The S6 module performs dynamic selective feature extraction on the sequence, and after scanning the information in each direction, the sequence is restored to a ganglion image of the same size as the input ganglion image through a scan merge operation;

[0087] The ganglion image output by the second branch is subjected to layer normalization processing again, and is multiplied element by element with the ganglion image output by the first branch to fuse the ganglion images output by the two branches;

[0088] The fused ganglion image is mixed with a linear layer and output the ganglion image after extracting global features through a residual connection.

[0089] As Figure 2 shown, in some embodiments, the gated attention module in the IG-Mamba module adopts a GatedAttention model. After inputting the ganglion image into the gated attention module, an attention map with the same shape as the input ganglion image is generated through depthwise separable convolution, and the feature information of unimportant regions transmitted in the fused features of the local feature extraction module and the global feature extraction module is suppressed, so that the model pays more attention to key feature information and improves the model's feature extraction ability.

[0090] To verify the automatic segmentation effect of the cell bodies of histological stained ganglion images by the IG-MambaUNet image segmentation model obtained by training with the training method of an IG-MambaUNet image segmentation model disclosed above, the following experiments are carried out:

[0091] 1. Obtain the dataset and preprocess the dataset:

[0092] Among them, according to the above training method, a dataset of images of superior cervical ganglion cells of monkeys is provided. In this dataset, the size of each original image is 6144*5120, and the resolution is 0.32μm*0.32μm*1μm;

[0093] Referring to the training method and application method of the above IG-MambaUNet image segmentation model, the original images are cropped. After cropping, the size of all images in the dataset is 512*512 pixels; 1103 images are randomly selected from a set of datasets, and labels are made using Amira software. Among them, 719 images are used for training, and the remaining 384 images are used for testing.

[0094] 2. Comparison of evaluation metrics and methods

[0095] The segmentation performance of the model was measured using Specificity (Spe), Sensitivity (Sen), Mean Intersection over Union (mIoU), Accuracy (Acc), Precision (Pre), and Dice Similarity Coefficient (DSC). Specificity measures the ability of the algorithm to correctly exclude non-neuronal regions, that is, the proportion of correct predictions among negative samples; Sensitivity represents the proportion of correctly identified true positive samples, reflecting the model's ability to identify neurons; the Intersection over Union quantifies the overlap between the prediction and the ground truth sample, and the higher the mIoU, the closer the segmentation result is to the ground truth; Accuracy represents the proportion of correct predictions in the overall prediction, reflecting the overall segmentation performance of the model; Precision represents the proportion of correct predictions among positive samples, used to evaluate the accuracy of the model in identifying positive samples; the DSC coefficient is used to measure the similarity between the prediction result and the ground truth annotation, and the higher the DSC, the closer the segmentation result of the model is to the actual situation.

[0096] The specific formulas for the above evaluation metrics are as follows:

[0097]

[0098]

[0099]

[0100]

[0101]

[0102]

[0103] Where: TP, TN, FP, and FN are the numbers of true positives, true negatives, false positives, and false negatives, respectively. To verify the effectiveness of the IG-MambaUNet image segmentation model, model training method, and its application method disclosed in the present invention, the application method of the IG-MambaUNet image segmentation model disclosed in the present invention and several mainstream segmentation methods in recent years were tested using the dataset of the processed images of the superior cervical ganglion neurons of monkeys in 1 above, and the test results shown in Table 1 were obtained according to the above evaluation metrics:

[0104] Table 1: Results of the comparative experiment;

[0105]

[0106] As can be seen from the test results in Table 1, the application method of the IG-MambaUNet image segmentation model disclosed in the present invention is significantly superior to the image segmentation methods of other models in the prior art in terms of various performances.

[0107] To further verify the effectiveness of the present solution, the application method of the IG-MambaUNet image segmentation model disclosed in the present invention and other mainstream segmentation networks in the prior art were compared again using the publicly available dataset Gland; among them, the training images provided by the Gland dataset are 85, and the test images are 80. The test results are shown in Table 2, where "-" indicates that the original paper did not provide the value of this indicator.

[0108] Table 2: Comparative experiment results;

[0109]

[0110]

[0111] The above test results show that the application method of the IG-MambaUNet image segmentation model disclosed in the present invention is efficient and feasible in medical image segmentation tasks compared with the existing mainstream segmentation methods, and at the same time further improves the segmentation accuracy, realizes a more refined segmentation of the target area, thus significantly improving the accuracy of segmentation.

[0112] There are still various implementation manners of the present invention. All technical solutions formed by equivalent transformation or equivalent substitution fall within the protection scope of the present invention.

Claims

1. IG-MambaUNet image segmentation model, which is based on the UNet network architecture and features: The invention comprises an encoder and a decoder which are symmetrically arranged in parallel, wherein the encoder comprises three first encoding modules which are sequentially connected from top to bottom and a second encoding module which is connected to the first encoding module at the bottom layer, the first encoding module at the top layer is used as input, and the input end of the first encoding module at the top layer is connected to a first convolution and maximum pooling operation module, the first encoding module is formed by coupling an IG-Mamba module with a Patch Merging module, the second encoding module comprises an IG-Mamba module, and the IG-Mamba module comprises a local feature extraction module, a global feature extraction module and a gated attention module; the decoder comprises three first encoding modules which are sequentially connected from bottom to top and a second encoding module which is connected to the first encoding module at the bottom layer, the first decoding module at the top layer is used as output, and the output end of the first decoding module at the top layer is connected to a second convolution and maximum pooling operation module, the first decoding module is formed by coupling a global feature extraction module with a Patch Expanding module, the second decoding module comprises a global feature extraction module, the first encoding module of each layer is jump-connected to the first decoding module which is symmetrically arranged therewith, and the output end of the second encoding module is connected to the input end of the second decoding module.

2. The training method of the IG-MambaUNet image segmentation model is characterized by: The following steps are involved: Obtain a dataset of histologically stained ganglion images, uniformly preprocess all ganglion images in the dataset, adjust the size of the ganglion images, and perform data enhancement on the ganglion images; Inputting the data set into the IG-MambaUNet image segmentation model as claimed in claim 1, and training the IG-MambaUNet image segmentation model; The IG-MambaUNet image segmentation model is further trained through the loss function.

3. The training method of the IG-MambaUNet image segmentation model according to claim 2, characterized in that: The data enhancement processing includes vertical flipping, horizontal flipping and random rotation.

4. The training method of the IG-MambaUNet image segmentation model according to claim 2, characterized in that: The loss functions include Dice loss function and cross entropy loss function, and weights of 1 and 1 are assigned to the two loss functions respectively. The initial learning rate is set to 0.001, the minimum learning rate is set to 0.00001, and the cosine annealing learning rate scheduler is used to adaptively adjust the learning rate.

5. An application method of the IG-MambaUNet image segmentation model is suitable for automatic segmentation of cell bodies in histologically stained ganglion images, characterized in that it includes the following steps: S1: preprocessing the ganglion image to be segmented, adjusting the size of the ganglion image, and performing data enhancement processing on the ganglion image; S2: inputting the data-enhanced ganglion image into the IG-MambaUNet image segmentation model obtained by training the IG-MambaUNet image segmentation model training method according to any one of claims 2 to 4, and automatically segmenting the ganglion image into cell bodies; Wherein, step S2 comprises: S21: inputting the data-enhanced ganglion image into the first convolution and maximum pooling operation module for preliminary feature extraction; S22: The ganglion image after the preliminary features are extracted is output after traversing the first encoding module and the second encoding module from top to bottom in sequence from the first encoding module at the top layer; S23: The ganglion image outputted from the first encoding module at the bottom layer traverses the second decoding module and the first decoding module in order from bottom to top from the first decoding module at the bottom layer, and then is outputted; S24: The ganglion image output from the first decoding module at the top layer is input into the second convolution and maximum pooling operation module, and the final result is output.

6. The application method of the IG-MambaUNet image segmentation model according to claim 5, characterized in that: In step S22, the working process of each of the first encoding modules includes the following steps: The ganglion images are input in parallel to the local feature extraction module and the global feature extraction module of the IG-Mamba module to extract features, and the ganglion images output by the two paths are fused; The ganglion image after fusion features is input into the gated attention module, the fusion features are enhanced by suppressing invalid information, and the ganglion image after enhanced features is output; The ganglion image with enhanced features is input into the Patch Merging module to increase the number of feature channels of the image and reduce the image feature size, and then output.

7. The application method of the IG-MambaUNet image segmentation model according to claim 5, characterized in that: In step S23, the working process of each of the first decoding modules includes the following steps: Input the ganglion image into the Patch Expanding module to reduce the number of feature channels of the ganglion image and increase the image feature size; The ganglion image is input into the global feature extraction module, the feature map of the ganglion image is linearly calculated, and an image close to semantic segmentation is output.

8. The application method of the IG-MambaUNet image segmentation model according to claim 5, characterized in that: The local feature extraction module is implemented based on a convolutional neural network, and its working process includes the following steps: Convolve and rearrange the input ganglion image, resize it to (C, HW), and obtain a preliminary feature image; Expand the number of channels of the preliminary feature image by four times so that its size becomes (4C, HW); The shared weight convolution is used to restore the dimension of the preliminary feature image, the preliminary feature image is restored to its original size through a rearrangement operation, and the preliminary feature image is output through a convolution operation again, and the preliminary feature image is added to the residual information; The preliminary feature image is divided into four parts in the channel dimension, and the multi-scale information of the image is obtained through dilated convolution; The splicing operation is performed on the channel dimension to restore the preliminary feature image to its original size, and the convolution operation is used to make the multi-scale feature information of the image interact.

9. The application method of the IG-MambaUNet image segmentation model according to claim 5, characterized in that: The working process of the global feature extraction module includes the following steps: The input ganglion image is processed by layer normalization and then input into two branches in parallel; In the first branch, the input ganglion image is passed through a linear layer and an activation function to output the ganglion image; In the second branch, the input ganglion image passes through the linear layer, depthwise separable convolution and activation function in sequence, and then enters the SS2D module to expand into a sequence along four different directions. The S6 module performs dynamic selective feature extraction on the sequence, and after scanning the information in each direction, the sequence is restored to a ganglion image of the same size as the input ganglion image through a scan merge operation; The ganglion image output by the second branch is processed again by layer normalization, and is element-wise multiplied with the ganglion image output by the first branch to fuse the ganglion images output by the two branches; The fused ganglion images are mixed with a linear layer and the ganglion images after global features are extracted are output through a residual connection.

10. The application method of the IG-MambaUNet image segmentation model according to claim 5, characterized in that: After the gated attention module inputs the ganglion image into the gated attention module, it generates an attention map with the same shape as the input ganglion image through depthwise separable convolution, and suppresses feature information of unimportant areas transmitted in the fusion features of the local feature extraction module and the global feature extraction module.

Citation Information

Patent Citations

  • A colonoscopy polyp image segmentation method based on CNN and Transformer fusion

    CN115018824B

  • Medical image segmentation method based on CNN-Transformer parallel encoder

    CN118297961B

Cited By

  • Target tracking method and system based on Mama visual hybrid module

    CN120298458A

  • Medical image segmentation method and imaging method based on Mama network

    CN120689296A

  • Medical image segmentation method based on mamba network and imaging method

    CN120689296B

  • Remote sensing image segmentation method and device

    CN120997496A

  • Remote sensing image segmentation method and device

    CN120997496B