Cell image semantic segmentation method and system based on MTransUnet
By introducing MCCT and GAM modules into the Unet network, combining Transformer structure and global attention mechanism, the problems of insufficient global information extraction capabilities and poor long-distance dependency processing in cell image semantic segmentation are solved, which significantly improves the accuracy and efficiency of segmentation.
Patent Information
- Application Number
- CN202510224837.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art has problems in the semantic segmentation of cell images, such as insufficient global information extraction capability, poor long-distance dependency processing, sensitivity to small sample data sets, and susceptibility to noise interference, resulting in low segmentation accuracy and efficiency.
Using the cellular image semantic segmentation method based on MTransUnet, by introducing MCCT and GAM modules into the Unet network, combining Transformer structure and global attention mechanism, multi-scale features are extracted and feature fusion capabilities are enhanced, and the robustness to noise is improved.
It significantly improves the accuracy and efficiency of semantic segmentation of cell images, can more accurately identify complex structural relationships in cell images, reduce segmentation errors, and perform more stably when processing small sample data sets.
Smart Images

Figure CN120071346A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and specifically relates to a method and system for semantic segmentation of cell images based on MTransUnet. Background Art
[0002] In biomedical research, the semantic segmentation of cell images is crucial for analyzing the morphology, structure, and function of cells. By accurately segmenting different components in cell images, such as cell nuclei, cytoplasm, etc., key information such as the number, size, and shape of cells can be obtained, providing strong support for disease diagnosis, drug research and development, and cell biology research.
[0003] In the field of medical image analysis, the semantic segmentation of cell images is of great significance for disease diagnosis, pathological research, etc. Currently, Unet and its variant networks have been widely used in the field of medical image segmentation. However, when directly applying these networks to cell image segmentation, there are many problems. The foreground objects in cell images have characteristics such as complex biological tissue structures, small feature differentiations between cells, large cell colony densities, and large numbers.
[0004] The encoder part of the Unet network uses convolutional operations for image feature extraction. Since the receptive field size of the convolutional layer is fixed, when processing large-sized images, the global information extraction ability is poor. For example, when segmenting cell images, it is difficult to capture the long-distance dependencies between cells, resulting in situations such as missed segmentation and missegmentation when dealing with cell clusters with complex structures.
[0005] The Unet decoder uses upsampling and convolutional layers for feature fusion and restoration. This method can only capture local features and cannot handle long-distance dependencies well. When processing complex structure images, the upsampling operation causes the information lost by the pooling layer to not be fully restored, blurring the spatial relationship between adjacent pixels, and making it difficult for the neural network to distinguish the small-sized structures of densely distributed cells in the image, thereby affecting the segmentation accuracy.
[0006] In addition, in terms of data processing, during the acquisition, processing, and transmission of cell images, they are easily affected by various interferences, doped with noise and artifacts, affecting the image quality and segmentation accuracy. Moreover, there are unstained cells in the original images, which will interfere with the learning process of the model. Existing semantic segmentation models are difficult to achieve ideal segmentation effects when dealing with cell images and cannot meet the needs of medical research and clinical diagnosis. Summary of the Invention
[0007] (1) Technical Problems to be Solved
[0008] In view of the deficiencies of the prior art, the present invention provides a method and system for semantic segmentation of cell images based on MTransUnet, which solves the problems existing in the prior art of semantic segmentation of cell images, such as insufficient ability to extract global information, poor processing of long-distance dependence relationships, sensitivity to small sample datasets, and susceptibility to noise interference, and improves the accuracy and efficiency of semantic segmentation of cell images.
[0009] (2) Technical solution
[0010] The present invention specifically adopts the following technical solutions to achieve the above objectives:
[0011] A method and system for semantic segmentation of cell images based on MTransUnet, comprising the following steps:
[0012] S1. Establish a cell image dataset;
[0013] S2. Construct an MTransUnet network model;
[0014] S3. Define a loss function;
[0015] S4. Use the established dataset to train the MTransUnet network model;
[0016] S5. Use the trained model to segment cell images.
[0017] Further, in step S1, the steps of establishing the dataset are as follows:
[0018] S1-1. Collect cell images, and for the noise in the images, use a bilateral filtering algorithm for denoising;
[0019] S1-2. Normalize the images, and through color space analysis, evaluate the separability of the probability density of regions of interest (ROIs) of different objects (such as cell nuclei, cytoplasm, unstained cells, and background) in color spaces such as RGB, HSV, and Lab. Remove the interference of unstained cells on model learning by setting pixel intensity thresholds, optimize the image contrast, and make stained cells and the background easier to distinguish;
[0020] S1-3. Divide the dataset into three parts: a training set, a test set, and a validation set.
[0021] Further, in step S2, the specific steps of constructing the MTransUnet network model are as follows:
[0022] S2-1. Incorporate MCCT (Multi-Scale Channel Cross-Fusion Transformer) into the Unet network, and then propose the MTransUnet model. By introducing branches of different scales before each pooling layer, features with different receptive fields and resolutions can be extracted on different branches. The feature maps extracted by the encoder-decoder are simultaneously input into the MCCT module, and a fully convolutional layer is used to fuse the upsampled and downsampled features, strengthening the semantic connection between the encoder and the decoder. This method adopts a parallel Transformer structure in the encoder part to extract local and global information in the image, and introduces a two-way interaction design of channel information and spatial information to promote the communication between local information and global information. At the same time, the multi-scale feature information in the encoder is connected to each layer operation of the decoder to make full use of semantic information and achieve more accurate cell segmentation;
[0023] S2-2. Introduce the global attention module GAM in the decoder part of the Unet network. The decoder is responsible for upsampling to restore the image resolution. GAM helps to fuse the high-dimensional features transmitted by the encoder and the upsampled features, strengthening the cross-dimensional interaction between features of different scales. When restoring the low-resolution feature map to high resolution, GAM enables the model to more effectively integrate context information and accurately map cell features back to the original image space, improving the segmentation accuracy.
[0024] Furthermore, in step S3, the loss function is a function used to evaluate the data similarity between the prediction result and the label data. The Dice loss function is used during the training process. The Dice loss function focuses on the degree of overlap of the segmentation region, can effectively guide the model to focus on accurately outlining the cell boundary, and the Dice loss function is a symmetric loss function, which can more balancedly consider all classes, making it easier for the model to learn the features of the classes with fewer numbers. Its expression is:
[0025]
[0026] In the above formula, y represents the distribution of the true label, represents the label predicted by the model, y i and respectively represent the probabilities of the i-th class in the true label and the model-predicted label. The value of the Dice loss function is between 0 and 1, and the smaller the value, the lower the degree of overlap between the model prediction and the true label.
[0027] Furthermore, in step S4, the steps of training the MTransUnet network model using the established dataset are as follows:
[0028] S4-1. Set the training parameters of the network, including: the input resolution of the dataset input to the network, batch size (batch_size), number of training epochs (epoch), optimizer used (optimizer), learning rate (Ir) and decay rate (weight_decay) of the optimizer, etc.;
[0029] S4-2. During the training process, evaluate the model performance on the test set once every several iterations, and save the model weight file with the best performance.
[0030] Further, the specific method of step S5 is:
[0031] Prepare the cell images to be detected, select the model and load the weight file, and then segment the cell images.
[0032] (III) Beneficial Effects
[0033] Compared with the prior art, the present invention provides a cell image semantic segmentation method and system based on MTransUnet, having the following beneficial effects:
[0034] 1. The cell image semantic segmentation method and system based on MTransUnet proposed by the present invention utilize the powerful global information capture ability of Transformer, combined with the efficient fusion of multi-scale features, enabling the model to accurately analyze the complex structural relationships in cell images, effectively reducing the segmentation error, and being able to accurately identify more target cell pixels. In addition, the GAM attention mechanism further enhances the model's focusing ability on key cell features, adaptively weighting the features of different channels and spatial positions, highlighting the important information related to cell segmentation, suppressing the interference of irrelevant information, and significantly improving the accuracy and detail of segmentation.
[0035] 2. The present invention utilizes the synergistic effect of the bilateral filtering algorithm and image normalization operation to greatly improve the quality of the input image. The bilateral filtering algorithm removes image noise while well preserving the edge information of cells, avoiding segmentation errors caused by noise interference. Image normalization removes unstained cells through color space analysis and optimizes the image contrast, making the different components of cells more prominent in the image, facilitating model learning and recognition. The preprocessed image provides clearer and more representative data for model training, improving the segmentation accuracy and stability of the model from the source. Description of the Drawings
[0036] Figure 1 It is a flowchart of the cell image semantic segmentation method and system based on MTransUnet of the present invention;
[0037] Figure 2Schematic diagram of the MTransUnet network structure of the present invention;
[0038] Figure 3 Schematic diagram of the GAM attention mechanism of the present invention. Detailed implementation manners
[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] Embodiment
[0041] As Figure 1 shown, a cell image semantic segmentation method based on MTransUnet proposed in an example of the present invention includes the following steps:
[0042] S1. Establish a cell image data set, and the specific steps are as follows:
[0043] S1-1. Collect cell images. For the noise in the images, use the bilateral filtering algorithm for denoising. Calculate the spatial domain kernel and the range domain kernel through two Gaussian functions respectively, and use the product of the two as the weight to perform weighted averaging on the neighborhood pixel values, smoothing the image while maintaining the image edges;
[0044] S1-2. Normalize the images. Through color space analysis, evaluate the separability of the probability density (ROIs) of the regions of interest of different objects (such as cell nuclei, cytoplasm, unstained cells, and background) in color spaces such as RGB, HSV, and Lab. It is found that the separability between stained cells and unstained cells in the H channel is the largest. Set the pixel intensity threshold p, and select p = 120. Replace the pixels with pixel values less than this threshold in the H channel (representing unstained cells) with background pixels, remove the interference of unstained cells on model learning, optimize the image contrast, and make the stained cells and the background easier to distinguish;
[0045] S1-3. Divide the data set into three parts: a training set, a test set, and a validation set, and the training set, the test set, and the validation set account for 60%, 20%, and 20% respectively.
[0046] Further, in step S2, the specific steps for constructing the MTransUnet network model are as follows:
[0047] S2-1. Incorporate MCCT (Multi-Scale Channel Cross-Fusion Transformer) into the Unet network and use the MCCT module to replace the skip connections. In the MCCT module, the feature maps at each level are first normalized (LN) to make the input data distribution more regular and less different. After normalization, the features at each level are divided into two paths. One path is directly input into the multi-head self-attention module as the query vector matrix for the self-attention mechanism operation. The other path concatenates the features at each level on the channel dimension and inputs them into the multi-head self-attention module as the key vector matrix for operation. The transpose of the key vector matrix is the query vector matrix, and the query vector, key vector, and value vector matrices perform cross-channel attention operations;
[0048] S2-2. Introduce the global attention module GAM in the decoder part of the Unet network. In GAM, the channel attention sub-module uses 3D permutation and multi-layer perceptron to retain information and enhance cross-dimensional dependencies. The spatial attention sub-module uses two convolutional layers to fuse spatial information, removes the pooling operation to retain the feature maps, and uses grouped convolution and channel shuffle to prevent a significant increase in parameters.
[0049] Furthermore, in step S3, the loss function is a function used to evaluate the similarity between the prediction result and the label data. The Dice loss function is used during the training process. The Dice loss function focuses on the overlap degree of the segmentation regions, can effectively guide the model to focus on accurately outlining the cell boundaries, and the Dice loss function is a symmetric loss function that can more balancedly consider all classes, making it easier for the model to learn the features of the classes with fewer numbers. Its expression is:
[0050]
[0051] In the above formula, y represents the distribution of the true label, represents the label predicted by the model, y i and respectively represent the probabilities of the i-th class in the true label and the model-predicted label. The value of the Dice loss function is between 0 and 1, and the smaller the value, the lower the coincidence degree between the model prediction and the true label.
[0052] Furthermore, in step S4, the steps of training the MTransUnet network model using the established dataset are as follows:
[0053] S4-1. Set the training parameters of the network. The input resolution and batch size (batch_size) of the dataset input to the network are set to 224×224 and 16 respectively. The number of training epochs is 40000, and the model weights with the best performance are saved. The SGD optimizer is used to train the model, with an initial learning rate (Ir) of 0.01 and a decay rate (weight_decay) of 0.0005;
[0054] S4-2. During the training process, every 5000 iterations, evaluate the model performance on the test set once, and save the model weight file with the best performance.
[0055] Furthermore, the specific method of step S5 is as follows:
[0056] Prepare the cell images to be detected. After selecting the model and loading the weight file, segment the cell images.
[0057] The present invention establishes the association between the encoder and the decoder in the Unet network structure through multi-scale channel cross fusion by Transformer, replaces the original skip connection to solve the semantic gap to improve the segmentation performance. Secondly, the GAM attention module is used in the encoder part, so that when the model segments aggregated cells, it can enhance the dependence between space and channels, and can better understand the overall structure of the cell population and the unique features of each cell. The present invention improves the segmentation accuracy of the model for cell images and expands the effective segmentation area.
[0058] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A cell image semantic segmentation method based on MTransUnet, characterized by: The following steps are involved: S1. Establish a cell image dataset; S2, build the MTransUnet network model; S3. Define the loss function; S4, use the established data set to train the MTransUnet network model; S5. Use the trained model to segment the cell image.
2. The cell image semantic segmentation method based on MTransUnet according to claim 1, characterized in that: In step S1, the steps of establishing a cell image data set are as follows: collecting cell images, performing denoising processing using a bilateral filtering algorithm, normalizing the images, and dividing the data set into a training set, a test set, and a validation set.
3. The cell image semantic segmentation method based on MTransUnet according to claim 1, characterized in that: In step S2, the specific steps of constructing the MTransUnet network model are as follows: S2-1. MCCT (Multi-scale Channel Cross-fusion Transformer) is added to the Unet network, and then the MTransUnet model is proposed. By introducing branches of different scales before each pooling layer, features with different receptive fields and resolutions can be extracted on different branches; S2-2. Introduce the global attention module GAM in the decoder part of the Unet network.
4. The cell image semantic segmentation method based on MTransUnet according to claim 1, characterized in that: In step S3, the loss function is a function used to evaluate the similarity between the prediction result and the label data, and the Dice loss function is used in the training process.
5. A cell image semantic segmentation system based on MTransUnet, characterized in that: include: The data processing module and model algorithm module are used to obtain cell image datasets, perform bilateral filtering denoising and normalization on the images, build the MTransUnet model based on the Unet network, use MCCT instead of skip connections, and integrate the GAM attention mechanism.