Image segmentation method based on improved TransUnet

By introducing CrissCross Attention, SimAM, ECA and SCOT modules into the medical image segmentation network, the problem of redundancy and large amount of calculation of jump connection information is solved, the accuracy and efficiency of image segmentation are improved, and global feature modeling and redundant feature filtering are realized.

CN120235887APending Publication Date: 2025-07-01GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510271740.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-09
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing medical image segmentation networks such as U-Net and TransUNet have problems such as redundancy of jump connection information and large amounts of computation in practical applications. The convolution operation is limited to the extraction of local context information, making it difficult to effectively deal with complex scenarios that require global semantic understanding.

Method used

A lightweight medical image segmentation network ES-TransUNet is proposed. By introducing the CrissCross Attention mechanism, it captures long-distance dependencies, uses SimAM and ECA modules to optimize the Transformer structure, uses Dynamic Upsampling for efficient upsampling, and introduces SCOT modules to filter redundant features in the jump connection part.

Benefits of technology

ES-TransUNet is better than the existing TransUNet model in terms of efficiency and performance, significantly improving the effect of medical image segmentation and achieving a good balance of performance and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235887A_ABST
    Figure CN120235887A_ABST
Patent Text Reader

Abstract

The invention discloses a medical image segmentation method based on an improved TransUNet. The medical image segmentation method comprises the following four steps: S1, acquiring an image and preprocessing the image; s2, constructing a medical image segmentation model based on the improved TransUNet; s4, training a medical image segmentation model of the improved TransUNet; and S5, testing the model. According to the medical image segmentation method based on the improved TransUNet, the improved TransUnet segmentation model is constructed, a Synapse data set is used for training and testing, the advantages of the method in segmentation precision, segmentation speed and calculation amount are verified, and the method plays an important role in improving the medical image segmentation performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision images, and specifically relates to an image segmentation method based on improved TransUnet. Background Art

[0002] In the field of medical image analysis, image segmentation is the basis for understanding and processing imaging data, especially with the rapid development of computer vision technology. Medical image segmentation can not only achieve accurate identification of anatomical structures, but also provide key information for doctors to help them make more accurate diagnosis and treatment decisions. Especially in CT images, accurate multi-organ segmentation can significantly improve the diagnostic accuracy of diseases and optimize the treatment plan for patients. With the continuous evolution of deep learning technology, especially the widespread application of convolutional neural networks (CNNs), traditional image processing methods have been surpassed by deep learning.

[0003] As a classic network in the field of medical image segmentation, the unique symmetric encoder-decoder architecture and skip connection design of U-Net enable the network to effectively fuse fine-grained features and high-level semantic features, thus improving the accuracy and precision of segmentation. However, the introduction of skip connections, while helping to retain important detail information to some extent, may also lead to the accumulation of redundant information, which in turn affects the segmentation accuracy. In addition, the local convolution operation in U-Net limits the ability to extract global context information, which may not perform ideally when dealing with some complex scenarios that require global semantic understanding.

[0004] The Transformer structure performs well in global feature modeling and has been gradually applied to the field of medical image segmentation. TransUNet combines Transformer and U-Net, and uses the self-attention mechanism to enhance the global feature modeling ability and improve the segmentation accuracy. However, TransUNet still faces the problems of redundant skip connection information and high computational complexity.

[0005] To address the limitations of convolutional extraction of local information and the redundancy in the skip connections of the U-shaped network, this paper proposes a lightweight medical image segmentation network, ES-TransUNet (Efficient Channel Attention and Simple-TransUNet). In the encoder part of this network, the CrissCross Attention (CCA) mechanism is introduced to capture long-range dependencies in the image, and the SimAM (Similarity-Aware Activation Module) and ECA (Efficient Channel Attention) modules are used to optimize the Transformer structure. In the decoder part, the Dynamic Upsampling (Dysample) sampler is adopted for efficient upsampling, combining low-level and high-level information to reduce the computational burden of dynamic convolution. In addition, the SCOT (Simple Contextual Transformer) block is introduced before the skip connection fusion to filter redundant features and enhance cross-channel fusion of information. Through these improvements, ES-TransUNet is superior to the existing TransUNet models in terms of efficiency and performance. Summary of the Invention

[0006] The purpose of the present invention is to provide a medical image segmentation method based on an improved TransUNet to solve the problems of large redundancy of skip connection information and high computational complexity in the practical application of the U-shaped network structure (such as U-Net) in the above-mentioned background technology. While improving the segmentation accuracy, the segmentation efficiency is also improved.

[0007] To achieve the above purpose, the technical solutions adopted by the present invention include the following steps:

[0008] S1: Collect images and preprocess the images;

[0009] S2: Construct a medical image segmentation model based on the improved TransUnet;

[0010] S3: Train the improved TransUnet model;

[0011] S4: Test the model.

[0012] The specific steps of the step S1 include the following steps:

[0013] S1.1: Select the images in the Synapse dataset as the dataset for this experiment, and divide the dataset into a training set and a test set according to a certain ratio;

[0014] S1.2: Resize the dataset images to a unified size of 224×224;

[0015] The specific steps of step S2 are as follows:

[0016] S2.1: Add the Criss-Cross Attion mechanism before each convolutional part of the encoder;

[0017] S2.2: Replace the multi-head attention (MSA) in the Transformer block with SimAM, and replace the multi-layer perceptron (MLP) with ECA-Net;

[0018] S2.3: Replace the bilinear interpolation upsampling in the decoder with Dysample sampling;

[0019] S2.4: Use the Simple Contextual Transformer (SCOT) module to filter redundant features in the skip connection part

[0020] The specific steps of step S3 are as follows:

[0021] S3.1: Set the training parameters

[0022] S3.2: Put the training set images in the dataset into the improved TransUnet image segmentation model for training;

[0023] S3.3: Train the model according to the set parameters, observe the change trend of the loss function, adjust the learning rate of the model training until the change of the loss function tends to be stable, and obtain the final trained model;

[0024] Step S4 uses the test set in the dataset to test the trained segmentation model based on the improved TransUnet, and evaluates the performance of the model through the test data.

[0025] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0026] In view of the problems of redundant skip connection information, insufficient utilization of local context information by convolution, and high computational complexity of Transformer in the existing U-shaped network structure, this invention proposes the ES-TransUNet medical image segmentation network. By introducing the Criss-Cross Attention (CCA) mechanism, it effectively captures long-range dependencies and realizes the comprehensive modeling of local and global features. It uses Dynamic Upsampling (Dysample) to improve the upsampling efficiency, enhance the detailed features, and uses the SCOTblock to filter redundant information in the skip connection and optimize the feature fusion. The experimental results show that ES-TransUNet achieves a good balance between performance and efficiency on the Synapse and ACDC datasets, significantly improving the segmentation effect. Description of the Drawings

[0027] Figure 1 It is a flowchart of an image segmentation method based on an improved TransUnet of the present invention.

[0028] Figure 2 It is a structure diagram of the CCA block.

[0029] Figure 3 It is a structure diagram of the improved Transformer

[0030] Figure 4 It is a structure diagram of the SCOTblock Detailed Implementation Manner

[0031] Example:

[0032] As Figure 1 shown, the technical solution of the present invention includes the following steps:

[0033] S1: Collect images and preprocess the images;

[0034] S2: Construct a medical image segmentation model based on the improved TransUnet;

[0035] S4: Train the improved TransUnet segmentation model;

[0036] S5: Test the model.

[0037] The specific steps of step S1 include the following steps:

[0038] S1.1: Select the images in the Synapse dataset as the dataset for this experiment, and divide the dataset into a training set and a test set according to a certain ratio;

[0039] S1.2: Resize the dataset images to a unified size of 224×224;

[0040] The specific steps of step S2 include the following steps:

[0041] S2.1: Add the Criss-Cross Attion mechanism before each convolutional part of the encoder;

[0042] Ordinary convolution operations apply the convolution kernel on the input image through a sliding window, which is limited to extracting local features because the convolution kernel can only perceive a small part of the image within its size range. To better capture long-range context dependencies, each position can interact with other positions in the same row; in the vertical path, self-attention is also applied to the features of each column to capture long-range dependencies in the vertical direction. Finally, by fusing the Criss-Cross features of the horizontal and vertical paths, the final output feature map is generated. This fusion not only ensures that each position can simultaneously focus on the information in the horizontal and vertical dimensions but also retains the details of the original feature map.

[0043] First, reduce the dimensionality of the output feature map of the backbone network through a convolutional layer to obtain the feature map , where ( is the number of channels, is the width, is the height), the feature goes through 3 1×1 convolutional operations to obtain , , , where , The number of channels is one-eighth of the original to reduce the computational load. In the Affinity operation, for each position in , extract the vector in, and then extract the set of vectors in the same row or column as the position , and the th position parameter is , and the Affinity operation formula is:

[0044] (1)

[0045] In the formula: is the vector extracted from , is the vector, is the th position parameter. After passing through the softmax activation, it obtains , For each position , we can obtain . Then perform the Aggregation operation to make the vector set multiply with , and finally add the original feature to obtain . The formula is as follows:

[0046] (2)

[0047] In the formula: is the activated vector, is the corresponding vector

[0048] S2.2: Replace the multi-head attention (MSA) in the Transformer block with SimAM, and replace the multi-layer perceptron (MLP) with ECA-Net;

[0049] The ES-Transformer structure first performs layer normalization on the input features to improve training stability, and then performs self-attention calculation on each position through SimAM attention. Compared with MSA, the calculation of SimAM is more simplified. The importance of neurons is evaluated by the magnitude of the energy function, further reducing the computational complexity. The energy function is as follows: S2.3: Adjust the number of prior boxes of the feature map;

[0050] (3)

[0051] In the formula: is the target neuron of the input feature, is other neurons, is the spatial dimension index, is the number of all neurons in one channel, and are binary labels.

[0052] (4)

[0053] In the formula: is the input feature, is the output feature, It is the minimum energy. The smaller the energy, the more important the neuron is. Multiply it with the original feature to enhance the feature, strengthening the important information of the output feature for subsequent information extraction. Then, after a layer normalization to maintain feature stability, it finally enters the ECA. First, without reducing the dimension, global average pooling is performed, and cross-channel interaction is carried out using 1×1 convolution, so as to maintain performance while reducing the computational cost.

[0054] S2.3: Replace the bilinear interpolation upsampling in the decoder with Dysample sampling;

[0055] The decoder is mainly composed of skip connection feature fusion, upsampling, 3×3 convolution block and segmentation head. Among them, the feature fusion part solves the problem of information redundancy in skip connections by introducing the SCoT Block. Traditionally, bilinear interpolation method is usually used for upsampling, but its accuracy is insufficient when dealing with high-frequency details or edges, which is likely to cause image blurring. In contrast, Dysample dynamically generates convolution kernels and adaptively adjusts the upsampling process according to the changes of the input feature map, so as to better retain detail and edge information and improve the quality of the upsampled features.

[0056] S2.4: Use the Simple Contextual Transformer (SCOT) module to filter redundant features in the skip connection part

[0057] TransUNet combines the Transformer and U-Net architectures, using skip connections to transfer features between the encoder and the decoder to restore resolution and details. However, simple skip connections may introduce redundant information, increase the computational burden and affect the segmentation accuracy. To solve this problem, CoTNet adopts a hybrid architecture, combining convolutional neural network (CNN) and Transformer to process local and global information simultaneously. By introducing the self-attention mechanism, CoTNet can enhance the understanding of long-range dependencies. However, the complexity of the Transformer brings a high computational burden. To optimize the CoTblock, the lightweight attention mechanism SimAM is introduced, which assigns weights to each pixel point by simulating the activation of neurons. This mechanism can highlight the important information in the input feature map and suppress redundant features. Thus, before CoTNet processing, the feature map has been preliminarily optimized and has higher quality.

[0058] In CoTNet, first use 1×1 convolution to obtain 、 、 Three different spaces. After adding position information to and 、 The product addition results in :

[0059] (5)

[0060] Wherein is the position information.

[0061] (6)

[0062] In the formula: is the output feature, where is normalized, and finally multiplied by to obtain the output result.

[0063] S3.1: Set the training parameters, adopt the stochastic gradient descent optimization algorithm for model training, set the batch size batch of training to 24, the weight decay to 1E - 4, the momentum parameter momentum to 0.9, the initial learning rate to 0.01, the number of iterations epoch to 150, the patch size to 16, the random seed to 1234, and the ViT model to use R50 - ViT - B_16. During the training process, adopt a hybrid loss function, including cross - entropy loss (Cross Entropy Loss, CE) and Dice loss. As shown in the following formula:

[0064] (7)

[0065] S3.2: Put the training set images in the dataset into the improved TransUnet segmentation model for training;

[0066] S3.3: Train the model according to the set parameters, adjust the learning rate of model training by observing the change trend of the loss function until the change of the loss function tends to be stable, and obtain the final trained model;

[0067] The step S5 uses the test set in the dataset to test the trained segmentation model based on the improved TransUnet. Its evaluation metrics include Dice coefficient, Hausdorff (Hausdorff Distance) distance, etc. Among them, the Dice coefficient is a statistical metric used to calculate the similarity between two samples. It is used to compare the overlap degree between the prediction result and the true label. The Dice coefficient calculation formula is as follows:

[0068] (8)

[0069] In the formula: and respectively represent two samples, is the set and the set The size of the intersection, that is, the size of their overlapping part

[0070] The Hausdorff Distance is a way to measure the maximum distance between two point sets. The Hausdorff distance formula is as follows:

[0071] (9)

[0072] In the formula: represents and The Hausdorff distance between the point sets The function is to find the maximum of two numbers is to take the maximum element in the set is to take the minimum element in the set represents the largest difference in distance between the two point sets

Claims

1. A medical image segmentation method based on improved TransUnet, characterized in that: The following steps are involved: S1: collect images and preprocess them; S2: Build a medical image segmentation model based on improved TransUnet; S3: training the improved TransUnet model; S4: Test the model.

2. The medical image segmentation method based on improved TransUnet according to claim 1, characterized in that: The step S1 specifically includes the following steps: S1.1: Select images from the Synapse dataset as the dataset for this experiment, and divide the dataset into a training set and a test set according to a certain ratio; S1.2: Resize the dataset images to a uniform size of 224×224.

3. The medical image segmentation method based on improved TransUnet according to claim 1, characterized in that: The step S2 specifically includes the following steps: S2.1: Add the Criss-Cross Attion mechanism before each convolution part of the encoder; S2.2: Replace the multi-head attention (MSA) in the Transformer block with SimAM, and replace the multi-layer perceptron (MLP) with ECA-Net; S2.3: Replace the bilinear interpolation upsampling in the decoder with Dysample sampling; S2.4: A Simple Contextual Transformer (SCOT) module is used in the skip connection part to filter redundant features.

4. The medical image segmentation method based on improved TransUnet according to claim 1, characterized in that: The step S4 specifically comprises the following steps: S3.1: Set the training parameters, use the stochastic gradient descent optimization algorithm for model training, set the training batch size to 24, the momentum parameter to 0.9, the initial learning rate to 0.01, and the number of epochs to 150; S3.2: Put the training set images in the dataset into the pruned improved TransUnet segmentation model for training; S3.3: Train the model according to the set parameters, and adjust the learning rate of the model training by observing the changing trend of the loss function until the loss function becomes stable to obtain the final training model.

5. The medical image segmentation method based on improved TransUnet according to claim 1, characterized in that: The step S5 uses the test set in the data set to test the trained segmentation model based on the improved TransUnet, and evaluates the performance of the model through the test data. The evaluation indicators include Dice coefficient, Hausdorff (Hausdorff Distance) distance, etc., wherein the Dice coefficient is a statistical indicator for calculating the similarity between two samples; it is used to compare the overlap between the predicted results and the true labels; and the Hausdorff distance (Hausdorff Distance) is an important concept for measuring the similarity between two groups of points, and is often used to calculate the "maximum and minimum differences" or "longest distance" of two point sets. The smaller the value, the higher the similarity.

Citation Information

Cited By

  • Steel plate surface defect detection method based on Swin-Transform network structure and electronic equipment

    CN121353763A