Method for medical image segmentation based on UNet of efficient Transform

By introducing TA-Block and LMs-Trans-Block modules in UNet, combining CNN and Transformer structures, the cross-spatial attention mechanism and multi-head self-attention mechanism are adopted, which solves the shortcomings of the existing Transformer structure in dealing with long-distance dependencies, and significantly improves the performance of medical image segmentation.

CN120147630APending Publication Date: 2025-06-13CHONGQING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510114594.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When handling long-distance dependencies, the existing Transformer structure cannot effectively balance the relationship between the context information and noise of long sequences, and the calculation and space complexity are high, and the ability to obtain more detailed information is lacking.

Method used

A UNet method based on efficient Transformer was designed. By introducing TA-Block and LMs-Trans-Block modules into the encoder and decoder, combining CNN and Transformer structures, the cross-space attention mechanism and the multi-head self-attention mechanism are adopted to optimize the loss function to improve segmentation performance.

Benefits of technology

It significantly improves the performance of medical image segmentation, reduces the impact of noise on signal extraction, enhances the network's ability to reconstruct spatial information, and achieves more accurate multi-objective medical image segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147630A_ABST
    Figure CN120147630A_ABST
Patent Text Reader

Abstract

The invention discloses a method for medical image segmentation based on UNet of an efficient Transform, and relates to the technical field of medical image processing. According to the invention, through the encoder structure and the Transform, extraction of detail information can be considered and a long-distance dependency relationship can be established, so that accurate segmentation of a multi-target medical image is achieved; an efficient cross-space attention mechanism is fused into Transform, so that the influence of noise on signal extraction is greatly reduced, and the reconstruction capability of the network on space information is enhanced; according to the module, the segmentation performance of the medical image is remarkably improved; meanwhile, TA-Block is introduced to reduce information loss of the model, a Focal loss function and Dice are introduced to be combined, so that the model is more focused on edge details to achieve a more accurate segmentation result, a large number of experiments are performed, and the result shows that compared with other methods, the method is effective, and an excellent effect is achieved on a plurality of data sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image processing, and particularly to a method for medical image segmentation using an efficient Transformer-based UNet. Background Art

[0002] Medical image segmentation is an important research direction in the fields of medical image processing and computer-aided intervention. This technology extracts key information such as the morphology, size, and structure of diseased organs or tissues through the processing and analysis of medical images, and is of great significance for clinical diagnosis, treatment decision-making, and surgical plan planning. Commonly used medical detection tools around us include: computed tomography (CT) imaging, ultrasound (Ultrasound) imaging, magnetic resonance (MRI) imaging, nuclear (Nuclear) imaging, etc. CT imaging has advantages such as multi-directional imaging and fast imaging. The characteristics of high efficiency and speed make this method the most widely used imaging method at present. However, CT images often contain problems such as noise or artifacts, blurred boundaries between tissues and organs, weak density differences between tissues and organs, complex anatomical structures and irregular shapes, and cross-overlapping between organs or tissues.

[0003] Although CNN-based methods have achieved good success in medical image segmentation, due to the limitations of the limited receptive field of convolutional kernels, they are unable to capture global and long-range semantic information and have an inherent inductive bias. Inspired by the transformative impact of the Transformer architecture in natural language processing (NLP), researchers have begun to apply this technology to computer vision tasks in order to overcome some limitations of convolutional neural networks (CNNs). The core of the Transformer architecture is the self-attention mechanism, which can process the embedding information of all positions in the input sequence in parallel, rather than sequentially. This mechanism enables the Transformer to efficiently capture long-range dependencies and adapt to input sequences of different lengths. In the field of image processing, the Vision Transformer, as a specific adaptation of this architecture, divides the input image into multiple fixed-size patches, converts each patch into a vector, and feeds it to the Transformer encoder for further processing. During the encoding stage, the self-attention mechanism establishes the relationships between patches, thereby capturing rich contextual information. Subsequently, the decoder or classifier utilizes these encoded features to complete tasks such as object detection and image segmentation. The introduction of the Vision Transformer not only injects a new perspective into image processing but also achieves results comparable to or better than those of traditional CNNs. Although the Transformer architecture performs well in processing global and long-range semantic information, due to the extensiveness of its self-attention mechanism, its computational efficiency is often affected and it lacks the capture of detailed information. To address this inefficiency issue, the Swin Transformer innovates a window self-attention mechanism that restricts attention to discrete windows in the image, thereby significantly reducing the computational complexity. However, this method somewhat limits the interaction between receptive fields. To overcome this, the CSWin Transformer proposes cross-shaped window (CSWin) self-attention, which can compute self-attention horizontally and vertically in parallel to obtain better results at a lower computational cost. In addition, the CSWin Transformer also introduces Local Enhanced Position Encoding (LePE), which imposes position information on each Transformer block. Different from previous position encoding methods, LePE directly manipulates the results of attention weights instead of adding them to the calculation of attention. LePE makes the CSWin Transformer more effective in object detection and image segmentation. To make up for the problem that the Transformer lacks the capture of detailed information, many studies have combined CNNs with transformer modules.TransUNet and LeViT-UNet integrate UNet with Transformers and achieve competitive results on abdominal multi-organ and heart segmentation datasets. In addition, some researchers have developed segmentation models using pure Transformers. SwinUNet uses Swin Transformer modules to build encoders and decoders in an architecture similar to UNet, and the performance is improved compared to TransUNet. However, this Swin Transformer-based segmentation method still has limitations in terms of receptive field interaction and is computationally expensive.

[0004] Although combining CNN with Transformer can take into account both the extraction of detailed information and the establishment of long-distance dependencies, problems also arise, such as the impact of noise and artifacts on the segmentation results. In particular, abdominal CT imaging contains no less than 20 types of tissues and organs. In addition to the segmentation target, there are often noise, multiple organs squeezing each other, and unclear boundaries. In addition to the segmentation target, there are many organs with similar morphology to the target organs. These noises will have a great impact on our segmentation results and are also the part that is most likely to produce incorrect segmentation during the segmentation process. In addition, there is the problem of artifacts from soft tissues. In existing models, such artifacts are often segmented as boundaries, which seriously affects the segmentation results.

[0005] In summary, the current research has the following problems:

[0006] (1) The existing Transformer structure cannot effectively balance the relationship between contextual information and noise in long sequences when dealing with long-distance dependencies;

[0007] (2) Due to the inequality between low-level information in the encoder and high-level information in the decoder, there is still a semantic gap between the encoder and decoder in the U-Net structure;

[0008] (3) The Transformer structure has high computational and spatial complexity and lacks the ability to obtain more detailed information.

[0009] Therefore, a new solution to the above problems needs to be proposed. Summary of the invention

[0010] The purpose of the present invention is to provide a method for medical image segmentation based on an efficient Transformer-based UNet to address the deficiencies of the prior art mentioned in the background technology.

[0011] To achieve the above object, the present invention provides the following technical solutions: A method for medical image segmentation using a UNet based on an efficient Transformer, at least including the following steps:

[0012] S1: Since the contrast between the image segmentation target and the noise in the dataset is low, the dataset is preprocessed for data augmentation and denoising. The dataset at least includes Synapse and ACDC;

[0013] S2: Design the model framework;

[0014] S3: Design the loss function for learning;

[0015] S4: Feed the processed dataset into the designed model for training, and save the weights with the best training effect;

[0016] S5: After the model training is completed, the model is tested and the effectiveness of the model is verified using public evaluation metrics.

[0017] Preferably, the preprocessing in S1 at least includes the following steps:

[0018] First, normalize the image, scale the pixel values to the range of [0,1] or [-1,1] to accelerate model convergence and improve training stability. At the same time, to ensure the consistency of the image size, the image is usually cropped or scaled to adjust to a fixed size;

[0019] Secondly, perform data augmentation. Increase data diversity by means of random rotation, translation, flipping, scaling, etc., so as to improve the generalization ability of the model. In addition, the input image and the corresponding label need to be aligned to ensure that they match under the same size;

[0020] Slice the 3D medical image, and extract individual 2D slices as the input;

[0021] In order to further improve the image quality, edge enhancement technology is sometimes applied to highlight the structural boundaries and enhance the image features. In medical images, images of different modalities are usually stacked into multi-channel inputs, and at the same time, standardization operations can also be selected, subtracting the mean and dividing by the standard deviation;

[0022] To improve the balance of the dataset, oversampling or undersampling techniques are selected according to specific circumstances to handle the class imbalance problem and ensure the diversity of the training data.

[0023] Preferably, the architecture of the model framework at least includes an encoder, a decoder, and skip connections;

[0024] Fuse CNN and Transformer together in the encoder stage to extract local and global information, and through the TA-

[0025] Block to further enrich the features, enabling the model to more effectively extract multi-faceted information during the encoding stage. The TA-Block is also known as the triple attention block, and the TA-Block is a feature extraction module;

[0026] Secondly, also introduce the TA-Block in the skip connection stage to enhance the input of low-level features to the decoder and improve the detail accuracy of image reconstruction;

[0027] Finally, build the LMs-Trans-Block module, which at least includes feature embedding, an efficient multi-scale attention module, and a multi-head self-attention module;

[0028] The efficient multi-scale attention is EMA, and the multi-head self-attention is MCA;

[0029] The LMs-Trans-Block module optimizes the Transformer, abandons the residual connection and the feed-forward neural network, and at the same time introduces EMA to increase the ability to reconstruct the spatial structure, enabling it to more efficiently establish long-range connection dependencies in this model.

[0030] Preferably, the application process of the LMs-Trans-Block module at least includes the following steps:

[0031] Assume the input data is F ∈ R S×C , where C is the number of channels, S is the image size S = H × W, and H and W are the height and width of the image respectively;

[0032] First, perform position information embedding, expressed as follows:

[0033] E i = F i + Pos i (1)

[0034] where i ∈ {1, 2, …, S}, Pos i is the position encoding dependent on F i , and E represents the input data of EMA, E ∈ R S×C ;

[0035] To enhance the ability of the Transformer module to obtain cross-space information of the feature map, introduce EMA to enhance the model's acquisition of spatial information from a multi-scale perspective;

[0036] To improve the computational efficiency of the module, feature grouping is first performed at the front end, and multi-scale features are obtained through a parallel processing strategy to achieve fast response. Moreover, to capture the dependencies between all channels and reduce the computational budget, cross-channel information interaction is modeled in the channel direction;

[0037] That is, EMA implements the use of three parallel paths to extract the attention weight descriptors of the grouped feature maps;

[0038] The first parallel branch stacks a single 3x3 kernel for capturing multi-scale feature representations. The representation of the first branch is as follows:

[0039]

[0040] O 1i = Reshape(K 3×3 (G i )) (3)

[0041]

[0042] where g n (·) represents dividing the features into n groups according to channels; C represents the number of channels of the input data; G represents the grouped feature set; G i represents the divided feature map, Reshape(·) represents reshaping the features; is the feature after convolution processing; K 3×3 represents performing a 3×3 convolution operation on the features while keeping the data shape unchanged; j represents the index of the channel; through a cross-spatial information aggregation method δ(·) in different spatial dimension directions, richer feature aggregation is achieved, and cross-channel information is modeled through global average pooling operation. Assuming the original input feature is X ∈ R C×H×W , it is represented as follows:

[0043]

[0044] where (m,n) represents any point in the C-channel feature map, and formula (5) indicates that global pooling means taking the average along the X and Y directions;

[0045] The other two branches respectively encode the channels along two spatial directions through global average pooling operations;

[0046] Calculate the weights of channel attention through matrix multiplication and weight the grouped features G i as follows:

[0047]

[0048] Ti = Norm(sig(exp(X wi ·X hi ·G i ))) (7)

[0049] T 1i = Reshape(T i ) (8)

[0050]

[0051] where δ_X(·) and δ_Y(·) are the average pooling operations along the X and Y directions respectively; cat(·) represents the concatenation operation, which concatenates the features pooled along the X and Y directions into a vector; split(·) represents splitting the features after 1×1 convolution and pooling concatenation into different parts; exp(·) represents the exponential operation, which helps to enhance the large values in the weighted features; sig(·) represents the Sigmoid activation function; Norm(·) represents the normalization operation; Reshape(T i ) represents reshaping the calculated channel attention weights; as the output result

[0052] Integrate the feature information of the three branches, calculate the cross - spatial attention weights through matrix multiplication and feature fusion, and the output of the smallest branch will be directly transformed into the corresponding dimension shape before the joint activation mechanism of the channel features. Finally, calculate the weighted average of the original input feature vectors, as shown below:

[0053] W i = reshape((O 2i ·T 1i )+(T 2i ·O 1i )) (10)

[0054] M i = sig(exp(W i ·G i ))#(11)

[0055] When obtaining features with multiple elements, then continue to use the multi - head self - attention in Transformer to improve the parallelism and computational efficiency of the model, as shown below:

[0056] L i = LN(M i ) (12)

[0057]

[0058] LMT o= Linear(Concat(Head 1 , Head 2 , …, Head h )) (14)

[0059] where LN(·) represents layer normalization, and M i represents the feature input to layer normalization; h represents the number of multi - heads, and h is set to 12; Q, K, V represent query, key, and value respectively, and Q i = L i W Qi , K i = L i W Ki , V i = L i W Vi ; represents calculating the similarity between query and key; d k represents the dimension of the subspace; Concat(Head 1 , Head 2 , …, Head h ) represents concatenating the outputs of multiple heads. Here, Head 1 , Head 2 , …, Head h are the outputs obtained from multiple - head self - attention; Linear(·) represents a linear transformation layer; LMT o represents the final output of multi - head self - attention.

[0060] Preferably, the TA - Block integrates the position, channel, and spatial features unique to images, enabling efficient feature extraction for the unique attributes of images. Especially in the U - Net architecture, the dedicated feature extraction function of the TA - Block is crucial.

[0061] The TA - Block includes at least a position attention module, a channel attention module, and a spatial attention module. The position attention module is the RPA, the channel attention module is the RCA, and the spatial attention module is the RFA.

[0062] Preferably, the RPA adopts a position attention mechanism, which is the PAM. The PAM is used to capture the position mapping relationship between any two points in the feature map, updating the feature parameters through the weighted sum of all position features, and determining the weight by the feature similarity between two positions. Therefore, the PAM can effectively extract the position mapping relationship between any two points in the feature map;

[0063] First, use the initial feature, X ∈ RC×H×W Denoted as, C represents the channel, H represents the height, and W represents the width;

[0064] Then, X is fed into convolutional layers with different weights, generating three new feature maps, namely X Q , X K and X V , and the size of each feature map is R C×H×W ;

[0065] Next, X K and X V are reshaped into R C×N , where N = H × W represents the number of pixels. A batch matrix multiplication is performed between X K and the transpose of X Q to calculate the energy scores between any two pixel positions of X K and X Q . Subsequently, a softmax layer is used to calculate the spatial attention map E ∈ R N×N ;

[0066] Then, the matrix X V is reshaped back into R N×N . A batch matrix multiplication is performed between X V and the transpose of E, and then the result is reshaped into R C×H×W ;

[0067] Finally, R C×H×W is multiplied by the parameter α, and the feature X is also multiplied by β element - by - element for residual connection to obtain the final output P ∈ R C×H×W .

[0068] The feature P is the weighted sum of the features across all positions and the original features. The feature P has global context features and aggregates context weighted based on pixel - level position information, which ensures the effective extraction of position features while maintaining global context information.

[0069] Preferably, the RCA adopts a channel attention mechanism, and the channel attention mechanism is the CAM. What is different between the CAM and the PAM is that the CAM directly reshapes the original feature X ∈ R C×H×W into R C×N , obtaining the matrix X a . Then, a batch matrix multiplication is performed between X b and its transpose. To suppress noise, the difference is calculated and adjusted by comparing with the maximum value of each row or column;

[0070] Subsequently, a softmax layer is applied to obtain the channel attention map S ∈ R C×H×W . The result is multiplied by the ω scale parameter and element - wise summed with X to obtain the final output I ∈ RC×H×W , finally, the final feature of each channel is the weighted sum of all channels and the original features, thus endowing the CAM with strong channel feature extraction ability.

[0071] Preferably, the RFA adopts a spatial attention mechanism, which is the SAM. The SAM extracts the importance weights of different features through the receptive field slider, prioritizes the extracted receptive field spatial features, and reconstructs the feature weights of the spatial features of the feature map in a more efficient way, improving the model's ability to extract regional spatial features;

[0072] First, the original feature X ∈ R C×H×W , after global average pooling and unfolding methods, its dimension becomes L rf ∈ R C×1×1 , and at the same time, a method for quickly extracting receptive field spatial features is adopted, that is, Group Conv, using a 3×3 convolution kernel to extract features, where each 3×3 size window represents a receptive field slider;

[0073] In this way, the original feature is mapped to a new feature, while reducing the computational amount of convolution operations, and improving the network's expression ability by increasing the number of output channels;

[0074] Specifically, the feature X ∈ R C×H×W is first grouped according to the number of input channels, then each group of features is input into different convolution kernel groups for convolution operations, and then a softmax operation is performed, and the output data is

[0075] The RFA solves the problem of insensitivity to spatial information differences caused by position changes by emphasizing the importance of different features in the receptive field slider and prioritizing the receptive field spatial features.

[0076] Preferably, the loss function in S4 includes at least the Dice loss function and the Focal loss function;

[0077] The mathematical expression of the Dice loss function is as follows:

[0078]

[0079] where N refers to the number of input pixels; t i,j and p i,j respectively represent the label value and the predicted value of the pixel at the position (i, j); T refers to the total number of segments included in the background; ∈ refers to the smoothing term, which aims to avoid the denominator or numerator being zero.

[0080] To make up for the deficiencies of the Dice loss function, the Focal loss function is introduced to guide the training of SSTrans-Net. The expressions of the cross-entropy loss function and the Focal loss function are as follows:

[0081]

[0082]

[0083] where p t is the classification probability of the current pixel; γ is a hyperparameter defined to control the intensity of punishing misclassified pixels, usually set to 2;

[0084] The Focal loss function assigns a lower loss weight to easily classified pixels and imposes a greater penalty on difficult-to-classify pixels. This mechanism can not only effectively address the problem of imbalance in the number of positive and negative samples but also alleviate the challenges brought by the imbalance in class distribution;

[0085] Different from the simple combination of cross-entropy (CE) and Dice loss functions used in most works, an improved method that combines Focal and Dice losses is used. The improved method that combines Focal and Dice losses not only focuses on the segmentation accuracy of the entire image but also emphasizes difficult-to-classify pixels more. The difficult-to-classify pixels at least include boundary regions and multi-scale pixels. The expression of the loss function for the overall loss function is as follows:

[0086] Loss = μDice Loss + φFocal Loss (18)

[0087] where μ and φ are weighting parameters, and the best experimental results are 0.6 and 0.4. In this way, high-quality long-range dependent features in different channels can be filtered, thus significantly improving the segmentation performance.

[0088] Preferably, the publicly disclosed evaluation metrics used in S5 are DSC and HD95.

[0089] Compared with the prior art, the beneficial effects of the present invention are:

[0090] The present invention can not only extract detailed information but also establish long-range dependencies through an encoder structure and Transformer, achieving precise segmentation of multi-object medical images. By integrating an efficient cross-space attention mechanism into Transformer, the influence of noise on signal extraction is greatly reduced, and the network's ability to reconstruct spatial information is enhanced. This module significantly improves the segmentation performance of medical images. At the same time, TA-Block is introduced to reduce information loss in the model. Additionally, the combination of the Focal loss function and Dice makes the model more focused on edge details to achieve more precise segmentation results. Through a large number of experiments, the results show that the proposed method is effective and achieves excellent results on multiple datasets compared with other methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0091] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0092] Figure 1 Schematic diagram of the overall model framework of the present invention;

[0093] Figure 2 Schematic diagram of the module LMs-Trans that effectively establishes long-range dependencies by integrating cross-space information of the present invention;

[0094] Figure 3 Schematic diagram of the module TA-Block for reducing information loss of the present invention;

[0095] Figure 4 Schematic diagram of CT medical image noise of the present invention;

[0096] Figure 5 Visual comparison diagram of the multi-organ segmentation results of the present invention and other models;

[0097] Figure 6 Visual comparison schematic diagram of the average DSC, average HD, and error bars (95% confidence interval) of DSC for each organ on the Synapse dataset of the present invention;

[0098] Figure 7 Visual comparison schematic diagram of the average DSC and error lines (95% confidence interval) of DSC for each heart structure on the ACDC database of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0099] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0100] In the field of abdominal multi-organ CT image segmentation, due to problems such as image blurring, low color contrast, and high noise in CT imaging, the organs present a non-rigid state, and there are also challenges such as adjacent organs blocking each other, squeezing each other, and unclear boundaries, making this topic highly valuable for research.

[0101] Current deep learning models cannot correctly identify and segment the target and noise artifacts, and there are also problems such as a large semantic difference between the encoder and decoder and serious information loss in the commonly used encoder-decoder segmentation structure, resulting in a low segmentation accuracy of the model.

[0102] We designed a method based on an efficient Transformer for UNet in medical image segmentation, which is a new U-shaped segmentation method, namely LMsT-TransUNet. This method redesigned the Transformer module and introduced it into UNet to reduce the noise interference of complex backgrounds and enhance the acquisition of cross-space information. In addition, a TA-Block was designed to reduce the semantic difference and information loss between the encoder and decoder. This method significantly improves the computational efficiency and the ability to reconstruct the dependence of cross-space information. In addition, our innovative Transformer structure uses the cross-space domain to establish context dependencies to enhance the multi-faceted acquisition of information to achieve an accurate edge segmentation effect.

[0103] A method based on an efficient Transformer for UNet in medical image segmentation includes at least the following steps:

[0104] S1: Since the contrast between the image segmentation target and noise in the dataset is low, the dataset is preprocessed for data augmentation and denoising. The dataset includes at least Synapse and ACDC;

[0105] S2: Design the model framework. In the encoder part, combine the CNN and Transformer structures to make full use of their advantages. CNN is good at capturing local detailed information, while Transformer can handle global dependencies. LMsT-TransUNet not only integrates CNN and Transformer in the encoder stage, but also introduces the TA-Block module. This module further enriches the feature extraction process, enabling the encoder to extract multi-faceted information from images more efficiently. The TA-Block helps the model capture fine-grained local information by mining specific features from the spatial and channel dimensions, while enhancing the ability to model global information. The skip connection part is optimized based on the traditional U-Net architecture and further introduces the TA-Block module. The skip connection combines the low-level features extracted in the encoder stage with the high-level features in the decoder, enabling the model to better retain detailed information during the image reconstruction process. The introduction of the TA-Block here makes the input of the low-level features more accurate, improving the detail accuracy of image reconstruction and the ability to reconstruct spatial information. The decoder part is one of the core parts of LMsT-TransUNet. To reduce the shallow noise and computational burden in Transformer, we optimize its structure, remove redundant modules, and introduce the EMA module (Exponential Moving Average). This module effectively enhances the model's ability to reconstruct the spatial structure by weighted averaging features. The EMA module helps optimize the structure of the residual connection and the feed-forward neural network, enabling the model to establish long-range dependencies more efficiently and improve the segmentation performance and robustness in complex backgrounds;

[0106] S3: Design the loss function for learning, using an improved method that combines Focal and Dice losses. This method not only focuses on the segmentation accuracy of the entire image, but also emphasizes difficult-to-classify pixels, such as boundary regions and multi-scale pixels. And introduce the weighted parameters μ and φ. In this way, LMsT-TransUNet can filter high-quality long-range dependent features in different channels, thus significantly improving the segmentation performance;

[0107] S4: Feed the processed dataset into the designed model for training, save the weights with the best training effect, and train the model using the SGD optimizer. The momentum is set to 0.9, the weight_decay is set to 0.0001, and the learning rate is set to 0.01. The number of iterations is 150 epochs, and the batch size is 24. And record the weights of the best round;

[0108] S5: After the model training is completed, test the model and use public evaluation metrics to verify the effectiveness of this model.

[0109] The preprocessing in S1 includes at least the following steps:

[0110] First, normalize the image, scale the pixel values to the range of [0, 1] or [-1, 1] to accelerate model convergence and improve training stability. At the same time, to ensure consistent image size, usually perform cropping or scaling operations on the image to adjust it to a fixed size;

[0111] Secondly, perform data augmentation to increase data diversity by means of random rotation, translation, flipping, scaling, etc., thereby improving the generalization ability of the model. In addition, the input image and the corresponding label (mask) need to be aligned to ensure they match at the same size;

[0112] Slice the 3D medical image and use the extracted single 2D slice as the input;

[0113] To further improve the image quality, edge enhancement technology is sometimes applied to highlight the structural boundaries and enhance the image features. In medical images, images of different modalities (such as CT, MRI) are usually stacked into multi-channel inputs, and at the same time, standardization operations can also be selected, subtracting the mean and dividing by the standard deviation;

[0114] To improve the balance of the dataset, oversampling or undersampling techniques are selected according to specific situations to handle the class imbalance problem and ensure the diversity of training data;

[0115] Through these preprocessing steps, the model can effectively process and analyze medical images and perform high-precision image segmentation tasks to ensure that the input data can better meet the training needs of the model.

[0116] Refer to Figure 1 The architecture of the model framework includes at least an encoder, a decoder, and skip connections;

[0117] In the encoder stage, fuse CNN and Transformer to extract local and global information, and further enrich the features through the TA-

[0118] Block, enabling the model to more effectively extract multi-faceted information during the encoding stage. The TA-Block is also called the triple attention block, and the TA-Block is a feature extraction module;

[0119] Secondly, in the skip connection stage, also introduce the TA-Block to enhance the input of low-level features to the decoder and improve the detail accuracy of image reconstruction;

[0120] Finally, build the LMs-Trans-Block module, and the LMs-Trans-Block module includes at least a feature embedding, an efficient multi-scale attention module, and a multi-head self-attention module;

[0121] Efficient multi-scale attention is EMA, and multi-head self-attention is MCA;

[0122] Figure 2 The illustration of the LMs-Trans-Block is shown, where (a) is the structure of the original Vision Transformer, and (b) is the modified Transformer structure. In the Transformer architecture, the residual connection refers to directly passing the input signal to the subsequent layers of the network. Although this can avoid over-compression or loss of features, it also leads to an increase in the impact of noise in the shallow features on the segmentation result. The feed-forward neural network consists of two fully connected layers and an activation function (such as ReLU), which maps and transforms features, performs non-linear transformation on information, and enhances the representation ability of the module. However, the increase in the number of calculations and parameters in this structure will also reduce the module's ability to reconstruct the spatial structure. Therefore, during the modification of the Transformer, we boldly abandoned the residual connection and the feed-forward neural network. The LMs-Trans-Block module optimizes the Transformer by abandoning the residual connection and the feed-forward neural network, and at the same time introduces EMA to increase the ability to reconstruct the spatial structure, enabling it to more efficiently establish long-range connection dependencies in this model.

[0123] The application process of the LMs-Trans-Block module at least includes the following steps:

[0124] Assume the input data is F ∈ R S×C , where C is the number of channels, S is the size of the picture S = H × W, and H and W are the height and width of the picture respectively;

[0125] First, perform position information embedding, which is expressed as follows:

[0126] E i = F i + Pos i (1)

[0127] where i ∈ {1, 2, …, S}, Pos i is the position encoding dependent on F i , and E represents the input data of EMA, E ∈ R S×C ;

[0128] To enhance the ability of the Transformer module to obtain cross-space information of the feature map, EMA is introduced to enhance the model's acquisition of spatial information from a multi-scale perspective;

[0129] To improve the computational efficiency of the module, feature grouping is performed at the front end, and multi-scale features are obtained through a parallel processing strategy to achieve fast response. Moreover, to capture the dependencies between all channels and reduce the computational budget, cross-channel information interaction is modeled in the channel direction;

[0130] That is, EMA implements the use of three parallel paths to extract the attention weight descriptors of the grouped feature maps;

[0131] The first parallel branch stacks a single 3x3 kernel for capturing multi-scale feature representations. The representation of the first branch is as follows:

[0132]

[0133] O 1i = Reshape(K 3×3 (G i )) (3)

[0134]

[0135] where g n (·) represents dividing the features into n groups according to channels; C represents the number of channels of the input data; G represents the set of grouped features; G i represents the divided feature map, Reshape(·) represents reshaping the features; is the feature after convolution processing; K 3×3 represents performing a 3×3 convolution operation on the features while keeping the data shape unchanged; j represents the channel index; through a cross-spatial information aggregation method δ(·) in different spatial dimension directions, richer feature aggregation is achieved, and cross-channel information is modeled through global average pooling operation. Assuming the original input feature is X ∈ R C×H×W , it is represented as follows:

[0136]

[0137] where (m,n) represents any point in the C-channel feature map, and formula (5) indicates that global pooling means taking the average along the X and Y directions;

[0138] The other two branches respectively encode the channels along two spatial directions through global average pooling operations;

[0139] Calculate the weights of channel attention through matrix multiplication and weight the grouped features G i as follows:

[0140]

[0141] Ti = Norm(sig(exp(X wi ·X hi ·G i ))) (7)

[0142] T 1i = Reshape(T i ) (8)

[0143]

[0144] where δ_X(·) and δ_Y(·) are the average pooling operations along the X and Y directions respectively; cat(·) represents the concatenation operation, which concatenates the features pooled along the X and Y directions into a vector; split(·) represents splitting the features after 1×1 convolution and pooling concatenation into different parts; exp(·) represents the exponential operation, which helps to enhance the large values in the weighted features; sig(·) represents the Sigmoid activation function; Norm(·) represents the normalization operation; Reshape(T i ) represents reshaping the calculated channel attention weights; as the output result

[0145] Integrating the feature information of the three branches, calculate the cross - spatial attention weights through matrix multiplication and feature fusion, and the output of the smallest branch will be directly transformed into the corresponding dimension shape before the joint activation mechanism of the channel features. Finally, calculate the weighted average of the original input feature vector, which is expressed as follows:

[0146] W i = reshape((O 2i ·T 1i )+(T 2i ·O 1i )) (10)

[0147] M i = sig(exp(W i ·G i ))#(11)

[0148] When obtaining the multi - feature, next continue to use the multi - head self - attention in Transformer to improve the parallelism and computational efficiency of the model, which is expressed as follows:

[0149] L i = LN(M i ) (12)

[0150]

[0151] LMT o= Linear(Concat(Head 1 , Head 2 , …, Head h )) (14)

[0152] where LN(·) represents layer normalization, and M i represents the feature input to layer normalization; h represents the number of heads, and h is set to 12; Q, K, and V represent query, key, and value respectively, and Q i = L i W Qi , K i = L i W Ki , V i = L i W Vi ; represents calculating the similarity between query and key; d k represents the dimension of the subspace; Concat(Head 1 , Head 2 , …, Head h ) represents concatenating the outputs of multiple heads. Here, Head 1 , Head 2 , …, Head h are the outputs obtained from multiple head self-attention; Linear(·) represents a linear transformation layer; LMT o represents the final output of the multi-head self-attention.

[0153] As Figure 3 shown, the TA-Block integrates the position, channel, and spatial features unique to images, enabling efficient feature extraction for the unique properties of images. Especially in the U-Net architecture, the dedicated feature extraction function of the TA-Block is crucial. Although the Transformer is good at extracting global features with its self-attention mechanism, its design is not specifically optimized for the specific properties of images. In contrast, the TA-Block performs well in extracting position and channel features and can capture a more detailed and accurate set of features. Therefore, we integrate the TA-Block into the encoder and skip connections to enhance the performance of the model in image segmentation tasks.

[0154] The TA-Block includes at least a position attention module, a channel attention module, and a spatial attention module. The position attention module is the RPA, the channel attention module is the RCA, and the spatial attention module is the RFA.

[0155] RPA adopts the position attention mechanism, which is the PAM. The PAM is used to capture the position mapping relationship between any two points in the feature map, update the feature parameters through the weighted sum of all position features, and determine the weights by the feature similarity between two positions. Therefore, the PAM can effectively extract the position mapping relationship between any two points in the feature map;

[0156] First, the initial feature, X ∈ R C×H×W is represented as, where C represents the channel, H represents the height, and W represents the width;

[0157] Then, X is sent into convolutional layers with different weights, generating three new feature maps, namely X Q 、X K and X V , and the size of each feature map is R C×H×W ;

[0158] Next, X K and X V are reshaped into R C×N , where N = H × W represents the number of pixels. A batch matrix multiplication is performed between X K and the transpose of X Q to calculate the energy scores between any two pixel positions of X K and X Q . Subsequently, a softmax layer is used to calculate the spatial attention map E ∈ R N×N ;

[0159] Then, the matrix X V is reshaped back into R N×N . A batch matrix multiplication is performed between X V and the transpose of E, and then the result is reshaped into R C×H×W ;

[0160] Finally, R C×H×W is multiplied by the parameter α, and the feature X is also multiplied by β element-wise for residual connection to obtain the final output P ∈ R C×H×W .

[0161] The feature P is the weighted sum of the features across all positions and the original feature. The feature P has global context features and aggregates context based on pixel-level position information, which ensures the effective extraction of position features while maintaining global context information.

[0162] RCA adopts the channel attention mechanism, which is the CAM. The difference between the CAM and the PAM is that the CAM directly reshapes the original feature X ∈ R C×H×W into R C×N , obtaining the matrix X a , and then between X bPerform batch matrix multiplication between it and its transpose. To suppress noise, calculate the difference and make adjustments by comparing with the maximum value of each row or column.

[0163] Subsequently, apply the softmax layer to obtain the channel attention map S ∈ R C×H×W , multiply the result by the ω scale parameter, and perform an element-wise summation operation with X to obtain the final output I ∈ R C×H×W . Finally, the final feature of each channel is the weighted sum of all channels and the original features, thus endowing CAM with strong channel feature extraction capabilities.

[0164] RFA adopts a spatial attention mechanism, namely the SAM. The SAM extracts the importance weights of different features through the receptive field slider, prioritizes the spatial features of the extracted receptive field, and reconstructs the feature weights of the spatial features of the feature map in a more efficient way to improve the model's ability to extract regional spatial features.

[0165] First, input the original feature X ∈ R C×H×W . After global average pooling and unfolding methods, its dimension becomes L rf ∈ R C×1×1 . At the same time, adopt a method for quickly extracting the spatial features of the receptive field, namely Group Conv, and use a 3×3 convolution kernel to extract features, where each 3×3 size window represents a receptive field slider.

[0166] In this way, map the original features to new features, while reducing the computational amount of convolution operations and enhancing the network's expressive ability by increasing the number of output channels.

[0167] Specifically, first group the feature X ∈ R C×H×W by the number of input channels, then input each group of features into different groups of convolution kernels for convolution operations, and then perform the softmax operation. The output data is

[0168] RFA solves the problem of insensitivity to spatial information differences caused by position changes by emphasizing the importance of different features in the receptive field slider and prioritizing the spatial features of the receptive field.

[0169] The loss function in S4 includes at least the Dice loss function and the Focal loss function.

[0170] The Dice loss function has certain limitations in the segmentation task. It ignores the imbalance between foreground and background pixels and simply calculates the overlap ratio between the prediction result and the ground truth label over the entire image. Therefore, the Dice loss function is difficult to capture the segmentation boundary information of pixels and the changes in the target shape, which leads to its insufficient performance in multi-object segmentation and multi-scale segmentation tasks. The mathematical expression of the Dice loss function is as follows:

[0171]

[0172] where N refers to the number of input pixels; t i,j and p i,j represent the label value and the predicted value of the pixel at the position (i, j) respectively; T refers to the total number of segments included in the background; ∈ is the smoothing term, which aims to avoid the denominator or numerator being zero.

[0173] To make up for the deficiencies of the Dice loss function, the Focal loss function is introduced to guide the training of SSTrans-Net. The Focal loss function was first applied in the field of object detection. It adjusts the loss weight by adding an exponential weighting term on the basis of the cross-entropy loss (CE). This weighting term is jointly determined by the classification probability of the pixel and the hyperparameter γ. The expressions of the cross-entropy loss function and the Focal loss function are as follows:

[0174]

[0175]

[0176] where p t is the classification probability of the current pixel; γ is the hyperparameter defined to control the intensity of punishing misclassified pixels, usually set to 2;

[0177] The Focal loss function assigns a lower loss weight to easily classified pixels and imposes a greater penalty on difficult-to-classify pixels. This mechanism can not only effectively address the problem of the imbalance in the number of positive and negative samples but also alleviate the challenges brought by the imbalance in class distribution;

[0178] Different from the simple combination of cross-entropy (CE) and Dice loss functions used in most works, an improved method that combines Focal and Dice losses is used. The improved method that combines Focal and Dice losses not only focuses on the segmentation accuracy of the entire image but also emphasizes more on difficult-to-classify pixels. Difficult-to-classify pixels at least include boundary regions and multi-scale pixels. The expression of the overall loss function is as follows:

[0179] Loss = μDice Loss + φFocalLoss (18)

[0180] Among them, μ and φ are weighted parameters, and the best results of the experiment are 0.6 and 0.4. In this way, high-quality long-range dependence features in different channels can be filtered, thus significantly improving the segmentation performance.

[0181] The publicly used evaluation metrics in S5 are DSC and HD95.

[0182] The method of the present invention will be further demonstrated in combination with data below, and reference can be made to Figures 4 - 7 ;

[0183] As shown in Table 1, the method we proposed on the Synapse dataset improved the average DSC and HD. At the same time, we show the average DSC, average HD, and error bars of DSC (95% confidence interval) for each organ in Figure 5 . Compared with TransUNet

[13] and Swin-UNet

[30] , our average DSC increased by 4.22% and 3.99% respectively, and the average HD increased by 12.83% and 2.69% respectively. It is worth noting that in the segmentation of the spleen and stomach, the DSC of LMsT-TrnasUNet is significantly higher than other segmentation methods. Different from other organs, the spleen and stomach have large individual variations and are prone to occluded parts, and our method achieved more accurate segmentation results, indicating that our LMsT-TrnasUNet provides higher segmentation accuracy in complex segmentation environments.

[0184] Table 1 presents a detailed comparison with the recent DSC and HD medical image segmentation methods on the Synapse dataset.

[0185]

[0186]

[0187] As shown in Table 2, our model (LMsT-TransUNet) outperforms other models in terms of overall segmentation accuracy and segmentation performance in each key region. In particular, it achieved the highest scores in the segmentation of the left ventricle and myocardium, demonstrating its strong performance advantages.

[0188] Table 2 presents a detailed comparison with the recent DSC and HD medical image segmentation methods on the Synapse dataset.

[0189] Model Backbone DSC↑ RV MYO LV our CNN+Trans 89.96 89.46 84.79 95.63 UNet CNN 87.55 87.1 80.63 94.92 Att-UNet CNN 86.75 87.58 79.2 93.47 Swin-UNet CNN+Trans 86.46 91.95 81.79 85.65 TransUNet CNN+Trans 88.55 93.41 84.77 87.47 UNet++ CNN 88.86 86.6 84.57 95.42

[0190] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in any respect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Accordingly, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

Claims

1. A method for medical image segmentation based on efficient Transformer UNet, characterized by: At least the following steps are included: S1: Since the image segmentation target in the data set has a low contrast with the noise, the data set is preprocessed for data enhancement and denoising, and the data set includes at least Synapse and ACDC; S2: Design model framework; S3: Design loss function for learning; S4: Send the processed data set to the designed model for training, and save the weight with the best training effect; S5: After the model training is completed, the model is tested and the effectiveness of the model is verified using public evaluation indicators.

2. The method for medical image segmentation based on efficient Transformer UNet according to claim 1, characterized in that: The preprocessing in S1 at least includes the following steps: First, the image is normalized and the pixel values ​​are scaled to the range of [0, 1] or [-1, 1] to accelerate model convergence and improve training stability. At the same time, in order to ensure the consistency of image size, the image is usually cropped or scaled to a fixed size. Secondly, data enhancement is performed to increase data diversity through random rotation, translation, flipping, scaling, etc., thereby improving the generalization ability of the model. In addition, the input image and the corresponding label need to be aligned to ensure that they match at the same size; Slice processing of 3D medical images by extracting a single 2D slice as input; To further improve image quality, edge enhancement techniques are sometimes used to highlight structural boundaries and enhance image features. In medical images, images of different modalities are usually stacked into multi-channel inputs, and can also be standardized by subtracting the mean and dividing by the standard deviation. In order to improve the balance of the data set, oversampling or undersampling technology is used to deal with the category imbalance problem according to the specific situation to ensure the diversity of training data.

3. The method for medical image segmentation based on efficient Transformer UNet according to claim 1, characterized in that: The architecture of the model framework includes at least an encoder, a decoder and a skip connection; In the encoder stage, CNN and Transformer are fused together to extract local and global information, and the features are further enriched through TA-Block, so that the model can more effectively extract various information in the encoding stage. The TA-Block is also called triple attention block, and the TA-Block is a feature extraction module; Secondly, TA-Block is also introduced in the skip connection stage to enhance the low-level feature input to the decoder and improve the detail accuracy of image reconstruction; Finally, build a LMs-Trans-Block module, which includes at least feature embedding, an efficient multi-scale attention module, and a multi-head self-attention module; The efficient multi-scale attention is EMA, and the multi-head self-attention is MCA; The LMs-Trans-Block module optimizes the Transformer by abandoning residual connections and feedforward neural networks, and introduces EMA to increase the ability to reconstruct spatial structures, so that it can more efficiently establish long connection dependencies in the model.

4. The method for medical image segmentation based on efficient Transformer UNet according to claim 3, characterized in that: The application process of the LMs-Trans-Block module includes at least the following steps: Assume that the input data is F∈R S×C , where C is the number of channels, S is the image size S = H × W, H, W are the height and width of the image respectively; First, the position information is embedded, as shown below: AND i =F i +Pos i (1) where i∈{1,,2,…,S}, Pos i It depends on F i Position encoding, E represents the input data of EMA E∈R S×C ; In order to enhance the ability of the Transformer module to obtain cross-spatial information of feature maps, EMA is introduced to enhance the model's acquisition of spatial information from a multi-scale perspective; In order to improve the computational efficiency of the module, feature grouping is first performed at the front end, and multi-scale features are obtained through a parallel processing strategy to achieve fast response. In addition, in order to capture the dependencies between all channels and reduce the computational budget, cross-channel information interaction is modeled in the channel direction. That is, EMA implements the attention weight descriptor that uses three parallel paths to extract group feature maps; The first parallel branch stacks a single 3x3 kernel to capture multi-scale feature representations. The representation of the first branch is as follows: O 1i = Reshape(K 3×3 (G i )) (3) where g n (·) indicates that the features are divided into n groups according to the channels; C indicates the number of channels of the input data; G indicates the feature set after grouping; G i represents the feature map after division, Reshape(·) means reshaping the feature; is the feature after convolution processing; K 3×3 Indicates a 3×3 convolution operation on the feature, while keeping the data shape unchanged; j represents the index of the channel; a cross-spatial information aggregation method δ(·) in different spatial dimensions is used to achieve richer feature aggregation, and the cross-channel information is modeled through a global average pooling operation. Assume that the original input feature is X∈R C×H×W , which is expressed as follows: Where (m,n) represents any point in the C channel feature map. Formula (5) indicates that global pooling refers to averaging along the X and Y directions; The other two branches encode the channels along two spatial directions through global average pooling operations respectively; The weight of the channel attention is calculated by matrix multiplication, and the grouped features G i The weighting is expressed as follows: T i = Norm(sig(exp(X wi ·X hi ·G i ))) (7) T 1i =Reshape(T i ) (8) Among them, δ_X(·) and δ_Y(·) refer to the average pooling operations along the X direction and the Y direction respectively; cat(·) represents the concatenation operation, which concatenates the features after pooling along the X direction and the Y direction into a vector; split(·) means dividing the features after 1×1 convolution and pooling concatenation into different parts; exp(·) represents the exponential operation, which helps to enhance the large values ​​in the weighted features; sig(·) represents the Sigmoid activation function; Norm(·) represents the normalization operation; Reshape(T i ) indicates that the calculated channel attention weights are reshaped; as the output result The feature information of the three branches is integrated, and the cross-space attention weights are calculated through matrix multiplication and feature fusion. The output of the smallest branch will be converted to the corresponding dimensional shape directly before the joint activation mechanism of the channel features. Finally, the weighted average of the original input feature vector is calculated, which is expressed as follows: W i =reshape((O 2i ·T 1i )+(T 2i ·O 1i )) (10) M i =sig(exp(W i ·G i ))#(11) When we get multivariate features, we continue to use the multi-head self-attention in Transformer to improve the parallelism and computational efficiency of the model, as shown below: L i =LN(M i ) (12) LMT o =Linear(Concat(Head1,Head2,…,Head h )) (14) Where LN(·) represents layer normalization, M i represents the normalized features input to the layer; h represents the number of heads, and h is set to 12; Q, K, and V represent query, key, and value, respectively, and Q i =L i W Qi , K i =L i W Ki , V i =L i W Vi ; Indicates calculating the similarity between query and key; d k Represents the dimension of the subspace; Concat(Head1,Head2,…,Head h ) means concatenating the outputs of multiple heads together, where Head1, Head2, …, Head h is the output obtained from multiple head self-attentions; Linear(·) represents a linear transformation layer; LMT o Represents the final output of multi-head self-attention.

5. The method for medical image segmentation based on efficient Transformer UNet according to claim 4, characterized in that: The TA-Block integrates the unique position, channel and spatial features of the image, so as to efficiently extract features based on the unique attributes of the image. In particular, the dedicated feature extraction function of the TA-Block is crucial in the U-Net architecture. The TA-Block includes at least a position attention module, a channel attention module and a space attention module, the position attention module is RPA, the channel attention module is RCA, and the space attention module is RFA.

6. The method for medical image segmentation based on efficient Transformer UNet according to claim 5, characterized in that: The RPA adopts a position attention mechanism, namely PAM, which is used to capture the position mapping relationship between any two points in the feature map. The feature parameters are updated by weighting all position features, and the weight is determined by the feature similarity between the two positions. Therefore, PAM can effectively extract the position mapping relationship between any two points in the feature map. First, the initial features are used, X∈R C×H×W It is expressed as, C represents channel, H represents height, and W represents width; Then, X is fed into convolutional layers with different weights to generate three new feature maps, namely X Q , X K and X V , the size of each feature map is R C×H×W ; Next, X K and X V Reshape to R C×N , where N = H × W represents the number of pixels, K and X Q Perform batch matrix multiplication between the transposes of X K and X Q The energy scores of any two pixel positions between are then calculated using a softmax layer to compute the spatial attention map E∈R N×N ; Then, the matrix X V Reshape to R N×N , in X V Perform batch matrix multiplication between E and the transpose of E, and then reshape the result into R C×H×W ; Finally, R C×H×W Multiply by parameter α, and perform residual connection on the feature X by β to obtain the final output P∈R C×H×W . Feature P is a weighted sum of features and original features across all positions. Feature P has global context features and weighted aggregated context based on pixel-level position information, which ensures the effective extraction of position features while maintaining global context information.

7. The method for medical image segmentation based on efficient Transformer UNet according to claim 6, characterized in that: The RCA adopts a channel attention mechanism, namely CAM. The difference between CAM and PAM is that CAM directly converts the original feature X∈R C×H×W Reshape to R C×N , we get the matrix X a , then in X b Perform batch matrix multiplication with its transpose, and to suppress noise, calculate the difference and adjust by comparing with the maximum value of each row or column; Then, a softmax layer is applied to obtain the channel attention map S∈R C×H×W , multiply the result by the ω scale parameter and perform an element-wise sum operation with X to obtain the final output I∈R C×H×W ,Finally, the final feature of each channel is the weighted sum of all channels and original features,,which gives CAM a powerful channel feature extraction capability.

8. The method for medical image segmentation based on efficient Transformer UNet according to claim 1, characterized in that: The RFA adopts a spatial attention mechanism, namely SAM, which extracts the importance weights of different features through the receptive field slider, prioritizes the extracted spatial features of the receptive field, and reconstructs the feature weights of the spatial features of the feature map in a more efficient way, thereby improving the model's ability to extract regional spatial features; First input the original feature X∈R C×H×W , after global average pooling and expansion method, its dimension becomes L rf ∈R C ×1×1 ,At the same time, a method for quickly extracting receptive field spatial features, namely Group Conv, is used to extract features using a 3×3 convolution kernel, where each 3×3 size window represents a receptive field slider; In this way, the original features are mapped to new features, which reduces the computational complexity of the convolution operation and improves the network's expressiveness by increasing the number of output channels. Specifically, the feature X∈R C×H×W First, group the features according to the number of input channels, then input each set of features into different convolution kernel groups for convolution operation, and then perform softmax operation. The output data is The RFA solves the problem of insensitivity to spatial information differences caused by position changes by emphasizing the importance of different features in the receptive field slider and giving priority to the receptive field spatial features.

9. The method for medical image segmentation based on efficient Transformer UNet according to claim 1, characterized in that: The loss function in S4 at least includes a Dice loss function and a Focal loss function; The mathematical expression of the Dice loss function is as follows: Where N is the number of input pixels; t i,j and p i,j denote the label value and prediction value of the pixel at position (i, j) respectively; T refers to the total number of segments contained in the background; ∈ refers to a smoothing term, which aims to avoid the denominator or numerator being zero. In order to make up for the shortcomings of the Dice loss function, the Focal loss function is introduced to guide the training of SSTrans-Net. The expressions of the cross entropy loss function and the Focal loss function are: Among them, p t is the classification probability of the current pixel; γ is a hyperparameter defined to control the intensity of penalizing misclassified pixels, usually set to 2; The Focal loss function assigns lower loss weights to pixels that are easy to classify and imposes greater penalties on pixels that are difficult to classify. This mechanism can not only effectively address the problem of imbalanced number of positive and negative samples, but also alleviate the challenges brought about by imbalanced category distribution. Unlike the simple combination of cross entropy and Dice loss functions used in most works, an improved method combining Focal and Dice losses is used. The improved method combining Focal and Dice losses not only focuses on the segmentation accuracy of the entire image, but also places more emphasis on pixels that are difficult to classify. The pixels that are difficult to classify include at least boundary areas and multi-scale pixels. The loss function expression of the overall loss function is proposed as follows: Loss=μDice Loss +φFocal Loss (18) Among them, μ and φ are weighting parameters. In this way, high-quality long-range dependent features in different channels can be filtered, thereby significantly improving the segmentation performance.

10. The method for medical image segmentation based on efficient Transformer UNet according to claim 1, characterized in that: The public evaluation indicators used in S5 are DSC and HD95.

Citation Information

Cited By

  • Diagnosis method and system for glenoid cavity and humeral head defect area

    CN120809145A