Semi-supervised medical image segmentation method based on improved Transform

Through the improved semi-supervised medical image segmentation method, the spherical embedded Transformer module and the morphological difference capture module are used, and combined with the adversarial generation network discriminator, the problems of high computing costs and high annotation costs in the prior art are solved, and efficient and accurate segmentation of medical image segmentation is achieved.

CN120259640APending Publication Date: 2025-07-04ZHONGYUAN ENGINEERING COLLEGE
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510113508.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing medical image segmentation technology has high computational cost and dependence on large-scale labeled datasets in capturing global context information. The annotation of medical images is expensive, and blurred images and complex boundary conditions increase the difficulty of segmentation.

Method used

Using a bilateral architecture based on U-shaped structure, the spherical embedded Transformer module is used to enhance the long-range dependency modeling capabilities, combined with the morphological difference capture module and the multi-scale information extraction module, an adversarial generation network discriminator is designed to evaluate the authenticity of pseudo-labels, and an exponential moving average strategy is used to improve model stability and generalization performance, and semi-supervised learning is performed.

Benefits of technology

In the case of limited resources, the accuracy and robustness of medical image segmentation are improved, the cost of data labeling is reduced, and the segmentation accuracy and model generalization ability are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259640A_ABST
    Figure CN120259640A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medicine, and particularly discloses a semi-supervised medical image segmentation method based on an improved Transform, which designs a sphere embedded improved Transform module, constructs a boundary enhancement module based on morphological difference, and aims to improve the medical image segmentation performance. According to the method, a semi-supervised learning framework is adopted, and the dependence on a large-scale annotated data set is reduced by effectively utilizing limited annotated data and rich unannotated data, so that the accurate segmentation of the medical image is realized under the condition of limited resources. The method is expected to reduce the cost and time of data annotation while improving the segmentation precision, and brings a new breakthrough to the field of medical image segmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical image segmentation, and particularly to a semi-supervised medical image segmentation method based on an improved Transformer. Background Art

[0002] Medical image segmentation technology is an advanced artificial intelligence method that can accurately identify and separate lesion regions from surrounding tissues. This technology can not only relieve the tedious work of manually annotating lesion regions, but also significantly improve the accuracy and efficiency of medical diagnosis, greatly improving the work efficiency of clinicians. In clinical diagnosis and treatment planning, medical image segmentation technology plays a crucial role. It can not only improve the quality of medical services, but also significantly improve the treatment effect of patients, having a profound impact on the medical field. Methods based on convolutional neural networks (CNNs) perform well in extracting local features, but they have limitations in capturing global context information. Although methods based on Transformer can capture global context information through their global perception ability, they face high computational costs and dependence on large-scale labeled datasets. In medical image segmentation, obtaining high-quality annotated data is costly and time-consuming. In addition, blurred images and complex boundary conditions pose additional challenges to the segmentation task, so the recognition and understanding of the segmentation target boundary are crucial for improving segmentation accuracy.

[0003] Based on this, there is now provided a semi-supervised medical image segmentation method based on an improved Transformer, which can eliminate the drawbacks of the prior art. Summary of the Invention

[0004] The purpose of the present invention is to provide a semi-supervised medical image segmentation method based on an improved Transformer, which solves the problems in the prior art.

[0005] To achieve the above purpose, the present invention provides the following technical solutions:

[0006] A semi-supervised medical image segmentation method based on improved Transformer. In this invention, a bilateral architecture Network-A (denoted as fA) and Network-B (denoted as fB) based on the U-shaped structure are adopted. The input image is fed into the bilateral architecture. fA serves as the teacher network and is pre-trained on ImageNet to provide stable feature representations for the model. fB serves as the student network and learns from the teacher network and is fine-tuned without any preprocessing. In both networks, the encoder part uses the designed Sphere-embedding transformer (SEFormer) module to extract key features. This module enhances the long-range dependence modeling ability of Transformer through spherical embedding and random rotation strategies. The skip connection part uses the designed Morphological difference catcher (MDC) module to capture morphological difference features at different scales in both networks, enhancing the feature expression ability and focusing on improving the representation of the boundary of the segmented target area.

[0007] The bottleneck part uses the designed Multi-scale information extractor (MIE) module to extract multi-scale information in both networks, improving the richness of the extracted features and enhancing the learning ability and generalization ability of the model.

[0008] The loss function includes the key pixel loss function L k 、the MDC boundary loss function L b 、the similarity loss function L sim 、the discriminator loss function L based on GAN Disc and the segmentation loss function L seg . The differences between the key pixels processed by the SEFormer module between fA and fB, the boundary information processed by the MDC module, the similarity between the prediction results of the two networks, the discrimination accuracy of the discriminator for pseudo-labels and real labels, and the difference between the model prediction result and the real result are calculated respectively, which are used to guide learning, optimize boundary detection, and promote collaborative learning, etc.

[0009] This invention designs a reliability discriminator based on the Generative adversarial network (GAN) to improve the performance of semi-supervised learning by evaluating the authenticity of the pseudo-labels generated by the model. The Exponential moving average (EMA) strategy is adopted to enhance the training stability and generalization performance of the model. Finally, the segmentation result is output using the trained network.

[0010] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0011] The present invention designs a reliability discriminator based on a Generative Adversarial Network (GAN), which improves the performance of semi-supervised learning by evaluating the authenticity of the pseudo-labels generated by the model. An Exponential Moving Average (EMA) strategy is adopted to enhance the training stability and generalization performance of the model. Finally, the segmentation result is output using the trained network. Description of the Drawings

[0012] Figure 1 It is the training execution flowchart of the present invention.

[0013] Figure 2 It is the structural diagram of SM-Net of the present invention.

[0014] Figure 3 It is the structural diagram of SEFormer of the present invention.

[0015] Figure 4 It is the structural diagram of MDC of the present invention.

[0016] Figure 5 It is the visualization display result diagram of the segmentation result of the present invention on the ISIC2017 dataset.

[0017] Figure 6 It is the heat map obtained by SM-Net using SEFormer on the ISIC2017 dataset of the present invention. Detailed Embodiments

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] Such as Figures 1-6As shown in the figure, a bilateral architecture Network-A (denoted as fA) and Network-B (denoted as fB) based on a U-shaped structure are adopted. The input image is fed into the bilateral architecture. fA serves as the teacher network and is pre-trained on ImageNet in advance to provide stable feature representations for the model. fB serves as the student network and learns from the teacher network and is fine-tuned without any preprocessing. In both networks, the encoder part uses the designed Sphere-embedding transformer (SEFormer) module to extract key features. This module enhances the long-range dependence modeling ability of the Transformer through spherical embedding and random rotation strategies. The skip connection part uses the designed Morphological difference catcher (MDC) module to capture morphological difference features at different scales in both networks, enhancing the feature expression ability and focusing on improving the representation of the boundary of the segmented target area.

[0020] The bottleneck part uses the designed Multi-scale information extractor (MIE) module to extract multi-scale information in both networks, improving the richness of the extracted features and enhancing the learning ability and generalization ability of the model.

[0021] The loss function includes the key pixel loss function L k , the MDC boundary loss function L b , the similarity loss function L sim , the discriminator loss function L Disc based on GAN, and the segmentation loss function L seg . The differences in key pixels after the SEFormer module processing between fA and fB, the differences in boundary information after the MDC module processing, the similarity between the prediction results of the two networks, the discriminative accuracy of the discriminator for pseudo-labels and real labels, and the differences between the model prediction results and the real results are calculated respectively, which are used to guide learning, optimize boundary detection, and promote collaborative learning, etc.

[0022] The present invention designs a reliability discriminator based on a Generative adversarial network (GAN) to improve the performance of semi-supervised learning by evaluating the authenticity of the pseudo-labels generated by the model. The Exponential moving average (EMA) strategy is adopted to enhance the training stability and generalization performance of the model. Finally, the segmentation result is output using the trained network;

[0023] The objective of the present invention is to achieve accurate segmentation of medical images under limited resources by effectively utilizing limited labeled data and abundant unlabeled data. To this end, a deep neural network structure SM-Net as shown in Figure 2 is designed. This network structure adopts a semi-supervised strategy and mainly includes an encoder, a decoder, and multiple key modules, such as SEFormer, MDC, MIE, and a GAN-based pseudo-label discriminator. In SM-Net, both labeled and unlabeled data are input into fA and fB. fA serves as the teacher network and provides stable feature extraction performance for the model through pre-training. fB serves as the student network, learns from network fA, and is fine-tuned. In the first three layers of the encoder in network fA and network fB, SEFormer is used to extract key features, improving the network's long-range dependency modeling ability. And a mask Mask is designed based on the SEFormer module. Through partial feature operations during the training process, data augmentation is performed. The mask uses the key pixel loss function L k , captures the differences in key pixels of the encoded features of the SEFormer module between fA and fB, and is used to guide learning. The designed MDC module is adopted on the skip connection. By capturing the feature differences at different scales in network fA and network fB, the feature expression ability of the network is enhanced. The boundary loss function L b is constructed to calculate the differences in the boundary information processed by the MDC module between fA and fB, and is used to optimize the target boundary representation. The designed MIE module is adopted in the bottleneck part of the encoder in network fA and network fB to extract multi-scale information, improving the richness of feature expression. SM-Net simultaneously inputs both labeled and unlabeled data into the GAN-based reliability discriminator to evaluate the authenticity of the pseudo-labels generated by network fA, thereby improving the performance of semi-supervised learning. At the same time, the similarity loss function L sim is introduced to minimize the differences between the fA and fB models. SM-Net can extract rich features from the input images and simultaneously learn on labeled and unlabeled data to improve the accuracy and robustness of segmentation.

[0024] The encoder retains the CNN-based method for local detail extraction, while the decoder utilizes these extracted features for reconstruction. This design of the CNN-based encoder and decoder enables the model to capture the local features of the input data and, through the upsampling process of the decoder, recover the output with rich details. The designed SEFormer module uses spherical embedding and random rotation strategies for key pixel extraction, thus effectively improving the network's ability to model long-range dependencies. To make full use of the rich semantic information in the highest-level encoded features, the present invention uses the designed MIE module to capture multi-scale information to enhance the learning and generalization ability of the model. The design of the MDC module focuses on paying attention to boundary information to address fuzzy and complex boundary problems and improve the high-precision boundary perception ability of the segmentation network. To make more effective use of unlabeled data, SM-Net introduces a GAN-based reliability discriminator, which can evaluate the authenticity of the generated pseudo-labels and further improve the performance of semi-supervised learning. When dealing with medical image segmentation tasks, SM-Net can not only improve the accuracy of segmentation but also enhance the utilization efficiency of the model for unlabeled data, which is particularly important in resource-constrained clinical environments;

[0025] The labeled data is represented as and the unlabeled data is represented as The entire dataset is represented as D = D l ∪ D u , where N >> M. The outputs of the encoder, SEFormer, MDC, and decoder are represented as E i , S i , M i and D i respectively;

[0026] To effectively obtain the global information in medical images, while filtering out irrelevant information and modeling long-range dependencies, the present invention designs the SEFormer module, as shown in Figure 3 . The SEFormer takes the output E i of the encoder as the input and uses a segmentation operation to process E iThe segmented features are mapped to the unit sphere through a spherical embedding operation, and each sub-region corresponds to a part of the surface on the sphere. This division strategy enables the points on the sphere to represent the corresponding regional features in the original feature map, thereby improving the feature expression ability while maintaining the topological structure between pixels in the original feature map. In order to achieve full interaction of global information, SEFormer introduces spherical embedding and multi-angle flipping strategies, which enhance the multi-angle representation ability of features through horizontal flipping, vertical flipping and diagonal flipping. Then the features are rotated at random angles using the random rotation operation to enhance the generalization ability of the model. By performing key point extraction operations on the rotated features, key information is obtained. These key points are used as queries (Q), keys (K) and values ​​(V) of the cross attention mechanism after linear mapping. By using Q, K and V for attention calculation, the feature representation of the key area is enhanced to obtain the fused feature E. fuse .

[0027] The specific implementation process is as follows:

[0028]

[0029] Where Φ() represents random rotation and key pixel extraction operations. H(), V(), and D() represent horizontal flip, vertical flip, and diagonal flip, respectively. Θ() represents the spherical embedding mapping function. R represents the Reshape operation. d represents E K Dimension.

[0030] In order to map the features into a sphere, the following spherical coordinates are constructed:

[0031]

[0032] Where θ and φ represent longitude and latitude respectively. H and W represent the height and width of the feature. i and j represent the feature index before spherical embedding.

[0033] Use coordinate transformation to transform spherical coordinates into three-dimensional Cartesian coordinates. The specific transformation formula is as follows:

[0034]

[0035] Cartesian coordinates are represented by (x, y, z), and are rotated around the z-axis. The rotated coordinates are as follows:

[0036]

[0037] Where Δ represents random rotation. x'=x, y'=y are the key point coordinates obtained.

[0038] For Efuse Adjust the weights of different channels using Efficient Channel Attention (ECA) to enhance the representation ability of important features. ECA takes E fuse as the input, first performs global average pooling to aggregate the information in the spatial dimension into a single global feature vector, and then uses a convolutional kernel adaptive selection strategy to perform a convolutional operation on the feature map. The result of the convolutional operation is input into the Sigmoid function, which maps it between 0 and 1 and is used to calculate the weight of each channel. The calculated weight vector is multiplied pointwise with the original feature map to obtain the ECA output. SEFormer simultaneously introduces Spatial Attention (SA) and residuals to enhance the feature representation ability. SA first calculates the Global Average Pooling (GAP) and Global Max Pooling (GMP) of the feature map. After the concatenation operation, a 1×1 convolution is used to adjust the number of channels of the feature map. After the Softmax operation, it provides additional information for channel attention, and then a residual connection is made with the original feature to enhance the feature representation ability of SEFormer;

[0039] The output S of SEFormer i The specific calculation process is as follows:

[0040]

[0041] Among them, GAP and GMP represent global average pooling and global max pooling respectively. Concat represents the concatenation operation. Conv 1×1 represents a 1×1 convolution. ECA is the efficient channel attention used to adjust the weights of different channels.

[0042] To make full use of the rich semantic information in the highest-level encoded features, the present invention uses the designed MIE module to capture multi-scale information to enhance the learning and generalization ability of the model. The encoded features output by the highest-level encoder are used as the input of MIE, and then multi-scale convolutional operations with a dilation rate of 2 and a convolutional kernel of 5×5, a dilation rate of 3 and a convolutional kernel of 3×3, a dilation rate of 2 and a convolutional kernel of 3×3, a dilation rate of 1 and a convolutional kernel of 3×3, and a dilation rate of 1 and a convolutional kernel of 1×1 are respectively performed. Then these features are concatenated, and finally output to the decoder through Batch Normalization (BN) and the activation function ReLU (Rectified Linear Unit, ReLU). In addition, SM-Net adopts the Exponential Moving Average (EMA) strategy to transfer parameters between fA and fB, reduce the fluctuations during each parameter update, optimize the training stability of the model, and improve the generalization performance.

[0043] In medical images, the accurate identification and segmentation of lesion regions are crucial for disease diagnosis and treatment planning. To address the efficient utilization of boundary information, the present invention designs an MDC module to improve the segmentation accuracy of lesion region boundaries by enhancing the boundary representation of the target.

[0044] The specific structure is as Figure 4 shown;

[0045] MDC respectively receives the output S i from SEFormer and the output D i from the decoder. i S is divided into two branches, and morphological dilation and erosion operations are respectively performed on them. These two operations help capture the boundary information in the image. After the morphological dilation operation and erosion operation, D e and E e are obtained respectively. Then, a set of separable convolutions (1×1; 3×1, 1×3; 5×1, 1×5) are used to process the features D e and E e . The specific process is as follows: D e is divided into three branches, and 5×1 and 1×5 convolutions (Conv 5×1 and Conv 1×5 ), 3×1 and 1×3 convolutions (Conv 3×1 and Conv 1×3 ), and 1×1 convolution (Conv 1×1 ) are respectively performed on them, and then they are concatenated to obtain X D ; E e is divided into three branches, and Conv 1×1 , Conv 3×1 , and Conv 1×3 , and Conv 5×1 , and Conv 1×5 are respectively performed on them, and then they are concatenated to obtain X E ; The specific formula is as follows:

[0046]

[0047] where X D , X E are respectively the outputs of the dilation and erosion operations.

[0048] Then, a Cross-branch difference module (CDM) is designed and inserted into the separable convolution. After dilation and erosion operations, the output of the first strip convolution is used as the input of CDM. Since the strip convolution extracts information from different receptive fields, there are differences between this information, and these differences contribute to the localization of the segmentation target boundary in medical images. In CDM, the Conv 1×1 outputs in the dilation branch are respectively subtracted from the Conv 3×1 and Conv 5×1 outputs in the erosion branch, and the Conv 1×1 output in the erosion branch is respectively subtracted from the Conv 3×1 and Conv 5×1 outputs in the dilation operation. By performing this subtraction operation on the two branches of CDM to extract differential information, and then through element-wise addition, batch normalization (BN), and ReLU activation operations, the differential information X Diff is obtained. The specific formula is as follows:

[0049]

[0050] where D e and E e represent the output features of the dilation and erosion operations respectively. Conv 3×1 and Conv 5×1 denote strip convolutions with kernel sizes of 3×1 and 5×1. Sub represents the element-wise subtraction operation. X Diff represents the output of CDM.

[0051] X Diff and X D , X E and D i are jointly input into the designed Multi-scale fusion module (MFM) for fusing features at different scales to form a rich feature representation M i . The specific design of the MFM module is as follows: On the one hand, the output D i of the decoder is used to learn Q, k, and V through linear projection, and the output X Diff of CDM is element-wise multiplied with Q. On the other hand, X D and X E are first element-wise subtracted to obtain X b , then passed through ReLU, followed by global average pooling and global max pooling operations respectively, and finally concatenated and passed through Sigmoid activation to obtain W b ; then W b is multiplied with Di The obtained V performs element-wise multiplication. The MFM module introduces difference information and boundaries into the attention calculation with Q, k, and V as additional weights to enhance the boundary expression in the feature map. In addition, depth-wise convolution and point-wise convolution are innovatively introduced into the multi-scale fusion module (MFM) to strengthen the channel and pixel information fusion of features. The specific calculation formula of MFM is as follows:

[0052]

[0053] where Sub represents element-wise subtraction. ε is the sigmoid function used to generate the boundary weight w b . Max and Avg represent the Max pooling and Average pooling operations respectively. * represents the multiplication operation. d k represents the dimension of K. DWC (Depth-wise convolution, DWC) and PWC (Point-wise convolution, PWC) represent depth-wise convolution and point-wise convolution respectively.

[0054] The MDC module effectively enhances the boundary representation of the segmented target region in medical images through the combination of morphological difference operations, separable convolutions, cross-branch difference modules, and multi-scale fusion modules, thereby improving the recognition and segmentation accuracy of the region of interest.

[0055] SM-Net adopts a semi-supervised learning strategy and uses limited labeled data and a large amount of unlabeled data for training. This strategy is implemented through a bilateral structure, including a pre-trained network fA (as the teacher) and an unpreprocessed network fB (as the student). In addition, a GAN-based pseudo-label reliability discriminator is introduced to evaluate the authenticity of the generated pseudo-labels, thereby improving the performance of semi-supervised learning. In the semi-supervised segmentation model, pseudo-labels are the results predicted by the model, and they are used as additional training data to improve the model performance. However, these pseudo-labels are usually accompanied by uncertainties, which may mislead the model training and thus seriously affect the model performance. To solve the above problems, the present invention designs a reliability discriminator based on the generative adversarial network (GAN) to guide the generator to generate high-quality pseudo-labels. In this framework, fA is regarded as the generator, and the pre-trained VGGNet16 network plus a fully connected layer and a Sigmoid activation function are used as the discriminator.

[0056] As Figure 2 shown, the unlabeled data and the labeled data Are used as the inputs of the GAN-based reliability discriminator. These inputs will correspond to the prediction results of fA and are concatenated. The concatenated result is fed into the discriminator of the GAN to calculate the reliability score Will be used as weights for adaptive adjustment. The loss of the segmentation result between fA and fB is as follows:

[0057]

[0058] The task of the generator fA is to continuously learn and generate pseudo-labels that are closer and closer to the true labels. The task of the discriminator is to distinguish the pseudo-labels. Through this adversarial mechanism, fA generates higher-quality pseudo-labels and improves its performance on unlabeled data. In addition, the discriminator provides feedback to the generator based on multi-level data, enabling the quality of the pseudo-labels to be continuously optimized. By introducing the adversarial learning mechanism, the generator is more rigorous when generating pseudo-labels, and at the same time the discriminator provides feedback to help the generator continuously improve, ultimately improving the performance of the model when dealing with unlabeled data.

[0059] In terms of the design of the loss function of SM-Net, on the one hand, after being processed by the SEFormer module, the key pixels are more prominent than other pixels. Therefore, the present invention constructs a key pixel loss function to measure the difference between the outputs of SEFormer in fA and fB. The specific method is as follows: First, the difference is normalized to the range of (0,1), and then a binary Mask operation is performed. Values less than 0.5 are set to 0, and values greater than or equal to 0.5 are set to 1. The specific formula is:

[0060]

[0061] Among them, P Key represents the normalized result, which consists of 0 and 1. N is the normalization and masking operation, which compresses the data to (0,1) and binarizes it. Att is the self-attention operation.

[0062] Using the Dice loss, the key pixel loss function L k is as follows:

[0063]

[0064] Among them, P key_i,a ∈(0,1) and P key_i,b ∈(0,1) respectively represent the key pixels of the masks Mask of fA and fB. The parameter δ is a small positive number to prevent the denominator from being zero. N = H×W is the number of pixels.

[0065] In addition, the boundary loss L b, the Dice loss function is used, and the specific formula is as follows:

[0066]

[0067] Among them, P b_i,a ∈(0,1) and P b_i,a ∈(0,1) represent the predicted boundary maps of fA and fB respectively.

[0068] To further promote the collaborative learning between fA and fB, a similarity loss function L sim is constructed between the prediction results of fA and fB, and the formula is as follows:

[0069]

[0070] Among them, p a_i and p b_i represent the predicted probabilities of fA and fB respectively, within the range of (0, 1). represents the reliability score of the pseudo-label obtained based on GAN.

[0071] To make the discriminator more suitable for the medical image segmentation scenario, the present invention designs a discriminant loss function L Disc to calculate the loss on the labeled data. Using the probability weight map of each pixel value being true in the prediction result output by the reliability discriminator based on GAN, the loss function L Disc is constructed, and the specific formula is as follows:

[0072]

[0073] Among them, r i represents the probability weight at the i-th index of the reliability distribution generated by the discriminator based on GAN, g i ∈{0,1} represents the true value, and p i represents the predicted value. Finally, the segmentation loss function is used to measure the difference between the segmentation result and the true value (GT). The segmentation loss L seg is used not only for fA but also for fB. The present invention uses binary cross-entropy to measure the segmentation loss, and the formula is as follows:

[0074]

[0075] Among them, g i ∈{0,1} and p i ∈(0,1) represent the true value and the predicted probability respectively.

[0076] Based on the above analysis, the total loss function can be expressed as follows:

[0077] L total= L seg+L sim +L Dis +α·L k +β·L b (36)

[0078] Wherein, α and β are balance coefficients between the key pixel loss and the boundary loss.

[0079] After obtaining the prediction loss, the Adam optimizer is used to optimize the segmentation network, and the optimized network parameters can be used for medical image segmentation. The results of the method of the present invention for medical image segmentation are shown below.

[0080] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A semi-supervised medical image segmentation method based on an improved Transformer, characterized in that: It includes the following steps: Step 1: A bilateral architecture fA and fB based on a U-shaped structure are adopted. The input image is fed into the bilateral architecture. fA serves as the teacher network, which is pre-trained on ImageNet to provide stable feature representations for the model. fB serves as the student network, which learns from the teacher network fA without any preprocessing and is fine-tuned. Step 2: In both networks, the encoder part uses the designed spherical embedding Transformer module SEFormer to extract key features. This module enhances the long-range dependency modeling ability of the Transformer while reducing the computational overhead through spherical embedding and random rotation strategies. The skip connection part uses the designed morphological difference capture module to capture morphological difference features at different scales in both networks, enhancing the feature expression ability and focusing on improving the representation of the boundaries of the segmented target regions. Step 3: The bottleneck part uses the designed multi-scale information extraction module to extract multi-scale information in both networks, improving the richness of the extracted features and enhancing the learning ability and generalization ability of the model. The loss function includes the key pixel loss function L k , the MDC boundary loss function L b , the similarity loss function L sim , the discriminator loss function L Disc based on GAN, and the segmentation loss function L seg , and respectively calculate the differences in key pixels processed by the SEFormer module between fA and fB, the differences in boundary information processed by the MDC module, the similarity between the prediction results of the two networks, the discrimination accuracy of the discriminator for pseudo-labels and real labels, and the differences between the model prediction results and the real results; Step 4: Design a reliability discriminator based on a generative adversarial network. By evaluating the authenticity of the pseudo-labels generated by the model, the performance of semi-supervised learning is improved. The exponential moving average strategy is used to enhance the training stability and generalization performance of the model. Finally, the segmentation results are output using the trained network.

2. The semi-supervised medical image segmentation method based on the improved Transformer according to claim 1, wherein, In step 3 when dealing with medical image segmentation tasks: The labeled data is represented by , and the unlabeled data is represented by . The entire dataset is represented as D = D l ∪ D u , where N >> M; the outputs of the encoder, SEFormer, MDC, and decoder are represented by E i , S i , M i , and D i respectively.

3. The semi-supervised medical image segmentation method based on the improved Transformer according to claim 1, wherein In step 3: SEFormer introduces spherical embedding and multi-angle flipping strategies. Through horizontal flipping, vertical flipping, and diagonal flipping, it enhances the multi-angle representation ability of features. Then, it uses random rotation operations to randomly rotate the features by a certain angle, enhancing the generalization ability of the model. By performing key-point extraction operations on the rotated features, key information is obtained. These key points are used as the query Q, key K, and value V of the cross-attention mechanism after linear mapping. By using Q, K, and V for attention calculation, the feature representation of the key region is enhanced, and the fused feature E is obtained fuse ; The specific implementation process is as follows: Among them, Φ(·) represents the random rotation and key pixel extraction operations; H(·), V(·), and D(·) represent horizontal flipping, vertical flipping, and diagonal flipping respectively; Θ(·) represents the spherical embedding mapping function; R represents the Reshape operation; d represents the dimension of E K ; To map the features into a sphere, the following spherical coordinates are constructed: where θ and φ represent longitude and latitude respectively; H and W represent the height and width of the features; i and j represent the feature indices before spherical embedding. The spherical coordinates are converted into three-dimensional Cartesian coordinates using coordinate transformation. The specific transformation formula is as follows: The Cartesian coordinates are represented by (x, y, z), and a rotation operation is performed with the z-axis as the central axis. The rotated coordinates are as follows: where Δ represents random rotation; x' = x, y' = y are the obtained key point coordinates. For E fuse Use efficient channel attention (ECA) to adjust the weights of different channels and enhance the representation ability of important features; ECA takes E fuse as the input, first performs global average pooling to aggregate the information in the spatial dimension into a single global feature vector, and then uses a convolutional kernel adaptive selection strategy to perform a convolutional operation on the feature map; the result of the convolutional operation is input into the Sigmoid function, which maps it between 0 and 1 for calculating the weight of each channel; the calculated weight vector is multiplied pointwise with the original feature map to obtain the ECA output; SEFormer simultaneously introduces spatial attention and residuals to improve the feature expression ability; SA first calculates the global average pooling and global maximum pooling of the feature map, and after a concatenation operation, uses a 1×1 convolution to adjust the number of channels of the feature map; after a Softmax operation, it provides additional information for channel attention, and then performs a residual connection with the original feature to improve the feature expression ability of SEFormer.

4. The output S of SEFormer in the semi-supervised medical image segmentation method based on the improved Transformer according to claim 3 i The specific calculation process is as follows: Among them, GAP and GMP represent global average pooling and global max pooling respectively; Concat represents the concatenation operation, and Conv 1×1 represents a 1×1 convolution; ECA is the efficient channel attention, which is used to adjust the weights of different channels.

5. The semi-supervised medical image segmentation method based on the improved Transformer according to claim 1, wherein: The MDC receives the output S from the SEFormer i and the output D of the decoder i ; divides S i into two branches, and performs morphological dilation and erosion operations respectively. These two operations help capture boundary information in the image; after the morphological dilation operation and the erosion operation, D e and E e are obtained respectively. Then, a group of separable convolutions (1×1; 3×1, 1×3; 5×1, 1×5) are used to process the features D e and E e ; the specific process is as follows: divides D e into three branches, and performs 5×1 and 1×5 convolutions (Conv 5×1 and Conv 1×5 ), 3×1 and 1×3 convolutions (Conv 3×1 and Conv 1×3 ), and 1×1 convolution (Conv 1×1 ) respectively, and then splices them to obtain X D ; divides E e into three branches, and performs Conv 1×1 , Conv 3×1 , and Conv 1×3 , and Conv 5×1 and Conv 1×5 respectively, and then splices them to obtain X E ; the specific formula is as follows: where X D and X E are the outputs of dilation and erosion operations, respectively; Then, a Cross-branch difference module (CDM) is designed and inserted into the separable convolution. The output of the first strip convolution after dilation and erosion operations is used as the input of CDM. Since the strip convolution extracts information from different receptive fields, there are differences between these information, and these differences are helpful for the localization of the segmentation target boundary in medical images. In CDM, the outputs of Conv 1×1 in the dilation branch are respectively subtracted from the outputs of Conv 3×1 and Conv 5×1 in the erosion branch, and the outputs of Conv 1×1 in the erosion branch are respectively subtracted from the outputs of Conv 3×1 and Conv 5×1 in the dilation operation. By performing this subtraction operation on the two branches of CDM to extract the differential information, and then through element-wise addition, batch normalization (BN) and ReLU activation operations, the differential information X Diff is obtained. The specific formula is as follows: Among them, D e and E e respectively represent the output features of dilation and erosion operations; Conv 3×1 and Conv 5×1 represent strip convolutions with kernel sizes of 3×1 and 5×1; Sub represents an element-wise subtraction operation; X Diff represents the output of CDM.

6. The semi-supervised medical image segmentation method based on the improved Transformer according to claim 1, characterized in that: Depthwise convolution and pointwise convolution are innovatively introduced into the multi-scale fusion module (MFM) to strengthen the channel and pixel point information fusion of the features. The specific calculation formula of MFM is as follows: wb = ε(Conv1×1(Concat(Max(ReLU(Xb)),Avg(ReLU(Xb))))) (11) Among them, Sub represents element-wise subtraction; ε is the sigmoid function used to generate the boundary weight w b ; Max and Avg represent Max pooling and Average pooling operations respectively; * represents multiplication operation; d k represents the dimension of K; DWC (Depth-wise convolution, DWC) and PWC (Point-wise convolution, PWC) represent depth-wise convolution and point-wise convolution respectively.

7. The semi-supervised medical image segmentation method based on the improved Transformer according to claim 1, wherein: Unlabeled data and labeled data are used as inputs to the GAN-based reliability discriminator; these inputs are concatenated with the corresponding prediction results of fA and and the concatenated result is fed into the discriminator of the GAN to calculate the reliability score will be used as a weight to adaptively adjust the loss of the segmentation result between fA and fB, and the specific formula is as follows: The task of the generator fA is to continuously learn and generate pseudo-labels that are increasingly close to the true labels, and the task of the discriminator is to distinguish the pseudo-labels.

8. The semi-supervised medical image segmentation method based on the improved Transformer according to claim 1, wherein Construct a loss function \(L\) using the probability weight map where each pixel value is true in the prediction result output by the GAN-based reliability discriminator Disc , and the specific formula is as follows: where r i represents the probability weight at the i-th index of the reliability distribution generated by the GAN-based discriminator, g i ∈ {0, 1} represents the true value, p i represents the predicted value; finally, a segmentation loss function is used to measure the difference between the segmentation result and the true value (GT); the segmentation loss L seg is used not only for fA but also for fB; the present invention uses binary cross-entropy to measure the segmentation loss, and the formula is as follows: where g i ∈ {0, 1} and p i ∈ (0, 1) represent the true value and the predicted probability respectively; Based on the above analysis, the total loss function can be expressed as follows: L total= L seg +L sim +L Disc +α·L k +β·L b (16) where α and β are the balance coefficients between the key pixel loss and the boundary loss. After obtaining the prediction loss, the Adam optimizer is used to optimize the segmentation network, and the optimized network parameters can be used for medical image segmentation.

Citation Information

Cited By

  • Boundary enhancement model for semi-supervised medical image segmentation

    CN121095274A

  • Semi-supervised medical image segmentation method and device of collaborative training network based on integrated attention correction, and medium

    CN121259317A

  • Medical operation automatic scoring method based on deep learning

    CN121725521A