Transfer Learning Method and System Integrating Swin Transformer and UNet for Segmentation Tasks

By fusing the transfer learning method of SwinTransformer and UNet, combined with the domain discriminant network of spatial position attention and gradient inversion layer, the problems of low sample utilization efficiency and inconsistent domain distribution in medical image segmentation are solved, and better segmentation and transfer learning effects are achieved in the target domain.

CN114511703BActive Publication Date: 2025-06-17SUZHOU YIZHIYING INTELLIGENT TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210076477.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-21
Publication Date
2025-06-17
Estimated Expiration
2042-01-21

AI Technical Summary

Technical Problem

Existing deep learning technologies are difficult to effectively utilize a small number of samples in medical image segmentation, and they face the problem of inconsistent distribution of source and target domains, resulting in poor performance in the target domain.

Method used

Using a transfer learning method that combines SwinTransformer and UNet, features of different scales are extracted through SwinTransformer and introduced into UNet's decoder, combining spatial position attention-weighted operation. At the same time, a domain discriminating network with a gradient inversion layer is introduced to determine whether the features come from the source domain or the target domain.

Benefits of technology

It improves the stability and accuracy of the segmentation effect in the target domain, can effectively utilize a small number of samples, and extract public feature spaces from the source domain and the target domain to achieve better transfer learning effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114511703B_ABST
    Figure CN114511703B_ABST
Patent Text Reader

Abstract

The present invention relates to a transfer learning method and system that integrates Swin Transformer and UNet for segmentation tasks, including: Step 1, input an image to be segmented with Nc channels into the feature extraction network F; Step 2, restore the vectorized features of different scales extracted by the feature extraction network F into feature maps of different scales; Step 3, input the feature maps of different scales into the segmentation network S to obtain the segmentation results of the segmentation objects required for both the source domain and the target domain; Step 4, input the feature maps of different scales into the domain discriminant network D to determine whether the features come from the source domain or the target domain and give corresponding labels; Step 5, calculate the segmentation loss part of the source domain training samples, the segmentation loss part of the target domain training samples, and the discriminant loss part of the domain discriminant network, and weight and superimpose the above three parts to obtain the overall loss; Step 6, by minimizing the overall loss, iteratively optimize until the overall loss meets the requirements, and complete the transfer learning process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of machine learning and image segmentation, and in particular to a migration learning method and system for segmentation tasks that integrates SwinTransformer and UNet. Background Art

[0002] With the rapid development of artificial intelligence technology in recent years, the automatic delineation of target and organs at risk (OAR) based on deep learning (DL) has made significant progress. Although the automatic delineation of target and organs at risk based on deep learning has achieved good results, the establishment of a reliable and robust segmentation model depends on large-scale data labeled according to consistent standards. Since the cost of obtaining large-scale data annotation is very high, it is difficult to obtain in many practical application scenarios, such as:

[0003] (1) High labeling cost: Since the expert group’s time is precious, they can usually only focus on discussing a few to a dozen patients, and cannot complete large-scale data labeling;

[0004] (2) Individualized target delineation: Different doctors in different hospitals have different target delineation details in practice. The universal target delineation effect cannot meet their needs. Therefore, it is necessary to establish an individualized target delineation model. However, it is difficult to obtain a large amount of data due to the limited number of patients treated by a single doctor.

[0005] (3) Radiotherapy for rare diseases: such as childhood neuroblastoma, which has a low incidence and difficult to obtain sufficient data.

[0006] Therefore, how to improve existing deep learning methods so that they can more efficiently utilize sample information and achieve similar learning effects with fewer samples is an urgent clinical need and a hot topic in current research.

[0007] In addition, most of the current supervised transfer learning technologies for medical image segmentation are mainly aimed at the situation where the source domain and the target domain have inconsistent distributions, that is, the source domain data type is consistent with the target domain data type, and the source domain and the target need to be segmented are the same, but due to the different image acquisition conditions of the source domain and the target domain, the distribution of images in the two domains is inconsistent, resulting in the problem that the source domain model performs poorly in the target domain. When the adversarial method based on the domain discriminator aligns the features of the source domain and the target domain or extracts domain-invariant features, the highest-level features are usually used, and features of different scales are not used for domain discrimination. Summary of the invention

[0008] To solve the above technical problems, the present invention combines the advantages of Swin Transforer and UNet. It uses Swin Transformer as the encoder and adopts different-scale skip connections in UNet to introduce the features of different scales extracted by Swin Transformer into the decoder based on convolutional operations. A spatial position attention weighting operation is also added to the decoder. In addition, a domain discriminant network D with a gradient reversal layer is adopted in this solution. The domain discriminant network takes the features of different scales extracted by Swin Transformer as input to determine whether the current features come from the source domain or the target domain. Introducing the domain discriminant network in the way of the gradient reversal layer is easier to train.

[0009] The technical solution of the present invention is: A transfer learning method for fusing Swin Transformer and UNet for segmentation tasks, including the following steps:

[0010] Step 1: Input the image data to be segmented into the feature extraction network. The feature extraction network F uses the multi-scale multi-window attention mechanism of Swin Transformer as the backbone network to extract feature vectors of different scales. The input image to be segmented may contain Nc channels.

[0011] Step 2: Restore the feature vectors of different scales obtained by the feature extraction network to feature maps of different scales.

[0012] Step 3: Input the feature maps of different scales into the segmentation network to obtain the segmentation results of the required segmentation objects in two domains, the source domain and the target domain. The segmentation network S is a UNet decoding network with a spatial attention mechanism added.

[0013] Step 4: Input the feature maps of different scales into the domain discriminant network to determine whether the feature maps come from the source domain or the target domain and give corresponding labels. The domain discriminant network D includes an improved UNet encoding network combined with the UNet skip connection method, as well as two-level fully connected layers and an output layer.

[0014] Step 5: Calculate the source domain segmentation loss part for the source domain required segmentation object results obtained by passing the source domain training samples through the feature extraction network F and the segmentation network S. Calculate the target domain segmentation loss part for the target domain required segmentation object results obtained by passing the target domain training samples through the feature extraction network F and the segmentation network S. Calculate the domain discriminant loss part for the domain label results obtained by passing the source domain and target domain training samples through the feature network F and the domain discriminant network D. Weight and superimpose the source domain segmentation loss part, the target domain segmentation loss part, and the domain discriminant loss part as the overall loss.

[0015] Step 6: By minimizing the overall loss, iteratively optimize the parameters in the feature extraction network F, the segmentation network S, and the domain discrimination network D until the overall loss meets the requirements, completing the transfer learning process.

[0016] According to another aspect of the present invention, a method for image segmentation based on transfer learning is also proposed, comprising the following steps:

[0017] Step 1: Combining the feature extraction network F trained by the aforementioned transfer learning method with the segmentation network S into a target domain segmentation network that can segment the target domain required segmentation objects;

[0018] Step 2: Input the target domain image to be segmented into the feature extraction network F;

[0019] Step 3: restore the multi-scale feature vector obtained by the feature extraction network F into a multi-scale feature map;

[0020] Step 4: Input the multi-scale feature map into the segmentation network S to obtain the segmentation result. Remove the segmentation result of the object required to be segmented in the source domain to obtain the segmentation result of the object to be segmented in the target domain.

[0021] According to another aspect of the present invention, a transfer learning system integrating Swin Transformer and UNet for segmentation tasks is also proposed, comprising:

[0022] An input module is used to input the image data to be segmented into a feature extraction network. The feature extraction network F uses the multi-scale and multi-window attention mechanism of Swin Transformer as the backbone network to extract feature vectors of different scales. The input image to be segmented contains Nc channels.

[0023] A restoration module is used to restore feature vectors of different scales obtained by the feature extraction network into feature maps of different scales;

[0024] An object segmentation module is used to input feature maps of different scales into a segmentation network S to obtain segmentation results of the required segmented objects in the source domain and the target domain. The segmentation network S is a UNet decoding network with a spatial attention mechanism added;

[0025] A domain judgment module is used to input feature maps of different scales into a domain discrimination network, judge whether the feature map is from the source domain or the target domain, and give a corresponding label. The domain discrimination network D includes an improved UNet encoding network combined with a UNet skip link method, and a two-level fully connected layer and an output layer;

[0026] The overall loss calculation module is used to calculate the source domain segmentation loss part from the source domain required segmentation object results obtained by passing the source domain training samples through the feature extraction network F and the segmentation network S, calculate the target domain segmentation loss part from the target domain required segmentation object results obtained by passing the target domain training samples through the feature extraction network F and the segmentation network S, calculate the domain discrimination loss part from the domain label results obtained by passing the source domain and target domain training samples through the feature network F and the domain discrimination network D, and weighted sum the source domain segmentation loss part, the target domain segmentation loss part, and the domain discrimination loss part into an overall loss;

[0027] The optimization module is used to iteratively optimize the parameters in the feature extraction network F, the segmentation network S, and the domain discrimination network D by minimizing the overall loss until the overall loss meets the requirements, thus completing the transfer learning process.

[0028] According to another aspect of the present invention, an image segmentation system based on transfer learning is further proposed, including:

[0029] The target domain segmentation network combination construction module is used to combine the feature extraction network F trained by the aforementioned transfer learning method with the segmentation network S to form a target domain segmentation network that can segment the required segmentation objects in the target domain;

[0030] The input module is used to input the target domain image to be segmented into the feature extraction network F; the feature extraction network F uses the multi-scale multi-window attention mechanism of Swin Transformer as the backbone network to extract feature vectors of different scales, and the input image to be segmented contains Nc channels;

[0031] The restoration module is used to restore the multi-scale feature vectors obtained by the feature extraction network F into multi-scale feature maps;

[0032] The object segmentation module is used to input the multi-scale feature maps into the segmentation network S to obtain the segmentation results of the required segmentation objects in both the source domain and the target domain. The segmentation network S is a UNet decoding network with a spatial attention mechanism added to obtain the segmentation results;

[0033] The elimination module is used to eliminate the segmentation results of the required segmentation objects in the source domain, thereby obtaining the segmentation results of the segmentation objects in the target domain.

[0034] Beneficial effects

[0035] 1. The transfer learning method mentioned in the present invention can handle the situation where there are obvious differences in the image acquisition objects of the source domain and the target domain, and the targets to be segmented in the source domain and the target domain are different. For example, the source domain consists of a large number of labeled CT data of adult mesenteries, while the target domain only contains labeled liver data of a dozen pediatric CT images. The present invention can transfer the knowledge of the source domain data to the target domain with greater task differences, improving and enhancing the segmentation effect of the target domain.

[0036] 2. In order to ensure the common feature space (domain-invariant features) that can be extracted in both the source domain and the target domain at different scales, the method of the present invention takes different-scale features as the input of the domain discriminant network, and designs the domain discriminant network structure by referring to the different-scale feature processing methods of UNet.

[0037] 3. The present invention uses the gradient reversal layer and the domain discriminant network part to describe the feature differences, and realizes the acquisition of the common feature space by minimizing the classification loss of the domain discriminator.

[0038] 4. The source domain and the target domain completely share the network to extract domain-invariant features, and the respective segmentation results of the source domain and the target domain are generated based on the common features only until the last output layer. Therefore, the organs segmented in the source domain and the target domain can be different.

[0039] 5. The segmentation network uses Swin Transformer as the encoder to obtain features with a larger receptive field; the decoding part with a spatial attention mechanism is added. In addition to using multi-scale information, it can suppress the situation where segmentation results appear in non-interested regions.

[0040] 6. The domain discriminant network takes multi-scale features as the input and adopts a network design similar to the UNet encoder network, increasing the network complexity of the domain discriminant network and making it more sensitive.

[0041] 7. The present invention suppresses the problem that the improvement of the segmentation effect of the target domain caused by sample imbalance is limited by amplifying the target domain data, so that the source domain and target domain samples used for training are the same in each round. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1A : Flowchart of a transfer learning method integrating Swin Transformer and UNet for a segmentation task according to the present invention;

[0043] Figure 1B : Flowchart of an image segmentation method based on transfer learning according to the present invention;

[0044] Figure 2 : Schematic diagram of the network structure of the system according to the present invention;

[0045] Figure 3: Diagram of the correspondence between each vector label and the sub-block;

[0046] Figure 4 : Composition of the Swin Transformer module;

[0047] Figure 5 : Schematic diagram of the relationship between pooling and position;

[0048] Figure 6 : Schematic diagram of the sub-block fusion operation;

[0049] Figure 7 : Sub-block recovery operation;

[0050] Figure 8 : Schematic diagram of the process of restoring features at the first scale;

[0051] Figure 9 : Schematic diagram of the spatial attention module;

[0052] Figure 10A : Block diagram of the transfer learning system that fuses Swin Transformer and UNet for the segmentation task according to the present invention;

[0053] Figure 10B : Schematic diagram of an image segmentation system based on transfer learning;

[0054] Figure 11 : Schematic diagram of a case with a slightly lower Dice of the target domain child liver transfer learning model; (a) Liver, (b) UNet-liver, (c) DANN-Liver;

[0055] Figure 12 : Comparison chart of the segmentation results of the target domain child liver test data. Detailed implementation manners

[0056] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0057] According to an embodiment of the present invention, as Figure 1A shown, for the case where the source domain has a large amount of data of a certain organ of a certain part that has been labeled (the unlabeled data can be optional), and the target domain needs to have a small amount of labeled data of another organ (the unlabeled data can be optional), the present invention proposes a transfer learning method that fuses Swin Transformer and UNet for the segmentation task, including the following steps:

[0058] Step 1: Input the image data to be segmented into the feature extraction network. The feature extraction network F uses the multi-scale and multi-window attention mechanism of Swin Transformer as the backbone network to extract feature vectors of different scales. The input image to be segmented can contain Nc channels;

[0059] Step 2: Restore the feature vectors of different scales obtained by the feature extraction network to feature maps of different scales;

[0060] Step 3: Input the feature maps of different scales into the segmentation network to obtain the segmentation results of the required segmentation objects in the source domain and the target domain. The segmentation network S is a UNet decoding network with a spatial attention mechanism added;

[0061] Step 4: Input the feature maps of different scales into the domain discriminant network to determine whether the feature maps come from the source domain or the target domain and give corresponding labels. The structure of the domain discriminant network D includes an improved UNet encoding network combined with the UNet skip connection method, as well as two-level fully connected layers and an output layer;

[0062] Step 5: Calculate the source domain segmentation loss part from the source domain required segmentation object results obtained by passing the source domain training samples through the feature extraction network F and the segmentation network S. Calculate the target domain segmentation loss part from the target domain required segmentation object results obtained by passing the target domain training samples through the feature extraction network F and the segmentation network S. Calculate the domain discriminant loss part from the domain label results obtained by passing the source domain and target domain training samples through the feature extraction network F and the domain discriminant network D. Weightedly superimpose the source domain segmentation loss part, the target domain segmentation loss part, and the domain discriminant loss part into the overall loss;

[0063] Step 6: By minimizing the overall loss, iteratively optimize the parameters in the feature extraction network F, the segmentation network S, and the domain discriminant network D until the overall loss meets the requirements, and complete the transfer learning process;

[0064] Step 7: After the network training is completed, combine the feature extraction network F and the segmentation network S into a target domain segmentation network that can segment the required segmentation objects in the target domain. Input the target domain image to be segmented into the feature extraction network F according to Step 1. Restore the multi-scale feature vectors obtained by the feature extraction network F to multi-scale feature maps according to Step 2. Input the multi-scale feature maps into the segmentation network S according to Step 3 to obtain the segmentation results. Exclude the segmentation results of the required segmentation objects in the source domain to obtain the segmentation results of the target domain segmentation objects.

[0065] According to an embodiment of the present invention, in step 1, the feature extraction network F selects the multi-scale and multi-window attention mechanism of Swin Transformer as the backbone network, and uses the maximum-minimum pooling method to achieve sub-block fusion;

[0066] The segmentation network S adopts the structure of the decoder of UNet. Considering that most human anatomical structures are similar, a spatial attention module is introduced to strengthen the features of the parts that may appear in the segmented part and suppress the features of the parts that cannot appear in the part to be segmented, and finally the segmentation results of two organs in the source domain and the target domain are obtained;

[0067] The domain discriminant network part uses a gradient reversal layer and the multi-scale methods in UNet and Swin Transformer. Taking the feature information of different scales obtained by the feature extraction network as input, it judges whether the features come from the source domain or the target domain and gives corresponding labels.

[0068] Since the adjacent organ anatomical structures in medical image segmentation also have important reference significance for the organ to be segmented, in the embodiment of the present invention, the input image can be not one image, but N images adjacent to the image to be segmented. As Figure 2 shown, the feature extraction network F first uses a sub-block division module to divide each input image into sub-blocks of 4×4 size, and each sub-block data is unfolded into a one-dimensional vector, generating a total of vectors. Then, through the linear embedding module, the shared linear transformation matrix W C×16N is multiplied by the vectors obtained by sub-block division to transform all vectors into a new set of vectors with a length of C. The feature extraction network performs feature extraction at 4 different scales, and 2 consecutive Swin Transformer modules are used for calculation at each scale, where H and W are the pixel height and width of the image respectively.

[0069] According to an embodiment of the present invention, optionally, the number of times the Swin Transformer module is applied at each scale is not limited to 2 times and can be adjusted as needed.

[0070] The input of Swin Transformer is a set of vectors z l (z l is composed of M l ×N l vectors of the same length, specifically expressed as The output is also a set of vectors z l+1 (z l+1 is composed of M l ×N l vectors, and the length of each vector is the same as the input zl The medium vector lengths are the same, and the specific representation is The corresponding relationship between each vector label and the sub-block is as Figure 3 shown.

[0071] As Figure 4 shown, the schematic diagram of the Swin Transformer module. The Swin Transformer first performs vector normalization within the layer, and then each vector is processed by the multi-head self-attention module within the standard window where the vector is located (the entire image is divided into several parts). The processing result is added to the input result; then it enters the second Transformer module, which is the multi-head self-attention module within the circularly shifted window, and other components are the same as those of the first Transformer module.

[0072] In one embodiment of the present invention, a group of vectors generated by the Swin Transformer of one scale can be merged into a new vector by sub-block fusion for 4 vectors within the adjacent 2×2 spatial range. As Figure 5 shown, the rectangular frame in the figure encloses the adjacent 4 vectors that need to be sub-block fused. Without loss of generality, assume that the vector corresponding to the upper left sub-graph in each frame is represented as the upper right is represented as the lower left is represented as the lower right is represented as As Figure 6 shown, the merging process calculates the maximum pooling (preserving the maximum value at the same position among the four vectors) and the average pooling result (preserving the average value at the same position among the four vectors) of the vectors to be fused, and connects these two pooling results together to generate a new vector with a length twice that of the original input vector, while the number of output vectors is reduced to one-fourth of the number of input vectors. The sub-block merging operation here is different from the downsampling used in the classical Swin Transformer network processing. The purpose of introducing the two poolings is to more comprehensively describe the differences of the 4 feature vectors used for merging. It replaces the original sub-block fusion method of fully linking the feature vectors of the sub-blocks to be fused, thereby achieving the purpose of reducing the length of the feature vectors after sub-block fusion and reducing the number of parameters of the Swin Transformer.

[0073] According to one embodiment of the present invention, the segmentation network S draws on the structure of the decoder in the UNet network. Since the feature extraction network outputs a group of vectors, it needs to go through as Figure 7The sub-block restoration operation shown restores it into images of different scales. According to the spatial positions corresponding to the vectors, the m×n vectors are rearranged to obtain an image sequence of m×n×C. m represents the number of sub-blocks divided in the height direction of the image at the current scale, and n represents the number of sub-blocks divided in the width direction of the image at the current scale. Taking the first scale as an example, as Figure 8 shown: At this time C is the length of the current feature vector and also the number of channels of the restored feature map. Since the sub-block restoration operation does not change the total number of parameters of the original feature vector, a C×1 vector, the total number of parameters is restored into C feature maps with a size of . The letter C represents the length of the input vector and also the number of channels after being restored into the image. After the restored image sequence is concatenated with the sequence of linearly upsampled low-level scale features, through convolution, the spatial attention module, and convolution operations, the feature result of the current scale is obtained, and after upsampling, it is sent to the high-level scale for processing. The features of the last scale are linearly upsampled by 4*4, restored to the size of the input image, and after two convolution operations, finally, the final segmentation result is generated via the Sigmoid layer. The segmentation includes the segmentation result of the source domain and the segmentation result of the target domain.

[0074] Regarding the spatial attention module introduced in the segmentation network, it judges spatially which regions are related to the target to be segmented. The features of the relevant regions will be enhanced by large spatial attention weights, while the features at positions unrelated to the region to be segmented will be weakened by small spatial attention weights.

[0075] As Figure 9 shown, the spatial attention module performs max and min pooling on the features of the input K channels. At the same time, after connecting the other 2-channel features obtained by 1*1 convolution processing for estimating the spatial weights, and then through a group of 1*1 convolution processing, a single-channel spatial weight is obtained. This weight result is multiplied point by point with the features of each channel of the input to obtain the feature result weighted by spatial attention.

[0076] According to an embodiment of the present invention, the domain discrimination network D also borrows the idea of multi-scale feature processing, takes the features of different scales obtained after sub-block recovery as input, and first passes through a gradient reversal layer. When the gradient reversal layer propagates forward, the data remains unchanged when multiplied by 1; in backpropagation, it needs to be multiplied by -λ (λ is a hyperparameter used to control the influence of the domain discriminator gradient on the optimization of the feature extraction network parameters). Since the feature extraction network is to confuse the domain discrimination network so that it cannot correctly distinguish whether the features belong to the source domain or the target domain, the parameter optimization goal is to increase the loss function of the domain discrimination network, which is opposite to the goal of the domain discrimination network. Therefore, this process is called the gradient reversal layer, that is, the gradient needs to be reversed during backpropagation. The input features of the high-level scale are processed by two convolutional operations, then processed by the max pooling layer into low-level scale features, and then connected to the low-level input features. After that, they enter the convolutional operation of the next level. The features processed through four different scales form a vector of length L1 after unfolding and dynamic average pooling operations, and the domain label result is output after passing through two fully connected networks.

[0077] The input image data is divided into four types, namely source domain labeled data (where the subscript L indicates having a label; the subscript S indicates the source domain; the superscript i indicates the i-th sample, taking an integer from 1 to N SL inclusive; represents the i-th segmented image to be segmented in the source domain with a label, represents the segmentation label of the object to be segmented corresponding to the i-th segmented image to be segmented in the source domain with a label, can represent the segmentation label of one segmentation object or the segmentation labels of multiple segmentation objects, N SL refers to the number of labeled samples in the source domain data), source domain unlabeled data (where the subscript U indicates no label; represents the i-th unlabeled segmented image in the source domain, N SU refers to the number of unlabeled samples in the source domain data), target domain labeled data (the subscript T indicates the target domain; represents the i-th segmented image to be segmented in the target domain with a label, represents the segmentation label of the object to be segmented corresponding to the i-th segmented image to be segmented in the target domain with a label, can be the segmentation label of one segmentation object or the segmentation labels of multiple segmentation objects, N TL refers to the number of labeled samples in the target domain data) and target domain unlabeled data ( represents the i-th unlabeled segmented image in the target domain, N TU refers to the number of unlabeled samples in the target domain data).

[0078] To ensure that the network can extract a common feature space from the source domain and target domain data, an equal number of source domain labeled data and target domain labeled data are required to be included in each batch of data. Since the amount of target domain data is very small, it is necessary to perform an amplification process to amplify the target domain data samples to the same number as the source domain target data samples. If there is unlabeled data in the source domain and target domain, the same number of unlabeled data from the source domain and target domain are selected for each batch specifically for training the domain discriminant network; when one domain has unlabeled data and the other domain does not, the labeled data can be used to replace the unlabeled data where there is no unlabeled data available; in order to make the number of source domain and target domain samples used as unlabeled data the same each time, some amplification processes also need to be done.

[0079] All training samples in a batch can be represented as labeled samples and unlabeled samples where d S represents the domain label of the source domain features and can be set to 0, d T represents the domain label of the target domain features and can be set to 1, and the empty set Φ represents no corresponding segmentation label. Since the transfer learning segmentation network will output both the source domain segmentation result and the target domain segmentation result simultaneously, the samples need to include two sets of segmentation labels for training. The source domain labeled samples are used to calculate the source domain loss of the segmentation network, and the target domain samples are used to calculate the target domain loss in the segmentation network; all samples have domain labels, so they can all be used to calculate the loss of the domain discrimination network.

[0080] The optimization objective of the transfer learning network model of the present invention is to minimize the total loss, including the source domain segmentation loss, the target domain segmentation loss, and the domain discrimination network classification loss:

[0081]

[0082] L Total represents the total loss, L Seg (x SL ,y SL ) represents the source domain segmentation loss, L Seg (x TL ,y TL ) represents the target domain segmentation loss, L Domain represents the domain discrimination network classification loss; where the hyperparameter λ is a positive real number used to control the proportion between the domain discrimination network loss and the segmentation loss, and the value setting can refer to the magnitudes of the two parts of the segmentation loss and the domain discrimination network classification loss. Currently, λ = 0.2 is used in the experiment. Among them, the segmentation loss can adopt binary cross entropy (BCE) loss, Dice loss, focal loss, etc., or a weighted combination of multiple loss functions. The classification loss of the domain discrimination network can be calculated using BCE.

[0083]

[0084] In the segmentation loss function of the source domain, the superscript 0 indicates taking the segmentation channel result of the source domain. Calculate the BCE loss L according to the source domain sample segmentation result and the source domain label BCE , the Dice result, and the Focal Loss L Focal . w BCE , w Dice , w Focal respectively represent the weights assigned to various losses and are hyperparameters selected according to actual training. Among them, the symbol represents the segmentation network in the form of a function, and θ S represents the parameters to be optimized by the segmentation network; the symbol represents the feature extraction network in the form of a function, represents the parameters to be optimized by the feature extraction network; represents the sample x SL after passing through the feature extraction network to obtain the feature map; The feature map passes through the segmentation network to obtain the segmentation results of all objects to be segmented in the source domain and the target domain, and the results of the channels of the objects to be segmented in the source domain are taken from them.

[0085] L BCE is defined as:

[0086]

[0087] represents the label result estimated by the network P for the input sample x; the letter y represents the true label. Since it is a binary result, y takes 0 or 1; N represents the total number of samples, the symbol i represents the sample label, and L BCE is an average statistical result of the difference between the estimated labels and the true labels of all samples. The smaller L BCE , the closer the estimated label is to the true label.

[0088] L Focal is a loss to suppress the impact of the imbalance between positive and negative samples. It is an improved BCE loss:

[0089]

[0090] γ is called the focusing parameter and is a non - negative real number; α t is also a non - negative real number.

[0091] Dice is a similarity metric function between labeled graphs, and its value ranges from [0, 1]. The larger the Dice value, the more similar the two sets are:

[0092]

[0093] It represents the segmentation result estimated by the network G for the input sample x of G(x). Each pixel position in the segmentation result has a label result; the letter y represents the true labeled segmentation result, and each pixel has a true label. G(x)·y represents the result of multiplying the segmentation result and the data corresponding to the true label pixel by pixel, and the result is still a graph. || represents the result of accumulating the data after taking the modulo at each pixel position, which is a non-negative real number.

[0094] The segmentation loss function for the target domain:

[0095]

[0096] Among them, the superscript 1 represents taking the segmentation channel result of the target domain, and the meanings of other parameters are the same as those defined in the source domain segmentation loss.

[0097]

[0098] Among them, the symbol represents the domain discriminant network in the form of a function, and θ D The symbol represents the parameters to be optimized by the domain discriminant network. is the feature map after passing through the domain discriminant network to obtain the result of the domain label. The loss function of the domain discriminant network is a comprehensive L BCE classification loss of various samples including labeled data in the source domain, labeled data in the target domain, unlabeled data in the source domain, and unlabeled data in the target domain.

[0099] Since the same number of source domain and target domain data are loaded simultaneously in each batch for parameter training, it can ensure that the network can find common features in the two domains, enabling the segmentation network to obtain good segmentation results in both the source domain and the target domain.

[0100] In an embodiment of the present invention, the number of source domain and target domain samples used for training in one batch is the same, which solves the problem of sample imbalance; at the same time, the segmentation results of the source domain and the target domain are output. The source domain samples only describe the source domain segmentation loss, and the target domain samples only describe the target domain segmentation loss; the target domain segmentation organ results of the source domain samples and the source domain segmentation organ results of the target domain samples are not used to calculate the segmentation loss.

[0101] According to another embodiment of the present invention, a method for image segmentation based on transfer learning is proposed, as Figure 1B shown, including the following steps:

[0102] Step 1: Combining the feature extraction network F trained by the aforementioned transfer learning method with the segmentation network S into a target domain segmentation network that can segment the target domain required segmentation objects;

[0103] Step 2: Input the target domain image to be segmented into the feature extraction network F;

[0104] Step 3: restore the multi-scale feature vector obtained by the feature extraction network F into a multi-scale feature map;

[0105] Step 4: Input the multi-scale feature map into the segmentation network S to obtain the segmentation result. Remove the segmentation result of the object required to be segmented in the source domain to obtain the segmentation result of the object to be segmented in the target domain.

[0106] According to another embodiment of the present invention, a transfer learning system integrating Swin Transformer and UNet for segmentation tasks is also proposed. Figure 10A As shown, including:

[0107] An input module is used to input the image data to be segmented into a feature extraction network. The feature extraction network F uses the multi-scale and multi-window attention mechanism of Swin Transformer as the backbone network to extract feature vectors of different scales. The input image to be segmented contains Nc channels.

[0108] A restoration module is used to restore feature vectors of different scales obtained by the feature extraction network into feature maps of different scales;

[0109] An object segmentation module is used to input feature maps of different scales into a segmentation network S to obtain segmentation results of the required segmented objects in the source domain and the target domain. The segmentation network S is a UNet decoding network with a spatial attention mechanism added;

[0110] A domain judgment module is used to input feature maps of different scales into a domain discrimination network, judge whether the feature map is from the source domain or the target domain, and give a corresponding label. The domain discrimination network D includes an improved UNet encoding network combined with a UNet skip link method, and a two-level fully connected layer and an output layer;

[0111] The overall loss calculation module is used to calculate the source domain segmentation loss part based on the source domain required segmentation object results obtained by passing the source domain training samples through the feature extraction network F and the segmentation network S, calculate the target domain segmentation loss part based on the target domain required segmentation object results obtained by passing the target domain training samples through the feature extraction network F and the segmentation network S, calculate the domain discrimination loss part based on the domain label results obtained by passing the source domain and target domain training samples through the feature network F and the domain discrimination network D, and weightedly superimpose the source domain segmentation loss part, the target domain segmentation loss part, and the domain discrimination loss part into an overall loss;

[0112] The optimization module is used to iteratively optimize the parameters in the feature extraction network F, the segmentation network S, and the domain discrimination network D by minimizing the overall loss until the overall loss meets the requirements, thus completing the transfer learning process.

[0113] According to another embodiment of the present invention, as Figure 10B shown, a further image segmentation system based on transfer learning is also proposed, including:

[0114] The target domain segmentation network combination construction module is used to combine the feature extraction network F trained by the aforementioned transfer learning method with the segmentation network S to form a target domain segmentation network that can segment the required segmentation objects in the target domain;

[0115] The input module is used to input the target domain image to be segmented into the feature extraction network F; the feature extraction network F uses the multi-scale multi-window attention mechanism of Swin Transformer as the backbone network to extract feature vectors of different scales, and the input image to be segmented contains Nc channels;

[0116] The restoration module is used to restore the multi-scale feature vectors obtained by the feature extraction network F into multi-scale feature maps;

[0117] The object segmentation module is used to input the multi-scale feature maps into the segmentation network S to obtain the segmentation results of the required segmentation objects in both the source domain and the target domain. The segmentation network S is a UNet decoding network with a spatial attention mechanism added to obtain the segmentation results;

[0118] The elimination module is used to eliminate the segmentation results of the required segmentation objects in the source domain, and thus obtain the segmentation results of the segmentation objects in the target domain.

[0119] To verify the proposed transfer learning method that combines Transformer and UNet for medical image segmentation, the CT-labeled peritoneal cavity data of adults was used as the source domain, and the CT-labeled liver data of children was used as the target domain to experiment on the effect of transferring the knowledge of the peritoneal cavity segmentation model to the segmentation of the liver. In this experiment, no unlabeled data from the source domain or the target domain was involved in the training. The source domain data from adult peritoneal cavities came from 73 patients, generating a total of 6,990 samples. The target domain data of children's livers included 20 cases, with a total of 1,772 samples. Additionally, the liver data of another 10 children was used as the test set to analyze the segmentation effect of the transfer model. First, the children's data samples were randomly translated, rotated, and scaled, and then amplified to 7,088 samples. The same number of samples as the source domain peritoneal cavity samples was extracted for the network training of transfer learning. In each round of training, the data was divided into training samples and validation samples in a 9:1 ratio. The training samples were used to optimize the network parameters, and the validation samples were used to evaluate the performance of the current model to find the current optimal estimate. Considering the use of the anatomical structure information of the surrounding tissues, three adjacent consecutive CT images were used as the input to estimate the segmentation result of the middle layer.

[0120] In the feature extraction network, the parameters were basically the same as those of the Swin Transformer network, with the number of channels C = 96, the number of heads in the multi-head self-attention module set to 32, the window size set to 7, and the expansion layer of the multi-layer perceptron set to 4. Only when fusing the sub-blocks, two pooling methods were used and then connected. The segmentation loss of the network consisted of the average of the BCE loss and the Dice loss. A total of 4 source domain samples and 4 target domain samples were loaded in one batch, and the samples were randomly rotated, translated, and scaled within a small range. The optimizer for network training was Adam, and the initial learning rate was set to 10 -4 . Subsequently, it was updated using the Exponential decay method, and a total of 100 rounds of training were performed to obtain the final segmentation model parameters. In the segmentation network and the domain discriminator network, the convolution operation used a 3×3 window, and the probability value p of the dropout layer was 0.2.

[0121] The CPU of the hardware training environment was an Intel(R) Core(TM) i9-10900K with a main frequency of 3.7GHz, equipped with two NVIDIA TITAN RTX, and the operating system.

[0122] To comparatively analyze the improvement of transfer learning on the segmentation effect of the target domain, a classical UNet segmentation model was directly trained using the augmented liver data of the target domain as a benchmark, and the segmentation results of the UNet model and the transfer learning model were evaluated on 10 test data of the target domain. From the 3D Dice results of the liver segmentation effects of different models on the 10 test sets of the target domain shown in Table 1, the Dice of the transfer learning liver segmentation model was higher for 6 patients, and slightly lower for 4 cases. Figure 11 The segmentation effect diagram with a slightly lower 3D Dice value of the transfer learning model is given (the overall volume of the liver in this case is relatively small). In the figure, (a) is the liver, (b) is the UNet-liver, and (c) is the DANN-Liver; the result (c) of the transfer learning model is also very close to the doctor's annotation (a), but the dice result is slightly lower. From the average Dice in Table 1, the statistical result of the transfer learning model is higher than that of the target domain. Figure 12 It is observed from the large liver segmentation comparison diagram shown that the segmentation contour (+ sign marked dotted line) of the transfer learning model is closer to the solid contour manually marked by the doctor.

[0123] Table 1 Dice values of the segmentation results of different models on the target domain test set

[0124]

[0125] In summary, transfer learning can use the knowledge provided by the segmentation of different-shaped organs in the source domain to improve the segmentation effect of training new organs in the small-sample target domain.

[0126] According to an embodiment of the present invention, optionally, using semi-supervised technology, applying an existing model to unlabeled target domain data, and screening the labeled results that meet the requirements as pseudo-labels to expand the target domain training samples, and repeating such self-learning to obtain the model of the target domain. However, errors may occur during the pseudo-label screening process, resulting in a training effect inferior to that of the supervised method.

[0127] According to an embodiment of the present invention, optionally, using registration technology to map a small amount of reference annotations to unlabeled target domain data to achieve an increase in the amount of labeled data in the target domain. However, due to the registration accuracy, it is impossible to ensure the accuracy of the annotation information obtained by registration.

[0128] Although the above-described illustrative specific embodiments of the present invention have been described to facilitate the understanding of the present invention by those skilled in the art, and it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

Claims

1. A transfer learning method integrating Swin Transformer and UNet for segmentation tasks, characterized in that, It includes the following steps: Step 1: Input the image data to be segmented into the feature extraction network F. The feature extraction network F uses the multi-scale and multi-window attention mechanism of Swin Transformer as the backbone network to extract feature vectors of different scales. The input image to be segmented contains Nc channels; Step 2: Restore the feature vectors of different scales obtained by the feature extraction network F into feature maps of different scales; Step 3: Input the feature maps of different scales into the segmentation network S to obtain the segmentation results of the required segmentation objects in the source domain and the target domain. The segmentation network S is a UNet decoding network with a spatial attention mechanism added; Step 4: Input the feature maps of different scales into the domain discriminant network D to determine whether the feature maps come from the source domain or the target domain and give corresponding labels. The domain discriminant network D includes an improved UNet encoding network combined with the UNet skip connection method, as well as two-level fully connected layers and an output layer; Step 5: Calculate the source domain segmentation loss part for the source domain required segmentation object results obtained by passing the source domain training samples through the feature extraction network F and the segmentation network S. Calculate the target domain segmentation loss part for the target domain required segmentation object results obtained by passing the target domain training samples through the feature extraction network F and the segmentation network S. Calculate the domain discriminant loss part for the domain label results obtained by passing the source domain and target domain training samples through the feature extraction network F and the domain discriminant network D. Weightedly superimpose the source domain segmentation loss part, the target domain segmentation loss part, and the domain discriminant loss part into an overall loss; Step 6: By minimizing the overall loss, iteratively optimize the parameters in the feature extraction network F, the segmentation network S, and the domain discriminant network D until the overall loss meets the requirements, and complete the transfer learning process.

2. The transfer learning method integrating Swin Transformer and UNet for segmentation tasks according to claim 1, characterized in that, The feature extraction network F uses the multi-scale and multi-window attention mechanism of Swin Transformer as the backbone network. The sub-block fusion method adopted by the feature extraction network F is formed by linking the maximum pooling result and the average pooling result of the feature vectors of the sub-blocks to be fused.

3. The transfer learning method integrating Swin Transformer and UNet for segmentation tasks according to claim 1, characterized in that, Step 1 further includes: Step 1.

1. First, the feature extraction network F uses the sub-block division module to divide each input image with a size of H×W into sub-blocks of size Np×Np, where Np is the width of the sub-block. Each sub-block data is unfolded into a one-dimensional vector with a size of Np×Np×Nc; Nc represents the number of channels of the input image, and a total of vectors are generated, Ns is the number of different scales, and H and W are the pixel height and pixel width of the image; Step 1.2: Multiply the vectors obtained by sub-block partitioning with the shared linear transformation matrix W through the linear embedding module, and transform all vectors into a new set of vectors with a length of C. The feature extraction network F uses the Swin Transformer module to perform feature extraction at Ns different scales. At each scale, 2 consecutive Swin Transformer modules are used for calculation. The input of the Swin Transformer is a set of vectors, and the output is also a set of vectors. First, vector normalization is performed within the layer, and for each vector, a multi-head self-attention module is processed within the standard window where the vector is located, and the processing result is added to the input result. Then, it enters the second Swin Transformer module, which is a multi-head self-attention module that cyclically shifts the window within the window. c×Np×Np×Nc Multiply with the vectors obtained by sub-block partitioning, and transform all vectors into a new set of vectors with a length of C. The feature extraction network F uses the SwinTransformer module to perform feature extraction at Ns different scales. At each scale, 2 consecutive SwinTransformer modules are used for calculation. The input of the Swin Transformer is a set of vectors, and the output is also a set of vectors. First, perform vector normalization within the layer. For each vector, perform a multi-head self-attention module process within the standard window where the vector is located, and add the processing result to the input result. Then, enter the second Swin Transformer module, which is a multi-head self-attention module that cyclically shifts the window within the window.

4. The transfer learning method integrating Swin Transformer and UNet for segmentation tasks according to claim 3, characterized in that, Step 1.2 further includes: A group of vectors generated by Swin Transformer at one scale are merged into a new vector by sub-block fusion, merging 4 vectors within the adjacent 2×2 spatial range into one new vector. In the merging process, calculate the maximum pooling and average pooling results of these four vectors, and connect these two pooling results together to generate a new vector with a length twice that of the original input vector. At the same time, the number of output vectors is reduced to one-fourth of the number of input vectors.

5. The transfer learning method integrating Swin Transformer and UNet for segmentation tasks according to claim 1, characterized in that, The segmentation network S adopts a UNet decoding network combined with a spatial attention mechanism. In the spatial attention mechanism, after the maximum pooling, minimum pooling, and the superposition of the results of 1×1 convolution among the channels of the input feature map X, through 1×1 convolution, batch normalization, and Sigmoid activation, the spatial weight is obtained. This weight is applied to each channel of the input feature map, so that the obtained feature results can highlight the information at the spatial positions related to the object to be segmented and suppress the information at the spatial positions unrelated to the object to be segmented.

6. The transfer learning method integrating Swin Transformer and UNet for segmentation tasks according to claim 1, characterized in that, The domain discriminator network D takes feature maps of different scales as input.

7. The transfer learning method integrating Swin Transformer and UNet for segmentation tasks according to claim 1, characterized in that, In the step 1, the input image data is divided into four types, namely: source domain labeled data where the subscript L represents having a label, the subscript S represents the source domain, and the superscript i represents the i-th sample, taking an integer between 1 and N SL inclusive, denotes the i-th image to be segmented with a label in the source domain, denotes the segmentation label of the object to be segmented corresponding to the i-th image to be segmented with a label in the source domain, which can be used to represent the segmentation label of a single segmentation object or the segmentation labels of multiple segmentation objects, N SL refers to the number of labeled samples in the source domain data; source domain unlabeled data where the subscript U represents having no label, denotes the i-th unlabeled image to be segmented in the source domain, N SU refers to the number of unlabeled samples in the source domain data; target domain labeled data the subscript T represents the target domain, denotes the i-th image to be segmented with a label in the target domain, denotes the segmentation label of the object to be segmented corresponding to the i-th image to be segmented with a label in the target domain, which can be used to represent the segmentation label of a single segmentation object or the segmentation labels of multiple segmentation objects, N TL refers to the number of labeled samples in the target domain data; target domain unlabeled data where denotes the i-th unlabeled image to be segmented in the target domain, N TU refers to the number of unlabeled samples in the target domain data; Each batch of input image data needs to contain an equal number of source domain labeled data and target domain labeled data. If the amount of target domain labeled data is less than that of the source domain, the target domain data samples need to be amplified through amplification processing to be the same as the number of source domain target data samples; The number of unlabeled data in the source domain and target domain input in each batch is not required.

8. The transfer learning method integrating Swin Transformer and UNet for segmentation tasks according to claim 7, characterized in that, In step 1, if there are unlabeled data in the source domain and the target domain, the same number of unlabeled data from the source domain and the target domain are selected for each batch specifically for training the domain discriminator network D; when one domain has unlabeled data and the other domain does not, the labeled data is used to replace the unlabeled data when there is no unlabeled data; in order to make the number of source domain and target domain samples used as unlabeled data the same each time, it is also necessary to perform data augmentation on the data of one of the domains.

9. The transfer learning method integrating Swin Transformer and UNet for segmentation tasks according to claim 1, characterized in that, Step 3 specifically includes the following steps: Step 3.1: The segmentation network adopts the decoder structure in the UNet network; for a group of vectors output by the feature extraction network, a sub-block restoration operation is required to restore them into images of different scales. According to the spatial positions corresponding to the vectors, m×n vectors are rearranged to obtain an image sequence of m×n×C, where the letter C represents the length of the input vector and also represents the number of channels after restoring to an image. Step 3.2: After the restored image sequence is connected with the sequence of linearly upsampled low-level scale features, through convolution, a spatial attention module, and convolution operations, the feature results of the current scale are obtained, and after upsampling, they are sent to the high-level scale for processing. Step 3.3: The features of the last scale are linearly upsampled by 4×4 to restore to the size of the input image. After two convolution operations, finally, the final segmentation result is generated through the Sigmoid layer. The final segmentation result includes the segmentation results of the source domain and the target domain.

10. A transfer learning method that fuses Swin Transformer and UNet for segmentation tasks, characterized in that, Step 3.2 further includes: The spatial attention module introduced in the segmentation network judges which regions are related to the object to be segmented in space. The features of the related regions will be strengthened by large spatial attention weights, while the features at the positions unrelated to the region to be segmented will be weakened by small spatial attention weights; the spatial attention module performs maximum and minimum pooling on the features of the input K channels, and at the same time, after connecting the features of another 2 channels obtained by 1×1 convolution processing for estimating the spatial weight, through a group of 1×1 convolution processing, a single-channel spatial weight is obtained. The spatial weight result is multiplied point by point with the features of each channel of the input to obtain the feature result weighted by spatial attention.

11. A transfer learning method that fuses Swin Transformer and UNet for segmentation tasks according to claim 1, characterized in that, Step 4 specifically includes the following steps: Step 4.1: The domain discriminant network D takes the features of different scales obtained after the sub-blocks are restored as inputs. First, it passes through a gradient reversal layer. When the gradient reversal layer propagates forward, the data remains unchanged after being multiplied by 1. During backpropagation, it needs to be multiplied by -λ, where λ is a positive real number and is used as a hyperparameter to control the influence of the domain discriminator gradient on the optimization of the feature extraction network parameters. Step 4.2: The input features of the high-level scale are processed by two convolutional operations and then by a max pooling layer to obtain low-level scale features. These are then concatenated with the low-level input features and enter the next-level convolutional operation. The features processed at four different scales form a vector of length L1 after unfolding and dynamic average pooling operations, and the domain label result is obtained after passing through two fully connected layers and an output layer.

12. An image segmentation method based on transfer learning, characterized in that, It includes the following steps: Step 1: Combine the feature extraction network F trained by the transfer learning method of any one of the foregoing claims 1-11 with the segmentation network S to form a target domain segmentation network capable of segmenting the segmentation object required for the target domain. Step 2: Input the image to be segmented in the target domain into the feature extraction network F. Step 3: Restore the multi-scale feature vector obtained by the feature extraction network F to a multi-scale feature map. Step 4: Input the multi-scale feature map into the segmentation network S to obtain a segmentation result. Exclude the segmentation result of the segmentation object required for the source domain to obtain the segmentation result of the segmentation object in the target domain.

13. A transfer learning system that fuses Swin Transformer and UNet for segmentation tasks, characterized in that, It includes: An input module for inputting the image data to be segmented into the feature extraction network F. The feature extraction network F uses the multi-scale multi-window attention mechanism of Swin Transformer as the backbone network to extract feature vectors of different scales. The input image to be segmented contains Nc channels. A restoration module for restoring the feature vectors of different scales obtained by the feature extraction network F to feature maps of different scales. An object segmentation module for inputting the feature maps of different scales into the segmentation network S to obtain the segmentation results of the segmentation objects required for both the source domain and the target domain. The segmentation network S is a UNet decoding network with a spatial attention mechanism added. A domain judgment module for inputting the feature maps of different scales into the domain discriminant network D to determine whether the feature maps come from the source domain or the target domain and give corresponding labels. The domain discriminant network D includes an improved UNet encoding network combined with the UNet skip connection method, as well as two fully connected layers and an output layer. An overall loss calculation module for calculating the source domain segmentation loss part from the source domain required segmentation object result obtained by passing the source domain training samples through the feature extraction network F and the segmentation network S, calculating the target domain segmentation loss part from the target domain required segmentation object result obtained by passing the target domain training samples through the feature extraction network F and the segmentation network S, calculating the domain discriminant loss part from the domain label result obtained by passing the source domain and target domain training samples through the feature extraction network F and the domain discriminant network D, and weighted adding the source domain segmentation loss part, the target domain segmentation loss part, and the domain discriminant loss part to obtain the overall loss. The optimization module is used to iteratively optimize the parameters in the feature extraction network F, the segmentation network S, and the domain discriminant network D by minimizing the overall loss until the overall loss meets the requirements, thus completing the transfer learning process.

14. An image segmentation system based on transfer learning, characterized in that, It includes: The target domain segmentation network combination construction module is used to combine the feature extraction network F and the segmentation network S trained by the transfer learning method of any one of the foregoing claims 1-11 into a target domain segmentation network that can segment the required segmentation objects in the target domain; The input module is used to input the image to be segmented in the target domain into the feature extraction network F; the feature extraction network F uses the multi-scale and multi-window attention mechanism of Swin Transformer as the backbone network to extract feature vectors of different scales, and the input image to be segmented contains Nc channels; The restoration module is used to restore the multi-scale feature vectors obtained by the feature extraction network F into multi-scale feature maps; The object segmentation module is used to input the multi-scale feature maps into the segmentation network S to obtain the segmentation results of the required segmentation objects in the source domain and the target domain. The segmentation network S is a UNet decoding network with a spatial attention mechanism added; The elimination module is used to eliminate the segmentation results of the required segmentation objects in the source domain, and thus the segmentation results of the target domain segmentation objects can be obtained.

Citation Information

Patent Citations

  • Label-free pancreatic image automatic segmentation system based on adversarial learning

    CN113870258A