An unsupervised domain adaptation medical image segmentation method based on style consistency

By using the SCUDA network and style consistency estimation method, the problems of large data requirements and poor generalization effect in pancreatic segmentation algorithms are solved, achieving more efficient pancreatic region segmentation and improving segmentation accuracy and completeness.

CN119273906BActive Publication Date: 2025-12-19SUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410628147.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-21
Publication Date
2025-12-19
Estimated Expiration
2044-05-21

AI Technical Summary

Technical Problem

Existing pancreas segmentation algorithms require a large amount of labeled data and have poor generalization performance between NCECT and CECT images, especially when identifying pancreas images with tissue adhesions or strong individual specificity.

Method used

An unsupervised adaptive medical image segmentation method based on style consistency is adopted. By constructing a SCUDA network, the first sub-network is used for image reconstruction and the second sub-network is used for segmentation. Combined with a phase consistency discriminator, style inconsistency estimation and a multi-stage style consistency module, the domain offset is reduced and the integrity of the target organ is improved.

Benefits of technology

The Recall, Precision, Jac, and DSC metrics for pancreatic segmentation were improved, enhancing the segmentation accuracy and integrity of the pancreatic region and resolving issues of inconsistent image style and texture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119273906B_ABST
    Figure CN119273906B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on style consistency unsupervised domain self-adapting medical image segmentation method and system, it is related to image processing technical field, including acquisition contains multiple source domain and target domain CT abdominal organ scan image construction original training set;SCUDA segmentation network is constructed, a local phase enhancement style fusion strategy is designed through first subnetwork to carry out image conversion and image reconstruction, find out inconsistent mapping in the second subnetwork through different style intermediate target domain synthetic image, and quantitative analysis pancreas area in the domain difference between target domain synthetic image and target domain real image, to further improve the integrity of pancreas area;And style consistency entropy is defined on target domain image, so as to focus more on inconsistent area, so as to further optimize the integrity of pancreas area.Two subnetworks are trained simultaneously to update parameters to obtain segmented image.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to an unsupervised domain adaptive medical image segmentation method and system based on style consistency. BACKGROUND

[0002] Chronic pancreatitis (CP) is a persistent inflammation of the pancreas that leads to fibrosis and permanent structural damage of the ductal stricture, and is often accompanied by the formation of stones, which can eventually lead to pancreatic fibrosis, calcification, and even pancreatic cancer, severely affecting the patient's quality of life and health status. Computed tomography imaging is currently the most important method for the examination of pancreatic diseases. Non-contrast-enhanced (NCECT) scan images are helpful in identifying the structure and density of the pancreas, and high-density lesions (calcification and pancreatic duct stones) in the pancreas appear as high-density shadows on NCE images, which form a clear contrast with the surrounding tissue, facilitating observation. Contrast-enhanced (CECT) scan images, on the other hand, can better observe the blood vessels and other tissue structures around the pancreas. Therefore, the combination of NCECT and CECT dual-phase diagnosis of the pancreas and its high-density lesions can provide important anatomical and pathological information for doctors. Traditional pancreatic segmentation algorithms usually require a large amount of labeled data for training. However, the pancreas of chronic pancreatitis patients has a very small proportion in CT image sequences, is strongly individual-specific in shape, and is adhered to surrounding tissue organs. Moreover, NCECT and CECT images have a non-uniform and non-alignable style and texture, so the generalization effect of the segmentation network between NCECT and CECT images is poor. Since the annotation task of multiple phases requires a lot of time and effort from doctors, unsupervised domain adaptation (UDA) segmentation from a labeled domain (source domain) to an unlabeled domain (target domain) has become an urgent problem to be solved.

[0003] Many studies on pattern recognition, image processing, and computer vision have focused on transferring images from one domain to another. These methods usually require a set of paired images that are identical in content and structure but different in style, assuming that there is a certain underlying relationship between the paired images of different styles. By learning these relationships, style transfer of images can be achieved. However, obtaining paired images is a challenging task and sometimes even impractical. Recently, inspired by generative adversarial networks (GAN) and cycle-consistent adversarial networks (CycleGAN), many studies have implemented unsupervised domain adaptation (UDA) in medical image analysis by aligning image appearances or latent features, without the need for paired data. SUMMARY

[0004] In view of the above-mentioned problems, the present application is proposed.

[0005] Therefore, the technical problem solved by this invention is that existing pancreas segmentation algorithms require a large amount of data, have poor generalization effect, and have problems in identifying pancreas images with tissue adhesion or high specificity.

[0006] To address the aforementioned technical problems, this invention provides the following technical solution: an unsupervised domain adaptive medical image segmentation method based on style consistency, comprising constructing a SCUDA network, performing image reconstruction through a first sub-network, and performing image segmentation through a second sub-network. The target domain image is then segmented. and source domain image The input is fed into the first sub-network to generate the target domain synthetic image. Source domain reconstructed image Source domain synthesized image Target domain reconstructed image . Target domain image Corresponding content features Source domain image Corresponding content features Image synthesized with target domain Corresponding content features The input is fed into the second sub-network to generate a segmentation prediction map. , and In the first sub-network, a phase consistency discriminator is used to distinguish the phase consistency features corresponding to source domain content features and target domain content features, thereby enhancing the domain-invariant content feature encoder. Source domain style encoder Style encoder for target domain The decoupling and removal of domain-specific features from the domain-invariant content feature encoder are performed. In the second sub-network, style inconsistency estimation is used to segment and predict target domain synthesized images with different styles, generating inconsistency mapping regions. Misclassified regions are measured to mitigate the domain offset between the synthesized target domain image and the true target domain image. Style consistency entropy, segmentation consistency loss, and segmentation loss are used in the second sub-network. Both sub-networks are trained simultaneously to update parameters and obtain the segmented image.

[0007] As a preferred embodiment of the style consistency-based unsupervised domain adaptive medical image segmentation method described in this invention, the first sub-network includes a domain-invariant content feature encoder. Style encoder for the target domain Source domain style encoder Image reconstruction decoder Local Phase Enhancement Style Fusion Module (LPSF), Phase Enhancement Style Fusion Module (PSF), and a discriminator for synthesized images from the source domain. Discriminator for synthesized images in the target domain For phase consistency discriminators .

[0008] The second sub-network includes a multi-stage style consistency module (MSSC) and an image segmenter (G).

[0009] It is the backbone network of ResNet34. and Adopt and With the same architecture, G uses a network with 5 convolutional layers. The first two convolutional layers consist of residual blocks, and each of the last three layers consists of a transposed convolution, a BN normalization layer, and a ReLU activation layer. It adopts the same architecture as G.

[0010] The discriminator uses a network with four convolutional layers and one fully connected layer. Each convolutional layer is followed by an IN normalization layer and a LeakyReLU activation layer.

[0011] As a preferred embodiment of the style consistency-based unsupervised adaptive medical image segmentation method of the present invention, wherein: the image reconstruction through the first sub-network includes reconstructing the source domain image in the first sub-network. and target domain image They were respectively sent to the domain-invariant content feature encoder Get from and At the same time, they were sent into and Extracting style features from the source domain and the style characteristics of the target domain ,Will and The LPSF module is used to generate local phase-enhanced content features in the source domain. Fusion features with local phase enhancement style of the target domain .

[0012] LPSF consists of the Local Content Feature Extraction (LCFE) submodule and the Phase Enhancement Style Fusion (PSF) module, which will... , and source domain tags Combined input into LCFE, utilizing from Extracting local vectors , is represented as:

[0013] ;

[0014] in, is the number of target vectors, is a reshape operation, is the number of channels of , and then is input into a pooling layer, followed by three convolution operations to obtain three dynamic convolution kernel parameters , and , the three dynamic convolution kernel parameters respectively perform convolution operations on , and then the features are unified to the same size as and spliced in the channel dimension to obtain the local content features of the source domain .

[0015] The LPSF module extracts the local content features of the source domain corresponding phase spectrum , and the PSF module enhances the features of the source domain image by adding the phase spectrum to , which is represented as:

[0016] ;

[0017] wherein, is the inverse fast Fourier transform.

[0018] The FFT operation is performed on the style features of the target domain , to extract the amplitude spectrum and the phase spectrum of , the target domain amplitude spectrum is low-pass filtered to remove the target domain noise, to obtain the pure target domain style, and the local phase enhanced style fusion features of the target domain are represented as:

[0019] ;

[0020] wherein, is a dot product operation, is a low-pass filter mask, indicates that convolution, average pooling and full connection operations are first performed between two inputs.

[0021] As a preferred scheme of the unsupervised domain adaptive medical image segmentation method based on style consistency according to the present application, wherein: the image reconstruction through the first subnetwork further includes generating a target domain synthetic image having the same structural content as the source domain image but the same style as the target domain image , and input decoder , simultaneously embedding into each layer of the decoder , generating target domain synthetic images with adaptive instance normalization , denoted as:

[0022] ;

[0023] wherein denotes a variance operation, denotes a mean operation, is an intermediate output feature obtained by an up-sampling operation in the decoder .

[0024] adopting an FFT operation on the style feature of the source domain , extracting amplitude spectrum and phase spectrum of , performing low-pass filtering on the source domain amplitude spectrum to remove source domain noise to obtain a pure source domain style, and a local phase enhanced style fusion feature of the source domain denoted as:

[0025] ;

[0026] inputting and into the LPSF module to generate a source domain reconstructed image , denoted as:

[0027] ;

[0028] wherein is the local phase enhanced style fusion feature of the source domain.

[0029] inputting and into the PSF module to generate a target domain phase enhanced content feature and a source domain phase enhanced style fusion feature , denoted as:

[0030] ;

[0031] ;

[0032] wherein is a phase spectrum obtained by performing FFT on , and are respectively obtained by performing FFT on The amplitude spectrum and the phase spectrum obtained by performing the FFT.

[0033] Source domain synthetic image Through the decoder Using AdaIN generation, input the decoder , while embedding into each layer of the decoder , to generate a source domain synthetic image with adaptive instance normalization , denoted as:

[0034] ;

[0035] Wherein, is the intermediate output feature obtained by the upsampling operation in the decoder .

[0036] Input and into the PSF module to generate target domain phase-enhanced content features and target domain phase-enhanced style fusion features , denoted as:

[0037] ;

[0038] ;

[0039] Target domain reconstructed image Through the decoder Using AdaIN generation, input the decoder , while embedding into each layer of the decoder , to generate a target domain reconstructed image with adaptive instance normalization , denoted as:

[0040] ;

[0041] Wherein, is the target domain phase-enhanced style fusion feature.

[0042] As a preferred scheme of the unsupervised domain adaptive medical image segmentation method based on style consistency consistency according to the application, wherein: a phase consistency discriminator is constructed to distinguish the phase consistency features corresponding to the source domain content features and the target domain content features, and the loss function corresponding to the discriminator during training is as follows:

[0043] ; ​​

[0044] wherein, denotes expectation, a discriminator for the target domain synthetic image is introduced Adversarial learning is performed, and a loss function corresponding to the discriminator during training is represented as:

[0045] ;

[0046] The source domain synthetic image and the target domain reconstructed image that do not lose domain-invariant information are generated, and a discriminator for the source domain synthetic image is introduced Adversarial learning is performed, and a loss function corresponding to the discriminator during training is represented as:

[0047] ;

[0048] A double reconstruction consistency loss is constructed and represented as:

[0049] .

[0050] As a preferred scheme of the unsupervised domain adaptive medical image segmentation method based on style consistency provided by the application, wherein: the second sub-network obtains an inconsistency map from the inconsistent features corresponding to the content features of the target domain synthetic image of different styles and the content features corresponding to the target domain real image, measures the difficult area and performs constraint, and the high-level features obtained by inputting the target domain synthetic image at different training stages into the encoder are represented as: wherein, denotes the number of epochs at the training stage, and in the training of the style conversion network, the synthetic target domain image at the first epoch is selected and represented as: wherein, is the total number of selected epochs, the style consistency in the style conversion process is measured by the multi-stage style consistency module, the MSSC module is composed of segmentation heads , and is used to generate the prediction probability of the organ of interest in different styles An inconsistency mapping map is obtained from the prediction probabilities in different styles and represented as:

[0051] ;

[0052] wherein, is the first class probability in , and is the first class probability in .

[0053] By utilizing style inconsistency, a multi-stage style consistency loss is defined The head, tail and edge of the organ of interest and the adjacent organ tissue region are focused on, and are represented as:

[0054] ;

[0055] Wherein, is the predicted probability of the i-th pixel of the j-th class in the source domain image , is the real label of the i-th pixel of the j-th class in the source domain image , is the real label of the i-th pixel of the j-th class in the source domain image , is the inconsistency map of the i-th pixel of the j-th class in the source domain image .

[0056] As a preferred scheme of the unsupervised domain adaptive medical image segmentation method based on style consistency provided by the application, wherein: the two sub-networks are trained and updated parameters to obtain the segmented image includes sending the target domain image into the pre-trained model under different styles to generate the predicted probability of the organ of interest , and generating the corresponding prediction result , if the foreground values of the same pixels in the i-th prediction result are the same, it is considered as a reliable prediction result .

[0057] The target domain image is sent into the domain-invariant content feature encoder which is being updated , and the generated is sent into the target domain image segmentation head which is being updated , to generate the predicted probability of the organ of interest and the prediction result , by performing exclusive or operation on the two prediction results and , the inconsistency map of the target domain image is obtained , the style consistency entropy loss is defined, and is represented as:

[0058] ;

[0059] Wherein, is the i-th class of the inconsistency map , is the i-th class of the inconsistency map​​​​​​ a pixel, is a predicted probability map of an organ of interest a pixel in the i-th class of the j-th image, a pixel in the i-th class of the j-th image, a pixel in the i-th class of the j-th image, is set to 1e-7.

[0060] reliable consistency region is obtained The source domain image and the target domain image are jointly trained, and the target domain image is combined with data enhancement and Dropout to input the target segmentation head to generate two predicted probability maps of the organ of interest and In the training process, the output segmentation consistency loss for the target domain image is represented as:

[0061] ;

[0062] The unlabeled target domain image is trained by a segmentation loss , represented as:

[0063] ;

[0064] wherein, is a pixel in the i-th class of the j-th image, a pixel in the i-th class of the j-th image, a pixel in the i-th class of the j-th image, a pixel in the i-th class of the j-th image, a pixel in the i-th class of the j-th image, a pixel in the i-th class of the j-th image, a pixel in the i-th class of the j-th image, a pixel in the i-th class of the j-th image. a pixel in the i-th class of the j-th image, a pixel in the i-th class of the j-th image, a pixel in the i-th class of the j-th image, a pixel in the i-th class of the j-th image. a pixel in the i-th class of the j-th image, a pixel in the i-th class of the j-th image. a pixel in the i-th class of the j-th image, a pixel in the i-th class of the j-th image.

[0065] A segmentation loss is defined, represented as:

[0066] ;

[0067] wherein, is a pixel in the i-th class of the j-th image, a pixel in the i-th class of the j-th image. a pixel in the i-th class of the j-th image, a pixel in the i-th class of the j-th image. a pixel in the i-th class of the j-th image, a pixel in the i-th class of the j-th image. a pixel in the i-th class of the j-th image, a pixel in the i-th class of the j-th image.

[0068] ​The overall style consistency segmentation loss is represented as:

[0069] ;

[0070] Segmentation of the target organ is performed by optimizing .

[0071] Another object of the present application is to provide an unsupervised domain adaptive medical image segmentation system based on style consistency, which can segment and predict the target domain synthetic image of different styles through a style inconsistency estimation method and generate an inconsistency mapping area to measure the easy-to-misclassify area, thereby reducing the domain shift between the target domain synthetic image and the real target domain image and improving the integrity of the target organ, and solving the problem of the current traditional pancreatic segmentation algorithm containing NCECT and CECT images which cannot be aligned and are uneven in style and texture.

[0072] As a preferred scheme of the unsupervised domain adaptive medical image segmentation system based on style consistency, the application comprises a data set construction module, an SCUDA segmentation network optimization module, and a parameter updating module. The data set construction module is used to collect CT abdominal organ scan images of multiple source domains and target domains to construct a primary training set. The SCUDA segmentation network optimization module is used to construct an SCUDA segmentation network, reconstruct images through a first subnetwork, and generate an inconsistency mapping area through a prediction image of different segmentation target domain images in a second subnetwork. The parameter updating module is used to simultaneously train and update parameters of the two subnetworks to obtain segmented images.

[0073] A computer device comprises a memory and a processor, the memory stores a computer program, and the processor executes the computer program to implement the steps of the unsupervised domain adaptive medical image segmentation method based on style consistency.

[0074] A computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the unsupervised domain adaptive medical image segmentation method based on style consistency.

[0075] The application provides a style consistency-based unsupervised domain adaptive medical image segmentation method and a style inconsistency estimation method. BRIEF DESCRIPTION OF DRAWINGS

[0076] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.

[0077] Figure 1 A whole flowchart of the style consistency-based unsupervised domain adaptive medical image segmentation method provided by the first embodiment of the application.

[0078] Figure 2 Another flowchart of the style consistency-based unsupervised domain adaptive medical image segmentation method provided by the first embodiment of the application.

[0079] Figure 3 A local phase enhancement style fusion (LPSF) module schematic diagram of the style consistency-based unsupervised domain adaptive medical image segmentation method provided by the first embodiment of the application.

[0080] Figure 4 A phase enhancement style fusion (PSF) module schematic diagram of the style consistency-based unsupervised domain adaptive medical image segmentation method provided by the first embodiment of the application.

[0081] Figure 5The target domain synthetic image of different training stages of a style-consistency-based unsupervised domain adaptive medical image segmentation method provided by the first embodiment of the present application.

[0082] Figure 6 The overall flowchart of a style-consistency-based unsupervised domain adaptive medical image segmentation system provided by the third embodiment of the present application. DETAILED DESCRIPTION

[0083] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the protection scope of the present application.

[0084] Embodiment 1

[0085] Reference Figures 1-5 For an embodiment of the present application, a style-consistency-based unsupervised domain adaptive medical image segmentation method is provided, comprising:

[0086] S1: Collect CT abdominal organ scan images containing multiple source domains and target domains to construct an original training set.

[0087] Further, with reference to Figure 1The figure is a whole flow chart of the present application. In this embodiment, the data collected by the present application contains 127 double-phase 3D abdominal CT images of patients with chronic pancreatitis. The axial size of each non-contrast enhanced (NCE) and contrast enhanced (CE, venous phase) CT image is 512x512, and the double-phase CT images of each patient are taken at different time points and in different scanning modes. In order to maintain the consistency of the image size in the two data sets, all the images are uniformly adjusted to 256x256. Due to the different acquisition methods of NCE CT and CECT images and the individual differences of patients, the number of cross-sectional images ranges from 35 to 120. The sagittal and coronal intervals of NCE CT and CECT images are [0.51, 0.79] mm, and the axial plane interval is 1 mm. A total of 9865 two-dimensional (2D) axial images of NCE CT and 9978 CECT images are obtained. Since it is difficult to label the high-density lesions in the pancreas (stones and calcifications in the pancreas) in the CECT image, only the labeling information of the pancreas in the CECT image is available, while the labeling of the pancreas and its high-density lesion area in the NCECT image is available. All the pancreas and high-density lesion areas are manually labeled by professional doctors, and the pixels in the pancreas area are labeled as 1, the pixels in the high-density lesion area are labeled as 2, and the other areas are labeled as 0. In order to prove that the method proposed in this chapter is more fair and reliable, a 2-fold cross-validation scheme is adopted, and there are 65 and 62 patients in each case number. The data of the ith (i=1, 2) fold are used as the evaluation set, and the remaining fold data are used as the training set. Two models are trained. The evaluation indexes of the two models on the respective evaluation sets are averaged as the final results. In terms of data preprocessing, first, the original intensity values of the CT images are truncated to [-100, 240] or [900, 1300] to enhance the details of the pancreas and its high-density lesion area. Second, all the data images are normalized to reduce the data differences caused by the medical image acquisition process and obtain a relatively uniform data set. Finally, due to the diversity of the location of the pancreas and its high-density lesions, data augmentation such as random flipping, random noise, and random rotation within [-30, 30] degrees is performed on the images before training to avoid model overfitting.

[0088] It should be noted that the present application uses the pytorch deep learning framework, and the torch version is 1.13.0. All training and verification processes are completed on an NVIDIA GeForce RTX 3090 graphics card with a 24G size of display memory. During the training process, the neural network reads data in a small batch method, and the batch size is set to 2. The experiments all use the Adaptive Moment Estimation (Adam) optimization algorithm, in which , , the initial learning rate is set to , the training period epochs is set to 120.

[0089] It should also be pointed out that the first subnetwork includes a domain-invariant content feature encoder , a style encoder of the target domain , a style encoder of the source domain , an image reconstruction decoder , a local phase enhanced style fusion module LPSF, a phase enhanced style fusion module PSF, a discriminator for the synthesized image of the source domain , a discriminator for the synthesized image of the target domain , a phase consistency discriminator .

[0090] The second subnetwork includes a multi-stage style consistency module MSSC and an image segmentor G,

[0091] is a backbone network of ResNet34, and adopt the same architecture, G adopts a network with 5 convolutional layers, the first two layers of convolution are composed of residual blocks, and each of the last three layers is composed of a transposed convolution, a BN normalization layer and a ReLU activation layer, adopt the same architecture as G; The discriminator adopts a network with four convolutional layers and one fully connected layer, and after convolution in each convolutional layer, an IN normalization layer and a LeakyReLU activation layer are performed.

[0092] S2: Construct the SCUDA segmentation network, reconstruct the image through the first subnetwork, and generate inconsistent mapping areas through the prediction map of different segmentation target domain images in the second subnetwork.

[0093] Further, the SCUDA segmentation network is constructed, including 3 encoders (a domain-invariant content feature encoder

[0094] , a style encoder of the source domain and a style encoder of the target domain ), 1 style conversion decoder , 1 local phase enhanced style fusion (LPSF) module, 1 phase enhanced style fusion (PSF) module, 2 discriminators for the synthesized images of the source domain and the target domain, a phase consistency discriminator , and a pancreas segmentation head G.

[0095] ​It should be noted that the image reconstruction by the first subnetwork includes sending the source domain image and the target domain image to the domain-invariant content feature encoder and the style feature extractor and respectively, and sending and to the style feature extractor and the style feature extractor respectively, and sending and to the LPSF module to generate the local phase enhanced content feature of the source domain and the local phase enhanced style fusion feature of the target domain .

[0096] The LPSF consists of a local content feature extraction sub-module LCFE and a phase enhanced style fusion module PSF, and and the label of the source domain are jointly input into the LCFE, and is extracted from to obtain a local vector , which is expressed as:

[0097] ;

[0098] wherein, is the number of target vectors, is a reshape operation, is the number of channels of , then is input into a pooling layer, followed by three convolution operations to obtain three dynamic convolution kernel parameters and , the three dynamic convolution kernel parameters are respectively convolved with , and then the features are unified to the same size as and spliced in the channel dimension to obtain the local content feature of the source domain .

[0099] The LPSF module extracts the local content feature of the source domain corresponding phase spectrum , and the PSF module enhances the features of the source domain image by adding the phase spectrum to , which is expressed as:

[0100] ;

[0101] wherein, It is the inverse fast Fourier transform.

[0102] Style features of the target domain Perform FFT operation to extract amplitude spectrum and phase spectrum For the target domain amplitude spectrum Low-pass filtering is performed to remove noise in the target domain, resulting in a clean target domain style, and local phase enhancement style fusion features in the target domain are also applied. Represented as:

[0103] ;

[0104] in, It's a dot product operation. It is the mask for a low-pass filter. This indicates that convolution, average pooling, and full connection operations are performed between the two inputs.

[0105] It should also be noted that image reconstruction through the first sub-network also includes generating an image with characteristics similar to the source domain image. Same structural content but with the target domain image Synthetic images of the same target domain ,Will Input Decoder At the same time Embedded into decoder In each layer, a target domain synthetic image with adaptive instance normalization is generated. , represented as:

[0106] ;

[0107] in, This indicates variance operations. Indicates average operation. In the decoder The intermediate output features are obtained by performing upsampling operations.

[0108] Stylistic features of the source domain Perform FFT operation to extract amplitude spectrum and phase spectrum For the source domain amplitude spectrum Low-pass filtering is performed to remove source domain noise, resulting in a clean source domain style, and local phase enhancement style fusion features in the source domain are also employed. Represented as: ;

[0109] Will and The input is fed into the LPSF module to generate the source domain reconstructed image. , is represented as:

[0110] ;

[0111] in, It is a local phase enhancement style fusion feature of the source domain.

[0112] Will and The input is fed into the PSF module to generate target domain phase-enhanced content features. Source domain phase enhancement style fusion features , is represented as:

[0113] ;

[0114] in, Through the The phase spectrum obtained by performing FFT and They are respectively through the The amplitude spectrum and phase spectrum obtained by performing FFT.

[0115] Source domain synthesized image via decoder Generate using AdaIN, Input Decoder At the same time Embedded into decoder In each layer, source domain synthesized images with adaptive instance normalization are generated. , is represented as:

[0116] ;

[0117] in, In the decoder The intermediate output features are obtained by performing upsampling operations.

[0118] Will and The input is fed into the PSF module to generate target domain phase-enhanced content features. Phase-enhanced style fusion features with target domain , is represented as:

[0119] ;

[0120] Target domain reconstructed image via decoder Generate using AdaIN, Input Decoder At the same time Embedded into decoder In each layer, a target domain reconstructed image with adaptive instance normalization is generated. , represented as: ;

[0121] in, It is a phase-enhanced style fusion feature in the target domain.

[0122] Furthermore, a phase consistency discriminant is constructed. To distinguish between phase consistency features corresponding to source domain content features and target domain content features, the loss function for the discriminator during training is as follows:

[0123] ;

[0124] in, To express the expectation, a discriminator for the synthesized image in the target domain is introduced. In adversarial learning, the loss function for the discriminator during training is expressed as:

[0125] ;

[0126] Generate source-domain synthesized images and target-domain reconstructed images without losing domain-invariant information, and introduce a discriminator for the source-domain synthesized image. In adversarial learning, the loss function for the discriminator during training is expressed as:

[0127] ;

[0128] The double reconstruction consistency loss is constructed and expressed as:

[0129] .

[0130] It should be noted that, due to the source domain image and target image The significant style differences between the target organ and surrounding tissues, as well as the contrast differences between the target organ and surrounding tissues, make it difficult for style transfer networks to quickly and accurately transfer the style of the target domain image to the target organ. This results in different target domain synthesized images at different training epochs. and the target domain real image Style differences still exist between them. Therefore, training using only target domain synthetic images generated in a specific period without adding additional constraints may result in the loss of certain parts of the target organ. To address this issue, an inconsistency estimation method is proposed to synthesize images from target domains with different styles. Inconsistency maps are obtained from the real images of the target domain to measure difficult regions, and then they are constrained to reduce the difference between the synthetic images of the target domain and the real images of the target domain, thereby improving the integrity of the target organ.

[0131] Furthermore, in the second sub-network, inconsistency maps are obtained from the content features corresponding to the target domain synthesized images of different styles and the content features corresponding to the target domain real images. These maps measure and constrain difficult regions. The target domain synthesized images from different training periods are then input into the encoder. The high-level features obtained are ,in This represents the number of epochs during training. In the training of the style transfer network, the epoch number is selected. Synthesized target domain images for each epoch, and represented as... ,in The total number of selected epochs is used to measure style consistency during style transfer through the multi-stage style consistency module. The MSSC module consists of... Each segment head Composition, Use Predicted probabilities of generating different styles of organs of interest An inconsistency mapping is obtained from the prediction probabilities of multiple different styles. Represented as:

[0132] ;

[0133] in, yes The first in Class probability, yes The first in Class probability.

[0134] By leveraging style inconsistency, a multi-stage style consistency loss is defined. Focusing on the head, tail, and edges of an organ, as well as adjacent organ tissue regions, is represented as follows:

[0135] ;

[0136] in, yes The Middle The class of The predicted probability of each pixel. It is a source domain image The true label, yes The Middle The class of The real label of each pixel yes The Middle The class of Inconsistency mapping of pixels.

[0137] S3: Two subnetworks are trained and updated simultaneously to obtain the segmented image.

[0138] Furthermore, due to the significant differences between different domains, there is always a certain gap between the intermediate synthesized target domain images and the real target domain images at different training periods. In certain regions of the organ of interest (ROI) in the source domain image, these regions can be easily transferred to the target domain and are therefore considered consistent regions, while other regions are difficult to transfer and are considered inconsistent regions. Therefore, the target domain image can be divided into consistent and inconsistent regions to further improve the integrity of the ROI using the target segmentation head.

[0139] It should be noted that the two sub-networks are trained and updated simultaneously to obtain the segmented image, including the target domain image. Generate predicted probabilities of organs of interest using segmentation heads of different styles from pre-trained models. And generate the corresponding prediction results. ,like If the foreground values ​​at the same pixel are the same in all prediction results, then the prediction result is considered reliable. .

[0140] target domain image Input into the constantly updated domain-invariant content feature encoder , generated The image is fed into the target domain segmentation head that is being updated. The predicted probability of generating the organ of interest. and prediction results By analyzing the two prediction results and Performing an XOR operation yields an inconsistent mapping of the target domain image. Define style consistency entropy loss as follows:

[0141] ;

[0142] in, It is an inconsistent mapping The Middle The first category 1 pixel, This is a predicted probability map of the organ of interest. The Middle The first category 1 pixel, The value is set to 1e-7.

[0143] Utilizing the obtained reliable consistency region Joint training is performed on source and target domain images. A combination of data augmentation and Dropout is used to enhance the target domain image... Input the target segmentation head and generate two predicted probability maps of the organ of interest. and During training, the segmentation consistency loss for the target domain image is expressed as:

[0144] ;

[0145] Training unlabeled target domain images using segmentation loss , represented as:

[0146] ;

[0147] in, yes The Middle The class of 1 pixel, for The first in The class of 1 pixel, for The first in The class of 1 pixel, for The first in The class of 1 pixel.

[0148] Define a segmentation loss , represented as:

[0149] ;

[0150] in, for The first in The class of 1 pixel, for The first in The class of 1 pixel.

[0151] The overall style consistency segmentation loss is expressed as:

[0152] ;

[0153] Through optimization To perform segmentation of the target organ.

[0154] Embodiment 2

[0155] In one embodiment of the present application, an unsupervised domain adaptation medical image segmentation method based on style consistency is provided. In order to verify the beneficial effects of the present application, economic benefit calculation and simulation experiments are used for scientific demonstration.

[0156] Firstly, in order to quantitatively evaluate the performance of the method proposed in the present application, the present application uses four evaluation indexes commonly used in the field of image segmentation, namely DSC coefficient (Dice-similarity coefficient, DSC), Jac coefficient (Jaccard similarity coefficient), precision (Precision) and recall (Recall), to measure the segmentation performance of each algorithm on the pancreas image. The calculation method of each index is as follows:

[0157] ;

[0158] TP (True Positive) represents the number of correctly predicted positive pixels, FP (False Positive) represents the number of incorrectly predicted positive pixels, and FN (False Negative) represents the number of incorrectly predicted negative pixels. DSC is an important evaluation index in the image segmentation task, which is used to measure the similarity between the prediction result and the standard result. The value close to 1 indicates high similarity. The Jac index represents the ratio of the intersection and union of the prediction result and the standard result. The precision measures the proportion of correctly predicted pixels in the pixels predicted as positive by the network, and the recall refers to the proportion of positive pixels in the standard that are correctly predicted by the network.

[0159] In order to verify the effectiveness of the proposed unsupervised domain adaptation medical image segmentation algorithm based on style consistency, the present application conducts an ablation experiment on the local phase enhancement style fusion method (LPSF), the phase consistency (PC) method, the multi-stage style inconsistency estimation (MSSC) method and the style consistency entropy (SCE) method. The framework without LPSF, PC, MSSC and SCE methods is taken as Baseline. Net3-1 is the result obtained by adding LPSF method based on Baseline, Net3-2 adds phase consistency method based on Net3-1, Net3-3 increases multi-stage style inconsistency estimation loss based on Net3-2, and Net3-4 introduces style consistency entropy loss at the end.

[0160] Take NCECT image as the source domain and CECT image as the target domain. After ablation experiment of the method of the application, Table 1 is obtained, wherein the LPSF method of Net3-1 aims to alleviate domain shift and produce locally enhanced target organs, and compared with Baseline, Recall, Jac and DSC are respectively improved by 2.58%, 2.12% and 1.88%. Net3-2 increases the PC method, further distinguishes the phase consistency between the content features of the source domain and the content features of the target domain, enhances the decoupling between the domain-invariant content feature encoder and the style encoder, and removes the domain-specific features in the domain-invariant content feature encoder, and compared with Baseline, Recall, Precision, Jac and DSC are respectively improved by 4.18%, 1.04%, 3.79% and 3.36%. In order to reduce the difference between the synthesized image of the target domain and the real image of the target domain, and improve the integrity of the target organ, the MSSC method is increased in Net3-3, and compared with Baseline, Recall, Precision, Jac and DSC are respectively improved by 2.42%, 5.06%, 5.25% and 4.44%. Finally, by encouraging the network to pay more attention to the area with high uncertainty to produce better segmentation results, the SCE method is increased in Net3-4, and compared with Baseline, Recall, Precision, Jac and DSC are respectively improved by 2.74%, 5.72%, 6.21% and 5.2%.

[0161] Table 1 Ablation experiment data table

[0162]

[0163] The application compares the segmentation results with the traditional unsupervised domain adaptive segmentation learning network, as shown in Table 2, and Table 2 is a comparison of the segmentation results of the application with the method of others.

[0164] Table 2 Multi-method data comparison table

[0165]

[0166] According to the above table analysis, DDANet, DSAN and DDFseg respectively align the features from the perspective of feature decoupling at the feature level and the predicted image level. However, in the domain adaptation segmentation task with severe modality difference, such as small organ segmentation, they tend to cause local deformation and a large number of false segmentation results. In contrast, DLaST combines feature decoupling learning with self-training and introduces pixel-level shape constraints to alleviate the local deformation problem and improve the domain adaptation performance. The DSC score of DLaST from NCECT to CECT is 65.62%, while the DSC score from CECT to NCECT is only 65.99%. In addition, MSYN and SABE perform low-frequency Fourier space exchange operations at the synthetic image level and the predicted image level, respectively, but due to the high requirement for data, the domain adaptation segmentation performance is poor. On the basis of low-frequency exchange, MeFDA and DOCR also introduce adversarial learning and entropy minimization to enhance the performance of the domain-invariant feature extractor. Specifically, the DSC score of DOCR from NCECT to CECT is 68.61%, while the DSC score from CECT to NCECT is only 68.24%. SASAN focuses on solving the domain adaptation problem by using attention in the image transformation process. In addition, PNP, SIFA and DSAN perform feature and image alignment, but in the small target pancreas segmentation task in this paper, all three methods have achieved poor results and have also produced some serious organ deformation in the image conversion process. Relatively speaking, CycleGAN uses image-level adversarial learning for cross-domain segmentation and obtains a DSC score of 70.47% from NCECT to CECT and 68.63% from CECT to NCECT. However, Synseg integrates the method of CycleGAN into an end-to-end framework, but does not achieve better results. Finally, UESM achieves good results due to the uncertainty-aware self-training strategy and the feature recalibration module. The DSC score of UESM from NCECT to CECT is 74.72%, while the DSC score from CECT to NCECT is 66.44%. The SCUDA segmentation method proposed in this paper is different from most of the above UDA methods. This method focuses on the local area of the target segmented organ and designs phase consistency and multi-stage style inconsistency to alleviate domain shift, and introduces style consistency entropy to effectively utilize unlabelled target domain data. Therefore, in the various indicators of the CT different phase cross-domain task, this method has shown excellent performance and is superior to other UDA methods.

[0167] Table 3 shows the ablation experiment data of the method of the present application, wherein the LPSF method of Net3-1 aims to alleviate domain shift and produce a locally enhanced target organ, and compared with the Baseline, the Precision, Jac and DSC are improved by 5.44%, 3.13% and 2.59% respectively. The PC method of Net3-2 further distinguishes the phase consistency between the content features of the source domain and the content features of the target domain, enhances the decoupling between the domain-invariant content feature encoder and the style encoder, and removes the domain-specific features in the domain-invariant content feature encoder, and compared with the Baseline, the Recall, Precision, Jac and DSC are improved by 1.74%, 1.04%, 3.76% and 3.63% respectively. In order to reduce the difference between the synthesized image of the target domain and the real image of the target domain, and improve the integrity of the target organ, the MSSC method is added in Net3-3, and compared with the Baseline, the Recall, Precision, Jac and DSC are improved by 2.19%, 6.61%, 6.19% and 5.32% respectively. Finally, by encouraging the network to pay more attention to the areas with high uncertainty to produce better segmentation results, the SCE method is added in Net3-4, and compared with the Baseline, the Recall, Precision, Jac and DSC are improved by 3.5%, 7.34%, 7.32% and 6.15% respectively.

[0168]

[0169] Table 4 shows the comparative experiment results of taking CECT images as the source domain and NCECT images as the target domain, and the optimal segmentation performance is indicated in bold. As can be seen from the table, the Recall, Precision, Jac and DSC scores of the full supervision model of the target domain are 76.39%, 78.66%, 64.29% and 75.53% respectively, and when the segmentation model of the source domain is directly used to test the target domain, the Recall, Precision, Jac and DSC scores are 29.98%, 45.38%, 25.21% and 32.39% respectively, which shows that there is a serious domain shift between NCECT and CECT images, which leads to a sharp decline in the performance of deep neural networks in the UDA image segmentation task. The Recall, Precision, Jac and DSC scores of the SCUDA method are 78.37%, 75.27%, 62.82% and 74.57% respectively, and especially in the DSC which is the most important indicator in the segmentation task, the SCUDA method has achieved great improvement and also obtained the best segmentation effect.

[0170] Table 4 Multi-method data comparison table

[0171]

[0172] Example 3

[0173] As Figure 6 An unsupervised domain adaptation medical image segmentation system based on style consistency is provided, and an embodiment of the application is shown, which comprises a data set construction module, an SCUDA segmentation network optimization module, and a parameter updating module.

[0174] The data set construction module is used to collect CT abdominal organ scan images containing multiple source domains and target domains to construct a primary training set; the SCUDA segmentation network optimization module is used to construct an SCUDA segmentation network, to perform image reconstruction through a first subnetwork, and to generate inconsistent mapping areas through different segmentation target domain images in a second subnetwork; and the parameter updating module is used to simultaneously train and update parameters of the two subnetworks to obtain segmented images.

[0175] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the application or the part of the technical solutions that essentially contribute to the prior art can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the embodiments of the method of the application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0176] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered a list of executable instructions for implementing logic functions, and can be specifically embodied in any computer-readable medium for use by an instruction execution system, device or apparatus, such as a computer-based system, a system including a processor, or other system that can fetch and execute instructions from the instruction execution system, device or apparatus, or in conjunction with these instructions execution system, device or apparatus. For the purpose of this specification, the "computer-readable medium" can be any device that can contain, store, communicate, propagate or transport programs for use by an instruction execution system, device or apparatus, or in conjunction with these instruction execution system, device or apparatus.

[0177] More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection (electronic) having one or more wires, a portable computer diskette (magnetic), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can also be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example, via an optical scanner, then compiled, interpreted, or otherwise processed, using an appropriate medium, into a computer program in a suitable language.

[0178] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the above embodiments, various steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, and in another embodiment, any of the following techniques, which are well known in the art, can be used to implement the application: a hybrid of the techniques mentioned above; a combination of one or more of the techniques mentioned above; or one or more other techniques suitable for use in the computer-based systems described above.

[0179] It should be understood that the above-described embodiments are merely illustrative of the application and should not be construed as limiting the same. Although the application has been described with reference to the preferred embodiments, persons skilled in the art will recognize that changes can be made in form and detail without departing from the spirit and the scope of the application. Therefore, the disclosed application should be understood to encompass all such options, modifications and alternatives.

Claims

1. An unsupervised domain adaptation medical image segmentation method based on style consistency, characterized in that, Comprise: The SCUDA network is constructed, image reconstruction is performed through the first sub-network, and image segmentation is performed through the second sub-network; a target domain image x t and a source domain image x s into a first subnetwork to generate a target domain synthetic image x s→t , a source domain reconstructed image x s→t→s , a source domain synthetic image x t→s , a target domain reconstructed image x t→s→t ; a target domain image x t corresponding content features a source domain image x s corresponding content features and a target domain synthetic image x s→t corresponding content features into a second subnetwork to generate segmentation predictions p t , p s and p s→t ; a phase consistency discriminator is used in the first subnetwork to distinguish phase consistency features corresponding to source domain content features and target domain content features, to enhance decoupling of domain-specific features from a domain-invariant content feature encoder E cont , a style encoder for the source domain and a style encoder for the target domain and remove domain-specific features from the domain-invariant content feature encoder. In the second sub-network, style inconsistency estimation is used to perform segmentation prediction on target domain synthetic images of different styles and generate inconsistency mapping areas, and the easy-to-misclassify areas are measured to reduce the domain shift between the target domain synthetic images and the real target domain images; In the second sub-network, style consistency entropy, segmentation consistency loss, and segmentation loss are used; The two sub-networks are simultaneously trained and updated to obtain segmented images; In the second sub-network, an inconsistency map is obtained from the content features corresponding to the synthesized images of the target domain of different styles and the content features corresponding to the real images of the target domain, the difficult regions are measured and constrained, and it is assumed that the high-level features obtained by inputting the synthesized target domain images of different training epochs to the encoder E ej are denoted as where j represents the epoch number of the training epoch, in the training of the style conversion network, the synthesized target domain image of the ejth epoch is selected and denoted as where N is the total number of selected epochs, the style consistency in the style conversion process is measured by the multi-stage style consistency module, and the MSSC module is composed of N segmentation heads G ej , and the predicted probability of the organ of interest in different styles is generated An inconsistency map C is obtained from a plurality of predicted probabilities of different styles s→t is denoted as:​ wherein, is C s→t the kth probability in, is the kth probability in ; By exploiting style inconsistency, a multi-stage style consistency loss is defined The head, tail and edges of the organ of interest, as well as adjacent organ tissue regions, are attended to, denoted as: wherein, is p s→t is the predicted probability of the lth pixel of the kth class in y s is the true label of the source domain image x s , is the true label of the lth pixel of the kth class in y s , is the inconsistency map of the lth pixel of the kth class in C s→t .

2. The style-consistency-based unsupervised domain adaptation medical image segmentation method of claim 1, wherein: The first sub-network comprises a domain-invariant content feature encoder E cont , a style encoder of the target domain , a style encoder of the source domain , an image reconstruction decoder Dec, a local phase-enhanced style fusion module LPSF, a phase-enhanced style fusion module PSF, a discriminator D for the synthesized image of the source domain S , a discriminator D for the synthesized image of the target domain T , a discriminator D for phase consistency pc ; The second sub-network comprises a multi-stage style consistency module MSSC and an image segmenter G, E cont is a ResNet34 backbone network, and adopts the same architecture as E cont G adopts a network with 5 convolutional layers, the first two layers of convolution are composed of residual blocks, and each of the last three layers is composed of a transposed convolution, a BN normalization layer and a ReLU activation layer, and Dec adopts the same architecture as G; The discriminator adopts a network with four convolutional layers and one fully connected layer, and after convolution in each convolutional layer, an IN normalization layer and a LeakyReLU activation layer are performed.

3. The style-consistency-based unsupervised domain adaptation medical image segmentation method of claim 2, wherein: The image reconstruction through the first subnetwork includes sending the source domain image x s and the target domain image x t to the domain-invariant content feature encoder E cont respectively to obtain and at the same time, sending and to the source domain style feature extractor E and the target domain style feature extractor E respectively to extract and to the LPSF module to generate the local phase enhanced content feature of the source domain and the local phase enhanced style fusion feature of the target domain The LPSF is composed of a local content feature extraction submodule LCFE and a phase enhancement style fusion module PSF. The LPSF is used for and the label of the source domain are jointly input into the LCFE, and the local vector O is extracted from the , and is expressed as: s ​ where N o is the number of target vectors, R is a reshape operation, C is the number of channels, then O s is input into a pooling layer, followed by three convolution operations to obtain three dynamic convolution kernel parameters ω c1 , ω c2 and ω c3 , respectively, which are used for convolution operations on , respectively, and then the features are unified to the same size as and spliced in the channel dimension to obtain the local content features of the source domain The LPSF module extracts local content features of the source domain using a fast Fourier transform (FFT) The corresponding phase spectrum The PSF module enhances the features of the source domain image by adding the phase spectrum and is represented as: wherein is an inverse fast Fourier transform; Style features of the target domain Taking FFT operation, extracting amplitude spectrum and phase spectrum Low-pass filtering the target domain amplitude spectrum to remove the target domain noise, obtaining the pure target domain style, and the local phase enhanced style fusion features of the target domain is expressed as: Wherein, is the dot product operation, M is the mask of the low-pass filter, F represents the convolution, average pooling and full connection operations between the two inputs.

4. The style-consistency-based unsupervised domain adaptation medical image segmentation method of claim 3, wherein: The image reconstruction through the first subnetwork further includes generating a target domain synthetic image x s with the same structural content but with the same style as the target domain image x t with the same structural content but with the same style as the target domain image x s→t , x to an input decoder Dec, while embedding into each layer of the decoder Dec, generating a target domain synthetic image x s→t with adaptive instance normalization, denoted as: wherein σ denotes a variance operation and μ denotes a mean operation, is an intermediate output feature obtained by an up-sampling operation performed in the decoder Dec. style features of the source domain Taking FFT operation, extracting amplitude spectrum and phase spectrum Low-pass filtering the amplitude spectrum of the source domain to remove the noise of the source domain, obtaining the pure source domain style, and the local phase enhanced style fusion features of the source domain is expressed as: wherein, is the local phase enhanced style fusion feature of the source domain; Will and The input is fed into the PSF module to generate target domain phase-enhanced content features. Source domain phase enhancement style fusion features Represented as: over the amplitude spectrum and phase spectrum obtained by performing FFT source domain synthetic image x t→s By using AdaIN generation with the decoder Dec, inputting the decoder Dec, while embedding into each layer of the decoder Dec, a source domain synthetic image x with adaptive instance normalization is generated t→s is represented as: wherein, is an intermediate output feature resulting from an up-sampling operation performed in the decoder Dec; Will and The input is fed into the PSF module to generate target domain phase-enhanced content features. Phase-enhanced style fusion features in the target domain Represented as: target domain reconstructed image x t→t By using AdaIN generation with the decoder Dec, input to the decoder Dec, while embedding into each layer of the decoder Dec, generating a target domain reconstructed image x with adaptive instance normalization t→t is represented as: wherein, is the target domain phase enhanced style fusion feature.

5. The style-consistency-based unsupervised domain adaptation medical image segmentation method of claim 4, wherein: Constructing the phase consistency discriminator D pc , which distinguishes the phase consistency features corresponding to the source domain content features and the target domain content features, the loss function corresponding to the discriminator during training is as follows: Wherein E represents the expectation, a discriminator D for the target domain synthetic image is introduced T The adversarial learning is performed, and a loss function corresponding to the discriminator during training is represented as: The source domain synthetic image and the target domain reconstructed image without losing domain invariant information are generated, and a discriminator D for the source domain synthetic image is introduced S Adversarial learning is performed, and a loss function corresponding to the discriminator during training is represented as: The double reconstruction consistency loss is constructed and is represented as:

6. The style-consistency-based unsupervised domain adaptation medical image segmentation method of claim 5, wherein: The two sub-networks simultaneously train and update parameters to obtain a segmented image, including obtaining a target domain image x t Generating a predicted probability of an organ of interest using a segmentation head in different styles in the pre-trained model And generating a corresponding prediction result If the foreground values of the same pixels in the N prediction results are the same, it is considered as a reliable prediction result The target domain image x t is sent into the constantly updated domain-invariant content feature encoder E cont , and the generated is sent into the constantly updated target domain image segmentation head G to generate the predicted probability p t of the organ of interest and the predicted result y, and the inconsistent mapping E of the target domain image is obtained by performing an exclusive or operation on the two predicted results y and , and the style consistency entropy loss is defined, expressed as: wherein e lk is the kth pixel of the lth class in the inconsistent mapping E, is the predicted probability map p t of the lth class in the kth pixel of the organ of interest, and the value of ε is set to 1e-7. The obtained reliable consistency region 1-E is used for joint training of the source domain image and the target domain image, and through combination of data enhancement and Dropout, the target domain image x t An input target segmentation head is generated to generate two prediction probability maps of the organ of interest And During the training process, the output segmentation consistency loss for the target domain image is represented as: Training unlabeled target domain images by segmentation loss is represented as: Among them, e lk It is the l-th pixel of class k in e. for The l-th pixel of the k-th class in the image. for The l-th pixel of the k-th class in the image. for The l-th pixel of the k-th class; define a segmentation loss is represented as: wherein, is p s the lth pixel of the kth class in the image, is y s the lth pixel of the kth class in the image; The overall style consistency segmentation loss is represented as: Segmentation of the target organ is performed by optimizing L total Segmentation of the target organ is performed by optimizing L 7. A system employing the style-consistency-based unsupervised domain adaptation medical image segmentation method according to any one of claims 1-6, characterized in that: The data set construction module, the SCUDA segmentation network optimization module, and the parameter updating module are included. The data set construction module is used to collect CT abdominal organ scan images of source domain and target domain to construct the original training set; The SCUDA segmentation network optimization module is used to construct the SCUDA segmentation network, perform image reconstruction through the first sub-network, and generate inconsistency mapping areas through the prediction images of different segmentation target domain images through the second sub-network; The parameter updating module is used to simultaneously train and update the parameters of the two sub-networks to obtain segmented images.

8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor executes the computer program to realize the steps of the style consistency-based unsupervised domain adaptive medical image segmentation method in any one of claims 1 to 6.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the steps of the style consistency-based unsupervised domain adaptive medical image segmentation method in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Cross-modal medical image segmentation system and method

    CN116721116A

  • Feature reconstruction-based unsupervised domain adaptive OCT image segmentation method and system

    CN116823851A