Medical Image Segmentation Method and System Based on Diffusion Model and Domain Adaptation
Through the medical image segmentation method based on diffusion model and domain adaptation, the pseudo-label, conditional diffusion model and deformation enhancement module are used to solve the problem of insufficient model generalization ability due to scarcity of labeling and inter-domain differences in 3D medical image segmentation, and a better target domain adaptation effect is achieved.
Patent Information
- Application Number
- CN202410897903.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-04
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-07-04
AI Technical Summary
The prior art is difficult to effectively solve the problem of unsupervised domain adaptation caused by scarcity of labeling and interdomain differences, especially in 3D medical imaging segmentation, the model performs poorly on image data in different medical centers.
The medical image segmentation method based on diffusion model and domain adaptation is adopted to gradually solve the differences between domains by generating pseudo-labels, conditional diffusion models and deformation enhancement modules, and enhance the model's generalization ability of the target domain.
This method generates image-mask pairs that conform to the characteristics of the target domain by generating pseudo-labels and conditional diffusion models, and combines the deformation enhancement module to increase the diversity of the generated images, optimizes model parameters, and improves their performance in the target domain.
Smart Images

Figure CN118799337B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and in particular, to a medical image segmentation method, system, computer device, and storage medium based on a diffusion model and domain adaptation. Background Art
[0002] Precise annotation of 3D medical images requires professional knowledge and a large amount of time, so high-quality voxel-level annotation data is very scarce; and due to the large differences that may exist in imaging equipment, imaging protocols, and patient groups in different medical centers, these inter-domain differences lead to poor performance of models trained in one center on the image data of another center, and the high-dimensional characteristics and complex structure of 3D medical image data increase the difficulty of annotation and analysis.
[0003] Currently, the following several methods are usually adopted to solve the above problems. The first one is (Teacher-Student Frameworks), where the teacher model is trained on the source domain (annotated data), and then the student model is adapted to the target domain (unannotated data) in an unsupervised manner. However, it requires designing and training two models, increasing the computational overhead and complexity, and although the teacher model performs well on the source domain, the student model may not perform well when adapting to the target domain, especially when the difference between the source domain and the target domain is large. The second one is to use a generative adversarial network (GANs) to make the generated target domain image data have a similar distribution to the source domain data, so as to achieve domain adaptation. However, due to the instability of the adversarial training process, it is easy to cause mode collapse, and training GANs requires a large amount of computing resources, especially for 3D data. The third one is semi-supervised learning (Semi-Supervised Learning), which jointly trains with a small amount of target domain annotation data and a large amount of source domain annotation data. However, although the annotation requirements are reduced, a certain amount of annotation data still needs to be provided on the target domain, which may be difficult to obtain in practical applications, and when the target domain annotation data is very small, the improvement of model performance is limited. The fourth one is feature alignment, which aligns the features of the source domain and the target domain to make the performance of the model on the target domain close to that on the source domain. However, effective feature alignment may require complex selection and transformation of data features, and only aligning at the feature level may not be able to completely eliminate the distribution differences between the source domain and the target domain. It can be seen that the above methods cannot effectively solve the unsupervised domain adaptation problem caused by scarce annotation and inter-domain differences. Summary of the Invention
[0004] Based on this, it is necessary to provide a medical image segmentation method, system, computer device, and storage medium based on diffusion models and domain adaptation for the above technical problems to solve at least one of the problems existing in the above prior art.
[0005] The embodiments of the present application are implemented as follows. In the first aspect, a medical image segmentation method based on diffusion models and domain adaptation is provided, including the following steps:
[0006] Collect target domain image data and source domain image data;
[0007] Generate pseudo-labels for unlabeled images in the target domain image data through a preset model to obtain target domain image-mask pairs;
[0008] Mask the source domain image data and its corresponding annotation to obtain source domain image-mask pairs, and perform deformation enhancement processing on the source domain image-mask pairs and the target domain image-mask pairs through a preset deformation enhancement module;
[0009] Input the deformed and enhanced source domain image-mask pairs and target domain image-mask pairs into a pre-constructed conditional diffusion model for iterative training to generate multiple new image-mask pairs;
[0010] Mix the new image-mask pairs with the source domain image-mask pairs to obtain a mixed dataset, and perform iterative training on a pre-constructed domain adaptation model through the mixed dataset, and perform segmentation processing on the image to be segmented based on the trained domain adaptation model to obtain a segmentation result.
[0011] In the second aspect, a medical image segmentation system based on diffusion models and domain adaptation is provided, including:
[0012] A data acquisition unit for collecting target domain image data and source domain image data;
[0013] A pseudo-label generation unit for generating pseudo-labels for unlabeled images in the target domain image data through a preset model to obtain target domain image-mask pairs;
[0014] A deformation enhancement processing unit for masking the source domain image data and its corresponding annotation to obtain source domain image-mask pairs, and performing deformation enhancement processing on the source domain image-mask pairs and the target domain image-mask pairs through a preset deformation enhancement module;
[0015] A conditional diffusion model training unit for inputting the deformed and enhanced source domain image-mask pairs and target domain image-mask pairs into a pre-constructed conditional diffusion model for iterative training to generate multiple new image-mask pairs;
[0016] A domain adaptation model training unit is configured to mix the new image-mask pairs with the source domain image-mask pairs to obtain a mixed dataset, iteratively train a pre-constructed domain adaptation model through the mixed dataset, and perform segmentation processing on the image to be segmented based on the trained domain adaptation model to obtain a segmentation result.
[0017] In a third aspect, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the above-mentioned medical image segmentation method based on a diffusion model and domain adaptation is implemented.
[0018] In a fourth aspect, a readable storage medium is provided. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the above-mentioned medical image segmentation method based on a diffusion model and domain adaptation.
[0019] For the above-mentioned medical image segmentation method, system, computer device, and storage medium based on a diffusion model and domain adaptation, the implementation of the method includes: collecting target domain image data and source domain image data; generating pseudo-labels for unlabeled images in the target domain image data through a preset model to obtain target domain image-mask pairs; masking the source domain image data with its corresponding annotation to obtain source domain image-mask pairs, and performing deformation enhancement processing on the source domain image-mask pairs and the target domain image-mask pairs through a preset deformation enhancement module; inputting the deformed and enhanced source domain image-mask pairs and target domain image-mask pairs into a pre-constructed conditional diffusion model for iterative training to generate multiple new image-mask pairs; mixing the new image-mask pairs with the source domain image-mask pairs to obtain a mixed dataset, iteratively training a pre-constructed domain adaptation model through the mixed dataset, and performing segmentation processing on the image to be segmented based on the trained domain adaptation model to obtain a segmentation result. In the embodiments of the present application, by generating pseudo-labels, preliminary annotation information can be provided for subsequent domain adaptation training. Through the conditional diffusion model, image-mask pairs conforming to the characteristics of the target domain can be generated according to the source domain data, enhancing the generalization ability of the model to the target domain. By performing deformation enhancement processing on the image-mask pairs generated by the conditional diffusion model through the deformation enhancement module, the diversity and deformation ability of the generated images can be increased to help the model better adapt to the changes in the target domain. Training the domain adaptation model by combining the source domain data and the generated image-mask pairs can optimize the parameters of the model and improve its performance in the target domain. Description of the Drawings
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments of the present invention. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0021] Figure 1 It is a schematic diagram of an application environment of a medical image segmentation method based on a diffusion model and domain adaptation in an embodiment of the present invention;
[0022] Figure 2 It is a schematic diagram of an application environment of a diffusion model prediction method in an embodiment of the present invention;
[0023] Figure 3 It is a flowchart of an implementation of a medical image segmentation method based on a diffusion model and domain adaptation in an embodiment of the present invention;
[0024] Figure 4 It is a schematic diagram of the structure of a medical image segmentation device based on a diffusion model and domain adaptation in an embodiment of the present invention;
[0025] Figure 5 It is a schematic diagram of a computer device in an embodiment of the present invention. Detailed implementation manners
[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0027] A medical image segmentation method based on a diffusion model and domain adaptation provided in this embodiment can be applied to an application environment such as Figure 1 . Specifically, it can include three parts. The first part is pseudo-label acquisition, the second part is the training and generation of the diffusion model, and the third part is unsupervised domain adaptation training. By inputting source domain image-mask pairs and unlabeled image data in the target domain, pseudo-labels of the unlabeled image data in the target domain are generated through a semi-supervised or weakly supervised model, and then, after deformation enhancement with the source domain image-mask pairs, they are input into the diffusion model together for iterative training of the diffusion model, and the new image-mask pairs generated by the diffusion model are input into the domain adaptation model for iterative training.
[0028] Among them, Ds represents the source domain data, Dt represents the target domain data, Ys represents the source domain label, Ypseudo represents the source domain pseudo-label, and Dg represents the generated data.
[0029] See Figure 2 , which provides an application environment diagram of a prediction method based on a diffusion model. After the source domain image and the target domain image are first subjected to deformation enhancement processing through a deformation enhancement module and then input into a conditional diffusion model for forward and backward propagation processing, a new image-mask pair is obtained.
[0030] Among them, y0 represents a label or a pseudo-label, x0 represents a source domain or a target domain image, x' represents the deformed image, t represents the number of diffusion times, and T represents the total number of diffusion times.
[0031] In one embodiment, as Figure 3 shown, a medical image segmentation method based on a diffusion model and domain adaptation is provided, including the following steps:
[0032] In step S110, target domain image data and source domain image data are collected;
[0033] In the embodiments of the present application, the target domain image data may include image data such as medical images and street view images, and the image data includes images without labeled tags. The source domain image data may come from a field different from the target domain, such as image data in the fields of aerial remote sensing and satellite remote sensing, etc., and it may be labeled with tags through manual annotation. The above source domain image and target domain image can be obtained from existing data sets, or can be docked with medical platforms, street control platforms, etc. to obtain corresponding image data.
[0034] In the embodiments of the present application, after collecting the target domain image data and the source domain image data, a training data set, a validation data set, and a test data set can be constructed for subsequent use.
[0035] In step S120, a pseudo-label is generated for the unlabeled images in the target domain image data through a preset model to obtain a target domain image-mask pair;
[0036] In the embodiments of the present application, the preset model can be a semi-supervised model or a weakly-supervised model. Taking the semi-supervised model as an example, consistency regularization and pseudo-labeling techniques can be used to generate pseudo-labels for the target domain images. Specifically, labeled sample data and unlabeled sample data, that is, target domain image data, can be collected to construct a convolutional neural network CNN. The convolutional neural network is initially trained through the labeled sample data, and the loss between the predicted result and the true result is calculated using consistency regularization. The convolutional neural network is trained based on this loss, and based on the trained convolutional neural network, pseudo-labels for the unlabeled sample data in the target domain image data are generated.
[0037] It is understandable that by masking the target domain image data with its corresponding annotation, a target domain image-mask pair is obtained. The target domain image can be any natural image, and the mask is usually a binary image representing the foreground and background.
[0038] In step S130, the source domain image data is masked with its corresponding annotation to obtain a source domain image-mask pair, and the source domain image-mask pair and the target domain image-mask pair are subjected to deformation enhancement processing through a preset deformation enhancement module;
[0039] In the embodiment of the present application, the deformation enhancement module can adopt a variety of different transformation methods to enhance the data, specifically including transformation methods such as affine transformations, perspective transformations, and non-linear warping. The specific deformation enhancement process is as follows: preprocess the source domain image-mask pair and the target domain image-mask pair, such as adjusting the image size and performing normalization processing, etc., to ensure that the sizes of the image and the mask are consistent, and the numerical range is suitable for subsequent model input. Then, any one of the above transformation methods can be selected to perform deformation processing on each preprocessed image-mask pair to generate multiple deformed and enhanced image-mask pairs, and then save them to a specified directory for subsequent model training. It should be understood that the file naming and saving path need to be clear for subsequent query and use.
[0040] Among them, the affine transformation specifically includes: rotation, which is used to randomly rotate the image and the mask by an angle. Scaling, which is used to randomly scale the image and the mask. Translation, which is used to translate the image and the mask in a random direction and distance. Shearing, which is used to perform a random shearing transformation on the image and the mask. The specific transformation process is as follows: randomly select a rotation angle, for example, between -30 degrees and 30 degrees, and use a rotation matrix to rotate the image and the mask, keeping the sizes of the image and the mask consistent. Then, randomly select a scaling factor, for example, between 0.8 and 1.2, and use the scaling factor to scale the rotated image and mask. Then, randomly select the translation distance and direction, and use a translation matrix to translate the scaled image and mask. Finally, randomly select a shearing angle to shear the translated image and mask to achieve the transformation.
[0041] Among them, the perspective transformation is to perform a random perspective transformation on the image and the mask to simulate different shooting angles. For example, randomly select the offsets of four corner points and use a perspective transformation matrix to transform the image and the mask to simulate different perspectives, thereby achieving deformation enhancement.
[0042] Among them, the non-linear distortion includes elastic deformation and noise perturbation. The elastic deformation is to randomly elastically deform the image and the mask, causing local distortion of the image to generate a random displacement field, and applying the displacement field to locally distort the image and the mask. The noise perturbation is to randomly perturb the image and the mask with noise, simulating various noise conditions, ensuring that the noise intensity is moderate and does not affect the semantic information of the image.
[0043] In step S140, the deformed and enhanced source domain image-mask pair and the target domain image-mask pair are input into a pre-constructed conditional diffusion model for iterative training to generate multiple new image-mask pairs.
[0044] In the embodiment of the present application, first, the deformed and enhanced source domain image-mask pair and the target domain image-mask pair are preprocessed, such as resizing, normalizing, etc., so that the sizes of the image and the mask are consistent, and the numerical range is suitable for model input. Then, an initial conditional diffusion model is constructed. Its network architecture may specifically include an input layer for receiving the preprocessed source domain image-mask pair and the target domain image-mask pair, an encoder layer for extracting features using convolutional layers, batch normalization, and activation functions; a decoder layer that passes multi-scale features through skip connections with the encoder layer, for reconstructing the image through deconvolutional layers or upsampling layers; and a diffusion module for introducing noise in each step of the diffusion process and guiding the denoising process through conditional information, that is, the mask. The deformed and enhanced source domain image-mask pair and the target domain image-mask pair are preprocessed by the above-constructed initial conditional diffusion model until the preset convergence conditions are met, such as the number of iterations reaches the preset number, or the loss value is less than the preset threshold, and the trained conditional diffusion model can be obtained.
[0045] It can be understood that in each iteration process of the conditional diffusion model, a batch of new image-mask pairs will be generated. The new image-mask pairs conform to the target domain features and can enhance the generalization ability of the model to the target domain.
[0046] In step S150, the new image-mask pairs are mixed with the source domain image-mask pairs to obtain a mixed data set, and the pre-constructed domain adaptation model is iteratively trained through the mixed data set, and the image to be segmented is segmented based on the trained domain adaptation model to obtain a segmentation result.
[0047] In the embodiments of the present application, the domain adaptation model solves the problem of different distributions between the source domain and the target domain. Through domain adaptation techniques, the model can achieve good performance on the target domain. Specifically, it can include adversarial domain adaptation, reconstruction-based domain adaptation, and statistical matching for domain adaptation. Among them, the adversarial domain adaptation model can use adversarial training methods to make the feature distributions of the source domain and the target domain as similar as possible. Common methods include DANN (Domain-Adversarial Neural Network) and ADDA (Adversarial Discriminative Domain Adaptation). The reconstruction-based domain adaptation can use autoencoders or generative adversarial networks (GANs) to convert source domain data into the style of the target domain, such as CycleGAN, UNIT, etc. The statistical matching can make the feature distributions of the source domain and the target domain match through methods such as the maximum mean discrepancy (MMD).
[0048] Specifically, taking the DANN model as an example, the network architecture of the domain adaptation model is described. It specifically includes a feature extractor, such as a convolutional neural network CNN, for extracting high-level features, a label classifier for performing classification tasks on the features extracted by the feature extractor, and a domain discriminator for distinguishing whether the input data comes from the source domain or the target domain. Through the above network architecture, the image-mask pairs in the mixed dataset are sequentially segmented until the preset convergence conditions are met, such as the number of iterations reaches the preset number, or the loss value is less than the preset threshold, and the trained domain adaptation model can be obtained.
[0049] Furthermore, by using the trained domain adaptation model to segment the image to be segmented, the segmentation result can be obtained.
[0050] In an embodiment of the present application, a medical image segmentation method based on a diffusion model and domain adaptation is provided, including: collecting target domain image data and source domain image data; generating pseudo-labels for unlabeled images in the target domain image data through a preset model to obtain target domain image-mask pairs; masking the source domain image data and its corresponding annotations to obtain source domain image-mask pairs, and performing deformation enhancement processing on the source domain image-mask pairs and the target domain image-mask pairs through a preset deformation enhancement module; inputting the deformed and enhanced source domain image-mask pairs and target domain image-mask pairs into a pre-constructed conditional diffusion model for iterative training to generate multiple new image-mask pairs; mixing the new image-mask pairs with the source domain image-mask pairs to obtain a mixed dataset, and iteratively training a pre-constructed domain adaptation model through the mixed dataset, and performing segmentation processing on the image to be segmented based on the trained domain adaptation model to obtain a segmentation result. In an embodiment of the present application, by generating pseudo-labels, preliminary annotation information can be provided for subsequent domain adaptation training. Through the conditional diffusion model, image-mask pairs conforming to the characteristics of the target domain can be generated according to the source domain data, enhancing the generalization ability of the model to the target domain. By performing deformation enhancement processing on the image-mask pairs generated by the conditional diffusion model through the deformation enhancement module, the diversity and deformation ability of the generated images can be increased to help the model better adapt to the changes in the target domain. Combining the source domain data and the generated image-mask pairs to train the domain adaptation model can optimize the parameters of the model and improve its performance in the target domain.
[0051] In an embodiment of the present application, the inputting the deformed and enhanced source domain image-mask pairs and target domain image-mask pairs into a pre-constructed conditional diffusion model for iterative training includes:
[0052] Performing preprocessing on the deformed and enhanced source domain image-mask pairs and target domain image-mask pairs to obtain preprocessed source domain image-mask pairs and target domain image-mask pairs;
[0053] Constructing a basic network and introducing a conditional diffusion model into the basic network to obtain an initial conditional diffusion model;
[0054] Inputting the preprocessed source domain image-mask pairs and target domain image-mask pairs into the initial conditional diffusion model for iterative training.
[0055] Specifically, preprocess the augmented source domain image-mask pairs and the target domain image-mask pairs, such as resizing, normalizing, etc., to ensure that the sizes of the images and masks are consistent and the numerical ranges are suitable for model input. Build a basic network, such as U-Net, which performs excellently in image generation tasks. U-Net consists of a downsampling path (encoder) and an upsampling path (decoder), with features passed through skip connections in the middle. Introduce a conditional diffusion model into the basic network. The conditional diffusion model is a generative model based on the diffusion process to obtain the initial conditional diffusion model. Then, input the preprocessed source domain image-mask pairs and target domain image-mask pairs into the initial conditional diffusion model and perform iterative training on the initial conditional diffusion model. Use the image-mask pairs of the source domain to train the initial conditional diffusion model to achieve high-quality image generation and pseudo-label generation.
[0056] Among them, the network architecture of the initial conditional diffusion model can specifically be:
[0057] Input layer: Used to obtain the preprocessed source domain image-mask pairs;
[0058] Encoder: Used to extract multi-scale features of the source domain image-mask pairs;
[0059] Decoder: There is a skip connection between the decoder and the encoder to transfer multi-scale features to the decoder;
[0060] Conditional diffusion module: Used to introduce noise in each step of the diffusion process and guide the denoising process through conditional information.
[0061] Through this input layer, the preprocessed source domain image-mask pairs and target domain image-mask pairs can be received. The image and mask are concatenated together as the input. The encoder may include multiple convolutional layers. By gradually downsampling the image, multi-scale features are extracted. After each convolutional layer, there are batch normalization and ReLU activation functions. The decoder includes multiple transposed convolutional layers and upsampling layers, which can gradually reconstruct the image. After each transposed convolutional layer, there are batch normalization and ReLU activation functions. Between each downsampling layer and the corresponding upsampling layer, the extracted features can be transferred through skip connections to retain high-resolution information. The conditional diffusion module, that is, the conditional diffusion model, can introduce noise in each step of the diffusion process and guide the denoising through conditional information, that is, the mask. Output the denoised image and compare it with the real image to calculate the loss value. Iteratively train the model based on this loss value.
[0062] In an embodiment of the present application, the step of inputting the preprocessed source domain image-mask pairs and the target domain image-mask pairs into the initial conditional diffusion model for iterative training includes:
[0063] Step a: Gradually add noise to the source domain image and the target domain image to obtain multiple noisy images;
[0064] Step b: When adding noise at each step, record the current noise level and the corresponding noisy image;
[0065] Step c: Using the masks corresponding to the source domain image and the target domain image as conditional information, start from the strongest noisy image and gradually denoise to reconstruct the source domain image and the target domain image;
[0066] Step d: During each denoising process, output the denoised image and calculate the loss value between the denoised image and the real image using a preset loss function;
[0067] Repeat the above steps a - d until the loss value meets the first preset loss condition to obtain a trained conditional diffusion model.
[0068] Specifically, the initial conditional diffusion model may include a forward diffusion process and a reverse diffusion process. The forward diffusion process refers to gradually adding noise to the source domain image or the target domain image, and the reverse diffusion process refers to gradually denoising to generate sample data. And during the diffusion process, the initial conditional diffusion model can receive the current noisy image and conditional information, that is, the mask, and predict the denoised image.
[0069] The training process of the initial conditional diffusion model is specifically as follows:
[0070] Forward diffusion process:
[0071] Gradually add noise to the source domain image and the target domain image to obtain multiple noisy images;
[0072] When adding noise at each step, record the current noise level and the corresponding noisy image;
[0073] Reverse diffusion process:
[0074] Using the masks corresponding to the source domain image and the target domain image as conditional information, start from the strongest noisy image and gradually denoise to reconstruct the source domain image and the target domain image;
[0075] It can be understood that each iteration process includes a forward diffusion process and a reverse diffusion process. The training data for each batch is randomly selected from the training set to ensure that the model can generalize to different samples.
[0076] After the above forward diffusion process and reverse diffusion process, a preset loss function can be used to calculate the loss value between the denoised image and the real image. Optionally, a mean squared error (MSE) loss function or other suitable generation loss function can be adopted to calculate the loss value between the denoised image and the original image. When the loss value is less than the preset threshold, the training is completed. By combining the above conditional information, it can be ensured that the generated image conforms to the structure of the mask.
[0077] Among them, an optimizer (such as Adam) can be used to update the model parameters and minimize the loss function.
[0078] In an embodiment of the present application, after the conditional diffusion model training is completed, the model performance can be evaluated through a validation set. The specific evaluation metrics can include PSNR, SSIM, etc., to measure the quality of the generated image. According to the evaluation results, the hyperparameters of the model can be adjusted correspondingly to optimize the model performance.
[0079] In an embodiment of the present application, the deformation enhancement processing of the source domain image-mask pair and the target domain image-mask pair by the preset deformation enhancement module includes:
[0080] Preprocess the source domain image-mask pair and the target domain image-mask pair to make the size and format of the target image consistent with its corresponding mask;
[0081] Through a preset deformation algorithm, perform deformation enhancement processing on the preprocessed source domain image-mask pair and target domain image-mask pair.
[0082] Specifically, the deformation enhancement module can use set transformation operations to enhance the dataset, which can specifically include transformation methods such as affine transformations, perspective transformations, and non-linear warping. The specific implementation process is as follows: Preprocess the source domain image-mask pair and the target domain image-mask pair, such as resizing the image, performing normalization processing, etc., to ensure that the size of the image and the mask is consistent and the numerical range is suitable for further processing. Then, any one of the above transformation methods can be selected to perform deformation processing on each preprocessed image-mask pair to generate multiple deformed and enhanced image-mask pairs, and then save them to a specified directory for subsequent model training.
[0083] It can be understood that the source domain image-mask pair and the target domain image-mask pair can be batch-processed for deformation enhancement according to batches, or any randomly selected image-mask pair can be processed for deformation enhancement.
[0084] In one embodiment of the present application, the deformation enhancement processing of the preprocessed source domain image-mask pair and target domain image-mask pair through a preset deformation algorithm includes:
[0085] Construct an initial affine matrix;
[0086] Construct a rotation matrix and update the rotation part of the initial affine matrix;
[0087] Add a translation vector and a scaling factor to the updated initial affine matrix to obtain a target affine matrix;
[0088] Generate an affine network based on the target affine matrix;
[0089] Construct a target displacement field, add the target displacement field to the affine network to obtain a transformation network;
[0090] Sample the preprocessed source domain image-mask pair and target domain image-mask pair through the transformation network to generate a transformed source domain image-mask pair and a transformed target domain image-mask pair.
[0091] Specifically, the process of transforming an image-mask pair through an affine transformation method is as follows: Define a rotation matrix r, a scaling factor sc, a translation amount sh, a window size w, a field size f, a displacement field scaling factor a, and a smoothing number sn, and use them as inputs. Initialize an affine matrix A. Use the method of converting a rotation vector represented by an angle axis to a rotation matrix and assign it to the rotation part of the affine matrix A to update the rotation part of the affine matrix A. Assign the value of the translation vector sh to the translation part of the affine matrix A. For each dimension d, assign the d-th component of sh to the translation part of A, that is, A[:, d, 3]. Apply the scaling factor sc to the scaling part of the affine matrix A, that is, for each axis a, multiply the d-th component of the scaling factor sc by the corresponding part A[:, d, a] of the affine matrix A. Then, remove the fourth column of the affine matrix A, only the transformation information of the first three dimensions is needed, and create a grid with the same size as the input image, and apply the grid to the affine matrix to generate a transformed network, that is, an affine network. Construct a target displacement field, add the target displacement field to the affine network to obtain a transformation network; Sample the preprocessed target image-mask pair through the transformation network to generate a transformed image-mask pair.
[0092] Furthermore, the construction of the target displacement field includes:
[0093] Construct an initial displacement field;
[0094] Adjust the initial displacement field using the scaling factor;
[0095] Perform multiple smoothing operations on the scaled initial displacement field using a preset convolution kernel;
[0096] Crop the smoothed initial displacement field to remove the padded part;
[0097] Adjust the cropped initial displacement field to the same size as the original volume to obtain the target displacement field.
[0098] Specifically, the padding size pad required for the displacement field can be calculated based on the window size w, usually w / 2. Then, an initial displacement field dz, dy, dx representing the displacement on each axis is created according to the window size w and the field size f. The initial value can be random noise or zero. By multiplying the initial displacement field by the scaling factor a, the scaling of the initial displacement field dz, dy, dx is achieved. Then, 3x3 convolution operations are used to smooth dz, dy, dx respectively to reduce noise. Remove the padded part from the smoothed dz, dy, dx to restore it to the original size. After adjusting dz, dy, dx to the same size as the input image using the interpolation method, add it to the affine network to create the final transformation network. Use the grid sampling method to remap the input image to the grid position in the transformation network to generate the transformed image to obtain the target image-mask pair.
[0099] In an embodiment of the present application, taking the new image-mask pair as the new target domain image-mask pair, the iterative training of the pre-constructed domain adaptation model by the mixed dataset includes:
[0100] Input the source domain image and the new target domain image in the mixed dataset into the feature extractor of the domain adaptation model to respectively extract the source domain high-level features and the target domain high-level features;
[0101] Input the source domain high-level features into the label classifier for prediction to obtain the predicted class label;
[0102] Calculate the classification loss value based on the predicted class label and the true label;
[0103] Input the source domain high-level features and the target domain high-level features into the domain discriminator to calculate the domain discrimination loss value;
[0104] Respectively perform weighted averaging on the classification loss value and the domain discrimination loss value to obtain the total loss value;
[0105] Based on the total loss value, perform iterative training until the total loss value meets the second preset loss condition, and the training is completed.
[0106] Specifically, the domain adaptation model can be the DANN model in adversarial domain adaptation. Its network architecture includes: a feature extractor for extracting high-level features from the input source domain images and the new target domain images, which can be a convolutional neural network (CNN) or other types of networks; a label classifier for classifying through the high-level features extracted by the feature extractor and predicting the class labels of the source domain data; and a domain discriminator for distinguishing whether the input data is source domain data or target domain image data. Through adversarial training, the features extracted by the feature extractor are made as similar as possible between the source domain and the target domain. The specific training process is as follows: the forward propagation process of the source domain data and the forward propagation process of the target domain. The forward propagation process of the source domain data specifically is to extract the high-level source domain features of the source domain image data through the feature extractor, input the extracted high-level domain features into the label classifier for prediction to obtain the predicted class labels, and calculate the classification loss value based on the predicted class labels and the true labels through a preset loss function, such as the cross-entropy loss function. The forward propagation process of the target domain is to extract the high-level target domain features of the target domain image data through the feature extractor, input the high-level target domain features and the high-level source domain features into the domain discriminator, set the domain label of the source domain features to 1 and the domain label of the target domain features to 0, and calculate the domain discrimination loss through a preset loss function, such as the binary cross-entropy loss function. Then, through adversarial training, the features extracted by the feature extractor are made as similar as possible between the source domain and the target domain. The gradient of the domain discriminator is backpropagated to the feature extractor so that it learns features that cannot distinguish between the source domain and the target domain. By respectively taking the weighted average of the obtained classification loss value and the domain discrimination loss value, the total loss value is obtained. Based on the total loss value, iterative training is carried out until the total loss value meets the second preset loss condition, and the training is completed. Through the above process, the domain adaptation model can effectively transfer knowledge between the source domain and the target domain and improve the performance of the target domain task. The DANN model makes the features extracted by the feature extractor as similar as possible between the two domains through adversarial training, thereby achieving domain adaptation.
[0107] Among them, the weights of the classification loss and the domain discrimination loss need to be balanced and can be dynamically adjusted.
[0108] It should be noted that optimizers such as Adam or SGD can be used to update the parameters of the feature extractor, the label classifier, and the domain discriminator. In each training round, the label classifier and the domain discriminator are alternately updated to ensure that the features extracted by the feature extractor are as similar as possible between the source domain and the target domain while maintaining a high classification accuracy.
[0109] In the embodiments of the present application, after the domain adaptation model is trained, the performance of the label classifier can be evaluated on a pre-constructed source domain validation set to determine the classification accuracy, and the transfer ability of the model can be evaluated on the target domain to ensure that the model has good performance on the target domain.
[0110] In an embodiment of the present application, after obtaining the actual segmentation result based on the trained domain adaptation model, it includes:
[0111] Construct a target domain test set based on the target domain image data, where the target domain test set includes target domain images and their corresponding segmentation masks;
[0112] Load the target domain images and their corresponding segmentation masks into the target segmentation model in batches through a data loader for prediction to obtain predicted segmentation regions;
[0113] Determine the similarity between the predicted segmentation region and the true segmentation region;
[0114] Evaluate the target segmentation model based on the similarity.
[0115] Specifically, a target domain test set can be constructed according to the collected target domain image data, which may include target domain images and their corresponding segmentation masks. Preprocess the test data, such as resizing, normalizing, etc., to keep the test data in the same format as the training data. Load the trained domain adaptation model, that is, the image segmentation model, from the model file saved in the training stage. Use a data loader (DataLoader) to load the test data into the domain adaptation model in batches for prediction. Obtain the predicted segmentation result, and evaluate the model through image segmentation evaluation metrics such as DICE Score and NSD Score. Among them, DICE Score is used to measure the similarity between the predicted segmentation result and the true segmentation result, and NSD Score is used to predict the distance between the predicted boundary and the true boundary. Taking DICE Score as an example, it can be specifically calculated by the following formula:
[0116] DICE(A, B)=2
[0117] Where A represents the predicted segmentation result and B represents the true segmentation result. When the DICE value is 1, the model is optimal; when the DICE value is 0, the model is the worst.
[0118] In the embodiments of the present application, by generating target domain pseudo-labels, preliminary annotation information can be provided for subsequent domain adaptation training. Through the conditional diffusion model, image-mask pairs that conform to the characteristics of the target domain can be generated based on source domain data, enhancing the generalization ability of the model to the target domain. Through the deformation enhancement module, the image-mask pairs generated by the conditional diffusion model are subjected to deformation enhancement processing, which can increase the diversity and deformation ability of the generated images to help the model better adapt to the changes in the target domain. By combining the source domain data and the generated image-mask pairs to train the domain adaptation model, the parameters of the model can be optimized to improve its performance in the target domain.
[0119] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0120] In one embodiment, a medical image segmentation system based on a diffusion model and domain adaptation is provided. The medical image segmentation system based on a diffusion model and domain adaptation corresponds one-to-one with the medical image segmentation method based on a diffusion model and domain adaptation in the above embodiments. As Figure 4 shown, the medical image segmentation system based on a diffusion model and domain adaptation includes a data acquisition unit 10, a pseudo-label generation unit 20, a deformation enhancement processing unit 30, a conditional diffusion model training unit 40, and a domain adaptation model training unit 50. The detailed description of each functional module is as follows:
[0121] The data acquisition unit 10 is configured to acquire target domain image data and source domain image data;
[0122] The pseudo-label generation unit 20 is configured to generate pseudo-labels for unlabeled images in the target domain image data through a preset model to obtain target domain image-mask pairs;
[0123] The deformation enhancement processing unit 30 is configured to mask the source domain image data and its corresponding annotation to obtain source domain image-mask pairs, and perform deformation enhancement processing on the source domain image-mask pairs and the target domain image-mask pairs through a preset deformation enhancement module;
[0124] The conditional diffusion model training unit 40 is configured to input the deformed and enhanced source domain image-mask pairs and the target domain image-mask pairs into a pre-constructed conditional diffusion model for iterative training to generate multiple new image-mask pairs;
[0125] The domain adaptation model training unit 50 is configured to mix the new image-mask pairs with the source domain image-mask pairs to obtain a mixed data set, and perform iterative training on a pre-constructed domain adaptation model through the mixed data set, and perform segmentation processing on the image to be segmented based on the trained domain adaptation model to obtain a segmentation result.
[0126] In an embodiment of the present application, the conditional diffusion model training unit 40 is configured to:
[0127] Preprocess the augmented source domain image-mask pair and the target domain image-mask pair to obtain the preprocessed source domain image-mask pair and the target domain image-mask pair;
[0128] Construct a basic network, and introduce a conditional diffusion model into the basic network to obtain an initial conditional diffusion model;
[0129] Input the preprocessed source domain image-mask pair and the target domain image-mask pair into the initial conditional diffusion model for iterative training.
[0130] In an embodiment of the present application, the conditional diffusion model training unit 40 is configured to:
[0131] Step a: Gradually increase noise in the source domain image and the target domain image to obtain multiple noisy images;
[0132] Step b: When increasing noise at each step, record the current noise level and the corresponding noisy image;
[0133] Step c: Use the masks corresponding to the source domain image and the target domain image as conditional information, and start from the strongest noisy image to gradually denoise and reconstruct the source domain image and the target domain image;
[0134] Step d: In the denoising process at each step, output the denoised image, and use a preset loss function to calculate the loss value between the denoised image and the real image;
[0135] Repeat the above steps a-d until the loss value meets the first preset loss condition, and obtain the trained conditional diffusion model.
[0136] In an embodiment of the present application, the initial conditional diffusion model includes:
[0137] Input layer: Used to obtain the preprocessed source domain image-mask pair and the target domain image-mask pair;
[0138] Encoder: Used to extract multi-scale features of the source domain image-mask pair and the target domain image-mask pair;
[0139] Decoder: The decoder is connected to the encoder through skip connections to transfer multi-scale features to the decoder;
[0140] Conditional diffusion module: Used to introduce noise in each diffusion process and guide the denoising process through conditional information.
[0141] In one embodiment of the present application, the deformation enhancement processing unit 30 is further configured to:
[0142] Preprocess the source domain image-mask pair and the target domain image-mask pair so that the size and format of the target image and its corresponding mask are consistent;
[0143] Perform deformation enhancement processing on the preprocessed source domain image-mask pair and target domain image-mask pair through a preset deformation algorithm.
[0144] In one embodiment of the present application, the deformation enhancement processing unit 30 is further configured to:
[0145] Construct an initial affine matrix;
[0146] Construct a rotation matrix and update the rotation part of the initial affine matrix;
[0147] Add a translation vector and a scaling factor to the updated initial affine matrix to obtain a target affine matrix;
[0148] Generate an affine network based on the target affine matrix;
[0149] Construct a target displacement field, add the target displacement field to the affine network to obtain a transformation network;
[0150] Sample the preprocessed target image-mask pair through the transformation network to generate a transformed target image-mask pair.
[0151] In one embodiment of the present application, the deformation enhancement processing unit 30 is further configured to:
[0152] Construct an initial displacement field;
[0153] Adjust the initial displacement field using a scaling factor;
[0154] Perform multiple smoothing processes on the scaled initial displacement field using a preset convolution kernel;
[0155] Crop the smoothed initial displacement field to remove the padding part;
[0156] Adjust the cropped initial displacement field to the same size as the original volume to obtain the target displacement field.
[0157] In one embodiment of the present application, the domain adaptation model training unit 50 is further configured to:
[0158] Input the source domain images and the new target domain images in the mixed dataset into the feature extractor of the domain adaptation model to extract source domain high-level features and target domain high-level features respectively;
[0159] Input the high-level features of the source domain into a label classifier for prediction to obtain a predicted class label;
[0160] Calculate a classification loss value based on the predicted class label and the true label;
[0161] Input the high-level features of the source domain and the high-level features of the target domain into a domain discriminator to calculate a domain discrimination loss value;
[0162] Respectively perform weighted averaging on the classification loss value and the domain discrimination loss value to obtain a total loss value;
[0163] Based on the total loss value, perform iterative training until the total loss value meets the second preset loss condition, at which point the training is completed.
[0164] In an embodiment of the present application, the system further includes a model evaluation unit for:
[0165] Construct a target domain test set based on the target domain image data, where the target domain test set includes target domain images and their corresponding segmentation masks;
[0166] Load the target domain images and their corresponding segmentation masks into the target segmentation model in batches through a data loader for prediction to obtain predicted segmentation regions;
[0167] Determine the similarity between the predicted segmentation region and the true segmentation region;
[0168] Evaluate the target segmentation model based on the similarity.
[0169] In the embodiments of the present application, by generating pseudo-labels, preliminary annotation information can be provided for subsequent domain adaptation training. Through a conditional diffusion model, image-mask pairs that conform to the characteristics of the target domain can be generated based on source domain data, enhancing the generalization ability of the model to the target domain. Through a deformation enhancement module, deformation enhancement processing is performed on the image-mask pairs generated by the conditional diffusion model, which can increase the diversity and deformation ability of the generated images to help the model better adapt to the changes in the target domain. By combining the source domain data and the generated image-mask pairs to train the domain adaptation model, the parameters of the model can be optimized to improve its performance in the target domain.
[0170] For the specific limitations of the medical image segmentation system based on the diffusion model and domain adaptation, reference may be made to the limitations of the medical image segmentation method based on the diffusion model and domain adaptation in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned medical image segmentation system based on the diffusion model and domain adaptation can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.
[0171] In one embodiment, a computer device is provided. The computer device can be a terminal device, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a readable storage medium. The readable storage medium stores computer-readable instructions. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer-readable instructions are executed by the processor, a medical image segmentation method based on the diffusion model and domain adaptation is implemented. The readable storage medium provided in this embodiment includes a non-volatile readable storage medium and a volatile readable storage medium.
[0172] In an embodiment of the present application, a computer device is provided, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the steps of the medical image segmentation method based on the diffusion model and domain adaptation as described above are implemented.
[0173] In an embodiment of the application, a readable storage medium is provided. The readable storage medium stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the steps of the medical image segmentation method based on the diffusion model and domain adaptation as described above are implemented.
[0174] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a non-volatile readable storage medium or a volatile readable storage medium. When the computer-readable instructions are executed, they can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0175] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0176] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A medical image segmentation method based on diffusion model and domain adaptation, characterized in that: The method comprises: Collecting target domain image data and source domain image data; Generate pseudo labels for unlabeled images in the target domain image data through a preset model to obtain target domain image-mask pairs; After masking the source domain image data and the corresponding annotations, a source domain image-mask pair is obtained, and deformation enhancement processing is performed on the source domain image-mask pair and the target domain image-mask pair by a preset deformation enhancement module; The deformed and enhanced source domain image-mask pair and the target domain image-mask pair are input into a pre-built conditional diffusion model for iterative training to generate a plurality of new image-mask pairs, wherein the conditional diffusion model includes a forward diffusion process and a reverse diffusion process, wherein the forward diffusion process refers to gradually adding noise to the source domain image or the target domain image, and the reverse diffusion process refers to gradually denoising, wherein in the diffusion process, the initial conditional diffusion model is used to receive the current noise image and condition information, and guide denoising based on the condition information; The new image-mask pair is mixed with the source domain image-mask pair to obtain a mixed data set, a pre-built domain adaptation model is iteratively trained through the mixed data set, and the image to be segmented is segmented based on the trained domain adaptation model to obtain a segmentation result.
2. The medical image segmentation method based on diffusion model and domain adaptation according to claim 1, characterized in that: The step of inputting the deformed and enhanced source domain image-mask pair and the target domain image-mask pair into a pre-built conditional diffusion model for iterative training includes: Preprocessing the deformed and enhanced source domain image-mask pair and the target domain image-mask pair to obtain a preprocessed source domain image-mask pair and the target domain image-mask pair; Constructing a basic network, and introducing a conditional diffusion model into the basic network to obtain an initial conditional diffusion model; The preprocessed source domain image-mask pair and the target domain image-mask pair are input into the initial conditional diffusion model for iterative training.
3. The medical image segmentation method based on diffusion model and domain adaptation according to claim 2, characterized in that: The step of inputting the preprocessed source domain image-mask pair and the target domain image-mask pair into the initial conditional diffusion model for iterative training includes: Step a: gradually add noise to the source domain image and the target domain image to obtain multiple noisy images; Step b: When adding noise at each step, record the current noise level and the corresponding noise image; Step c: using the masks corresponding to the source domain image and the target domain image as conditional information, starting from the image with the strongest noise, gradually denoising, and reconstructing the source domain image and the target domain image; Step d: In each denoising process, the denoised image is output, and the loss value between the denoised image and the real image is calculated using a preset loss function; Repeat steps ad above until the loss value meets the first preset loss condition, and obtain a trained conditional diffusion model.
4. The medical image segmentation method based on diffusion model and domain adaptation according to claim 3, characterized in that: The initial condition diffusion model comprises: Input layer: used to obtain the preprocessed source domain image-mask pair and the target domain image-mask pair; Encoder: used to extract multi-scale features of the source domain image-mask pair and the target domain image-mask pair; Decoder: The decoder is connected to the encoder via a jump connection to transmit multi-scale features to the decoder; Conditional diffusion module: used to introduce noise in each diffusion step and guide the denoising process through conditional information.
5. The medical image segmentation method based on diffusion model and domain adaptation according to claim 1, characterized in that: The performing deformation enhancement processing on the source domain image-mask pair and the target domain image-mask pair by using a preset deformation enhancement module includes: Preprocessing the source domain image-mask pair and the target domain image-mask pair so that the target image and its corresponding mask have the same size and format; The pre-processed source domain image-mask pair and the target domain image-mask pair are deformed and enhanced by a preset deformation algorithm.
6. The medical image segmentation method based on diffusion model and domain adaptation according to claim 5, characterized in that: The method of performing deformation enhancement processing on the pre-processed source domain image-mask pair and the target domain image-mask pair by using a preset deformation algorithm includes: Construct the initial affine matrix; Constructing a rotation matrix and updating the rotation part of the initial affine matrix; Adding the translation vector and the scaling factor to the updated initial affine matrix to obtain a target affine matrix; Based on the target affine matrix, generating an affine network; Constructing a target displacement field, and adding the target displacement field to the affine network to obtain a transformation network; The preprocessed source domain image-mask pair and the target domain image-mask pair are sampled through the transformation network to generate transformed source domain image-mask pair and target domain image-mask pair.
7. The medical image segmentation method based on diffusion model and domain adaptation according to claim 6, characterized in that: The constructing of the target displacement field comprises: Construct the initial displacement field; adjusting the initial displacement field using a scaling factor; Using a preset convolution kernel to perform multiple smoothing processes on the scaled initial displacement field; Clipping the smoothed initial displacement field to remove the filling part; The cropped initial displacement field is adjusted to be the same as the original volume to obtain the target displacement field.
8. The medical image segmentation method based on diffusion model and domain adaptation according to claim 1, characterized in that: The step of using the new image-mask pair as a new target domain image-mask pair and iteratively training the pre-built domain adaptation model using the mixed data set includes: Inputting the source domain image and the new target domain image in the mixed data set into the feature extractor of the domain adaptation model to extract source domain high-level features and target domain high-level features respectively; Inputting the source domain high-level features into a label classifier for prediction to obtain a predicted category label; Calculate the classification loss value based on the predicted category label and the true label; Inputting the source domain high-level features and the target domain high-level features into a domain discriminator, and calculating a domain discrimination loss value; Taking a weighted average of the classification loss value and the domain discrimination loss value respectively to obtain a total loss value; Based on the total loss value, iterative training is performed until the total loss value meets the second preset loss condition, and the training is completed.
9. The medical image segmentation method based on diffusion model and domain adaptation according to claim 1, characterized in that: After obtaining the actual segmentation result based on the trained domain adaptation model, the following steps are included: Based on the target domain image data, construct a target domain test set, wherein the target domain test set includes a target domain image and its corresponding segmentation mask; The target domain image and its corresponding segmentation mask are loaded into the target segmentation model in batches through a data loader for prediction to obtain a predicted segmentation region; Determining the similarity between the predicted segmented region and the actual segmented region; Based on the similarity, the target segmentation model is evaluated.
10. A medical image segmentation system based on diffusion model and domain adaptation, characterized in that: The system comprises: A data acquisition unit, used for acquiring target domain image data and source domain image data; A pseudo label generating unit, used to generate pseudo labels for unlabeled images in the target domain image data through a preset model to obtain a target domain image-mask pair; A deformation enhancement processing unit, configured to obtain a source domain image-mask pair after masking the source domain image data and the corresponding annotation, and perform deformation enhancement processing on the source domain image-mask pair and the target domain image-mask pair through a preset deformation enhancement module; A conditional diffusion model training unit, used for inputting the deformed and enhanced source domain image-mask pair and the target domain image-mask pair into a pre-built conditional diffusion model for iterative training to generate a plurality of new image-mask pairs, wherein the conditional diffusion model includes a forward diffusion process and a reverse diffusion process, wherein the forward diffusion process refers to gradually adding noise to the source domain image or the target domain image, and the reverse diffusion process refers to gradually denoising, and in the diffusion process, the initial conditional diffusion model is used for receiving the current noise image and condition information, and guiding denoising based on the condition information; The domain adaptation model training unit is used to mix the new image-mask pair with the source domain image-mask pair to obtain a mixed data set, iteratively train the pre-built domain adaptation model through the mixed data set, and perform segmentation processing on the image to be segmented based on the trained domain adaptation model to obtain a segmentation result.
Citation Information
Patent Citations
Non-supervision SAR image ship target detection method based on domain adaptation
CN116503732A
Unsupervised domain adaptive medical image segmentation method and system based on shape guidance
CN117934494A