A Cross-Domain Adaptive Spinal CT Image Segmentation Method

By using diffusion model for cross-domain adaptive processing in spinal CT image segmentation, the problem of domain differences in multi-source CT data sets is solved, and higher segmentation accuracy and cross-domain adaptability are achieved, while reducing data annotation and training costs.

CN119888237BActive Publication Date: 2025-06-17JIANGSU SHIYU INTELLIGENT MEDICAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510356764.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-17
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

In the prior art, when processing domain differences in multi-source CT data sets, it is difficult to achieve effective cross-domain adaptation, resulting in poor segmentation effect of target domain data.

Method used

The forward diffusion and reverse denoising process is carried out using the diffusion model. By gradually adding noise and using the encoder, decoder and domain generalization module, cross-domain feature alignment and segmentation label recovery are achieved.

Benefits of technology

The segmentation accuracy between multi-source CT data sets is significantly improved, especially in the target domain that has not been seen, the segmentation effect is better than traditional methods, and the dependence on a large amount of labeled data is reduced, and the cost of data labeling and model training is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888237B_ABST
    Figure CN119888237B_ABST
Patent Text Reader

Abstract

The present invention provides a cross-domain adaptive spinal CT image segmentation method, comprising the following steps: performing a forward diffusion process on data; constructing a model architecture; performing a reverse denoising process; and optimizing a loss function. Beneficial effects of the present invention: It can improve working performance: The segmentation accuracy between multi-source CT data sets of this framework is greatly improved, and particularly good cross-domain adaptation can be achieved on unseen target domains, and the segmentation effect is better than that of traditional methods; It can reduce production costs and energy consumption: By using a diffusion model and a contrast learning mechanism, this solution can effectively reduce the dependence on a large number of labeled data sets from other sources, and reduce the costs of data annotation and model training; It can increase stability: Through domain-guided asymmetric contrast learning, the cross-domain feature alignment effect is more stable, reducing the performance fluctuations of the model on different data sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical image processing, and particularly relates to a cross-domain adaptive spinal CT image segmentation method. Background Art

[0002] In the prior art, spinal segmentation usually relies on traditional convolutional neural network (CNN) models, such as U-Net, FCN, etc., to extract features from CT images and perform segmentation. However, these methods usually rely on a large amount of labeled data and there are significant domain differences between different datasets. Especially when the training data comes from different hospitals or different devices, the performance of existing methods often degrades. In addition, cross-domain adaptation techniques based on transfer learning have been proposed to alleviate this data bias problem, but they often have the following deficiencies:

[0003] 1. Traditional transfer learning methods are difficult to handle the domain differences of multi-source datasets.

[0004] 2. The feature alignment effect of cross-domain adaptation is poor, resulting in poor segmentation effect of target domain data.

[0005] 3. There is a lack of an efficient processing mechanism for multi-source data fusion and denoising processes in diffusion models in the prior art, resulting in unstable performance. Summary of the Invention

[0006] In view of this, the present invention aims to propose a cross-domain adaptive spinal CT image segmentation method to adopt a diffusion model to handle the domain difference problem of multi-source CT datasets.

[0007] To achieve the above object, the technical solution of the present invention is realized as follows:

[0008] A cross-domain adaptive spinal CT image segmentation method includes the following steps:

[0009] Step 1: Perform a forward diffusion process on the data;

[0010] Step 2: Construct a model architecture based on the noise data processed in Step 1;

[0011] Step 3: Perform a reverse denoising process based on the model architecture constructed in Step 2;

[0012] Step 4: Optimize the loss function for the noise data processed in Step 3;

[0013] In Step 1, performing a forward diffusion process on the data includes:

[0014] Initializing the segmentation label data:

[0015] Among them, the goal is to gradually transform multi-source domain segmentation label data into a noise distribution; the input is the original segmentation label data ; the output is a noise sequence ;

[0016] Define the noise intensity:

[0017] At each time step , set the noise intensity , to control the noise addition rate;

[0018] Gradually add Gaussian noise:

[0019] According to the properties of the Markov chain, generate noise data through the following formula:

[0020] ;

[0021] Among them, represents the scaling term that retains the previous step's data; represents the noise variance added in the current step, represents the Gaussian distribution;

[0022] The final distribution converges:

[0023] After steps, the data distribution approximates the standard Gaussian distribution:

[0024] ;

[0025] Among them, is the noise data at time step , is the noise intensity at time step , is the identity matrix.

[0026] Furthermore, in step 2, the model architecture is constructed, including:

[0027] Adopt an encoder 、an encoder 、a decoder to construct the model architecture;

[0028] Among them, the encoder processes the noise input, and the noise input includes the noise data and the time step ;

[0029] The encoder is used to extract the multi-scale features of the noise data and combine the time embedding to encode the time correlation;

[0030] Among them, the encoder Process the CT images and annotations, and for the input feature map perform

[0031] Adaptive frequency domain enhancement:

[0032] For the input feature map perform a fast Fourier transform:

[0033] ;

[0034] Wherein, is the transformed frequency domain representation, and are the frequency coordinates in the frequency domain, is the complex exponential function in the Fourier transform, is the pixel value of the input image in the spatial domain, and are the horizontal and vertical coordinates of the pixel in the spatial domain respectively, represents the number of feature map channels, represents the spatial height of the feature map, represents the spatial width of the feature map.

[0035] Furthermore, in step 3, in the reverse denoising process, the goal is to gradually recover the original segmentation label from the noisy data ; the input is the noisy data , the conditional information ; the output is the reconstructed segmentation label ; the denoising process specifically includes:

[0036] Reverse conditional probability modeling;

[0037] Construct a domain generalization module;

[0038] Perform denoising iterations.

[0039] Furthermore, the reverse conditional probability modeling includes:

[0040] For each time step , estimate the conditional distribution:

[0041] ;

[0042] Wherein: and represent the predicted mean and variance respectively, is the conditional information.

[0043] Furthermore, constructing the domain generalization module includes:

[0044] Domain feature contrast learning:

[0045] Extraction encoder features and the source domain feature library ;

[0046] Calculate the maximum mean discrepancy (MMD) to measure the domain difference:

[0047] ;

[0048] Select the most similar domain features and the least similar domain features to construct the contrastive loss:

[0049] ;

[0050] where is the contrastive learning loss, : cosine similarity; : temperature parameter;

[0051] Domain classification loss:

[0052] Optimize the domain classification task through cross-entropy:

[0053] ;

[0054] where is the multi-class cross-entropy loss function for optimizing the domain classification task, is the number of samples in the dataset, is the true label of the th sample, is the probability that the model predicts that the th sample belongs to the true class when given the input feature represents the logarithmic function for calculating the cross-entropy loss.

[0055] Furthermore, the denoising iteration includes:

[0056] Starting from , perform the following operations step by step until :

[0057] Input and to the conditional U-Net to predict and ;

[0058] Sample ;

[0059] Apply the domain generalization module to update the feature representation.

[0060] Furthermore, in step 4, the objective of loss function optimization is to jointly optimize denoising, segmentation, and domain generalization tasks; the optimization process specifically includes:

[0061] Reverse diffusion MSE loss:

[0062] Minimize the difference between the predicted noise and the true noise:

[0063] ;

[0064] Segmentation cross-entropy loss:

[0065] Optimize the final segmentation result:

[0066] ;

[0067] Total loss function:

[0068] ;

[0069] Among them, represents the balance weight, is the mean square error MSE loss, is the total number of diffusion steps, is the true distribution of the original data at the previous step given the current noise and the conditional information ; is the predicted distribution obtained by the model through learning indicating the conditional probability of predicting the previous time step from the current noise ; represents the squared error, calculating the difference between the predicted distribution and the true distribution, is the loss function of the segmentation task, represents the spatial size of the image, is the number of classes, is the image in the class true label, is the probability that the model predicts the image in the class ;

[0070] Compared with the prior art, the cross-domain adaptive spinal CT image segmentation method described in the present invention has the following advantages:

[0071] (1) For the cross-domain adaptive spinal CT image segmentation method described in the present invention, the improvement of working performance: the segmentation accuracy of this framework between multi-source CT data sets is greatly improved, especially in the unseen target domain, it can achieve better cross-domain adaptation, and the segmentation effect is better than traditional methods.

[0072] (2) For a cross - domain adaptive spinal CT image segmentation method described in the present invention, reduction of production cost and energy loss: By using a diffusion model and a contrast learning mechanism, this solution can effectively reduce the dependence on a large number of labeled data from other sources, and reduce the costs of data annotation and model training.

[0073] (3) For a cross - domain adaptive spinal CT image segmentation method described in the present invention, increase in stability: Through domain - guided asymmetric contrast learning, the cross - domain feature alignment effect is more stable, reducing the performance fluctuations of the model on different data sets. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0075] Figure 1 It is a schematic diagram of the overall method flow described in the embodiment of the present invention;

[0076] Figure 2 It is a schematic diagram of the segmentation result of the present invention described in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0077] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0078] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0079] In the description of the present invention, it should be noted that unless otherwise clearly specified and defined, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations. The present invention will be described in detail below with reference to the drawings and in combination with embodiments.

[0080] As Figures 1 to 2 shown, a cross-domain adaptive spinal CT image segmentation method includes the following steps:

[0081] Step 1: Forward diffusion process (adding noise)

[0082] Objective: Gradually transform the multi-source domain segmentation label data into a noise distribution;

[0083] Input: Original segmentation label data ;

[0084] Output: Noise sequence ;

[0085] 1. Define the noise intensity:

[0086] At each time step , set the noise intensity to control the noise addition rate.

[0087] 2. Gradually add Gaussian noise:

[0088] According to the properties of the Markov chain, generate noise data through the following formula:

[0089] ;

[0090] where : Scaling term for retaining the previous step data; : Variance of the noise added in the current step, represents the Gaussian distribution.

[0091] 3. Convergence of the final distribution:

[0092] After steps, the data distribution approaches the standard Gaussian distribution:

[0093] ;

[0094] Step 2: Model architecture construction (conditional U-Net)

[0095] Objective: Design a denoising network that supports cross-domain feature alignment and conditional reconstruction;

[0096] Components: Dual encoders and a decoder ;

[0097] 1. Encoder (Noise input processing):

[0098] Input: Noisy data and time steps ;

[0099] Function: Extract multi-scale features of the noisy data and encode the time correlation by combining with time embedding.

[0100] 2. Encoder (CT image and annotation processing):

[0101] Input: Paired CT images and segmentation labels (conditional information )

[0102] Adaptive frequency domain enhancement:

[0103] Perform a fast Fourier transform (FFT) on the input feature map :

[0104] ;

[0105] It can enhance the frequency domain features and then inverse transform them back to the spatial domain to improve the feature expression ability. Among them, is the transformed frequency domain representation, and are the frequency coordinates in the frequency domain, is the complex exponential function in the Fourier transform, is the pixel value of the input image in the spatial domain, and are the horizontal and vertical coordinates of the pixel in the spatial domain respectively, represents the number of channels of the feature map, represents the spatial height of the feature map, represents the spatial width of the feature map.

[0106] Step 3: Inverse denoising process (cross-domain generalization reconstruction)

[0107] Objective: Gradually recover the original segmentation label from the noisy data ; ;

[0108] Input: Noisy data , conditional information ;

[0109] Output: Reconstructed segmentation labels ;

[0110] 1. Inverse conditional probability modeling:

[0111] For each time step , estimate the conditional distribution:

[0112] ;

[0113] Where: Denote the mean and variance predicted by the neural network;

[0114] 2. Domain generalization module:

[0115] Domain feature contrastive learning:

[0116] Extract the features of the encoder and the source domain feature library ;

[0117] Calculate the maximum mean discrepancy (MMD) to measure the domain difference:

[0118] ;

[0119] Select the most similar domain feature and the least similar domain feature , and construct the contrastive loss:

[0120] ;

[0121] : Cosine similarity; : Temperature parameter;

[0122] Domain classification loss:

[0123] Optimize the domain classification task through cross-entropy:

[0124] ;

[0125] 3. Denoising iteration:

[0126] Starting from , perform the following operations step by step until :

[0127] Input and to the conditional U-Net, and predict and ;

[0128] Sample ;

[0129] The application domain generalization module updates the feature representation.

[0130] Step 4: Loss function optimization

[0131] Objective: Jointly optimize denoising, segmentation, and domain generalization tasks;

[0132] 1. Reverse diffusion MSE loss:

[0133] Minimize the difference between the predicted noise and the true noise:

[0134] ;

[0135] 2. Segmentation cross-entropy loss:

[0136] Optimize the final segmentation result:

[0137] ;

[0138] 3. Total loss function:

[0139] ;

[0140] Among them, represents the balance weight, is the mean square error MSE loss, is the total number of diffusion steps, is the true distribution of the original data at the previous step given the current noise and the conditional information , is the predicted distribution obtained by the model through learning , representing the conditional probability of predicting the previous time step from the current noise , represents the squared error, calculating the difference between the predicted distribution and the true distribution, is the loss function of the segmentation task, represents the spatial size of the image, is the number of classes, is the image at the class true label, is the model's prediction of the image at the class probability.

[0141] The application field of the present invention mainly relates to medical image processing, especially spinal segmentation tasks. This solution is applicable to CT image datasets from multiple source domains, aiming to solve the domain difference problem between different CT datasets and achieve the generalization ability for segmentation tasks in unseen target domains. It can be applied to fields such as medical imaging, medical image analysis, computer-aided diagnosis, etc., especially in medical image automatic segmentation and cross-domain model adaptation.

[0142] Advantages of the present invention:

[0143] Improvement in working performance: The segmentation accuracy of this framework between multi-source CT datasets is greatly improved. Especially in unseen target domains, it can achieve better cross-domain adaptation, and the segmentation effect is better than traditional methods.

[0144] Reduction in production cost and energy loss: By using the diffusion model and contrast learning mechanism, this solution can effectively reduce the dependence on a large amount of labeled data from other sources, reducing the costs of data annotation and model training.

[0145] Increase in stability: Through domain-guided asymmetric contrast learning, the cross-domain feature alignment effect is more stable, reducing the performance fluctuations of the model on different datasets.

[0146] In addition, it should be noted that:

[0147] 1. Application of the diffusion model in cross-domain adaptation

[0148] This solution utilizes the powerful representation ability of the diffusion model to recover the spinal segmentation label by gradually denoising during the reverse diffusion process. During the reverse diffusion process, a domain classifier and a domain generalization module are introduced, enabling the model not only to gradually recover the noisy data but also to effectively perform cross-domain adaptation on the source domain data, thus performing effective segmentation tasks on unseen target domains.

[0149] Forward diffusion process: By gradually adding noise, the source domain data is transformed into a noise distribution close to the standard Gaussian distribution, and the model learns the feature representation of the source domain through this noisy data during the training process.

[0150] Reverse denoising reconstruction process: The model estimates the conditional probability to gradually remove the noise and recover the original segmentation label.

[0151] In this way, the diffusion model can achieve cross-domain generalization without target domain data, enabling the segmentation model to not only focus on the source domain data during the training process but also learn the feature information shared between different source domains, thereby enhancing the cross-domain adaptation ability.

[0152] 2. Domain classifier and domain generalization module

[0153] Domain Classifier: By classifying the input features the model can identify from which source domain the data comes. This is the first step in achieving cross-domain adaptation, helping the model understand the feature differences between different source domains.

[0154] Maximum Mean Discrepancy (MMD): By using MMD to measure the difference in feature distributions between source domains, the model tries to minimize the difference in source domain feature distributions during training, making the features of different source domains have similar distributions. This helps the model learn domain-invariant feature representations, thereby improving the model's generalization ability on the target domain.

[0155] 3. Based on the condition architecture, combined with the multi-head attention mechanism and time embedding, further enhances the denoising ability.

[0156] And use the frequency domain enhancement method. In the encoder use the adaptive frequency domain enhancement method. With the help of the Fourier transform, the image is transformed from the spatial domain to the frequency domain, further refining the feature extraction, especially outstanding in dealing with high-frequency noise and retaining detail information. Through the weighted sum optimization of the low-frequency and high-frequency components, the model can effectively retain the global structure and local details of the image.

[0157] Example 1

[0158] This solution proposes a cross-domain adaptive spinal segmentation framework based on the diffusion model, aiming to solve the domain difference problem between multi-source CT datasets and achieve the generalization ability for the segmentation task of unseen target domains. The core idea is to utilize the powerful representation ability of the diffusion model. During the reverse diffusion process, a domain classifier is used to learn domain-invariant feature representations from multiple source domains. By training features at each denoising step to match the closest feature distributions between source domains, the effect of adapting to domain differences is achieved, realizing effective cross-domain generalization without target domain data.

[0159] 1. Forward Diffusion

[0160] During the forward diffusion process, the segmentation label data from multiple source domains is gradually transformed into a noise distribution by gradually adding Gaussian noise. Assume the original segmentation label data is , and gradually adding Gaussian noise generates , until the final approximates the standard Gaussian distribution. This process is represented by the following formula:

[0161] ;

[0162] where, is the data representation at time step , is the time step The noise intensity, is the identity matrix.

[0163] After steps, the original data is transformed into an approximately isotropic Gaussian distribution , expressed as:

[0164] ;

[0165] The reverse denoising reconstruction process is the core of the diffusion model and the key to achieving cross - domain generalization. By gradually removing noise to restore structural information, the generated results gradually approach the original segmentation labels. Let represent the noise image at the th step. The model reconstructs the image step by step from the previous time step by estimating the conditional probability . The conditional information encodes the paired CT image and segmentation label, while represents the current time step, guiding the denoising process at each step. This conditional probability distribution is defined as:

[0166] ;

[0167] where and represent the predicted mean and variance respectively, which are parameterized by neural networks to adapt to reconstructing structural information at different noise levels. In the reverse denoising process, at each step from to 1, a domain generalization module is introduced. Through the domain - guided asymmetric contrast learning mechanism, the model can dynamically perceive the domain characteristics of the current input data and perform feature alignment to match the feature distribution of the current domain, thus ensuring domain alignment throughout the denoising process and obtaining better segmentation performance on unseen target domains.

[0168] 2. Reverse Denoising

[0169] The denoising model is based on a conditional U - Net architecture, incorporating a multi - head attention mechanism and time embedding to enhance denoising ability and cross - domain feature alignment. The model consists of two encoders and as well as the decoder . Encoder processes the noisy input, while encoder processes the paired CT image and its segmentation label. Through their collaborative action, encoders and Extract consistent representations from multi-source domain features to achieve a powerful denoising effect. The reverse denoising reconstruction process is the core of the diffusion model and the key to achieving cross-domain generalization. By gradually removing noise to restore structural information, the generated results gradually approximate the original segmentation labels. Let denote the noisy image at the -th step. The model reconstructs the image from the previous time step by estimating the conditional probability . The conditional information encodes the paired CT image and segmentation label, while represents the current time step, guiding the denoising process at each step. This conditional probability distribution is defined as:

[0170] ;

[0171] where and represent the predicted mean and variance respectively, which are parameterized by a neural network to adapt to reconstructing structural information at different noise levels. In the reverse denoising process, at each step from to 1, a domain generalization module is introduced. Through the domain-guided asymmetric contrast learning mechanism, the model can dynamically perceive the domain characteristics of the current input data and perform feature alignment to match the feature distribution of the current domain, thereby ensuring domain alignment throughout the denoising process and obtaining better segmentation performance on unseen target domains.

[0172] The denoising model is based on a conditional U-Net architecture, combined with a multi-head attention mechanism and temporal embedding to enhance the denoising ability and cross-domain feature alignment. The model consists of two encoders and and a decoder . Encoder processes the noisy input, while encoder processes the paired CT image and its segmentation label. Through their collaborative action, encoders and extract consistent representations from multi-source domain features to achieve a powerful denoising effect. In each downsampling module of encoder , an adaptive frequency domain enhancement method is adopted to further improve the feature extraction ability. Specifically, the input feature map is first transformed to the frequency domain through a two-dimensional fast Fourier transform (FFT) to obtain the frequency domain representation :

[0173] ;

[0174] is the transformed frequency domain representation, and is the frequency coordinate in the frequency domain, represents the number of channels of the feature map, represents the spatial height of the feature map, represents the spatial width of the feature map.

[0175] is the complex exponential function in the Fourier transform, used to calculate the influence of each frequency component on the image pixels.

[0176] is the pixel value of the input image in the spatial domain, and are the horizontal and vertical coordinates of the pixel in the spatial domain respectively.

[0177] In the frequency domain, the features are decomposed into low-frequency components and high-frequency components , and the decomposition is based on the adaptive frequency segmentation function :

[0178] ;

[0179] is the weighting matrix of the high-frequency components.

[0180] is of norm, aiming to encourage sparsity and reduce irrelevant high-frequency components.

[0181] is a hyperparameter used to control the strength of the sparsity regularization.

[0182] is the gradient of the high-frequency component of norm, used to strengthen the retention of edges and details to ensure that high-frequency information is not over-smoothed. The intermediate value is a learnable hyperparameter that dynamically controls the frequency decomposition boundary.

[0183] Low-frequency component weighting matrix is generated by the Gaussian kernel to enhance the global feature stability; the high-frequency components are optimized by sparse regularization, and the weighted frequency-domain features are reconstructed back to the spatial domain by the inverse FFT to form the enhanced feature map.

[0184] 3. Domain generation module

[0185] To further enhance the model's adaptability to multi-source domain features, a domain generalization module is designed and a domain-guided asymmetric contrast learning (DG-ACL) mechanism is introduced. This mechanism improves the generalization ability of cross-domain features by dynamically mapping input features to the most similar known multi-source domain feature representations and ensures domain consistency during the reverse diffusion process. In the encoder The output features of each layer are processed by adaptive average pooling (AAP) to obtain a global low-dimensional feature representation which captures the global distribution information of the multi-source domain. These aggregated features are then fed into a domain classifier consisting of a fully connected layer, an MLP layer, and a Softmax layer. The role of the domain classifier is to explicitly partition the feature distributions of the multi-source domain and calculate the probability distribution of the input features belonging to each domain:

[0186] ;

[0187] where is the probability that the input feature belongs to the th class given the feature representation . That is, it represents the probability that the input feature belongs to a certain domain.

[0188] are the features obtained after the encoder processes, representing the global features of the input data. These features are processed by adaptive pooling (AAP) and converted into a lower-dimensional global representation where is the dimension of the feature. is a function processed by a multi-layer perceptron (MLP) network to extract the feature representation related to the domain . Here are the weights in the MLP. is the exponential function used to transform the features so that the output of the classifier conforms to the probability distribution. is the normalization term to ensure that the sum of probabilities for all classes is 1. is the number of source domains, i.e., the total number of possible classes. Among them, is the number of source domains, is a function parameterized by the MLP layer to extract the feature representation of the domain .

[0189] To further reduce the differences in feature distributions between source domains and enhance cross-domain generalization ability, the maximum mean discrepancy (MMD) is introduced to measure the difference in feature distributions between domains:

[0190] ;

[0191] in:

[0192] Represents the Maximum Mean Discrepancy metric, which is used to measure the difference between two distributions. Here, it is used to measure the difference in feature distribution between source domains.

[0193] and is the distribution of the two source domains.

[0194] and They are respectively distributed from and Features sampled in .

[0195] and It is the feature representation calculated by the MLP network.

[0196] Represents expectation and calculates the mean of sample features.

[0197] is the Hopf distance (H-space distance), which is used to measure the difference between two distributions.

[0198] Select the domain feature with the highest similarity As a positive sample, the lowest As negative samples, construct contrastive learning loss:

[0199] ;

[0200] It is a contrastive learning loss, which is used to narrow the feature representation distance of similar domains and push the feature representation distance of other domains away.

[0201] is the input feature and the most similar domain features The similarity measure between them is usually cosine similarity or Euclidean distance. Here is the domain feature that is most similar to the input feature.

[0202] is an exponential function used to convert similarity values ​​into probabilities.

[0203] is the temperature parameter that controls the smoothness of the similarity metric. This will make the similarity difference more significant.

[0204] is a normalization term that ensures the sum of probabilities of all domain features is 1.

[0205] The contrastive learning loss enhances the model's discrimination ability for different domain features by maximizing the distance between positive samples (the most similar domain features) and minimizing the distance between negative samples (other domain features).

[0206] Through this mechanism, the distance between the input features and the most similar domain is reduced, while the distances to other domain features are increased, enhancing the cross-domain feature discrimination ability.

[0207] 4. Model Training

[0208] During the model training process, the domain generalization module uses the multi-class cross-entropy loss to optimize domain classification:

[0209] ;

[0210] where:

[0211] is the multi-class cross-entropy loss function used to optimize the domain classification task.

[0212] is the number of samples in the dataset.

[0213] is the true label (target domain class label) of the

[0214] th sample. is the probability that the model predicts the th sample belongs to the true class when given the input feature Here, is the feature representation extracted by the encoder

[0215] represents the logarithmic function for calculating the cross-entropy loss.

[0216] Combined with the contrastive loss, the total domain generalization loss is obtained:

[0217] ;

[0218] is the total domain generalization loss, which combines the domain classification loss and the contrastive learning loss

[0219] is the domain classification loss, which has been explained above.

[0220] is the contrastive learning loss, aiming to enhance the cross-domain feature discrimination ability of the model by minimizing the distance between the input features and the most similar domain.

[0221] During the reverse diffusion process, the model predicts the forward noise through the mean squared error (MSE) loss, and the optimization objective is:

[0222] ;

[0223] is the mean squared error (MSE) loss, used to train the model to predict the forward noise through the reverse diffusion process.

[0224] is the total number of diffusion steps.

[0225] is the given current noise and conditional information when, the true distribution of the original data at the previous step.

[0226] is the predicted distribution learned by the model, representing the conditional probability of predicting the previous time step from the current noise .

[0227] represents the squared error, calculating the difference between the predicted distribution and the true distribution.

[0228] The segmentation task uses the multi-class cross-entropy loss:

[0229] ;

[0230] is the loss function for the segmentation task, usually using the multi-class cross-entropy loss to optimize the segmentation results.

[0231] represents the spatial size of the image (the product of height and width).

[0232] is the number of classes.

[0233] is the image in the class true label (such as the label of whether the spine exists).

[0234] is the model's prediction of the image in the class probability.

[0235] such as Figure 2This is the segmentation result diagram of the present invention. The present invention uses a diffusion model, which is an innovative part in the prior art. The prior art more relies on traditional CNN models or transfer learning methods, while the diffusion model provides a more effective cross-domain adaptation ability through the reverse denoising process.

[0236] The present invention introduces an asymmetric contrast learning mechanism. By dynamically perceiving the domain characteristics of the input data and optimizing feature alignment, the performance of cross-domain segmentation is significantly improved.

[0237] Compared with the prior art, the present invention can better handle the differences between multiple source domains and can still achieve excellent cross-domain generalization effects without target domain data.

[0238] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A cross-domain adaptive spine CT image segmentation method, characterized by: The following steps are involved: Step 1: Perform forward diffusion process on the data; Step 2: Build the model architecture based on the noise data processed in step 1; Step 3: Perform reverse denoising process based on the model architecture built in step 2; Step 4: Optimize the loss function of the noise data processed in step 3; In step 1, the data is forward diffused, including: Initialize segmentation label data: The goal is to gradually transform the multi-source domain segmentation label data into a noise distribution; the input is the original segmentation label data x0; the output is the noise sequence x1, x2, ..., x T ; Define the noise intensity: At each time step t∈{1,2,…,T}, set the noise intensity β t ∈(0,1), controls the noise adding rate; gradually adds Gaussian noise: According to the Markov chain properties, noise data is generated by the following formula: in, Represents the scaling term that retains the previous step data; β t I represents the noise variance added in the current step, represents Gaussian distribution; The final distribution converges to: After T steps, the data distribution approaches the standard Gaussian distribution: Among them, x t is the noise data at time step t, β t ∈(0,1) is the noise intensity at time step t, I is the identity matrix; In step 2, the model architecture is built, including: The model architecture is constructed using encoder E0, encoder E1, and decoder D0; The encoder E0 processes the noise input, which includes the noise data x t and time step t; Encoder E0 is used to extract multi-scale features of noise data and combine time embedding to encode time correlation; among them, encoder E1 processes CT images and annotations and inputs feature maps Perform adaptive frequency domain enhancement: For the input feature map Perform a fast Fourier transform: Among them, F(u,v) is the frequency domain representation after conversion, u and v are the frequency coordinates in the frequency domain, It is the complex exponential function in Fourier transform, X(x,y) is the pixel value of the input image in the spatial domain, x and y are the horizontal and vertical coordinates of the pixel in the spatial domain, C represents the number of feature map channels, H represents the spatial height of the feature map, and W represents the spatial width of the feature map; In step 3, the reverse denoising process is performed, with the goal of T Gradually restore the original segmentation label x0; the input is the noise data x T , condition information c ct ; The output is the reconstructed segmentation label x0; The denoising process specifically includes: Inverse conditional probability modeling; Build a domain generalization module; Perform denoising iterations; Inverse conditional probability modeling, including: For each time step t=T,T-1,…,1, estimate the conditional distribution: Where: μ θ (x t ,c ct ,t) and They represent the mean and variance of the prediction, c ct For conditional information; build a domain generalization module, including: Domain Feature Contrastive Learning: Extract the features of encoder E0 and source domain feature library Calculate the maximum mean difference (MMD) to measure the difference between domains: Select the most similar domain feature Z max and the most dissimilar domain feature Z min , construct the contrast loss: in, is the contrastive learning loss, S(·): cosine similarity; τ: temperature parameter; Domain classification loss: Optimizing domain classification tasks via cross entropy: in, is a multi-class cross-extraction loss function used to optimize domain classification tasks, N is the number of samples in the dataset, and y i is the true label of the i-th sample, is the model's prediction given the input features When the i-th sample belongs to the true category y i The probability of , log represents the logarithmic function, and the cross-extraction loss is calculated; Denoising iterations, including: From x T Initially, the following operations are performed step by step until t=1: Input x t and c ct To conditional U-Net, predict μ θ and σ θ ; sampling The domain generalization module is applied to update the feature representation.

2. The cross-domain adaptive spine CT image segmentation method according to claim 1, characterized in that: In step 4, the goal of loss function optimization is to jointly optimize denoising, segmentation and domain generalization tasks; the optimization process specifically includes: Backward diffusion MSE loss: Minimize the difference between predicted noise and true noise: Segmentation cross entropy loss: Optimize the final segmentation result: Total loss function: Among them, λ1,λ2 represent the balance weights, is the mean square error MSE loss, T is the total number of diffusion steps, q(x t-1 ∣x t ,c t ) is given the current noise x t and condition information c t When the original data is actually distributed in the previous step, p(x t-1 ∣x t ,c ct ) is the predicted distribution obtained by the model through learning, which means that from the current noise x t Predict the previous time step x t-1 The conditional probability of ||·|| 2 represents the squared error, which calculates the difference between the predicted distribution and the true distribution. is the loss function of the segmentation task, H×W represents the spatial size of the image, K is the number of categories, and y label,k is the image x i The true label on category k, p(y label,k ∣x i ) is the model prediction image x i The probability of being in class k.

Citation Information

Patent Citations

  • Hyperspectral image multi-source domain adaptive classification method based on diffusion model

    CN118247668A

  • Image restoration method of enhanced conditional diffusion model based on dual-domain interactive Transform

    CN118537249A