A robust medical image segmentation method based on multi-modal diffusion model
Patent Information
- Application Number
- CN202410256193.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-06
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2044-03-06
AI Technical Summary
评估者之间的诊断水平和经验的不同会导致评估者间的差异,而同一评估者在不同时间对同一图像区域进行分割时,也会出现内部差异
Smart Images

Figure CN118038052B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence technology, specifically relating to a differential medical image segmentation method based on a multimodal diffusion model. Background Technology
[0002] Medical image segmentation is a crucial tool for diagnosing diseases and assessing tumor boundaries. Currently, deep learning-based medical image segmentation methods often incorporate linguistic modalities to improve accuracy. However, these methods rely solely on a single evaluator's interpretation of multimodal data, introducing individual evaluator bias. In clinical practice, multiple evaluators are typically employed to perform image segmentation collaboratively, reducing bias introduced by personal preferences and mitigating the impact of individual variability. While multi-evaluator learning strategies show promise in reducing segmentation errors, their application in multimodal medical image segmentation has been limited due to inconsistencies between and within evaluators. Differences in diagnostic skill and experience among evaluators lead to inter-evaluator variability, and even internal variability can occur when the same evaluator segments the same image region at different times. Summary of the Invention
[0003] To overcome the shortcomings of existing technologies, this invention proposes a differential medical image segmentation method based on a multimodal diffusion model. This method achieves lesion segmentation in medical images. The specific steps are as follows:
[0004] Step 1: Preprocessing of experimental data; preprocessing of the collected medical image data.
[0005] Step 2: Construct the Medical Image Segmentation Network MMDSN;
[0006] Step 3: Train the MMDSN network model;
[0007] Step 4: Use the trained MMDSN network model for multi-evaluator inference.
[0008] Step 1 specifically includes the following steps:
[0009] Step 1.1: Crop the medical image;
[0010] Step 1.2 Perform data augmentation on the cropped image;
[0011] Step 1.3 Divide the dataset into training set, validation set, and test set.
[0012] Step 2 includes the following steps:
[0013] Step 2.1 Construct a text encoder to extract semantic information from the input text;
[0014] For the input text information, we first perform word segmentation to obtain a text vector T. Then, the text vector T is processed by a text encoder to extract features and obtain a high-dimensional representation of the text. The specific structure of the text encoder is as follows:
[0015] First, the text vector T obtained after word segmentation is converted into a fixed-size vector through the embedding layer. At the same time, each vector is added with position encoding to obtain the embedding vector.
[0016] Furthermore, the embedding vector passes through a Transformer layer, where each token pays attention to all other tokens in the input sequence and calculates a weighted sum, with the weights reflecting the importance of other tokens to the current token.
[0017] Furthermore, the output of the Transformer layer is normalized to obtain a high-dimensional representation of the text. This indicates that semantic information in the text has been captured.
[0018] Step 2.2: Add noise to the image forward;
[0019] The input medical image mask x0 is perturbed by Gaussian random noise, and noise is added iteratively to blur and make the data samples unclear and uncertain. The noise addition process is as follows:
[0020]
[0021] Where β t It is used to adjust the variance of Gaussian noise, where I is the identity matrix, t is the time step, and x is the variance. t This is the image after adding noise for t time steps using the original image mask x0. Alternatively, the noisy image at time step t can be obtained directly using x0, as shown below:
[0022]
[0023]
[0024] Where α t =1-β t ,and It is standard Gaussian random noise.
[0025] Step 2.3 Construct the image feature extraction head;
[0026] Given an input medical image X and a noisy image mask x at t time steps. t First, the image is concatenated along the channel dimension. Then, the concatenated image is fed into the image feature extraction head to obtain Y0. The image feature extraction head consists of convolutional blocks.
[0027] Step 2.4 Construct a U-shaped visual Transformer branch, which consists of four Transformer encoder layers and four Transformer decoder layers;
[0028] Specifically, image features Y0 and text features The input image Y0 is fed into a first-layer visual Transformer encoder, which consists of visual Transformers. First, the input image Y0 is divided into N equal-sized image patches. Then, each image patch is flattened and transformed into a D-dimensional vector through a linear layer. Y0 is then coupled with positional encodings and text features. The values are summed, and then self-attention is calculated. After layer normalization, the output of each encoder layer is obtained, denoted as Y. 1_down ,Y 2_down ,Y 3_down ,Y 4_down .
[0029] Furthermore, the output feature Y of the fourth layer encoder 4_down It will sequentially pass through the fourth, third, second, and first decoder layers. Each decoder consists of bilinear interpolation layers and convolutional layers. Each decoder layer will have output features, denoted as...
[0030] Step 2.5 Construct the U-shaped network branches;
[0031] This branch consists of a four-layer encoder, decoder, and vision graph module, used to predict the segmentation mask at the current time step t. Specifically, the input image feature Y0 first passes through four encoder layers, each consisting of residual blocks and downsampling blocks. The residual block structure is as follows:
[0032] First, for the input temporal embedding t, i.e., the number of noise-adding steps, it first passes through the SiLU activation function and a linear layer to obtain the temporal vector. The input image feature Y0 passes through a group normalization layer, a SiLU activation function layer, and a convolutional layer to obtain the output feature. Then, the output feature is added to the temporal vector and passed through another group normalization layer, a SiLU activation function layer, and a convolutional layer to obtain the final output feature. Finally, the feature is downsampled to obtain the output feature of the encoder layer. Each encoder layer has an output feature, denoted as Z. 1_down Z 2_down Z 3_down Z 4_down .
[0033] U-shaped network branch fourth layer encoder output characteristics Output features of the fourth layer decoder of the U-shaped visual Transformer Before being fed into the fourth layer decoder of the U-shaped network branch, the feature is fed into the visual language graph module for feature fusion. Here, C and N are the dimensional representations of the features. The specific structure of the visual language graph module is as follows:
[0034] First, calculate Z. 4_down and The affinity matrix between them is represented as follows:
[0035]
[0036] in It is a learnable weight matrix. It is an affinity matrix.
[0037] Furthermore, regarding the affinity matrix Standardize, then Z 4_down and Feature extraction is performed using a graph convolutional neural network, as shown below:
[0038]
[0039]
[0040] Here, `concat` is a concatenation operation along the channel dimension, and `GCN` is a graph convolutional neural network. Then Z... 4_new With Y 4_new Next, the channel dimensions are concatenated, and the concatenated features are fed into the fourth layer decoder of the U-shaped network branch to obtain... Then and The image is then fed back into the visual language graph module for feature fusion. The resulting fused features are then fed into the third layer decoder of the U-shaped network branch. After four decoding operations, the original image mask predicted at time step t is obtained. The decoder consists of residual blocks and upsampling layers.
[0041] Step 2.6 Latent Gaussian distribution modeling: We will predict the image mask at time step t. After being stitched together with the medical image X along the channel dimension, it is fed into a prior distribution mapping function f. θ This function maps features to a Gaussian distribution with a mean of 1 / 2. variance is It is expressed as follows:
[0042]
[0043] Where z q It is a prior Gaussian distribution, fθ It is the prior distribution mapping function, which consists of convolutional layers and generates a prior Gaussian distribution.
[0044] Furthermore, we concatenate the original image mask x0 with the medical image X along the channel dimension, and then feed the concatenation into a posterior distribution mapping function f. η This function maps the features to a Gaussian distribution with mean μ(x0,X; f). η )∈R N The variance is σ(x0,X; f η )∈R N×N This is expressed as follows:
[0045]
[0046] Where z p It is a posterior Gaussian distribution, f η It is the posterior distribution mapping function, which consists of convolutional layers and generates a posterior Gaussian distribution.
[0047] Step 3 includes the following steps:
[0048] Step 3.1 Calculate the loss function of MMDSN. The first loss function is the predicted mask. The mean square error between the actual mask x0 and the real mask x0 is expressed as follows:
[0049]
[0050] Where x0 is the actual mask. It is the mask predicted at the t-th time step.
[0051] The second loss function is the variational lower bound loss of the diffusion model, expressed as follows:
[0052]
[0053]
[0054]
[0055]
[0056] in It is the total variational lower bound loss. It is the variational lower bound loss at the t-th time step. and It is the variational lower bound loss for the initial and last time steps, D. KL This is the KL divergence calculation.
[0057] The third loss function is the loss function for modeling the latent Gaussian distribution, expressed as follows:
[0058]
[0059] Where D KL It is the KL divergence calculation function, z p It is a posterior latent Gaussian distribution, z q It is a prior latent Gaussian distribution. It is the loss function for modeling the potential Gaussian distribution.
[0060] The final loss function is obtained by adding the three loss functions:
[0061]
[0062] in It is the final loss function.
[0063] Step 3.2 Use the AdamW optimizer during training;
[0064] Step 4 includes the following steps:
[0065] Step 4.1 Sampling of evaluator distribution;
[0066] When multiple evaluators segment the same image, differences exist among them due to variations in experience or skill level. We sample M distributions from a random distribution to simulate the differences among the M evaluators, as shown below:
[0067]
[0068] Where r represents the index of the rater, and q(g|r) is the distribution of the r-th rater sampled.
[0069] Step 4.2 Output the prediction mask for each evaluator;
[0070] We feed the sampled random noise q(g|r) and the original image X into the MMDSN network to iteratively predict the segmentation mask. The evaluator generates a segmentation mask at each time step. Ultimately, the mask generated by each evaluator is obtained by weighting the predicted mask exponents at each time step, as shown below:
[0071]
[0072] in It is the original image mask predicted at time step t, and α is the weight. It is the segmentation result of the kth evaluator.
[0073] Step 4.3 Merge the prediction masks from all evaluators
[0074] The segmentation results from M evaluators are processed by a multi-evaluator consensus module to obtain a final, unique segmentation mask. The multi-evaluator consensus module is represented as follows:
[0075]
[0076]
[0077] in It is the segmentation result of evaluator k at position (i,j), and S is the threshold. It is the final segmentation mask that summarizes the results from M evaluators. Attached Figure Description
[0078] Figure 1 This is a network structure diagram of MMDSN.
[0079] Figure 2 This demonstrates the practical application effectiveness of MMDSN. Detailed Implementation
[0080] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0081] To address the challenges of medical image segmentation, we propose a differentially resistant medical image segmentation method based on a multimodal diffusion model. Specifically, this network utilizes deep learning techniques and a multimodal diffusion model to accurately segment lesions. First, we introduce a text encoder to extract semantic information from the text. Then, we introduce a visual language graph module for multimodal feature extraction and fusion. Next, we introduce a latent Gaussian distribution to model and constrain the discrepancies between evaluators. Finally, the prediction results from each evaluator at multiple time steps are fed into a multi-evaluator consensus module to obtain the final prediction mask.
[0082] Example 1: Preprocessing of experimental data.
[0083] (1) Cropping medical images.
[0084] (2) Perform data augmentation on the cropped image.
[0085] (3) Divide the dataset into training set, validation set and test set.
[0086] Example 2: Constructing the MMDSN network model.
[0087] (1) Construct a text encoder to extract semantic information from the input text.
[0088] (2) Image forward noise addition: Gaussian random noise perturbation is applied to the input medical image mask x0.
[0089] (3) Construct an image feature extraction head.
[0090] (4) Construct a U-shaped visual Transformer branch, which consists of a four-layer Transformer encoder and a four-layer Transformer decoder.
[0091] (5) Construct a U-shaped network branch, which consists of a four-layer encoder, decoder and visual language graph module.
[0092] (6) Model the potential Gaussian distribution.
[0093] Example 3: Training the MMDSN network model.
[0094] (1) Calculate the loss function of MMDSN. The loss function of MMDSN is composed of MSE, variational lower bound and latent Gaussian distribution modeling.
[0095] (2) The AdamW optimizer was used to optimize MMDSN.
[0096] Example 4 utilizes a trained MMDSN network model for multi-evaluator inference.
[0097] (1) Evaluation of the distribution of the evaluators.
[0098] (2) Output the prediction mask for each evaluator.
[0099] (3) Integrate the prediction masks of all evaluators.
Claims
1. A differentially resistant medical image segmentation method based on a multimodal diffusion model, characterized in that... Includes the following steps: Step 1: Preprocessing of experimental data; preprocessing of the collected medical image data. Step 2: Construct the Medical Image Segmentation Network MMDSN; Step 2.1 Construct a text encoder to extract semantic information from the input text; Step 2.2 Forward noise addition to the image, using the input medical image mask. The data will be disturbed by Gaussian random noise, and the noise will be added iteratively, making the data samples blurry and uncertain. The noise addition process is as follows: in It is used to adjust the variance of Gaussian noise. It is the identity matrix. It is a time step. It is the original image mask. Add noise Images after one time step; Step 2.3 Constructing an image feature extraction head for the input medical image. and noise Noisy image mask at each time step First, the image is stitched together along the channel dimension. Then, the stitched image is fed into an image feature extraction head to obtain the desired result. The image feature extraction head consists of convolutional blocks; Step 2.4 Construct a U-shaped visual Transformer branch, which consists of four Transformer encoder layers and four Transformer decoder layers; Step 2.5 Construct a U-shaped network branch, which consists of four Transformer encoder layers and four Transformer decoder layers; Step 2.6 Modeling the latent Gaussian distribution; Step 3: Train the MMDSN network model; Step 4: Use the trained MMDSN network model for multi-evaluator inference.
2. The anti-difference medical image segmentation method based on a multimodal diffusion model according to claim 1, characterized in that... Step 2.5 is implemented as follows: The U-shaped network branch consists of a four-layer encoder, decoder, and visual graph module, used to predict the current i-th... Segmentation mask at each time step Specifically, the input image features It will first pass through four encoder layers, each consisting of residual blocks and downsampling blocks; the structure of the residual block is as follows: Firstly, regarding the time embedding of the input... The noise-adding steps, i.e., the time vector, are obtained by first passing through the SiLU activation function and a linear layer; the input image features... The output features are obtained by sequentially passing through a group normalization layer, a SiLU activation function layer, and a convolutional layer. Then, the output features are added to the time vector and passed through another group normalization layer, a SiLU activation function layer, and a convolutional layer to obtain the final output features. final The features are then downsampled to obtain the output features of the encoder layer. Each encoder layer will have output features, denoted as... ; U-shaped network branch fourth layer encoder output characteristics Output features of the fourth layer decoder of the U-shaped visual Transformer Before being fed into the fourth layer decoder of the U-shaped network branch, the feature is fed into the visual language graph module for feature fusion. Here, C and N are the dimensional representations of the features. The specific structure of the visual language graph module is as follows: First calculate and The affinity matrix between them is represented as follows: in It is a learnable weight matrix. It is an affinity matrix; Furthermore, regarding the affinity matrix Standardize, and then and Feature extraction is performed using a graph convolutional neural network, as shown below: in It is a splicing operation at the channel level. It is a graph convolutional neural network; then and Next, the channel dimensions are concatenated, and the concatenated features are fed into the fourth layer decoder of the U-shaped network branch to obtain... ,Then and The data is then fed back into the visual language graph module for feature fusion. The resulting fused features are then fed into the third layer decoder of the U-shaped network branch. After four decoding operations, the final result is obtained. Original image mask predicted at each time step ; The decoder consists of residual blocks and upsampling layers.
3. The anti-difference medical image segmentation method based on a multimodal diffusion model according to claim 1, characterized in that... Step 2.6 is implemented as follows: The first Image mask predicted at each time step With medical imaging After concatenating the channels, the data is fed into a prior distribution mapping function. This function maps features to a Gaussian distribution with a mean of 1 / 2. The variance is ; indicates the following: in It is a priori Gaussian distribution. It is the prior distribution mapping function, which consists of convolutional layers and generates a prior Gaussian distribution; Furthermore, mask the actual original image. With medical imaging After concatenating the channels, the data is fed into a posterior distribution mapping function. This function maps features to a Gaussian distribution with a mean of 1 / 2. The variance is ; indicates the following: in It is a posterior Gaussian distribution. It is the posterior distribution mapping function, which consists of convolutional layers and generates a posterior Gaussian distribution.
4. The anti-difference medical image segmentation method based on a multimodal diffusion model according to claim 1, characterized in that... Step 3 includes the following steps: Step 3.1 Calculate the loss function of MMDSN. The first loss function is the predicted mask. and the real mask The mean square error between them is expressed as follows: in It is the real mask. It is the mask predicted at the t-th time step. It is the mean squared error loss function; The second loss function is the variational lower bound loss of the diffusion model, expressed as follows: in It is the total variational lower bound loss. It is the variational lower bound loss at the t-th time step; and It is the variational lower bound loss for the initial and last time steps. This is the KL divergence calculation; The third loss function is the loss function for modeling the latent Gaussian distribution, expressed as follows: in It is the KL divergence calculation function. It is a posterior latent Gaussian distribution. It is a prior latent Gaussian distribution. It is the loss function for modeling the latent Gaussian distribution; The final loss function is obtained by adding the three loss functions: in It is the final loss function; Step 3.2 Use the AdamW optimizer during training.
5. The anti-difference medical image segmentation method based on a multimodal diffusion model according to claim 1, characterized in that... Step 4 includes the following steps: Step 4.1 Sampling of evaluator distribution; When multiple evaluators segment the same image, differences exist among them due to variations in experience or skill level; sampling from a random distribution... A distribution, simulation The differences among the evaluators are represented as follows: in The serial number representing the rater It is the sampling number Distribution of raters; Step 4.2 Output the prediction mask for each evaluator; Sampled random noise and original images The data is fed into an MMDSN network to iteratively predict the segmentation mask, and the evaluator generates a segmentation mask at each time step. Finally, the mask generated by each evaluator is obtained by weighting the predicted mask exponents at each time step, as shown below: in It is the first The original image mask predicted at each time step, It's weight. It is the segmentation result of the kth evaluator; Step 4.3 Merge the prediction masks from all evaluators The segmentation results from each evaluator are processed by a multi-evaluator consensus module to obtain a final, unique segmentation mask. The multi-evaluator consensus module is represented as follows: in It is the segmentation result of evaluator k at position (i,j). It is a threshold. It is a summary The final segmentation mask for each evaluator.