Cross-modal ultrasonic image segmentation method based on conditional diffusion model and transfer learning
Through the method based on conditional diffusion model and transfer learning, the problems of unclear boundaries, poor generalization ability, insufficient data and cross-modal adaptation in ultrasonic image segmentation are solved, and the ultrasonic image segmentation effect with high precision, fast and clear boundaries are achieved.
Patent Information
- Application Number
- CN202510370540.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-27
AI Technical Summary
In the automatic segmentation of ultrasound images of brain and thoracic abdomen, there are problems such as unclear boundary segmentation, poor generalization ability of model, insufficient standard data, and cross-modal data adaptation, which makes the segmentation task difficult.
A cross-modal ultrasound image segmentation method based on conditional diffusion model and transfer learning is adopted to generate high-quality labeled three-dimensional simulated ultrasound data sets through conditional diffusion model, and a hybrid attention segmentation model is trained in combination with transfer learning technology to improve the adaptability and segmentation accuracy of the model.
It achieves the effect of high accuracy, fast speed and clear boundaries on the segmentation of lesions in the three-dimensional ultrasound data set, solves the problems of scarcity of data and cross-modal adaptation, and improves the generalization and segmentation performance of the model.
Smart Images

Figure CN120163984A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and particularly to a cross-modal ultrasound image segmentation method based on conditional diffusion models and transfer learning. Background Art
[0002] As a non-invasive imaging technology, ultrasound images are widely used in the diagnosis of various diseases such as brain tumors, pneumonia, and pleural effusion in multiple parts including the brain, chest, and abdomen due to their advantages of real-time, high safety, and low cost. However, the quality of ultrasound images is limited by imaging equipment, patient position, and the operator's skills, and usually has problems such as high noise, low contrast, and unclear anatomical structures, resulting in a difficult segmentation task for ultrasound images in the diagnosis of brain and chest and abdominal diseases. Ultrasound image segmentation can not only accurately locate the lesion area but also provide important information about the shape, size, and function of organs, assisting in the early diagnosis and treatment decision-making of diseases.
[0003] In recent years, deep learning techniques, especially semantic segmentation algorithms based on convolutional neural networks (CNNs), have been widely applied in medical image segmentation tasks of modalities such as MRI (Magnetic Resonance Imaging), CT (Computed Tomography), and US (Ultrasound), and significant progress has been made. However, there are still the following major problems in the automatic segmentation of brain and chest and abdominal ultrasound images: (1) unclear boundary segmentation: The noise in ultrasound images is relatively large, and the tissue boundaries are not obvious, especially in the imaging of deep structures in the chest and abdomen or the brain, which makes it difficult for traditional image processing and segmentation algorithms to accurately segment the boundaries of the target area. (2) Poor model generalization ability: The anatomical structures of the brain and chest and abdomen are complex, and there are significant individual differences among patients. Especially in different pathological states, the shape and size of organs may change significantly. Such diversity requires that the ultrasound image segmentation model has strong adaptability and generalization ability. (3) Insufficient standard data: Since the annotation of ultrasound images requires expert manual annotation, and the ultrasound image sections and angles of each patient are different, it results in scarce training data and incomplete annotation, further increasing the difficulty of the segmentation task. (4) Cross-modal data adaptation problem: Brain and chest and abdominal ultrasound images usually have significant differences from other modality medical image data such as CT and MRI. Existing deep learning models based on a single modality are difficult to directly extract consistent features from images of different modalities, which limits the joint analysis and utilization of multi-modal images.
[0004] In response to the above challenges, diffusion models and transfer learning techniques have shown potential in the field of ultrasound image segmentation in recent years. Diffusion models effectively suppress noise interference and improve the accuracy of image boundary segmentation by gradually reconstructing or generating high-quality feature distributions. At the same time, transfer learning can significantly alleviate the problem of scarce labeled data by fine-tuning the target task using large-scale pre-trained models.
[0005] Therefore, developing an ultrasound image segmentation algorithm based on diffusion models and transfer learning will be able to effectively expand the labeled dataset, improve the model's adaptability to complex anatomical structures and multi-modal features, and achieve precise and clearly bounded segmentation of ultrasound images, which is of great significance for the diagnosis and treatment of diseases in the brain, chest, abdomen, etc. Summary of the Invention
[0006] Aiming at the deficiencies of the prior art, the technical problem to be solved by the present invention is to provide a cross-modal ultrasound image segmentation method based on conditional diffusion models and transfer learning.
[0007] The technical solution of the present invention to solve the above technical problem is to provide a cross-modal ultrasound image segmentation method based on conditional diffusion models and transfer learning, characterized in that the method comprises the following steps:
[0008] Step 1: Obtain a three-dimensional original modality dataset and a three-dimensional target modality dataset of the same organ or lesion, perform data preprocessing to adapt to the inputs of the conditional diffusion model and the hybrid attention segmentation model, obtain the preprocessed dataset, and divide the preprocessed dataset into a three-dimensional original modality training set, a three-dimensional original modality validation set, a three-dimensional original modality test set, a three-dimensional target modality training set, a three-dimensional target modality validation set, and a three-dimensional target modality test set; the three-dimensional target modality dataset is a three-dimensional real ultrasound dataset composed of three-dimensional real ultrasound images;
[0009] Step 2: Construct a conditional diffusion model, set training parameters based on the three-dimensional original modality training set and the three-dimensional target modality training set obtained in Step 1, and use dynamic learning rate adjustment and noise regulation to fully train the conditional diffusion model. After training, obtain the trained conditional diffusion model and the weight parameters of the conditional denoising network TU-Net in the trained conditional diffusion model;
[0010] Step 3: Use the trained conditional diffusion model in Step 2 to infer each three-dimensional original modality image in the preprocessed three-dimensional original modality dataset obtained in Step 1, batch translate it into a labeled three-dimensional simulated ultrasound image, and then form a labeled three-dimensional simulated ultrasound dataset, and divide the labeled three-dimensional simulated ultrasound dataset into a training set and a test set;
[0011] Step 4: Transfer the weight parameters of the conditional denoising network TU-Net in the trained conditional diffusion model obtained in Step 2 to the hybrid attention segmentation network, and use the training set in Step 3 to pre-train the hybrid attention segmentation network to obtain the pre-trained hybrid attention segmentation network;
[0012] Step 5: Use the method of transfer learning to fine-tune the parameters of the pre-trained hybrid attention segmentation network obtained in Step 4 with the three-dimensional target modality training set obtained in Step 1 to adapt to the three-dimensional real ultrasound data. Take the fine-tuned model with the best performance on the three-dimensional target modality validation set in Step 1 after the fine-tuning training as the hybrid attention segmentation model; then use the hybrid attention segmentation model to process the three-dimensional target modality test set in Step 1 to obtain the segmentation result.
[0013] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0014] (1) The present invention uses a conditional diffusion model for cross-modal translation to generate a high-quality labeled three-dimensional simulated ultrasound data set. At the same time, the transfer technology is used to combine the generated three-dimensional simulated ultrasound data set with the conditional diffusion model and the hybrid attention segmentation model to achieve accurate, fast, and efficient training of the hybrid attention segmentation model, and the lesion segmentation in the three-dimensional ultrasound data set reaches the effects of high accuracy, fast speed, and clear boundary.
[0015] (2) The present invention can directly perform cross-modal image translation on three-dimensional medical images. By combining multiple modal data sets, the conditional diffusion model can directly translate multiple three-dimensional original modal images into three-dimensional simulated ultrasound images.
[0016] (3) The three-dimensional simulated ultrasound data set generated by the conditional diffusion model of the present invention provides a large amount of pre-training data for the downstream segmentation task, solves the problems of overfitting and poor effect in model training caused by insufficient three-dimensional ultrasound images. At the same time, it effectively alleviates the differences between multi-modal data, making the model use the effect of multi-modal data transfer better.
[0017] (4) By migrating the weights of the conditional diffusion model to the hybrid attention segmentation model, the present invention establishes a connection between the conditional diffusion model and the semantic segmentation model. Through shared encoder feature extraction, the transfer learning of the conditional diffusion model to the hybrid attention segmentation model is further realized, improving the training speed and segmentation accuracy of the hybrid attention segmentation model.
[0018] (5) Through the fine-tuning strategy, the present invention locally fine-tunes the hybrid attention segmentation model on the three-dimensional real ultrasound data set, fully improving the generalization of the model, so that the model can be applied to various three-dimensional ultrasound image segmentations. Description of the Drawings
[0019] Figure 1 is the overall flowchart of the present invention;
[0020] Figure 2 is the schematic diagram of the conditional diffusion model of the present invention;
[0021] Figure 3 is the schematic diagram of the structure of the conditional denoising network TU-Net of the present invention;
[0022] Figure 4 is the schematic diagram of the principle of the segmentation method of the present invention;
[0023] Figure 5 is the original brain ultrasound image of an embodiment of the present invention;
[0024] Figure 6 is of the present invention Figure 5 brain tumor segmentation result diagram;
[0025] Figure 7 is the original brain ultrasound image of another embodiment of the present invention;
[0026] Figure 8 is of the present invention Figure 7 brain tumor segmentation result diagram. Specific Embodiments
[0027] The following gives specific embodiments of the present invention. The specific embodiments are only used to further illustrate the present invention in detail and do not limit the protection scope of the present invention.
[0028] The present invention provides a cross-modal ultrasound image segmentation method based on a conditional diffusion model and transfer learning (hereinafter referred to as the method), and the method includes the following steps:
[0029] Step 1: Obtain a three-dimensional original modality dataset and a three-dimensional target modality dataset of the same organ or lesion from a publicly available medical image dataset (in this embodiment, the Resect publicly available dataset is used), perform data preprocessing to adapt to the inputs of the conditional diffusion model and the hybrid attention segmentation model, obtain the preprocessed dataset, and divide the preprocessed dataset into a three-dimensional original modality training set, a three-dimensional original modality validation set, a three-dimensional original modality test set, a three-dimensional target modality training set, a three-dimensional target modality validation set, and a three-dimensional target modality test set;
[0030] Preferably, in step 1, the three-dimensional original modality dataset is composed of three-dimensional original modality images, and the three-dimensional original modality images are three-dimensional MRI images or three-dimensional CT images; the three-dimensional target modality dataset is a three-dimensional real ultrasound dataset composed of three-dimensional real ultrasound images.
[0031] Preferably, in step 1, the training set, the validation set, and the test set are divided according to a ratio of 7:2:1.
[0032] Preferably, in step 1, the specific steps of data preprocessing are as follows:
[0033] S11. Use the Gaussian filtering algorithm to remove artifacts and noise in the three-dimensional original modal dataset and the three-dimensional target modal dataset;
[0034] S12. Taking the three-dimensional target modal dataset as the standard, adjust the image origin and coordinate system direction of the three-dimensional original modal dataset to be the same as those of the three-dimensional target modal dataset, and adjust the voxel spacing and number of slices in the three-dimensional original modal dataset to be the same as those of the three-dimensional target modal dataset;
[0035] In this embodiment, the voxel spacing is uniformly 0.5mm×0.5mm×0.5mm, and the number of slices is uniformly 256×256×256.
[0036] S13. Use the organ or lesion labels corresponding to the three-dimensional original modal dataset to determine the field of view range of the three-dimensional original modal dataset, and then crop the three-dimensional original modal dataset based on the organ or lesion label information according to the field of view range;
[0037] In this embodiment, the specific operation for determining the field of view range of the three-dimensional original modal dataset is: taking the center of the lesion as the center point, and at the same time expanding 5mm outward with the window size of the three-dimensional target modal dataset for cropping, and the cropped area is used as the field of view range of the three-dimensional original modal dataset.
[0038] S14. Normalize the three-dimensional target modal dataset and the cropped three-dimensional original modal dataset so that their intensity values are adapted to the input of the conditional diffusion model;
[0039] In this embodiment, the initial intensity value of the three-dimensional target modal dataset is 0 - 150, and the initial intensity value of the three-dimensional original modal dataset is -200 - 1500, which are uniformly normalized to -1 - 1.
[0040] S15. Perform data augmentation on the normalized three-dimensional original modal dataset and the data in the three-dimensional target modal dataset used as the training set.
[0041] Preferably, in step S15, the data augmentation uses flipping with a 50% probability and Gaussian blur.
[0042] Step 2. Construct a conditional diffusion model, set corresponding training parameters based on the three-dimensional original modal training set and the three-dimensional target modal training set obtained in step 1, and use the dynamic learning rate adjustment and noise adjustment methods to fully train the conditional diffusion model. After training, obtain the trained conditional diffusion model and the weight parameters of the conditional denoising network TU-Net in the trained conditional diffusion model;
[0043] Preferably, in step 2, the overall structure of the conditional diffusion model is as follows Figure 2 shown, which consists of a forward diffusion process and a reverse diffusion process carried out in sequence; in the forward diffusion process, the three-dimensional original modal images in the three-dimensional original modal training set are input, and Gaussian noise is added step by step in time to generate a completely Gaussian noise image and the process-added noise images with different noise levels; in the reverse diffusion process, starting from the completely Gaussian noise image, the conditional denoising network TU-Net is guided by conditional information to perform denoising (removing noise) step by step in time, improving the texture and semantic consistency of the generated image until a completely denoised image is generated.
[0044] Preferably, in step 2, in the forward diffusion process, the process of adding Gaussian noise step by step in time is as follows:
[0045]
[0046] In formula (1), x0 represents the input image; x t is the noise-added image at time step t; I is the identity matrix; where α t = 1 - β t , β t represents the noise variance adjustment coefficient at time t; ∈ represents Gaussian noise that follows a zero mean and unit variance.
[0047] Preferably, in step 2, in the reverse diffusion process, the process of denoising and restoring the image step by step in time is as follows:
[0048]
[0049] In formula (2), L represents conditional information; μ θ , ∑ θ respectively represent the mean and variance of the noise predicted by the neural network; ∈ θ (x t , t) represents the noise predicted by the neural network at time step t; x t-1 represents the denoised image at time step t - 1.
[0050] Preferably, in step 2, the training parameters include the diffusion time step and the number of training epochs. In this embodiment, the time steps of both the forward diffusion process and the reverse diffusion process are set to 1000, and the number of training epochs is 500.
[0051] Preferably, in step 2, the dynamic learning rate adjustment adopts cosine annealing, and the noise adjustment adopts a linear scheduling method.
[0052] Preferably, in step 2, the structure of the conditional denoising network TU-Net is as follows Figure 3As shown, it includes five layers of parallel encoders and five layers of shared decoders; the five layers of parallel encoders are sequentially connected in order, and the five layers of shared decoders are sequentially connected in order; the output of the 5th layer parallel encoder is transmitted to the 1st layer shared decoder;
[0053] The output of the 1st layer parallel encoder is simultaneously transmitted to the 4th layer shared decoder through a skip connection, the output of the 2nd layer parallel encoder is simultaneously transmitted to the 3rd layer shared decoder through a skip connection, the output of the 3rd layer parallel encoder is simultaneously transmitted to the 2nd layer shared decoder through a skip connection, and the output of the 4th layer parallel encoder is simultaneously transmitted to the 1st layer shared decoder through a skip connection;
[0054] The structures of the 1st to 4th layer parallel encoders are the same, and each consists of a CNN encoding block, a conditional information embedding module, a Transformer encoding block, and a feature fusion module (CFIM); among them, the CNN encoding block, the conditional information embedding module, and the Transformer encoding block are parallel; the input enters the CNN encoding block, the conditional information embedding module, and the Transformer encoding block; the conditional information embedding module processes the conditional information and outputs the CNN local feature weight w CNN and the Transformer local feature weight w Tr , which are used to modulate the outputs of the CNN encoding block and the Transformer encoding block respectively; the feature fusion module fuses the output feature F CNN from the CNN encoding block with the output feature F Tr from the Transformer encoding block to obtain a fused feature map as the output of each layer of parallel encoder;
[0055] The 5th layer parallel encoder is a bottleneck layer, including two convolutional blocks. After each convolution, Batch Normalization is used for normalization, then the Sigmoid activation function is used for activation, and then downsampling is performed once; preferably, each convolutional block uses a 3×3×3 convolutional kernel, and both padding and stride are set to 1; the downsampling uses 2×2×2 max pooling;
[0056] The structures of the 1st to 5th layer shared decoders are the same. First, upsampling is performed once, and then two convolutional blocks are used. After each convolution, Batch Normalization is used for normalization, and then the Sigmoid activation function is used for activation; preferably, the upsampling uses trilinear interpolation three times; each convolutional block uses a 3×3×3 convolutional kernel, and both padding and stride are set to 1.
[0057] Preferably, in step 2, in the CNN encoding block, first two convolutional blocks are adopted. After each convolution, BatchNormalization is used for normalization, then the Sigmoid activation function is used for activation, and then downsampling is performed once. More preferably, each convolutional block uses a 3×3×3 convolutional kernel, and both padding and stride are set to 1; the downsampling uses 2×2×2 max pooling;
[0058] In the Transformer encoding block, first a self-attention mechanism modulated by conditional information is adopted, and then LayerNorm is used for normalization;
[0059] The conditional information embedding module adopts a convolutional block; more preferably, the convolutional block uses a 3×3×3 convolutional kernel;
[0060] The feature fusion module uses a 1×1×1 convolution for feature alignment, then performs feature concatenation, and finally uses a 3×3×3 convolution for feature extraction.
[0061] Preferably, in step 2, in the Transformer encoding block, the specific implementation of the self-attention mechanism modulated by conditional information is as follows:
[0062]
[0063] In formula (3), Q = XW Q , K = XW K , V = XW V ; W Q , W K , W V are learnable weight matrices; X is the input feature; d k is the dimension of the key; F MI is the weight matrix of the dynamic modulation multi-head attention.
[0064] Preferably, in step 2, the feature fusion module is expressed as:
[0065] F CFIM = Conv3×3×3([w CNN ·Conv1×1×1(F CNN ) + w Tr ·Conv1×1×1(F Tr )]) (4)
[0067] In formula (4), Conv1×1×1 represents a convolution operation with a convolutional kernel size of 1×1×1 for feature alignment; [·] represents a feature concatenation operation, and Conv3×3×3 is used to integrate the concatenated features.
[0068] Preferably, in step 2, the specific steps of denoising at each time step of the reverse diffusion process are as follows:
[0069] (1) Calculate the local mutual information between the process-added noise image of the current time step with dimensions of n×a×b×c (1×256×256×256 in this embodiment) and the three-dimensional real ultrasound image as conditional information; then input the process-added noise image of the current time step and the conditional information of the same dimension into the first-layer parallel encoder of the conditional denoising network TU-Net; the conditional information enters the conditional information embedding module of the first-layer parallel encoder to generate the CNN local feature weight w CNN and the Transformer local feature weight w Tr ; the process-added noise image of the current time step simultaneously enters the CNN encoding block and the Transformer encoding block of the first-layer parallel encoder. After two convolutions in the CNN encoding block, a feature map with dimensions of c×a / 2×b / 2×c / 2 (64×128×128×128 in this embodiment) is output. After passing through the self-attention mechanism in the Transformer encoding block, a feature map with dimensions of c×a / 2×b / 2×c / 2 is output; then the output feature F CNN of the CNN encoding block and the output feature F Tr of the Transformer encoding block are concatenated in the feature fusion module, and the output feature map with dimensions of c×a / 2×b / 2×c / 2 is the output of the first-layer parallel encoder;
[0070] (2) Calculate the local mutual information between the feature map output by the first-layer parallel encoder and the three-dimensional real ultrasound image as conditional information; send the conditional information and the feature map output by the first-layer parallel encoder to the second-layer parallel encoder, and the second-layer parallel encoder outputs a feature map with dimensions of 2c×a / 4×b / 4×c / 4 (128×64×64×64 in this embodiment);
[0071] (3) Calculate the local mutual information between the feature map output by the second-layer parallel encoder and the three-dimensional real ultrasound image as conditional information; send the conditional information and the feature map output by the second-layer parallel encoder to the third-layer parallel encoder, and the third-layer parallel encoder outputs a feature map with dimensions of 4c×a / 8×b / 8×c / 8 (256×32×32×32 in this embodiment);
[0072] (4) Calculate the local mutual information between the feature map output by the third-layer parallel encoder and the three-dimensional real ultrasound image as conditional information; send the conditional information and the feature map output by the third-layer parallel encoder to the fourth-layer parallel encoder, and the fourth-layer parallel encoder outputs a feature map with dimensions of 8c×a / 16×b / 16×c / 16 (512×16×16×16 in this embodiment);
[0073] (5) Calculate the local mutual information between the feature map output by the fourth-layer parallel encoder and the three-dimensional real ultrasonic image as conditional information; send the conditional information and the feature map output by the fourth-layer parallel encoder into the bottleneck layer. After two 3×3×3 convolutions in the bottleneck layer, a feature map with a dimension of 16c×a / 32×b / 32×c / 32 (1024×8×8×8 in this embodiment) is output;
[0074] (6) Send the feature map output by the bottleneck layer into the first-layer shared decoder. First, perform an upsampling, then pass through two convolutional blocks. After each convolution, use Batch Normalization for normalization, then use the Sigmoid activation function for activation, and then splice it with the feature map output by the fourth-layer parallel encoder through a skip connection, and output a feature map with a dimension of 8c×a / 16×b / 16×c / 16 (512×16×16×16 in this embodiment);
[0075] (7) Input the feature map output by the first-layer shared decoder into the second-layer shared decoder, successively perform an upsampling, two convolutions, one normalization, and one activation, and then splice it with the feature map output by the third-layer parallel encoder through a skip connection, and output a feature map with a dimension of 4c×a / 8×b / 8×c / 8 (256×32×32×32 in this embodiment);
[0076] (8) Input the feature map output by the second-layer shared decoder into the third-layer shared decoder, successively perform an upsampling, two convolutions, one normalization, and one activation, and then splice it with the feature map output by the second-layer parallel encoder through a skip connection, and output a feature map with a dimension of 2c×a / 4×b / 4×c / 4 (128×64×64×64 in this embodiment);
[0077] (9) Input the feature map output by the third-layer shared decoder into the fourth-layer shared decoder, successively perform an upsampling, two convolutions, one normalization, and one activation, and then splice it with the feature map output by the first-layer parallel encoder through a skip connection, and output a feature map with a dimension of c×a / 2×b / 2×c / 2 (64×128×128×128 in this embodiment);
[0078] (10) Input the feature map output by the fourth-layer shared decoder into the fifth-layer shared decoder, successively perform an upsampling, two convolutions, one normalization, and one activation, and output a feature map with a dimension of c×a×b×c (64×256×256×256 in this embodiment);
[0079] (11) The feature map output by the shared decoder of the fifth layer is dimensionally reduced through a 1×1×1 convolution to output a denoised image of the current time step of n×a×b×c (1×256×256×256 in this embodiment).
[0080] Preferably, in step 2, the conditional information used in the conditional denoising network TU-Net is local mutual information (LMI); in the training phase, by calculating the local mutual information, guiding the conditional diffusion model to learn the texture feature information of the target modality, adjusting the noise distribution in the process of adding noise to the image, and finally training the conditional diffusion model to generate an image with the same style as the three-dimensional target modality image; the implementation of local mutual information is as follows:
[0081]
[0082] In formula (5), p δ (x, y) is the joint probability density function on the domain δ; are respectively the marginal probability density functions on the domains δ xi and δ yj .
[0083] Preferably, in step 2, the denoising loss function of the conditional diffusion model combined with local mutual information is expressed as:
[0084]
[0085] In formula (6), LMI norm (F; F, F t ) is the local mutual information used to measure the target modality F and the noisy image of the current time step F t after normalization; S θ (F t , LMI, t) is the predicted noise of the conditional denoising network combining mutual information and time step t; ∈ is the Gaussian noise of the current time step; λ is the loss weight of the local mutual information term; by minimizing the loss function, continuously optimize the denoising ability of the model, and update all parameters in the conditional diffusion model in the reverse direction.
[0086] Step 3: Use the trained conditional diffusion model in step 2 to perform inference on each three-dimensional original modality image in the preprocessed three-dimensional original modality dataset obtained in step 1, batch translate it into a three-dimensional simulated ultrasound image with labels, and then form a three-dimensional simulated ultrasound dataset with labels, and divide the three-dimensional simulated ultrasound dataset with labels into a training set and a test set;
[0087] Preferably, in step 3, in the inference process, three-dimensional original modal images are used to calculate local mutual information to guide the conditional diffusion model. The overall semantic features of the three-dimensional original modal images are learned to further guide the prediction process of the conditional denoising network, so that the finally generated three-dimensional simulated ultrasound images are both close to the texture features of the target modality and retain the semantic features of the original modality.
[0088] Step 4: Transfer the weight parameters of the conditional denoising network TU-Net in the trained conditional diffusion model obtained in step 2 to the hybrid attention segmentation network, pre-train the hybrid attention segmentation network using the training set in step 3, and test the pre-training effect using the test set in step 3 to obtain the pre-trained hybrid attention segmentation network.
[0089] Preferably, in step 4, the hybrid attention segmentation network has the same structure as the conditional denoising network TU-Net, except that: a conditional information embedding module is not set in the parallel encoders of the first to fourth layers.
[0090] Preferably, in step 4, the parallel encoder of the hybrid attention segmentation network is initialized using the weight parameters of the conditional denoising network in the trained conditional diffusion model obtained in step 2, and the trained weight parameters are assigned to the hybrid attention segmentation network, so that the ability of the encoders of both the conditional denoising network and the hybrid attention segmentation network to extract low-level and intermediate features is directly shared, improving the adaptability of the hybrid attention segmentation network to cross-modalities, enhancing the feature extraction ability of the encoder of the hybrid attention segmentation network for three-dimensional ultrasound images, accelerating convergence at the same time, reducing the probability of overfitting, and finally improving the segmentation effect of the hybrid attention segmentation network.
[0091] Preferably, in step 4, the Transformer encoding block of the parallel encoder of the hybrid attention segmentation network captures global features through the self-attention mechanism; the self-attention mechanism is implemented as:
[0092]
[0093] In Equation (7), Q = XW Q , K = XW K , V = XW V ; W Q , W K, W V are learnable weight matrices; X is the input feature; d k is the dimension of the key.
[0094] Preferably, in step 4, the structure of the feature fusion module in the hybrid attention segmentation network is the same as that in the conditional denoising network TU-Net, with the only difference being that the values of the feature weights of the CNN and the feature weights of the Transformer are fixed and satisfy the sum of the two being equal to 1. In this embodiment, the feature weights of the CNN and the feature weights of the Transformer are both set to 0.5.
[0095] Preferably, in step 4, the loss function of the hybrid attention segmentation network is expressed as:
[0096] L seg =αL Dice +βL Edge
[0097]
[0098] In formula (8), L Dice is the Dice loss, which can improve the segmentation accuracy of small target regions; L Edge is the edge-aware loss, which focuses on the matching of edge regions and improves the boundary segmentation accuracy; p i is the predicted segmentation probability; g i is the ground truth label; N is the total number of pixels; represents the Sobel gradient operator; N e represents the number of edge pixels; ∥·∥ 2 represents the square of the two-norm; L seg is the total loss function; α and β are weighting coefficients, and in this embodiment, both α and β are set to 1; the overall region and edge segmentation effects of the hybrid attention segmentation network are optimized by using the joint loss function.
[0099] Step 5: Use the method of transfer learning to fine-tune the parameters of the pre-trained hybrid attention segmentation network obtained in step 4 with the three-dimensional target modality training set obtained in step 1 to adapt to the three-dimensional real ultrasound data, and use the fine-tuned model with the best performance on the three-dimensional target modality validation set in step 1 after the fine-tuning training is completed as the hybrid attention segmentation model; then use the hybrid attention segmentation model to process the three-dimensional target modality test set in step 1 to obtain the segmentation result.
[0100] Preferably, in step 5, the fine-tuning training is specifically: training and updating the parameters of the CNN encoding block of the pre-trained hybrid attention segmentation network obtained in step 4, freezing the parameters of the Transformer encoding block, and adjusting the network's ability to extract local features to make the network more adaptable to the three-dimensional real ultrasound data.
[0101] By Figure 5 and Figure 6 comparison, Figure 7And Figure 8 As can be seen from the comparison, the present invention can accurately segment ultrasonic images to obtain a segmentation map with clear boundaries.
[0102] What is not described in the present invention is applicable to the prior art.
Claims
1. A cross-modal ultrasound image segmentation method based on conditional diffusion model and transfer learning, characterized in that: The method comprises the following steps: Step 1, obtaining a three-dimensional original modality dataset and a three-dimensional target modality dataset of the same organ or lesion, performing data preprocessing to adapt the input of the conditional diffusion model and the hybrid attention segmentation model, obtaining a preprocessed dataset, and dividing the preprocessed dataset into a three-dimensional original modality training set, a three-dimensional original modality verification set, a three-dimensional original modality test set, a three-dimensional target modality training set, a three-dimensional target modality verification set, and a three-dimensional target modality test set; the three-dimensional target modality dataset is a three-dimensional real ultrasound dataset, consisting of three-dimensional real ultrasound images; Step 2: construct a conditional diffusion model, set training parameters based on the three-dimensional original modality training set and the three-dimensional target modality training set obtained in step 1, and use dynamic learning rate adjustment and noise adjustment to fully train the conditional diffusion model. After the training is completed, obtain the trained conditional diffusion model and the weight parameters of the conditional denoising network TU-Net in the trained conditional diffusion model; Step 3: Use the trained conditional diffusion model of step 2 to infer each 3D original modality image in the preprocessed 3D original modality data set obtained in step 1, and translate them into labeled 3D simulated ultrasound images in batches, thereby forming a labeled 3D simulated ultrasound data set, and divide the labeled 3D simulated ultrasound data set into a training set and a test set; Step 4: Migrate the weight parameters of the conditional denoising network TU-Net in the trained conditional diffusion model obtained in step 2 to the hybrid attention segmentation network, and use the training set in step 3 to pre-train the hybrid attention segmentation network to obtain the pre-trained hybrid attention segmentation network; Step 5, using the transfer learning method, fine-tune the parameters of the pre-trained hybrid attention segmentation network obtained in step 4 through the three-dimensional target modality training set obtained in step 1 to adapt to the three-dimensional real ultrasound data, and use the fine-tuned model that performs best on the three-dimensional target modality verification set in step 1 after the fine-tuning training is completed as the hybrid attention segmentation model; then use the hybrid attention segmentation model to process the three-dimensional target modality test set in step 1 to obtain the segmentation result.
2. The cross-modal ultrasound image segmentation method based on conditional diffusion model and transfer learning according to claim 1, characterized in that: In step 1, the three-dimensional original modality data set consists of three-dimensional original modality images, and the three-dimensional original modality images are three-dimensional MRI images or three-dimensional CT images; In step 1, the specific steps of data preprocessing are as follows: S11, using a Gaussian filtering algorithm to remove artifacts and noise in the three-dimensional original modality data set and the three-dimensional target modality data set; S12, taking the 3D target modality dataset as a standard, adjusting the image origin and coordinate system direction of the 3D original modality dataset to be the same as those of the 3D target modality dataset, and adjusting the voxel spacing and number of slices in the 3D original modality dataset to be the same as those of the 3D target modality dataset; S13, using the organ or lesion label corresponding to the three-dimensional original modality dataset to determine the field of view of the three-dimensional original modality dataset, and then cutting the three-dimensional original modality dataset based on the organ or lesion label information according to the field of view; S14, normalizing the three-dimensional target modal data set and the cropped three-dimensional original modal data set so that their intensity values are adapted to the input of the conditional diffusion model; S15. Perform data enhancement on the data in the normalized three-dimensional original modal data set and the three-dimensional target modal data set used as training sets.
3. The cross-modal ultrasound image segmentation method based on conditional diffusion model and transfer learning according to claim 1, characterized in that: In step 2, the conditional diffusion model consists of a forward diffusion process and a backward diffusion process performed in sequence; the forward diffusion process inputs the 3D original modality image in the 3D original modality training set, and generates a complete Gaussian noise image and process-noised images with different noise levels by adding Gaussian noise step by step; the backward diffusion process takes the complete Gaussian noise image as the starting point, and guides the conditional denoising network TU-Net to denoise step by step through conditional information, thereby improving the texture and semantic consistency of the generated image until a completely denoised image is generated; In step 2, during the forward diffusion process, the process of adding Gaussian noise step by time is as follows: In formula (1), x0 represents the input image; x t is the noisy image at time step t; I is the identity matrix; where α t =1-β t , β t represents the noise variance adjustment coefficient at time t; ∈ represents Gaussian noise with zero mean and unit variance; In step 2, during the reverse diffusion process, the denoising process is as follows: In formula (2), L represents conditional information; μ θ ,∑ θ Respectively represent the mean and variance of the neural network prediction noise; ∈ θ (x t , t) represents the noise of the neural network when predicting time step t; x t-1 represents the denoised image at time step t-1.
4. The cross-modal ultrasound image segmentation method based on conditional diffusion model and transfer learning according to claim 1, characterized in that: In step 2, the training parameters include the diffusion time step and the number of training rounds; In step 2, the dynamic learning rate adjustment adopts cosine annealing, and the noise adjustment adopts linear scheduling.
5. The cross-modal ultrasound image segmentation method based on conditional diffusion model and transfer learning according to claim 1, characterized in that: In step 2, the conditional denoising network TU-Net includes five layers of parallel encoders and five layers of shared decoders; the five layers of parallel encoders are connected sequentially, and the five layers of shared decoders are connected sequentially; the output of the fifth layer of parallel encoder is transmitted to the first layer of shared decoder; The output of the 1st layer parallel encoder is simultaneously transmitted to the 4th layer shared decoder through a skip connection, the output of the 2nd layer parallel encoder is simultaneously transmitted to the 3rd layer shared decoder through a skip connection, the output of the 3rd layer parallel encoder is simultaneously transmitted to the 2nd layer shared decoder through a skip connection, and the output of the 4th layer parallel encoder is simultaneously transmitted to the 1st layer shared decoder through a skip connection; The structures of the 1st to 4th layer parallel encoders are the same, and they are all composed of a CNN encoding block, a conditional information embedding module, a Transformer encoding block and a feature fusion module. The CNN encoding block, the conditional information embedding module and the Transformer encoding block are parallel. The input enters the CNN encoding block, the conditional information embedding module and the Transformer encoding block. The conditional information embedding module processes the conditional information and outputs the CNN local feature weight w. CNN and Transformer local feature weight w Tr , which are used to modulate the outputs of the CNN encoding block and the Transformer encoding block respectively; feature The fusion module combines the output features F from the CNN encoding block CNN And the output feature F of the Transformer encoding block Tr Perform fusion and obtain the fused feature map as the output of each layer of parallel encoder; The fifth parallel encoder is a bottleneck layer, which includes two convolution blocks. Batch Normalization is used for normalization after each convolution, and then the Sigmoid activation function is used for activation, and then downsampling is performed once. The structures of the shared decoders in the 1st to 5th layers are the same. They all perform upsampling first, then use two convolution blocks, use Batch Normalization for normalization after each convolution, and then use the Sigmoid activation function for activation.
6. The cross-modal ultrasound image segmentation method based on conditional diffusion model and transfer learning according to claim 4, characterized in that: In step 2, in the CNN encoding block, two convolution blocks are first used. BatchNormalization is used for normalization after each convolution, and then the Sigmoid activation function is used for activation, and then downsampling is performed; In the Transformer encoding block, a self-attention mechanism modulated by conditional information is first used, and then normalized using LayerNorm; The conditional information embedding module adopts a convolutional block; The feature fusion module uses 1×1×1 convolution for feature alignment, followed by feature concatenation, and finally uses 3×3×3 convolution for feature extraction; In step 2, the specific implementation of the self-attention mechanism modulated by conditional information in the Transformer encoding block is: In formula (3), Q = XW Q , K = XW K , V = XW V ; W Q , W K , W V is a learnable weight matrix; X is the input feature; d k is the dimension of the key; F MI is the weight matrix for dynamically modulating multi-head attention; In step 2, the feature fusion module is expressed as: F CFIM =Conv3×3×3([w CNN ·Conv1×1×1(F CNN )+w Tr ·Conv1×1×1(F Tr )]) (4) In formula (4), Conv1×1×1 represents the convolution operation with a convolution kernel size of 1×1×1, which is used for feature alignment; [·] represents the feature concatenation operation, and Conv3×3×3 is used to integrate the concatenated features.
7. The cross-modal ultrasound image segmentation method based on conditional diffusion model and transfer learning according to claim 3, characterized in that: In step 2, the specific steps of denoising at each time step of the reverse diffusion process are as follows: (1) Calculate the local mutual information between the process-noised image of the current time step and the three-dimensional true ultrasound image with a dimension of n×a×b×c as conditional information; Then, the process-noised image and conditional information of the current time step are input into the first-layer parallel encoder of the conditional denoising network TU-Net; the conditional information enters the conditional information embedding module of the first-layer parallel encoder to generate the CNN local feature weight w cNN and Transformer local feature weight w Tr ; The process-noised image of the current time step enters the CNN encoding block and the Transformer encoding block of the first-layer parallel encoder at the same time. After two convolutions in the CNN encoding block, the output feature map of the dimension is c×a / 2×b / 2×c / 2. After the self-attention mechanism in the Transformer encoding block, the output feature map of the dimension is c×a / 2×b / 2×c / 2. Then the output feature map F of the CNN encoding block is CNN And the output feature F of the Transformer encoding block Tr After concatenation in the feature fusion module, the output feature map with the dimension of c×a / 2×b / 2×c / 2 is the output of the first layer parallel encoder; (2) Calculating the local mutual information between the feature map output by the first-layer parallel encoder and the three-dimensional real ultrasound image as conditional information; sending the conditional information and the feature map output by the first-layer parallel encoder to the second-layer parallel encoder, and the second-layer parallel encoder outputs a feature map with a dimension of 2c×a / 4×b / 4×c / 4; (3) Calculating the local mutual information between the feature map output by the second-layer parallel encoder and the three-dimensional real ultrasound image as conditional information; sending the conditional information and the feature map output by the second-layer parallel encoder to the third-layer parallel encoder, and the third-layer parallel encoder outputs a feature map with a dimension of 4c×a / 8×b / 8×c / 8; (4) Calculating the local mutual information between the feature map output by the third-layer parallel encoder and the three-dimensional real ultrasound image as conditional information; sending the conditional information and the feature map output by the third-layer parallel encoder to the fourth-layer parallel encoder, and the fourth-layer parallel encoder outputs a feature map with a dimension of 8c×a / 16×b / 16×c / 16; (5) Calculate the local mutual information between the feature map output by the 4th parallel encoder and the 3D real ultrasound image as conditional information; send the conditional information and the feature map output by the 4th parallel encoder to the bottleneck layer. After two convolutions at the bottleneck layer, the feature map with a dimension of 16c×a / 32×b / 32×c / 32 is output; (6) The feature map output by the bottleneck layer is sent to the first-layer shared decoder, which is first upsampled once and then passes through two convolution blocks. Batch Normalization is used after each convolution, and then the Sigmoid activation function is used for activation. Then, it is concatenated with the feature map output by the fourth-layer parallel encoder through a skip connection, and the output dimension is a feature map of 8c×a / 16×b / 16×c / 16. (7) The feature map output by the first-layer shared decoder is input into the second-layer shared decoder, and is sequentially upsampled, convolved twice, normalized once, and activated once. It is then concatenated with the feature map output by the third-layer parallel encoder through a skip connection, and the output dimension is a feature map of 4c×a / 8×b / 8×c / 8. (8) The feature map output by the second-layer shared decoder is input into the third-layer shared decoder, and is sequentially upsampled, convolved twice, normalized once, and activated once. It is then concatenated with the feature map output by the second-layer parallel encoder through a skip connection, and the output feature map is a 2c×a / 4×b / 4×c / 4 feature map. (9) The feature map output by the third-layer shared decoder is input into the fourth-layer shared decoder, and is sequentially upsampled, convolved twice, normalized once, and activated once. It is then concatenated with the feature map output by the first-layer parallel encoder through a skip connection, and the output dimension is a feature map of c×a / 2×b / 2×c / 2. (10) The feature map output by the 4th layer shared decoder is input into the 5th layer shared decoder, and is sequentially upsampled once, convolved twice, normalized once, and activated once, and the output dimension is a feature map of c×a×b×c; (11) The feature map output by the 5th layer shared decoder is reduced in dimension by 1×1×1 convolution, and the denoised image of the current time step of n×a×b×c is output.
8. The cross-modal ultrasound image segmentation method based on conditional diffusion model and transfer learning according to claim 3, characterized in that: In step 2, the conditional information used in the conditional denoising network TU-Net is the local mutual information. In the training phase, the conditional diffusion model is trained to generate an image with the same style as the 3D target modality image by calculating the local mutual information. The local mutual information is realized as follows: In formula (5), p δ (x, y) is the joint probability density function on the domain δ; In the field δ xi and δ yj The marginal probability density function of ; In step 2, the denoising loss function of the conditional diffusion model combined with local mutual information is expressed as: In formula (6), LMI norm (F; F, F t ) is normalized to measure the target mode F and the current time step F t The local mutual information between noisy images; S θ (F t , LMI, t) is the conditional denoising network combined with mutual information and the prediction noise of time step t; ∈ is the Gaussian noise of the current time step; λ is the loss weight of the local mutual information term; the denoising ability of the model is continuously optimized by minimizing the loss function, and all parameters in the conditional diffusion model are updated in reverse.
9. The cross-modal ultrasound image segmentation method based on conditional diffusion model and transfer learning according to claim 1, characterized in that: In step 4, the structure of the hybrid attention segmentation network is the same as that of the conditional denoising network TU-Net, with the only difference being that no conditional information embedding module is set in the 1st to 4th layer parallel encoders; In step 4, the parallel encoder of the hybrid attention segmentation network is initialized using the weight parameters of the conditional denoising network in the trained conditional diffusion model obtained in step 2, and the trained weight parameters are assigned to the hybrid attention segmentation network, so that the encoders of the conditional denoising network and the hybrid attention segmentation network can directly share the ability to extract low-level and intermediate features, improve the adaptability of the hybrid attention segmentation network to cross-modality, and enhance the feature extraction ability of the encoder of the hybrid attention segmentation network for three-dimensional ultrasound images, while accelerating convergence, reducing the probability of overfitting, and ultimately improving the segmentation effect of the hybrid attention segmentation network; In step 4, the Transformer encoding block of the parallel encoder of the hybrid attention segmentation network captures global features through the self-attention mechanism; the self-attention mechanism is implemented as: In formula (7), Q = XW Q , K = XW K , V = XW v ; W q , W k , W v is a learnable weight matrix; X is the input feature; d K is the dimension of the key; In step 4, the structure of the feature fusion module in the hybrid attention segmentation network is the same as the feature fusion module in the conditional denoising network TU-Net, the only difference is that the feature weights of CNN and Transformer are fixed and the sum of the two is equal to 1; In step 4, the loss function of the hybrid attention segmentation network is expressed as: L seg =αL Dice +βL Edge In formula (8), L Dice is Dice loss; L Edge is the edge perception loss; p i is the predicted segmentation probability; g i is the true label; N is the total number of pixels; represents the Sobel gradient operator; N e Indicates the number of edge pixels; ||·|| 2 represents the square of the two norm; L seg is the total loss function; α and β are weighting coefficients; the overall area and edge segmentation effect of the hybrid attention segmentation network are optimized by using the joint loss function.
10. The cross-modal ultrasound image segmentation method based on conditional diffusion model and transfer learning according to claim 1, characterized in that: In step 5, the fine-tuning training is specifically: training and updating the parameters of the CNN encoding block of the pre-trained hybrid attention segmentation network obtained in step 4, and freezing the parameters of the Transformer encoding block.
Citation Information
Patent Citations
Lung CT image segmentation method based on Transform and convolutional neural network
CN116739985A
Diffusion model, multi-scale and attention module medical ultrasonic image segmentation method
CN119180826A
Method for two-dimensional nuclear magnetic resonance diffusion ordered spectroscopy based on deep learning
US20240019515A1
Cited By
Wearable signal cross-user automatic monitoring method based on external distribution detection and diffusion model
CN121234131A