A pseudo-CT cross-modal conversion method and system
By registering and feature fusion of CT and MR images, and combining LLM analyzing text reports, we drive the adversarial diffusion model to generate pseudo-CT images, solving the accuracy and efficiency problems of cross-modal conversion in the existing technology, and achieving rapid generation of high-fidelity fake CT and high-precision support for radiotherapy planning.
Patent Information
- Application Number
- CN202510473259.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The prior art has anatomical distortion, artifacts, insufficient fidelity, difficulty in dealing with nonlinear intensity differences, and lack of real-time regulation capabilities in the cross-modal conversion of MR to pseudo-CT, which is difficult to meet the high-precision requirements of radiotherapy planning.
By rigid and deformation registration of CT images and MR images, key information is extracted and encoded into condition vectors available for diffusion model, combined with LLM to analyze clinical text reports, perform multimodal feature fusion, and drive the adversarial diffusion model to generate high-fidelity and false CT images. The training process of the diffusion model is divided into two stages, optimized condition generation ability and semantic alignment ability.
It realizes the rapid and accurate generation of high-fidelity and false CT images, improves the efficiency and quality of image generation, and meets the high-precision needs of radiotherapy planning.
Smart Images

Figure CN119991751B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pseudo-CT cross-modal conversion, and particularly relates to a pseudo-CT cross-modal conversion method and system. Background Art
[0002] In radiotherapy planning, CT images are the gold standard for dose calculation, but they have the risk of ionizing radiation, and some patients (such as children, pregnant women) need to avoid repeated CT scans. Although MRI has no radiation and excellent soft tissue contrast, it lacks electron density information and cannot be directly used for dose calculation. Therefore, developing MR-to-pseudo-CT cross-modal conversion technology has become the core requirement for precise radiotherapy.
[0003] The current mainstream methods have the following limitations: Methods based on generative adversarial networks (GANs) (such as CycleGAN, pix2pix): The generated images are prone to anatomical structure distortion or artifacts, especially in complex lesion areas (such as bone-soft tissue boundaries) with insufficient fidelity; Traditional registration methods (such as Elastix): Rely on rigid / non-rigid transformations, difficult to handle non-linear intensity differences between multi-modalities, and unable to fuse clinical text semantic information; Single-modal conditional generation models: Ignore key semantic descriptions in clinical reports (such as lesion location, morphology), resulting in the disconnection between the generated pseudo-CT and the individual characteristics of patients; Lack of dynamic control: Existing methods lack the ability to real-time regulate the generation process (such as adjusting the generation granularity according to lesion complexity), making it difficult to meet the high-precision requirements of radiotherapy planning. Summary of the Invention
[0004] Based on this, the embodiments of the present invention provide a pseudo-CT cross-modal conversion method and system, aiming to accurately and quickly generate pseudo-CT images.
[0005] The first aspect of the embodiments of the present invention provides a pseudo-CT cross-modal conversion method, the method comprising:
[0006] Performing rigid and deformable image registration on the obtained CT image and MR image, and aligning the CT image and the MR image;
[0007] Performing text annotation processing on the aligned MR image, and extracting key information, wherein the key information at least includes anatomical structure, lesion characteristics, and spatial relationship;
[0008] Encoding the extracted key information into a conditional vector available for a diffusion model, and at the same time, performing multi-modal feature fusion on the text and the corresponding MR image, wherein the text of the key information is converted into a structured conditional vector, i.e., text embedding, through the LLM;
[0009] Build a diffusion model that guides and counteracts the diffusion process through boundary contour information to improve the generation efficiency and quality of images. Additionally, inject the conditional vector into the diffusion model through a cross-attention layer;
[0010] Input the result of multimodal feature fusion into the diffusion model to train the diffusion model. The training process is divided into two stages. In the first stage, only optimize the conditional generation ability of the diffusion model. In the second stage, simultaneously optimize the semantic alignment ability between the diffusion model and the LLM;
[0011] Input the MR image to be converted into the trained diffusion model to output a pseudo-CT image.
[0012] Furthermore, the steps for multimodal feature fusion of the text and the corresponding MR image include:
[0013] Use PCA or an autoencoder to compress the high-dimensional text embedding to a dimension that matches the MR image features, and align the text embedding with the visual features of the corresponding MR image in the latent space through contrastive learning. Among them, if the text contains anatomical location information, convert the anatomical location information into a 3D coordinate mask.
[0014] Furthermore, the latent space alignment adopts a multimodal feature fusion mechanism, expressed as:
[0015]
[0016] Among them, is the cosine similarity, is the temperature parameter, is the positive sample image feature, is the text embedding, is the k-th image feature, and N is the total number of samples in the batch, is the contrast loss value.
[0017] Furthermore, in the diffusion model, insert a dynamic convolution module in the skip connection layer to enhance the multi-scale feature fusion ability, and at the same time use spectral normalization to stabilize the adversarial training.
[0018] Furthermore, the calculation formula in the diffusion model is:
[0019]
[0020] Among them, is the noise scheduling parameter, is the conditional vector of the combination of the text and the MR image, and are output by the noise prediction network, is the mean of the inverse process, is the covariance of the inverse process, is the probability distribution of the inverse process, is the noisy image at the t-th step, is the noisy image at the (t - 1)-th step, is the transition probability, and t is the time step, is the cumulative product result of the first t steps, N(•) is the Gaussian distribution with mean and covariance β t I , I is the identity matrix, is the noise prediction network, with input , t, and , and outputs the predicted noise, is the average value of the cumulative product result of the noise schedule parameters for the first t steps, is the average value of the cumulative product result of the noise schedule parameters for the first (t - 1) steps, is the product symbol, is the noise schedule parameter related to the variable s, where the variable s ranges from 1 to t.
[0021] Furthermore, in the step of injecting the conditional vector into the diffusion model through the cross-attention layer, the text input is used to guide the diffusion model, and the diffusion process is optimized through CLIP embedding or Latent Guidance.
[0022] Furthermore, the total loss function of the first stage is:
[0023]
[0024] where, is the total loss function of the first stage, is the noise prediction loss, is the conditional distribution constraint loss, is the adversarial loss, is the CLIP loss, , , and are the initial weights of the first stage respectively.
[0025] Furthermore, the total loss function of the second stage is:
[0026]
[0027] where, is the total loss function of the second stage, is the noise prediction loss, is the conditional distribution constraint loss, is the adversarial loss, is the CLIP loss, is the text reconstruction loss, , , and is the weight adjusted in the second stage, is the text description extracted by the reverse encoder in the generated pseudo-CT, is the original real text description, is the generator network taking the MR image x MR and the conditional vector c as inputs, is the image encoder, is the text encoder.
[0028] The second aspect of the embodiments of the present invention provides a pseudo-CT cross-modal conversion system for implementing the pseudo-CT cross-modal conversion method provided in the first aspect. The system includes:
[0029] A registration module for performing rigid and deformable image registration on the acquired CT image and MR image, and aligning the CT image and MR image;
[0030] An extraction module for performing text annotation processing on the aligned MR image and extracting key information, where the key information at least includes anatomical structures, lesion characteristics, and spatial relationships;
[0031] An encoding module for encoding the extracted key information into a conditional vector available for the diffusion model. Meanwhile, multi-modal feature fusion is performed on the text and the corresponding MR image. Among them, the key information text is converted into a structured conditional vector, i.e., text embedding, through the LLM;
[0032] A model establishment module for establishing a diffusion model. The diffusion model guides and counteracts the diffusion process through boundary contour information to improve the generation efficiency and quality of the image. In addition, the conditional vector is injected into the diffusion model through a cross-attention layer;
[0033] A training module for inputting the result of multi-modal feature fusion into the diffusion model to train the diffusion model. The training process is divided into two stages. In the first stage, only the conditional generation ability of the diffusion model is optimized. In the second stage, the semantic alignment ability of both the diffusion model and the LLM is optimized simultaneously;
[0034] An input module for inputting the MR image to be converted into the trained diffusion model and outputting a pseudo-CT image.
[0035] A pseudo-CT cross-modal conversion method and system provided in an embodiment of the present invention. After registering and aligning CT images and MR images, the method uses an LLM to parse clinical text reports, extracts anatomical structure descriptions, lesion characteristics, and spatial relationships, and generates a structured conditional vector. Subsequently, the text semantic embedding is fused with the MR image features to drive an adversarial diffusion model to generate high-fidelity pseudo-CT. Among them, the training process of the diffusion model is divided into two stages. In the first stage, only the conditional generation ability of the diffusion model is optimized, and in the second stage, the semantic alignment ability of both the diffusion model and the LLM is optimized. Finally, the MR image to be converted is input into the trained diffusion model to quickly output an accurate pseudo-CT image. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 FIG. is a flowchart of an implementation of a pseudo-CT cross-modal conversion method provided in Embodiment 1 of the present invention;
[0037] Figure 2 FIG. is a structural block diagram of a pseudo-CT cross-modal conversion system provided in Embodiment 2 of the present invention;
[0038] Figure 3 FIG. is a structural block diagram of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0040] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.
[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0042] Embodiment 1
[0043] Please refer to Figure 1 ,Figure 1 The figure shows an implementation flowchart of a pseudo-CT cross-modal conversion method provided in Embodiment 1 of the present invention. The pseudo-CT cross-modal conversion method specifically includes steps S01 to S06.
[0044] In step S01, the obtained CT image and MR image are subjected to rigid and deformable image registration, and the CT image and MR image are aligned.
[0045] Specifically, 200 cases of CT images and MR images of pelvic tumor patients are collected. It can be understood that the CT images and MR images contain data information. Then, rigid and deformable image registration is performed, and the CT image and MR image are aligned to obtain a Dice similarity coefficient greater than 0.98, indicating good spatial correspondence between the modalities.
[0046] In step S02, the aligned MR image is subjected to text annotation processing, and key information is extracted. Among them, the key information at least includes anatomical structure, lesion characteristics, and spatial relationship.
[0047] The process of text annotation processing is as follows:
[0048] 1) Remove sensitive information such as patient name and ID number (regular expression matching) for de-identification; replace the hospital number with an anonymous identifier (such as `Patient_001`);
[0049] 2) Use a medical ontology library to convert non-standard text descriptions into standard terms, such as the original text: "Occupation in the right temporal lobe of the brain" → Standardization: "Occupying lesion in the right temporal lobe";
[0050] 3) Segment the text by anatomical region (such as separating "brain description" from "abdomen description"); use punctuation marks such as full stops and semicolons to divide independent semantic units.
[0051] Furthermore, use BioBERT to parse the imaging data to clarify the key information (such as organ position, lesion characteristics, soft tissue names, bone boundaries, etc.). Specifically, entity recognition is performed through a domain-specific BioBERT. For example: anatomical sites (such as "right temporal lobe", "liver", "bone", "kidney", etc.); quantitative parameters (such as "2.1 cm × 1.8 cm"). Based on rule templates (such as SemRep) or LLM fine-tuning models, the association between anatomical structures and outer contours is constructed (such as "right kidney → boundary position").
[0052] In step S03, the extracted key information is encoded into a conditional vector available for the diffusion model. At the same time, multi-modal feature fusion is performed on the text and the corresponding MR image. Among them, the text of the key information is converted into a structured conditional vector, that is, text embedding, through the LLM.
[0053] It should be noted that the key information encoding can be completed through the embedding layer of the LLM (such as the `[CLS]` vector of BioBERT or entity-level embedding). Further, the steps of multi-modal feature fusion of the text and the corresponding MR image include:
[0054] Use PCA or autoencoder to compress the high-dimensional text embedding to a dimension matching the MR image features, and align the text embedding with the visual features of the corresponding MR image in the latent space through contrastive learning. Among them, if the text contains anatomical location information, the anatomical location information is converted into a 3D coordinate mask (registered with the MR image space). Specifically, the latent space alignment adopts a multi-modal feature fusion mechanism, which is expressed as:
[0055]
[0056] Among them, is the cosine similarity, is the temperature parameter, is the positive sample image feature, is the text embedding, is the k-th image feature, N is the total number of samples in the batch, is the contrastive loss value.
[0057] To ensure the accuracy of training the diffusion model, the association between text and image data is confirmed. Specifically, it is automatically filtered based on the `Body Part Examined` field in the DICOM metadata to ensure that the anatomical region described in the text matches the MR image ROI (Region of Interest) (for example, when describing "liver disease", the MR must include an abdominal scan), and the radiologist confirms the correspondence between the text description and the image lesion (such as the boundary position and size of the anatomical structure).
[0058] Step S04, establish a diffusion model. The diffusion model guides and counteracts the diffusion process through boundary contour information to improve the generation efficiency and quality of the image. In addition, the conditional vector is injected into the diffusion model through a cross-attention layer.
[0059] In this embodiment, a novel boundary contour information-guided adversarial diffusion model is proposed as the main framework for generating synthetic CT images from MR images, which combines the advantages of diffusion models and generative adversarial networks (GANs), and improves the generation efficiency and quality of images through boundary contour information guidance and adversarial diffusion processes ( ). In addition, a dynamic convolution module is inserted into the skip connection layer to enhance the multi-scale feature fusion ability, and spectral normalization is used to stabilize the adversarial training. The specific mathematical formula is as follows, where the effective adversarial diffusion forward noise addition and inverse process are expressed as:
[0060]
[0061] Among them, is the noise scheduling parameter, which controls the noise intensity at the t-th step (0 < < 1), is the conditional vector of the text and MR image combination, and are output by the noise prediction network, is the mean of the inverse process, is the covariance of the inverse process, is the probability distribution of the inverse process, modeled by the neural network parameter θ, is the noisy image at the t-th step, is the noisy image at the (t - 1)-th step, is the transition probability, and t is the time step, is the cumulative product result of the previous t steps, and N(•) is the Gaussian distribution with mean and covariance β t I , I is the identity matrix to simplify the covariance structure and ensure that the noise is evenly distributed in all spatial dimensions, is the noise prediction network. In the inverse process, by predicting the noise and calculating the mean , it guides the generation of high-fidelity pseudo-CT images. The input , t, and are input, and the predicted noise is output, is the average value of the cumulative product result of the noise scheduling parameters of the previous t steps, is the average value of the cumulative product result of the noise scheduling parameters of the previous (t - 1) steps, is the product symbol, is the noise scheduling parameter related to the variable s, and the variable s changes from 1 to t.
[0062] It should be noted that in the main framework, the value of t is significantly greater than 1. Considering that t is much larger than 1, the large-step backward process lacks a closed-form expression. To solve this problem, the present invention introduces an original domain traditional mapper to obtain the complex transition probability in the large-step conditional diffusion model. The idea of an adversarial projector with a cycle-consistent structure is borrowed and incorporated into the diffusion process. In the generator Gen θ takes the input x t of the source modality y and the newly introduced corresponding boundary mask mask as inputs, , extracts intermediate features for step-by-step denoising, is the generated pseudo-CT image, m is the boundary mask, and x is the input image of the source modality. Additionally, a learnable embedding mask calculated based on time t is incorporated into these extracted features. Then, a discriminative factor is introduced to distinguish the estimated denoised distribution generated by Gen θ from the actual denoised samples. Meanwhile, a time embedding is introduced as a bias term for the feature map.
[0063] To generate the target synthetic CT (sCT, pseudo-CT), the original MRI image with the same anatomical structure can be used as prior information for the inverse diffusion step. The source image corresponding to each sCT image in the training dataset is estimated through non-diffusion and diffusion processes. Previous diffusion-based image conversion methods usually rely on global feature matching to achieve comprehensive image conversion. However, due to the inclusion of a large number of irrelevant background regions, this method often leads to low feature learning efficiency, thereby affecting the synthesis result. To address this issue, the present invention proposes to explicitly integrate body contour guidance into the feature learning framework. Specifically, boundary guidance is integrated into the conditional generation process of its diffusion and non-diffusion networks, effectively reducing the influence of background noise and ensuring that the model learning focuses on the human body region. Through this efficient and selective feature learning paradigm, the model can capture key features such as internal anatomical details and external body contours, significantly improving the feature learning effect of internal anatomical synthesis, and thus obtaining more accurate and reliable results. To achieve unsupervised learning, a cycle consistency loss is utilized to compare the real target data with its generated data. During the diffusion process, the reconstructed image is called the synthetic target image; in the non-diffusion module, the estimated source image is mapped to the target domain through the generator. The diffusion and non-diffusion modules are jointly trained without a pre-training process. To calculate and , when the sampling network parameterizes the denoising distribution , the operation of the generation distribution is as follows:
[0064]
[0065] where H θ (•) is the parameterized denoising conditional distribution, representing the conditional probability distribution of generating the image at the t−k step given the current image state x t , the conditional input y, and the body contour mask mask. Understandably, x t−k is the denoised target image state at the t−k step in the diffusion process, x t is the current image state at the t step in the diffusion process, y is the conditional input, usually the source domain image (such as the original MRI image), providing anatomical structure prior information, mask is the body contour mask, used to explicitly guide the generation process, restricting the model to focus on the human body region and reducing background noise interference, Gen θis the generator network, responsible for generating intermediate images based on conditional inputs , L (•) is the log-likelihood function, which measures the matching degree between the generated image and the target distribution x t−k . The model diffusion module is used to estimate and synthesize target domain images from the raw domain data output by the non-diffusion module. To achieve this goal, two traditional adversarial diffusion methods are adopted, each with its own discriminator. At each step of the reverse process, the generator first makes a deterministic estimate of the denoised target image, and then uses the denoising distribution specific to each image modality to synthesize the target image.
[0066] Furthermore, text inputs are used to guide the diffusion model, and the diffusion process is optimized through CLIP embeddings or Latent Guidance. Among them, the text description is converted into an embedding vector using a CLIP text encoder or an LLM (such as BioBERT) . A cross-attention module is inserted into each downsampling and upsampling layer of the U-Net to align the text embeddings with the image features:
[0067]
[0068] where is the image feature projection, is the text conditional projection. The visual features of the generated pseudo-CT are constrained by the CLIP image encoder to align with the text semantics, enhancing conditional control:
[0069]
[0070] where is the CLIP image encoder, is the CLIP text encoder.
[0071] Step S05, input the result of multi-modal feature fusion into the diffusion model to train the diffusion model. The training process is divided into two stages. In the first stage, only the conditional generation ability of the diffusion model is optimized. In the second stage, the semantic alignment ability of both the diffusion model and the LLM is optimized.
[0072] It should be noted that during the initial training stage of the diffusion model (LLM parameters are frozen), only the conditional generation ability of the diffusion model is optimized, that is, the first stage. At this time, the LLM serves as a fixed text encoder and does not participate in backpropagation. During the end-to-end joint training stage (unfreezing some LLM parameters), the semantic alignment ability of both the diffusion model and the LLM is optimized, that is, the second stage.
[0073] Specifically, the total loss function in the first stage is:
[0074] (1)
[0075] Among them, is the total loss function of the first stage, is the noise prediction loss, is the conditional distribution constraint loss, is the adversarial loss, is the CLIP loss, , , and are the initial weights of the first stage respectively. Further, the formulas for the noise prediction loss and the conditional distribution constraint loss are:
[0076]
[0077] where x 0 is the real CT image, is the real noise, p(c) is the standard normal distribution, the constraint condition vector distribution, is the mathematical expectation of x 0 , t and . D KL is the divergence, used to measure the difference between two probability distributions, is the probability distribution of under the given and text. In addition, the adversarial loss generator, discriminator loss, and perceptual loss integrated in the diffusion model are:
[0078]
[0079] where is the feature of the i-th layer of the pre-trained VGG network, is the expected score that the sample generated by the generator is judged to be true by the discriminator, is to take the expectation of the joint distribution of x t and t. D is the discriminator, G is the generator, is to calculate the logarithmic probability of the discriminator D for , is to calculate the logarithmic probability that the discriminator judges the forged sample generated by the generator to be false. In this embodiment, , , , and are 1, 0.1, 0.5, 0.2, and 0.05 respectively.
[0080] The total loss function of the second stage is:
[0081] (2)
[0082]
[0083] Among them, is the total loss function of the second stage, is the noise prediction loss, is the conditional distribution constraint loss, is the adversarial loss, is the CLIP loss, is the text reconstruction loss, , , and are the adjusted weights in the second stage, is the text description extracted by the reverse encoder in the generated pseudo-CT, is the original real text description, is the generator network, taking the MR image x MR and the conditional vector c (such as noise or other control parameters) as inputs, and outputting the synthesized pseudo-CT image, is the image encoder, mapping the generated pseudo-CT image to the text feature space, is the text encoder, encoding the real text description into a feature vector.
[0084] It should be noted that the diffusion model is finally trained with mixed precision, using FP16 precision to accelerate the calculation, and scaling the gradients (Scale = 1024) to prevent underflow. When Inf / NaN appears in the gradients, the dynamic loss scaling strategy is adopted to automatically reduce the scaling factor. Finally, the conditional total loss is integrated for the joint optimization of the deep semantic alignment of the LLM and the diffusion model, that is, the total loss function of the second stage.
[0085] It can be understood that when performing end-to-end fine-tuning, the text reconstruction loss is introduced to constrain that the generated pseudo-CT image can reconstruct the original text description through the reverse model, thereby enhancing the alignment between the text and the image. The total loss in formula (2) is in the subsequent joint training or fine-tuning stage, when the LLM and the diffusion model are optimized together, and the text reconstruction loss is additionally added. Formula (1) is for the preliminary training of the diffusion model, and the latter is the end-to-end optimization after integrating the LLM. In the second stage, some LLM layers are unfrozen, and the text reconstruction loss is added to optimize the entire system in an end-to-end manner. The two total loss functions correspond to different training stages and optimization objectives respectively. The total loss in formula (2) is introduced in the end-to-end fine-tuning stage of formula (1), adding the text reconstruction loss to strengthen the consistency between the text and the generated image. Through the above strategies, the LLM and the diffusion model achieve deep cooperation, while ensuring the generation quality, improving the calculation efficiency through dynamic control, and meeting the high-precision and real-time requirements of pseudo-CT for precise radiotherapy.
[0086] Through a large number of ablation experiments and quantitative analyses, all relevant parameters were reasonably determined through multiple adjustment processes. The proposed deep network architecture was implemented based on PyTorch, and the training and validation processes were carried out on a computer equipped with an Nvidia RTXA6000 GPU (48G video memory), which supports CUDA acceleration and can meet the computational requirements of the diffusion model. All networks were trained using the Adam optimizer to minimize the L1 loss function. The generator and discriminator were trained in an alternating manner, with a batch size of 1. The loss weights η Φ and η θ were both set to 0.5, and the diffusion step number T / k was set to 4. The network was trained for a total of 200 epochs, with a batch size of 1 and a diffusion step number of 8. The initial learning rate was set to 0.0001 in the first 100 epochs, and it was halved every 100 epochs thereafter. During the training process, the trade-off parameter was set to 1. Using the trained generation model, it takes approximately 30 hours of computing time. Generating sCT for new cases based on MR data only takes a fraction of a second.
[0087] Step S06, input the MR image to be converted into the trained diffusion model, and output a pseudo-CT image.
[0088] It should be noted that after generating the pseudo-CT image, the LLM model is used to automatically generate a report to explain the key anatomical or pathological features in the generation result. Let the LLM analyze the generated pseudo-CT, automatically detect artifacts, and combine the LLM to generate correction suggestions to guide the secondary optimization of the diffusion model. Quantitative evaluations include: using indicators such as SSIM (structural similarity), MAE (mean absolute error), PSNR (peak signal-to-noise ratio), and Dice coefficient (organ segmentation consistency). Qualitative evaluations include the evaluation of the anatomical rationality and clinical usability of the pseudo-CT by radiologists.
[0089] In summary, the pseudo-CT cross-modal conversion method in the above embodiments of the present invention. After registering and aligning the CT image and the MR image, the method uses the LLM to parse the clinical text report, extracts anatomical structure descriptions, lesion characteristics, and spatial relationships, generates a structured conditional vector. Subsequently, the text semantic embedding is fused with the MR image features to drive the adversarial diffusion model to generate a high-fidelity pseudo-CT. Among them, the training process of the diffusion model is divided into two stages. In the first stage, only the conditional generation ability of the diffusion model is optimized, and in the second stage, the semantic alignment ability of both the diffusion model and the LLM is optimized. Finally, the MR image to be converted is input into the trained diffusion model to quickly output an accurate pseudo-CT image.
[0090] Embodiment 2
[0091] Please refer to Figure 2 ,Figure 2 It is a structural block diagram of a pseudo-CT cross-modal conversion system provided in the second embodiment of the present invention. The pseudo-CT cross-modal conversion system 200 includes: a registration module 21, an extraction module 22, an encoding module 23, a model establishment module 24, a training module 25, and an input module 26, where:
[0092] The registration module 21 is used to perform rigid and deformable image registration on the acquired CT image and MR image, and align the CT image and MR image;
[0093] The extraction module 22 is used to perform text annotation processing on the aligned MR image and extract key information, where the key information at least includes anatomical structure, lesion characteristics, and spatial relationship;
[0094] The encoding module 23 is used to encode the extracted key information into a conditional vector available for the diffusion model. At the same time, multi-modal feature fusion is performed on the text and the corresponding MR image. Among them, the text of the key information is converted into a structured conditional vector, that is, text embedding, through the LLM. Specifically, PCA or an autoencoder is used to compress the high-dimensional text embedding to a dimension matching the MR image features, and contrastive learning is used to align the text embedding with the visual features of the corresponding MR image in the latent space. Among them, if the text contains anatomical location information, the anatomical location information is converted into a 3D coordinate mask. In addition, the latent space alignment adopts a multi-modal feature fusion mechanism, which is expressed as:
[0095]
[0096] Among them, is the cosine similarity, is the temperature parameter, is the positive sample image feature, is the text embedding, is the k-th image feature, N is the total number of samples in the batch, is the contrast loss value;
[0097] The model establishment module 24 is used to establish a diffusion model. The diffusion model guides and counteracts the diffusion process through boundary contour information to improve the generation efficiency and quality of the image. In addition, the conditional vector is injected into the diffusion model through a cross-attention layer. In the diffusion model, a dynamic convolution module is inserted into the skip connection layer to enhance the multi-scale feature fusion ability, and spectral normalization is used to stabilize the adversarial training. The calculation formula in the diffusion model is:
[0098]
[0099] Among them, is the noise scheduling parameter, is the conditional vector for the combination of text and MR images, and is output by the noise prediction network, is the mean of the reverse process, is the covariance of the reverse process, is the probability distribution of the reverse process, is the noisy image at the t-th step, is the noisy image at the (t - 1)-th step, is the transition probability, and t is the time step, is the cumulative product result of the first t steps, and N(•) is the Gaussian distribution with mean and covariance β t I , I is the identity matrix, is the noise prediction network, with input , t, and , and outputs the predicted noise, is the average of the cumulative product results of the noise scheduling parameters for the first t steps, is the average of the cumulative product results of the noise scheduling parameters for the first (t - 1) steps, is the product symbol, is the noise scheduling parameter related to the variable s, where the variable s ranges from 1 to t. Additionally, the text input is used to guide the diffusion model, and the diffusion process is optimized through CLIP embedding or Latent Guidance;
[0100] The training module 25 is used to input the result of multi-modal feature fusion into the diffusion model to train the diffusion model. The training process is divided into two stages. In the first stage, only the conditional generation ability of the diffusion model is optimized, and in the second stage, the semantic alignment ability of both the diffusion model and the LLM is optimized. The total loss function in the first stage is:
[0101]
[0102] where, is the total loss function in the first stage, is the noise prediction loss, is the conditional distribution constraint loss, is the adversarial loss, is the perceptual loss, is the CLIP loss, , , , and are the initial weights in the first stage respectively. The total loss function in the second stage is:
[0103]
[0104] Among them, is the total loss function of the second stage, is the noise prediction loss, is the conditional distribution constraint loss, is the adversarial loss, is the CLIP loss, is the text reconstruction loss, , , and are the adjusted weights in the second stage, is the text description extracted by the reverse encoder in the generated pseudo-CT, is the original real text description, is the generator network with the MR image x MR and the conditional vector c as inputs, is the image encoder, is the text encoder;
[0105] The input module 26 is used to input the MR image to be converted into the trained diffusion model and output a pseudo-CT image.
[0106] Embodiment III
[0107] On the other hand, the present invention also proposes an electronic device. Please refer to Figure 3 , which shows the electronic device in Embodiment III of the present invention, including a memory 20, a processor 10, and a computer program 30 stored in the memory and executable on the processor. When the processor 10 executes the computer program 30, the pseudo-CT cross-modal conversion method as described above is implemented.
[0108] Among them, in some embodiments, the processor 10 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips, and is used to run the program code stored in the memory 20 or process data, such as executing an access restriction program, etc.
[0109] Among them, the memory 20 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 20 can be an internal storage unit of the electronic device in some embodiments, such as the hard disk of the electronic device. The memory 20 can also be an external storage device of the electronic device in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a FlashCard, etc. equipped on the electronic device. Further, the memory 20 can also include both an internal storage unit and an external storage device of the electronic device. The memory 20 can be used not only to store application software and various types of data of the electronic device, but also to temporarily store the data that has been output or will be output.
[0110] It should be noted that Figure 3 The structure shown does not constitute a limitation on the electronic device. In other embodiments, the electronic device may include fewer or more components than shown in the figure, or combine certain components, or have a different component arrangement.
[0111] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the pseudo-CT cross-modal conversion method as described above.
[0112] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device.
[0113] More specific examples (a non-exhaustive list) of computer-readable media include the following: electrical connections (electronic devices) having one or more wirings, portable computer diskettes (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber devices, and portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing it as appropriate, and then storing it in a computer memory.
[0114] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and the like.
[0115] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0116] The above embodiments merely represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.
Claims
1. A pseudo-CT cross-modality conversion method, characterized in that: The method comprises: Performing rigid and deformable image registration of the acquired CT image with the MR image, and aligning the CT image with the MR image; Performing text annotation processing on the aligned MR images and extracting key information, wherein the key information at least includes anatomical structure, lesion characteristics and spatial relationship; The extracted key information is encoded into a conditional vector that can be used by the diffusion model. At the same time, the text and the corresponding MR image are multimodal feature fused, where the text of the key information is converted into a structured conditional vector, namely, text embedding, through LLM; Establishing a diffusion model, wherein the diffusion model guides and counteracts the diffusion process through boundary contour information to improve the generation efficiency and quality of the image, and further, injecting the conditional vector into the diffusion model through a cross attention layer; The result of multimodal feature fusion is input into the diffusion model to train the diffusion model. The training process is divided into two stages. The first stage only optimizes the conditional generation ability of the diffusion model, and the second stage simultaneously optimizes the semantic alignment ability of the diffusion model and the LLM. The MR image to be converted is input into the trained diffusion model and a pseudo CT image is output.
2. The pseudo-CT cross-modality conversion method according to claim 1, characterized in that: The step of performing multimodal feature fusion of the text and the corresponding MR image comprises: PCA or autoencoder is used to compress high-dimensional text embeddings to dimensions that match MR image features, and contrastive learning is used to align the text embeddings with the visual features of the corresponding MR images in the latent space. If the text contains anatomical location information, the anatomical location information is converted into a 3D coordinate mask.
3. The pseudo-CT cross-modality conversion method according to claim 2, characterized in that: The latent space alignment adopts a multimodal feature fusion mechanism, which is expressed as: in, is the cosine similarity, is the temperature parameter, is the positive sample image feature, For text embedding, is the kth image feature, N is the total number of samples in the batch, is the comparison loss value.
4. The pseudo-CT cross-modality conversion method according to claim 3, characterized in that: In the diffusion model, a dynamic convolution module is inserted into the jump connection layer to enhance the multi-scale feature fusion capability, and spectral normalization is used to stabilize adversarial training.
5. The pseudo-CT cross-modality conversion method according to claim 4, characterized in that: The calculation formula in the diffusion model is: in, is the noise scheduling parameter, is the conditional vector of the combination of text and MR image, and Output from the noise prediction network, is the mean of the inverse process, is the covariance of the inverse process, is the probability distribution of the inverse process, is the noisy image at step t, is the noisy image at step t-1, is the transition probability, t is the time step, is the cumulative product result of the previous t steps, N(•) is a Gaussian distribution with a mean of , the covariance is β t I , I is the identity matrix, For the noise prediction network, input , t, and , output predicted noise, is the average value of the cumulative product of the noise scheduling parameters in the previous t steps, is the average value of the cumulative product of the noise scheduling parameters in the previous t-1 steps, is the multiplication symbol, is the noise scheduling parameter associated with the variable s, and the variable s varies from 1 to t.
6. The pseudo-CT cross-modality conversion method according to claim 5, characterized in that: The step of injecting the conditional vector into the diffusion model through a cross-attention layer, guiding the diffusion model using text input, and optimizing the diffusion process through CLIP embedding or Latent Guidance.
7. The pseudo-CT cross-modality conversion method according to claim 6, characterized in that: The total loss function of the first stage is: in, is the total loss function of the first stage, is the noise prediction loss, is the conditional distribution constraint loss, To combat losses, is the perceived loss, is the CLIP loss, , , , as well as are the initial weights of the first stage respectively.
8. The pseudo-CT cross-modality conversion method according to claim 7, characterized in that: The total loss function of the second stage is: in, is the total loss function of the second stage, is the noise prediction loss, is the conditional distribution constraint loss, To combat losses, is the CLIP loss, is the text reconstruction loss, , , as well as is the weight after adjustment in the second stage, To generate the text description extracted by the reverse encoder in pseudo CT, is the original real text description, The MR image x MR and conditional vector c as input to the generator network, is the image encoder, A text encoder.
9. A pseudo-CT cross-modality conversion system, characterized in that: For implementing the pseudo-CT cross-modality conversion method according to any one of claims 1 to 8, the system comprises: a registration module for performing rigid and deformable image registration of the acquired CT image with the MR image and aligning the CT image with the MR image; An extraction module, used to perform text annotation processing on the aligned MR images and extract key information, wherein the key information at least includes anatomical structure, lesion characteristics and spatial relationship; The encoding module is used to encode the extracted key information into a conditional vector that can be used by the diffusion model. At the same time, the text and the corresponding MR image are multimodal feature fused. The text of the key information is converted into a structured conditional vector, namely, text embedding, through LLM. A model building module, used for building a diffusion model, wherein the diffusion model guides and counteracts the diffusion process through boundary contour information to improve the generation efficiency and quality of the image, and further, injects the conditional vector into the diffusion model through a cross attention layer; A training module, used for inputting the result of multimodal feature fusion into the diffusion model to train the diffusion model. The training process is divided into two stages. The first stage only optimizes the conditional generation ability of the diffusion model, and the second stage simultaneously optimizes the semantic alignment ability of the diffusion model and the LLM. The input module is used to input the MR image to be converted into the trained diffusion model and output a pseudo CT image.
Citation Information
Patent Citations
Multimodal image registration method and device using diffusion model, and medium
CN116402865A
PPG generation ECG cross-modal generation method based on diffusion model
CN118333107A