Pseudo-CT cross-modal conversion method and system

Through the combination of the adversarial diffusion model and LLM, cross-modal conversion from MR to pseudo-CT is performed, which solves the problems of easy distortion of image generation, insufficient fidelity and inability to fuse clinical text semantic information in the prior art, and realizes efficient and accurate pseudo-CT image generation, meeting the high-precision requirements of radiotherapy planning.

CN119991751AActive Publication Date: 2025-05-13江西省肿瘤医院(江西省第二人民医院 江西省癌症中心)

Patent Information

Application Number
CN202510473259.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The prior art has the potential for anatomical distortion or artifacts in the cross-modal conversion of MR to pseudoCT. In particular, in complex lesion areas, it is difficult to deal with nonlinear intensity differences between multimodals, it is impossible to integrate clinical text semantic information, lack real-time regulation of the generation process, and it is difficult to meet the high-precision requirements of radiotherapy planning.

Method used

By registering the CT image with the MR image for rigid and deformation images, key information is extracted and encoded into condition vectors available for the diffusion model, combined with LLM to analyze clinical text reports, perform multimodal feature fusion, and drive the adversarial diffusion model to generate high-fidelity and false CT images. The training process of the diffusion model is divided into two stages: the first stage optimizes the condition generation ability, and the second stage optimizes the semantic alignment ability.

Benefits of technology

It realizes the rapid and accurate generation of high-fidelity and false CT images, improves the efficiency and quality of image generation, meets the high-precision requirements of radiotherapy planning, and can regulate the generation process in real time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991751A_ABST
    Figure CN119991751A_ABST
Patent Text Reader

Abstract

The invention provides a pseudo-CT cross-modal conversion method and system, and the method comprises the steps: carrying out the registration and alignment of a CT image and an MR image, employing an LLM to analyze a clinical text report, extracting the description of an anatomical structure, lesion features and a spatial relation, generating a structured condition vector, carrying out the fusion of text semantic embedding and MR image features, and obtaining a pseudo-CT image. The adversarial diffusion model is driven to generate a high-fidelity pseudo CT, the training process of the diffusion model is divided into two stages, only the condition generation capacity of the diffusion model is optimized in the first stage, the semantic alignment capacity of the diffusion model and the LLM is optimized in the second stage at the same time, finally, the MR image to be converted is input into the trained diffusion model, and an accurate pseudo CT image is rapidly output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of pseudo-CT cross-modality conversion, and in particular relates to a pseudo-CT cross-modality conversion method and system. Background Art

[0002] In radiotherapy planning, CT imaging is the gold standard for dose calculation, but it carries the risk of ionizing radiation, and some patients (such as children and pregnant women) need to avoid repeated CT scans. Although MRI is radiation-free and has excellent soft tissue contrast, it lacks electron density information and cannot be used directly for dose calculation. Therefore, the development of cross-modality conversion technology from MR to pseudo-CT has become a core requirement for precision radiotherapy.

[0003] The current mainstream methods have the following limitations: Generative Adversarial Network (GAN)-based methods (such as CycleGAN and pix2pix): generated images are prone to anatomical distortion or artifacts, especially in complex lesion areas (such as bone-soft tissue boundaries), where the fidelity is insufficient; Traditional registration methods (such as Elastix): rely on rigid / non-rigid transformations, have difficulty handling nonlinear intensity differences between multiple modalities, and cannot fuse clinical text semantic information; Single-modal conditional generation models: ignore key semantic descriptions in clinical reports (such as lesion location and morphology), resulting in the generation of pseudo-CT that is out of touch with the individual characteristics of the patient; Lack of dynamic control: Existing methods lack the ability to regulate the generation process in real time (such as adjusting the generation granularity according to the complexity of the lesion), making it difficult to meet the high-precision requirements of radiotherapy planning. Summary of the invention

[0004] Based on this, an embodiment of the present invention provides a pseudo CT cross-modality conversion method and system, aiming to generate pseudo CT images accurately and quickly.

[0005] A first aspect of an embodiment of the present invention provides a pseudo CT cross-modality conversion method, the method comprising: Performing rigid and deformable image registration of the acquired CT image with the MR image, and aligning the CT image with the MR image; Performing text annotation processing on the aligned MR images and extracting key information, wherein the key information at least includes anatomical structure, lesion characteristics and spatial relationship; The extracted key information is encoded into a conditional vector that can be used by the diffusion model. At the same time, the text and the corresponding MR image are multimodal feature fused, where the text of the key information is converted into a structured conditional vector, namely, text embedding, through LLM; Establishing a diffusion model, wherein the diffusion model guides and counteracts the diffusion process through boundary contour information to improve the generation efficiency and quality of the image, and further, injecting the conditional vector into the diffusion model through a cross attention layer; The result of multimodal feature fusion is input into the diffusion model to train the diffusion model. The training process is divided into two stages. The first stage only optimizes the conditional generation ability of the diffusion model, and the second stage simultaneously optimizes the semantic alignment ability of the diffusion model and the LLM. The MR image to be converted is input into the trained diffusion model and a pseudo CT image is output.

[0006] Furthermore, the step of performing multimodal feature fusion of the text and the corresponding MR image includes: PCA or autoencoder is used to compress high-dimensional text embeddings to dimensions that match MR image features, and contrastive learning is used to align the text embeddings with the visual features of the corresponding MR images in the latent space. If the text contains anatomical location information, the anatomical location information is converted into a 3D coordinate mask.

[0007] Furthermore, latent space alignment adopts a multimodal feature fusion mechanism, which is expressed as:

[0008] in, is the cosine similarity, is the temperature parameter, is the positive sample image feature, For text embedding, is the kth image feature, N is the total number of samples in the batch, is the comparison loss value.

[0009] Furthermore, in the diffusion model, a dynamic convolution module is inserted into the jump connection layer to enhance the multi-scale feature fusion capability, and spectral normalization is used to stabilize adversarial training.

[0010] Furthermore, the calculation formula in the diffusion model is:

[0011] in, is the noise scheduling parameter, is the conditional vector of the combination of text and MR image, and Output from the noise prediction network, is the mean of the inverse process, is the covariance of the inverse process, is the probability distribution of the inverse process, is the noisy image at step t, is the noisy image at step t-1, is the transition probability, t is the time step, is the cumulative product result of the previous t steps, N(•) is a Gaussian distribution with a mean of , the covariance isβ t I , I is the identity matrix, For the noise prediction network, input , t, and , output predicted noise, is the average value of the cumulative product of the noise scheduling parameters in the previous t steps, is the average value of the cumulative product of the noise scheduling parameters in the previous t-1 steps, is the multiplication symbol, is the noise scheduling parameter associated with the variable s, and the variable s varies from 1 to t.

[0012] Furthermore, the conditional vector is injected into the diffusion model through a cross-attention layer, the diffusion model is guided by text input, and the diffusion process is optimized by CLIP embedding or Latent Guidance.

[0013] Furthermore, the total loss function of the first stage is:

[0014] in, is the total loss function of the first stage, is the noise prediction loss, is the conditional distribution constraint loss, To combat losses, is the CLIP loss, , , as well as are the initial weights of the first stage respectively.

[0015] Furthermore, the total loss function of the second stage is:

[0016] in, is the total loss function of the second stage, is the noise prediction loss, is the conditional distribution constraint loss, To combat losses, is the CLIP loss, for text reconstruction loss, , , as well as is the weight after adjustment in the second stage, To generate the text description extracted by the reverse encoder in pseudo CT, is the original real text description, The MR image x MR and conditional vector c as input generator network, is the image encoder, A text encoder.

[0017] A second aspect of an embodiment of the present invention provides a pseudo CT cross-modality conversion system, which is used to implement the pseudo CT cross-modality conversion method provided in the first aspect, and the system includes: a registration module for performing rigid and deformable image registration of the acquired CT image with the MR image and aligning the CT image with the MR image; An extraction module, used to perform text annotation processing on the aligned MR images and extract key information, wherein the key information at least includes anatomical structure, lesion characteristics and spatial relationship; The encoding module is used to encode the extracted key information into a conditional vector that can be used by the diffusion model. At the same time, the text and the corresponding MR image are multimodal feature fused. The text of the key information is converted into a structured conditional vector, namely, text embedding, through LLM. A model building module, used for building a diffusion model, wherein the diffusion model guides and counteracts the diffusion process through boundary contour information to improve the generation efficiency and quality of the image, and further, injects the conditional vector into the diffusion model through a cross attention layer; A training module, used for inputting the result of multimodal feature fusion into the diffusion model to train the diffusion model. The training process is divided into two stages. The first stage only optimizes the conditional generation ability of the diffusion model, and the second stage simultaneously optimizes the semantic alignment ability of the diffusion model and the LLM. The input module is used to input the MR image to be converted into the trained diffusion model and output a pseudo CT image.

[0018] A pseudo-CT cross-modal conversion method and system are provided in an embodiment of the present invention. The method aligns the CT image with the MR image, uses LLM to parse the clinical text report, extracts the anatomical structure description, lesion characteristics and spatial relationship, and generates a structured conditional vector. Subsequently, the text semantic embedding is fused with the MR image features to drive the adversarial diffusion model to generate a high-fidelity pseudo-CT. The training process of the diffusion model is divided into two stages. The first stage only optimizes the conditional generation ability of the diffusion model, and the second stage simultaneously optimizes the semantic alignment ability of the diffusion model and the LLM. Finally, the MR image to be converted is input into the trained diffusion model to quickly output an accurate pseudo-CT image. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A flowchart of a pseudo CT cross-modality conversion method provided in the first embodiment of the present invention; Figure 2 A structural block diagram of a pseudo CT cross-modality conversion system provided in Embodiment 2 of the present invention; Figure 3 This is a structural block diagram of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION

[0020] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.

[0021] It should be noted that when an element is referred to as being "fixed to" another element, it may be directly on the other element or there may be a central element. When an element is considered to be "connected to" another element, it may be directly connected to the other element or there may be a central element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0023] Embodiment 1 See also Figure 1 , Figure 1 A flowchart of an implementation of a pseudo CT cross-modality conversion method provided in the first embodiment of the present invention is shown. The pseudo CT cross-modality conversion method specifically includes steps S01 to S06.

[0024] Step S01 : performing rigid and deformation image registration on the acquired CT image and the MR image, and aligning the CT image with the MR image.

[0025] Specifically, CT images and MR images of 200 patients with pelvic tumors were collected. It can be understood that the CT images and MR images contain data information. Rigid and deformable image registration was then performed, and the CT images were aligned with the MR images to obtain a Dice similarity coefficient greater than 0.98, indicating that there is good spatial correspondence between the modalities.

[0026] Step S02, performing text annotation processing on the aligned MR images and extracting key information, wherein the key information at least includes anatomical structure, lesion characteristics and spatial relationship.

[0027] The process of text annotation is as follows: 1) Delete sensitive information such as patient name and ID number (regular expression matching) to de-anonymize; replace hospital number with anonymous identifier (such as `Patient_001`); 2) Use the medical ontology library to convert non-standard text descriptions into standardized terms, such as the original text: "right temporal lobe space-occupying lesion" → standardized: "right temporal lobe space-occupying lesion"; 3) Segment the text by anatomical region (e.g., “brain description” is separated from “abdomen description”); use punctuation marks such as periods and semicolons to separate independent semantic units.

[0028] Furthermore, BioBERT is used to parse imaging data and clarify key information (such as organ location, lesion characteristics, soft tissue name, bone boundary, etc.). Specifically, the domain-specific BioBERT is used for entity recognition, such as: anatomical parts (such as "right temporal lobe", "liver", "bone", "kidney", etc.); quantitative parameters (such as "2.1cm×1.8cm"). The association between anatomical structure and outer contour is constructed based on rule templates (such as SemRep) or LLM fine-tuning models (such as "right kidney→boundary position").

[0029] Step S03, encode the extracted key information into a conditional vector available to the diffusion model, and at the same time, perform multimodal feature fusion on the text and the corresponding MR image, wherein the text of the key information is converted into a structured conditional vector, i.e., text embedding, through LLM.

[0030] It should be noted that the key information encoding can be completed through the embedding layer of LLM (such as the `[CLS]` vector or entity-level embedding of BioBERT). Furthermore, the step of multimodal feature fusion of the text and the corresponding MR image includes: PCA or autoencoder is used to compress high-dimensional text embedding to a dimension that matches the MR image features, and contrastive learning is used to align the text embedding with the visual features of the corresponding MR image in the latent space. If the text contains anatomical location information, the anatomical location information is converted into a 3D coordinate mask (aligned with the MR image space). Specifically, the latent space alignment adopts a multimodal feature fusion mechanism, which is expressed as:

[0031] in, is the cosine similarity, is the temperature parameter, is the positive sample image feature, For text embedding, is the kth image feature, N is the total number of samples in the batch, is the comparison loss value.

[0032] In order to ensure the accuracy of the trained diffusion model, the association between text and image data will be confirmed. Specifically, automatic filtering is performed based on the `Body Part Examined` field in the DICOM metadata to ensure that the anatomical region described in the text matches the MR image ROI (Region of Interest) (for example, when describing "liver disease", MR must include abdominal scans), and the radiologist confirms the correspondence between the text description and the image lesion (such as the location and size of the anatomical structure boundary).

[0033] Step S04, establishing a diffusion model, wherein the diffusion model guides and counteracts the diffusion process through boundary contour information to improve the generation efficiency and quality of the image, and further, injects the conditional vector into the diffusion model through a cross attention layer.

[0034] In this embodiment, a novel adversarial diffusion model based on boundary contour information guidance is proposed as the main framework for generating synthetic CT images from MR images. The advantages of the diffusion model and the adversarial generative network (GAN) are combined to guide the boundary contour information and the adversarial diffusion process ( ), improve the generation efficiency and quality of images. In addition, a dynamic convolution module is inserted in the jump connection layer to enhance the multi-scale feature fusion capability, and spectral normalization is used to stabilize adversarial training. The specific mathematical formula is as follows, where the effective adversarial diffusion forward noise addition and inverse process are expressed as:

[0035] in, is the noise scheduling parameter, controlling the noise intensity of the tth step (0< <1), is the conditional vector of the combination of text and MR image, and Output from the noise prediction network, is the mean of the inverse process, is the covariance of the inverse process, is the probability distribution of the inverse process, modeled by the neural network parameters θ, is the noisy image at step t, is the noisy image at step t-1, is the transition probability, t is the time step, is the cumulative product result of the previous t steps, N(•) is a Gaussian distribution with a mean of , the covariance is β t I , Iis the identity matrix to simplify the covariance structure and ensure that the noise is evenly distributed in all spatial dimensions. For the noise prediction network, in the inverse process, by predicting the noise and calculating the mean , guide the generation of high-fidelity pseudo CT images, input , t, and , output predicted noise, is the average value of the cumulative product of the noise scheduling parameters in the previous t steps, is the average value of the cumulative product of the noise scheduling parameters in the previous t-1 steps, is the multiplication symbol, is the noise scheduling parameter associated with the variable s, and the variable s varies from 1 to t.

[0036] It should be noted that in the main framework, the value of t is significantly greater than 1. Considering that t is much larger than 1, the large-step backward process lacks a closed-form expression. To solve this problem, the present invention introduces a primitive domain traditional mapper to obtain the complex transition probability in the large-step conditional diffusion model. The idea of ​​adversarial projectors with cycle-consistent structures is borrowed and integrated into the diffusion process. θ Input x with source modality y t and the newly introduced corresponding boundary mask mask as input, , extract intermediate features for step-by-step denoising, is the generated pseudo CT image, m is the boundary mask, and x is the input image of the source modality. In addition, the learnable embedding mask calculated based on time t is incorporated into these extracted features. Then, a discriminant factor is introduced to distinguish the features generated by Gen θ The generated estimated denoised distribution and actual denoised samples. At the same time, time embedding is introduced as a bias term for feature mapping.

[0037] In order to generate the target synthetic CT (sCT, pseudo-CT), the original MRI image with the same anatomical structure can be used as the prior information for the inverse diffusion step. The source image corresponding to each sCT image in the training dataset is estimated through non-diffusion and diffusion processes. Previous diffusion-based image conversion methods usually rely on global feature matching to achieve comprehensive image conversion, but this method often leads to inefficient feature learning due to the inclusion of a large number of irrelevant background regions, which in turn affects the synthesis results. To solve this problem, the present invention proposes to explicitly integrate body contour guidance into the feature learning framework. Specifically, the boundary guidance is integrated into the conditional generation process of its diffusion and non-diffusion networks, which effectively reduces the influence of background noise and ensures that the model learning is focused on the human body area. Through this efficient and selective feature learning paradigm, the model can capture key features such as internal anatomical details and external body contours, significantly improving the feature learning effect of internal anatomical synthesis, thereby obtaining more accurate and reliable results. In order to achieve unsupervised learning, the real target data is compared with its generated data using cycle consistency loss. In the diffusion process, the reconstructed image is called the synthetic target image; in the non-diffusion module, the estimated source image is mapped to the target domain through the generator. The diffusion and non-diffusion modules are jointly trained without the need for a pre-training process. In order to calculate and , when the sampling network parameterizes the denoising distribution When , the operation of generating the distribution is as follows:

[0038] Among them, H θ (•) is the parameterized denoising conditional distribution, which means that given the current image state x t , conditional input y and body contour mask mask, generate the conditional probability distribution of the t-kth step image. It can be understood that x t−k is the denoising target image state at the t-kth step in the diffusion process, x t is the current image state at the tth step in the diffusion process, y is the conditional input, usually the source domain image (such as the original MRI image), providing anatomical structure prior information, mask is the body contour mask, which is used to explicitly guide the generation process, limit the model to focus on the human body area, and reduce background noise interference. θ The generator network is responsible for generating intermediate images based on conditional input , L (•) is the log-likelihood function, which measures the probability of generating an image With the target distribution x t−kThe model diffusion module is used to estimate and synthesize the target domain image from the original domain data output by the non-diffusion module. To achieve this goal, two traditional adversarial diffusion methods are used, each with its own discriminator. In each step of the inverse process, the generator first makes a deterministic estimate of the denoised target image, and then uses the denoised distribution unique to each image modality to synthesize the target image.

[0039] Furthermore, the diffusion model is guided by text input and the diffusion process is optimized by CLIP embedding or Latent Guidance, where the text description is converted into an embedding vector using CLIP text encoder or LLM (such as BioBERT) Insert a cross-attention module in each downsampling and upsampling layer of U-Net to align text embedding with image features:

[0040] in, , is the image feature projection, , which is a text conditional projection. The visual features of the pseudo-CT generated by the CLIP image encoder are aligned with the text semantics to enhance the conditional control:

[0041] in, For the CLIP image encoder, For the CLIP text encoder.

[0042] Step S05, inputting the result of multimodal feature fusion into the diffusion model to train the diffusion model. The training process is divided into two stages. The first stage only optimizes the conditional generation ability of the diffusion model, and the second stage simultaneously optimizes the semantic alignment ability of the diffusion model and the LLM.

[0043] It should be noted that in the initial training phase of the diffusion model (LLM parameter freezing), only the conditional generation capability of the diffusion model is optimized, i.e., the first phase. At this time, LLM is a fixed text encoder and does not participate in back propagation. In the end-to-end joint training phase (unfreezing some LLM parameters), the semantic alignment capability of the diffusion model and LLM is optimized at the same time, i.e., the second phase.

[0044] Specifically, the total loss function of the first stage is: (1) in, is the total loss function of the first stage, is the noise prediction loss, is the conditional distribution constraint loss, To combat losses, is the CLIP loss, , , as well as are the initial weights of the first stage respectively. Furthermore, the formulas for noise prediction loss and conditional distribution constraint loss are:

[0045] Among them, x0 is the real CT image, is the real noise, p(c) is the standard normal distribution, the constraint vector distribution, For x0, t and The mathematical expectation of KL is the divergence, which is used to measure the difference between two probability distributions. For a given Under the conditions of and text, The probability distribution of . In addition, the adversarial loss generator, discriminator loss and perceptual loss integrated in the diffusion model are:

[0046] in, is the i-th layer feature of the pre-trained VGG network, The expected score for the samples generated by the generator to be judged as true by the discriminator, For x t The joint distribution expectation of and t is obtained, D is the discriminator, G is the generator, For Calculate the log probability of the discriminator D, The logarithmic probability that the discriminator judges the forged sample generated by the generator as fake is calculated. In this embodiment, , , , as well as They are 1, 0.1, 0.5, 0.2 and 0.05 respectively.

[0047] The total loss function of the second stage is: (2)

[0048] in, is the total loss function of the second stage, is the noise prediction loss, is the conditional distribution constraint loss, To combat losses, is the CLIP loss, for text reconstruction loss, , , as well as is the weight after adjustment in the second stage, To generate the text description extracted by the reverse encoder in pseudo CT, is the original real text description, is the generator network, with MR image x MR and condition vector c (such as noise or other control parameters) as input, and output a synthetic pseudo CT image, is an image encoder that maps the generated pseudo CT image to the text feature space. is a text encoder that encodes the true text description into a feature vector.

[0049] It should be noted that the diffusion model is finally trained with mixed precision, using FP16 precision to accelerate calculations, and scaling the gradient (Scale=1024) to prevent underflow. When the gradient appears Inf / NaN, a dynamic loss scaling strategy is used to automatically reduce the scaling factor. Finally, the conditional total loss is integrated to jointly optimize the deep semantic alignment of the LLM and the diffusion model, which is the total loss function of the second stage.

[0050] It can be understood that when performing end-to-end fine-tuning, a text reconstruction loss is introduced to constrain the generated pseudo-CT image to be able to reconstruct the original text description through the inverse model, thereby enhancing the alignment between text and image. The total loss in formula (2) is the additional text reconstruction loss when the LLM and the diffusion model are optimized together in the subsequent joint training or fine-tuning stage. Formula (1) is the initial training diffusion model, and the latter is the end-to-end optimization after integrating LLM. In the second stage, some LLM layers are unfrozen and text reconstruction loss is added to optimize the entire system in an end-to-end manner. The two total loss functions correspond to different training stages and optimization objectives. The total loss in formula (2) is introduced in the end-to-end fine-tuning stage of formula (1), and the text reconstruction loss is added to strengthen the consistency between text and generated images. Through the above strategy, LLM and the diffusion model achieve deep collaboration, while ensuring the generation quality, improving the computational efficiency through dynamic control, and meeting the high precision and real-time requirements of pseudo-CT for precision radiotherapy.

[0051] Through a large number of ablation experiments and quantitative analysis, all relevant parameters have been reasonably determined through multiple adjustment processes. The proposed deep network architecture is implemented based on PyTorch. The training and validation processes are carried out on a computer equipped with an Nvidia RTXA6000 GPU (48G video memory), which supports CUDA acceleration and can meet the computational requirements of the diffusion model. All networks are trained using the Adam optimizer to minimize the L1 loss function. The generator and discriminator are trained in an alternating manner, with a batch size of 1. The loss weight η Φand η θ are all set to 0.5, and the number of diffusion steps T / k is set to 4. The network was trained for a total of 200 epochs, with a batch size of 1 and a diffusion step of 8. The initial learning rate was set to 0.0001 for the first 100 epochs, and the learning rate was halved every 100 epochs thereafter. During training, the trade-off parameter was set to 1. Using the trained generative model, approximately 30 hours of computational time was required. Generating sCT for a new case based on MR data takes only a few tenths of a second.

[0052] Step S06: input the MR image to be converted into the trained diffusion model, and output a pseudo CT image.

[0053] It should be noted that after the pseudo-CT image is generated, the LLM model is used to automatically generate a report to explain the key anatomical or pathological features in the generated results. Let LLM analyze the generated pseudo-CT, automatically detect artifacts (Artifacts) and combine LLM to generate correction suggestions to guide the secondary optimization of the diffusion model. Quantitative evaluation includes: using indicators such as SSIM (structural similarity), MAE (mean absolute error), PSNR (peak signal-to-noise ratio) Dice coefficient (organ segmentation consistency). Qualitative evaluation includes radiologists evaluating the anatomical rationality and clinical usability of the pseudo-CT.

[0054] In summary, the pseudo-CT cross-modal conversion method in the above-mentioned embodiment of the present invention aligns the CT image with the MR image, uses LLM to parse the clinical text report, extracts the anatomical structure description, lesion characteristics and spatial relationship, and generates a structured conditional vector. Subsequently, the text semantic embedding is fused with the MR image features to drive the adversarial diffusion model to generate a high-fidelity pseudo-CT. The training process of the diffusion model is divided into two stages. The first stage only optimizes the conditional generation ability of the diffusion model, and the second stage simultaneously optimizes the semantic alignment ability of the diffusion model and the LLM. Finally, the MR image to be converted is input into the trained diffusion model to quickly output an accurate pseudo-CT image.

[0055] Embodiment 2 See also Figure 2 , Figure 2 : is a structural block diagram of a pseudo CT cross-modal conversion system provided in Embodiment 2 of the present invention. The pseudo CT cross-modal conversion system 200 includes: a registration module 21, an extraction module 22, an encoding module 23, a model building module 24, a training module 25 and an input module 26, wherein: A registration module 21, for performing rigid and deformable image registration of the acquired CT image and the MR image, and aligning the CT image with the MR image; An extraction module 22, configured to perform text annotation processing on the aligned MR images and extract key information, wherein the key information at least includes anatomical structure, lesion characteristics and spatial relationship; The encoding module 23 is used to encode the extracted key information into a conditional vector that can be used by the diffusion model. At the same time, the text and the corresponding MR image are subjected to multimodal feature fusion, wherein the text of the key information is converted into a structured conditional vector, i.e., text embedding, by LLM. Specifically, PCA or autoencoder is used to compress the high-dimensional text embedding to a dimension that matches the MR image features, and the text embedding is aligned with the visual features of the corresponding MR image in the latent space by contrastive learning. If the text contains anatomical position information, the anatomical position information is converted into a 3D coordinate mask. In addition, the latent space alignment adopts a multimodal feature fusion mechanism, which is expressed as:

[0056] in, is the cosine similarity, is the temperature parameter, is the positive sample image feature, For text embedding, is the kth image feature, N is the total number of samples in the batch, is the comparison loss value; The model building module 24 is used to build a diffusion model. The diffusion model guides and counteracts the diffusion process through boundary contour information to improve the generation efficiency and quality of the image. In addition, the conditional vector is injected into the diffusion model through the cross attention layer. In the diffusion model, a dynamic convolution module is inserted in the jump connection layer to enhance the multi-scale feature fusion capability. At the same time, spectral normalization is used to stabilize the adversarial training. The calculation formula in the diffusion model is:

[0057] in, is the noise scheduling parameter, is the conditional vector of the combination of text and MR image, and Output from the noise prediction network, is the mean of the inverse process, is the covariance of the inverse process, is the probability distribution of the inverse process, is the noisy image at step t, is the noisy image at step t-1, is the transition probability, t is the time step, is the cumulative product result of the previous t steps, N(•) is a Gaussian distribution with a mean of , the covariance is βt I , I is the identity matrix, For the noise prediction network, input , t, and , output predicted noise, is the average value of the cumulative product of the noise scheduling parameters in the previous t steps, is the average value of the cumulative product of the noise scheduling parameters in the previous t-1 steps, is the multiplication symbol, is the noise scheduling parameter associated with the variable s, which varies from 1 to t. In addition, the diffusion model is guided by text input and the diffusion process is optimized by CLIP embedding or Latent Guidance; The training module 25 is used to input the result of multimodal feature fusion into the diffusion model to train the diffusion model. The training process is divided into two stages. The first stage only optimizes the conditional generation ability of the diffusion model, and the second stage simultaneously optimizes the semantic alignment ability of the diffusion model and the LLM. The total loss function of the first stage is:

[0058] in, is the total loss function of the first stage, is the noise prediction loss, is the conditional distribution constraint loss, To combat losses, is the perceived loss, is the CLIP loss, , , , as well as are the initial weights of the first stage respectively, and the total loss function of the second stage is:

[0059] in, is the total loss function of the second stage, is the noise prediction loss, is the conditional distribution constraint loss, To combat losses, is the CLIP loss, for text reconstruction loss, , , as well as is the weight after adjustment in the second stage, To generate the text description extracted by the reverse encoder in pseudo CT, is the original real text description, The MR image x MR and conditional vector c as input generator network, is the image encoder, is a text encoder; The input module 26 is used to input the MR image to be converted into the trained diffusion model and output a pseudo CT image.

[0060] Embodiment 3 Another aspect of the present invention provides an electronic device, see Figure 3 , shown is an electronic device in Embodiment 3 of the present invention, comprising a memory 20, a processor 10, and a computer program 30 stored in the memory and executable on the processor, wherein the processor 10 implements the pseudo-CT cross-modality conversion method as described above when executing the computer program 30.

[0061] In some embodiments, the processor 10 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor or other data processing chip, used to run program codes or process data stored in the memory 20, such as executing access restriction programs.

[0062] The memory 20 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 20 may be an internal storage unit of an electronic device, such as a hard disk of the electronic device. In other embodiments, the memory 20 may also be an external storage device of an electronic device, such as a plug-in hard disk equipped on the electronic device, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), etc. Further, the memory 20 may also include both an internal storage unit and an external storage device of the electronic device. The memory 20 may be used not only to store application software and various types of data of the electronic device, but also to temporarily store data that has been output or is to be output.

[0063] It should be pointed out that Figure 3 The structure shown does not constitute a limitation on the electronic device. In other embodiments, the electronic device may include fewer or more components than those shown in the figure, or combine certain components, or arrange the components differently.

[0064] The embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the pseudo-CT cross-modality conversion method as described above is implemented.

[0065] Those skilled in the art will appreciate that the logic and / or steps represented in the flowchart or otherwise described herein, for example, may be considered as an ordered list of executable instructions for implementing logical functions, and may be specifically implemented in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For purposes of this specification, "computer-readable medium" may be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0066] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.

[0067] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or a combination thereof: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0068] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0069] The above embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the attached claims.

Claims

1. A pseudo-CT cross-modality conversion method, characterized in that: The method comprises: Performing rigid and deformable image registration of the acquired CT image with the MR image, and aligning the CT image with the MR image; Performing text annotation processing on the aligned MR images and extracting key information, wherein the key information at least includes anatomical structure, lesion characteristics and spatial relationship; The extracted key information is encoded into a conditional vector that can be used by the diffusion model. At the same time, the text and the corresponding MR image are multimodal feature fused, where the text of the key information is converted into a structured conditional vector, namely, text embedding, through LLM; Establishing a diffusion model, wherein the diffusion model guides and counteracts the diffusion process through boundary contour information to improve the generation efficiency and quality of the image, and further, injecting the conditional vector into the diffusion model through a cross attention layer; The result of multimodal feature fusion is input into the diffusion model to train the diffusion model. The training process is divided into two stages. The first stage only optimizes the conditional generation ability of the diffusion model, and the second stage simultaneously optimizes the semantic alignment ability of the diffusion model and the LLM. The MR image to be converted is input into the trained diffusion model and a pseudo CT image is output.

2. The pseudo-CT cross-modality conversion method according to claim 1, characterized in that: The step of performing multimodal feature fusion of the text and the corresponding MR image comprises: PCA or autoencoder is used to compress high-dimensional text embeddings to dimensions that match MR image features, and contrastive learning is used to align the text embeddings with the visual features of the corresponding MR images in the latent space. If the text contains anatomical location information, the anatomical location information is converted into a 3D coordinate mask.

3. The pseudo-CT cross-modality conversion method according to claim 2, characterized in that: The latent space alignment adopts a multimodal feature fusion mechanism, which is expressed as: in, is the cosine similarity, is the temperature parameter, is the positive sample image feature, For text embedding, is the kth image feature, N is the total number of samples in the batch, is the comparison loss value.

4. The pseudo-CT cross-modality conversion method according to claim 3, characterized in that: In the diffusion model, a dynamic convolution module is inserted into the jump connection layer to enhance the multi-scale feature fusion capability, and spectral normalization is used to stabilize adversarial training.

5. The pseudo-CT cross-modality conversion method according to claim 4, characterized in that: The calculation formula in the diffusion model is: in, is the noise scheduling parameter, is the conditional vector of the combination of text and MR image, and Output from the noise prediction network, is the mean of the inverse process, is the covariance of the inverse process, is the probability distribution of the inverse process, is the noisy image at step t, is the noisy image at step t-1, is the transition probability, t is the time step, is the cumulative product result of the previous t steps, N(•) is a Gaussian distribution with a mean of , the covariance is β t I , I is the identity matrix, For the noise prediction network, input , t, and , output predicted noise, is the average value of the cumulative product of the noise scheduling parameters in the previous t steps, is the average value of the cumulative product of the noise scheduling parameters in the previous t-1 steps, is the multiplication symbol, is the noise scheduling parameter associated with the variable s, and the variable s varies from 1 to t.

6. The pseudo-CT cross-modality conversion method according to claim 5, characterized in that: The step of injecting the conditional vector into the diffusion model through a cross-attention layer, guiding the diffusion model using text input, and optimizing the diffusion process through CLIP embedding or Latent Guidance.

7. The pseudo-CT cross-modality conversion method according to claim 6, characterized in that: The total loss function of the first stage is: in, is the total loss function of the first stage, is the noise prediction loss, is the conditional distribution constraint loss, To combat losses, is the perceived loss, is the CLIP loss, , , , as well as are the initial weights of the first stage respectively.

8. The pseudo-CT cross-modality conversion method according to claim 7, characterized in that: The total loss function of the second stage is: in, is the total loss function of the second stage, is the noise prediction loss, is the conditional distribution constraint loss, To combat losses, is the CLIP loss, is the text reconstruction loss, , , as well as is the weight after adjustment in the second stage, To generate the text description extracted by the reverse encoder in pseudo CT, is the original real text description, The MR image x MR and conditional vector c as input to the generator network, is the image encoder, A text encoder.

9. A pseudo-CT cross-modality conversion system, characterized in that: For implementing the pseudo-CT cross-modality conversion method according to any one of claims 1 to 8, the system comprises: a registration module for performing rigid and deformable image registration of the acquired CT image with the MR image and aligning the CT image with the MR image; An extraction module, used to perform text annotation processing on the aligned MR images and extract key information, wherein the key information at least includes anatomical structure, lesion characteristics and spatial relationship; The encoding module is used to encode the extracted key information into a conditional vector that can be used by the diffusion model. At the same time, the text and the corresponding MR image are multimodal feature fused. The text of the key information is converted into a structured conditional vector, namely, text embedding, through LLM. A model building module, used for building a diffusion model, wherein the diffusion model guides and counteracts the diffusion process through boundary contour information to improve the generation efficiency and quality of the image, and further, injects the conditional vector into the diffusion model through a cross attention layer; A training module, used for inputting the result of multimodal feature fusion into the diffusion model to train the diffusion model. The training process is divided into two stages. The first stage only optimizes the conditional generation ability of the diffusion model, and the second stage simultaneously optimizes the semantic alignment ability of the diffusion model and the LLM. The input module is used to input the MR image to be converted into the trained diffusion model and output a pseudo CT image.

Citation Information

Patent Citations

  • Multimodal image registration method and device using diffusion model, and medium

    CN116402865A

  • Image synthesis method and device, electronic equipment and storage medium

    CN117036184A

  • Diffusion image generation method and system based on retrieval and segmentation enhancement

    CN117725247A

  • PPG generation ECG cross-modal generation method based on diffusion model

    CN118333107A

  • Unsupervised pseudo-CT adversarial diffusion model construction method and system, medium and equipment

    CN118864288A

Cited By

  • Pseudo-CT cross-modal conversion method and system of multi-modal feature coupling diffusion model

    CN121120367A

  • Pseudo ct cross-modal conversion method and system based on multi-modal feature coupled diffusion model

    CN121120367B