Pseudo ct cross-modal conversion method and system based on multi-modal feature coupled diffusion model
By fusing frequency domain and textual semantic features of MR and CT images using a multimodal feature coupling diffusion model, the problem of insufficient accuracy in pseudo-CT image generation in existing technologies is solved, achieving high-precision pseudo-CT image generation that meets the needs of radiotherapy planning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-13
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, cross-modal conversion methods from MR images to pseudo-CT images are insufficient in utilizing frequency domain features, resulting in limited physical fidelity of the generated pseudo-CT images. This is especially true in complex regions where anatomical distortion or artifacts are severe, failing to meet the high-precision requirements of radiotherapy planning.
By performing rigid and non-rigid registration on MR and CT images, and combining frequency domain and textual semantic features, a structured conditional vector is generated. Multimodal feature coupling is performed in the diffusion model, and dynamic weighted fusion is carried out using a cross-attention mechanism and a frequency domain mapping function to generate a pseudo-CT image.
It improves the generation accuracy and alignment precision of pseudo-CT images, ensures the optimization of high-frequency details and low-frequency contours, and generates pseudo-CT images that fit the patient's anatomical characteristics, meeting the high-precision requirements of radiotherapy planning.
Smart Images

Figure CN121120367B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a pseudo-CT cross-modal conversion method and system based on a multimodal feature coupling diffusion model. Background Technology
[0002] In radiotherapy planning, CT imaging is the gold standard for dose calculation, but it carries a significant risk of ionizing radiation, especially for sensitive populations (such as children and pregnant women), requiring strict avoidance of repeated scans. While MRI offers the advantages of being radiation-free and providing excellent soft tissue contrast, it lacks electron density information and cannot be directly used for dose calculation. Therefore, developing high-precision cross-modal conversion technology from MR images to pseudo-CT images has become a core requirement for precision radiotherapy.
[0003] While existing technologies employ methods that combine spatial domain image features and textual semantic features to perform cross-modal conversion of pseudo-CT images for relatively accurate and rapid generation of pseudo-CT images, these methods suffer from significant limitations in utilizing frequency domain features. The frequency domain details and global morphological features (such as organ contour periodicity) of anatomical structures in MR images are not fully exploited, resulting in limited physical fidelity of the generated pseudo-CT images.
[0004] Furthermore, the existing method has the following drawbacks: it still produces anatomical distortions or artifacts in complex areas (such as bone-soft tissue boundaries). The root cause is that the spatial domain generation process has insufficient modeling ability for frequency domain features and cannot coordinate the optimization of high-frequency details and low-frequency contours; it also lacks dynamic control and adaptive perception of frequency domain features (such as wavelet energy distribution and spectral attenuation characteristics), resulting in a disconnect between the generated results and the patient's anatomical characteristics, making it difficult to meet the high-precision real-time requirements of radiotherapy planning. Summary of the Invention
[0005] Based on this, the purpose of this invention is to provide a pseudo-CT cross-modal conversion method and system for a multimodal feature coupling diffusion model, aiming to solve the problem that there is a lack of a high-precision and high-accuracy pseudo-CT cross-modal conversion method and system for a multimodal feature coupling diffusion model in the prior art.
[0006] A pseudo-CT cross-modal conversion method for a multimodal feature coupling diffusion model according to an embodiment of the present invention includes:
[0007] The wavelet energy spectrum and FFT logarithmic spectrum of the acquired MR images are pre-calculated, and the global anatomical structure of the acquired MR and CT images is initially aligned by rigid registration. Furthermore, the B-spline deformation field is optimized under the guidance of FFT phase spectrum to complete the non-rigid registration of MR and CT images, thereby aligning the CT and MR images.
[0008] The clinical reports corresponding to the aligned MR images are subjected to joint text semantic and frequency domain annotation processing.
[0009] Multi-dimensional key information is extracted and encoded. After encoding, the dimensions are unified and dynamically weighted and fused to generate a structured conditional vector. In the skip connection layer of the diffusion model, the multimodal features in the structured conditional vector are associated through a cross-attention mechanism combined with a frequency domain mapping function.
[0010] The pre-defined diffusion model is trained based on the aligned image, and the MR image to be converted is input into the trained diffusion model to output a pseudo-CT image.
[0011] Furthermore, the pseudo-CT cross-modal conversion method of the multimodal feature coupling diffusion model according to the above embodiments of the present invention may also have the following additional technical features:
[0012] Furthermore, the steps of pre-calculating a frequency domain reference datum for the acquired MR images, performing rigid registration of the acquired CT images and MR images, and performing non-rigid image registration based on frequency domain constraints according to the frequency domain reference datum include:
[0013] Pre-calculate wavelet energy maps and FFT logarithmic maps for the acquired MR images;
[0014] The global anatomical structure was initially aligned by rigid registration, and the B-spline deformation field was optimized under the guidance of FFT phase spectrum to complete non-rigid registration.
[0015] Determine whether the Dice coefficients and wavelet mutual information of the registered CT and MR images are greater than the corresponding preset thresholds;
[0016] If so, confirm that the registered CT and MR images meet the requirements and complete the registration.
[0017] Furthermore, the steps for jointly annotating the clinical reports corresponding to the aligned MR images using text semantics and frequency domain include:
[0018] The clinical reports corresponding to the aligned MR images are regularized, and sensitive information is filtered and replaced with anonymous identifiers;
[0019] The processed clinical report text descriptions are converted into standardized terms, and the standardized terms are bound to predefined frequency domain patterns to establish a semantic frequency domain mapping rule base.
[0020] The text, which is standardized and associated with frequency domain features, is divided into paragraphs based on organ systems, and frequency domain markers are inserted in key regions of the text to facilitate subsequent feature extraction.
[0021] Furthermore, the steps of extracting and encoding multi-dimensional key information, unifying the dimensions after encoding, and dynamically weighting and fusing the data to generate structured conditional vectors include:
[0022] The processed clinical reports are processed using LLM to obtain text semantic embedding vectors;
[0023] High-frequency detail feature vectors are extracted by performing a three-level discrete wavelet transform on the aligned MR image;
[0024] Fast Fourier Transform is performed on the aligned MR image to extract global low-frequency morphological feature vectors;
[0025] A structured conditional vector is constructed based on the text semantic embedding vector, high-frequency detail feature vector, and global low-frequency morphological feature vector;
[0026] The formula for the structured condition vector is:
[0027]
[0028] in, For text semantic embedding vectors, These are high-frequency detail feature vectors. This is a global low-frequency morphological feature vector. , and For the corresponding fusion weights.
[0029] Furthermore, the cross-attention mechanism is expressed as follows:
[0030]
[0031] in, Q The query vector is derived from image features. K and V A key-value vector derived from a condition vector. c , is the dimension of the key vector.
[0032] Furthermore, the mathematical formula corresponding to the preset diffusion model is:
[0033]
[0034]
[0035] in, For the first t The image after adding noise. For the first t 1 The image of the step, where I is the identity matrix. For noise scheduling parameters,α t To control the proportion of the original signal retained, c Let be the condition vector, and be The normal distribution (Gaussian distribution) controls the size of the jump step during the reverse process. For parameterized denoising distribution, For a given previous step image Current step image under the condition The distribution For the image at the current step Given the conditional vector c, predict the preceding... k Step Image The distribution k Adjust parameters for step size. For text vectors, For boundary mask information, For wavelet transform, This is a Fourier transform.
[0036] Furthermore, after the step of inputting the MR image to be converted into the trained diffusion model to output a pseudo-CT image, the following steps are included:
[0037] Key features are output and artifacts are detected in pseudo-CT images generated by LLM analysis;
[0038] Based on LLM, correction suggestions are generated according to artifacts and key features to guide the optimization of the diffusion model.
[0039] Another object of the present invention is a pseudo-CT cross-modal conversion system for a multimodal feature coupling diffusion model, the system comprising:
[0040] The image alignment module is used to pre-calculate the frequency domain reference benchmark for the acquired MR image, perform rigid registration of the acquired CT image and MR image, and perform non-rigid image registration based on the frequency domain reference benchmark to align the CT image and MR image.
[0041] The annotation module is used to perform joint text semantic and frequency domain annotation processing on the clinical reports corresponding to the aligned MR images;
[0042] The vector construction module is used to extract and encode multi-dimensional key information. After encoding, the dimensions are unified and dynamically weighted and fused to generate a structured conditional vector. In the skip connection layer of the diffusion model, the multimodal features in the structured conditional vector are associated through a cross-attention mechanism combined with a frequency domain mapping function.
[0043] The image generation module is used to train a preset diffusion model based on the aligned image, and input the MR image to be converted into the trained diffusion model to output a pseudo-CT image.
[0044] Another objective of this invention is to provide a storage medium storing a computer program that, when executed by a processor, implements the steps of the pseudo-CT cross-modal conversion method of the multimodal feature coupling diffusion model described above.
[0045] Another objective of this invention is to provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the pseudo-CT cross-modal conversion method of the multimodal feature coupling diffusion model described above.
[0046] This invention calculates a frequency domain reference basis for MR images, enabling rigid registration of CT and MR images while simultaneously performing non-rigid registration under frequency domain constraints. This improves the efficiency and accuracy of alignment. Subsequently, semantic parsing of clinical reports and joint semantic and frequency domain annotation result in encoded and dimensionally unified feature vectors encompassing semantic, frequency, and spatial domains. These vectors are then weighted and fused to obtain structured conditional vectors, simultaneously ensuring optimization of high-frequency details and low-frequency contours. Furthermore, semantic, frequency, and spatial features are injected into the diffusion model through a hybrid mechanism, enabling dynamic conditional input and enhancing adaptive perception of frequency domain features. This results in pseudo-CT images that closely match the patient's anatomical characteristics. Therefore, this invention addresses the lack of a high-precision and accurate multimodal feature coupling diffusion model for pseudo-CT cross-modal conversion in existing technologies. Attached Figure Description
[0047] Figure 1 This is a flowchart of the pseudo-CT cross-modal conversion method of the multimodal feature coupling diffusion model in the first embodiment of the present invention;
[0048] Figure 2 This is a schematic diagram of the results of the pseudo-CT cross-modal conversion system of the multimodal feature coupling diffusion model in the second embodiment of the present invention;
[0049] Figure 3 This is a schematic diagram of the structure of the electronic device in the third embodiment of the present invention;
[0050] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation
[0051] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.
[0052] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0053] Example 1
[0054] Please see Figure 1 The figure shows a pseudo-CT cross-modal conversion method for a multimodal feature coupling diffusion model in the first embodiment of the present invention. The method specifically includes steps S01-S04.
[0055] S01, pre-calculate the wavelet energy spectrum and FFT logarithmic spectrum of the acquired MR image, and preliminarily align the global anatomical structure of the acquired MR image and CT image through rigid registration, and optimize the B-spline deformation field under the guidance of FFT phase spectrum to complete the non-rigid registration of MR image and CT image, thereby aligning CT image and MR image.
[0056] In practical implementation, MR images can be processed using a three-level Db4 wavelet decomposition to obtain wavelet energy maps. Db4 (Daubechies 4 wavelet) possesses tight support and good detail preservation capabilities; the three-level decomposition can extract high-frequency details (such as bone edge sharpness) layer by layer, providing "detail anchors" for subsequent frequency domain feature matching. Meanwhile, the FFT logarithmic amplitude spectrum is used to transform the MR image from the spatial domain to the frequency domain. The logarithmic transformation enhances low-frequency signals, corresponding to the periodic features of the organ's global contour, avoiding high-frequency noise interference, and forming "global morphological anchors." Furthermore, during registration, translation and rotation adjustments to the global positions of MR and CT images can initially align large organs, but cannot handle soft tissue deformation. Then, the B-spline deformation field is optimized under the guidance of the FFT phase spectrum. The FFT phase spectrum contains structural position information of the image, which constrains the deformation direction, ensuring alignment of key areas such as bone or soft tissue boundaries. Finally, Dice coefficients and wavelet mutual information are used to measure the overlap of organ segmentation masks and the similarity of frequency domain features, respectively, to ensure spatial structural consistency and frequency domain information alignment. In some alternative embodiments, the CT and MR images corresponding to the Dice coefficients need to be greater than 0.98 and the wavelet mutual information needs to be greater than 0.95 to meet the registration requirements.
[0057] S02, perform joint annotation processing of text semantics and frequency domain on the clinical report corresponding to the aligned MR image.
[0058] Specifically, perform regular expression on the clinical report corresponding to the aligned MR image, filter sensitive information and replace it with anonymous identifiers; convert the text description of the processed clinical report into standardized terms, and bind the standardized terms with predefined frequency domain patterns to establish a semantic frequency domain mapping rule library; divide the text of the standardized and frequency domain-feature-associated text into text paragraphs according to organ systems, and insert frequency domain markers in key text areas for subsequent feature extraction. In specific implementation, sensitive information is processed through regular expression to ensure patient privacy, and then non-standard descriptions (such as "space-occupying lesion in the right temporal lobe of the brain") are converted into standard terms (such as "space-occupying lesion in the right temporal lobe") through a medical ontology library (such as UMLS), and then associated with a preset frequency domain pattern (such as "space-occupying lesion" corresponding to "local entropy of wavelet HH component > 2.5") to ensure accurate matching of semantics and frequency domain. Finally, the text paragraphs are split according to organ systems (such as pelvic cavity / abdominal cavity), and frequency domain markers (such as <Wavelet_Region = acetabulum>) are inserted to achieve the binding from text paragraphs to anatomical regions and then to frequency domain features, avoiding confusion of frequency domain features of different organs.
[0059] S03, extract multi-dimensional key information and encode it. After encoding, unify the dimensions and perform dynamic weighted fusion to generate a structured conditional vector, and use the cross-attention mechanism in the skip connection layer of the diffusion model to combine with the frequency domain mapping function to associate the multi-modal features in the structured conditional vector.
[0060] Specifically, process the processed clinical report through an LLM (Large Language Model, LLM, i.e., large language model) to obtain a text semantic embedding vector; perform three-level discrete wavelet transform on the aligned MR image to extract high-frequency detail feature vectors; perform fast Fourier transform on the aligned MR image to extract global low-frequency morphological feature vectors; construct a structured conditional vector based on the text semantic embedding vector, high-frequency detail feature vectors, and global low-frequency morphological feature vectors;
[0061] The formula for the structured conditional vector is:
[0062]
[0063] where is the text semantic embedding vector, is the high-frequency detail feature vector, is the global low-frequency morphological feature vector, , and For the corresponding fusion weights.
[0064] In practical implementation, information extraction can be performed through semantic, spatial, and frequency domain layers. Specifically, semantic layer extraction uses pre-trained medical LLM (such as BioBERT) to identify entities in the text (such as "right ischium", "bone metastasis", "diameter 3cm"), leveraging its advantage of pre-training medical corpora to accurately capture professional terms in radiology reports and avoid semantic biases common in general LLMs. Spatial layer extraction constructs an "anatomical structure → three-dimensional coordinate mapping" based on rule templates (such as "right ischium → [x1,y1,z1]:[x2,y2,z2]"), transforming anatomical locations in the text into regions in the image coordinate system, achieving "semantic → spatial" association. Frequency domain layer extraction explicitly extracts implicit frequency domain parameters from the text (such as "sharp bone edge → wavelet high-frequency band energy threshold ≥0.8"); simultaneously, a fine-tuned LLM is used to predict typical spectral features of the lesion area (such as "FFT radial distribution of bone metastasis") to supplement the deficiencies of explicit parameters and cover complex pathological scenarios. The semantic embedding vector of the text is directly determined by semantic layer extraction, and the frequency domain parameter constraints extracted by frequency domain layer are used as the screening criteria to assist in the extraction of the third-level discrete wavelet transform and fast Fourier transform of MR images.
[0065] Furthermore, to further enhance semantic and frequency domain alignment, a frequency domain mapping function is introduced. and :
[0066]
[0067] in, The hidden state vector output by LLM; It is a wavelet feature mapper that maps text semantics to the wavelet frequency domain space; This is an FFT feature mapper that maps text semantics to the spectral space. A frequency domain mapping function is used to constrain and adjust high-frequency detail feature vectors and low-frequency morphological feature vectors to enhance semantic and frequency domain alignment.
[0068] S04, train the preset diffusion model based on the aligned image, and input the MR image to be converted into the trained diffusion model to output a pseudo-CT image.
[0069] Specifically, a contour-guided adversarial diffusion model is proposed as the main framework for generating synthetic CT images from MRI. It combines the advantages of diffusion models and Generative Adversarial Networks (GANs), improving image generation efficiency and quality through boundary contour information guidance and adversarial diffusion processes. Furthermore, dynamic convolutional modules are inserted into skip connection layers to enhance multi-scale feature fusion capabilities, and spectral normalization is used to stabilize adversarial training. The specific mathematical formulas are as follows, effectively counteracting forward noise addition and the inverse process of diffusion:
[0070]
[0071]
[0072] in, For the first t The image after adding noise. For the first t 1 The image of the step, where I is the identity matrix. For noise scheduling parameters, α t To control the proportion of the original signal retained, c Let be the condition vector, and be The normal distribution (Gaussian distribution) controls the size of the jump step during the reverse process. For parameterized denoising distribution, For a given previous step image Current step image under the condition The distribution For the image at the current step Given the conditional vector c, predict the preceding... k Step Image The distribution k Adjust parameters for step size. For text vectors, For boundary mask information, For wavelet transform, This is a Fourier transform.
[0073] In our main framework, the value of t is significantly greater than 1. Given that t is much larger than 1, the large-step back process lacks a closed-form expression. To address this, we introduce a primitive domain conventional mapper to obtain the complex transition probabilities in the large-step conditional diffusion model. We borrow the idea of an adversarial projector with a cyclic consistent structure and incorporate it into the diffusion process. In our generator... With the input of the source mode y And the newly introduced corresponding boundary mask as input, Intermediate features are extracted for progressive denoising. Furthermore, a learnable embedding mask calculated over time t is incorporated into these extracted features. Then, a discriminant factor is introduced to distinguish features derived from... The generated estimated denoised distribution and the actual denoised samples are used. Temporal embedding is also introduced as a bias term in the feature mapping.
[0074] To generate the target synthetic CT image, the original MRI image with the same anatomical structure can be used as prior information for the inverse diffusion step. The source image corresponding to each sCT image in the training dataset is estimated through both non-diffusion and diffusion processes. Previous diffusion-based image transformation methods typically rely on global feature matching to achieve comprehensive image transformation. However, this method often suffers from inefficient feature learning due to the inclusion of numerous irrelevant background regions, thus affecting the synthesis results. To address this issue, this paper proposes explicitly integrating body contour guidance into the feature learning framework. Specifically, integrating boundary guidance into the conditional generation process of both diffusion and non-diffusion networks effectively reduces the influence of background noise, ensuring that model learning focuses on the human body region. Through this efficient and selective feature learning paradigm, the model can capture key features such as internal anatomical details and external body contours, significantly improving the feature learning effect of internal anatomical synthesis and thus obtaining more accurate and reliable results. To achieve unsupervised learning, a cycle consistency loss is used to compare the real target data with its generated data. During the diffusion process, the reconstructed image is called the synthetic target image; in the non-diffusion module, the estimated source image is mapped to the target domain through the generator. The diffusion and non-diffusion modules are trained jointly without a pre-training process.
[0075] The model diffusion module is used to estimate and synthesize the target domain image from the raw domain data output by the non-diffusion module. To achieve this, two traditional adversarial diffusion methods are employed, each with its own discriminator. In each step of the inverse process, the generator first performs a deterministic estimate of the denoised target image, and then uses the denoising distribution specific to each image mode to synthesize the target image.
[0076] Furthermore, to achieve fine-grained control of the generation process by multimodal conditions, this framework designs a cross-modal attention injection mechanism that dynamically fuses text embeddings, frequency domain features, and image features. The text embeddings generated by LLM are used as conditions and injected into a contour-guided adversarial diffusion model through a cross-attention layer. The diffusion model is guided by text input: "Generate a pseudo-CT scan with clear bone structures and softtissue contrast," with the diffusion process optimized via CLIP embedding or Latent Guidance. The text description is converted into embedding vectors using a CLIP text encoder or LLM (such as BioBERT). Cross-attention modules are inserted into each downsampling and upsampling layer of the U-Net to align the text embeddings with image features.
[0077]
[0078] in, Q The query vector is derived from image features. K and V A key-value vector derived from a condition vector. c , is the dimension of the key vector.
[0079] Furthermore, by constraining the visual features of the generated pseudo-CT images with the semantics of the text through the CLIP image encoder, the consistency of the generated images and text conditions in the semantic space is enhanced through the pre-trained CLIP model.
[0080]
[0081] in, For CLIP image encoder; For CLIP text encoder, For MR images, The text semantic vector is used. Wavelet features and FFT features are fused through feature concatenation and 1×1 convolution, and then injected into the skip connection layer to enhance the preservation of high-frequency details and low-frequency contours.
[0082] Furthermore, following this step, the process includes outputting key features and detecting artifacts from the generated pseudo-CT images through LLM analysis; based on the LLM, correction suggestions are generated according to the artifacts and key features to guide the optimization of the diffusion model. By automatically detecting artifacts and guiding secondary optimization, the quality of pseudo-CT images is improved, and manual intervention is reduced.
[0083] Furthermore, based on the MAE, PSNR, and SSIM metrics of the entire test cohort, a comprehensive quantitative comparison was performed between the pseudo-CT images generated by this method and real CT images. All metrics were calculated in three-dimensional volume space, and the results are as follows:
[0084]
[0085] The results show that although most deep learning methods can effectively suppress noise and achieve reasonable synthesis quality, traditional GAN-based models have significant limitations. These models rely on mirror-symmetric adversarial architectures supplemented by auxiliary loss functions to maintain structural consistency, but inevitably sacrifice Henness unit (HU) accuracy and introduce artifacts into the generated images. Among the other comparative methods, the SynDiff model exhibits superior performance compared to CUT and F-LseSim, achieving lower MAE and RMSE values, as well as a higher SSIM exponent. Most importantly, the two diffusion model-based methods (SynDiff and the novel method proposed in this proposal) achieve the best overall quantitative results. Specifically, the model proposed in this proposal achieves the lowest reconstruction error and the highest image fidelity index. This result significantly outperforms existing mainstream solutions. This performance improvement fully demonstrates the effectiveness of the proposed model framework: the proposed multimodal feature-coupled diffusion model can more accurately model the complex mapping relationship of tissue density while maintaining the invariance of key anatomical structures. The significant improvement in voxel-level HU accuracy and structural similarity is directly attributed to the introduction of semantic guidance and frequency-spatial domain joint modeling mechanisms. Furthermore, the quantitative results above are highly consistent with the qualitative visual assessments conducted by radiologists, further confirming the excellent robustness and reliability of this method in clinical pseudo-CT image generation tasks.
[0086] In summary, the pseudo-CT cross-modal conversion method based on the multimodal feature coupling diffusion model in the above embodiments of the present invention improves the efficiency and accuracy of alignment by calculating the frequency domain reference basis of the MR image, enabling rigid registration of the CT and MR images while simultaneously performing non-rigid registration under frequency domain constraints. Furthermore, by performing semantic parsing and joint annotation of the clinical report, the annotated information is encoded and dimensionally unified to obtain feature vectors containing semantic, frequency, and spatial domains. These vectors are then weighted and fused to obtain structured conditional vectors, thereby simultaneously ensuring the optimization of high-frequency details and low-frequency contours. The semantic, frequency, and spatial features are injected into the diffusion model through a hybrid mechanism, achieving dynamic conditional input and improving the adaptive perception capability of frequency domain features, resulting in pseudo-CT images that closely match the patient's anatomical characteristics. Therefore, the present invention solves the problem of the lack of a high-precision and accurate pseudo-CT cross-modal conversion method and system based on the multimodal feature coupling diffusion model in the prior art.
[0087] Example 2
[0088] Please see Figure 2 The diagram shows the structural block diagram of the pseudo-CT cross-modal conversion system based on the multimodal feature coupling diffusion model proposed in the second embodiment of the present invention. This pseudo-CT image cross-modal conversion system 200 based on the multimodal feature coupling diffusion model includes: an image alignment module 21, an annotation module 22, a vector construction module 23, and an image generation module 24, wherein:
[0089] The image alignment module 21 is used to pre-calculate the frequency domain reference benchmark for the acquired MR image, perform rigid registration of the acquired CT image and MR image, and perform non-rigid image registration based on the frequency domain constraints of the frequency domain reference benchmark to align the CT image and MR image.
[0090] Annotation module 22 is used to perform joint text semantic and frequency domain annotation processing on the clinical reports corresponding to the aligned MR images;
[0091] The vector construction module 23 is used to extract and encode multi-dimensional key information, unify the dimensions after encoding, and perform dynamic weighted fusion to generate a structured conditional vector. In the skip connection layer of the diffusion model, the multimodal features in the structured conditional vector are associated through a cross-attention mechanism combined with a frequency domain mapping function.
[0092] The image generation module 24 is used to train a preset diffusion model based on the aligned image, and input the MR image to be converted into the trained diffusion model to output a pseudo-CT image.
[0093] Example 3
[0094] In another aspect, the present invention also proposes an electronic device, please refer to [link to relevant documentation]. Figure 3 The diagram shows an electronic device according to the third embodiment of the present invention, including a memory 20, a processor 10, and a computer program 30 stored in the memory and executable on the processor. When the processor 10 executes the computer program 30, it implements the pseudo-CT cross-modal conversion method of the multimodal feature coupling diffusion model as described above.
[0095] In some embodiments, the processor 10 may be a central processing unit (CPU), controller, microcontroller, microprocessor or other data processing chip, used to run program code stored in memory 20 or process data, such as executing access restriction programs.
[0096] The memory 20 includes at least one type of readable storage medium, such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 20 can be an internal storage unit of an electronic device, such as the hard disk of the electronic device. In other embodiments, the memory 20 can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. Furthermore, the memory 20 can include both internal and external storage units of the electronic device. The memory 20 can be used not only to store application software and various types of data of the electronic device, but also to temporarily store data that has been output or will be output.
[0097] It should be pointed out that, Figure 3 The structure shown does not constitute a limitation on the electronic device. In other embodiments, the electronic device may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0098] This invention also proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the pseudo-CT cross-modal conversion method of the multimodal feature coupling diffusion model as described above.
[0099] Those skilled in the art will understand that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can mean any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0100] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0101] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0102] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0103] The above embodiments merely illustrate several implementation methods of the present invention, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this patent should be determined by the appended claims.
Claims
1. A pseudo-CT cross-modal conversion method based on a multimodal feature coupling diffusion model, characterized in that, The method includes: The wavelet energy spectrum and FFT logarithmic spectrum of the acquired MR images are pre-calculated, and the global anatomical structure of the acquired MR and CT images is initially aligned by rigid registration. Furthermore, the B-spline deformation field is optimized under the guidance of FFT phase spectrum to complete the non-rigid registration of MR and CT images, thereby aligning the CT and MR images. The clinical reports corresponding to the aligned MR images are subjected to joint text semantic and frequency domain annotation processing. Multi-dimensional key information is extracted and encoded. After encoding, the dimensions are unified and dynamically weighted and fused to generate a structured conditional vector. In the skip connection layer of the diffusion model, the multimodal features in the structured conditional vector are associated through a cross-attention mechanism combined with a frequency domain mapping function. The pre-defined diffusion model is trained based on the aligned image, and the MR image to be converted is input into the trained diffusion model to output a pseudo-CT image. The steps for performing joint text semantic and frequency domain annotation on the clinical report corresponding to the aligned MR image include: The clinical reports corresponding to the aligned MR images are regularized, and sensitive information is filtered and replaced with anonymous identifiers; The processed clinical report text descriptions are converted into standardized terms, and the standardized terms are bound to predefined frequency domain patterns to establish a semantic frequency domain mapping rule base. The text, which is standardized and associated with frequency domain features, is divided into paragraphs based on organ systems, and frequency domain markers are inserted in key regions of the text to facilitate subsequent feature extraction. The steps of extracting and encoding multi-dimensional key information, unifying the dimensions after encoding, and dynamically weighting and fusing the data to generate a structured conditional vector include: The processed clinical reports are processed using LLM to obtain text semantic embedding vectors; High-frequency detail feature vectors are extracted by performing a three-level discrete wavelet transform on the aligned MR image; Fast Fourier Transform is performed on the aligned MR image to extract global low-frequency morphological feature vectors; A structured conditional vector is constructed based on the text semantic embedding vector, high-frequency detail feature vector, and global low-frequency morphological feature vector; The formula for the structured condition vector is: in, For text semantic embedding vectors, These are high-frequency detail feature vectors. This is a global low-frequency morphological feature vector. , and For the corresponding fusion weights.
2. The pseudo-CT cross-modal conversion method of the multimodal feature coupling diffusion model according to claim 1, characterized in that, The cross-attention mechanism is represented as follows: in, Q The query vector is derived from image features. K and V A key-value vector derived from a condition vector. c , is the dimension of the key vector.
3. The pseudo-CT cross-modal conversion method of the multimodal feature coupling diffusion model according to claim 2, characterized in that, The mathematical formula corresponding to the preset diffusion model is: in, For the first t The image after adding noise. For the first t 1 The image of the step, where I is the identity matrix. For noise scheduling parameters, To control the proportion of the original signal retained, c For conditional vectors, It follows a Gaussian distribution. For parameterized denoising distribution, For a given previous step image Current step image under the condition The distribution For the image at the current step Given the conditional vector c, predict the image of the previous step. The distribution For text vectors, For boundary mask information, For wavelet transform, This is a Fourier transform.
4. The pseudo-CT cross-modal conversion method of the multimodal feature coupling diffusion model according to claim 3, characterized in that, After the step of inputting the MR image to be converted into the trained diffusion model to output the pseudo-CT image, the following steps are included: Key features are output and artifacts are detected in pseudo-CT images generated by LLM analysis; Based on LLM, correction suggestions are generated according to artifacts and key features to guide the optimization of the diffusion model.
5. A pseudo-CT cross-modal conversion system based on a multimodal feature coupling diffusion model, characterized in that, A pseudo-CT cross-modal conversion method for implementing a multimodal feature coupling diffusion model as described in any one of claims 1 to 4, the system comprising: The image alignment module is used to pre-calculate the frequency domain reference benchmark for the acquired MR image, perform rigid registration of the acquired CT image and MR image, and perform non-rigid image registration by combining the frequency domain constraints based on the frequency domain reference benchmark, so as to align the CT image and MR image. The annotation module is used to perform joint text semantic and frequency domain annotation processing on the clinical reports corresponding to the aligned MR images; The vector construction module is used to extract and encode multi-dimensional key information. After encoding, the dimensions are unified and dynamically weighted and fused to generate a structured conditional vector. In the skip connection layer of the diffusion model, the multimodal features in the structured conditional vector are associated through a cross-attention mechanism combined with a frequency domain mapping function. The image generation module is used to train a preset diffusion model based on the aligned image, and input the MR image to be converted into the trained diffusion model to output a pseudo-CT image.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the pseudo-CT cross-modal conversion method of the multimodal feature coupling diffusion model as described in any one of claims 1 to 4.
7. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the pseudo-CT cross-modal conversion method of the multimodal feature coupling diffusion model as described in any one of claims 1-4.
Citation Information
Patent Citations
Pseudo-CT cross-modal conversion method and system
CN119991751A