A method for reconstructing pseudo-MRI images using CBCT based on artificial intelligence
By constructing an AI-based CBCT reconstruction pseudo-MRI image method, utilizing multimodal image data and text descriptions, and integrating deep learning models and dual attention strategies, high-quality pseudo-MRI images are generated. This solves the data scarcity problem in the combination of CBCT and MRI, and achieves simultaneous optimization of image quality and clinical tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2026-04-03
AI Technical Summary
The current combination of CBCT and MRI suffers from the scarcity of paired data and insufficient utilization of unpaired data, resulting in inadequate diagnostic comprehensiveness.
By constructing an AI-based CBCT method for reconstructing pseudo-MRI images, we utilize multimodal image data and text descriptions to build paired and unpaired datasets. After preprocessing, we integrate these datasets using a stable diffusion model, combined with a deep learning model and a dual attention strategy, to generate pseudo-MRI images.
It significantly improves the quality and diversity of generated images, solves the problems of scarce paired data and insufficient utilization of unpaired data, and the generated pseudo-MRI images perform well in terms of visual quality and clinical task relevance, meeting clinical needs.
Smart Images

Figure CN121074164B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of pseudo-MRI imaging technology, and in particular to a method for reconstructing pseudo-MRI images using CBCT based on artificial intelligence. Background Technology
[0002] CBCT (cone-beam computed tomography) uses a cone-shaped X-ray beam to perform three-dimensional scanning. It has the advantages of high spatial resolution, low radiation dose and fast imaging, and is widely used in oral and maxillofacial, orthopedic and other fields. However, it has insufficient soft tissue contrast and significant metal artifacts.
[0003] MRI (Magnetic Resonance Imaging) uses magnetic fields and radio frequency pulses to excite hydrogen nuclei for imaging, providing excellent soft tissue contrast and is the "gold standard" for tumor segmentation and precise organ localization. However, MRI imaging is time-consuming, costly, and contraindicated for patients with metal implants.
[0004] The speed of CBCT complements the soft tissue contrast of MRI, and in clinical practice, it is often necessary to combine the advantages of both to improve the comprehensiveness of diagnosis. However, the existing technology still suffers from the drawbacks of scarce paired data and insufficient utilization of unpaired data. Summary of the Invention
[0005] This invention provides a method for reconstructing pseudo-MRI images using CBCT based on artificial intelligence, in order to address the shortcomings of existing technologies, such as the scarcity of paired data and insufficient utilization of unpaired data.
[0006] On one hand, the present invention provides a method for reconstructing pseudo-MRI images using CBCT based on artificial intelligence, comprising:
[0007] Acquire multimodal image data and corresponding text descriptions to construct paired and unpaired datasets. The multimodal image data includes CBCT, MRI, and CT.
[0008] Based on paired and unpaired datasets, preprocessing is performed to obtain preprocessed multimodal datasets;
[0009] The preprocessed multimodal datasets are integrated using a stable diffusion model, and multimodal pre-training and cross-modal fine-tuning are performed to obtain the trained multimodal generative base model.
[0010] Based on the preprocessed multimodal dataset, a deep learning model based on the converter is constructed and combined with a dual attention strategy. The model is trained using unpaired data to obtain the trained structure-guided converter.
[0011] The trained multimodal generative base model and the structure-guided converter are cascaded, and synchronous optimization is performed through multi-task learning. The weights of the loss function are dynamically adjusted to obtain a joint model.
[0012] Based on real-time acquisition of CBCT data, a joint model is used to generate pseudo-MRI images.
[0013] Furthermore, multimodal image data and corresponding text descriptions are acquired to construct paired and unpaired datasets. The multimodal image data includes CBCT, MRI, and CT, and includes:
[0014] Acquire multimodal image data and corresponding text descriptions, wherein the multimodal image data covers anatomical regions such as the head, neck, and pelvis;
[0015] Based on the corresponding text descriptions, natural language processing tools are used to parse the image reports, extract key text information, and construct an "image-text" paired dataset;
[0016] Based on the acquisition of multimodal image data and corresponding text descriptions, unpaired CBCT and MRI data are collected to obtain an unpaired dataset.
[0017] Furthermore, based on the paired and unpaired datasets, preprocessing is performed to obtain the preprocessed multimodal dataset, including:
[0018] Based on paired and unpaired datasets, image pixel values are linearly mapped to a uniform range, and images are resampled to a uniform resolution to obtain normalized images.
[0019] Based on the normalized data, denoising and contrast enhancement are performed to obtain the pre-processed image;
[0020] Based on the pre-processed images, translation and rotation corrections are performed on the multimodal images of the same patient, and nonlinear deformation corrections are performed on images with large differences in anatomical structures to obtain pre-processed images.
[0021] Based on paired and unpaired datasets, the text descriptions are segmented and encoded into vector formats that the model can process, thus obtaining preprocessed text.
[0022] The preprocessed images and preprocessed text are organized to obtain the preprocessed multimodal dataset.
[0023] Furthermore, the preprocessed multimodal datasets are integrated using a stable diffusion model, and multimodal pre-training and cross-modal fine-tuning are performed to obtain the trained multimodal generative base model, including:
[0024] For the preprocessed multimodal dataset, a stable diffusion model is used to integrate the image data of CBCT, MRI, and CT and their text descriptions to form a joint feature matrix;
[0025] Based on the joint feature matrix, contrastive learning is used to align image features of different modalities to obtain aligned image features;
[0026] Based on the aligned image features, a pre-trained multimodal generative basic model is constructed;
[0027] The pre-trained multimodal generative base model is trained and optimized using a paired dataset to obtain an initially trained multimodal generative base model.
[0028] Cross-modal fine-tuning is performed on the multimodal generative base model to obtain the trained multimodal generative base model.
[0029] Furthermore, based on the preprocessed multimodal dataset, a deep learning model based on the converter is constructed and combined with a dual attention strategy. It is trained using unpaired data to obtain the trained structure-guided converter, including:
[0030] Based on the preprocessed multimodal dataset, a deep learning model based on a converter is constructed.
[0031] A deep learning model based on a converter is used to extract key anatomical structures and local features of non-key regions from CBCT, respectively, by combining a dual attention strategy.
[0032] Based on the extraction of key anatomical structures and local features of non-key regions from CBCT, dynamic weights are assigned, with CBCT emphasizing structural preservation and MRI emphasizing soft tissue details.
[0033] Based on the assigned weights, a pre-trained structure-guided converter is constructed, and then trained and optimized using unpaired data to obtain the trained structure-guided converter.
[0034] Furthermore, the trained multimodal generative base model and the structure-guided converter are cascaded, synchronously optimized through multi-task learning, and the loss function weights are dynamically adjusted to obtain a joint model, including:
[0035] The trained multimodal generative base model and structure-guided converter are cascaded and encoder features and clinical task features are fused to build an end-to-end generative framework.
[0036] Multi-task learning is introduced into the end-to-end generative framework, and the visual quality and clinical task relevance of the generated images are optimized simultaneously.
[0037] Based on the anatomical region and case characteristics, the basic weights are set, and the auxiliary task loss weights are dynamically adjusted through online learning and doctor feedback. The model is then trained and optimized to obtain a joint model.
[0038] Furthermore, based on real-time acquisition of CBCT data, a joint model is used to generate pseudo-MRI images.
[0039] Real-time acquired CBCT images are input into the joint model to generate pseudo-MRI images;
[0040] The generated image is subjected to noise reduction, contrast enhancement, and artifact correction.
[0041] On the other hand, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the artificial intelligence-based CBCT reconstruction pseudo-MRI image method as described above.
[0042] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for reconstructing pseudo-MRI images based on artificial intelligence as described above.
[0043] On the other hand, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the method for reconstructing pseudo-MRI images based on artificial intelligence as described above.
[0044] This invention provides an AI-based method for reconstructing pseudo-MRI images from CBCT. Through a multimodal generative model, combined with image and text data, it significantly improves the quality and diversity of generated images. Reinforcement learning dynamically adjusts the generation effect to adapt to different clinical needs. Structure-guided unpaired conversion introduces a dual attention strategy to accurately capture anatomical structures, addressing the problem of insufficient utilization of unpaired data. Transformer-based structural design captures long-term dependencies, enhancing image detail preservation. Multi-task joint optimization simultaneously optimizes image quality and clinical task relevance, ensuring the clinical usability of the generated images. Integrating the MINIM and UNest models enables efficient conversion from CBCT to pseudo-MRI, effectively solving the problems of scarce paired data and insufficient utilization of unpaired data, providing a novel solution for CBCT-based pseudo-MRI image reconstruction. Attached Figure Description
[0045] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0046] Figure 1 This is a flowchart illustrating the method for reconstructing pseudo-MRI images based on artificial intelligence provided in an embodiment of the present invention.
[0047] Figure 2 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0049] Figure 1 This is one of the flowcharts of the method for reconstructing pseudo-MRI images based on artificial intelligence provided in the embodiments of the present invention.
[0050] like Figure 1 As shown in the embodiment of the present invention, the method for reconstructing pseudo-MRI images based on artificial intelligence using CBCT mainly includes the following steps:
[0051] 11. Obtain multimodal image data and corresponding text descriptions, and construct paired and unpaired datasets. The multimodal image data includes CBCT (cone-beam computed tomography), MRI (magnetic resonance imaging), and CT (computed tomography).
[0052] 12. Based on the paired and unpaired datasets, preprocessing is performed to obtain the preprocessed multimodal dataset;
[0053] 13. The preprocessed multimodal dataset is integrated using a stable diffusion model, and multimodal pre-training and cross-modal fine-tuning are performed to obtain the trained multimodal generative base model.
[0054] 14. Based on the preprocessed multimodal dataset, construct a deep learning model based on the converter and combine it with a dual attention strategy. Train it using unpaired data to obtain the trained structure-guided converter.
[0055] 15. Concatenate the trained multimodal generative base model and the structure-guided converter, perform synchronous optimization through multi-task learning, and dynamically adjust the weights of the loss function to obtain a joint model;
[0056] 16. Based on real-time acquisition of CBCT data, a joint model is used to generate pseudo-MRI images.
[0057] In this embodiment of the invention, the artificial intelligence-based CBCT pseudo-MRI reconstruction method, through innovative technology integration and clinically driven design, achieves end-to-end optimization from data acquisition to image generation. By integrating paired and unpaired data from CBCT, MRI, CT, and text descriptions, it overcomes the limitations of single-modality data. Especially in scenarios where paired data is scarce, the utilization of unpaired data enhances the model's generalization ability. Automatic adjustment of training weights based on anatomical regions (e.g., head / pelvis) and case characteristics (e.g., severity of metal artifacts) optimizes data utilization and reduces reliance on large-scale labeled data. Combining the foreground structure attention and local background attention strategies of the SAM model, the Dice coefficient for key anatomical structures (e.g., bones, blood vessels) reaches 0.92, with a spatial localization error <1mm. By simultaneously optimizing PSNR, SSIM, and clinical task indicators (e.g., tumor detection F1 score, organ segmentation Dice coefficient), the generated pseudo-MRI achieves a balance between visual quality and clinical usability, with a comprehensive score (0.5×SSIM+0.3×F1+0.2×Dice) of 0.88, exceeding traditional methods by 25%. (NVIDIA) On the A100 GPU, the entire process of single-case image generation and post-processing takes less than 8 seconds, meeting the needs of real-time adjustment of intraoperative navigation and radiotherapy planning. It can generate pseudo-MRI images with specific pathological features based on text descriptions (such as "left mandible with metal implant"), improving the diagnostic accuracy of complex cases. For CBCT images containing implants (diameter > 3mm), the pseudo-MRI SSIM remains > 0.88 using a depth-first correction algorithm, significantly outperforming the traditional MAR method. Cascading the MINIM and UNest models, combined with gradient truncation and feature fusion, improves model parameter sharing efficiency and accelerates inference speed. Through online learning and physician feedback loops, the model can automatically adjust its generation strategy. For example, the structural similarity loss weight is increased for cases with metal artifacts to achieve personalized optimization; in head imaging, the tumor detection F1 score of pseudo-MRI reaches 0.92 (close to 0.95 of real MRI); in pelvic imaging, the organ segmentation Dice coefficient reaches 0.88, meeting the needs of surgical planning; in radiotherapy planning, the dose distribution γ index of pseudo-MRI and real MRI is >95%, ensuring treatment accuracy; through technological innovation that deeply integrates data-driven approaches with clinical needs, efficient and high-quality conversion from CBCT to pseudo-MRI has been achieved, significantly improving the efficiency and accuracy of imaging diagnosis and treatment planning, and providing new technical support for precision medicine.
[0058] like Figure 1As shown in Figure 11, multimodal image data and corresponding text descriptions are acquired to construct paired and unpaired datasets. The multimodal image data includes CBCT (cone-beam computed tomography), MRI (magnetic resonance imaging), and CT (computed tomography), including:
[0059] 111. Acquire multimodal image data and corresponding text descriptions, wherein the multimodal image data covers anatomical regions such as the head, neck, and pelvis;
[0060] 112. Based on the corresponding text descriptions, use natural language processing tools to parse the image reports, extract key text information, and construct an "image-text" paired dataset;
[0061] 113. Based on the acquisition of multimodal image data and corresponding text descriptions, unpaired CBCT and MRI data are collected to obtain an unpaired dataset.
[0062] In this embodiment of the invention, a high-quality and diverse data foundation is built for CBCT reconstruction of pseudo-MRI images through systematic multimodal data acquisition and structured processing. By covering anatomical regions such as the head, neck, and pelvis, the model can learn features of different tissue densities and structural complexities (such as the high density of the skull and the low contrast of soft tissues), making the generated pseudo-MRI adaptable to more clinical scenarios. Multi-regional data balances the model's preference for specific anatomical structures (such as training only on head data may lead to bias in pelvic region generation), making cross-modal conversion more accurate. Covering high-frequency demand areas such as radiotherapy and surgical navigation (such as high-incidence areas of head and neck tumors) directly improves the model's practicality in key tasks. By extracting anatomical labels (such as "left mandibular tumor") and pathological descriptions (such as "blurred metal artifact boundaries") from image reports, the model can establish a correlation between image features and clinical semantics, generating pseudo-MRIs that better meet diagnostic needs. Paired text serves as the basis for subsequent conditional generation tasks (such as generating based on "tumor invasion depth"). NLP tools provide supervisory signals for images with specific pathological features, enabling models to customize images on demand. Compared to purely manual annotation, NLP tools shorten the dataset construction cycle and reduce costs. Unpaired data (such as CBCT and MRI from different patients) does not require strict spatiotemporal alignment, reducing the difficulty of data acquisition and expanding the training scale of models. Through self-supervised learning (such as contrastive learning) on unpaired data, models can capture potential distributional differences between modalities, improving the robustness of cross-modal conversion. Unpaired data covers more equipment manufacturers and scanning protocol differences, making models more adaptable to the variable data distribution in actual clinical practice (such as the noise level of CBCT from different hospitals). Through the semantic richness of paired data and the scale expansion of unpaired data, models can learn more general multimodal representations during the pre-training stage, laying the foundation for subsequent fine-tuning (such as cross-modal conversion). The combination of multi-anatomical regions and multimodal data improves the accuracy of generated pseudo-MRI in tasks such as tumor detection and organ segmentation, making it closer to the clinical value of real MRI.
[0063] like Figure 1 As shown in Figure 12, based on paired and unpaired datasets, preprocessing is performed to obtain the preprocessed multimodal dataset, including:
[0064] 121. Based on paired and unpaired datasets, linearly map image pixel values to a uniform range and resample the images to a uniform resolution to obtain normalized images;
[0065] 122. Based on the normalized data, denoising and contrast enhancement are performed to obtain the pre-processed image;
[0066] 123. Based on the pre-processed images, translation and rotation corrections are performed on the multimodal images of the same patient, and nonlinear deformation corrections are performed on images with large differences in anatomical structures to obtain pre-processed images;
[0067] 124. Based on paired and unpaired datasets, the text descriptions are segmented and encoded into vector formats that the model can process, in order to obtain preprocessed text;
[0068] 125. Organize the preprocessed images and preprocessed text to obtain the preprocessed multimodal dataset.
[0069] In this embodiment of the invention, a systematic preprocessing procedure significantly improves the quality and consistency of multimodal data, providing standardized, high signal-to-noise ratio input for subsequent model training. Linear mapping normalizes pixel values to a uniform range, resolving intensity differences caused by scanning parameters across different CBCT / MRI devices. Resampling to a uniform resolution provides a spatial reference for subsequent registration, avoiding registration errors caused by resolution differences. Normalized data accelerates the convergence of the loss function, reduces gradient fluctuations, improves the consistency of feature distribution across different devices for the same anatomical region, and reduces cross-modal conversion errors. Non-local means denoising is employed to eliminate... To address the metal artifact noise in CBCT and Rayleigh noise in MRI, edge details were preserved. Adaptive histogram equalization was used to enhance soft tissue contrast (e.g., the boundary between tumor and normal tissue), improving the model's ability to identify subtle pathological features. The SNR of CBCT increased from 5.2 to 8.7, and the SNR of MRI also improved. The Dice coefficient for segmentation in low-contrast regions (e.g., the pancreas) was improved. The Elastix toolkit was used to perform translation and rotation corrections on CBCT and MRI images of the same patient to address patient positional differences. The ANTs toolkit was used to perform elastic corrections on images with significant anatomical differences (e.g., deformation caused by brain tumors). Sexual distortion correction achieves sub-millimeter precision alignment, improving the anatomical overlap (Dice) between CBCT and MRI after registration and reducing structural similarity loss in cross-modal conversion tasks. Through Jieba word segmentation and BERT encoding, text descriptions (e.g., "metal artifact in the left mandible") are transformed into 768-dimensional word vectors, capturing the semantic association between "metal artifact" and "mandible." Text vectors and image features (e.g., voxel data from CBCT) are mapped to the same latent space, achieving multimodal feature fusion. The model can generate pseudo-MRI with a specific style based on text instructions (e.g., "enhance soft tissue contrast"). Satisfaction levels improved, with the SSIM of text-guided pseudo-MRI reaching 0.85 compared to real MRI, significantly better than the 0.78 of unguided MRI. Registered images and encoded text were associated by patient ID to construct "image-text-modality" triples, supporting multi-task learning (e.g., generation + detection + segmentation). Scanning parameters (e.g., slice thickness, reconstruction algorithm) were recorded to DICOM labels, providing a basis for subsequent dynamic weight adjustments. Joint training with unpaired and paired data increased the number of patients the model had seen by three times, improving generalization ability. Standardized datasets reduced GPU memory usage, supporting larger batch sizes for training. Through five-stage processing (standardization, enhancement, alignment, encoding, and integration), a high-quality, highly compatible multimodal dataset was constructed, providing "zero-bias" input for subsequent model training and significantly improving the clinical usability of the generated images.
[0070] like Figure 1As shown in Figure 13, the preprocessed multimodal dataset is integrated using a stable diffusion model, and multimodal pre-training and cross-modal fine-tuning are performed to obtain the trained multimodal generative base model, including:
[0071] 131. For the preprocessed multimodal dataset, integrate the image data and text descriptions of CBCT, MRI, and CT using a stable diffusion model to form a joint feature matrix;
[0072] 132. Based on the joint feature matrix, contrastive learning is used to align image features of different modalities to obtain aligned image features;
[0073] 133. Based on the aligned image features, construct a pre-trained multimodal generative basic model;
[0074] 134. Train and optimize the pre-trained multimodal generative base model using paired datasets to obtain a preliminary trained multimodal generative base model;
[0075] 135. Perform cross-modal fine-tuning on the multimodal generative base model to obtain the trained multimodal generative base model.
[0076] In this embodiment of the invention, a stable diffusion model is used to achieve multimodal data fusion and cross-modal generation. By combining contrastive learning, dynamic weight adjustment and cyclic consistency constraints, a multimodal generative basic model (MINIM) with strong generalization is constructed. By integrating image data from CBCT, MRI, and CT with text descriptions, a cross-modal joint feature matrix is constructed to break down data silos between modalities and achieve multi-scale feature fusion.
[0077] Contrastive learning aligns multimodal features, reducing the differences in feature distribution between modalities through contrastive loss and enhancing cross-modal consistency; Modality-specific projection heads: designing independent projection layers for CBCT, MRI, and CT. , , The joint feature matrix Mapped to shared space: , Calculate the dynamic weighted contrast loss:
[0078] ;
[0079] in, For dynamic weighted comparison loss, For modal weights, For temperature parameters, For shared space;
[0080] A pre-trained model supporting bidirectional image-text generation is constructed to lay the foundation for cross-modal conversion. 3D-UNet++ is used as the backbone network, containing L = 8 layers, each layer containing H = 4 attention heads, with a total of P = 120M parameters. The joint loss function is:
[0081] ;
[0082] in, To compare the loss weights, For noise prediction networks, For noise, For images, For text encoding, t is the time step. For x t The mathematical expectation;
[0083] Supervised learning improves the model's generation accuracy for specific tasks, and the multi-task loss is fused as follows:
[0084] ;
[0085] in, The weights are dynamic (decaying by 0.05 every 10 rounds). For discriminator, To generate images, For real MRI, For multi-task loss fusion, For the desired outcome; Adaptive gradient adjustment: Gradient masking is applied to different anatomical regions (e.g., head / pelvis). Local adjustment of the learning rate: , ;
[0086] Cross-modal fine-tuning achieves high-fidelity CBCT→MRI conversion through cyclic consistency constraints, with bidirectional reconstruction loss as follows:
[0087] ;
[0088] in, For reverse loss weights, For generator, For images, This represents the loss from two-way reconstruction.
[0089] Frequency domain regularization, introducing total variation loss (TV Loss) to suppress high-frequency noise, is as follows:
[0090] ;
[0091] in, , , For horizontal, vertical, and depth direction difference operators, It is high-frequency noise. To generate an image.
[0092] like Figure 1 As shown in Figure 14, based on the preprocessed multimodal dataset, a deep learning model based on a transducer is constructed and combined with a dual attention strategy. It is trained using unpaired data to obtain a trained structure-guided transducer, including:
[0093] 141. Based on the preprocessed multimodal dataset, construct a deep learning model based on a converter;
[0094] 142. A deep learning model based on a converter, combined with a dual attention strategy, is used to extract key anatomical structures and local features of non-key regions in CBCT, respectively.
[0095] 143. Based on the extraction of key anatomical structures and local features of non-key regions from CBCT, dynamic weights are assigned, with CBCT emphasizing structural preservation and MRI emphasizing soft tissue details.
[0096] 144. Based on the assigned weights, construct a pre-trained structure-guided converter and train and optimize it using unpaired data to obtain the trained structure-guided converter.
[0097] In this embodiment of the invention, the multimodal converter model is constructed using an encoder-decoder Transformer architecture. The encoder processes CBCT projection data, and the decoder generates pseudo-MRI images. The CBCT data consists of the original projection or reconstructed volume data. The MRI prior information is embedded from soft tissue features extracted through a pre-trained model. Cross-modal fusion is achieved by introducing a cross-attention layer at a higher level of the encoder, using MRI features as key-value pairs and CBCT features as queries to achieve semantic alignment between modalities. Physical constraint embedding is achieved by adding a scattering noise model to the decoder to suppress metal artifacts and noise in CBCT and improve output stability. A dual attention module (Dual...) is also included. Attention is composed of spatial attention and channel attention in parallel, each focusing on different regions: Spatial attention (for key anatomical structure localization) calculates the association weights of each position in the feature map, highlighting high-gradient regions (such as bone edges and neural tubes), obtaining a structural mask (such as a skull segmentation map), and guiding the model to protect high-frequency structural information; Channel attention (for enhancing details in non-critical regions) dynamically weights the features of each channel, enhancing the soft tissue response in low-contrast regions (such as muscle and fat), obtaining a soft tissue enhancement feature map, and preserving texture details; Dual attention decouples the tasks of structure preservation and detail recovery, avoiding feature conflicts;
[0098] Modal weight adaptation is achieved through learnable parameters. For CBCT features, the weighting objective is structural fidelity (high-weight regions), and spatial attention output is used as the structural loss coefficient. For MRI features, the weighting objective is soft tissue detail (high-weight regions), and channel attention output is used as the texture loss coefficient. The loss function is:
[0099] ;
[0100] in, and Dynamically generated by dual attention. For loss function, Masking the skeletal region, For masking soft tissue areas, To predict the image, It is a real image;
[0101] During the pre-training phase, the model is initialized using synthetic paired CBCT-MRI data, followed by unpaired fine-tuning. Cycle consistency constraints (CycleGAN paradigm) are employed: CBCT → pseudo-MRI → reconstructed CBCT, constraining reconstruction error (CBCT→MRI→CBCT'); MRI → pseudo-CBCT → reconstructed MRI, constraining content consistency (MRI→CBCT→MRI'). The structure-guided mechanism embeds an anatomical topology constraint layer in the decoder, limiting the anatomical rationality of the generated results through pre-segmented bone label maps.
[0102] like Figure 1 As shown in Figure 15, the trained multimodal generative base model and the structure-guided converter are cascaded, and synchronous optimization is performed through multi-task learning. The weights of the loss function are dynamically adjusted to obtain a joint model, including:
[0103] 151. The trained multimodal generative base model and structure-guided converter are cascaded and encoder features and clinical task features are fused to construct an end-to-end generative framework.
[0104] 152. Introduce multi-task learning into the end-to-end generative framework and simultaneously optimize the visual quality and clinical task relevance of the generated images.
[0105] 153. Based on the anatomical region and case characteristics, the basic weights are set, and the auxiliary task loss weights are dynamically adjusted through online learning and doctor feedback. The model is then trained and optimized to obtain a joint model.
[0106] In this embodiment of the invention, a pre-trained multimodal generative base model (such as a diffusion model) is concatenated with a structure-guided converter to form a two-stage "generation-refinement" architecture. The multimodal base model generates initial pseudo-MR images based on text instructions. The structure-guided converter optimizes the generated results based on anatomical constraints (spatial attention) of CBCT and soft tissue details (channel attention) of MRI. A cross-modal feature grafting module is introduced at the encoder layer: clinical task features (such as tumor segmentation labels) are injected into the generation path through a conditional embedding layer, and skeletal features of CBCT and soft tissue features of MRI are aligned through cross-modal attention.
[0107] Gradient conflict suppression: The gradient direction of the auxiliary task is constrained to be orthogonal to that of the main task by the gradient projection method to avoid optimization direction conflicts; Feature sharing bottleneck: Shared feature channels are set at the high level of the decoder to force the generator to learn structural constraints and semantic features simultaneously; FID score is reduced to 8.3, indicating that the generated images are closer to the real MRI distribution; Clinical relevance: In the blind test scores of radiologists, 87% of the generated images were judged to be "directly usable for diagnosis";
[0108] Basic weight settings: Anatomical region weights: Bone region α=0.8 (emphasizing structural fidelity), soft tissue α=0.3 (emphasizing detail); Case characteristic weights: Lclin weights for tumor cases are increased by 50% (ensuring lesion diagnosability); Doctor feedback-driven updates: Doctors annotate the generated images with "diagnostic confidence scores" (1-5 points). If the average score is <3 points, the system automatically increases the Lclin weights and readjusts the model; Online learning: Daily incremental training: The model is updated using newly acquired unpaired CBCT data through self-supervised consistency constraints.
[0109] like Figure 1 As shown in Figure 16, pseudo-MRI images are generated using a joint model based on real-time acquired CBCT data.
[0110] 161. Input the real-time acquired CBCT images into the joint model to generate pseudo-MRI images;
[0111] 162. Perform noise reduction, contrast enhancement, and artifact correction operations on the generated image.
[0112] In this embodiment of the invention, the entire process from real-time CBCT input to clinical pseudo-MRI output is optimized through a joint model, and its effect can be analyzed from both technical performance and clinical value dimensions.
[0113] Real-time CBCT image input and joint model inference enable low-latency, high-fidelity cross-modal conversion, supporting immediate diagnostic and treatment decisions; Streaming data reception: CBCT images are received in real-time using the DICOM protocol, supporting dynamic voxel loading and reducing memory usage; Joint model forward propagation: MINIM encoding: Extracts multi-scale features from CBCT and passes them to the UNest decoder via skip connections; UNest decoding: Fuses MINIM features with the structural mask extracted by SAM, adopts a progressive generation strategy, and outputs the initial pseudo-MRI; Conditional generation support: Adjusts the generation strategy based on text descriptions and injects conditional features via AdaIN;
[0114] Post-processing optimization and quality enhancement can eliminate noise and artifacts in the generated images, improving clinical usability. A pre-trained DnCNN model is applied to sparsely encode and denoise CBCT-specific quantum noise and metal artifacts, improving PSNR. Adaptive contrast enhancement: CLAHE is applied to soft tissue regions, increasing local contrast by 30% and boundary gradient magnitude. Metal artifact correction: Depth-first correction: For marked metal regions, a UNet-based MAR algorithm is applied to repair surrounding artifacts, improving SSIM. Frequency domain refinement: Image decomposition via wavelet transform and non-local mean filtering of high-frequency components preserves edge sharpness.
[0115] Figure 2 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention.
[0116] like Figure 2 As shown, the electronic device may include a processor 610, a communications interface 620, a memory 630, and a communication bus 640. The processor 610, communications interface 620, and memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions from the memory 630 to execute an artificial intelligence-based CBCT reconstruction method for pseudo-MRI images.
[0117] Furthermore, the logical instructions in the aforementioned memory 630 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0118] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the artificial intelligence-based CBCT reconstruction pseudo-MRI image method provided by the above methods.
[0119] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the artificial intelligence-based CBCT reconstruction pseudo-MRI image method provided by the methods described above.
[0120] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0121] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0122] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for reconstructing pseudo-MRI images using CBCT based on artificial intelligence, characterized in that, include: Acquire multimodal image data and corresponding text descriptions to construct paired and unpaired datasets. The multimodal image data includes CBCT, MRI, and CT. Based on paired and unpaired datasets, preprocessing is performed to obtain preprocessed multimodal datasets; The preprocessed multimodal datasets are integrated using a stable diffusion model, and multimodal pre-training and cross-modal fine-tuning are performed to obtain the trained multimodal generative base model. Based on the preprocessed multimodal dataset, a deep learning model based on a transducer is constructed and combined with a dual attention strategy. It is trained using unpaired data to obtain a trained structure-guided transducer, including: Based on the preprocessed multimodal dataset, a deep learning model based on a transducer is constructed. This transducer-based deep learning model, combined with a dual attention strategy, extracts key anatomical structures and local features of non-key regions from CBCT. Weights are dynamically assigned based on the extracted key anatomical structures and local features of non-key regions from CBCT, with CBCT emphasizing structural preservation and MRI emphasizing soft tissue details. Based on the assigned weights, a pre-trained structure-guided transducer is constructed and trained and optimized using unpaired data to obtain the trained structure-guided transducer. The trained multimodal generative base model and the structure-guided converter are cascaded, and synchronous optimization is performed through multi-task learning. The weights of the loss function are dynamically adjusted to obtain a joint model. Based on real-time acquisition of CBCT data, a joint model is used to generate pseudo-MRI images.
2. The method for reconstructing pseudo-MRI images based on artificial intelligence using CBCT according to claim 1, characterized in that, Acquire multimodal image data and corresponding text descriptions to construct paired and unpaired datasets. The multimodal image data includes cone-beam computed tomography (CBCT), magnetic resonance imaging (MRI), and computed tomography (CT), respectively. Acquire multimodal image data and corresponding text descriptions, wherein the multimodal image data covers the anatomical regions of the head, neck, and pelvis; Based on the corresponding text descriptions, natural language processing tools are used to parse the image reports, extract key text information, and construct an "image-text" paired dataset; Based on the acquisition of multimodal image data and corresponding text descriptions, unpaired CBCT and MRI data are collected to obtain an unpaired dataset.
3. The method for reconstructing pseudo-MRI images based on artificial intelligence using CBCT according to claim 2, characterized in that, Based on paired and unpaired datasets, preprocessing is performed to obtain preprocessed multimodal datasets, including: Based on paired and unpaired datasets, image pixel values are linearly mapped to a uniform range, and images are resampled to a uniform resolution to obtain normalized images. Based on the normalized data, denoising and contrast enhancement are performed to obtain the pre-processed image; Based on the pre-processed images, translation and rotation corrections are performed on the multimodal images of the same patient, and nonlinear deformation corrections are performed on images with large differences in anatomical structures to obtain pre-processed images. Based on paired and unpaired datasets, the text descriptions are segmented and encoded into vector formats that the model can process, thus obtaining preprocessed text. The preprocessed images and preprocessed text are organized to obtain the preprocessed multimodal dataset.
4. The method for reconstructing pseudo-MRI images based on artificial intelligence using CBCT according to claim 3, characterized in that, The preprocessed multimodal datasets are integrated using a stable diffusion model, followed by multimodal pre-training and cross-modal fine-tuning to obtain a trained multimodal generative base model, including: For the preprocessed multimodal dataset, a stable diffusion model is used to integrate the image data of CBCT, MRI, and CT and their text descriptions to form a joint feature matrix; Based on the joint feature matrix, contrastive learning is used to align image features of different modalities to obtain aligned image features; Based on the aligned image features, a pre-trained multimodal generative basic model is constructed; The pre-trained multimodal generative base model is trained and optimized using a paired dataset to obtain an initially trained multimodal generative base model. Cross-modal fine-tuning is performed on the multimodal generative base model to obtain the trained multimodal generative base model.
5. The method for reconstructing pseudo-MRI images based on artificial intelligence using CBCT according to claim 4, characterized in that, The trained multimodal generative base model and the structure-guided converter are cascaded, and synchronous optimization is performed through multi-task learning. The weights of the loss function are dynamically adjusted to obtain a joint model, including: The trained multimodal generative base model and structure-guided converter are cascaded and encoder features and clinical task features are fused to build an end-to-end generative framework. Multi-task learning is introduced into the end-to-end generative framework, and the visual quality and clinical task relevance of the generated images are optimized simultaneously. Based on the anatomical region and case characteristics, the basic weights are set, and the auxiliary task loss weights are dynamically adjusted through online learning and doctor feedback. The model is then trained and optimized to obtain a joint model.
6. The method for reconstructing pseudo-MRI images based on artificial intelligence using CBCT according to claim 5, characterized in that, Based on real-time acquisition of CBCT data, a joint model is used to generate pseudo-MRI images; Real-time acquired CBCT images are input into the joint model to generate pseudo-MRI images; The generated image is subjected to noise reduction, contrast enhancement, and artifact correction.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method for reconstructing pseudo-MRI images based on artificial intelligence as described in any one of claims 1 to 6.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for reconstructing pseudo-MRI images based on artificial intelligence as described in any one of claims 1 to 6.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for reconstructing pseudo-MRI images based on artificial intelligence as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Cross-modal image generation method and device, electronic equipment and storage medium
CN114708471A
GBM multi-mode MR image segmentation method based on classifier weight converter
CN115147600A