A method for synthesizing coronary artery CCTA images and an electronic device

CN122510089APending Publication Date: 2026-08-04RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
Filing Date
2026-03-18
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

该现有专利申请存在难以满足临床对冠脉斑块识别、狭窄程度评估的高精度要求的问题

Benefits of technology

1)本发明通过设计预处理方法,构建适配冠状动脉影像的潜在扩散模型 LDM,采用图像与文本双重条件引导去噪 U-Net 进行全参数微调训练,实现从无对比剂 NCCT 图像到高分辨率 SynCCTA 图像的精准合成;其中预处理阶段的质量筛选、心脏区域分割与非刚性配准提升了训练数据质量与一致性,多尺度交叉注意力网络增强了特征融合与细节还原能力,反向扩散去噪过程保证了生成影像的结构保真度,最终输出的 SynCCTA 图像可满足冠状动脉疾病辅助诊断的临床需求,同时避免了对比剂相关风险,提升了检查的安全性与效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122510089A_ABST
    Figure CN122510089A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of synthesis method and electronic equipment of coronary artery CCTA image, comprising: obtaining the image of paired NCCT and CCTA, after pretreatment, obtain high consistency image training pair;LDM model suitable for adapting coronary artery image synthesis is constructed, the LDM model includes frozen variational autoencoder, the denoising U-Net of trainable based on cross attention mechanism, and decoder, the denoising U-Net integrates ResNet2D module, Transformer2D module and multi-scale cross attention layer;After pre-processing, the NCCT image to be processed is input into the LDM model trained, and the latent representation of CCTA image is generated by the reverse diffusion denoising process in latent space, and finally the high-resolution Syn-CCTA image is output by decoder.Compared with prior art, the present application has the advantages of no contrast agent to realize coronary artery CCTA image with auxiliary diagnostic value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of medical image processing and artificial intelligence technology, and in particular to a method and electronic device for synthesizing contrast-free coronary CCTA images. Background Technology

[0002] Coronary artery disease (CAD) is one of the leading causes of morbidity and mortality worldwide. Its pathological basis is the accumulation of atherosclerotic plaques within the coronary arteries, leading to narrowing or blockage of the lumen and ultimately causing myocardial ischemia, hypoxia, or even myocardial infarction. Although invasive coronary angiography remains an important auxiliary reference for diagnosing CAD, its invasiveness, high cost, and limitations in rapid screening restrict its widespread application.

[0003] In contrast, coronary computed tomography angiography (CCTA) has become a preferred non-invasive diagnostic tool. CCTA provides high-resolution three-dimensional images that accurately depict the anatomy of the coronary arteries and related lesions (such as stenosis and non-calcified plaques), and can be used to assess fractional flow reserve. However, CCTA relies on intravenous injection of iodine-containing contrast agents, which poses a risk of adverse reactions, such as contrast-induced nephropathy or allergic reactions, in certain populations, such as the elderly, diabetic patients, and individuals with pre-existing kidney disease. Therefore, the development of innovative technologies that can reduce or even eliminate contrast agent dependence has become an important frontier in the field of medical imaging.

[0004] In recent years, the rapid development of AI, especially deep learning generative models such as generative adversarial networks, variational autoencoders, and diffusion models, has brought revolutionary changes to the field of medical imaging. These models have achieved significant success in image transformation, low-dose CT denoising, and accelerated magnetic resonance imaging. Existing research has attempted to utilize these models to generate synthetic enhanced images from non-contrast computed tomography (NCCT) scans, demonstrating good performance in CT and MRI imaging of areas such as the brain and abdomen.

[0005] However, applying such technologies to coronary arteries presents even greater challenges. The coronary artery structure is extremely complex, containing numerous branches and small vessels, and the continuous beating and breathing of the heart pose significant difficulties for image acquisition and registration. These complexities make accurately generating high-fidelity Syn-CCTA (Synthetic Coronary Computed Tomography Angiography) images from NCCT an extremely challenging task, and current technologies have yet to provide a reliable, accurate, and multi-center validated solution.

[0006] Therefore, there is an urgent clinical need for a new technology that can generate CCTA images with diagnostic value without the use of contrast agents, thereby providing a safer and more economical auxiliary reference for the screening and diagnosis of CAD.

[0007] A search revealed Chinese invention patent application publication number CN119515807A, which discloses a multi-stage conversion method from non-iodine contrast agent CT images to coronary artery CTA images. The method includes: acquiring matched NCCT and CTA volumes, performing data preprocessing, and constructing a training dataset; constructing NCCT encoder-decoder networks and CTA encoder-decoder networks respectively, and pre-training them based on the training dataset; constructing and training a marker reconstruction network, using the NCCT tokens obtained by the trained encoder as conditions to reconstruct masked low-resolution CTA markers, and then reconstructing the low-resolution CTA volume using the trained decoder; constructing and training a super-resolution network to generate high-resolution CTA images from low-resolution CTA images; and using the trained multi-stage model to generate high-resolution CTA volumes from low-resolution NCCT volumes. The advantages of this invention are: avoiding the posterior collapse problem, faster decoding speed during inference, and a more stable training process; in the low-resolution stage, the model can learn multi-scale contextual information, while in the high-resolution stage it focuses on restoring details, retaining more fine structure in the final output, which can meet clinical needs. The existing patent application has the problem of failing to meet the high accuracy requirements of clinical practice for coronary plaque identification and stenosis assessment.

[0008] How to generate coronary CCTA images with auxiliary diagnostic value without contrast agent has become a technical problem that needs to be solved. Summary of the Invention

[0009] The purpose of this invention is to overcome the shortcomings of the prior art by providing a method and electronic device for synthesizing coronary CCTA images.

[0010] The objective of this invention can be achieved through the following technical solutions: According to one aspect of the present invention, a method for synthesizing coronary CCTA images is provided, the method comprising: Acquire paired NCCT and CCTA images, and perform quality screening, cardiac region-specific segmentation, and non-rigid registration preprocessing on the images in sequence to obtain highly consistent image training pairs; An LDM model adapted for coronary artery image synthesis is constructed. The LDM model includes a frozen variational autoencoder, a trainable denoising U-Net based on cross-attention mechanism, and a decoder. The denoising U-Net integrates a ResNet2D module, a Transformer2D module, and a multi-scale cross-attention layer. The image training pairs are encoded by a variational autoencoder and then concatenated in the channel dimension. The text embedding vector converted by the text encoder is used as a dual conditional input to the denoising U-Net to perform full parameter fine-tuning training on the LDM model, so that the model learns to gradually denoise and reconstruct the latent representation of the corresponding CCTA image from the latent representation of the NCCT image and the noise. The NCCT image to be processed is preprocessed and then input into the trained LDM model. The latent representation of the CCTA image is generated in the latent space through the back diffusion denoising process. Finally, the high-resolution Syn-CCTA image is output through the decoder.

[0011] As a preferred technical solution, the quality screening includes: Images containing tachycardia, arrhythmia, or motion artifacts were excluded. Images with low signal-to-noise ratios and a standard deviation of Henle units greater than 40 within the aortic lumen were excluded. Images of the main coronary artery and the left and right coronary arteries with Henle unit values ​​below 200 were excluded. And images excluding those containing coronary stents.

[0012] As a preferred technical solution, the heart region-specific segmentation is as follows: two independent nnU-Net models are constructed to adapt to the Henle unit differences of NCCT and CCTA images, respectively. After training with labeled gold standard data of heart contours, the NCCT and CCTA images are automatically segmented, and the heart region is extracted as the region of interest.

[0013] As a preferred technical solution, the non-rigid registration adopts a symmetric normalization algorithm, with the CCTA image as the fixed reference image and the NCCT image as the moving image. For images with inconsistent slice thickness, the thin-slice CCTA image is first resampled to match the slice thickness of the NCCT image. Then, mutual information is used as a similarity measure. Finally, the image resolution is unified through linear interpolation to establish a pixel-level spatial correspondence.

[0014] As a preferred technical solution, the multi-scale cross-attention layer includes a CrossAttnDown2D layer, a CrossAttnMid2D layer, and a CrossAttnUp2D layer, which correspond to the downsampling stage, bottleneck layer, and upsampling stage of the denoising U-Net, respectively, to realize the interactive fusion of text embedding vectors and latent image features at multiple scales.

[0015] As a preferred technical solution, the text embedding vector is obtained by converting text conditions using the CLIP text encoder, where the text conditions are textual information characterizing the quality or features of coronary CT angiography images.

[0016] As a preferred technical solution, the reverse diffusion denoising process is as follows: starting from random Gaussian noise in the latent space, under the dual guidance of the NCCT image latent representation and the text embedding vector, a clear CCTA image latent representation is gradually generated through multi-step iterative denoising.

[0017] As a preferred technical solution, the paired NCCT and CCTA images include internal queue data and external queue data. The internal queue data is used for model training and validation, and the external queue data is used for model generalization ability testing.

[0018] According to a second aspect of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the program to implement the method described thereon.

[0019] According to a third aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method described thereon.

[0020] Compared with the prior art, the present invention has the following beneficial effects: 1) This invention designs a preprocessing method to construct a latent diffusion model (LDM) adapted to coronary artery images. It employs a dual-condition guided denoising U-Net with both image and text for full-parameter fine-tuning training, achieving accurate synthesis from contrast-free NCCT images to high-resolution SynCCTA images. The quality screening, cardiac region segmentation, and non-rigid registration in the preprocessing stage improve the quality and consistency of the training data. The multi-scale cross-attention network enhances feature fusion and detail restoration capabilities. The backdiffusion denoising process ensures the structural fidelity of the generated images. The final SynCCTA images can meet the clinical needs of auxiliary diagnosis of coronary artery diseases, while avoiding contrast-related risks and improving the safety and efficiency of the examination.

[0021] 2) By limiting and excluding motion artifacts, low signal-to-noise ratio images, and high-risk population data, this invention effectively reduces the impact of image noise and motion interference on the synthesis results, improves the quality stability and clinical applicability of the generated images, reduces the risk of misdiagnosis in subsequent diagnoses, and enables the model to maintain stable image generation results in complex clinical scenarios.

[0022] 3) Two-branch nnU-Net networks were constructed for the cardiac region to process NCCT and CCTA images and extract features. This fully explores the unique features and complementary information of the two types of images, enhances key anatomical features while preserving details of the cardiac region, improves the targeting and completeness of feature extraction, provides more accurate and richer feature support for subsequent image fusion, and thus improves the quality of fused images and their clinical application value.

[0023] 4) This invention uses a symmetric normalization algorithm to perform non-rigid registration of NCCT and CCTA images, which can effectively correct the nonlinear deformation caused by cardiac pulsation and respiratory motion between two scans, thereby establishing the most accurate pixel-level correspondence between NCCT and CCTA, making the fused image more closely resemble the real anatomical structure and adapting to the clinical needs of coronary artery disease diagnosis.

[0024] 5) This invention constructs a cross-attention module based on a cross-attention mechanism, which includes CrossAttnDown2D, CrossAttnMid2D, and CrossAttnUp2D. It embeds the downsampling, intermediate bottleneck, and upsampling stages of the -Net network, enabling precise interaction and fusion of NCCT and CCTA image features in multi-scale space. This strengthens the correlation between features, improves the structural consistency and detail integrity of the fused image, enhances the network's ability to capture features in the coronary artery region, and optimizes the image fusion effect. Attached Figure Description

[0025] Figure 1 This is a schematic flowchart of the contrast agent-free coronary CCTA image synthesis method of the present invention; Figure 2 This is a schematic diagram of the data preprocessing process in this invention; Figure 3 This is a schematic diagram of the LDM model in this invention; Figure 4 This is a schematic diagram of the inference process using the LDM model in this invention; Figure 5 This is a schematic diagram of the U-Net denoising process based on the cross-attention mechanism in this invention; Figure 6 This is a visual comparison diagram of the Syn-CCTA image generated by this invention and the real CCTA image in a two-dimensional axial view; Figure 7 This is a comparison image of the Syn-CCTA image generated by this invention and the real CCTA image under a curved surface reconstruction view; Figure 8 This is a comparative illustration of the performance indicators of Syn-CCTA in plaque type and stenosis classification according to the present invention; Figure 9 This is a schematic diagram comparing the Syn-CCTA of the present invention with the real CCTA in a classification task. Detailed Implementation

[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0027] This invention aims to provide a method and electronic device for generating Syn-CCTA images from NCCT images based on deep learning. Utilizing an advanced diffusion model, and through rigorous data preprocessing and model training, this invention achieves high-fidelity image generation, and its clinical diagnostic value has been verified.

[0028] Example 1 This embodiment relates to a contrast agent-free coronary CCTA image synthesis method, such as... Figure 1 ,include: Step 1, Data Acquisition and Preprocessing: Acquire paired NCCT and real CCTA scan image pairs, and perform a rigorous preprocessing procedure on these images, including: 1) Image quality screening to exclude low-quality images caused by factors such as tachycardia, arrhythmia, motion artifacts, or poor signal quality; 2) Heart region segmentation: Using a pre-trained neural network model (such as nnU-Net), the entire heart region is automatically and accurately segmented from NCCT and CCTA images as the Region of Interest (ROI) to focus the model's attention and improve training efficiency. 3) Image registration: A non-rigid registration algorithm (such as the symmetric normalization algorithm) is used to accurately align the NCCT image with the corresponding CCTA image to correct problems such as cardiac phase mismatch and provide pixel-level correspondence for supervised learning.

[0029] Step 2, Deep Learning Model Construction and Training: An advanced Latent Diffusion Model (LDM) is adopted. Specifically, the pre-trained stable latent diffusion model is fine-tuned with all parameters to adapt to the medical image synthesis task from NCCT to CCTA images.

[0030] The model's architecture comprises a variational autoencoder and a U-Net-based denoising network. During training, NCCT images are encoded into a low-dimensional latent space using the variational autoencoder and then concatenated with the similarly encoded CCTA latent representation. Furthermore, the model receives textual cues as additional conditions, incorporating semantic information into the image generation process through a cross-attention mechanism. Guided by this dual image-text condition, the model learns to progressively denoise and reconstruct the corresponding CCTA latent representation from the NCCT latent representation and noise.

[0031] Training strategies based on both image and text conditions, along with a rigorous multi-stage preprocessing process, are key to achieving high-fidelity coronary artery image generation.

[0032] Step 3, Syn-CCTA Image Generation and Inference: An arbitrary new NCCT image is input into the trained latent diffusion model. After preprocessing and variational autoencoder encoding, a synthetic CCTA latent representation is generated in the latent space through the back-diffusion (denoising) process of the latent diffusion model. Finally, it is restored to a high-resolution Syn-CCTA image through the variational autoencoder decoder.

[0033] Step 4, Performance Evaluation and Validation: A comprehensive performance evaluation of the generated Syn-CCTA images is performed, including: 1) Objective image quality evaluation, using indicators such as the Structural Similarity Index Measure (SSIM) and the Descein coefficient to quantify the similarity between Syn-CCTA and real CCTA in terms of structure and spatial overlap; 2) Diagnostic performance validation, analyzing Syn-CCTA images to evaluate their accuracy in classifying plaque types (none, calcified, non-calcified, mixed) and grading the degree of vascular stenosis (none, slight, mild, moderate, severe), and comparing them with diagnostic reports from radiologists based on real CCTA.

[0034] This invention innovatively applies advanced LDM to the highly challenging task of generating Syn-CCTA from NCCT, providing a completely non-invasive alternative to CCTA. It avoids the use of iodine contrast agents, thereby eliminating the associated risks of adverse reactions. It is particularly suitable for patients with renal insufficiency, a history of contrast agent allergy, or those unwilling to undergo invasive examinations. It has extremely high clinical application value and safety. Since there is no need to inject contrast agents and wait for them to distribute in the body, it simplifies the scanning process, improves examination efficiency, reduces patient waiting time, and can reduce medical costs and examination time.

[0035] Example 2 This embodiment also relates to a method for synthesizing coronary CCTA images without contrast agent, including: Data acquisition and preprocessing, the data preprocessing process is as follows: Figure 2 We collected paired NCCT and CCTA scan data from 4,522 patients across four different medical centers (one internal center and three external centers). The internal cohort, containing 3,007 patients, was used for model training, validation, and internal testing. The three external cohorts, totaling 1,515 patients, were used for testing the model's external generalization ability.

[0036] To ensure the quality of model training, strict image selection criteria were established. Excluded images include: (1) Poor image quality due to tachycardia, arrhythmia, irregular breathing or motion artifacts; (2) Low signal-to-noise ratio, i.e., the standard deviation of the Henle unit value in the aortic lumen is greater than 40; (3) The Henle unit values ​​of the lumen of the main coronary artery, left coronary artery, and right coronary artery are less than 200; (4) Coronary artery stents are present; (5) Due to cardiac phase mismatch, the CCTA and NCCT image registration failed.

[0037] Specifically, this invention employs a rigorous manual visual inspection process for assessing registration quality. Registered NCCT and CCTA images are simultaneously loaded into an ITK-SNAP viewer. The examiner manually traces the coronary arteries on the CCTA and verifies whether the corresponding anatomical structures on the NCCT maintain spatial consistency across consecutive slices. This visual assessment undergoes three iterations; any cases exhibiting misalignment or poor correspondence in any round are excluded to ensure that only high-quality aligned image pairs are used for subsequent analysis. After screening, 1005 cases were retained in the internal cohort, and 75, 106, and 284 cases were retained in the external cohorts, respectively.

[0038] Heart Segmentation: To ensure the model focuses on the heart region, this invention employs a pre-trained nnU-Net model to automatically segment the entire heart in NCCT and CCTA images. First, the heart outline is accurately delineated on 200 high-quality images as the gold standard. Then, two independent nnU-Net models are trained using these 200 labeled images, one specifically for NCCT and the other for CCTA, to accommodate the inherent differences in Henlein unit values ​​between the two modalities. Training employs single-fold cross-validation for 750 epochs and utilizes a 3D full-resolution U-Net architecture. After training, the nnU-Net model is applied to all remaining datasets for batch automatic heart segmentation, and the segmented heart regions are used as Regions of Interest (ROIs) for subsequent processing.

[0039] Image registration: Accurate image alignment is crucial for the success of supervised learning. This invention employs an advanced symmetric normalization algorithm for non-rigid registration, implemented using the Advanced Normalization Tools (ANTs) software library. In this process, the CCTA image is used as a fixed reference image, and the NCCT image as a moving corresponding image. This effectively suppresses redundant image information, highlights key features of the coronary artery region, and improves the structural similarity and detail reproduction of the fused image through feature alignment and weight optimization. This reduces image distortion and artifacts, making the fused image more closely resemble the actual anatomical structure and meeting the clinical needs of coronary artery disease diagnosis.

[0040] The registration process includes: First, for queues with inconsistent slice thicknesses (internal queue and external queue 1), thin-slice CCTA (≤1mm) images are resampled to match the slice thickness of thick-slice NCCT (2mm-3mm); this step is unnecessary for queues with consistent slice thicknesses. Next, a symmetric normalization algorithm is applied to transform the model, using mutual information as a similarity metric, with a gradient step size of 0.2 and a smoothing parameter of 3. This is continuously adjusted based on the similarity level to ensure complete alignment of blood vessel and tissue positions. The calculated transformation is applied to the moving NCCT image, generating a distorted image precisely aligned with the fixed CCTA image. After alignment, the resolution of the two images is unified, and linear interpolation is used to bring the CCTA image to the same clarity as the NCCT. This method effectively corrects for nonlinear deformation caused by cardiac pulsation and respiratory movements between scans, thereby establishing the most accurate pixel-level correspondence possible between NCCT and CCTA.

[0041] This invention modifies and fine-tunes the open-source stable diffusion model v1.5, whose architecture is as follows: Figure 3 As shown, it contains three core modules: encoder, denoising U-Net, and decoder.

[0042] Overall Generation Process Overview: This model takes NCCT images and textual conditions as input. First, a pre-trained and frozen encoder compresses the NCCT image into a conditional latent image; simultaneously, a CLIP text encoder converts the textual conditions into text embeddings. These two sets of conditional information are then fed into a forward diffusion process, working in conjunction with a trainable denoising U-Net to iteratively reconstruct the latent representation. Finally, the decoder decodes the optimized latent representation into a high-resolution Syn-CCTA image.

[0043] Denoising U-Net structure: such as Figure 5 The U-Net employs a classic encoder-decoder architecture, comprising downsampling and upsampling blocks, ResNet2D blocks, Transformer2D layers, and multi-scale cross-attention layers (CrossAttn, including CrossAttnDown2D / Mid2D / Up2D). Its input is a noisy latent representation (64×64×8×B), which undergoes multiple levels of downsampling (to 8×8×1280) before entering a bottleneck layer. It then undergoes upsampling to restore the spatial dimensions, outputting a denoised latent representation (64×64×4×B). The cross-attention mechanism allows text embeddings and image latent features to interact at multiple scales, enabling text-guided image generation. The denoising U-Net architecture employs a hierarchical denoising process of "downsampling-bottleneck layer-upsampling," adapting to the extraction and restoration of fine-grained features such as coronary artery micro-branches and plaques, achieving multi-scale feature interaction.

[0044] Inference workflow such as Figure 4 During the inference phase, given the NCCT image and text conditions, the latent diffusion model first extracts the latent representation through the encoder, combines it with the text embedding input into the pre-trained diffusion model, and gradually denoises and generates a refined latent representation through the backdiffusion process. Finally, the decoder restores it to the final Syn-CCTA image.

[0045] This invention constructs highly consistent training data pairs by segmenting cardiac regions from NCCT and CCTA images, performing multimodal non-rigid registration, and rigorous quality screening. Based on this, a frozen encoder and a multi-scale text-guided mechanism are introduced using a stable diffusion v1.5 architecture to achieve high-fidelity, pathologically controllable synthesis from NCCT to Syn-CCTA. Compared to the original model, this significantly enhances anatomical fidelity and clinical semantic editing capabilities.

[0046] The model of this invention employs dual-condition input. The image condition is a preprocessed NCCT image, encoded by a variational autoencoder and concatenated along the channel dimension with a noisy CCTA latent representation, before being input into U-Net. The text condition is a descriptive text, such as "a high-quality coronary CT angiography image," which is converted into an embedding vector by a text encoder. This dual-condition mechanism enables the model to learn specific transformations from NCCT to CCTA, while utilizing textual cues to control the style and quality of the generated images.

[0047] Training Process: This invention fully fine-tunes the entire U-Net portion of the pre-trained stable diffusion model. The training dataset comes from the internal queue in Example 2, with 703 examples used for training and 151 examples for validation. The training objective is to minimize the mean squared error between the noise predicted by the U-Net and the noise actually added during forward diffusion. Training uses the AdamW optimizer with a learning rate of 1e-5 and a batch size of 8. Training is performed on four NVIDIA A100 GPUs for approximately 100 epochs until the loss converges on the validation set.

[0048] Example 3 This embodiment also relates to a contrast agent-free coronary CCTA image synthesis method. The following describes how to generate Syn-CCTA images using a trained LDM model and how to objectively evaluate their quality, including: (3.1) Inference Process: For a new NCCT image, it first undergoes the same preprocessing steps (segmentation and registration) as during training. Then, it is input into the trained LDM model. The model starts with random Gaussian noise in the latent space and, guided by the NCCT latent representation and textual conditions, gradually generates a clear CCTA latent representation through a series of (e.g., 50 steps) back-diffusion (denoising) steps. Finally, the decoder converts this latent representation into the final Syn-CCTA image.

[0049] (3.2) Objective quality assessment: In order to quantify the fidelity of the generated image, the present invention calculates the following indicators for quality assessment: a. Dessian coefficient: Coronary arteries were segmented on Syn-CCTA and real CCTA, and the overlap of the segmentation masks was calculated. The results showed that the Dessian coefficient of coronary arteries exceeded 0.80 in all test cohorts, and even reached over 0.85 in the internal test set, indicating that the generated vascular structure was highly consistent with the real blood vessel in spatial location.

[0050] b. Structural Similarity Index (SSIM): This index measures the similarity of Syn-CCTA to real CCTA in terms of structure, brightness, and contrast. Throughout the cardiac region, the SSIM score ranges from 0.731 to 0.787; in the more challenging coronary artery region, the SSIM score ranges from 0.701 to 0.747. These values ​​indicate that the generated image is structurally highly similar to the real image.

[0051] Figure 6 The images show a comparison of two-dimensional images from several cases, which clearly demonstrate that Syn-CCTA (middle) successfully "restores" the contrast-enhanced blood vessels from NCCT (left), with an effect very close to that of real CCTA (right).

[0052] This invention also verified the practical efficacy of Syn-CCTA in the diagnosis of coronary artery disease by simulating clinical application scenarios.

[0053] Verification method: This invention uses a commercially approved fully automated AI diagnostic platform that can automatically analyze CCTA images and output analysis reports on plaques and stenosis.

[0054] The verification process is as follows: First, the AI ​​platform was applied to real CCTA images, and its analysis results were compared with clinical reports issued by experienced radiologists, verifying the high accuracy of the AI ​​platform itself (e.g., a weighted Kappa coefficient greater than 0.95 compared to the doctor's diagnosis). Then, the AI ​​platform was applied to Syn-CCTA images generated by this invention, and its diagnostic results were compared with the analysis reports of radiologists on real CCTA images of the same case.

[0055] Assessment metrics: The assessment focused on two core clinical tasks: plaque type classification (divided into four categories: none, calcified, non-calcified, and mixed) and stenosis severity grading (divided into five levels: none (0%), minor (1-24%), mild (25-49%), moderate (50-69%), and severe (≥70%)). Statistical indicators used included weighted / unweighted Cohen's Kappa coefficient, weighted F1 score, and multivariate Matthews correlation coefficient.

[0056] The verification results are as follows Figure 8 As shown, it includes: a) Severity assessment: In the two external cohorts, the weighted Kappa coefficients of Syn-CCTA analysis results and true CCTA were 0.801 and 0.828, respectively, indicating that the two had a “substantial” to “near-perfect” consistency, with weighted F1 scores of 0.861 and 0.892, respectively.

[0057] b) Plaque type assessment: Unweighted Kappa coefficients were 0.735 and 0.776, respectively, showing a similar "substantial" consistency, while weighted F1 scores were 0.878 and 0.899, respectively.

[0058] c) Performance analysis: such as Figure 8 and Figure 9 As shown, Syn-CCTA performs well in identifying "plaque-free" and "calcified plaques," with high F1 scores. It also shows relatively high accuracy in identifying "mixed plaques." However, a major challenge lies in identifying "non-calcified plaques," with relatively low precision and recall. This may be because the density difference between non-calcified plaques and surrounding tissues is not significant on NCCT, making it difficult for the model to learn their characteristics. Nevertheless, as... Figure 4 As shown, Syn-CCTA can provide valuable morphological information on curved reconstructed images, even for narrowing caused by non-calcified plaques.

[0059] In summary, the method provided by this invention can reliably generate high-quality Syn-CCTA images from conventional NCCT images, demonstrating high consistency with the gold standard CCTA in key diagnostic indicators of coronary artery disease. Although there are certain limitations in identifying non-calcified plaques, its overall performance indicates that this invention has great potential for clinical application and is expected to become a safe, effective, and economical new tool for CAD screening and diagnosis.

[0060] This invention demonstrates the high fidelity of its generated images and the reliability of its diagnostic results through training and validation on multi-center, large-scale datasets. The results show that the generated Syn-CCTA exhibits high consistency with real CCTA in plaque identification and stenosis assessment, demonstrating its significant potential as a clinical diagnostic tool.

[0061] Example 4 The electronic device of this invention includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) or loaded from a storage unit into random access memory (RAM). The RAM may also store various programs and data required for device operation. The CPU, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0062] Multiple components in the device are connected to the I / O interface, including: input units such as keyboards and mice; output units such as various types of displays and speakers; storage units such as disks and optical discs; and communication units such as network interface cards (NICs), modems, and wireless transceivers. The communication unit allows the device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0063] The processing unit performs the various methods and processes described above. For example, in some embodiments, the methods may be implemented as computer software programs tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on the device via ROM and / or a communication unit. When the computer program is loaded into RAM and executed by the CPU, one or more steps of the methods described above may be performed. Alternatively, in other embodiments, the CPU may be configured to execute the methods by any other suitable means (e.g., by means of firmware).

[0064] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0065] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0066] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0067] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for synthesizing contrast-free coronary CCTA images, characterized in that, The method includes: Acquire paired NCCT and CCTA images, and perform quality screening, cardiac region-specific segmentation, and non-rigid registration preprocessing on the images in sequence to obtain highly consistent image training pairs; An LDM model adapted for coronary artery image synthesis is constructed. The LDM model includes a frozen variational autoencoder, a trainable denoising U-Net based on cross-attention mechanism, and a decoder. The denoising U-Net integrates a ResNet2D module, a Transformer2D module, and a multi-scale cross-attention layer. The image training pairs are encoded by a variational autoencoder and then concatenated in the channel dimension. The text embedding vector converted by the text encoder is used as a dual conditional input to the denoising U-Net to perform full parameter fine-tuning training on the LDM model, so that the model learns to gradually denoise and reconstruct the latent representation of the corresponding CCTA image from the latent representation of the NCCT image and the noise. The NCCT image to be processed is preprocessed and then input into the trained LDM model. The latent representation of the CCTA image is generated in the latent space through the back diffusion denoising process. Finally, the high-resolution Syn-CCTA image is output through the decoder.

2. The method for synthesizing contrast-free coronary CCTA images according to claim 1, characterized in that, The quality screening includes: Images containing tachycardia, arrhythmia, or motion artifacts were excluded. Images with low signal-to-noise ratios and a standard deviation of Henle units greater than 40 within the aortic lumen were excluded. Images of the main coronary artery and the left and right coronary arteries with Henle unit values ​​below 200 were excluded. And images excluding those containing coronary stents.

3. The method for synthesizing contrast-free coronary CCTA images according to claim 1, characterized in that, The heart region-specific segmentation involves constructing two independent nnU-Net models, each adapted to the Henle unit differences in NCCT and CCTA images. After training with labeled gold standard heart contour data, the models automatically segment the NCCT and CCTA images, extracting the heart region as the region of interest.

4. The method for synthesizing contrast-free coronary CCTA images according to claim 1, characterized in that, The non-rigid registration adopts a symmetric normalization algorithm, using the CCTA image as a fixed reference image and the NCCT image as a moving image. For images with inconsistent slice thickness, the thin-slice CCTA image is first resampled to match the slice thickness of the NCCT image. Then, mutual information is used as a similarity measure. Finally, the image resolution is unified through linear interpolation to establish a pixel-level spatial correspondence.

5. The method for synthesizing contrast-free coronary CCTA images according to claim 1, characterized in that, The multi-scale cross-attention layer includes the CrossAttnDown2D layer, the CrossAttnMid2D layer, and the CrossAttnUp2D layer, which correspond to the downsampling stage, bottleneck layer, and upsampling stage of the denoising U-Net, respectively, to achieve interactive fusion of text embedding vectors and latent image features at multiple scales.

6. The method for synthesizing contrast-free coronary CCTA images according to claim 1, characterized in that, The text embedding vector is obtained by converting text conditions using the CLIP text encoder. The text conditions are textual information that characterizes the quality or features of coronary CT angiography images.

7. The method for synthesizing contrast-free coronary CCTA images according to claim 1, characterized in that, The reverse diffusion denoising process is as follows: In the latent space, starting from random Gaussian noise, under the dual guidance of the NCCT image latent representation and the text embedding vector, a clear CCTA image latent representation is gradually generated through multi-step iterative denoising.

8. The method for synthesizing contrast-free coronary CCTA images according to claim 1, characterized in that, The paired NCCT and CCTA images include internal queue data and external queue data. The internal queue data is used for model training and validation, and the external queue data is used for testing the model's generalization ability.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 8.