Multi-tracer PET (Polyethylene Terephthalate) cross-modal synthesis system based on diffusion Transformer
Through a multi-tracer PET cross-modal synthesis system based on diffusion Transformer, high-quality multi-tracer PET images are synthesized from MRI images, solving the problems of poor image quality and insufficient clinical practicality assessment in the prior art, and achieving a fast and accurate Alzheimer's diagnosis.
Patent Information
- Application Number
- CN202510116484.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has problems with poor image quality and lack of clinical practicality assessment in the generation of high-quality multitracer PET images, especially in revealing the interrelationships between Alzheimer's-related A-T-N biomarkers.
Using a multi-tracer PET cross-modal synthesis system based on diffusion Transformer, a multi-tracer PET image is synthesized from structural MRI images by designing a GenPET model, and the MRI is converted into a PET image of a specified tracer type using the hidden space diffusion Transformer model.
The generated PET images have high quality, strong detection ability, fast inference speed, can effectively capture pathophysiological changes, and provide a non-invasive and high-quality alternative method in clinical practice, which improves the diagnosis level of Alzheimer's disease.
Smart Images

Figure CN119991959A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of medical image processing, and in particular relates to a multi-tracer PET cross-modality synthesis system. Background Art
[0002] Alzheimer's disease (AD) is the leading cause of dementia, with approximately 55 million people suffering from the disease worldwide, a number that is expected to triple by 2050. As the global elderly population increases, the annual cost of AD medical care and long-term care is expected to exceed $1 trillion without effective treatments, highlighting the urgent need for early and accurate detection to enable timely management and intervention.
[0003] Neuroimaging techniques such as positron emission tomography (PET) and magnetic resonance imaging (MRI) can directly and objectively observe brain changes and are therefore widely used for early detection of Alzheimer's disease [9],
[10] ,
[11] . The advantage of PET is that pathophysiological changes can be observed before clinical onset, as well as structural neurodegeneration that can be observed on MRI
[12] ,
[13] . Specifically, PET uses amyloid-beta (Aβ)
[14] and tau
[10] radiotracers to image molecular biomarkers associated with AD, and fluorodeoxyglucose (FDG)
[11] to image metabolic activities associated with neurodegeneration. This multi-tracer PET enables neuroradiologists to comprehensively assess spatiotemporal pathophysiological changes within the diagnostic framework of amyloid-tau-neurodegeneration (ATN) biomarkers
[15] . However, limited accessibility, high cost, and radiation exposure to PET imaging have hampered its widespread use in clinical practice, thus emphasizing the need for alternative methods to acquire PET imaging data.
[0004] Deep generative artificial intelligence (AI) models have enabled the synthesis of cross-modality, tracer-specific PET (e.g., Aβ
[14] , tau
[10] , and FDG tracers
[11] ) from T1-weighted (T1w) MRI. However, most of these studies (1) have unsatisfactory image quality of synthesized PET and are unable to display various pathophysiological changes; and (2) lack systematic evaluation of the clinical utility of synthesized PET. In addition, these methods generally focus on single-tracer PET synthesis and ignore the interrelationships between ATN biomarkers revealed by multi-tracer PET. Summary of the invention
[0005] The purpose of the present invention is to provide a multi-tracer PET cross-modality synthesis system based on diffusion transformer, which has high PET image quality, strong detection capability and fast reasoning speed.
[0006] The multi-tracer PET cross-modal synthesis system based on diffusion transformer proposed in the present invention first designs a novel generative diffusion transformer model, denoted as GenPET, which is used to synthesize multi-tracer PET images from structural MRI images, such as Figure 2 As shown;
[0007] The design of GenPET is based on two basic conditions: (1) MRI provides a large amount of biological information for PET synthesis
[11] ,
[13] ,
[14] ; (2) the mapping relationship between structure and pathophysiology of specific tracers can be understood from paired MRI and multi-tracer PET data
[16] ,
[17] . Therefore, the present invention first maps all modality images into a unified latent representation space (also called latent space), and uses the latent space diffusion Transformer model to synthesize MRI into PET of a specified tracer type; that is, GenPET is obtained.
[0008] Specifically, GenPET consists of two main stages: the first stage is unified latent space representation learning, and the second stage is cross-modal synthesis in the unified latent space.
[0009] In the first stage, a 3D encoder [4] is used to encode MRI and multi-tracer PET images into a unified latent space, and a modality discrimination loss is designed to ensure the discrimination of different modalities in the latent space.
[0010] In the second stage, the latent space diffusion Tansformer model is used to realize the latent space feature conversion from MRI to specific tracer PET.
[0011] Further:
[0012] The task of synthesizing 3D multi-tracer PET images from structural T1w MRI is formulated as a deep learning framework. Given an MRIM∈R D×H×W (D, H, and W represent the depth, height, and width of the three-dimensional image), and the specific PET tracer type τ∈{Aβ,Tau,FDG} to be synthesized. The goal is to use the proposed GenPET(·) to synthesize PET images of a specific tracer, expressed as:
[0013]
[0014] Among them, Aβ, Tau, and FDG are designated tracer types for PET, which are used to designate imaging of AD-related biomarkers or neurological metabolic changes. Synthesized PET images It should be compared with the real PET image P τ Specifically, GenPET(·) integrates a unified 3D VAE[4] to encode MRI and multi-tracer PET images into a unified latent space. It also integrates a latent space diffusion Transformer network to convert MRI latent space representation into tracer-specific PET latent space features. For the specific network architecture, see Figure 3 and 5 As shown. Among them:
[0015] 3D Unified VAE It consists of an encoder ε(·) and a decoder , its architecture can be found in Figure 3 As shown in Figure 2. The unified encoder ε(·) adopts a multi-scale architecture with output channel sizes of 16, 64, 96, and 32 respectively. The unified decoder The input channel sizes are used in reverse order, i.e., 32, 96, 64, and 16, respectively. Both the encoder and decoder contain two residual blocks per layer to ensure strong feature extraction. The channel dimension of the final intrinsic space is set to 32 to balance the complexity and compactness of the model. Group Normalization (GroupNorm) of eight groups is applied across layers. In addition, a multi-head self-attention mechanism is introduced in the deepest layers of the encoder and decoder (where the channel size is 32) to enhance the relationship between features at different spatial locations. It should be noted that the size of the original image space is 1×192×224×192, while the unified latent space is compressed to 32×24×28×24.
[0016] In addition, the present invention also introduces a discriminator Dis(·) and a pre-trained 4-way 3D modality discriminant residual network (ResNet) [6] Participate in the training process of 3D unified VAE network, such as Figure 4 As shown in the figure, the discriminator Dis(·) processes real or synthetic PET images through four convolutional layers and outputs a block of size 19×23×19 to distinguish true from false. In order to unify images of different modalities into a shared latent space and enhance the separation between their encoded latent representations, a pre-trained 4-way modality discriminant residual network (ResNet) is used. [6] For the corresponding modal type τ m ∈{MRI,Aβ,Tau,FDG} for classification.
[0017] The latent diffusion Transformer network uses the encoded MRI latent space features and PET tracer type as input to synthesize the corresponding PET latent space features. Its architecture is shown in Figure 5 As shown in the figure. Specifically, the encoded MRI and intermediate variable features in the diffusion process are divided into patches with a spatial resolution size of 2×2×2. These patches are then concatenated and passed through N=13 latent space diffusion Transformer blocks with a dimension of 1536 for each token to estimate the noise added during the diffusion process. This joint input strategy promotes the seamless integration of self- and cross-attention mechanisms, allowing the latent space features of intermediate diffusion variables and MRI to work independently while achieving dynamic interaction. In addition, this approach helps to mitigate the stochasticity of the diffusion process, ensure that the synthesized PET images closely correspond to the reference MRI images, and enhance the model's ability to capture various pathophysiological changes. The established attention mechanism uses 24 multi-heads for sub-dimensional feature calculation and fusion, and uses a fast attention mechanism (FlashAttention) to improve training and inference efficiency. It is worth noting that the present invention can synthesize multi-tracer PET images in all ATN stages by simply specifying the required PET tracer type. The tracer type, denoted as τ∈{Aβ,Tau,FDG}, is used to dynamically adjust the feature distribution of each layer of the Transformer through the AdaIN module (normalization module).
[0018] The present invention also introduces a series of learnable parameters, Initialized to 0.5, it is used for linear interpolation in skip connections, connecting the output of the i-th Transformer block to the input of the (Ni) block, ensuring efficient training and effective feature integration. In addition, the LayerNorm operation is further optimized to improve the overall efficiency of training and inference.
[0019] Further:
[0020] In the first stage, i.e., the unified latent space representation learning stage, the 3D VAE proposed in this invention It is used to learn latent space representations of MRI and multi-tracer PET in a unified latent space; it consists of an encoder ε(·), a decoder in:
[0021] The encoder ε(·) is used to project MRI and multi-tracer PET images into a unified latent space of size c×d×h×w; the decoder Used to reconstruct the image back to the original space. Here, c, d, h, and w represent the channel, depth, height, and width of the latent space representation, respectively.
[0022] During the training phase, the neural image x and its corresponding modality type τm ∈{MRI,Aβ,Tau,FDG} are randomly sampled. Following the standard latent diffusion training strategy [1], the proposed unified 3D VAE is trained using a combination of l1 reconstruction loss, perceptual loss, adversarial loss [3], KL divergence loss [4] and a novel modality discrimination loss; where:
[0023] The l1 reconstruction loss provides voxel-level supervision and is defined as follows:
[0024]
[0025] Real and synthetic PET images were unfolded along three 3D axes (axial, sagittal, and coronal), and the slice numbers for each axis were denoted as I, J, and K, respectively.
[0026] The perceptual loss is calculated using the pre-trained 2D LPIPS metric LPIPS(·,·), which is based on a pre-trained 2D squeeze-and-excite network [5] and is defined as:
[0027]
[0028] The adversarial loss is formulated in least squares form to distinguish between real and synthetic patches:
[0029]
[0030] The KL divergence loss constrains the unified latent space to approximate the standard normal distribution:
[0031]
[0032] Where d is the dimension of the latent space, μ i and σ i are the mean and standard deviation of the i-th latent space dimension.
[0033] In order to unify images of different modalities into a shared latent space and enhance the separation between their encoded latent representations, a pre-trained 4-way 3D modality discriminative residual network (ResNet) is used. [6] For the corresponding modal type τ m ∈{MRI,Aβ,Tau,FDG} for classification. The corresponding modality discrimination loss is defined as follows:
[0034]
[0035] Among them, τ m is a real modal type, It is a three-dimensional modality discrimination ResNet For reconstructed images The predicted modality type, i.e. The final objective function for optimizing the 3D unified VAE is defined as follows:
[0036]
[0037] Among them, λ per , GAN , KL and λ MDis is a hyperparameter that balances the weights of these five objective functions. Specifically, it can be set to 0.1, 0.1, 1×10 -6 and 1, following the standard latent space diffusion training paradigm. Note that the 3D unified VAE can be repeatedly optimized and the discriminator Dis(·), maximizing the objective in formula (4).
[0038] In the second stage, i.e., the cross-modal synthesis stage in the unified latent space, the pre-trained 3D unified encoder ε(·) and unified decoder (·) and frozen weights are used to further train the latent space diffusion Transformrer network. Thus, cross-modality synthesis from MRI to multi-tracer PET is achieved in a unified latent space. The entire diffusion process can be expressed as:
[0039]
[0040] where time step t∈{1,2,…,T} and T=1000, ε(M)∈R c×d×h×w and ε(P τ )∈R c×d×h×w denote the encoded MRI and specific tracer PET latent space representations respectively. The t-th latent variable z1,…,z T With size c×d×h×w, the forward process and the reverse process are denoted by q(·) and p respectively. θ (·)express.
[0041] In the forward diffusion process, following the standard forward equation [7],
[18] ,
[19] , the Gaussian noise is gradually is introduced into the encoded PET latent space representation z0 with any tracer type τ, thereby obtaining the corresponding t-th latent variable z t ∈R c×d×h×w :
[0042]
[0043] in β s is a predefined variance value
[20] .
[0044] In the diffusion reverse process, the encoded MRI latent space feature ε(M) and the diffused t-time intermediate noise latent space feature z t Segmented into segments, connected and fed into the latent space diffusion Transformer network At the same time, the specified tracer type τ is injected through the AdaIN module to predict the Gaussian noise added at time step t, that is, . The joint input of MRI encoding and diffuse intermediate noise latent features enables seamless integration of self-attention and cross-attention mechanisms, allowing them to function independently while achieving dynamic interaction. In addition, we also adopt linear interpolation, flash attention and efficient layer normalization strategies in skip connections to achieve efficient training and effective feature integration. It is worth noting that GenPET can synthesize multi-tracer PET by simply controlling the input tracer type. Latent Space Diffusion Transformer Network Optimization using noise-based mean square error loss:
[0045]
[0046] For inference, the Gaussian noise The t-th latent space variable z is sampled and connected to the encoded MRI latent space feature ε(M). Given an arbitrary synthetic tracer type τ, the present invention uses the denoised diffusion latent model (DDIM) as a diffusion sampler to iteratively synthesize the t-th latent space variable z with 50 time steps. t , until the initial number of time steps is reached. Finally, the unified decoder is used The initial latent variable for the synthesis Decode and synthesize the corresponding PET image P τ ,Right now The total time required to synthesize any tracer PET is approximately 5 seconds.
[0047] GenPET was trained, validated, and tested on a dataset of 869 MRI-PET samples from three local clinical centers (Huashan Hospital Main Campus, Shanghai Oriental Hospital, and the Sixth People's Hospital Affiliated to Shanghai Jiao Tong University). At the same time, GenPET was visually validated using 106 samples from a local clinical center (Huashan Hospital Hongqiao Campus) and diagnostically validated using 1207 samples from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database. The imaging quality and diagnostic utility of the multi-tracer PET synthesized by GenPET were systematically evaluated through visual Turing tests by neuroradiologists, hierarchical quantitative indicators, clinical correlations, and radiomics analysis. Extensive experimental results show that GenPET provides a non-invasive, high-quality, and effective alternative to traditional multi-tracer PET, which is expected to address the problem of scarce imaging resources in underserved areas and improve the diagnosis of neurodegenerative diseases.
[0048] Existing MRI-to-PET studies for diagnosing AD
[10] ,
[11] ,
[12] ,
[13] ,
[14] usually use CNN or GAN technology to convert a single PET tracer. In contrast, the GenPET introduced in this paper combines the advantages of diffusion and transformer technology to explore the unified latent space of MRI and multi-tracer PET, which is significantly better than the existing SOTA MRI-to-PET and other cross-modal synthesis methods. GenPET has the following new features:
[0049] (1) GenPET integrates Aβ, tau, and FDG PET tracers into a unified training framework, providing a simple and powerful solution for accurately capturing tracer-specific pathophysiological changes;
[0050] (2) GenPET uses the powerful image modeling capabilities of latent spatial diffusion mechanisms to effectively capture complex and diverse pathophysiological patterns, thereby generating high-quality multi-tracer PET images;
[0051] (3) GenPET uses the Transformer architecture to enhance the global understanding of the anatomical relationships between different brain regions, and its built-in attention mechanism ensures that the generated PET images are closely related to the input MRI images;
[0052] (4) GenPET does not require all three PET tracers for a single subject during training, which reduces the burden of data collection;
[0053] (5) GenPET has a high inference efficiency. It can synthesize any tracer PET image from MRI data in about five seconds, which is much faster than the actual PET acquisition time. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is the technical roadmap of the present invention.
[0055] Figure 2 This is the structural diagram of the GenPET model proposed in the present invention.
[0056] Figure 3 This is a diagram of the 3D VAE network architecture proposed in the present invention.
[0057] Figure 4 This is the 3D VAE network training diagram proposed in this invention.
[0058] Figure 5 This is the proposed latent space diffusion Transformer network architecture diagram.
[0059] Figure 6 Qualitative results and visual Turing test results of synthetic multi-tracer PET images in the internal test cohort. (ac) Synthetic Aβ, tau, and FDG PET images and corresponding real and T1w MRI scans. (dg) Visual Turing test results of synthetic multi-tracer PET images evaluated by two neuroradiologists with 6 and 10 years of experience, respectively, for real versus synthetic discrimination, image structural fidelity, subjective visual lesion conspicuity, and diagnostic confidence; the latter three tasks were evaluated on a 4-point scale, with higher scores indicating better performance.
[0060] Figure 7 Hierarchical quantitative results of synthetic multi-tracer PET were evaluated at the region of interest (ROI), lobar, and whole-brain voxel levels. (ab) Box plot results of mean absolute error (MAE) and Pearson correlation of 40 ROIs and specific tracer meta-ROIs in the internal test cohort, with red boxes highlighting meta-ROI results. (cd) MAE and Pearson correlation results of different lobes from the internal test cohort, respectively. (e) Whole-brain voxel-level MAE and SSIM results in the internal test cohort. (f) Whole-brain voxel-level MAE and SSIM results from an external local clinical center.
[0061] Figure 8The results of the correlation between synthetic multi-tracer PET and various AD-related clinical characteristics. Among them, (a) the correlation between brain regions and meta-ROI SUVR values and MMSE scores in the internal test set. (b) The correlation between meta-ROI SUVR values and MMSE scores in the external ADNI group. (cd) The correlation between meta-ROI SUVR values and years of education in the internal test group and the external ADNI group, including regression lines and 95% confidence intervals. (ef) Inter-group comparison of meta-ROI SUVR values and apolipoprotein E4 (APOE4) gene status in the internal test group and the ADNI group. The internal results include real and synthetic multi-tracer PET, while the external ADNI results are based only on synthetic multi-tracer PET.
[0062] Fig. 9 Diagnostic results of synthetic multi-tracer PET. (ac) Internal stratification results of meta-ROI SUVR values from real and synthetic Aβ, tau and FDG PET images in different clinical subgroups. (dg) COG in ADNI dataset NC Task (differentiation of NC from MCI and AD cases), COG MCI Task (differentiating MCI from NC and AD cases), COG AD External receiver operating characteristic (ROC) and precision-recall (PR) curves for the task (differentiating AD from NC and MCI cases) and the MCI progression task (differentiating sMCI from pMCI cases).
[0063] Fig.10 Qualitative comparison of AβPET synthesized for GenPET and other SOTA competing generation methods. Two cases randomly selected from the internal test cohort. (a) Normal cognition (NC) case with negative Aβ imaging (b) Alzheimer's disease (AD) case with positive Aβ+ imaging.
[0064] Fig.11 Qualitative comparison of tau PET synthesized for GenPET and other SOTA competing generation methods. Two cases randomly selected from the internal test cohort. (a) Normal cognition (NC) case with negative Aβ imaging (b) Alzheimer's disease (AD) case with positive Aβ+ imaging.
[0065] Fig.12 Qualitative results comparison of FDG PET synthesized for GenPET and other SOTA competing generation methods. Two cases randomly selected from the internal test cohort. (a) Normal cognition (NC) case with negative Aβ imaging (b) Alzheimer's disease (AD) case with positive Aβ+ imaging.
[0066] Fig.13Synthetic Aβ PET qualitative results for six cases randomly selected from an external clinical center.
[0067] Fig.14 Composite tau PET qualitative results for six cases randomly selected from an external clinical center.
[0068] Fig.15 Composite FDG PET qualitative results of six cases randomly selected from an external clinical center. DETAILED DESCRIPTION
[0069] The following further demonstrates the internal and external data sets used for model development and validation. The internal data sets used for GenPET model training, validation and testing come from three local clinical centers: Huashan Hospital Main Campus, Shanghai Oriental Hospital and Shanghai Jiaotong University Affiliated Sixth People's Hospital. The external data sets include another local clinical center (Huashan Hospital Hongqiao Campus) and ADNI database. The specific implementation technology route is as follows: Figure 1 shown.
[0070] The present invention adopts the "training, internal testing and external validation" paradigm for specific research, and there are no overlapping experimental subjects between these data sets. Given that most of the internal MRI-PET paired samples include both AβPET and tau or FDG PET, we first selected 30% of the participants who underwent MRI-AβPET-tau PET and MRI-AβPET-FDG PET, divided them into validation sets (10%) and test sets (20%), and the remaining samples were used for training. This division method ensures that the ratio of training sets, validation sets, and test sets for different tracer PETs is approximately 7:1:2, with no overlap of participants. At the same time, data augmentation is applied to the training data, including random flipping along the x-axis, y-axis, and z-axis, and a translation of ±30 voxels in each direction to improve the robustness of the model. All deep learning-based methods in the present invention are implemented using the PyTorch library and trained on a single NVIDIA A10080G GPU. The 3D VAE network involved in GenPET ( Figure 3 ) and latent space diffusion Transformer network ( Figure 5 ) was initialized using the Kaiming method and optimized using the Adam algorithm. To select the best network weights during training, an early stopping method based on internal validation performance was used. After 10 epochs of warm-up, a half-cycle cosine learning rate adjustment strategy was used, starting from 2×10 -5 We start with an initial learning rate of 1 and a batch size of 2 to train the 3D unified VAE. For the latent space diffusion Transformer model, we use 1×10 -5A fixed learning rate of is used and the batch size is 8. In addition, the exponential moving average method (EMA) is applied to update the weights of the unified VAE and latent space diffusion Transformer networks.
[0071] In the specific implementation process, the Gaussian noise is first The t-th latent space variable z is sampled and connected with the MRI latent space feature ε(M) encoded by the 3D VAE. Given an arbitrary synthetic tracer type τ, the present invention uses the denoised diffusion latent model (DDIM) as a diffusion sampler to iteratively synthesize the t-th latent space variable z through the latent space diffusion Transformer model with 50 time steps. t , until the initial number of time steps is reached. Finally, the decoder of the 3D VAE is used The initial latent variable for the synthesis Decode and synthesize the PET image P of the corresponding specific tracer τ , thus carrying out the corresponding synthetic multi-tracer PET analysis.
[0072] Experimental Example 1: Synthetic multi-tracer PET reveals spatiotemporal sequential pathophysiological changes with higher structural fidelity
[0073] The image quality of synthetic multi-tracer PET was qualitatively evaluated in an internal test cohort that divided the samples into four subgroups: AβPET-negative cognitive normal (NC), AβPET-positive cognitive normal (NC), Aβ+ mild cognitive impairment (MCI), and Aβ+ Alzheimer's disease (AD). The PET images synthesized by GenPET are not only visually very similar to real images, but also can effectively capture the pathophysiological changes of different tracers and subgroups ( Figure 6 ac). Synthetic multi-tracer PET images reveal the spatiotemporal order of changes associated with the progression of attention deficit disorder: Synthetic Aβ PET images depict early amyloid deposition at the initial stage ( Figure 6 a), Synthetic tau positron emission tomography shows a regional progressive accumulation from the medial temporal lobe to the neocortex ( Figure 6 b), Synthetic FDG PET image highlights the metabolic reductions primarily associated with advanced AD ( Figure 6 c).
[0074] The present invention also conducted a blind visual Turing test to comprehensively evaluate the quality of synthetic and real multi-tracer PET images in terms of real and synthetic identification, image structure fidelity, subjective visual lesion clarity and diagnostic confidence. Two neuroradiologists with 6 and 10 years of experience, respectively, independently evaluated all images, and the agreement kappa
[28] between neuroradiologists ranged from [0.421, 0.777]. The identification results of real and synthetic ( Figure 6 d) shows that 88.4% of synthetic AβPET, 27.3% of synthetic tau PET, and 89.2% of synthetic FDG PET were identified as real PET, highlighting the powerful synthesis capability of GenPET. Figure 6 e), using a 4-point system, the average score of the synthetic PET was significantly better than that of the real PET (Aβ: 3.81 vs 3.51, p = 2.69e-3; tau: 3.00 vs 2.80, p = 6.58e-3; FDG: 3.89 vs 3.42, p = 7.00e-7, Wilcoxon Signed-Rank test), indicating that the multi-tracer PET synthesized by GenPET can better locate pathophysiological changes.
[0075] 84.8% of synthetic Aβ PET, 38.0% of synthetic tau PET, and 93.9% of synthetic FDG PET achieved the same or better subjective visual lesion conspicuity scores as their real counterparts, with 21.7% of synthetic tau PET imaging scores significantly reduced from higher (4 or 3) to lower (2 or 1) ( Figure 6 f). Similarly, 84.2% of synthetic AβPET, 62.0% of synthetic tauPET, and 97.7% of synthetic FDG PET achieved the same or better diagnostic confidence scores, with 17.8% of synthetic tauPET showing a substantial decrease ( Figure 6 g). These results suggest that the coexistence of FDG PET and structural MRI facilitates the synthesis of metabolic changes in FDG PET during the neurodegenerative (N) stage. The synthesis of pathological changes in Aβ and tau PET is more challenging, especially tau PET due to its higher specificity.
[0076] Experimental Example 2: The image quality of synthetic Aβ and FDG PET is comparable to that of real images
[0077] We quantitatively evaluated the image quality of multi-tracer PET images synthesized by GenPET in an internal test cohort at the region of interest (ROI), lobar, and whole-brain voxel levels.
[0078] In the quantitative analysis at the ROI level, 40 ROIs and tracer-specific meta-ROIs were selected to compare the SUVR values between the synthetic multi-tracer PET images and the real images, measured by mean absolute error (MAE) and Pearson correlation (r). Figure 7 a, b) show that the quality of synthetic Aβ and FDG PET images is comparable to that of real images (average MAE / r = 0.123 / 0.890 for Aβ and 0.086 / 0.904 for FDG), while synthetic tauPET has poor image quality due to its high specificity (average MAE / r = 0.137 / 0.666 for tau). When zooming in on individual ROIs, it was found that ROIs with lower MAE values or higher r values corresponded to higher structure-pathophysiological correlations, for example, the medial superior frontal gyrus (Frontal_Sup_Medial, MAE of Aβ was 0.095±0.082, r = 0.944±0.026; MAE of tau was 0.071±0.075, r = 0.841±0.019; MAE of FDG was 0.056±0.046, r = 0.960±0.012). In contrast, ROIs with larger differences corresponded to more tracer-specific pathophysiological changes, such as the inferior temporal gyrus (Temporal_Inf, MAE of Aβ was 0.147±0.127, r=0.856±0.061; MAE of Tau was 0.220±0.337, r=0.480±0.012). In contrast, ROIs with larger differences corresponded to pathophysiological changes with stronger tracer specificity, such as the inferior temporal gyrus (Temporal_Inf, MAE for Aβ was 0.147±0.127, r=0.856±0.061; MAE for tau was 0.220±0.337, r=0.480±0.159; MAE for FDG was 0.086±0.072, r=0.828±0.061) and the inferior parietal gyrus (Parietal_Inf, MAE for Aβ was 0.138±0.129, r=0.849±0.095). The MAE of FDG was 0.086 ± 0.072 and r = 0.828 ± 0.061, the MAE of Parietal_Inf was 0.138 ± 0.129 and r = 0.849 ± 0.095 for Aβ, the MAE of ParaHippocampal was 0.163 ± 0.188 and r = 0.567 ± 0.123, and the MAE of Angular was 0.110 ± 0.083 and r = 0.868 ± 0.049 for FDG. Figure 7a, b) show that the MAE of synthetic FDG is the lowest, which is 0.089±0.072, and the MAE of synthetic Aβ is the highest, r=0.917±0.034, while the MAE of synthetic tau is 0.193±0.158 and r=0.592±0.084.
[0079] In order to perform leaf-level quantitative analysis, the present invention divides these 40 ROIs into five lobes for comparison. The results of the five lobes ( Figure 7 c, d) show that the image quality of synthetic Aβ and FDG PET is comparable to that of real images (mean MAE / r = 0.086 / 0.907 for Aβ and 0.124 / 0.884 for FDG), while the results of synthetic tau PET are not ideal (mean MAE / r = 0.161 / 0.617). When zooming in on individual lobes, we found that the frontal lobe consistently showed lower MAE and higher r among the different tracers, suggesting a stronger structural pathophysiological relevance in the AD stage (MAE 0.102 ± 0.089, r = 0.931 ± 0.029 for Aβ; MAE 0.098 ± 0.139, r = 0.805 ± 0.093 for tau; MAE 0.069 ± 0.051, r = 0.948 ± 0.013 for FDG). In contrast, the temporal lobe showed great pathophysiological variability, especially for tau (MAE 0.161 ± 0.175, r = 0.625 ± 0.128), while FDG PET still showed strong synthesis performance (MAE 0.081 ± 0.065, r = 0.889 ± 0.029). Aβ PET faced challenges in the occipital lobe (MAE 0.140 ± 0.120, r = 0.850 ± 0.079), which was due to heterogeneous changes in the brain lobe and sparse Aβ deposition.
[0080] For whole-brain voxel-wise quantitative analysis, MAE and the structural similarity index measure (SSIM) were used for comparison. Figure 7e) showed that GenPET significantly ranked the image quality of synthetic PET as follows: FDG>Aβ>tau (average MAE / SSIM=0.0361 / 0.921 for FDG, average MAE / SSIM=0.0602 / 0.894 for Aβ, average MAE / SSIM=0.0617 / 0.872 for tau; F=11.6, p=1.92e-5 for MAE, F=4.01, p=2.11e-2 for SSIM; one-way ANOVA, Holm-Sidak post-test). In addition, we compare GenPET with eight SOTA generation methods, including CNN-based models (UNet
[21] , DenseUNet
[10] , and SwinUNetr
[22] ), GAN-based models (Pix2pix
[23] , EaGAN
[24] , and ShareGAN v2
[11] ), and diffusion-based models (Latent Diffusion Model (LDM)
[19] and Adaptive LDM (ALDM)
[25] ). Figure 7 e) shows that GenPET achieves the best synthesis performance, significantly surpassing all competing methods with low MAE of 0.0361 to 0.0617 and high SSIM of 0.872 to 0.921 (all p<0.05, one-sided Wilcoxon Signed-Rank test). We also show qualitative results compared with the internal dataset of the SOTA method ( Figure 10-12 ), the specific internal whole-brain quantitative performance comparison results are shown in Table 1, and the present invention is the best.
[0081] The generalizability of GenPET was further validated at an external local clinical center. Results at the whole-brain voxel level ( Figure 7 f) shows the same conclusion as the internal test cohort, but the performance is slightly reduced due to differences in resolution and acquisition equipment. The corresponding external dataset qualitative results are as follows Figure 13-15 As shown, the specific external whole-brain quantitative performance comparison results are shown in Table 2, and the present invention is the best.
[0082] Experimental Example 3: Correlation of synthetic multi-tracer PET with multiple AD-related clinical features
[0083] We evaluated the correlations between synthetic PET and various AD-related clinical features including cognitive scale scores, educational background, and genetic predisposition to assess whether these correlations were consistent with their true counterparts on both internal tests and the external ADNI dataset. For cognitive scores measured by the Mini-Mental State Examination (MMSE), we found significant negative correlations with meta-ROI SUVR values for synthetic Aβ and tauPET, and positive correlations with SUVR values for synthetic FDG (all p<0.05, t-test), which were consistent with their true counterparts on both internal tests and the external ADNI dataset ( Figure 8 a,b). Correlation analysis of internal test data ( Figure 8 a) showed that synthetic tau PET had higher specificity than synthetic FDG and Aβ PET in reflecting cognitive changes, as reflected by larger absolute correlation coefficients and more significant regions, while synthetic FDG PET performed poorly. Correlation analysis of external ADNI data ( Figure 8 b) shows that the meta-ROI SUVR values of synthetic tau PET showed good absolute correlation coefficients among different tracers, while the meta-ROI SUVR of synthetic Aβ began to increase in the early stage (MMSE=25), and the meta-ROI SUVR of synthetic FDG began to decrease in the late stage (MMSE=22).
[0084] In terms of education level, we found that the meta-ROI SUVR values of synthetic Aβ and tauPET were significantly negatively correlated with education level, while FDG was positively correlated with education level in both the internal test cohort and the external ADNI cohort (p < 0.05, t-test), even though the mean years of education in the internal test group and the external ADNI group were 11.5 and 15.9 years, respectively, which were consistent with the real group ( Figure 8 c, d). For AD-related genetic risk factors using the apolipoprotein E4 (APOE4) allele
[36] , we found that APOE4 carriers showed significantly larger meta-ROI SUVR values in synthetic Aβ and tauPET and smaller SUVR values in synthetic FDG PET compared with non-carriers in both the internal and external ADNI cohorts (p < 0.05, Mann-Whitney U test), which is consistent with the actual situation ( Figure 8 e,f).
[0085] Experimental Example 4: Synthetic multi-tracer PET can stratify AD subgroups, with tau PET being more discriminatory
[0086] We performed a meta-ROI SUVR subgroup stratification analysis on the internal test set, which was divided into four subgroups: NC Aβ-, NC Aβ+, MCI Aβ+, and AD Aβ+. Fig. 9 ac) showed that the tracer-specific meta-ROISUVR value changes of synthetic PET were consistent with those of real PET: Aβ and tau PET increased with the progression of AD stage, while FDG decreased. When zooming in on individual tracers, the significant subgroup differences of synthetic tau PET were consistent with those observed with real tracers and outperformed other tracers ( Fig. 9 b), which highlights the superior discriminatory power of synthetic tau PET. After careful comparison with real tracers, we found that synthetic AβPET did not clearly distinguish between the MCI Aβ+ subgroup and the AD Aβ+ subgroup, and synthetic FDG PET did not clearly distinguish between the NC Aβ- subgroup and the MCI Aβ+ subgroup. However, synthetic FDG PET was able to clearly distinguish between the MCI Aβ+ subgroup and the AD Aβ+ subgroup (p=3.92e-2, one-way ANOVA, Holm-Sidak post-test), which was not clearly observed in real PET images. We note that synthetic multi-tracer PET mainly captures population-level changes, such as metabolic changes detected by synthetic AβPET positivity in early AD and synthetic FDG PET in late disease.
[0087] Experimental Example 5: Synthetic multi-tracer PET can improve the diagnostic performance of MRI-based AD-related tasks
[0088] We performed radiomics analysis on various AD-related tasks on the external ADNI dataset to evaluate the diagnostic performance of synthetic PET. Specifically, we designed four AD-related tasks: (1) COG NC Task, used to distinguish NC from MCI / AD; (2) COG MCI Task, used to distinguish between MCI and NC / AD; (3) COG ADtask, used to distinguish AD from NC / MCI; and (4) an MCI progression task, used to distinguish stable MCI (sMCI) from progressive MCI (pMCI) within 36 months. For all tasks, five experimental methods combining different modalities were used: MRI only, MRI + synthetic AβPET, MRI + synthetic Tau PET, MRI + synthetic FDG PET, and MRI + synthetic multi-tracer PET. 310 MRI features and 41 SUVR values from 40 ROIs pre-extracted from FreeSurfer
[38] were used, as well as one tracer-specific meta-ROI for each PET tracer. For a fair comparison of different methods, Lasso regression was used to select 60 features for radiomics analysis, where random forest classification with 5-fold cross validation was used
[26] . The area under the receiver operating characteristic curve (AUC) and the area under the precision-recall curve (AP) were used to evaluate the diagnostic performance.
[0089] Radiomics analysis of the four tasks ( Fig. 9 dg) showed that the multi-tracer PET synthesized by GenPET (MRI+All) can significantly and continuously improve the diagnostic performance of pure MRI (all p < 0.05, one-sided Wilcoxon signed rank test). In different combinations of MRI and multi-tracer PET, the MRI+synthetic tau PET model (MRI+Tau) showed superior diagnostic performance and significantly improved the diagnostic performance of pure MRI, with an AUC increase of 18.1% (COG MCI ), AP improved by 16.5% (MCI progression). In four AD-related diagnostic tasks, the AUC / AP scores of MRI+Tau were 0.890 [95% confidence interval (CI): 0.879, 0.901] / 0.904 [CI: 0.895, 0.913], 0.783 [CI: 0.771, 0.805] / 0.897 [CI: 0.887, 0.908], 0.902 [CI: 0.887, 0.912] / 0.964 [CI: 0.953, 0.973], and 0.783 [CI: 0.776, 0.810] / 0.635 [CI: 0.613, 0.647], respectively. The MRI+synthetic AβPET model (MRI+Aβ) was significantly superior to the COG in terms of AUC / AP. NC The performance in the task was competitive, with AUC / AP scores of 0.887 [CI: 0.875, 0.895] / 0.899 [CI: 0.889, 0.907]. At the same time, the magnetic resonance imaging + synthetic FDG PET model (MRI+FDG) performed well in COG MCI , COG ADThe performances in the three tasks of , and MCI progression prediction were competitive, with corresponding AUC / AP scores of 0.771 [CI: 0.745, 0.793] / 0.889 [CI: 0.882, 0.900], 0.894 [CI: 0.885, 0.907] / 0.962 [CI: 0.953, 0.970], and 0.783 [CI: 0.765, 0.799] / 0.600 [CI: 0.577, 0.614]. In summary, the diagnostic results of the external ADNI dataset validated the effectiveness of GenPET synthesizing multi-tracer PET.
[0090] Table 1. Comparison of internal whole-brain voxel-level quantitative performance of GenPET and other SOTA competitive generation methods
[0091]
[0092]
[0093] Table 2. Comparison of external whole-brain voxel-level quantitative performance of GenPET and other SOTA competitive generation methods
[0094]
[0095] References
[0096] [1]Z.Fei, M.Fan, C.Yu, and J.Huang, "Scalable Diffusion Models with StateSpace Backbone," Feb.08, 2024, arXiv:arXiv:2402.05608.
[0097] [2] J. Johnson, A. Alahi, and L. Fei-Fei, "Perceptual Losses for Real-TimeStyle Transfer and Super-Resolution," arXiv:1603.08155[cs], Mar. 2016.
[0098] [3] T.Karras, S.Laine, and T.Aila, "A Style-Based Generator Architecture for Generative Adversarial Networks," p.10.
[0099] [4]D.P.Kingma and M.Welling,“Auto-Encoding Variational Bayes,”arXiv:1312.6114[cs,stat],May 2014.
[0100] [5]J.Hu,L.Shen,S.Albanie,G.Sun,and E.Wu,“Squeeze-and-ExcitationNetworks,”arXiv:1709.01507[cs],May 2019.
[0101] [6]S.Korolev,A.Safiullin,M.Belyaev,and Y.Dodonova,“Residual and PlainConvolutionalNeural Networks for 3D Brain MRI Classification,”Jan.23,2017,arXiv:arXiv:1701.06643.
[0102] [7]W.Peebles and S.Xie,“Scalable Diffusion Models with Transformers,”in 2023IEEE / CVFInternational Conference on Computer Vision(ICCV),Paris,France:IEEE,Oct.2023,pp.4172–4182.
[0103] [8]X.Huang and S.Belongie,“Arbitrary Style Transfer in Real-Time withAdaptive InstanceNormalization,”in 2017IEEEInternational Conference onComputer Vision(ICCV),Venice:IEEE,Oct.2017,pp.1510–1519.
[0104] [9]W.Cui et al.,“Bilinear pooling and metric learning network forearly Alzheimer’s diseaseidentification with FDG-PET images,”Nov.09,2021,arXiv:arXiv:2111.04985.
[0105]
[10] J.Lee,“Synthesizing images of tau pathology from cross-modalneuroimaging using deeplearning”.
[0106]
[11] C.Wang et al.,“Joint learning framework of cross-modal synthesisand diagnosis forAlzheimer’s disease by mining underlying shared modalityinformation,”Med.ImageAnal.,vol.91,p.103032,2024.
[0107]
[12] Y.Chen,Y.Pan,Y.Xia,andY.Yuan,“Disentangle First,Then Distill:AUnified Framework forMissing Modality Imputation and Alzheimer’s DiseaseDiagnosis,”IEEE Trans.Med.Imag.,pp.1–1,2023.
[0108]
[13] Y.Pan,M.Liu,C.Lian,Y.Xia,and D.Shen,“Spatially-Constrained FisherRepresentation forBrain Disease Identification With Incomplete Multi-ModalNeuroimages,”IEEE Trans.Med.Imag.,vol.39,no.9,pp.2965–2975,Sep.2020.
[0109]
[14] Z.Ou,Y.Pan,Y.Li,F.Xie,Q.Guo,and D.Shen,“Synthesizing Aβ-Pet ViaAn Image AndLabel Conditioning Latent Diffusion Model For DetectingAmyloidStatus,”in ICASSP 2024-2024IEEE International Conference on Acoustics,Speechand Signal Processing(ICASSP),Seoul,Korea,Republic of:IEEE,Apr.2024,pp.6610–6614.
[0110]
[15] R.C.Petersen et al.,“Alzheimer’s Disease Neuroimaging Initiative(ADNI):Clinicalcharacterization,”Neurology,vol.74,no.3,pp.201–209,Jan.2010.
[0111]
[16] T.Wang and X.Yang,“Take CT,get PET free:AI-powered breakthroughin lung cancerdiagnosis and prognosis,”CellReportsMedicine,vol.5,no.4,p.101486,Apr.2024.
[0112]
[17] M.Salehjahromi et al.,“Synthetic PET from CT improves diagnosisand prognosis for lungcancer:Proofofconcept,”CellReports Medicine,vol.5,no.3,p.101463,Mar.2024.
[0113]
[18] P.Esser et al.,“Scaling Rectified Flow Transformers for High-Resolution Image Synthesis”.
[0114]
[19] R.Rombach,A.Blattmann,D.Lorenz,P.Esser,and B.Ommer,“High-Resolution ImageSynthesis with Latent Diffusion Models,”Apr.13,2022,arXiv:arXiv:2112.10752.
[0115]
[20] A.Nichol and P.Dhariwal,“Improved Denoising DiffusionProbabilistic Models,”Feb.18,2021,arXiv:arXiv:2102.09672.
[0116]
[21] O.Ronneberger,P.Fischer,and T.Brox,“U-Net:Convolutional Networksfor BiomedicalImage Segmentation,”arXiv:1505.04597[cs],May 2015.
[0117]
[22] “[]swin UNETR:Swin transformers for semantic segmentation ofbrain tumors in MRIimages,”arXiv:2201.01266.
[0118]
[23] Y.Pang,J.Lin,T.Qin,and Z.Chen,“Image-to-Image Translation:MethodsandApplications,”Jul.03,2021,arXiv:2101.08629.
[0119]
[24] “Ea-GANs:Edge-aware generative adversarial networks for cross-modality MR imagesynthesis|IEEEjournals&magazine|IEEE xplore.”.
[0120]
[25] J.Kim and H.Park,“Adaptive Latent Diffusion Model for 3D MedicalImage to ImageTranslation:Multi-Modal Magnetic Resonance Imaging Study”.
[0121]
[26] P.J.Moore,T.J.Lyons,J.Gallacher,and A.D.N.Initiative,“Randomforest prediction ofAlzheimer’s disease using pairwise selection from timeseries data,”PloS one,vol.14,no.2,p.e0211558,2019。
Claims
1. A multi-tracer PET cross-modal synthesis system based on diffusion transformer, characterized in that: First, we design a generative diffusion Transformer model, denoted as GenPET, which is used to synthesize multi-tracer PET images from structural MRI images. First, all modality images are mapped to a unified latent representation space, and the latent space diffusion Transformer model is used to synthesize MRI into PET of the specified tracer type, that is, GenPET is obtained; Specifically, GenPET consists of two stages: the first stage is unified latent space representation learning, and the second stage is cross-modal synthesis in the unified latent space; In the first stage, a three-dimensional encoder is used to encode MRI and multi-tracer PET images into a unified latent space, and a modality discrimination loss is designed to ensure the discrimination of different modalities in the latent space; In the second stage, the latent space diffusion Tansformer model is used to realize the latent space feature conversion from MRI to specific tracer PET.
2. The multi-tracer PET cross-modality synthesis system according to claim 1, characterized in that: The task of synthesizing 3D multi-tracer PET images from structural T1w MRI is formulated as a deep learning framework; given an MRI M ∈ R D×H×W , and the specific PET tracer type τ∈{Aβ,Tau,FDG} to be synthesized, the goal is to use the proposed GenPET(·) to synthesize PET images of a specific tracer, expressed as: Where D, H, and W represent the depth, height, and width of the three-dimensional image; Aβ, Tau, and FDG are designated tracer types for PET, which are used to designate imaging of AD-related biomarkers or neurological metabolic changes; synthetic PET images Compared with the real PET image P τ conform to; Specifically, GenPET(·) integrates a unified 3D VAE to encode MRI and multi-tracer PET images into a unified latent space. It also integrates a latent space diffusion Transformer network to convert MRI latent space representation into tracer-specific PET latent space features. 3D Unified VAE It consists of an encoder ε(·) and a decoder The unified encoder ε(·) adopts a multi-scale architecture with output channel sizes of 16, 64, 96, and 32; the unified decoder The input channel sizes are used in reverse order, i.e., 32, 96, 64, and 16 respectively; both the encoder and decoder contain two residual blocks per layer to ensure strong feature extraction; the channel dimension of the final intrinsic space is set to 32 to balance the complexity and compactness of the model; group normalization of eight groups is applied across layers; in addition, a multi-head self-attention mechanism is introduced in the deepest layers of the encoder and decoder to enhance the relationship between features at different spatial positions; In addition, a discriminator Dis(·) and a pre-trained 4-way 3D modality discriminative residual network (ResNet) are used. Participate in the training process of the 3D unified VAE network; the discriminator Dis(·) processes real or synthetic PET images through four convolutional layers and outputs corresponding blocks of size 19×23×19 for distinguishing true from false; in order to unify images of different modalities into a shared latent space and enhance the separation between their encoded latent representations, a pre-trained 4-way 3D modality discriminative residual network (ResNet) is used For the corresponding mode type τ m ∈{MRI,Aβ,Tau,FDG} classification; The latent space diffusion Transformer network uses the encoded MRI latent space features and PET tracer types as input to synthesize the corresponding PET latent space features. Specifically, the encoded MRI and intermediate variable features in the diffusion process are divided into patches with a spatial resolution size of 2×2×2; these patches are then connected and passed through N=13 latent space diffusion Transformer blocks, with the dimension of each token being 1536 to estimate the added noise in the diffusion process; the established attention mechanism uses 24 multi-heads for sub-dimensional feature calculation and fusion, and uses a fast attention mechanism to improve training and inference efficiency.
3. The multi-tracer PET cross-modal synthesis system according to claim 2, characterized in that: In the first stage, i.e., the unified latent space representation learning stage, 3D VAE is used to Learning latent space representations for MRI and multi-tracer PET in a unified latent space; 3D VAE It consists of an encoder ε(·), a decoder in: The encoder ε(·) is used to project MRI and multi-tracer PET images into a unified latent space of size c×d×h×w; the decoder Used to reconstruct the image back to the original space; here, c, d, h, and w represent the channel, depth, height, and width of the latent space representation, respectively; During the training phase, the neural image x and its corresponding modality type τ m ∈{MRI,Aβ,Tau,FDG} are randomly sampled; Following the standard latent diffusion training strategy, the proposed unified 3D VAE is trained using a combination of l1 reconstruction loss, perceptual loss, adversarial loss, KL divergence loss and a novel modality discrimination loss; The l1 reconstruction loss provides voxel-level supervision and is defined as follows: Real and synthetic PET images were unfolded along three 3D axes (axial, sagittal, and coronal), and the slice numbers of each axis were recorded as I, J, and K, respectively; The perceptual loss is calculated using the pre-trained 2D LPIPS metric LPIPS(·,·), which is based on a pre-trained 2D squeeze-and-excite network and is defined as: The adversarial loss is formulated in least squares form to distinguish between real and synthetic patches: The KL divergence loss constrains the unified latent space to approximate the standard normal distribution: Where d is the dimension of the latent space, μ i and σ i is the mean and standard deviation of the i-th latent space dimension; The modality discrimination loss is defined as follows: Among them, τ m is a real modal type, It is three-dimensional mode discrimination For reconstructed images The predicted modality type, i.e. The final objective function for optimizing the 3D unified VAE is defined as follows: Among them, λ per , GAN , KL and λ MDis is a hyperparameter that balances the weights of these five objective functions; Repeatedly optimize 3D unified VAE and the discriminator Dis(·), maximizing the objective in formula (4).
4. The multi-tracer PET cross-modal synthesis system according to claim 3, characterized in that: In the second stage, i.e., the cross-modal synthesis stage in the unified latent space, the pre-trained 3D unified encoder ε(·) and unified decoder (·) and frozen weights are used to further train the latent space diffusion Transformrer network. Thus, cross-modality synthesis from MRI to multi-tracer PET is achieved in a unified latent space; the entire diffusion process is expressed as: where time step t∈{1,2,…,T} and T=1000, ε(M)∈R c×d×h×w and ε(P τ )∈R c×d×h×w denote the latent space representation of the encoded MRI and the specific tracer PET, respectively; the t-th latent variable z1,…,z T With size c×d×h×w, the forward process and the reverse process are denoted by q(·) and p respectively. θ (·)express; In the forward diffusion process, following the standard forward equation, Gaussian noise is gradually is introduced into the encoded PET latent space representation z0 with any tracer type τ, thereby obtaining the corresponding t-th latent variable z t ∈R c ×d×h×w : in, β s is a predefined variance value; In the diffusion reverse process, the encoded MRI latent space feature ε(M) and the diffused t-time intermediate noise latent space feature z t Segmented into segments, connected and fed into the latent space diffusion Transformer network At the same time, the specified tracer type τ is injected through the AdaIN module to predict the Gaussian noise added at time step t, that is, The joint input of magnetic resonance imaging encoding and diffuse intermediate noise latent features realizes the seamless integration of self-attention and cross-attention mechanisms, allowing them to work independently while achieving dynamic interaction; in addition, linear interpolation, flash attention and efficient layer normalization strategies are adopted in skip connections to achieve efficient training and effective feature integration; here, by controlling the input tracer type, GenPET can synthesize multi-tracer PET; latent space diffusion Transformrer network Optimization using noise-based mean square error loss: For inference, the Gaussian noise Sampling is performed and connected with the encoded MRI latent space feature ε(M); given an arbitrary synthetic tracer type τ, the denoised diffusion latent model (DDIM) is used as a diffusion sampler to iteratively synthesize the t-th latent space variable z with 50 time steps t , until the initial number of time steps is reached; finally, the unified decoder is used The initial latent variable for the synthesis Decode and synthesize the corresponding PET image P τ ,Right now
Citation Information
Cited By
Virtual organ simulation method and system based on three-dimensional stacking space transcriptome data
CN120977370A
A three-dimensional stacked space transcriptome data-based virtual organ simulation method and system
CN120977370B