A system and method for synthesis of pet images from multimodal mr images
The synthesis of PET images from multimodal MR images using GANs addresses the limitations of PET scanning by providing a cost-effective and non-invasive alternative for diagnostic imaging, enhancing accessibility and reducing radiation exposure.
Patent Information
- Application Number
- PCT/IB2024/058081
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-12
- Filing Date
- 2024-08-20
- Publication Date
- 2025-12-18
AI Technical Summary
PET scanning faces challenges such as high cost, radiation exposure, invasiveness, limited availability, and low spatial resolution, making it less accessible and feasible for widespread use in medical imaging.
A system and method using multimodal MR images and Generative Adversarial Networks (GANs) to synthesize PET images, leveraging existing MR data to create synthetic PET images that mimic real PET scans, reducing the need for actual PET scans.
Enables cost-effective, non-invasive, and accessible PET imaging by generating diagnostic-quality PET images from MR scans, thereby reducing unnecessary PET scans and enhancing patient care.
Smart Images

Figure IB2024058081_18122025_PF_FP_ABST
Abstract
Description
[0001] TITLE: A SYSTEM AND METHOD FOR THE SYNTHESIS OF PET IMAGES FROM MULTIMODAL MR IMAGES
[0002] DESCRIPTION
[0003] FIELD OF INVENTION
[0004] The present invention belongs to the field of neuroimaging. More particularly, it relates to a method of obtaining synthesized PET images from a single or a plurality of multimodal MRI images.
[0005] BACKGROUND
[0006] Medical imaging plays a pivotal role in diagnosing neurological diseases, offering intricate visualizations of brain structures and functions essential for accurate diagnosis. Techniques such as Magnetic Resonance Imaging (MRI) and Positron Emission Tomography (PET) scanning are particularly useful in detecting pathological depositions such as those seen in diseases like Alzheimer's disease (AD), facilitating early intervention and more effective disease management. While both MRI and PET imaging are valuable, MRI is more frequently recommended by healthcare practitioners due to its wider availability, feasibility, non-invasive nature, and comparatively lower cost. Due to the cost, limited availability, and invasiveness of PET scanning, it is essential to thoroughly evaluate their necessity to prevent unnecessary medical expenses, procedures, and inconvenience for patients and the healthcare industry. In recent years, artificial intelligence (Al) has emerged as a promising tool in medical imaging, assisting clinicians in time-consuming, resource intensive tasks such as pathology detection, disease diagnosis and disease progression prediction. Recent advances in the field of generative Al have made the generation of synthetic medical images possible. Harnessing the capabilities of Al could therefore not only streamline the utilization of PET scanning, but can also make this mode of medical imaging available to areas devoid of PET facilities, reduce costs and risks associated with PET, therefore optimizing its use in enhancing patient care.
[0007] In the field of medical imaging, scans are broadly classified into structural imaging and functional imaging scans. Structural imaging focuses on visualizing anatomical details and morphology, employing techniques such as T1 -weighted imaging, computed tomography (CT), and diffusion-weighted imaging to provide images with high spatial resolution of tissues and organs. In contrast, functional imaging is used to observe and measure the activity or function of brain tissue, with methods like functional MRI (fMRI), PET, and Single-photon Emission Computed Tomography (SPECT) capturing dynamic processes and metabolic changes. Functional imaging techniques provide good temporal resolution i.e. a measure of activity with time, but tend to have relatively lower spatial resolution. While structural imaging reveals the physical structure and location of abnormalities, functional imaging excels in identifying areas associated with specific tasks, studying brain activity, or detecting conditions with altered metabolic or haemodynamic function. In research and clinical applications, the combination of both types of imaging techniques offers a comprehensive understanding of both the structure and function of the human body.
[0008] PET and MRI scanning is recommended by the healthcare providers to check for signs of health conditions such as cancer including breast cancer, lung cancer and thyroid cancer or neurological disorders such as tumours, epilepsy and neurodegenerative diseases such as Alzheimer’s disease (AD). AD is currently one of the leading causes of dementia worldwide, a disease primarily linked with modem day lifestyle and increasing longevity. Given that the onset is closely related to higher age, the disease burden is expected to increase with the aging population and increased life expectancy. AD is a progressive neurodegenerative disorder that is pathologically characterized by a build-up of misfolded proteins namely, extracellular amyloid plaques and intracellular neurofibrillary tangles (NFT). It manifests clinically with a significant decline in cognitive functions, with the most prominent deficits exhibited in the memory domain that worsen with disease progression. It is characterized by progressive global and regional brain atrophy that is most pronounced in the medial temporal lobes (MTL). This atrophy has largely been attributed to the neuropathological deposits observed in AD. A definitive AD diagnosis can only be made following histopathological confirmation of AD pathology, postmortem.
[0009] The presence of AD pathology can be confirmed in-vivo using PET imaging or cerebrospinal fluid markers. Neuropathological hallmarks of AD namely, amyloid plaques and NFTs can be detected using specific PET tracers. Positron Emission Tomography, commonly known as PET, is a functional medical imaging technique that allows physicians and researchers to observe metabolic and biochemical processes in the body. It is particularly valuable in the fields of oncology, neurology, and cardiology. For instance, Fluorodeoxyglucose (FDG) PET is useful in cancer diagnosis and staging. Since cancer cells often have higher metabolic rates than normal cells they have a higher binding affinity to the injected radiotracer. This high metabolic activity appears as "hot spots" on the PET images, which aids in identifying and locating tumours. Similarly, in AD, injection of radiotracers that bind to amyloid and tau can highlight neuropathological depositions typical of AD in the brain.
[0010] Due to their low spatial resolution, PET scans are often used in conjunction with other imaging techniques, such as CT (Computed Tomography) or MRI (Magnetic Resonance Imaging), to provide a more comprehensive view of the structures and functions within the body. The combined use of PET-CT scans, for example, allows for both anatomical and metabolic information to be obtained simultaneously. In order to measure brain amyloid load using PET, radioactive substances such as Amyloid PET tracers such as Florbetapir (Amyvid), Florbetaben (Neuraceq) or Flutemetamol (Vizamyl) are used to detect and visualize amyloid plaques, i.e. abnormal clumps of proteins, specifically beta-amyloid, that accumulate in the brains of individuals with certain neurodegenerative conditions. PET imaging with amyloid tracers allows clinicians and researchers to observe the presence and distribution of amyloid plaques in the brain. In Tau-PET imaging, tracers that detect tau pathology include but are not limited to Flortaucipir and THK tracers. These tracers are used to visualize and quantify tau pathology in the brain; wherein the said tracers are radio labelled compounds that bind specifically to tau protein aggregates, allowing for the detection of abnormal tau deposition in neurodegenerative diseases.
[0011] In comparison to PET, which visualizes metabolic activity, MRI excels in providing detailed anatomical information, offering complementary perspectives for comprehensive diagnostic evaluation and treatment monitoring. MRI, or Magnetic Resonance Imaging, is a sophisticated medical imaging technique that utilizes strong magnetic fields to generate detailed images of the internal structures of the body. Unlike traditional X-rays or CT scans, MRI scanning does not use ionizing radiation, making it safer for frequent use. It provides high-resolution images of soft tissues, such as the brain, muscles, and organs, offering invaluable insights into various medical conditions, including tumours, injuries, and neurological disorders. By exploiting the magnetic properties of hydrogen atoms in water molecules within the body, MRI creates cross- sectional images that aid in diagnosis and treatment planning with remarkable clarity and precision. Although PET scanning techniques are powerful, favoured and widely used in medical imaging, it has some drawbacks and limitations. Some of the main drawbacks associated with PET scanning include: Radiation Exposure: PET scanning involves the use of radioactive tracers, which emit positrons. This radiation exposure is minimal, but still carries certain risks, especially for pregnant women and young children. It is recommended by physicians to image pathologies in the brain as the benefits of PET scanning outweigh the risks.
[0012] Limited Half-life of Radiotracers: Radiotracers used in PET scanning have relatively short halflives, ranging from minutes to a couple of hours. This limited half-life imposes time constraints on the scanning process. This time sensitivity can be logistically challenging, as any delays in the process of transporting the radiotracers from the production site to administration can compromise the effectiveness of the PET scan. As a result, the scheduling and coordination of radiotracer production, distribution, and patient scanning must be meticulously managed to maximize the utility of the radiotracer and obtain accurate diagnostic information.
[0013] Invasiveness: PET scanning is an invasive technique as it uses radioactive tracers which may be administered intravenously.
[0014] Cost and Tracer Synthesis: PET scanning involves the use of radioactive tracers that need to be synthesized in a cyclotron, a specialized and expensive particle accelerator. The production of these tracers requires skilled personnel, dedicated facilities, and sophisticated equipment, contributing significantly to the overall cost of PET scanning. The expenses associated with tracer production can limit the availability of a wide range of radiotracers, potentially restricting the types of studies that can be performed using PET scanning. This aspect adds an extra layer of complexity to the cost considerations associated with PET scanning.
[0015] Spatial Resolution: Compared to some other medical imaging techniques like MRI, PET has relatively lower spatial resolution, thereby making it challenging to precisely localize small abnormalities or lesions within the body. Although PET-MR seaming has become available in recent years, it is far more expensive than acquiring a PET-CT scan, which is why PET-CT is the recommended imaging modality in practice.
[0016] Limited Availability: Accessibility to PET scanners is limited to certain regions; which can cause delays in obtaining PET scans, especially in critical situations.
[0017] Tracer Availability: The production and availability of specific radiotracers can be a limiting factor in some cases. Developing new tracers is a complex process, and not all diseases or conditions have well-established tracers for PET scanning. Patient Preparation: Patients undergoing PET scanning often need to fast for several hours before the scan, or may need to discontinue certain medications thereby causing an inconvenience to patients.
[0018] Limited Soft Tissue Contrast: While PET seaming is excellent for capturing metabolic processes, it may not provide as much contrast for certain soft tissues as other imaging modalities like MRI.
[0019] Regardless of the numerous disadvantages of PET scanning, it has great clinical utility and cannot currently be replaced by a substitute method. The present invention therefore attempts to address some of the challenges faced in PET scanning.
[0020] Deep learning has revolutionized medical imaging by significantly enhancing the accuracy and efficiency of diagnostic processes. The use of various Convolutional Neural Networks and Generative Adversarial Networks (GANs), has enabled the development of advanced image recognition systems that can detect and classify diseases with high precision. These networks are capable of analysing complex medical images, such as X-rays, MRIs, and CT scans, often surpassing human experts in speed and accuracy. Furthermore, GANs are particularly useful in generating high-quality synthetic medical images, thus improving the robustness of medical imaging models. Moreover, transfer learning, another key technique in deep learning, has expedited model training by leveraging pre-trained networks on large-scale datasets, such as ImageNet, and fine-tuning them for medical imaging tasks. By transferring knowledge from more general domains to medical imaging, this technique has helped address the challenge of limited availability of medical imaging data for model training. Transfer learning has accelerated the development of robust and effective medical imaging models using smaller datasets, enabling more precise and timely diagnoses across various medical specialties. The synergy between CNNs, GANs and transfer learning has propelled medical imaging into a new era of innovation, paving the way for enhanced patient care and outcomes. The current invention aims to draw on this synergy to enable the synthesis of PET images using multimodal MR images, a more cost-effective and non-invasive mode of medical imaging in order to reduce the burden of challenges faced in PET imaging.
[0021] KR20180097214A relates to a method and apparatus for predicting a recurrence site and generating PET images using a single or plurality of multimodal MRI including a T1 image, a TICE image, a T2 image, a FLAIR image, and an ADC image; but it fails to specify what type of PET paterns will be predicted. Additionally, this method uses a convolutional neural network to generate PET images whereas the proposed method uses GANs to synthesize PET images.
[0022] KR102456338B1 discloses a device for predicting the positive rate of PET comprising a data acquisition part that acquires MR images and patient information of a patient; and a processor that predicts the positive rate of PET of the patient based on the acquired MR images namely, Tl-weighted and FLAIR MR images. This device calculates PET positivity rate of amyloid by segmenting the brain into specific brain regions whereas the proposed device aims to synthesize a PET image that could be used to calculate the amyloid PET positivity rate. Additionally, the proposed invention aims to approximate tau and amyloid accumulation, which provides a more comprehensive view of AD pathology in contrast to this device which only calculates amyloid positivity rate. Furthermore, this device does not use GANs.
[0023] WO2023153674A1 discloses an MRI-based method for predicting a degree of accumulation of a biomarker associated with Alzheimer's disease and a companion diagnosis method using same; wherein the method is used to predict biomarker accumulation value of amyloid and tau for each region of the target brain. The method uses both MRI and PET images to calculate biomarker accumulation whereas the present invention aims to synthesize a PET image for one or more MR images.
[0024] None of the prior arts disclose a method that provides tangible PET images synthesized using the multimodal MR images; such that it overcomes the drawbacks of the available prior arts. Therefore, a substantive system and method to synthesize PET images is required, thereby enabling the medical practitioners to decide if a PET scan is required.
[0025] DEFINITIONS
[0026] The expression “system” used hereinafter in this specification refers to an ecosystem comprising a user, input / output device, input, processing unit and output.
[0027] The expression “user” used hereinafter in this specification refers to, but is not limited to any individual, natural entity or a medical practitioner who initiates any step, procedure, workflow by providing input; or is a recipient of the tangible output. The expressions “input devices” or “output devices” used hereinafter in this specification refers to, but is not limited to a computer, mobile device, iPad, screen, scanners, pen drives, keyboards or mouse; that allows the user either to provide or receive instructions into a computer readable format or visual interpretation.
[0028] The expression “input” used hereinafter in this specification refers to, but is not limited to documents, image, set of images, MR images, PET images in a computer readable / acceptable format such as DICOM (Digital Imaging and Communications in Medicine), NIFTI, BIDS etc.
[0029] The expression “output” used hereinafter in this specification refers to, but is not limited to images, sets of images, synthetic PET images in a computer readable / acceptable format such as DICOM (Digital Imaging and Communications in Medicine).
[0030] The expression “processing unit” used hereinafter in this specification includes, but is not limited to an application, module, platform, website or webpage accessible through a cloud network or internet, embedded in a system; such that it follows a specific workflow to provide a final output by processing a set of data.
[0031] The expression “multimodal” used hereinafter in this specification refers to the utilization of multiple modalities of medical imaging to gather complementary information about the structure and function of the brain or other anatomical regions; where the said techniques can include, but are not limited to structural MRI, functional MRI (fMRI), diffusion MRI (dMRI), magnetic resonance spectroscopy (MRS), and others; whereby combining different modalities, researchers and clinicians can gain a more comprehensive understanding of brain anatomy, connectivity, and function, leading to improved diagnosis and treatment planning for various neurological conditions.
[0032] The expression “MR Images” or “MRI” used hereinafter in this specification refers to Magnetic Resonance Images, where the detailed visual representations of internal body structures are generated using a powerful magnetic field and radio waves. These images provide valuable structural and functional information about soft tissues, organs, and bones within the body without the use of ionizing radiation. MR images are created by measuring the response of hydrogen atoms in the body to the magnetic field, producing high-resolution cross-sectional images that can be used for medical diagnosis, treatment planning, and research purposes. The expression “PET images” used hereinafter in this specification is Positron Emission Tomography and refers to three-dimensional representations of the distribution of positronemitting radiopharmaceuticals within the body. PET imaging involves injecting a radioactive tracer into the body, which emits positrons. When these positrons encounter electrons in the body, they annihilate, releasing gamma rays. Detectors in the PET scanner surrounding the body then capture these gamma rays, allowing the reconstruction of images that depict the concentration of the radiotracer in various tissues and organs. PET images provide useful information such as metabolic activity, tissue function, and molecular processes, aiding in the diagnosis and management of various diseases, including cancer, neurological disorders, and cardiovascular conditions.
[0033] The expression “GANs” used hereinafter in this specification refer to Generative Adversarial Networks (GANs), which is a type of deep learning architecture that plays a pivotal role in generating synthetic medical images that closely resemble real patient data. GANs can create diverse and realistic images, addressing the scarcity of labelled medical data for training purposes.
[0034] The expression “CNNs” used hereinafter in this specification refers to convolutional neural networks (CNNs), which is a deep learning neural network designed for processing structured arrays of data such as images. Convolutional neural networks are widely used in computer vision and have become the state of the art for many visual applications such as image classification, and have also found success in natural language processing for text classification.
[0035] OBJECTS OF THE INVENTION
[0036] A primary object of the present invention is to provide a system and method for the synthesis of PET images from multimodal MR images.
[0037] Yet another object of the present invention is to provide a method that uses previously existing MRIs data from the patients to create synthetic PET images that could have diagnostic utility.
[0038] Yet another object of the present invention is to provide a cost-effective, accessible and non- invasive method of observing PET processes. Yet another object of the present invention is to allow the user to determine the need for PET scans prior to the actual scan, thereby reducing the need for unnecessary PET scanning and providing a method to determine the cases where PET scanning is required and potentially increasing valid use of the same.
[0039] SUMMARY OF THE INVENTION:
[0040] Before the present invention is described, it is to be understood that present invention is not limited to particular methodologies and materials described, as these may vary as per the person skilled in the art. It is also to be understood that the terminology used in the description is for the purpose of describing the particular embodiments only, and is not intended to limit the scope of the present invention.
[0041] The present invention describes a system and method for synthesizing PET images from multimodal MR images. The system comprises a user, input, input / output device(s), a processing unit and an output. The processing unit uses multimodal MR images, sent as input by the user in a computer acceptable format such as Digital Imaging and Communications in Medicine (DICOM) format, that passes through a trained deep convolutional generative adversarial network (DCGAN) workflow in order to provide an output in the form of synthesized PET image(s).
[0042] According to an aspect of the present invention, the method for synthesizing PET images from multimodal MR images is proposed across three stages namely model training, validation and deployment. The model training stage includes creating a dataset, determining the dataset size, preprocessing of PET and MR images in the dataset, train-test split comprising of a training set for model training that uses a DCGAN and a test set to assess the capability of the model to generalize to unseen data in order to produce synthetic PET images from multimodal MRI data. Validation is carried out on the test set that includes application of image similarity matrices such as Mean Squared Error (MSE), Structural Similarity Index (SSIM) or Peak Signal to Noise Ratio (PSNR) to evaluate image similarity between real and synthetic PET images, assessing image quality and comparing extracted Standardized Uptake Value Ratio (SUVr) scores between real and synthetic PET images. Finally, model deployment is carried out to obtain synthesized PET images from new multimodal MR images by loading the trained DCGAN model onto the system memory and using it with necessary infrastructure like servers, cloud services, or edge devices.
[0043] BRIEF DESCRIPTION OF THE DRAWINGS
[0044] A complete understanding of the present invention may be inferred from the following detailed description which is to be taken in conjugation with the accompanying drawings. The accompanying drawings, which are incorporated into and constitute a part of the specification, illustrate one or more embodiments of the present invention and, together with the detailed description, they serve to explain the principles and implementations of the invention.
[0045] Fig. 1 depicts the Model training workflow that illustrates the steps followed by the processing unit in the model training phase
[0046] Fig. 2. depicts the Model deployment workflow that illustrates the steps followed by the processing unit in the model deployment phase and
[0047] Fig. 3. depicts the system and methodology flowchart that illustrates the method followed by the processing unit of the present system.
[0048] DETAILED DESCRIPTION
[0049] Before the present invention is described, it is to be understood that this invention is not limited to particular methodologies described as these may vary with the person skilled in the art. It is also to be understood that the terminology used in the description is for the purpose of describing the particular embodiments only, and is not intended to limit the scope of the present invention. Throughout this specification, the word “comprise”, or variations such as “comprises” or “comprising”, will be understood to imply the inclusion of a stated element, integer or step, or group of elements, integers or steps, but not the exclusion of any other element, integer or step, or group of elements, integers or steps. The use of the expression “at least” or “at least one” suggests the use of one or more elements or ingredients or quantities, as the use may be in the embodiment of the invention to achieve one or more of the desired objects or results. Various embodiments of the present invention are described below. It is, however, noted that the present invention is not limited to these embodiments, but rather the intention is that the modifications that are apparent are also included.
[0050] In order to overcome the challenges posed by PET scanning such as high cost, radiation exposure to patients and feasibility and availability of PET imaging, the present invention proposes a technique that could help identify individuals who need to undergo PET scanning using multimodal MR images, a more cost-effective, accessible and non-invasive method of imaging compared to PET. Additionally, MRI is collected as part of the routine clinical procedure for patients with suspected diseases like AD. Therefore, this invention could provide additional information from previously existing data from the patients, that uses MR images to generate PET images thereby specifying the type of predicted PET patterns without the requirement of collecting additional patient information.
[0051] The present invention relates to a method for synthesis of a Positron Emission Tomography (PET) image from multimodal Magnetic Resonance (MR) images (2). The three stages in the said method comprise the model training stage, validation stage and model deployment stage. In the model training stage, there are several steps that include creation of the dataset, determining the dataset size, preprocessing PET and MR images (2), train-test split, model training using a DCGAN workflow (1) and synthesizing a PET image. This method is depicted in Figure 1 : The model training workflow. The validation stage is performed to assess similarity between synthesized and real PET images (16) and clinical utility of synthesized PET images. Steps in the validation stage include image similarity matrices, image quality assessment and comparison of SUVr scores between the synthetic (15) and real PET images (16). In the final stage or the model deployment stage, the trained model is packaged to be used by medical professionals to obtain synthesized PET images from a single or plurality of multimodal MR images (2). In the deployment stage, steps include setting up the environment using a container such as Docker, loading the trained generator, preprocessing the input MR images (2), processing the input (22) images by passing them through the loaded trained DCGAN and generating synthetic PET images (15). The detailed steps are as follows:
[0052] A. Model Training: a. Creation of dataset:
[0053] In the model training stage, data is provided as input (22) in the form of pairs of images comprising a single or plurality of multimodal MR images (2) including, but not limited to T1 -weighted MRI, T2 Fluid Attenuated Inversion Recovery (FLAIR), Arterial Spin Labelling, Diffusion-Weighted Imaging (DWI), functional MRI; and a corresponding PET image (16) including, but not limited to images acquired after injecting radiotracers that highlight specific biological markers or processes, such as those confirming disease pathology for each sample (patient). The said images are read from a suitable format such as DICOM or Nifti to train the model. The datasets include samples of patients at various stages of disease, therefore having varied levels of pathological burden, as well as of cognitively normal individuals, confirmed by an experienced neurologist, which collectively provides the model training stage with a range of cases having varying levels of the disease pathology; thereby enabling the model to synthesize more realistic PET images from unseen MRI scans. b. Determining the dataset size:
[0054] Any method with model training specifications requires substantial amounts of data to perform better. However, the exact amount of data required can vary significantly in order to be sufficient to train generative models depending on the demand and applicability of the method so as to provide an accurate result. Dataset size may vary between a range of 1000 - 2000; preferably 1000 image sets; each set including a pair of MR image(s) and the corresponding PET image of the same sample (patient). c. Preprocessing PET and MR images:
[0055] The PET images (16) are pre-processed before they are provided as input (22) to model training in order to ensure spatial alignment of brain structures with the MR image (s) (2) and to improve signal -to-noise ratio of PET images. The various steps involved in preprocessing the PET image may include but are not limited to skull stripping, image registration, image reorientation, attenuation correction, motion correction, image reconstruction, spatial smoothing, normalization, intensity normalization, correction for partial volume effects.
[0056] The MR images (2) may be pre-processed to remove any other components in the MR image that may hamper model training. The various steps involved in preprocessing the MR image may include but are not limited to image reorientation, skull stripping, probabilistic tissue segmentation, normalization, spatial smoothing. d. Train-test split:
[0057] The dataset determined in the previous step is split into training and test sets in a ratio of 90: 10 respectively; such that the training set is used to train the model and the test set is used to assess the capability of the model to generalize to unseen data and to assess the clinical utility of the synthesized PET image. The train -test split thus employs the sub-steps as follows: 1. Training Sets:
[0058] In the train-test split sub-step, 90 percent of the dataset from the determined dataset size is considered as the training set for model training using a Deep Convolutional Generative Adversarial Network (DCGAN) and transfer learning; thereby allowing the system (20) to pass new MRI data through the trained model to synthesize PET images.
[0059] 1.1. Model Training using DCGANs:
[0060] In this sub-step the said determined datasets of images are used to train the model to generate PET images using pairs of pre-processed multimodal MR images (2) and PET images of the sample; using Deep Convolutional Generative Adversarial Networks (DCGANs) (1); as referred to in Fig. 1. The present invention may use Generative Adversarial Networks (GANs) including, but not limited to Bidirectional Mapping GANs (BMGAN) and Brain-PET GANs (BPGAN), Conditional GANs (cGANs) or Cycle GANs; that are used for generating synthetic, realistic images; preferred in medical imaging for synthesizing high-quality images from available data (input) such as MR images (2) to generate PET images while preserving the detailed brain structures. The GANs operate through two main components: a generator and a discriminator; where the generator creates images aimed to be indistinguishable from real images, while the discriminator evaluates them against real data, refining the generator's output iteratively such that this adversarial process enhances the quality and realism of synthetic images.
[0061] 1.1.1 Generator Network:
[0062] The DCGAN in the current method works by first passing the training data through a generator network (3) employing a U-Net architecture to synthesize realistic PET images using an encoder-decoder path. The generator network (3) works by applying a series of convolutional and pooling layers to the input MR images (2), wherein the convolutional layers (4, 6) extract feature maps and downsample (5, 7) the input (22) by sliding a small filter referred to as a pooling window over the image and estimating the dot product between the filter and the input (22) to form pooled feature maps. The same step is repeated multiple times to form pooling layers. The pooling layers then downsample (5,7) the output (25) of the convolutional layers (4,6) to reduce the dimensionality of the data. Each convolutional layer is typically followed by batch normalization and a rectified linear unit (ReLU) activation function to introduce non- linearity and aid in feature extraction. Max pooling layers are often used to further downsample (5,7) the feature maps, reducing their spatial resolution while increasing their depth. As the input (22) image passes through successive convolutional layers (4,6) and downsampling (5,7) operations, the encoder path progressively extracts hierarchical features. At the end of the encoder path, the feature maps typically have reduced spatial dimensions but increased depth. These feature maps contain semantically rich information about the input (22) image, capturing both low-level and high-level features relevant to the task at hand. The output (25) of the encoder path serves as the input (22) to the bridge (8), connecting the encoder and decoder paths in the U-Net architecture.
[0063] At the bottom of the U-shaped architecture, a bridge (8) connects the encoder and decoder paths. The bridge (8) typically consists of one or more convolutional layers (4,6) that retain the learned features while transitioning from the encoding to the decoding stage. Skip connections (13, 14), also known as shortcut connections connect the encoder and decoder at each corresponding layer to promote information flow and preserve high-resolution features facilitating the propagation of information from earlier layers to later layers.
[0064] The decoder path begins with upsampling operations to gradually restore the spatial dimensions of the feature maps. Each upsampling step (9, 11) is usually followed by a concatenation operation, where feature maps from the encoder path are combined with those from the decoder path. Convolutional layers (10, 12) with decreasing filter sizes are applied to refine the features and generate the final output (25). Batch normalization and ReLU activation functions are applied after each convolutional layer (10, 12) to maintain non-linearity and stabilize training. The final layer of the decoder path typically consists of a convolutional layer (12) with a sigmoid activation function, which outputs (25) a generated image (15). These generated images are meant to mimic real images from the dataset and capture their visual characteristics.
[0065] Transfer learning may be incorporated in the model training stage using pre-trained networks such as Resnet-50, VGG-19; such that the pre-leamed features are extracted from previous models by freezing the weights of the earlier layers in the pretrained networks and only training the final few layers in order to achieve the specific task. Transfer learning is thus incorporated in the earlier layers of the generator network (3) that enables the use of learned general features such as edges and textures. In the case of image models this provides a good starting point, while the later layers adapt to the specific features of the new task. In this case, generating synthetic PET images (15) from multimodal MR images (2).
[0066] 1.1.2 Discriminator Network:
[0067] Batches of synthetic (15) and real (16) PET images are then passed through the discriminator network (17) in order to train the discriminator network (17) and to compute generator loss (19) and discriminator loss (18), respectively. The generator loss (19) and discriminator loss (18) are computed with generator and discriminator output and these two losses are used to train the model. Aggregate losses from real and generated images are used to update the generator and discriminator weights for model training . The discriminator network (17) typically consists of convolutional layers followed by dense layers, culminating in a binary classification output. The discriminator evaluates each synthesized image and predicts whether it is synthesized or real based on its learned criteria. Its architecture typically comprises convolutional layers with Leaky ReLU activations to introduce non-linearity and prevent saturation. Notably, feature maps from the encoder path of the generator network (3) are concatenated with corresponding upsampled feature maps in the decoder path, aiding discrimination. This fusion of information from different scales contributes to improved image quality and sharper details.
[0068] 1.1.3 Training Loop:
[0069] The training loop encompasses the iterative process of training both the generator (3) and discriminator (17) networks. MR images (2)serve as input (22) to the generator network (3) to generate synthetic PET images (15). The discriminator network (17) evaluates batches of real and generated images, computing losses accordingly. The total discriminator loss (18) is computed as the sum of losses from real (16) and synthetic (15) images, and backpropagation is performed to update discriminator weights. The generator's loss is then computed based on the discriminator's feedback, and its weights are updated accordingly. This process continues for a predefined number of epochs or until the generator synthesizes PET images of satisfactory quality.
[0070] 2. Test Set: Ten percent of the dataset from the determined dataset size is considered as the test set which is used to assess the capability of the model to use unseen MRI data in order to produce synthetic PET images (15) with high similarity to the real PET images (16).
[0071] B. Validation:
[0072] Followed by the model training using DCGANs, a validation step is performed to ensure that the model is well-generalized to synthesize high quality PET images by applying the trained DCGAN workflow (1) to the test set in order to synthesize PET images. This stage is primarily implemented to assess the similarity between the real and synthetic PET images (15). The validation is performed using the test set by applying a 3 -fold process including application of image similarity matrices, image quality assessment and comparison of Standard Uptake Value ratios (SUVr). It is to be noted that the three modes of comparing image similarity between the real and synthetic images can be performed parallelly as these methods are independent of each other.
[0073] 1. The image similarity matrices applied herein to compare the similarity between real and synthetic PET images (15) include but are not limited to: a. Mean Squared Error (MSE): MSE measures the average squared difference between corresponding pixel intensities of the real and synthetic PET images (15). Lower MSE or Lower root mean square error (RMSE) values indicate higher similarity where the RMSE values close to 0 indicate perfect similarity. RMSE values depend on the scale of the pixel values. b. Structural Similarity Index (SSIM): SSIM compares the structural information of the real and synthetic PET images (15) thereby considering the luminance, contrast, and structure similarity between the two images. SSIM values range from -1 to 1, where 1 indicates perfect similarity and the values above 0.9 are i.e. values closer to 1 are considered to represent higher similarity and as an acceptable measure of SSIM. c. Peak Signal-to-Noise Ratio (PSNR): PSNR measures the ratio between the maximum possible power of a signal and the power of corrupting noise, used in image processing to measure the quality of a compressed image. Higher PSNR values indicate higher similarity whereas values closer to infinity indicate perfect similarity; where the PSNR values with an acceptable threshold may be variable, preferably above 30 dB represents good quality.
[0074] 2. Image quality assessment:
[0075] Image quality assessment is conducted through a single blind study where radiologists visually inspect the images and rate the quality of both, real and synthetic PET images (15). The images are displayed to the radiologist in a random order, thereby allowing image quality assessment that aids in evaluating the clinical utility of the synthetic PET images (15).
[0076] 3. SUVr scores:
[0077] As another form of validation, a quantitative measure, Standardized Uptake Value Ratio (SUVr) scores from the PET images are calculated to assess the accumulation of radiotracer uptake in specific regions of the brain or body to estimate the level of target biological process / pathology accumulation; where SUVr compares the concentration of the tracer in a target region to a reference region, typically a region with minimal specific binding of the tracer. By normalizing uptake values in this manner, SUVr accounts for individual differences in tracer metabolism and injection dose, providing a standardized measure of tracer retention. For the validation of the invention, SUVr scores are calculated for the real PET image (16) and the synthetic PET image (15) in order to assess the clinical utility of the synthetic PET images (15).
[0078] It is to be noted that the SUVr scores are particularly valuable in neuroimaging studies of various diseases, where they help quantify the extent of pathological accumulations, aiding in disease diagnosis, monitoring progression, and evaluating treatment efficacy. A 10% difference in SUVr scores is considered as acceptable i.e. an assumption is made that 90% of the radiotracer uptake is captured in the synthetic PET image (15) compared to the real one.
[0079] C. Model Deployment:
[0080] Once the model has been trained and validated and the processing unit (24) is ready, when a user (21) inputs new MRI data (22) to the processing unit (24) in the system (20) through input device (23), it enables the workflow for the synthesis of PET images (25) from multimodal MR images (22). The model is deployed by first setting up the environment which includes but is not limited to setting up the necessary infrastructure, such as servers, cloud services, or edge devices. The trained model, along with its dependencies (libraries, configuration files, etc.), is packaged for deployment. This may be done using containers (like Docker) to ensure consistency across different environments. It is to be noted that the trained model or the processing unit (24) excludes the discriminator network (17) as the discriminator network (17) is used to train the generator network (3) in the model training stage. The trained model is then loaded onto the system memory. Depending on the deployment environment, this could involve loading from a file system, database, or cloud storage. The next step includes creating perceptually realistic synthetic PET images (15) while preserving detailed brain structures. The processing unit (24) of the system (20) uses the multimodal MR images (2), sent as input (22) by the user (21) in a computer acceptable format such as Digital Imaging and Communications in Medicine (DICOM) format. The DICOM files are transformed into a format suitable for the generator network (3). This might involve converting them to a specific image size and ensuring they are within a specific intensity range. No additional processing related to adversarial training (like the discriminator (17)) occurs here. The preprocessed MR images are fed into the trained generator network (3) or processing unit (24). The trained generator network (3), likely consisting of multiple convolutional layers (4,6), extracts relevant features from the MR images (2). These features are then processed by the core network to create a new internal representation suitable for generating synthetic PET images (15). Upsampling or decoding layers progressively increase the detail of the synthetic PET image (15) based on the learned patterns from the training stage. The trained generator network (3) then produces a synthetic image on the output device (23) that may resemble a realistic PET image from the MR image(s) input (22) while potentially containing new information or variations based on the training data.
[0081] Working example:
[0082] “Synthesis of PET images for Alzheimer’s disease pathology from multimodal MR images”
[0083] It is convenient to understand the specific features of the invention through a working example described herewith in a stepwise manner. An example disclosing the method of synthesizing PET images from multimodal MR images is provided wherein the workflow / steps are given hereinafter. It is to be noted that the example is to be considered for the purpose of understanding the method, and is not limited to the synthesis of PET images concerned with a specific disease or biological process.
[0084] The method includes the steps of:
[0085] A. Model training: a. Creation of a dataset:
[0086] In case of Alzheimer’s Disease (AD), a set of MR and PET images of the same sample (patient) having AD or cognitively normal individuals are provided in DICOM formats as input through at least one input device wherein the radiotracers specific to amyloid and tau proteins are injected intravenously so that for each PET image included in the dataset, an experienced neurologist has confirmed the presence or absence of AD pathology i.e. Amyloid plaques and neurofibrillary tangles. This dataset includes patients with mild to moderate AD and cognitively normal individuals (confirmed by an experienced neurologist), which provides the model training stage with a range of cases that have varying levels of AD pathological depositions (amyloid plaques and NFTs); thereby enabling the model to synthesize PET images from unseen MR images.. b. Determining the dataset size:
[0087] 1000 image sets; each set including a pair of multimodal MR and PET image(s) of the same sample (patient); where the neurologist has confirmed the presence or absence of AD pathology i.e. Amyloid plaques and neurofibrillary tangles are provided as input to the model training stage. c. Image preprocessing:
[0088] The PET and MR images will be pre-processed before they are provided as input to the model training stage in order to ensure spatial alignment of brain structures with the MR images and to improve the signal to noise ratio of PET images. The various steps involved in preprocessing the PET image may include but are not limited to skull stripping, image reorientation, image registration, attenuation correction, motion correction, image reconstruction, spatial smoothing, normalization, intensity normalization, correction for partial volume effects. The MR images may be pre-processed to remove any other components in the MR image that may hamper model training. The various steps involved in preprocessing the MR image may include but are not limited to image reorientation, skull stripping, probabilistic tissue segmentation, normalization, spatial smoothing. d. Train test split:
[0089] The determined dataset of the previous step is split into two sets as follows:
[0090] 1. The training set consists of 900 sets of images; each set including a pair of preprocessed MR image(s) and corresponding preprocessed PET images provided for the model training using the Deep Convolutional General Adversarial Network (DCGAN) workflow described in FIG. 3, Model Training to synthesize Brain Amyloid and Tau PET images from MR images.
[0091] 2. The test set comprises 100 sets of images; each set including a pair of a pre- processed single or plurality of multimodal MR image(s) and corresponding pre-processed PET image that are used to evaluate the model's ability to synthesize realistic PET images from unseen examples.
[0092] B. Validation:
[0093] Followed by the model training, a validation step is performed to ensure that the model is well-generalized to synthesize PET images that resemble the pathological patterns confirming AD; by applying the trained DCGAN workflow to the test set. The validation is performed using a 3 -fold process that involves but is not limited to application of image similarity matrices, image quality assessment and comparison of SUVr ratios calculated from real and synthetic PET images. Details of the validation methodologies are mentioned in the previous paragraphs in the detailed description. It is to be noted that these 3 folds can be performed parallelly and the validation procedures are independent of each other.
[0094] C. Model Deployment:
[0095] Once the model has been trained and validated and the processing unit is ready, when a user (21) inputs new data to the processing unit it enables the workflow for synthesis of PET images from multimodal MR images. The model is deployed by first setting up the environment which includes but is not limited to setting up the necessary infrastructure, such as servers, cloud services, or edge devices. The trained model, along with its dependencies (libraries, configuration files, etc.), is packaged for deployment. This may be done using containers (like Docker) to ensure consistency across different devices. The trained model is then loaded onto the device memory. Depending on the deployment environment, this could involve loading from a file system, database, or cloud storage.
[0096] The next step includes creating perceptually realistic synthetic PET images while preserving detailed brain structures. The processing unit of the system (20) uses the multimodal MR image(s), sent as input by the user (21) in a computer acceptable format such as Digital Imaging and Communications in Medicine (DICOM) format, passes the images through the trained DCGAN workflow (1) and provides an output (25) in the form of a PET image.
[0097] Advantages:
[0098] The above invention offers several advantages such as making PET scanning more efficient, acting as an aid in screening individuals who actually need a PET scan thereby saving the patient unnecessary costs and radiation exposure. This invention also potentially allows a method for increasing the use of PET scanning by screening candidates who may be in need of a PET scan. As MRI scanning is less expensive and non-invasive compared to PET scanning and are used as inputs to synthesize PET images, the present invention could aid in reducing harm to the patient’s health, costs for the patient, hospitals and national healthcare systems as it provides a cost-effective, accessible and non-invasive method of imaging. A further advantage of the invention would be acquiring additional information from data that has already been collected in routine clinical procedures as MRI scanning is routinely performed in the clinic for patients with suspected AD.
[0099] While considerable emphasis has been placed herein on the specific elements of the preferred embodiment, it will be appreciated that many alterations can be made and that many modifications can be made in preferred embodiment without departing from the principles of the invention. These and other changes in the preferred embodiments of the invention will be apparent to those skilled in the art from the disclosure herein, whereby it is to be distinctly understood that the foregoing descriptive matter is to be interpreted merely as illustrative of the invention and not as a limitation.
Claims
CLAIMS:I claim,1. A system and method for the synthesis of PET images from multimodal MR images characterized in that the system (20) comprises of a user (21), input / output device(s) (23) and processing unit (24) where the user (21) provides an input (22) using at least one input device (23) such that the processing unit (24) provide an output (25) in the form of synthetic PET images (15) through an output device (23); the method comprises of three stages namely, model training stage, validation stage and model deployment stage such that, a. model training stage comprises the steps of creation of the dataset in the form of pairs of images comprising a single or plurality of multimodal MR images (2) and corresponding single or plurality of PET images (16); determining the dataset size; preprocessing PET(16) and MR images(2); train-test split; and model training using a DCGAN workflow (1) that consists of a generator network (3) that by using an encoderdecoder path and applying a series of convolutional and pooling layers generates synthetic PET images (15) and a discriminator network (17) that evaluates synthetic (15) and real (16) PET images and predicts whether it is synthesized or real based on its learned criteria using the discriminator (18) and generator loss (19); and b. validation stage is performed to assess similarity between synthetic and real PET images (16) and clinical utility of synthesized PET images and steps include image similarity matrices, image quality assessment and comparison of SUVr scores between the synthetic and real PET images (16); and c. in model deployment stage, the trained model is loaded and deployed to obtain synthetic PET images (15) from a single or plurality of multimodal MR images as input (22) and the steps include setting up the environment using a container, loading the trained generator, preprocessing the input MR images (22), processing the input imagesby passing them through the loaded trained DCGAN in the processing unit (24) and generating synthetic PET images (15).
2. The system and method as claimed in claim 1, wherein the multimodal MR images (1) include but are not limited to T1 -weighted MRI, T2 Fluid Attenuated Inversion Recovery (FLAIR), Arterial Spin Labelling, Diffusion-Weighted Tensor Imaging (DWI), resting-state functional MRI; and the PET images (16) include images acquired after injecting radiotracers specific to a biological process that confirms the presence of said disease for each patient sample.
3. The system and method as claimed in claim 1, wherein the dataset size ranges between 1000 - 2000, preferably 1000 image sets and each set includes a pair of MR image (2) and the corresponding PET image (16) of the same patient sample.
4. The system and method as claimed in claim 1, wherein the PET images (16) are pre- processed in order to ensure spatial alignment of brain structures with the MR image(s) and to improve signal-to-noise ratio of PET images and it includes the steps of skull stripping, image registration, image reorientation, attenuation correction, motion correction, image reconstruction, spatial smoothing, normalization, intensity normalization, correction for partial volume effects.
5. The system and method as claimed in claim 1, wherein the dataset is split into training and test sets in a ratio of 90: 10 respectively; such that the training set of 90 percent dataset is used to train the model by using a Deep Convolutional Generative Adversarial Network DCGAN and transfer learning and the test set of 10 percent is used to assess the capability of the model to generalize to unseen data and to assess the clinical utility of the synthesized PET image.
6. The system and method as claimed in claim 1, wherein the generator network (3) using an encoder -decoder path, works by applying a series of convolutional and pooling layers to the input PET (16) and MR images (2), wherein the steps includea. extracting feature maps by convolutional layers (4, 6) and downsample (5, 7) the input (22) by sliding a small filter referred to as a pooling window over the image and estimating the dot product between the filter and the input to form pooled feature maps. b. repeating the step a. to form pooling layers which then downsample (5,7) the output of the convolutional layers (4,6) to reduce the dimensionality of the data and each convolutional layer is followed by batch normalization and a rectified linear unit (ReLU) activation function to introduce non-linearity and aid in feature extraction. Max pooling layers are often used to further downsample (5,7) the feature maps, reducing their spatial resolution while increasing their depth; c. passing the input image through successive convolutional layers (4,6) and downsampling (5,7) operations, such that the encoder path progressively extracts hierarchical features; d. generating the feature maps that have reduced spatial dimensions but increased depth and contain semantically rich information about the input image, capturing both low- level and high-level features such that the output of the encoder path serves as the input to the bridge network (8), connecting the encoder and decoder paths in the U-Net architecture; e. transitioning the feature maps through the bridge (8) that connects the encoder and decoder paths as the bridge (8) consists of one or more convolutional layers (4,6) that retain the learned features while transitioning from the encoding to the decoding stage and skip connections (13, 14) connect the encoder and decoder at each corresponding layer to promote information flow and preserve high-resolution features facilitating the propagation of information from earlier layers to later layers; f. employing decoder path where each upsampling step (9, 11) is followed by a concatenation operation, where feature maps from the encoder path are combined with those from the decoder path and convolutional layers (10, 12) with decreasing filter sizes are applied to refine the features and generate the final output (25); g. applying Batch normalization and ReLU activation functions after each convolutional layer (10, 12) to maintain non-linearity and stabilize training and the finallayer of the decoder path consists of a convolutional layer (12) with a sigmoid activation function, which outputs a synthetic PET image (15).
7. The system and method as claimed in claim 1, wherein the batches of synthetic (15) and real (16) images are then passed through the discriminator network (17) in order to train the discriminator network (17) and compute discriminator loss (18) and generator loss (19) and the discriminator network (17) consists of convolutional layers followed by dense layers, culminating in a binary classification output such that the discriminator evaluates each synthesized image and predicts whether it is synthesized or real based on its learned criteria.
8. The system and method as claimed in claim 1, wherein the validation includes a. application of image similarity matrices such as Mean Squared Error MSE where lower root mean square error (RMSE) values indicate higher similarity and the RMSE values close to 0 indicate perfect similarity, Structural Similarity Index (SSIM) where values range from -1 to 1, and 1 indicates perfect similarity and the values closer to 1 are considered to represent higher similarity and as an acceptable measure of SSIM; or Peak Signal to Noise Ratio where the PSNR values preferably above 30 dB represents good quality; b. assessing image quality ofthe synthetic PET images (15) by allowing the radiologists to rate the quality of both, real and synthesized PET images; where the images are displayed to the radiologist in a random order; c. calculating the Standardized Uptake Value Ratio (SUVr) scores from the synthetic and real PET images (16) to compare the extent to which the synthetic PET images (15) can assess the accumulation of radiopharmaceutical tracers in specific regions of the brain or body where 10% difference in SUVr scores is considered as acceptable.
9. The system and method as claimed in claim 1, wherein the model deployment comprises the steps of: a. setting up the environment which includes but is not limited to setting up the necessary infrastructure, such as servers, cloud services, or edge devices;b. packaging the trained model with only generator network (3), along with its dependencies like libraries, configuration files, for deployment using containers like docker to ensure consistency across different environments; c. loading the trained model onto the system memory from a file system, database, or cloud storage; d. pre-processing the new multimodal MR images, sent as input (22) by the user (21) in a computer acceptable format such as digital imaging and communications in medicine DICOM format by the processing unit (24) of the system (20) and transforming the DICOM files into a format suitable for the generator network (3); e. feeding the pre-processed MR images into the trained generator network (3) in the processing unit (24) that consists of multiple convolutional layers (4,6) and extracts relevant features from the MR images which are then processed by the core network to create a new internal representation; f. generating synthetic PET image as output (25) by the trained generator network (3) on the output device (23) that resemble a realistic PET image from the MR image (22) input while containing new information or variations based on the training data.
Citation Information
Patent Citations
Systems and Methods for Synthetic Medical Image Generation
US20200311932A1
Contextual image translation using neural networks
US20210374947A1
Cited By
Domain adaptation method from multi-modal synthetic image to real image
CN121746161A