Generating synthetic images

Using a GAN to generate synthetic CT images from MR data addresses the limitations of CT scans in radiotherapy planning, enhancing soft-tissue visualization and enabling MRI-only workflows by providing accurate electron density information.

GB2636045APending Publication Date: 2025-06-11ELEKTA AB
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
GB2023012835
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-08-22
Publication Date
2025-06-11

AI Technical Summary

Technical Problem

Current radiotherapy planning methods rely on CT scans, which expose patients to radiation and provide poor soft tissue contrast, while MR scans lack necessary tissue attenuation information for accurate dose calculations, complicating MRI-only radiotherapy workflows.

Method used

A generative adversarial network (GAN) is used to create synthetic CT (sCT) images from MR images, providing accurate tissue attenuation information without additional radiation exposure, enhancing soft-tissue visualization and enabling MR-only radiotherapy workflows.

Benefits of technology

The method generates high-quality synthetic CT images with proper Hounsfield Units (HU) for electron density calculations, reducing patient radiation exposure and improving treatment planning accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for generating a synthetic medical image, comprising using a generative model, optionally trained via a generative adversarial network (GAN), to generate the synthetic image based on a first
Need to check novelty before this filing date? Find Prior Art

Description

This disclosure relates to generating synthetic images, and in particular to generating synthetic images based on a first image of a first imaging modality, for example to generate synthetic CT or CBCT images based on a magnetic resonance image. Background Radiotherapy can be described as the use of ionising radiation, such as X-rays, to treat a human or animal body. Radiotherapy is commonly used to treat cancer, for example to treat tumours within the body of a patient or subject. In such treatments, ionising radiation is used to irradiate, and thus destroy or damage, cells which form part of the tumour. Medical imaging has become increasingly important in the diagnosis and treatment of oncological patients, particularly in radiotherapy (RT). Before radiotherapy treatment, a treatment plan is created to determine how and where the radiation should be applied. Typically, such a treatment plan is created with the assistance of medical imaging technology. For example, a CT scan may be taken of the patient in order to produce a three-dimensional image of the area to be treated. The three-dimensional image allows the treatment planner to observe and analyse the target region and identify surrounding tissues. Different structures within the body of the patient, for example bone, lung, muscle, etc., will attenuate and absorb radiation to differing degrees based on their respective densities. In other words, different tissues within the human body have different radiodensities, and hence attenuate and / or absorb radiation to differing degrees. Bone is an example of a particularly radiodense or radiopaque tissue. In contrast, soft tissue, for example lung tissue, is radiolucent. The radiodensity of various tissues can be quantified using, for example, the Hourisfield scale in a manner which is known to the skilled person. In order to plan the radiotherapy treatment, it is necessary to obtain information regarding the radiodensity of not only the target region, such as a tumour, but also surrounding tissues and any regions of the body through which the radiation treatment beam will pass. A CT image not only provides information about the patient geometry by providing a three-dimensional image of the patient, but also provides information regarding the radiodensities of the different tissues and structure within the patient's body. A CT scan typically produces a three-dimensional image comprised of voxels, where each voxel is assigned a CT value. Each voxel is associated with a particular location within the patient's body, and the CT-values of the voxels together describe the radiodensity of tissues within the patient's body. The CT-values are indicative of the attenuation properties of tissues within the patient's body. CT-values can be expressed in Hounsfield units, and are directly related to the electron density information which is needed for radiation dose calculations. However, a CT scan involves irradiating the patient from a number of angles in order to produce the three-dimensional image. Therefore, a disadvantage of CT scanning is that it adds a radiation dose to the patient even before treatment has started. Also, while CT scans can provide the necessary information regarding tissue densities for radiation treatment planning, CT scans provide poor soft tissue contrast. This makes it difficult for a treatment planner to distinguish between certain kinds of soft tissue. For example, it is very difficult to see a tumour in a prostate CT scan, because the tumour and the prostate have very similar densities and attenuation properties and thus look similar, or even identical, in a CT image. By comparison, obtaining magnetic resonance (MR) images does not involve exposing a patient to ionising radiation, and hence does not provide any dose to the patient. MR images provide good soft tissue contrast, allowing a treatment planner to better distinguish between, for example, tumour tissue and prostate tissue. The disadvantage of MR scanning is that it does not indicate the attenuation properties of tissues within the subject, i.e. the patient's tissue radiodensity information, which is required to create a radiation treatment plan. It is already possible to combine MR and CT data to facilitate the treatment planning process. It is known that MR images and CT images may be obtained independently and then aligned with each other, for example by aligning various discernible locations of features of interest in both images. Such alignment, or fusing, may involve either rigid or deformable adjustment of the MR image to align with the CT image to produce a pair of co-aligned images. However, such alignment, or registration, still requires a CT image to be taken, thereby providing a radiation dose to a patient, and this approach takes time, increases workload, and sometimes has sub-optimal accuracy. A certain degree of progress in the field has been made by creating so-called 'pseudo-CT', or synthetic CT (sCT) images. These are images similar to those obtained from a CT scan, i.e. which comprise tissue radiodensity information, but which are derived primarily or entirely from MR data such as an MR image of the patient. Recently, MRI-only based radiotherapy has been proposed to simplify and speed up the workflow, thereby decreasing patients' exposure to radiation. MRI-only RT may also reduce overall treatment costs and workload, and eliminate residual registration errors when using both imaging modalities. Additionally, the development of MRI-only techniques can be beneficial for MRI-guided RT such as can be accomplished using an MR-linac, for example. However, an issue with known approaches, and an obstacle to introducing MRI-only RT, is the lack of accurate tissue attenuation information required for accurate dose calculations. The present invention seeks to address these and other disadvantages encountered in the prior art. Summary An invention is set out in the independent claims. Optional features are set out in the dependent claims. Figures Specific embodiments are now described, by way of example only, with reference to the drawings, in which: Figure 1 depicts a radiotherapy device or apparatus according to the present disclosure; Figure 2 depicts a method of generating a synthetic image according to the present disclosure; Figure 3 depicts a method of training a model in accordance with the present disclosure; Figure 4a depicts a GAN model in accordance with the present disclosure; Figure 4b depicts an architecture for a generator in accordance with the present disclosure; Figures 5a-c depict the results of a multiple channel approach in accordance with the present disclosure, for example to eliminate the 'staircase effect'; Figures 6a-b show representative examples of MR-to-sCT translation; Figure 7 depicts a system according to the present disclosure; Figure 8 depicts a storage medium according to the present disclosure. Detailed Description The methods of generating a synthetic CT described herein are particularly beneficial for MR-guided external beam radiotherapy (MRgRT). This application area can strongly benefit from Al-powered image-to-image translation, for generating synthetic CT (sCT) contrast from MR images. This process can be referred to as MR-to-sCT. MR-to-sCT is an enabler for MR-only radiotherapy workflows. Such workflow could avoid the use of ionizing radiation imaging, while providing radiation-oncologists (RO) with enhanced soft-tissue visualization (with MRI) alongside co-registered sCT contrast with proper Hounsfield Units (HU) to infer electron density (ED) values for RT dose calculation. Moreover, MR-to-sCT can significantly enhance current online plan adaptation techniques in MRgRT. For example, current online adaptative workflows with Unity MR-Linac (Elekta AB, Stockholm, Sweden) require manual or semi-automatic contouring of different structures on the daily MRI, including all bony anatomy. Bulk electron density values are then assigned to each contoured structure as part of the plan adaptation process. MR-to-sCT could provide more realistic HU and ED values, while significantly reducing the burden of re-contouring all bony anatomy on each daily MRI. Figure 1 depicts an MR-linac radiotherapy device suitable for delivering, and configured to deliver, a beam of radiation to a patient during radiotherapy treatment. The device and its constituent components will be described generally for the purpose of providing useful accompanying information for the present invention. The device depicted in figure 1 is in accordance with the present disclosure and is suitable for use with the disclosed systems and apparatuses. While the device in figure 1 is an MR-linac, the generation of synthetic images, in particular sCT images, may enable workflows on any radiotherapy device, even one without an MR imager, for example a traditional linac device. The device 100 comprises both MR imaging apparatus 112 and radiotherapy (RT) apparatus which may comprise a linac device. The MR imaging apparatus 112 is shown in cross-section in the diagram. In operation, the MR scanner produces MR images of the patient, and the linac device produces and shapes a beam of radiation and directs it toward a target region within a patient's body in accordance with a radiotherapy treatment plan. The depicted device does not have the usual 'housing' which would cover the MR imaging apparatus 112 and RT apparatus in a commercial setting such as a hospital. The MR-linac device depicted in figure 1 comprises a source of radiofrequency waves 102, a waveguide 104, a source of electrons 106, a source of radiation 106, a collimator 108 such as a multi-leaf collimator configured to collimate and shape the beam, MR imaging apparatus 112, and a patient support surface 114. In use, the device may also comprise a housing (not shown) which, together with the ring-shaped gantry, defines a bore. The moveable support surface 114 can be used to move a patient, or other subject, into the bore when an MR scan and / or when radiotherapy is to commence. The MR imaging apparatus 112, RT apparatus, and a subject support surface actuator are communicatively coupled to a controller or processor. The controller is also communicatively coupled to a memory device comprising computer-executable instructions which may be executed by the controller. The radiation source may comprise a beam generation system. For a linac, the beam generation system may comprise a source of RF energy 102, an electron gun 106, and a waveguide 104. The radiation source is attached to the rotatable gantry 116 so as to rotate with the gantry 116. In this way, the radiation source is rotatable around the patient so that the treatment beam 110 can be applied from different angles around the gantry 116. In a preferred implementation, the gantry is continuously rotatable. In other words, the gantry can be rotated by 360 degrees around the patient, and in fact can continue to be rotated past 360 degrees. The gantry may be ring-shaped. In other words, the gantry may be a ring-gantry. The source 102 of radiofrequency waves, such as a magnetron, is configured to produce radiofrequency waves. The source 102 of radiofrequency waves is coupled to the waveguide 104 via circulator 118, and is configured to pulse radiofrequency waves into the waveguide 104. Radiofrequency waves may pass from the source 102 of radiofrequency waves through an RF input window and into an RF input connecting pipe or tube. A source of electrons 106, such as an electron gun, is also coupled to the waveguide 104 and is configured to inject electrons into the waveguide 104. In the electron gun 106, electrons are thermionically emitted from a cathode filament as the filament is heated. The temperature of the filament controls the number of electrons injected. The injection of electrons into the waveguide 104 is synchronised with the pumping of the radiofrequency waves into the waveguide 104. The design and operation of the radiofrequency wave source 102, electron source and the waveguide 104 is such that the radiofrequency waves accelerate the electrons to very high energies as the electrons propagate through the waveguide 104. The electrons may travel toward a heavy metal target which may comprise, for example, tungsten. The source of radiation is configured to direct a beam 110 of therapeutic radiation toward a patient positioned on the patient support surface 114. The source of radiation may comprise a heavy metal target toward which the high energy electrons exiting the waveguide are directed. When the electrons strike the target, X-rays are produced in a variety of directions. A primary collimator may block X-rays travelling in certain directions and pass only forward travelling X-rays to produce a treatment beam 110. The X-rays may be filtered and may pass through one or more ion chambers for dose measuring. The beam can be shaped in various ways by beam-shaping apparatus, for example by using a multi-leaf collimator 108, before it passes into the patient as part of radiotherapy treatment. The subject or patient support surface 114 is configured to move between a first position substantially outside the bore, and a second position substantially inside the bore. In the first position, a patient or subject can mount the patient support surface. The support surface 114, and patient, can then be moved inside the bore, to the second position, in order for the patient to be imaged by the MR imaging apparatus 112 and / or imaged or treated using the RT apparatus. The movement of the patient support surface is effected and controlled by a subject support surface actuator, which may be described as an actuation mechanism. The actuation mechanism is configured to move the subject support surface in a direction parallel to, and defined by, the central axis of the bore. The terms subject and patient are used interchangeably herein such that the subject support surface can also be described as a patient support surface. The subject support surface may also be referred to as a moveable or adjustable couch or table. The radiotherapy apparatus / device depicted in figure 1 also comprises MR imaging apparatus 112. The MR imaging apparatus 112 is configured to obtain images of a subject positioned, i.e. located, on the subject support surface 114. The MR imaging apparatus 112 may also be referred to as the MR imager. The MR imaging apparatus 112 may be a conventional MR imaging apparatus operating in a known manner to obtain MR data, for example MR images. The skilled person will appreciate that such a MR imaging apparatus 112 may comprise a primary magnet, one or more gradient coils, one or more receive coils, and an RF pulse applicator. The operation of the MR imaging apparatus is controlled by the controller. The controller is a computer, processor, or other processing apparatus. The controller may be formed by several discrete processors; for example, the controller may comprise an MR imaging apparatus processor, which controls the MR imaging apparatus 110; an RT apparatus processor, which controls the operation of the RT apparatus; and a subject support surface processor which controls the operation and actuation of the subject support surface. The controller is communicatively coupled to a memory, e.g. a computer readable medium. The controller is described in greater detail elsewhere herein, but in summary is arranged and configured to implement the methods of the present disclosure. Figure 2 is a flowchart depicting a method 200 according to the present disclosure at a high level. The method is for generating a synthetic image. At block 210, a first image is received. The first image is of a first imaging modality, and e.g. may have been acquired using imaging apparatus of a first imaging modality. The first imaging modality may be MR, such that the first image is an MR image. The first image depicts an anatomical region of a subject. For example, the first image may depict the pelvis or brain of a patient. At the time of imaging, the patient may be positioned on the subject / patient support surface of a radiotherapy machine such as an MR-linac. Alternatively, at block 210, first imaging data may be acquired, where the data is representative of the first image. At block 220, a synthetic image is generated using a generative model. The synthetic image corresponds to the first image, and may for example depict the same anatomical region of the patient, or at least appear to. The synthetic image corresponds to a second imaging modality. The synthetic image may take the form of a second imaging modality image. The synthetic image may represent, or be similar to an image of the second imaging modality, without having been acquired using imaging apparatus of the second imaging modality. The second imaging modality may be any of computed tomography (CT) and cone-beam computed tomography (CBCT). For example, the image may be an sCT image. In an implementation, generating the synthetic image may comprise generating synthetic imaging data using the generative model, where the synthetic imaging data is representative of a region of the input (first) image. For example, the input image may be subdivided into a plurality of image samples, each sample representing a region of the input image, and these samples can be presented as input to the generative model in order to generate corresponding regions of the synthetic image. This process can be repeated multiple times in order to generate synthetic imaging data representing multiple different, overlapping regions of the input image. The method may then comprise aggregating the results. For example, where the input is a 3D image and the synthetic imaging data comprises synthetic CT values generated for voxels of the input image, aggregating the results may comprise calculating an average synthetic CT value from each of the synthetic CT values generated by the model. Figure 3 depicts an example training method 300 for training the generative model. As shown in figure 3, the generative model may be trained via a generative adversarial network (GAN). The GAN may comprise a first generator convolutional network and a first discriminator convolutional network. At block 310, a training set of images is received, which comprises images of a first imaging modality and images of a second imaging modality. These images may be real images. The images may comprise real MR images (either T1 orT2 weighted, or a mixture of both) and real CT images. These images may be paired, such that each pair of images comprises a real MR and real CT image, the images being registered to one another. The CT and MR images of a particular pair of images depict the same patient, and in particular depict the same anatomical region of that patient to enable the images to be registered. However, different image pairs may depict different patients. Suitable training data is discussed in greater detail below. At blocks 320 and 330, generative and discriminative models are trained in a GAN. At block 320, a generative model, such as a generator convolutional network, is trained to generate synthetic images based on input training images of the first imaging modality, for example based on MR images. These synthetic images correspond to the second imaging modality. The generator network's 'goal' is to create new data instances in the form of synthetic CT images that resemble the real CT images in the training dataset. At block 330, a discriminative model such as a discriminator convolutional network is configured to discriminate between the synthesized images generated by the generative model, and the actual images comprised within the training data. The discriminator network's 'goal' is to distinguish between real CT images in the training dataset and fake, synthetic images generated by the generator. In this way, the output of the generative and discriminative networks can be used to train each other in an adversarial manner, as is understood by the skilled person. Block 320 may comprise providing, as an input to the generative model, images or imaging data based on the images of the first imaging modality, and receiving estimated synthetic imaging data as an output from the generative model, it can then be determined, using the discriminative model in the GAN, at least one difference between the resulting estimated synthetic imaging data and real imaging data based on the images of the second imaging modaiity. One or more parameters of the generative model can then be updated based on the determined at least one difference. This process can be performed iteratively until a stopping criterion is reached, for example a stopping criterion relating to the number of passes through the iterative process, or relating to the determined "at least one difference", e.g. a similarity metric between synthetic and real imaging data is reached, the method comprises outputting the trained model; At block 350, a trained generative model is output following the adversarial training process, and may be used to generate sCT or other synthetic images in the manner depicted in figure 2. The skilled person will understand that, where reference is made to an "image" herein, particularly with reference to the input or output of a machine learning model such as a generative or discriminative model, this should be considered to encompass the use of parts, regions or patches of an image. It should be understood that where reference is made to an "image" herein, this encompasses the use of "imaging data" such as images samples and patches, which incorporate pixel and voxel values and associated values such as CT scores associated with pixels and / or voxels. Reference is made to the methods of both figures 2 and 3. The first image may be a three-dimensional (3D) image comprising a plurality of voxels. As is known to the skilled person, a 3D image can be subdivided in a number of ways, for example into a plurality of 2D image slices arranged perpendicular to any of the three image axes. These image axes may define image planes. The 3D image may therefore be divided into a first plurality of 2D image slices arranged perpendicular to a first image axis, and into two different pluralities of 2D image slices, each arranged perpendicular to a second and third image axis respectively. As described in more detail below with respect to the disclosed case study implementing the disclosed method, methods of the present disclosure may involve using a multi-channel "2.5D" approach in order to improve the translation results in comparison to applying a 2D machine learning model independently to each 2D image slice. Using this approach, generating a synthetic image may comprise providing, as an input to the generative model, a plurality of consecutive 2D image patches. Each 2D image patch depicts a 2D region in the first image. When considered together, the plurality of consecutive 2D image patches may be referred to as an image sample. The consecutive 2D image patches are arranged perpendicular to a first image axis. The sample size may be, for example, 5 consecutive 2D image patches, where each 2D patch comprises 192x192 voxels. This image sample can be described as a stack of 5 2D image patches. Following this approach, when the model is trained e.g. according to the method of figure 3, the method may comprise randomly sampling groups of 5 consecutive slices of imaging data and supplying them to the generator and / or the discriminator. At inference time, for example when an sCT is to be generated according to the method of figure 2, the first (input) image is subdivided into patches and / or samples that overlap both in-plane and thru-plane. In this implementation, it is possible to consider three example image samples, or three 'stacks' of 2D image patches. Each stack of 2D image patches comprises a plurality of consecutive, adjacent 2D image patches. The first and second stack of 2D image patches are aligned along the first image axis, and the third stack of 2D image patches is aligned along a second (different) image axis. These axes may be perpendicular to one another. The first stack and the second stack overlap, such that a first subset of voxels is comprised within both the first stack and the second stack of 2D image patches. Similarly, the first stack and the third stack overlap, such that a second subset of voxels is comprised within both the first stack and the third stack of 2D image patches. When each of these stacks is passed through the generator network to generate first, second, and third synthetic image data, the voxels which are common to multiple 'stacks' of image patches will each have multiple values associated therewith. The 'final' values for each voxel in the synthetic image are an aggregation of each value generated in the aggregation stage. For example, the final voxel value may be an average value. In a particular implementation, the aggregation involves performing a weighted average, where the weight for any voxel value of a particular image patch is based on its proximity to the centre of the image patch. This means that a central pixel / voxel in the central image slice of an image sample will have a high weighting, and a pixel / voxel at the edge of an image sample will have a low weighting. Figure 4a depicts an example machine learning model that comprises a generative model in accordance with the present disclosure. The model can be trained using training data, for example training data as discussed above with respect to block 310 of figure 3. The training data comprises a series of paired input images, for example co-registered MRI and CT images of the same patient. The generator, which may also be described as a generator network, takes as input one or multiple (consecutive) real MRI slices and generates as output one or multiple (consecutive) synthetic CT slices. For example, a standard residual U-Net may be used for the generator network. The discriminator, which may also be describes as a discriminator network, takes as input one or multiple real MRI slices and one or multiple real or synthetic CT slices. The goal of the discriminator is to guess the supplied CT images as either real or fake (synthetic). Once the network is trained, the generator network is used for inference. Figure 4b depicts a generator network which may be employed in certain implementations of the present disclosure. The depicted generator is a Resllnet generator. The skilled person will understand that the filter sizes are variable; and in a preferred implementation the filter sizes are not in line with those depicted in figure 4b but as are as follows: Input layer: and first layer 64x5x152x152 instead of 16x112x112x112. Second layer: 128x5x76x76 Third layer: 256x5x38x38 Fourth layer: 512x5x19x19 The discriminator network is not shown, but may be a patch discriminator network and implemented similarly to the generator's encoder portion. A specific case study and implementation of the methods and machine learning models described herein will now be described for the purposes of aiding understanding. In the specific implementation, the method is for MR-to-sCT image translation using paired training data. The method is based on the Pix2Pix conditional GAN architecture. A multi-channel (2.5D) approach is used to improve translation results thru-plane in comparison to applying a 2D model independently on each slice, while keeping inference time small in comparison to a full 3D approach. Separate models were trained for both brain (Tl-weighted) and pelvis (Tl- and T2-weighted) using already paired data. According to this approach, models have been validated using 60 validation subjects. Image similarity metrics obtained during the validation phase are: mean absolute error (MAE) of 64.27 ±14.15, peak signal-to-noise ratio (PSNR) of 28.64 ± 1.77, structure similarity index (SSIM)of 0.872 ±0.032. In terms of prior art for this case study, one method for generating realistic sCT is to use deformable image registration of the patient planning CT with the daily MRI. However, deformable registration across two different modalities is limited in accuracy, especially for highly elastic structures like the bladder. Al methods based on Generative Adversarial Networks (GANs) can be used for image-to-image translation in medical imaging applications. For example, conditional GANs have been used to train models with paired and accurately registered training images. Alternatively, CycleGANs can be used when training data is unpaired. As such, they can be trained using unmatched training samples that are only available in one modality (MR or CT). This case study describes a method to enable MR-to-sCT translation, with image data depicting both pelvis and brain tumor sites. All training samples used for this case study are paired, and for this reason a conditional GAN architecture has been adopted. Section 2 below provides implementation details about the method. Section 3 describes strategies for hyperparameter optimization and summarizes the results obtained during a validation phase. Section 4 discusses specific design decisions. 2 Method 2.1 Training Data Fully anonymized data was collected from 3 different sites: Radboud University Medical Center, University Medical Center Utrecht, University Medical Center Groningen. The case study described herein consists of 3 phases: training phase, validation phase and test phase. For the training phase, a set of 180 brain MRI / CT pairs and 180 pelvis MRI / CT pairs was used. The MR and CT data were rigidly aligned. An evaluation mask was also used. All brain MRIs were Tl-weighted. 120 pelvis MRIs were Tl-weighted. The remaining 60 pelvis MRIs were T2-weighted. For the validation phase, a set of 30 brain MRI images and 30 pelvis MRI images was provided to the participants. The corresponding CT images were not used. 2.2 2D Conditional GAN In this implementation, a 2D conditional GAN (Pix2Pix) network is employed as follows. Let G: X -> Y be a generator that translates images from domain X to domain Y, and D be a discriminator network trained to distinguish between true and synthetized images in domain Y. The original Pix2Pix objective is: £(G,D)=£c^(G,D) + Mfi(G) (1) where £GAN is the GAN loss and LR is a regression loss measuring the difference between the output translation and the (aligned) ground truth. In the present implementation, three possible objective functions can be used for £GAN : classic (described, for example, in Goodfellow, I.: NIPS 2016 Tutorial: Generative Adversarial Networks, http: / / arxiv.org / abs / 1701.00160, (2017) which is incorporated herein by reference); least squares (described, for example, in Mao, X., Li, Q., Xie, H., Lau, R.Y.K., Wang, Z., Smolley, S.P.: Least Squares Generative Adversarial Networks, http: / / arxiv.org / abs / 1611.04076, (2017) which is incorporated herein by reference); and Hinge (described, for example, in Lim, J.H., Ye, J.C.: Geometric GAN, http: / / arxiv.org / abs / 1705.02894, (2017) which is incorporated herein by reference). A LI loss term is used for : ER = E^GCx) - y\\± (2) The generator network is implemented with a Resllnet (described, for example, in Zhang, Z., Liu, Q., Wang, Y.: Road Extraction by Deep Residual U-Net. IEEE Geosci. Re-mote Sensing Lett. 15, 749-753 (2018) https: / / doi.org / 10.1109 / LGRS.2018.2802944, which is incorporated herein by reference). The number of layers and filters at each layer is fully parametrizable. The discriminator network is implemented similarly to the encoding part of the ResUnet. Both the generator and discriminator network can optionally use spectral normalization following each convolutional layer, and instance normalization is used in place of group normalization as shown in figure 4b. Data augmentation can optionally be used, and may for example be randomly run on-the-fly during the training loop. Supported data augmentation includes: affine transformation, synthetic multiplicative bias field, blurring, sharpening, Gamma contrast change, linear intensity transform. Note that the affine transformation is applied to both MR and CT images. All other forms of data augmentation are applied to the MR images only. The model is implemented in Python using the PyTorch library. Optimization is carried out using Adam with pi = 0.5 and P2 = 0.999. A slow-moving exponential moving average (EMA) of the generator parameters was tracked during training (with a = 0.999, weights updated at every batch) and used as the final model for inference. A complete list of network hyperparameters is provided in Table 1 below. 2.3 Image Pre-Processing The voxel resolution for model training may be chosen as lxlxl mm for brain model and lxlx2.5 mm for the pelvis. If necessary, MR and CT images may be resampled to the model resolution and back to native resolution (during inference). MR and CT intensities are linearly rescaled to a range of [-1, +1], with a source range determined using percentiles for MR and a fixed source range of [-1000, +2200] HU or [-1000, +3000] HU to support metal artifacts. In the case study, models trained with the full range [-1000, +3000] were slightly less accurate because 1 / 4 of the intensity range [+2000, +3000] is reserved for unusual intensities. One strategy according to the present disclosure is to train one network with the full range and one with a narrower range, and use the content of the first network for intensities >2200 (if any) and the content of the latter network otherwise. During training, patches with a predefined sample size are randomly drawn from the training set images and randomly augmented. During inference, the test image is subdivided into overlapping patches and the model output combined via a weighted average - when combining the output of multiple overlapping patches, higher weight is given if a pixel is near the center of the patch and lower 5 weight if the pixel is close to the edge of a patch. Table 1. List of model hyperparameters ! Hyperparameter 1 ^GAN Description Type of adversarial loss function to use (either Classic, Least Squares or Hinge) Recommended value Least Squares j Learning rate Learning rate for the generator (G) and discriminator (D). The learning rate will linearly decrease to zero starting when half the number of iterations has been completed. G: le-4 D: 5e-5 j Num filters j Num disc updates Number of filters per layer in the generator (G) and discriminator (D) Number of discriminator updates per generator update G: 64,128,256,512 D: 64,128, 256,512,512 2 I Spectral norm Specify if spectral normalization layers are used in both the discriminator and generator True networks True j LI weight (X) Weight of LI loss term 50 j Voxel size Voxel size in mm Brain: Ixlxl mm Pelvis: 1x1x2.5 mm I Sample size The size of a sample (in voxels) used during training and inference. For a 2D network, the 3rd dimension is the channel dimension (c.f. Section 2.4) 192x192x5 j Batch size Number of samples per batch 16 ! Num batches per iteration Number of batches per iteration 300 j Num iterations Total number of iterations 2500 i Data augmentation Option to turn on data augmentation True 2.4 Multi-Channel Implementation One problem with applying a 2D neural network independently on each axial slice of a 3D image is the 10 so-called "staircase effect" that can appear in other orientations (sagittal and coronal). This is shown in figure 5a. One possible mitigation is to train a 3D network. However, this can lead to larger training and inference time, and higher GPU memory consumption in comparison to a 2D approach. Instead, a novel 2.5D approach is used, where five consecutive slices of data are supplied to the network, using the channel dimension. As illustrated in figure 5b, this approach significantly reduces the staircase effect without the drawbacks of a full 3D approach. At training time, no change is required other than supplying randomly sampled groups of 5 consecutive slices of data. At inference time, the test image is subdivided into patches that overlap both in-plane and thru-plane. The model output is averaged as previously described in section 2.3. 3 Results 3.1 Strategy for hyperparameter tuning During the training phase, the data was split for each anatomy (brain and pelvis) into a training set (75% of data) and a tuning set (25% of data). The training set is used to train the network and the tuning set is used for computing the image similarity metrics - mean absolute error (MAE), peak signal-to-noise (PSNR) and structure-similarity index (SSIM) - at the end of each training run. We started with an initial set of hyperparameters and then launched several trainings for the brain anatomy while varying one parameter at a time. Once a new set of best hyperparameters was found, this procedure was repeated a second time to obtain a final set of hyperparameters. Finally, a small hyperparameter search was performed on the pelvis anatomy, but no changes were found necessary other than the voxel size for keeping the training time reasonable. 3.2 Validation phase results Once the validation phase started, brain and pelvis models were trained independently, using 100% of the training data and using best found hyperparameters. The tables of figures 6a and 6b show the results obtained during validation. The MAE is 64.27 ± 14.15 the PSNR is 28.64 ± 1.77 and the SSIM is 0.872 ± 0.032. Representative examples of brain and pelvis translations are shown in Fig. 6a. For each model, six networks were trained with the hyperparameters described in Table 1. The output of all six networks were combined with ensemble averaging. The average inference time for one network is approximately 20 seconds for brain images and 30 seconds for pelvis images on a workstation equipped with a NVIDIA V100 16GB GPU card. With ensemble averaging, the inference time scales linearly with the number of networks composing the model. Figures 6a and 6b show representative examples of MR-to-sCT translation on the validation set. Images with best, median and worst MAE are shown. Synthetic CT images are displayed with different window / level settings to emphasis brain and / or soft tissue (middle image) and bone (right image). Table 2. Final results of validation phase MAE PSNR SSIM Mean 64.27 ±14.15 28.64 ± 1.77 0.872 ±0.032 Min 32.75 24.60 0.789 25pc 55.73 27.56 0.855 50pc 62.28 28.66 0.873 75pc 72.58 29.38 0.890 Max 109.88 34.58 0.969 4 Discussion The results presented in section 3 indicate that a model based on the pix2pix architecture is suitable for MR-to-sCT image translation. Given that this is a paired method, accurate image registration is necessary between each MR / CT pairs. Rigid data alignment was used. Different pulse sequences were used for the pelvis MRI (2 / 3 of the training data wasTl-weighted, 1 / 3 was T2-weighted). Additional accuracy could be obtained by training different models for Tl-weighted and T2-weighted MRI. The maximum inference time was set to 15 minutes per image, which is acceptable for offline use, but not suitable for online adaptive workflows. In this work, ensemble averaging leads to small improvements on the overall image similarity metrics. However, the computational cost increases linearly with the number of networks in the model, with possibly very limited impacton the dosimetry. For the pelvis tumor site, gas pockets in the rectum may not translate well into the sCT. This is because transient gas pocket information is not paired in the training data (i.e. MR and CT scans were not taken simultaneously). Moreover, they are difficult to visualize in MRI. Consequently, gas pockets may be hallucinated or missed out in the sCT. One mitigation is to mask all gas pockets in the training CTs with an average HU value. This may prevent the model from hallucinating air pockets during inference and, to some extent, would be consistent with some clinical workflows (large gas pockets are sometimes manually segmented and assigned an average HU / ED). However, during online imaging, gas pockets can provide useful information and radiation therapists may decide to wait for it to pass before starting the treatment. The present application discloses the use of a 2.5D (multi-channel) pix2pix network for training MR-to-sCT models and discloses a case study which demonstrate results for both brain (Tl-weighted) and pelvis (multi-contrast) anatomies. Separate models were provided in the case study for the different tumor sites as it led to superior image similarity metrics. Figure 7 illustrates a block diagram of one implementation of a system 700, for example a radiotherapy system. The radiotherapy system 700 comprises a computing system 710 within which a set of instructions, for causing the computing system 710 to perform any one or more of the methods discussed herein, may be executed. The computing system 710 shall be taken to include any number or collection of machines, e.g. computing device(s), that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods discussed herein. That is, hardware and / or software may be provided in a single computing device, or distributed across a plurality of computing devices in the computing system. In some implementations, one or more elements of the computing system may be connected (e.g., networked) to other machines, for example in a Local Area Network (LAN), an intranet, an extranet, or the Internet. One or more elements of the computing system may operate in the capacity of a server or a client machine in a client-server network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. One or more elements of the computing system may be a personal computer (PC), a tablet computer, a set-top box (STB), a Personal Digital Assistant (PDA), a cellular telephone, a web appliance, a server, a network router, switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that machine. The computing system 710 includes controller circuitry 711 and a memory 713 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM), etc.). The memory 713 may comprise a static memory (e.g., flash memory, static random access memory (SRAM), etc.), and / or a secondary memory (e.g., a data storage device), which communicate with each other via a bus (not shown). Controller circuitry 711 represents one or more general-purpose processors such as a microprocessor, central processing unit, accelerated processing units, or the like. More particularly, the controller circuitry 711 may comprise a complex instruction set computing (CISC) microprocessor, reduced instruction set computing (RISC) microprocessor, very long instruction word (VLIW) microprocessor, processor implementing other instruction sets, or processors implementing a combination of instruction sets. Controller circuitry 711 may also include one or more special-purpose processing devices such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), network processor, or the like. One or more processors of the controller circuitry may have a multicore design. Controller circuitry 711 is configured to execute the processing logic for performing the operations and steps discussed herein. The computing system 710 may further include a network interface circuitry 718. The computing system 710 may be communicatively coupled to an input device 720 and / or an output device 730, via input / output circuitry 717. In some implementations, the input device 720 and / or the output device 730 may be elements of the computing system 710. The input device 720 may include an alphanumeric input device (e.g., a keyboard or touchscreen), a cursor control device (e.g., a mouse or touchscreen), an audio device such as a microphone, and / or a haptic input device. The output device 730 may include an audio device such as a speaker, a video display unit (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)), and / or a haptic output device. In some implementations, the input device 720 and the output device 730 may be provided as a single device, or as separate devices. In some implementations, computing system 710 includes training circuitry 718. The training circuitry 718 is configured to train a method of generating a synthetic image; for example using the method described above with respect to figure 3. The model may comprise a deep neural network (DNN), such as a convolutional neural network (CNN) and / or recurrent neural network (RNN). Training circuitry 718 may be configured to access training data and / or testing data from memory 713 or from a remote data source, for example via network interface circuitry 715. In some examples, training data and / or testing data may be obtained from an external component, such as image acquisition device 740 and / or treatment device 750. In some implementations, training circuitry 718 may be used to update, verify and / or maintain the model for generating a synthetic image. In some implementations, the computing system 710 may comprise image processing circuitry 719. Image processing circuitry 719 may be configured to process image data 780 (e.g. images, or imaging data), such as medical images obtained from one or more imaging data sources, a treatment device 750 and / or an image acquisition device 740. Image processing circuitry 719 may be configured to process, or pre-process, image data. For example, image processing circuitry 719 may convert received image data into a particular format, size, resolution or the like. In some implementations, image processing circuitry 719 may be combined with controller circuitry 711. In some implementations, the radiotherapy system 700 may further comprise an image acquisition device 740 and / or a treatment device 750, such as those disclosed herein in the examples of Figure 1. The image acquisition device 740 and the treatment device 750 may be provided as a single device. In some implementations, treatment device 750 is configured to perform imaging, for example in addition to providing treatment and / or during treatment. Image acquisition device 740 may be configured to perform positron emission tomography (PET), computed tomography (CT), but in a preferred implementation is configured to generate magnetic resonance (MR) images. Image acquisition device 740 may be configured to output image data 780, which may be accessed by computing system 710. Treatment device 750 may be configured to output treatment data 760, which may be accessed by computing system 710. Computing system 710 may be configured to access or obtain treatment data 760, planning data 770 and / or image data 780. Treatment data 760 may be obtained from an internal data source (e.g. from memory 713) or from an external data source, such as treatment device 750 or an external database. Planning data 770 may be obtained from memory 713 and / or from an external source, such as a planning database. Planning data 770 may comprise information obtained from one or more of the image acquisition device 740 and the treatment device 750. The various methods described above may be implemented by a computer program. The computer program may include computer code (e.g. instructions) 810 arranged to instruct a computer to perform the functions of one or more of the various methods described above. The steps of the methods described above may be performed in any suitable order. The computer program and / or the code 810 for performing such methods may be provided to an apparatus, such as a computer, on one or more computer readable media or, more generally, a computer program product 800)), depicted in Figure 8. The computer readable media may be transitory or non-transitory. The one or more computer readable media 800 could be, for example, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, or a propagation medium for data transmission, for example for downloading the code over the Internet. Alternatively, the one or more computer readable media could take the form of one or more physical computer readable media such as semiconductor or solid state memory, magnetic tape, a removable computer diskette, a random access memory (RAM), a read-only memory (ROM), a rigid magnetic disc, and an optical disk, such as a CD-ROM, CD-R / W or DVD. The instructions 810 may also reside, completely or at least partially, within the memory 713 and / or within the controller circuitry 711 during execution thereof by the computing system 710, the memory 713 and the controller circuitry 711 also constituting computer-readable storage media. In an implementation, the modules, components and other features described herein can be implemented as discrete components or integrated in the functionality of hardware components such as ASICS, FPGAs, DSPs or similar devices. A "hardware component" is a tangible (e.g., non-transitory) physical component (e.g., a set of one or more processors) capable of performing certain operations and may be configured or arranged in a certain physical manner. A hardware component may include dedicated circuitry or logic that is permanently configured to perform certain operations. A hardware component may comprise a special-purpose processor, such as an FPGA or an ASIC. A hardware component may also include programmable logic or circuitry that is temporarily configured by software to perform certain operations. In addition, the modules and components can be implemented as firmware or functional circuitry within hardware devices. Further, the modules and components can be implemented in any combination of hardware devices and software components, or only in software (e.g., code stored or otherwise embodied in a machine-readable medium or in a transmission medium). Unless specifically stated otherwise, as apparent from the following discussion, it is appreciated that throughout the description, discussions utilizing terms such as "receiving", "determining", "comparing ", "enabling", "maintaining," "identifying" or the like, refer to the actions and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (electronic) quantities within the computer system's registers and memories into other data similarly represented as physical quantities within the computer system memories or registers or other such information storage, transmission or display devices. It is to be understood that the above description is intended to be illustrative, and not restrictive. Many other implementations will be apparent to those of skill in the art upon reading and understanding the above description. Although the present disclosure has been described with reference to specific example implementations, it will be recognized that the disclosure is not limited to the implementations described, but can be practiced with modification and alteration within the spirit and scope of the appended claims. Accordingly, the specification and drawings are to be regarded in an illustrative sense rather than a restrictive sense. The scope of the disclosure should, therefore, be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.

Claims

1. A computer-implemented method for generating a synthetic image based on a first image of a first imaging modality, the method comprising:receiving the first image, the first image depicting an anatomical region of a subject; and generating, using a generative mode!, a synthetic image corresponding to the first image, wherein the synthetic image corresponds to a second imaging modality.

2. The method of claim 1, wherein the first imaging modality is a magnetic resonance (MR) image.3, The method of any preceding claim, wherein the second imaging modality is any of computed tomography (CT) and cone-beam computed tomography (CBCT); and / or wherein the synthetic image is a synthetic CT (sCT) image.

4. The method of any preceding claim, wherein the generative model has been trained via a generative adversarial network (GAN).

5. The method of claim 4, wherein the GAN comprises a first generator convolutional network and a first discriminator convolutional network.

6. The method of claim 5, wherein the first generator convolutional network is trained to generate synthetic images based on input training images of the first imaging modality, and the first discriminator network is configured to discriminate between synthesized imaging data corresponding to the second imaging modality and actual imaging data corresponding to the second imaging modality.

7. The method of any preceding claim, wherein the first image is a three-dimensional (3D) image comprising a plurality of voxels; and wherein generating the synthetic image comprises providing, as an input to the generative model, an image sample comprising a plurality of consecutive 2D image patches; wherein the consecutive 2D image patches in the image sample are arranged perpendicular to a first image axis and each 2D image patch depicts a 2D region in the first image.

8. The method of claim 7, the method further comprising: providing a first set of image samples as an input to the generative model to obtain a model output for each image sample, and aggregating the model outputs to generate the synthetic image.

9. The method of claim 8, wherein the first set of image samples comprises a first image sample with a first plurality of consecutive 2D image patches and a second image sample with a second plurality of consecutive 2D image patches, wherein the first and second plurality of consecutive 2D image patches overlap with one another such that the voxels of a first subset of voxels are comprised within both the first and the second image sample.

10. The method of claim 8 or claim 9, further comprising providing a second set of image samples to the generative model., wherein the image samples in the second set of image samples each comprise consecutive 2D image patches arranged perpendicular to a second image axis.

11. The method of claim 10, wherein the second set of image samples comprises a third image sample comprising a third plurality of consecutive 2D image patches, wherein the third plurality of consecutive 2D image patches and the first plurality of consecutive 2D image patches overlap with one another such that the voxels of a second subset of voxels are comprised within both the first and the third image sample.

12. The method of any of ciaims 8 to 11, wherein the method further comprises: subdividing the first image into the first and / or the second set of image samples.

13. The method of any of claims 8 to 12, wherein the model outputs for each image sample comprise a synthetic CT value for each voxel of the image sample.

14. The method of claim 13, wherein aggregating the model outputs comprises, for each voxel of the first and / or the second image samples, calculating an average of the synthetic CT values available for that voxel.

15. The method of claim 14, wherein the average is a weighted average, the weight of each synthetic CT value being based on a proximity from a centre of its image sample.

16. The method of any preceding claim, wherein the generative model comprises a ResUnet generator.

17. A computer-implemented method for training a generative model to generate a synthetic image, the method comprising performing a process comprising:receiving training data comprising images of a first imaging modality and images of a second imaging modality: andproviding, as an input to a generative mode! in a generative adversarial network (GAN), imaging data based on the images of the first imaging modality, and receiving estimated synthetic imaging data as an output from the generative mode!;determine, using a discriminative model in the GAM, at least one difference between the estimated synthetic imaging data and real imaging data based on the images of the second imaging modality; andupdate one or more parameters of the generative model based on the determined at least one difference;the method further comprising iteratively performing the process until, when a stopping criterion is reached, the method comprises outputting the trained model;wherein, after being trained, the generative model is configured to receive an image of the 5 first modality as an input, and output a synthetic image corresponding to the second imagingmodality.

18. A system for generating a synthetic image, the system comprising:processing circuitry comprising at least one processor; and10 a computer-readable medium comprising instructions, which when executed by the at feastone processor, cause the processor to perform the method of any preceding claim.

19. A computer-readable medium comprising computer-readable instructions which, when executed by at least one processor, cause the at least one processor to perform the method 15 of any of claims 1 to 17.

Citation Information

Patent Citations

  • Neural network for generating synthetic medical images

    US20190362522A1

  • Systems and methods for image processing

    US20200211209A1

  • Apriori guidance network for multitask medical image synthesis

    US20230031910A1

  • System and method for CT image synthesis from MRI using generative sub-image synthesis

    WO2015175852A1