Conditional diffusion-based image reconstruction

GB2636718BActive Publication Date: 2026-08-21ELEKTA AB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
GB2023019531
Authority / Receiving Office
GB · GB
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2026-08-21
Estimated Expiration
2043-12-19

Smart Images

  • Figure 00000001_0000
    Figure 00000001_0000
  • Figure 00000001_0001
    Figure 00000001_0001
  • Figure 00000001_0002
    Figure 00000001_0002
Patent Text Reader

Abstract

A method for reconstructing a 3D image comprises: receiving 2D image data 260, the 2D image data being associated with one or more 2D medical images; obtaining a latent variable, the latent variable i
Need to check novelty before this filing date? Find Prior Art

Description

Field The present disclosure relates to 3D image reconstruction. More specifically, the present disclosure relates to a method for 3D image reconstruction, a method fortraining a conditional diffusion-based denoising network for use in 3D image reconstruction, and devices and computer-readable media of the same. Background Medical image analysis plays an important role in many medical aspects, including radiotherapy. Radiotherapy is commonly used for the treatment of tumours, such as cancer. Image registration is a fundamental component of medical image analysis, playing a pivotal role in various applications such as image guidance, motion tracking, segmentation, dose accumulation, and image reconstruction. A key objective of medical image registration is to determine optimal spatial transformations that align anatomical structures between two or more images. Said alignment is important for myriad reasons, including ensuring that a radiation dose is optimally delivered to a tumour, and to mitigate the risk of delivering radiation to healthy tissue. By improving medical image registration, the accuracy and precision of treatments such as radiotherapy, radiosurgery, endoscopy and the like are improved. Methods for registering images of the same modality (called unimodal) often differ from those dealing with images of different modalities (called multimodal). Their flexibility and degree of freedom categorize methods into rigid, affine, and deformable image registration. The dimensionality of images is also significant. 3D to 2D registration involves a 2D image which is represented as a slice (e.g. ultrasound, cine-MRI, or the like) and excluding projection data (e.g. x-ray). In particular, 3D to 2D registration methods are challenging due to the sparse information obtained in the fixed 2D image domain, which cause complications in the mapping of a moving 3D image volume. Due to the higher temporal resolution requirements, spatial resolution often suffers, resulting in a sparser representation of the intra-interventional image, typically one or a few 2D images. Current methods for reconstructing 3D images from 2D intrafractional images include estimating a deformable vector field (DVF) from a static pre-treatment image to a time sequence and reconstructing an image by aligning the static image with the desired time point. There is therefore a need to provide reliable 3D-2D image registration methods, for example for reconstructing a full 3D anatomy from an intra-interventional image. Implementations of the present disclosure seeks to address these and other problems encountered in the prior art. Summary Aspects and features of the present invention are described in the accompanying claims. According to an aspect, the present disclosure provides a computer-implemented method for reconstructing a 3D image, the method comprising: receiving 2D image data, the 2D image data being associated with one or more 2D medical images; obtaining a latent variable, the latent variable initially being a noisy variable; reducing a dimensionality of the 2D image data to obtain encoded image data; obtaining, based on the encoded image data, latent conditional information, the latent conditional information comprising a latent representation of the 2D image data; performing a denoising process in latent space using a trained conditional diffusion-based denoising network; wherein performing the denoising process comprises, for each iterative step of a plurality of iterative steps: inputting, into the trained denoising network, the latent variable, the latent conditional information, and a noise rate for said iterative step; estimating a noise level for said iterative step using the latent conditional information, the latent variable and the noise rate for said iterative step; and updating the latent variable for the next iterative step using the estimated noise level; and after performing the denoising process, decoding the latent variable to obtain 3D image data. According to a further aspect, a computer-readable medium is provided which comprises instructions that, when executed by one or more processors, cause the processors to carry out any of the methods disclosed herein. According to a further aspect, the present disclosure provides a computing device comprising the computer-readable medium. According to a further aspect, the present disclosure provides a computer-implemented method of training a conditional diffusion-based denoising network for use in image reconstruction, the method comprising: obtaining ground truth 3D image data; reducing a dimensionality of the ground truth 3D image data to obtain encoded ground truth 3D image data; obtaining a latent variable, the latent variable initially being a latent representation of the encoded ground truth 3D image data in latent space; performing a diffusion process to the latent variable, the diffusion process comprising, at each iterative step of a plurality of iterative steps, updating the latent variable by adding noise to the latent variable in accordance with a variance schedule, the variance schedule defining a noise rate and a noise level for each iterative step; obtaining ground truth 2D image data corresponding to the ground truth 3D image data; reducing a dimensionality of the ground truth 2D image data to obtain encoded 2D image data; obtaining, based on the encoded ground truth 2D image data, latent conditional information, the latent conditional information comprising a latent representation of the encoded ground truth 2D image data; training the conditional diffusion-based denoising network by: for each of a plurality of iterative steps: inputting, into the denoising network, the latent variable, the latent conditional information, and a noise rate for said iterative step, the noise rate for said iterative step being known from the variance schedule; estimating a noise level for said iterative step using the latent conditional information, the latent variable and the noise rate for said iterative step; updating the latent variable for the next iterative step using the estimated noise level; and minimising a loss function based on comparing the estimated noise level to the noise level according to the variance schedule. According to a further aspect, a computer-readable medium is provided which comprises instructions that, when executed by one or more processors, cause the processors to carry out any of the training methods disclosed herein. According to a further aspect, the present disclosure provides a computing device comprising the computer-readable medium. Brief Description of the Drawings Specific implementations are described below by way of example only and with reference to the accompanying drawings in which: Fig. 1 depicts a 3D image and several 2D views thereof; Fig. 2 depicts an architecture for training a model according to embodiments of the present disclosure; Fig. 3 depicts an architecture for reconstructing 3D image data according to embodiments of the present disclosure; Fig. 4 illustrates a diffusion-based denoising network according to embodiments of the present disclosure; Fig. 5 illustrates a transformer block of a diffusion-based denoising network according to embodiments of the present disclosure; Figs. 6a and 6b depict flowcharts of a method of training a model according to embodiments of the present disclosure; Figs. 7a and 7b depict flowcharts of a method of reconstructing a 3D image according to embodiments of the present disclosure; Fig. 8 depicts a block diagram of one implementation of a radiotherapy system according to embodiments of the present disclosure; Fig. 9 depicts a block diagram of one implementation of a computer-readable medium according to the present disclosure. Detailed Description Overview Aspects of the disclosure will be described below. In overview, and without limitation, the present application relates to a method for reconstructing a 3D image from one or more 2D images (or image data corresponding to one or more 2D images), and a method of training a conditional diffusion-based denoising network for use in 3D image reconstruction. The methods described herein may improve the accuracy of 3D image reconstruction from sparse 2D views. Fig. 1 depicts an exemplary 3D image and several corresponding 2D views. As shown in Fig. 1, a 3D image 1 of a liver is shown. Also shown in Fig. 1 are two 2D slices 2, 3 of the liver depicted in 3D image 1.2D slice 2 may be a 2D slice of the 3D image 1 in the coronal orientation. 2D slice 3 may be a 2D slice of the 3D image 1 in the sagittal plane. The 2D slices 2, 3 may be taken from the 3D image 1, for example when training the denoising network (as will be described with reference to Figs. 2 and 6). The 2D slices may be, for example, cine-MRI or ultrasound images, but are not limited thereto. Reference is made throughout to 2D and 3D images. It is to be understood that in all methods and embodiments, a 2D or 3D image may be used to generate a respective 2D or 3D DVF for use in the methods described herein. Fig. 2 depicts an architecture fortraining a model according to embodiments of the present disclosure. As illustrated in Fig. 2, ground truth 3D image data 210 is provided. In some embodiments, the ground truth 3D image data 210 is a 3D DVF. In some embodiments, the ground truth 3D image 210 may be a 3D medical image, such as a 3D MRI image, a 3D CT image or the like, and a ground truth 3D DVF may be generated from the ground truth 3D image 210. 3D images, such as 3D DVFs, have a high dimensionality. In some implementations, the dimensionality of the ground truth 3D image data 210 may be reduced. The dimensionality of the ground truth 3D image data may be reduced using a linear dimensional reduction technique, such as dynamic mode decomposition (DMD), or a non-linear dimensional reduction technique, such as a variational autoencoder (VAE). The ground truth 3D image data may be encoded using an encoder 220 to obtain a latent variable, z0 230. In some implementations, the encoder 220 may be a linear encoder, and may be based on principal component analysis (PCA). The PCA is an effective technique for compressing high-dimensionality image data, such as DVFs, and demonstrates a significant cumulative variance capture in a small number of components. As such, the latent variable z0 230 may be given through the projections of the initial C principal components, for example as shown in the below equation: z0 = [z^,... ,z^c)] e Kc, z^1 = (xp3D - p^wj In the above equation, p^ e [Rdim^3D represents the estimated mean of the 3D image data, and a)c e [Rdimy>3D denotes the principal component associated with c-th highest eigenvalue, Ac. Given z0, the image data may be reconstructed using the estimated mean p^ and the principal components cdc. c Wc C=1 Noise is added to the ground truth 3D image data by means of a diffusion process 240. Reducing the dimensionality of the ground truth 3D image data using, for example, encoder 220 may enable the diffusion process 240 to be run in the latent space. During training of the conditional diffusion-based denoising network, the diffusion process 240 may be run in the latent space. During the diffusion process 240, noise is added to the latent variable in an iterative manner. This may also be referred to as the forward process 240. That is, for each of a plurality of iterative steps, an amount of noise is added to the latent variable. At each iterative step t, an amount of noise is added to the latent variable zt_r to obtain latent variable zt for T steps. As such, the latent variable z becomes increasingly noisy with every step. The noise added at each step may be Gaussian noise. The amount of noise added for a step t may be referred to as the noise level, et for step t. The noise rate, ? / t, at each step may be defined by a variance scheduler ( / ^-scheduler). This addition of noise may define a Markov chain. For each step in the forward process, the sample (e.g. zt) loses distinguishable features, and, assuming Gaussian noise is added at each step, the sample (e.g. zr) may approach an isotropic Gaussian distribution. The noisy sample zT 250, generated by the final addition of noise (e.g. step T), is thus the output of the diffusion process 240. It is to be noted that, in some embodiments, the generation of a noisy sample zT corresponding to a ground truth 3D image data may be done previously. That is, the training method itself may comprise accessing previously generated noisy samples corresponding to ground truth 3D image data, rather than generating the noisy samples within the training method. In the training process, the noisy sample zT 250 generated by the diffusion process 240 is input to the conditional diffusion-based denoising network 300. One or more 2D images, or DVFs, 260 corresponding to the ground truth 3D image data (e.g. ground truth 3D images or ground truth 3D DVFs) are also used train the denoising network 300. The one or more 2D image data 260 may be referred to as conditional information. The one or more 2D images or DVFs 260 may correspond to one or more 2D slices of a 3D ground truth image, in one or more orientations. For example, one or more 2D slices may be in the same orientation. In some examples, a plurality of 2D slices may be in two or more orientations (such as transverse and coronal orientations, transverse and sagittal orientations, or coronal and sagittal orientations). The 2D images or DVFs 260 may also be referred to as sparse representations. The dimensionality of the one or more 2D images or DVFs 260 may be reduced using any dimensionality reduction technique. In particular, the dimensionality of the one or more 2D images or 2D DVFs may be reduced using a linear dimensional reduction technique, such as dynamic mode decomposition (DMD), or a non-linear dimensional reduction technique, such as a variational autoencoder (VAE). For example, the one or more 2D images or DVFs 260 may be encoded using encoder 220a. Encoder 220a may be identical to, or the same as, encoder 220. Encoder 220a may be a linear encoder, and may use principal component analysis. Encoder 220a may be a separate entity as encoder 220, but may be functionally similar, as described above. The reduced dimensionality of the one or more 2D images or DVFs may be the same as, or compatible with, the reduced dimensionality of the ground truth 3D image data. Encoder 220a may reduce the dimensionality of the 2D image data 260. In some implementations, the same principal components that are used to reduce the dimensionality of the ground truth 3D image data 210 may be also used to derive a latent representation (e.g. a representation in the latent space) of the 2D image data 260 and to reduce the dimensionality of the 2D image data 260. In other implementations, however, a different embedding may be used to reduce the dimensionality of the 2D image data. Zcond = ^cond “ / ¾) O where ip^id e Bdim^3D denotes the sparse representation, O represents the element-wise multiplication and m e IRdim^3D,mj e {0,1} is a binary mask indicating the location of the sparse information. Throughout this disclosure, variables z are used to indicate variables in the latent space. The outputs of encoder 220a (e.g. the latent representations of the 2D image data 260) are input into the denoising network 300, as will be described with reference to Figs. 4 and 5. The denoising network 300 is configured to denoise the latent representation zT using the latent representations of the 2D image data 260 (zeond) and time-related information, such as a time scalar t, or noise rate information r / t. The noise rate information r)t may be defined by the ^-scheduler used in the forward process to add noise. The denoising network 300 computes estimates of the noise level et at each step t (from 0 to T). Most diffusion models use a U-Net backbone which is effective in capturing both local and global features in imaging data. However, the latent variables used herein exhibit a lack of spatial dependence. Consequently, the denoising network 300 comprises a transformer-based multiplayer perceptual network (MPL), which will be described in more detail with reference to Figs. 4 and 5. At each step t from T to 0, the denoising network 300 generates a denoised latent representation, e.g. from z'T_1 270 to z'o which is then used by the denoising network 300 in the subsequent step t - 1, to ultimately generate latent representation z'o 280. Latent representation z'o 280 may be input to a decoder 290 to reconstruct an estimated 3D image data 295, such as an estimated 3D DVF or an estimated 3D image, from the latent representation z'o 280. The decoder 290 may correspond to the encoder 220 - e.g. the decoder 290 may correspond to the PCA used by the encoder 220 in order to reconstruct the estimated 3D image data 295. The estimated 3D image data 295 may be compared to the ground truth 3D image data 210. The denoising network 300 is trained by minimising a loss function. The loss function may be based on the dimensionality reduction technique used to reduce the dimensionality of the conditional information and / or the ground truth 3D image data. For example, given the PCA model used to derive the latent representations of the ground truth 3D image data and the 2D image data 260, the diffusion loss function may be modified by introducing a weighted norm to the noise error: ^our ^Z0'zcond't[^£t zcond'?lt)lla> ] where eo(c) a 1 is monotonically decreasing for larger components. The first dimensions, e.g. principal components, may contain higher variances of the data, and have a greater impact on the final estimation. By weighting the loss, the noise error in lower dimensions is penalised more heavily. Minimising the loss function comprises tuning parameters and hyperparameters of the denoising network 300. Once minimised, the corresponding parameters and hyperparameters may be stored and the denoising network 300 is trained. The loss function minimisation may be achieved using any known optimisation method. Fig. 3 depicts an architecture for reconstructing 3D image data according to embodiments of the present disclosure. The diagram of Fig. 3 relates closely to that of Fig. 2. However, in Fig. 3, the denoising network 300 is already trained. That is, the architecture depicted in Fig. 3 represents the system for generating a 3D image (or a 3D DVF) from one or more 2D images (or one or more 2D DVFs). For the sake of succinctness, components described with reference to Fig. 2 will be merely referred to here, and will not be described in full again. The denoising network 300 takes as input both a noisy variable zr 250 and latent representations zcond of 2D image data 260. The latent representations zcond of 2D image data 260 may also be referred to as conditional information zcond. The noisy latent variable zT may be generated from noise, such as Gaussian noise. 2D image data 260 are used to generate estimated 3D image data 295. For example, one or more 2D images 260 may be used to generate an estimated 3D image. As another example, one or more 2D DVFs may be used to generate an estimated 3D DVF. As discussed previously with reference to Fig. 2, the 2D image data 260 may correspond to one or more 2D slices, such as a cine-MRI or an ultrasound, in one or more orientations. That is, the one or more 2D slices may be in the same orientation, or may correspond to multiple different orientations. The dimensionality of the 2D image data 260 is reduced using any dimensionality reduction technique, for example by using an encoder 220a. Examples of dimensionality reduction techniques include a linear dimensional reduction technique, such as dynamic mode decomposition (DMD), or a non-linear dimensional reduction technique, such as a variational autoencoder (VAE). The 2D image data 260 may be input into encoder 220a, described with reference to Fig. 2. Encoder 220a may reduce the dimensionality of the 2D image data 260 and may generate latent representations zcond. As described previously, encoder 220a may be a linear encoder and may use principal component analysis. The encoder 220a may reduce the dimensionality of the 2D image data 260 to have dimensions compatible with or the same as the dimensionality of the noisy latent variable zT. The noisy variable zT 250 and latent representations zcond of 2D image data 260 are input into the denoising network 300, which is trained to estimate a noise level at each step t of steps T to 0 using the noisy latent variable zt, the latent representations zcond and noise rate information r / t. The noise rate information r)t may be defined by the ^-scheduler. The functioning of the denoising network 300 will be described in more detail with reference to Figs. 4 and 5. At each step t from T to 0, the denoising network 300 generates a denoised latent representation, e.g. from z'T_1 270 to z'o which is then used by the denoising network 300 in the subsequent step t - 1, to ultimately generate latent representation z'o 280. Latent representation z'o 280 may be input to a decoder 290 to reconstruct an estimated 3D image data 295 from the latent representation z'o 280. The estimated 3D image data 295 may be output to a user, and / or stored and / or saved. Fig. 4 illustrates a diffusion-based denoising network 300 according to embodiments of the present disclosure. It is to be understood that the architecture described herein is merely exemplary, and the methods described with reference to Figs. 6 and 7 may be performed on other architectures. The diffusion-based denoising network 300 may have any network architecture which estimates a noise level based on the noisy latent variable zt, the noise rate r)t and the conditional information zcond. Fig. 4 shows a plurality of blocks 330-1 to 330-n. Each block takes two inputs - one relates to the noisy latent variable and one related to the current step t. The first input 310 may be a combination of the noisy latent variable zT 250 and the latent conditional information zcond (e.g. corresponding to 2D image data 260, such as 2D images or 2D DVFs). The first input 310 may be generated using concatenation-based conditioning such as in-context concatenation, or in-context conditioning, for aligning the conditional information zcond with the noisy latent variable zt. In some embodiments, conditional scaling may be used instead of concatenationbased conditioning. The second input 320 may be based on the noise rate r)t given from the ^-scheduler for each step t. The noise rate r)t may be given by r)t = ^ / 1 - where = nLi at and at = ~ Pt- The noise rate r)t may be embedded and conditioned in the model using, for example, Feature-wise Linear Modulation layers (FiLM): FiLM(zt | rit) = zt (1 + / (¾)) + for a feature variable zt. Other methods of embedding the noise rate r)t which may be used instead of FiLM include concatenation, scaling, shifting, cross-attention and the like. Each block 330 may be a transformer block, also referred to as a transform block. The transformer block is described in further detail with reference to Fig. 5. Each block 330 may update the latent variable information (e.g. the combination of zt and zeond) and may pass the updated latent variable information to the subsequent block. After the last block 330-n, the updated latent variable information (e.g. the latent variable information updated by each block, including block 330-n) may be passed to a linear layer 340 which is configured to estimate an amount of noise et. The estimated noise et approximates the noise added during the forward process and denoises the noisy latent variable zt by removing the estimated noise et from the noisy latent variable zt to obtain the latent variable for the subsequent step, zt_r. Fig. 5 illustrates a transformer block 330 of a diffusion-based denoising network according to embodiments of the present disclosure. It is to be understood that the architecture described herein is merely exemplary, and the methods described with reference to Figs. 6 and 7 may be performed on other architectures. The diffusion-based denoising network 300 may have any network architecture which estimates a noise level based on the noisy latent variable zt, the noise rate r)t and the conditional information zcond. A transformer block 330 may comprise a plurality of Res-Net (residual network) blocks 332, which may include noise rate conditioning, followed by a multi-head self-attention layer 334. Each block 300 may comprise a MLP 332 including a FiLM conditioning layer followed by a multi-head self-attention layer 334. As illustrated in Fig. 5, the MLP 332 (e.g. the Res-Net blocks) may comprise a plurality of sequential components and / or techniques, such as linear layer 332-1, dropout 332-2, layer normalisation 332-3, element-wise multiplication 332-4, addition element 332-5, SiLU (sigmoid linear unit) activation layer 332-6, and a further addition element 332-7. In particular, the output of the layer normalisation 332-3, which may be referred to as feature variable zt, may be combined with the information derived from the noise rate r)t in the form of zt(l + y(? / t)) by the element-wise multiplication element 332-4, using the FiLM layer within the MLP 332. More specifically, this term is determined by embedding the noise rate information using the FiLM layer. Additionally, the first addition element 332-5 may combine the output of the element-wise multiplication element 332-4 with v(? / t), also using the FiLM layer within the MLP 332 (e.g. by embedding the noise rate information). Although the use of FiLM conditioning is described in detail herein, it is to be understood that the disclosure is not limited thereto. Other conditioning techniques are also envisaged, and provide a similar function as the FiLM conditioning. The output of the SiLU activation layer 332-6 may be combined with the combined latent information (e.g. the output of block 310) at the second addition element 332-7. The output of the second addition element 332-7 may be passed to the next Res-Net block 332. After the last Res-Net block 332, the output may be passed to the self-attention block 334. The embedded noise rate information, e.g. the output of block 320, may subsequently be processed by a SiLU activation layer 332-8 and a linear layer 332-9. The multi-head self-attention layer 334 may comprise a plurality of components as illustrated. The plurality of components may comprise multiple linear layers arranged in parallel 334-1 a, 334-1 b, 334-1c, a scaled dot-product attention mechanism 334-2, a concatenation layer 334-3, a linear layer 334-4, a dropout 334-5, an addition element 334-6 and a layer normalisation 334-7. The addition element 334-6 may combine the output from the Res-Net blocks 332 with the output from the dropout 334-5. The output from the addition element 334-6 may be passed as input to the layer normalisation 334-7. The output of the layer normalisation 334-7 may be passed to the subsequent block 330-2. The output of the layer normalisation 334-7 of the final block 330-n may be passed as an input to the linear layer 340, which may then be used to derive the estimated amount of noise et as defined previously. Fig. 6a depicts a flowchart of a method 600 of training a model according to embodiments. The method 600 may be performed by a computing device 810. Method 600 comprises, in an operation entitled “OBTAINING GROUND TRUTH 3D IMAGE DATA”, obtaining 610 ground truth 3D image data corresponding to at least one ground truth 3D image or DVF. Obtaining 610 the ground truth 3D image data may comprise obtaining a 3D image, or generating a 3D DVF from a 3D image, such as an MRI image or a CT image. The at least one ground truth 3D DVF may be obtained from local storage, remote storage, cloud storage, or the like. The ground truth 3D image data may be obtained by generating, either by the computing device 810 or by another device or system, 3D image data from a 3D image retrieved from local storage, remote storage, cloud storage or the like. Method 600 further comprises, in an operation entitled “REDUCING DIMENSIONALITY”, reducing 620 the dimensionality of the ground truth 3D image data to obtain encoded ground truth 3D image data. It is to be understood that the term “encoded” ground truth 3D image data may refer to ground truth 3D image data whose dimensionality has been reduced, regardless of whether an encoder was used. In some implementations, the dimensionality of the ground truth 3D image data may be reduced by an encoder. The dimensionality of the ground truth 3D image data may be reduced using any dimensionality reduction techniques, for example a linear dimensional reduction technique, such as dynamic mode decomposition (DMD), or a non-linear dimensional reduction technique, such as a variational autoencoder (VAE). The encoder may be a linear encoder. The (e.g. linear) encoder may be based on principal component analysis (PCA). Method 600 further comprises, in an operation entitled “OBTAINING LATENT VARIABLE”, obtaining 630 a latent variable. The latent variable is initially a latent representation of the encoded ground truth 3D image data in latent space. Obtaining 630 the latent variable in the first instance may also be referred to as initialising the latent variable. Method 600 further comprises, in an operation entitled “PERFORMING A DIFFUSION PROCESS TO THE LATENT VARIABLE”, performing 640 a diffusion process to the latent variable. The diffusion process comprises, at each iterative step of a plurality of iterative steps, updating the latent variable by adding noise to the latent variable in accordance with a variance schedule. The variance schedule defines a noise rate for each iterative step. The noise level may be the amount of noise added to the latent variable for a particular iterative step. The number of iterative steps may be known. The iterative steps may be referred to as time steps. The variance schedule may define the number of iterative steps. By adding noise iteratively in this manner, the latent variable becomes more noisy with each iterative step. The noise added to the latent variable may be Gaussian noise. Method 600 further comprises, in an operation entitled “OBTAINING GROUND TRUTH 2D IMAGE DATA”, obtaining 650 ground truth 2D image data corresponding to the ground truth 3D image data. The ground truth 2D image data may be one or more 2D images, or one or more 2D DVFs. For example, the ground truth 2D image data may correspond to 2D slices of the 3D image corresponding to the ground truth 3D image data. The ground truth 2D image data may correspond to 2D slices in one or more orientations. In some embodiments, the 2D slices may be in two or more orientations, such as transverse and coronal, transverse and sagittal, or coronal and sagittal. In some embodiments, the 2D slices may be in a single orientation. As a further example, the ground truth 2D image data may be registered against the ground truth 3D image data. Method 600 further comprises, in an operation entitled “REDUCING DIMENSIONALITY OF GROUND TRUTH 2D IMAGE DATA”, reducing 660 the dimensionality of the ground truth 2D image data to obtain encoded ground truth 2D image data. Although the term “encoded” ground truth 2D image data is used, it is not necessary that an encoder is used. The encoded ground truth 2D image data may refer to the ground truth 2D image data with reduced dimensionality. The dimensionality of the ground truth 2D image data may be reduced using any known dimensionality reduction techniques. Examples of dimensionality reduction techniques include a linear dimensional reduction technique, such as dynamic mode decomposition (DMD), or a non-linear dimensional reduction technique, such as a variational autoencoder (VAE). The encoder may be a linear encoder. The (e.g. linear) encoder may be based on principal component analysis (PCA). The encoder used to reduce the dimensionality of the ground truth 2D image data may be the same encoder as the encoder used to reduce the dimensionality of the ground truth 3D image data. The dimensionality of the ground truth 2D image data may be compatible with the dimensionality of the latent variable based on the ground truth 3D image data. Method 600 further comprises, in an operation entitled “OBTAINING LATENT CONDITIONAL INFORMATION”, obtaining 670 latent conditional information based on the encoded ground truth 2D image data. The latent conditional information may comprise a latent representation of the encoded ground truth 2D image data. Method 600 further comprises, in an operation entitled “TRAINING THE CONDITIONAL DIFFUSIONBASED DENOISING NETWORK”, training 680 the conditional diffusion-based denoising network 300. The sub-method of training the conditional diffusion-based denoising network is described in more detail with reference to Fig. 6b. As illustrated in Fig. 6b, operation 680 comprises sub-operations 682 (“INPUTTING THE LATENT VARIABLE, LATENT CONDITIONAL INFORMATION AND NOISE RATE”), 684 (“ESTIMATING NOISE LEVEL”), 686 (“UPDATING LATENT VARIABLE”) and 686 (“MINIMISING LOSS FUNCTION”) are performed iteratively, e.g. for each iterative step of a plurality of iterative steps. Operation 680 comprises, in a sub-operation entitled “INPUTTING THE LATENT VARIABLE, LATENT CONDITIONAL INFORMATION AND NOISE RATE”, inputting 682 the latent variable, the latent conditional information, and a noise rate for said iterative step, into the denoising network. The noise rate for said iterative step may be known from the variance schedule. The number of iterative steps of the plurality of iterative steps may be known from the variance schedule. Operation 680 comprises, in a sub-operation entitled “ESTIMATING NOISE LEVEL”, estimating 684 a noise level for the iterative step using the latent conditional information, the latent variable and the noise rate for the iterative step. The noise level may be estimated by a linear layer of the denoising network after the last transformer block of the denoising network, as previously described. Operation 680 comprises, in a sub-operation entitled “UPDATING THE LATENT VARIABLE”, updating 686 the latent variable for the next iterative step using the estimated noise level from sub-operation 684. For example, updating the latent variable may comprise subtracting or removing the estimated noise level from the latent variable. Operation 680 further comprises, in a sub-operation entitled “MINIMISE LOSS FUNCTION”, minimising 688 a loss function based on a comparison of the estimated noise level and the noise level according to the variance schedule. The loss function may be a diffusion loss function as described previously in this disclosure. A method of using a conditional diffusion-based denoising network trained in accordance with method 600 is envisaged. Fig. 7a depicts a flowchart of a method 700 of reconstructing a 3D image according to embodiments of the present disclosure. The method 700 may be performed by a computing device 810. Method 700 comprises, in an operation entitled “RECEIVING 2D IMAGE DATA”, receiving 710 2D image data associated with one or more 2D medical images. The 2D image data may comprise one or more 2D DVFs or one or more 2D images. Method 700 comprises, in an operation entitled “OBTAINING LATENT VARIABLE”, obtaining 720 a latent variable. The latent variable is initially a noisy variable. In some embodiments, obtaining the latent variable may comprise generating the latent variable, for example from Gaussian noise. Method 700 comprises, in an operation entitled “REDUCING DIMENSIONALITY OF THE 2D IMAGE DATA”, reducing 730 a dimensionality of the 2D image data to obtain encoded 2D image data. Although the term “encoded” 2D image data is used, it is not necessary that an encoder is used. The encoded 2D image data may refer to the 2D image data with reduced dimensionality. The dimensionality of the 2D image data may be reduced using any known dimensionality reduction techniques. Examples of dimensionality reduction techniques include a linear dimensional reduction technique, such as dynamic mode decomposition (DMD), or a non-linear dimensional reduction technique, such as a variational autoencoder (VAE). The encoder may be a linear encoder. The (e.g. linear) encoder may be based on principal component analysis (PCA). Method 700 comprises, in an operation entitled “OBTAINING LATENT CONDITIONAL INFORMATION”, obtaining 740 latent conditional information based on the encoded 2D image data. The latent conditional information comprises a latent representation of the 2D image data. Method 700 comprises, in an operation entitled “PERFORMING DENOISING PROCESS”, performing 750 a denoising process in latent space using a trained conditional diffusion-based denoising network. The trained conditional diffusion-based denoising network may be the denoising network trained according to method 600. Details of the denoising process will now be described with reference to Fig. 7b. As shown in Fig. 7b, performing 750 the denoising process comprises several sub-operations performed for a plurality of iterative steps. Performing 750 the denoising process comprises, for each iterative step of the plurality of iterative steps, sub-operations 752 (“INPUTTING LATENT VARIABLE, LATENT CONDITIONAL INFORMATION, NOISE RATE”) and 754 (“ESTIMATING A NOISE LEVEL”) and 756 (“UPDATING THE LATENT VARIABLE”). Operation 750 comprises, in a sub-operation entitled “INPUTTING LATENT VARIABLE, LATENT CONDITIONAL INFORMATION, NOISE RATE”, inputting 752 the latent conditional information, the latent variable and a noise rate for the iterative step into the trained conditional diffusion-based denoising network. The noise rate may be known from a variance schedule. The variance schedule may indicate a noise rate as described above. Operation 750 comprises, in a sub-operation entitled “ESTIMATING NOISE LEVEL”, estimating 754 a noise level for the iterative step using the latent conditional information, the latent variable and the noise rate for the iterative step. The noise level may be estimated as described previously in this disclosure. Operation 750 comprises, in a sub-operation entitled “UPDATING THE LATENT VARIABLE”, updating 756 the latent variable for the next iterative step using the estimated noise level. For example, updating the latent variable may comprise denoising the latent variable (e.g. subtracting or removing the estimated noise level from the latent variable). Returning to Fig. 7a, after performing 750 the denoising process (e.g. after performing the suboperations 752, 754 and 756 for the plurality of iterative steps), the method 700 further comprises, in an operation entitled “DECODING LATENT VARIABLE”, decoding 760 the latent variable to obtain 3D image data. The 3D image data may be a 3D DVF or a 3D image. Decoding 760 the latent variable may be performed using a decoder. Method 700 may further comprise outputting the decoded 3D image data, for example to a screen. Additionally or alternatively, method 700 may further comprise storing the decoded 3D image data, e.g. in local storage, remote storage and / or cloud storage. Fig. 8 illustrates a block diagram of one implementation of a radiotherapy system 800. The radiotherapy system 800 comprises a computing system 810 within which a set of instructions, for causing the computing system 810 to perform any one or more of the methods discussed herein, may be executed. The computing system 810 shall be taken to include any number or collection of machines, e.g. computing device(s), that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods discussed herein. That is, hardware and / or software may be provided in a single computing device, or distributed across a plurality of computing devices in the computing system. In some implementations, one or more elements of the computing system may be connected (e.g., networked) to other machines, for example in a Local Area Network (LAN), an intranet, an extranet, or the Internet. One or more elements of the computing system may operate in the capacity of a server or a client machine in a client-server network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. One or more elements of the computing system may be a personal computer (PC), a tablet computer, a set-top box (STB), a Personal Digital Assistant (PDA), a cellular telephone, a web appliance, a server, a network router, switch or bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that machine. The computing system 810 includes controller circuitry 811 and a memory 813 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM), etc.). The memory 813 may comprise a static memory (e.g., flash memory, static random access memory (SRAM), etc.), and / or a secondary memory (e.g., a data storage device), which communicate with each other via a bus (not shown). Controller circuitry 811 represents one or more general-purpose processors such as a microprocessor, central processing unit, accelerated processing units, or the like. More particularly, the controller circuitry 811 may comprise a complex instruction set computing (CISC) microprocessor, reduced instruction set computing (RISC) microprocessor, very long instruction word (VLIW) microprocessor, processor implementing other instruction sets, or processors implementing a combination of instruction sets. Controller circuitry 811 may also include one or more special-purpose processing devices such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), network processor, or the like. One or more processors of the controller circuitry may have a multicore design. Controller circuitry 811 is configured to execute the processing logic for performing the operations and steps discussed herein. The computing system 810 may further include a network interface circuitry 818. The computing system 810 may be communicatively coupled to an input device 820 and / or an output device 830, via input / output circuitry 817. In some implementations, the input device 820 and / or the output device 830 may be elements of the computing system 810. The input device 820 may include an alphanumeric input device (e.g., a keyboard or touchscreen), a cursor control device (e.g., a mouse or touchscreen), an audio device such as a microphone, and / or a haptic input device. The output device 830 may include an audio device such as a speaker, a video display unit (e.g., a liquid crystal display (LCD) or a cathode ray tube (CRT)), and / or a haptic output device. In some implementations, the input device 820 and the output device 830 may be provided as a single device, or as separate devices. In some implementations, computing system 810 includes training circuitry 818. The training circuitry 818 is configured to train the conditional diffusion-based denoising network 300 as described with reference to Figs. 2 and 6. The model may comprise a deep neural network (DNN), such as a convolutional neural network (CNN) and / or recurrent neural network (RNN). Training circuitry 818 may be configured to execute instructions to train a model that can be used to generate a 3D DVF, as described with reference to Figs. 3, 4, 5 and 7. Training circuitry 818 may be configured to access training data and / or testing data from memory 813 or from a remote data source, for example via network interface circuitry 815. In some examples, training data and / or testing data may be obtained from an external component, such as image acquisition device 840 and / or treatment device 850. In some implementations, training circuitry 818 may be used to update, verify and / or maintain the denoising network 300. In some implementations, the computing system 810 may comprise image processing circuitry 819. Image processing circuitry 819 may be configured to process image data 880 (e.g. images, or imaging data), such as medical images obtained from one or more imaging data sources, a treatment device 850 and / or an image acquisition device 840. Image processing circuitry 819 may be configured to process, or pre-process, image data. For example, image processing circuitry 819 may convert received image data into a particular format, size, resolution or the like. In some implementations, image processing circuitry 819 may be combined with controller circuitry 811. For example, the image processing circuitry 819 may be configured to generate DVFs from medical images, such as 3D DVFs and 2D DVFs. In some implementations, the radiotherapy system 800 may further comprise an image acquisition device 840 and / or a treatment device 850, for example to obtain medical images for use in training the denoising network 300 (e.g. 3D medical images corresponding to the ground truth 3D DVF and / or 2D medical images corresponding to the sparse representations 260) and / or for use in generating a 3D DVF from sparse representations 260. That is, the image acquisition device 840 may be configured to obtain 2D slices from which 2D DVFs may be generated. The image acquisition device 840 and the treatment device 850 may be provided as a single device. In some implementations, treatment device 850 is configured to perform imaging, for example in addition to providing treatment and / or during treatment. The treatment device 850 comprises the main radiation delivery components of the radiotherapy system. Image acquisition device 840 may be configured to perform positron emission tomography (PET), computed tomography (CT), magnetic resonance imaging (MRI), for example. Image acquisition device 840 may be configured to output image data 880, which may be accessed by computing system 810. Treatment device 850 may be configured to output treatment data 860, which may be accessed by computing system 810. Computing system 810 may be configured to access or obtain treatment data 860, planning data 870 and / or image data 880. Treatment data 860 may be obtained from an internal data source (e.g. from memory 813) or from an external data source, such as treatment device 850 or an external database. Planning data 870 may be obtained from memory 813 and / or from an external source, such as a planning database. Planning data 870 may comprise information obtained from one or more of the image acquisition device 840 and the treatment device 850. The various methods described above may be implemented by a computer program. The computer program may include computer code (e.g. instructions) 910 arranged to instruct a computer to perform the functions of one or more of the various methods described above. The steps of the methods described above may be performed in any suitable order. For example, steps 610 and 660 of method 600 may be performed in any order, simultaneously or substantially simultaneously. The computer program and / or the code 910 for performing such methods may be provided to an apparatus, such as a computer, on one or more computer readable media or, more generally, a computer program product 900)), depicted in Fig. 9. The computer readable media may be transitory or non-transitory. The one or more computer readable media 900 could be, for example, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, or a propagation medium for data transmission, for example for downloading the code over the Internet. Alternatively, the one or more computer readable media could take the form of one or more physical computer readable media such as semiconductor or solid state memory, magnetic tape, a removable computer diskette, a random access memory (RAM), a read-only memory (ROM), a rigid magnetic disc, and an optical disk, such as a CD-ROM, CD-R / W or DVD. The instructions 910 may also reside, completely or at least partially, within the memory 813 and / or within the controller circuitry 811 during execution thereof by the computing system 810, the memory 813 and the controller circuitry 811 also constituting computer-readable storage media. In an implementation, the modules, components and other features described herein can be implemented as discrete components or integrated in the functionality of hardware components such as ASICS, FPGAs, DSPs or similar devices. A “hardware component” is a tangible (e.g., non-transitory) physical component (e.g., a set of one or more processors) capable of performing certain operations and may be configured or arranged in a certain physical manner. A hardware component may include dedicated circuitry or logic that is permanently configured to perform certain operations. A hardware component may comprise a special-purpose processor, such as an FPGA or an ASIC. A hardware component may also include programmable logic or circuitry that is temporarily configured by software to perform certain operations. In addition, the modules and components can be implemented as firmware or functional circuitry within hardware devices. Further, the modules and components can be implemented in any combination of hardware devices and software components, or only in software (e.g., code stored or otherwise embodied in a machine-readable medium or in a transmission medium). Unless specifically stated otherwise, as apparent from the following discussion, it is appreciated that throughout the description, discussions utilizing terms such as "receiving”, “determining”, “comparing”, “enabling”, “maintaining,” “identifying,” or the like, refer to the actions and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (electronic) quantities within the computer system's registers and memories into other data similarly represented as physical quantities within the computer system memories or registers or other such information storage, transmission or display devices. It is to be understood that the above description is intended to be illustrative, and not restrictive. Many other implementations will be apparent to those of skill in the art upon reading and understanding the above description. Although the present disclosure has been described with reference to specific example implementations, it will be recognized that the disclosure is not limited to the implementations described, but can be practiced with modification and alteration within the spirit and scope of the appended claims. Accordingly, the specification and drawings are to be regarded in an illustrative sense rather than a restrictive sense. The scope of the disclosure should, therefore, be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.

Claims

1. A computer-implemented method for reconstructing a 3D image, the method comprising: receiving 2D image data, the 2D image data being associated with one or more 2D medical images;obtaining a latent variable, the latent variable initially being a noisy variable;reducing a dimensionality of the 2D image data to obtain encoded image data; obtaining, based on the encoded image data, latent conditional information, the latent conditional information comprising a latent representation of the 2D image data;performing a denoising process in latent space using a trained conditional diffusion-based denoising network;wherein performing the denoising process comprises, for each iterative step of a plurality of iterative steps:o inputting, into the trained denoising network, the latent variable, the latent conditional information, and a noise rate for said iterative step;o estimating a noise level for said iterative step using the latent conditional information, the latent variable and the noise rate for said iterative step; ando updating the latent variable for the next iterative step using the estimated noise level; andafter performing the denoising process, decoding the latent variable to obtain 3D image data.

2. The method of claim 1, wherein the 3D image data is a 3D DVF, and / or wherein the 2D image data is 2D DVF data.

3. The method of claim 1 or claim 2, wherein the 2D image data comprises image data corresponding to a plurality of 2D images in at least two orientations.

4. The method of any one of claims 1 to 3, wherein the 2D image data comprises image data corresponding to a plurality of 2D images in a single orientation.

5. The method of any preceding claim, wherein reducing the dimensionality of the 2D image data is performed by a linear encoder.

6. The method of claim 5, wherein the linear encoder is based on principal component analysis (PCA).

7. The method of any preceding claim, wherein the noise rate for each iterative step is obtained from a variance schedule, the variance schedule specifying the noise rate for each iterative step.

8. The method of any preceding claim, wherein the 2D image data comprises image data corresponding to a respective at least one 2D Cine-MRI image.

9. The method of any preceding claim, wherein the conditional diffusion-based denoising network comprises a linear layer and a plurality of transformer blocks, each transformer block comprising a multi-layer perceptual network (MLP) including a feature-wise linear modulation (FiLM) conditioning layer, followed by a multi-head self-attention layer;wherein the FiLM conditioning layer embeds and conditions the noise rate for a current iterative step; and wherein each MLP preferably comprises a plurality of residual neural networks.

10. The method of any preceding claim, wherein inputting the latent variable and the latent conditional information into the trained denoising network comprises concatenating the latent variable and the latent conditional information.

11. A computer-readable medium comprising computer-executable instructions configured to perform the method of any preceding claim.

12. A computing device comprising the computer-readable medium of claim 11.

13. A computer-implemented method of training a conditional diffusion-based denoising network for use in image reconstruction, the method comprising:obtaining ground truth 3D image data;reducing a dimensionality of the ground truth 3D image data to obtain encoded ground truth 3D image data;obtaining a latent variable, the latent variable initially being a latent representation of the encoded ground truth 3D image data in latent space;performing a diffusion process to the latent variable, the diffusion process comprising, at each iterative step of a plurality of iterative steps, updating the latent variable by adding noise to the latent variable in accordance with a variance schedule, the variance schedule defining a noise rate and a noise level for each iterative step;obtaining ground truth 2D image data corresponding to the ground truth 3D image data;reducing a dimensionality of the ground truth 2D image data to obtain encoded 2D image data;obtaining, based on the encoded ground truth 2D image data, latent conditional information, the latent conditional information comprising a latent representation of the encoded ground truth 2D image data;training the conditional diffusion-based denoising network by:o for each of a plurality of iterative steps:■ inputting, into the denoising network, the latent variable, the latent conditional information, and a noise rate for said iterative step, the noise rate for said iterative step being known from the variance schedule;■ estimating a noise level for said iterative step using the latent conditional information, the latent variable and the noise rate for said iterative step;■ updating the latent variable for the next iterative step using the estimated noise level; and■ minimising a loss function based on comparing the estimated noise level to the noise level according to the variance schedule.

14. A computer-readable medium comprising computer-readable instructions configured to perform the method of claim 13.

15. A computing device comprising the computer-readable medium of claim 14.