Method and system for 3D modelling of internal anatomy from 2d internal anatomical images

A deep learning-based framework for 3D coronary artery reconstruction from 2D X-ray images addresses automation and accuracy issues, enabling precise vFFR calculations and reducing manual intervention in coronary interventions.

WO2026022451A1PCT designated stage Publication Date: 2026-01-29KINGS COLLEGE LONDON
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2025/051507
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-26
Filing Date
2025-07-09
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Current methods for reconstructing 3D coronary artery anatomy from 2D X-ray images are not fully automated, prone to noise and patient motion, and require manual intervention, leading to significant uncertainty in virtual fractional flow reserve (vFFR) estimates.

Method used

A novel deep learning framework using an encoder-decoder neural network with implicit field decoders and a surface-focusing loss function, combined with data synthesis, to probabilistically reconstruct 3D coronary anatomy from multiple 2D X-ray images, enabling accurate and efficient vFFR calculations without physical sensors.

Benefits of technology

The method provides non-invasive, accurate, and efficient 3D reconstruction of coronary anatomy, allowing for precise vFFR measurements, reducing reliance on manual processes and improving treatment decision-making in coronary interventions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2025051507_29012026_PF_FP_ABST
    Figure GB2025051507_29012026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a method and system that can generate a virtual three-dimensional model of internal anatomy of an imaged subject, based on a plurality of two-dimensional medical images of the internal anatomy, such as X-ray images, obtained from slightly different viewpoints. The 3D model is then obtained using a probabilistic reconstruction technique, which takes an unconditional distribution of internal anatomies, including "normal" anatomies of healthy people, as well as diseased or otherwise impossible anatomies, and uses the obtained 2D images to guide a 3D reconstruction to a maximum likelihood internal anatomy that is most supported by the obtained 2D images from the whole unconditional distribution of anatomies. The maximum likelihood 3D reconstruction can then be output as a 3D model of the internal anatomy of the subject, or can be used as the basis of 2D projections to generate reconstructed 2D images of the 3D reconstruction onto any imaging plane, thus giving reconstructed 2D images of the internal anatomy from any directional viewpoint.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] P604279PC00 Method and System for 3D Modelling of Internal Anatomy from 2D Internal Anatomical Images Technical Field The present disclosure relates to the virtual modelling or reconstruction in three dimensions of internal anatomical structures of a subject based on two-dimensional internal anatomical images of the subject. In some embodiments the two- dimensional images may be two or more x-ray images, taken from slightly different fields of view of the subject. In particular examples the anatomy is the heart of a subject, and the structure that is modelled is the coronary artery. Background to the Invention Coronary artery disease (CAD) occurs when pathological changes to the inner surface of the coronary arteries reduce the oxygen supply to the heart muscle, leading to chest pain, myocardial infarction, and eventually death. CAD is one of the highest causes of global death by non-communicable disease. At present, coronary artery disease is managed with a combination of medical therapy, surgery, andpercutaneous coronary intervention [1], the latter of which is a minimally invasiveprocedure that is amongst the most frequently performed in the UK and elsewhere [2]. It involves the insertion via catheter of a metal stent into the lumen of the coronary artery at the site of the lesion to restore blood flow through the artery. The effectiveness of such an intervention depends on correctly locating, evaluating, and treating one or more flow-limiting stenoses or, alternatively, deciding to exit the procedure and continue with medical therapy only [3]. The clinical team typically rely on two sensors to aid their decision-making in these tasks: first, a 2D X-ray imaging apparatus that captures and displays video for real-time visual inspection; and second, a pressure-sensitive probe in a catheter that provides readouts of local blood pressure from which the fractional flow reserve (FFR), a functional measure of disease severity, can be calculated. Using a combination of these two sensors is the current gold standard method for real-time decision-making during percutaneous coronary intervention, but due to clinical practicalities pressure-wire measurement is usually neglected [4], which significantly impacts the effectiveness of treatment [5]. However, a well-chosen set of 2D images captured during the procedure contains sufficient geometrical information to recover the 3D anatomy which can in turn be used to estimate the P604279PC00 pressure reading without the need for a physical wire measurement. This computed haemodynamic measurement is called the virtual fractional flow reserve (vFFR) and— when using either 3D CT images or manually-processed coronary angiograms—hasbeen demonstrated to be an effective indicator for clinical decision making [6, 7, 8,9, 10] and has become routine practice in the UK

[0011] and elsewhere

[0012] . However,there is currently no clinically adopted method to calculate vFFR directly from 2D X-ray angiography images. Prior Art Classical angiographic reconstructionThis task of reconstruction is classically treated as a geometric problem in which a3D-to-2D image formation process must be ‘undone’. It is in this way considered asan inverse problem in which an underlying anatomy produces observations in theform of several angiography images with known camera pose and hence a uniquereconstruction can be determined using physical and geometric constraints.Precisely, those constraints respectively enforce the straight-line trajectory of image-forming photons, and the deterministic projections of a 3D shape onto a set of known2D image planes. The most widely employed X-ray reconstruction method incardiology is tomographic back-projection, where a dense set of 2D views isreconstructed into a 3D attenuation volume. This is the foundation of ComputedTomography (CT) imaging, but has also been applied to rotational angiography,usually by first removing 2D noise by applying an automatic vesselness filter such astubularity

[0013] or manual maps such as a binary vessel segmentation

[0014] anddistance to centreline field

[0015] , then backprojecting the set of filtered views into a3D volume. The relative sparsity of views in rotational angiography when comparedto CT means the reconstructed volume is noisy— compounded by cardiac,respiratory, and table motion—and requires subsequent 3D filtering to isolate thevessels. The direct application of back-projection into an attenuation volume meansthese methods can be done automatically, but their performance degradessignificantly as the number of 2D views decreases and ultimately require manualpost-processing to attain clinically useful reconstructions

[0016] .Due to the non-robustness of the direct back-projection method, most current X-rayangiography methods avoid reconstructing the full image volume via tomography.Instead, back-projection is applied to a limited set of salient landmark pixels fromeach image to optimise the reconstruction more specifically to the vessel structure. P604279PC00To the best of our knowledge, in all current methods the salient points are vesselcentrelines, landmarks such as bifurcations, or both, because any point within thevessel volume that is not on the centreline cannot be guaranteed to be uniquelyidentifiable in a set of X-ray images with arbitrary perspectives. The trade-off ofapplying a more selective backprojection is that these salient features must first beidentified and also then linked view-to-view in a process of correspondence matching.In its most basic form, both of these steps can be done manually

[0017] —which is amethod used in even recent work

[0018] —but can be reduced to a semi-manual taskby specifying a small number of points-of-interest followed by minimum-energypathfinding

[0019] or fast-marching level set

[0020] . Fully automatic centreline extractionin two dimensions has been demonstrated in X-ray coronary angiography using askeletonisation algorithm

[0021] , though the work does not include 3D reconstructionof the centrelines. Automatic centreline detection has also been demonstrated inretinal imaging

[0022] and CT coronary imaging [23, 24] using a neural network, andin X-ray angiograms by deriving centreline information extracted from CT data

[0025] .There has been no end-to-end automated reconstruction applied to X-ray coronaryangiograms alone.Following the selection of 2D salient points, epipolar backprojection is performed byfinding the intersection of these image-space points when projected into 3D world-space from the set of images given the known poses of the camera in each view (seeFigure 1). As shown in Figure 1, epipolar back-projection is a method for recovering3D information from two 2D images of different views. As shown in Figure 1 a) The3D coordinate y appears as the image coordinate x to the camera with centre C andx′ to the camera with centre C′. The points x, x′, and y lie in a common planecalled the epipolar plane, which can be constructed from the points y, C, and C′.When x, x′, C, and C′ are known, epipolar back-projection recovers the location ofy by finding the intersection of the ray projected between C and x backward and theray projected between C′ and x′ backward. In Figure 1 b), in reality, thecorresponding point x′ for an image landmark x is unknown. In this case, an epipolarline is formed in the second image from the known x, C, and C′, which passesthrough the baseline e. If C′ is known with certainty, then the true correspondingpoint x’t must lie on this line. Correspondence matching is used to select a pointalong this line that is most likely to be the same image feature visible in x. If x′ t iscorrectly selected, then the back-projection method is used to calculate the true 3D coordinate of the visible feature, yt. However, incorrect correspondence matching— P604279PC00 returning false points x’f1 or x’f2, for example—will recover the wrong 3D coordinate, such as yf1 or yf2.Within epipolar back projection if the centrelines are used as salient features, theresult will be a 3D centreline map, whereas if simpler landmarks such as bifurcationpoints are used, instead those landmark points will have been located within 3Dworld-space. In a perfectly calibrated system with no inter-view motion, this methodwill produce accurate and precise reconstructions, but in reality there is a finitedegree of precision of camera pose and salient-pixel labelling, in addition to cardiac,respiratory, and other motion between each view. The result is that the projectedpixels from different views do not intersect in 3D space, resulting in a failure toascribe a third dimensional value to the pixel point. While in bi-plane angiography,where two orthogonal images are captured simultaneously, this effect can benegligible

[0026] , this poses a significant challenge in monoplane angiography.For this reason, a large amount of work has been done to make the epipolar back-projection more robust to clinical conditions. In

[0027] , the vector defining theminimum common perpendicular of each epipolar line is used as an objective that isminimised iteratively through rigid transformation of each camera in the scene, whichis representative of the standard iterative optimisation of pose to address this issue[28, 20], though non-iterative methods such as least-squares have also beendemonstrated

[0018] . In

[0029] , a non-rigid optimisation is added to the rigid objectivein order to better account for respiratory and patient motion, but requires additionalmanual landmark pixel labelling.Forward approaches have also been shown, including a graphcuts method whichiteratively labels centreline pixels with a discretised depth until maximum agreementbetween views is achieved

[0021] . This can be considered a form of earlier forwardprojection methods, defined by their optimisation of an initial 3D centreline. This isposed as a dual-energy algorithm, with an external force applied to the currentgeometry by its disagreement with 2D observations, balanced by an internal forcethat discourages discontinuities in the geometry. The exemplary technique is theactive contour method

[0030] which has been applied to coronary artery reconstruction

[0031] . While the advantage of this approach is that the external energy can becalculated from the raw image intensities, which avoids the need for thepreprocessing tasks necessary in back-projection techniques, the initial geometrymust be derived from manual initial control points

[0016] . These earlier manual forward P604279PC00methods have given way to hybrid forward-backward methods that seek to combinethe robustness of forward methods with the explicit geometric constraints defined inbackward formulations

[0018] .Though in many cases the above techniques demonstrate a good degree of success,these traditional geometric methods do not fully solve the problem. Firstly, currentmonoplane reconstruction methods rely on at least some manual computer-basedtask, ranging in complexity from full vessel segmentation

[0014] , to the selection ofseveral landmarks

[0018] , to selecting a seed region for centreline extraction followingan automatic pre-processing routine

[0020] . Secondly, they are either not robust tonoise and miscalibration

[0017] or require an online optimisation algorithm that iscomputationally-expensive and dependent on good starting conditions to converge[20, 29, 18]. Finally, even under idealised conditions, existing classical reconstructionmethods introduce significant uncertainty into the derived vFFR estimate. Empiricalmeasurements of the uncertainty arising from these methods show thatreconstruction error exceeds that of pressure-wire FFR measurement (standarddeviation: ±0.015

[0032] ): in

[0033] , Solanki et al. show via phantom study on bothcurved aluminium rods and a patient-specific 3D printed reference model thatepipolar back-projection via semi-manual correspondence matching without motionartefacts results in vFFR error of ±0.06.Machine learning-based reconstruction for medical imagingReconstruction of coronary arteries from 2D X-ray using machine learning is anascent field of research. In Iyer et al.

[0034] , a set of 2 or 3 binarised, syntheticangiography images are reconstructed using a multistage neural network, whichpredicts an M × N matrix of centreline points and their radii, where M is a fixednumber of coronary vessels and N is a fixed number of centreline points per vessel.While the mean average error of the vessel centreline is low (0.83 ± 0.29 mm) thefully-connected architecture means the network can only reconstruct 4-vessel trees.Furthermore, it uses binary images as input and hence requires an additionalsegmentation preprocessing step.Wang et al.

[0035] demonstrate a method where a GAN generates plausible coronaryanatomies with 2D discriminator versus ground-truth 2D segmentations to ultimatelyproduce a voxel grid representation from an unseen set of 2D segmentations. Theyachieve an Intersection-over-Union of 0.718 when applied to an augmented set ofjust 44 clinically-derived coronary anatomies and there is no indication of how the P604279PC00reconstruction quality is affected by an imperfect segmentation, as would be the casein an end-to-end system. A segmentation-first approach has also been applied tobiplane neurovascular reconstruction

[0036] , with a generative-adversarial network(GAN) added to improve reconstruction quality. In this case though, reconstructionis performed slice-by-slice with ground-truth supervision from 3D magneticresonance images.Hwang et al.

[0037] use a multi-stage process with intermediate 2D segmentation thatalso demonstrates integration into a vFFR quantification process, achieving areconstruction that shows two-view correspondence error that is similar to humanerror and vFFR error of < 2.5%. The only machine learning component of theirmethod is a vessel start point and endpoint detection neural network implementedas a UNet. The predicted terminal points of the vessel are used to constrain anautomatic template deformation process based on classical back-projection with anenergy term calculated from the Frangi-filtered 2D angiography images. While intheory this approach could be used to build up the full coronary tree vessel-by-vessel,only single-vessel reconstruction of the main coronary arteries is demonstrated. Onehindrance of the extension to multi-vessel is that the templates used during 3Dreconstruction are produced per-vessel and hence in the smaller coronary vessels,where more intersubject variability arises, it is not trivial to conceive of a method toproduce robust templates.To find other examples of deep reconstruction in medical imaging, it is necessary tolook beyond angiography. In

[0038] , reconstruction of the thorax from a single chestX-ray is achieved using a 2D-to-3D UNet, in which the 2D image is first upsampledchannel-wise into a 3D volume using transposed 3D convolutional layers. In order toaccount for the sparse information following the dimensional-upsampling process, aprobabilistic approach is adopted in order to learn to produce a maximumlikelihoodgeometry which is again in a voxel grid form. A similar probabilistic approach isfollowed in

[0039] but a 2D-to- 3D conditional variational autoencoder is used in placeof the upsampling UNet to reconstruct a myocardium volume from three orthogonalMRI slices. Ultrasound imaging is, like X-ray fluoroscopy, a widely-used, natively 2Dmodality in medical imaging and 3D reconstruction is an active field of ultrasoundresearch. In

[0040] , a Pix2Vox++ architecture is used to reconstruct myocardiumvolume from synthetically-generated 2D views, while in

[0041] a conditional variationalautoencoder is used to improve robustness to noise when reconstructing voxel gridsof fetal skulls from three orthogonal views. Though voxel representations are the P604279PC00most widely used geometry representation in ultrasound reconstruction, other formshave been shown to work as well, such as an implicit representation of the fetal brain

[0042] , which is shown to achieve superior resolution when compared to discrete forms.In summary, while there are many examples of the application of deep learning tothe broader field of coronary X-ray angiogram image processing such assegmentation, as well as some headway into reconstruction tasks of medical images,there is a very limited basis upon which to further develop a solution to the automaticreconstruction of the coronary arteries. We also note that the small number ofprevious demonstrations of deep learning in coronary angiography reconstructionlook to replace individual components of existing, classical reconstruction pipelineswhich, as we demonstrated in section 2.1, is usually performed by first identifyinglandmark features or vessel regions in each 2D image and then applying a ray-projecting process to identify a corresponding 3D point with which to gradually forma discrete geometry of the vessel tree. However, as we will show in section 2.3, wefind that this formulation is at odds with the broader methodologies of deep learningreconstruction in other fields.In

[0062] , a vision transformer model is integrated as an encoder network, specificallydesigned to direct and enhance attention to particular regions within the X-ray image. The reconstructed meshes exhibit a similar topology to the ground truth organvolume, demonstrating the ability of the model in accurately capturing the 3Dstructure from a 2D image. The effectiveness of the model is evaluated on lung X-rays using several metrics, including volumetric Intersection over Union (IoU). Themodel outperforms the state-of-the-art method in the literature for lungs(DeepOrganNet) by about 7-9% achieving IoU’s between 0.892-0.942 versus DeepOrganNet’s IoU of 0.815-0.888.In

[0063] , e a benchmarking framework is proposed that evaluates tasks relevant toreal-world clinical scenarios, including reconstruction of fractured bones, bones with implants, robustness to population shift, and error in estimating clinical parameters.The open-source platform provides reference implementations of 8 models (many ofwhose implementations were not publicly available), APIs to easily collect and preprocess 6 public datasets, and the implementation of automatic clinical parameter and landmark extraction methods. An extensive evaluation of 8 2D-3D models on equal footing using 6 public datasets is presented comprising images for fourdifferent anatomies. The results show that attention-based methods that captureglobal spatial relationships tend to perform better across all anatomies and datasets; P604279PC00 performance on clinically relevant subgroups may be overestimated without disaggregated reporting; ribs are substantially more difficult to reconstruct compared to femur, hip and spine; and the dice score improvement does not always bring a corresponding improvement in the automatic estimation of clinically relevant parameters. Machine learning-based reconstruction in other fieldsThe application of neural networks in 2D-to-3D reconstruction is a diverse and activediscipline. Applications include human pose estimation

[0043] , aerial photogrammetry

[0044] , civil engineering project management

[0045] , reconstructing indoor scenes

[0046] ,and reconstructing single objects

[0047] .Voxel grid reconstruction: With the success of 2D image processing neuralnetworks following the adoption of convolutional architectures, an obvious candidatefor 3D reconstruction is one which generalises the 2D pixel grid method to a 3D voxelgrid. With what is widely considered to have been the first such method to have beenpublished, Choy et al.

[0048] use a fully convolutional architecture to reconstruct adense voxel map from one or more uncalibrated photographic images. In order tocombine the latent embeddings from multiple viewpoints, the authors use a longshort-term memory module after the encoder, which is chosen due to its behaviourupon receiving additional information alongside previously exposed information—inthe case of the optical multiview experiment, this is appropriate due to the occlusionof parts of the geometry according to optical light physics. The decoder then recoversthe higher-dimensional output from the combined latent embeddings viadeconvolution and unpooling and the network is trained via supervision with respectto a ground-truth voxel grid using a voxel-wise binary cross-entropy loss function.Though much of the publication focuses on the recurrent aspect of the method, itserves to prove that even a single 2D image can provide enough signal about the 3Dgeometry for the neural network to learn a procedure to recover the lost informationto some degree. However, a key limitation of recovering a fully-dense voxel grid isthat the memory requirements scale at least cubically, which necessarily limits themaximum grid resolution.Pointcloud reconstruction: Another foundational methodology for reconstructionwith neural networks is the family of pointcloud decoders, which trace a lineage toclassical depth map methods where the pixel grid of depth values is simply an image P604279PC00of a set of 3D points. For this reason, pointclouds have been used for a long time inrobotic tasks such as aerial robotics

[0049] and autonomous driving

[0050] . Whereasgrid-based methods have a discrete ground-truth that naturally lends to element-wise losses such as binary cross-entropy or p-norms, pointcloud reconstruction mustbe invariant to permutation of points. As well as introducing the first example of deeplearning for pointcloud reconstruction, Fan et al.

[0051] investigated two permutationinvariant loss functions—Chamfer distance and earth movers’ distance—finding thatthe former better preserves fine surface details, though at the expense of producinga greater number of spurious outlier points. While pointcloud reconstruction canresult in high-quality pointclouds, their primary drawback is that they do not containthe necessary structural information to recover a mesh without additional, lossy post-processing.Implicit field reconstruction: An implicit field is a continuous scalar fieldcontaining geometric information, for example occupancy or signed-distance to thenearest surface. It is represented as a mathematical equation which maps a 3Dcoordinate to a scalar value, with the primary advantage that inference can beperformed with linear time / memory complexity with respect to the 3D resolution andhence a significantly higher quality reconstruction is feasible than is possible withvoxel grid reconstruction.The successful application of neural networks to implicit field reconstruction wasintroduced in short succession by two independent works in 2018. In both works, theimplicit field is the occupancy field and is framed as a pointwise labelling task, butthe methods employed are quite different. Chen and Zhang

[0052] concatenate eachquery point with a dz-dimensional image embedding to produce a dz + 3 input vectorwhich is then fed through a contracting, fully-connected decoder in order to outputa scalar occupancy probability. The input vector is reintroduced unchanged in severalof the intermediate layers to improve the resultant accuracy and an L2 loss is used.In contrast, Occupancy Network

[0053] uses a classification-specific binary cross-entropy loss and utilises conditional batch normalisation in order to merge aconditioning observation latent embedding with an arbitrarily large number of points,which greatly reduces the number of repeated passes needed to quantify theoccupancy field during both training and inference.The first signed distance field neural network was employed by Park et al.

[0054] ,though they do not specifically investigate the image-to-geometry reconstruction P604279PC00task and instead focus on learning latent spaces for arbitrarily large shape datasetsvia their ‘auto-decoder’ method. In their experiments they use a clamped L1-normin order to focus learning near the surface, which is shown in their supplementarymaterial to be influential on the reconstruction performance. The conditional batchnormalisation decoder of ONet has since been employed in a signed distance functionin

[0055] , which also adds a multistep single image encoder that first generates adepth, silhouette, and normal maps estimates via a UNet and then jointly embedsthe intermediate group into a single latent vector.Summary of the InventionThis work specifically addresses the first stage of image-derived haemodynamicestimation in coronary angiography: the extraction of a 3D anatomy from multiple2D X-ray images in a process known as reconstruction. Particular challenges stillexist in current state-of-the-art reconstruction methods, including the lack of a fully-automated process, slow processing times, and sensitivity to noise and patientmotion. We present a novel framework for coronary angiography reconstructionwhich already shows promise in addressing each of these issues, by recasting theproblem of angiographic reconstruction into one more aligned with modern deeplearning methods beyond the medical imaging discipline. In particular, we presentfour novel aspects in our design: first, a generalised probabilistic framework forangiographic reconstruction with deep learning; second, the use of implicit fielddecoders to represent the coronary anatomy; third, a novel, surface-focusing lossfunction to improve the performance of implicit field methods for thin, sparsestructure; and fourth, a data synthesis routine, including the ability to render X-rayimages on-the-fly during training.In view of the above, from one aspect there is provided a method of generating a 3Dvirtual model of internal anatomy of a subject, comprising: receiving a plurality of2D medical images of the internal anatomy of the subject that is to be modelled, the 2D medical images having been obtained from at least 2 different imaging angles or planes; inputting the plurality of 2D medical images into a trained encoder-decoder neural network, the encoder-decoder neural network having been trained to reconstruct a feasible estimate for the internal anatomy shown in the plurality of 2D medical images from a universe of possible 3D anatomical models; finding, using the encoder-decoder neural network, a best estimate 3D anatomy from the universe of possible 3D anatomical models that probabilistically most closely matches the internal anatomy of the subject as represented in the plurality of 2D medical images P604279PC00 input to the encoder-decoder neural network; and outputting the best estimate 3D anatomy as representative of the internal anatomy of the subject. This method of generating a 3D virtual model of internal anatomy of a subject using2D medical images has several advantages. It is non-invasive, meaning that it doesnot require any physical pressure or other sensor within the coronary artery. It isalso accurate, as it uses a probabilistic technique to determine the maximumlikelihood 3D model of internal anatomy from multiple input 2D images. Additionally,it is efficient, as it allows for the calculation of computed haemodynamicmeasurements, such as the virtual fractional flow reserve (vFFR), directly from 2D X-ray angiography images.In one example, the finding comprises: encoding, using an encoder neural network,the 2D medical image data into latent space data; and decoding, using a decoderneural network, the latent space data into a virtual 3D model of the subject’s internalanatomy; wherein the latent space data is of a lower dimensionality than the 2Dmedical image data and the virtual 3D model. In one example, the encoder neural network has been trained to encode the received 2D imagery into a maximum-information latent code z, being the latent space data,wherein the latent code z is a low-dimensional latent embedding of observed data inthe input 2D image sequence, based on a stochastic image forming process of the 2D input images. One advantage of the encoder neural network being trained to encode the received 2D imagery into a maximum-information latent code z is that it allows the encoder to capture the most important information from the 2D imagery and represent it in a compact, low-dimensional form. This latent code z is a low- dimensional latent embedding of the observed data in the input 2D image sequence, based on a stochastic image forming process of the 2D input images. This means that the encoder can efficiently represent the most important information from the 2D imagery in a way that can be easily used by the decoder neural network to reconstruct a 3D model of the subject's internal anatomy. In one example, preferably the latent space data is input into the decoder neural network to condition the decoder on the observed data in the 2D input images, the decoder having been trained to learn the distribution of feasible internal anatomies, including healthy, diseased, and impossible anatomies as the universe of possible 3D internal anatomical models, resulting in an output 3D reconstruction that is the best P604279PC00estimate 3D reconstruction given the input 2D images. Inputting the latent spacedata into the decoder neural network to condition the decoder on the observed data in the 2D input images has the advantage of allowing the decoder to use its trainingon the distribution of feasible internal anatomies, including healthy, diseased, andimpossible anatomies, to produce an output 3D reconstruction that is the best estimate 3D reconstruction given the input 2D images. This means that the decoder can use its knowledge of the distribution of possible anatomies to produce a more accurate and reliable 3D reconstruction of the subject's internal anatomy In one example the decoder neural network uses an implicit field reconstruction torepresent the output best effort 3D reconstruction. One advantage of using animplicit field reconstruction to represent the output best effort 3D reconstruction is that it allows for the representation of the output best effort 3D reconstruction in a continuous scalar field containing geometric information, such as occupancy or signed-distance to the nearest surface. This means that the decoder can efficiently represent the 3D reconstruction in a way that can be easily used for further processing, such as fluid dynamics modeling or visualization. In one example, the latent space data is such that it suppresses any 3D reconstruction by the decoder which is implausible given the available information in the 2D input images. One example further comprises: generating a synthetic 2-dimensional image of the internal anatomy of the subject at a desired imaging plane by projecting the best effort 3D model of the internal anatomy to the desired imaging plane.Another example further comprises transforming the best effort 3D model into asurface and / or volume mesh model. In one example wherein the virtual 3D model is of the coronary artery, the method further comprises: using the 3D model, computing one or more haemodynamicmeasurements that relating to the coronary artery of the subject. Preferably the oneor more haemodynamic measures is a virtual Fractional Flow Reserve (vFFR) measurement of the coronary artery of the subject. One advantage of using the 3D model to compute one or more haemodynamic measurements, such as the virtual Fractional Flow Reserve (vFFR) measurement of the coronary artery of the subject,is that it allows for non-invasive testing. This means that no pressure sensor is P604279PC00 physically needed within the coronary artery, which can more readily inform a physician about whether a stent needs to be inserted into the artery or not, or to guide other treatment options.From another aspect, the present disclosure also describes a method of training anencoder-decoder neural network, comprising: receiving a plurality of 2D medicalimages of the internal anatomy of a plurality of subjects, the 2D medical images eachhaving been obtained from at least 2 different imaging angles or planes; and trainingthe encoder-decoder neural network, wherein training the encoder-decoder neuralnetwork comprises: training an encoder neural network to convert the 2D medicalimage data into latent space data; and training a decoder neural network to reconstruct a virtual 3D model of the subject’s internal anatomy from the latent space data, wherein the latent space data is of a lower dimensionality than the 2D medical image data and the virtual 3D model. In one example training the encoder neural network comprises training the encoder to encode the received 2D imagery into a maximum-information latent code z, beingthe latent space data, wherein the latent code z is a low-dimensional latentembedding of observed data in the input 2D image sequence, based on a stochastic image forming process of the 2D input images.Another example further comprises training the decoder to learn the unconditionaldistribution of internal anatomies, including healthy, diseased, and impossible anatomies as the universe of possible virtual 3D internal anatomical models, resultingin an output 3D reconstruction that is a best effort 3D reconstruction given a set ofinput 2D images. In one example the decoder neural network uses an implicit field reconstruction to determine the output best effort 3D reconstruction. In one example the latent space data is such that it suppresses any 3D reconstruction by the decoder which is implausible given the available information in the 2D input images.Another example further comprises increasing resolution of the virtual 3D model ofthe subject’s internal anatomy near to a surface of the anatomy by modifying a loss function to increase the weighting of points in the virtual 3D model close to the P604279PC00surface compared to points that are further away from the surface. In one examplethe increase in the weighting of points close to a surface effectively balances the set of points in terms of points falling within and outside of the anatomy.One example further comprises undertaking 3D data augmentations as part of theprocedure of training of the encoder-decoder model. The data augmentations includeone or more selected from the group comprising: motion modelling, contrast agenttransport, disease modelling, and arbitrary 3D affine transformations.Another aspect provides a computer system, comprising:a. one or more processors;b. computer readable memory; andc. a computer readable medium storing computer readable instructionsthat when executed by the processor cause the processor to operate in accordance with the method of any of the preceding aspects.A further aspect provides a medical imaging system comprising:a. a medical imager capable of producing medical images of internalanatomy of a subject to be imaged; and b. a computer system according to the above aspect, the computersystem being arranged such that in use it receives 2D medical images of the internal anatomy of the subject from the medical imager, and processes those 2D images according to the method of any of the preceding aspects to produce a virtual 3D model of the internalanatomy of the subject. In one example the medical imager is an X ray medical imaging system, and more preferably a C-arm X ray medical imaging system. Another aspect of the present disclosure provides a computer-implemented methodof generating synthetic x-ray images using a pseudo-angiographic imaging model,the method comprising modelling the formation of an X-ray image as a Poissonprocess in which photons leaving a modelled photon source have some probability ofbeing occluded before reaching a modelled detector, the probability being implicitlydefined according to an event density that is derived by volume sampling an arbitrary mesh. P604279PC00 Brief Description of the Drawings Embodiments of the invention will now be further described by way of example only and with reference to the accompanying drawings, wherein like reference numerals refer to like parts, and wherein:Figure 1: A diagram of Epipolar back-projection as a method for recovering 3Dinformation from two 2D images of different views.Figure 2: An overview of the data synthesis procedure. First, a directed acyclic graphrepresentation of a coronary artery network is generated. A surface mesh is thengenerated, which can either be rendered into a 2D pseudo-angiography image orfurther processed into another 3D representation such as a pointcloud or voxel grid.Figure 3: Example synthetic anatomies, shown with increasing complexity from left-to-right. From top-to-bottom are shown five random samples for each synthesisroutine.Figure 4: Example synthetic angiogram renderings with variable vessel opacity.Figure 5: Examples of occupancy field reconstructions of each class in the Vesseldataset. There are three columns showing, from left-to-right, the bifurcation, doublebifurcation, and coronary categories of the Vessel dataset. Each column is comprisedof three sub-columns showing, from left to right, the two input images, threeorthogonal perspectives of the rendered model prediction, and three orthogonalperspectives of the rendered ground-truth reconstruction. There are five rows, eachshowing a randomly-selected example from the test set.Figure 6: Examples of signed-distance field reconstructions of each class in the Vesseldataset. There are three columns showing, from left-to-right, the bifurcation, doublebifurcation, and coronary categories of the Vessel dataset. Each column is comprisedof three sub-columns showing, from left to right, the two input images, threeorthogonal perspectives of the rendered model prediction, and three orthogonalperspectives of the rendered ground-truth reconstruction. There are five rows, eachshowing a randomly-selected example from the test set. Figure 7: The effect of the α hyperparameter of Hyp-L2 and corresponding clampedL2 on reconstruction quality of the SDFNet decoder in terms of Hausdorff Distance(HD), Chamfer Distance (CD), F1 scores at 0.01, 0.02, and 0.05 normalised lengthunits, and Intersection-over-Union (IoU). P604279PC00 Figure 8: Reconstruction quality of Coronary:Vessel geometries by the ONet decoderwith 2nd-viewpoint angular uncertainty in terms of Hausdorff Distance (HD),Chamfer Distance (CD), F1 score at 0.01, 0.02, and 0.05 normalised length units,and Intersection-over-Union (IoU). The dotted line indicates model performancewithout any uncertainty. Figure 9: Reconstruction quality of Coronary:Vessel geometries by the ONet decoderwith a rigid translation motion model applied, in terms of Hausdorff Distance (HD),Chamfer Distance (CD), F1 score at 0.01, 0.02, and 0.05 normalised length units, and Intersection-over-Union (IoU). Figure 10: Reconstruction quality of Coronary:Vessel geometries by the ONetdecoder with a rigid rotation motion model applied, in terms of Hausdorff Distance(HD), Chamfer Distance (CD), F1 score at 0.01, 0.02, and 0.05 normalised lengthunits, and Intersection-over-Union (IoU).Figure 11: Examples of occupancy field reconstructions of each class in the Vesseldataset when using inconsistent view-points without angle-encodings provided.There are three columns showing, from left-to-right, the bifurcation, doublebifurcation, and coronary categories of the Vessel dataset. Each column is comprisedof three sub-columns showing, from left to right, the two input images, threeorthogonal perspectives of the rendered model prediction, and three orthogonalperspectives of the rendered ground-truth reconstruction. There are five rows, eachshowing a randomly-selected example from the test set. Figure 12: Reconstruction quality of Coronary:Vessel geometries by the ONetdecoder for different view angle separations between input images in terms ofHausdorff Distance (HD), Chamfer Distance (CD), F1 score at 0.01, 0.02, and 0.05normalised length units, and Intersection-over-Union (IoU).Figure 13: a diagram representing an overview of a system according to embodiments of the disclosure. Figure 14: a diagram representing an overview of a system according to further embodiments of the disclosure. Figure 15: a system block diagram of an example computer system forming the basis of a system of embodiments of the disclosure. P604279PC00Figure 16: a diagram illustrating the probabilistic framework for reconstruction usedin embodiments of the disclosure.Figure 17: a diagram illustrating how Implicit field reconstruction is used inembodiments of the disclosure.Figure 18: a diagram showing the operation of the surface-focused loss function usedin embodiments of the disclosure.Figure 19: a diagram illustrating a pseudo-angiographic imaging model used togenerate synthetic x-ray images. It shows an overview of the stochastic model usedto generate synthetic X-ray images. The formation of an X-ray image is modelled as a Poisson process in which photons leaving the source have some probability of being occluded before reaching the detector. Rather than requiring an explicit formulation of this probability, it is implicitly defined according to an ‘event density’ that can be derived by volume sampling an arbitrary mesh.Figure 20: a diagram showing a 3D data augmentation routine used in embodimentsof the disclosure. Figure 21: a method flow diagram illustrating an over-arching method undertaken by embodiments of the disclosure. Figure 22: a method flow diagram illustrating a method of generating a 3D virtual model of internal anatomy of a subject undertaken by embodiments of the disclosure.Figure 23: a method flow diagram illustrating a method of training an encoder-decoder neural network undertaken by embodiments of the disclosure.Overview of the Embodiments Embodiments of the present disclosure provide a method and system that can generate a virtual three-dimensional model of internal anatomy of an imaged subject, based on a plurality of two-dimensional medical images of the internal anatomy, such as X-ray images, obtained from slightly different viewpoints. The 3D model is then obtained using a probabilistic reconstruction technique, which takes an unconditional distribution of internal anatomies, including “normal” anatomies of healthy people, as well as diseased or otherwise impossible anatomies, and uses the obtained 2D images to guide a 3D reconstruction to a maximum likelihood internal anatomy that is most supported by the obtained 2D images from the whole P604279PC00 unconditional distribution of anatomies. The maximum likelihood 3D reconstruction can then be output as a 3D model of the internal anatomy of the subject, or can be used as the basis of 2D projections to generate reconstructed 2D images of the 3D reconstruction onto any imaging plane, thus giving reconstructed 2D images of theinternal anatomy from any directional viewpoint.Figures 16 and 17 illustrate the process in more detail, with Figure 16 showing aprobabilistic framework for 3D reconstruction of the internal anatomy, from input 2Dimages of the same anatomy. In the examples described herein the examples are focussed on the reconstruction of the coronary artery anatomy of the imaged subject, although it should be noted that examples of the present disclosure are not limited to reconstruction of the coronary artery anatomy only, and any internal anatomy of an imaged subject may be reconstructed. As shown in Figure 16, the main task is the 3D reconstruction of internal anatomy ofan imaged subject, such as coronary artery anatomy, from several (two or more) 2Dimages obtained of the internal anatomy of the subject. The image sequence (forexample an angiogram sequence in the case of the coronary artery) captured of asingle patient is fed into an encoder-decoder neural network, with the 3D anatomyrecovered as the output. At the centre of the network, the latent code z is the mostcompact representation of the anatomy. That is, in operation, embodiments of thepresent disclosure receive plural 2D images of the internal anatomy of the subject,and input those images into an encoder network which has been trained to encodethe received imagery into a maximum-information latent code z. The latent code z isa low-dimensional latent embedding of the observed data in the input 2D image sequence, which is based on the stochastic image forming process of the 2D inputimages. The latent code z is then used in the decoder to condition the decoder onthe observed 2D images. In this respect, the decoder has been trained to learn theunconditional distribution of internal anatomies, including healthy and diseasedanatomies, and impossible anatomies. This distribution contains information about what ‘normal’ hearts look like, what diseased hearts look like, as well as whichanatomies are impossible. The latent code z suppresses any 3D reconstruction bythe decoder which is implausible given the available information in the 2D images,resulting in an output 3D reconstruction that is the maximum likelihood 3Dreconstruction given the input 2D images. With respect to the latent code z, as noted above, this takes as its basis the stochastic image forming process of a 2D X-ray image. Figure 19 illustrates this is more detail, P604279PC00 providing an overview of the stochastic model used to generate synthetic X-ray images. The formation of an X-ray image is modelled as a Poisson process in which photons leaving the source have some probability of being occluded before reaching the detector. Rather than requiring an explicit formulation of this probability, it is implicitly defined according to an ‘event density’ that can be derived by volume sampling an arbitrary mesh. Figure 21 shows the whole 3D reconstruction and synthetic 2D imaging from thereconstruction process further, with reference to Figures 13, 14 and 15. Figures 13and 14 show two different imaging systems forming embodiments of the disclosure. In Figure 13 a subject 136 to be imaged is resting supine on an imaging table. An X- ray C-arm 132 is provided to take X ray images of the subject 136. The C-arm is moveable between different positions to allow plural X-ray images to be taken of subject 136 from slightly different angles. The plurality of 2D images should be obtained of the subject from slightly different viewing angles, but the difference in viewing angle may be as little as 4 degrees. A larger angle will usually, however, give improved results. The plural X ray images from the C-arm 132 are fed to an X-ray signal processing computer 138, which processes the plural images to generate a 3D model of the subject’s internal anatomy based on the probabilistic reconstruction technique described herein to obtain the maximum likelihood 3D reconstruction of the subject’s 135 internal anatomy. The obtained 3D representation (and any associated 2D projections thereof) 140 is then fed to a display 134, for display to a user. In Figure 14, the same elements are present, but with the difference that here the X-ray signal processing computer 138 is located remotely from X ray imaging apparatus 132 and the display 134, and more specifically is in communication with the imaging apparatus 132 and the display 134 via a network 142, such as the Internet. This allows the signal processing and probabilistic 3D reconstruction to be performed separately from the original 2D image acquisition, and hence permitsimplementation of the probabilistic reconstruction as a service that can be providedto multiple hospitals, clinics, and other medical imaging organisations at the same time, without requiring dedicated processing means and expertise in operation at each site. Figure 15 illustrates in more detail the X-ray processing computer 138, whether located remotely from the image acquisition means or not. In this respect, X-rayprocessing computer 138 comprises a CPU 150, memory such as RAM 152, an input P604279PC00 output interface 156, and a network interface 154, all of which are communicablewith each other via a central bus 158. Also provided is a non-volatile computerreadable storage medium 160, which may be, for example, a hard disc drive, or solid-state storage such as an NVME drive, or the like, upon which is stored variousprograms to allow the embodiment of the present disclosure to be operated, as wellas the input data used by and output data generated by the programs. In thisrespect, firstly provided is control program 162, which provides overarching controlof the computer system 138, and acts to coordinate the input and output data, aswell as the operation of all of the other programs. Next, specific to the presentembodiment, an encoder neural network 164 is provided, as well as a decoder neural network 166. Further provided is an implicit field reconstruction program 168, and a surface focused loss function program 170. A 3D data augmentation program 172 isalso provided. The operations of the encoder neural network 164, decoder neuralnetwork 166, implicit field reconstruction program 168, surface focused loss function program 170, and 3D data augmentation program 172 will be described later.Also stored on the non volatile storage 160 is the input data required for the systemto operate, in the form of a plurality of X-ray images 174. The plurality of X-rayimages 174 are received from the X-ray C-arm 132, either directly, or via the network142, as described previously. The X-ray images 174 are X-ray images of the subject136, with each image taken from a slightly different field of view, as described previously. At least two such images are required for 3D reconstruction, but more than two images can also be used, to obtain greater accuracy. The X-ray images 174are used as input images to the encoder neural network 164 with the output fromthe decoder neural network 166 being the maximum likelihood probabilistic 3D model176. As explained previously, the 3D model 176 is the 3D model which has themaximum likelihood probabilistically of representing the shape of the subject’sinternal anatomy, given the plural X-ray images 174 that are input into the encodernetwork 164. Having obtained the maximum likelihood 3D model 176, it is thenstraightforward to obtain projections of the 3D model into a plurality of different imaging planes, to give reconstructed 2D images in different imaging planes of the maximum likelihood 3D model that has been determined. As shown in Figure 15, a plurality of such 2D images 178 can be obtained from the generated maximum likelihood 3D model 176. Figure 21 illustrates the overarching method of embodiments of the present disclosure. Firstly, at step 21.2, 2D X-ray images of this subject are captured. As explained previously, a plurality of X-ray images of the subject are captured, from a P604279PC00slightly different angle to each other, to give a slightly different field of view in eachimage. A field of view difference of 4° or more is sufficient, and the difference canbe in any axis of movement of the X ray C arm (i.e. the C arm can be rotated in adirection longitudinally along the imaging table (as shown in the Figures), or laterallyacross it (which would be in and out of the plane of the Figures). The plurality of 2DX-ray images that have been captured are then input at 21.4 into the encoder-decoder neural networks to find the maximum likelihood 3D anatomy model that fitsthe 2D images of the subject that have been input. Once the maximum likelihood 3Danatomy model has been found, a 3D point cloud model can be generated at s.21.6 to represent the maximum likelihood 3D anatomy of the subject. Once the 3D point cloud model has been generated, the model can optionally be augmented byincluding surface mesh representations (s.21.8), as well as undertaking otheraugmentation operations (described later), and various 2D images as desired can be formed by casting the 3D model to any desired 2D imaging plane, at s.21.10. With the above, therefore, a maximum likelihood probabilistic technique is provided to generate 3D internal anatomy models from 2D medical imagery of the same anatomy. The generated 3D internal anatomy models represent the 3D anatomy which has the maximum likelihood of matching the input 2D images, matched froman entire universe of possible 3D anatomies, including anatomies of healthy subjects,diseased subjects, as well as impossible anatomies that could not exist in real life.Once the maximum likelihood 3D anatomy has been obtained, it can be augmentedfor example using surface mesh representations, and can then be used to generate2D images of the internal anatomy from any angle, by casting the 3D anatomy modelto any desired 2D imaging plane.Referring again to Figure 21, how the encoder- decoder network finds the maximumlikelihood 3D anatomy is shown in step 21.42 and 21.44. That is, in step 21.42 animplicit field representation is generated of the 3D anatomy, as described in more detail later. Moreover, in step 21.44 once a 3D anatomy model has been obtained, regions close to the surface of the 3D anatomy are weighted as more important than regions far from the surface, using a surface focussed loss function. Further details of these aspects will be described next, with reference to Figures 17 to 20. Turning first to Figure 17, this figure shows the implicit field reconstruction that isperformed by the decoder network 166. As shown in Figure 17.1, implicit fieldreconstruction is performed as a pointwise inference in Cartesian space. For any pointp and latent code z, the decoder predicts a useful quantity, such as occupancy or P604279PC00signed distance. Figure 17.2 illustrates that parallel computing can be leveraged forefficient inference, for example by predicting implicit field values across a grid (shown here in 2D for ease of illustration). Multipass inference allows for iterative enhancement of the reconstructed volume by increasing resolution in regions near asurface., as shown in Figure 17.3. Following the above, but not shown explicitly inthe figures, a “marching cubes” algorithm is then used to transform the inferredimplicit field values into a surface mesh for further processing (e.g. fluid dynamicsmodelling). “Marching cubes” algorithms are known already in the art (see e.g.Lorensen, W. E.; Cline, Harvey E. (1987). "Marching cubes: A high resolution 3dsurface construction algorithm". ACM Computer Graphics. Vol 21 No(4): 163–169.).Figure 18 shows an example of the operation of the surface focussed loss function which acts to increase the weightings or regions close to the surface at the expense of regions farther from the surface. Figure 18.1 illustrates the problem, in that thin structures such as blood vessels cause sparsity in the bounding volume. This means that empty regions dominate optimisation during model training, which reduces training efficiency and hence leads to underfitting and poor trained-modelperformance at inference time. The solution to this problem is shown in Figure 18.2,wherein modifying the loss function to increase the weighting of points close to a surface effectively balances the pointset during training, which significantly improves performance when reconstructing sparse geometries using implicit field methods. Once a maximum likelihood probabilistic 3D model has been generated, improved performance can be obtained by undertaking some data augmentation on the model, as described previously with respect to Figure 21. Data augmentation improves performance of computer vision models by reducing overfitting. Typically, this is done by stochastically applying transformations to input images in 2D space, such as scaling, cropping, shearing, and warping. 3D data augmentation can be used to further improve the performance of 2D-to-3D reconstruction models. In the present example, once a maximum likelihood of probabilistic 3D model has been obtained it can be represented as a point cloud, which can then in turn be usedas an input to 3D data augmentations. The point cloud can – for example – begenerated from a watertight surface mesh representation offline. An example suchpoint cloud is shown in Figure 20.1. Having obtained this point cloud representation, 3D data augmentations can then be applied efficiently to this point cloud on-the-fly during training using optimised matrix operations on a GPU. Example augmentations P604279PC00 that may be made include motion modelling, contrast agent transport, and arbitrary 3D affine transformations, as shown in Figure 20.2.Following 3D data augmentation, an image is formed by casting the point cloud to aplane using efficient matrix operations on the GPU, as shown in Figure 20.3. Multipledifferent 2D images can be obtained by casting the point cloud to any desired different imaging plane. With the above arrangement, therefore, an accurate 3D model of internal anatomy can be probabilistically determined from multiple input 2D images. Once the 3D model has been obtained it can be augmented if necessary, and 2D images of the anatomy can be generated by casting the 3D model onto any desired imaging plane. Moreover, the 3D model itself can also be used directly to find further clinical measurements relating to the subject. For example, where the 3D model is of acoronary artery, the found maximum likelihood 3D model can then be used to findcomputed haemodynamic measurements such as the virtual fractional flow reserve (vFFR). As noted previously in the introduction, use of the vFFR has been demonstrated to be an effective indicator for clinical decision making [6, 7, 8, 9, 10] and has become routine practice in the UK

[0011] and elsewhere

[0012] , but heretofore there has been no clinically-adopted method to calculate vFFR directly from 2D X- ray angiography images. The present disclosure addresses that precise issue, and permits calculation of the vFFR from the found maximum likelihood probabilistic model (where the model is of the coronary artery). The finding of such vFFR measurements using the probabilistic techniques of the present disclosure thus permits for non-invasive testing (no pressure sensor is physically needed within the coronary artery) to find vFFR, which in turn can more readily inform a physician about whether a stent is needed to be inserted into the artery or not, or to guide other treatment options. It will be appreciated that embodiments described herein relate to both the training of an encoder-decoder neural network and to the use of (also referred to herein as deployment of) such an encoder-decoder neural network. While Figure 21 shows an overarching method incorporating aspects of both, it will be recognised that the training and deployment phases can be performed separately, as part of distinct processes. For example, the training could be performed on one device (e.g., in a remote server) while the deployment could be performed on a different device (e.g., a local device). P604279PC00 The training phase typically involves data collection, neural network training, a surface-focused loss function, and the use of at least one data augmentation process, whereas the deployment phase typically pertains to image acquisition, single-pass 2D to 3D reconstruction, the implicit field, and clinical application. The two procedures, while related, can be divided, in generalised terms, as follows: Training Phase: 1. Data Collection and Preprocessing: Synthetic and real 2D X-ray images ofcoronary arteries are collected. Synthetic data is, for example, generated using physiological angiogenesis methods to create a variety of anatomical models. 2. Neural Network Training: An encoder-decoder neural network is trained. Theencoder converts 2D images into latent space representations, capturing salient features of the anatomy. The decoder reconstructs the 3D anatomyfrom these latent representations. 3. Loss Function Optimisation: A surface-focused loss function is used toprioritise the accuracy of surface structures over volumetric consistency, addressing the sparse nature of vascular structures. 4. Data Augmentation: On-the-fly data augmentation techniques, includingmotion modelling and affine transformations, are applied during training to improve the robustness of the network to clinical conditions. Figure 22 shows a generalised method 2200 relating to the training phase, which is performed by certain embodiments of the present disclosure. In particular, Figure 22 shows a method 2200 of generating a 3D virtual model of internal anatomy of asubject, the method 2200 comprising: receiving (s22.1) a plurality of 2D medicalimages of the internal anatomy of the subject that is to be modelled, the 2D medical images having been obtained from at least 2 different imaging angles orplanes; inputting (s22.2) the plurality of 2D medical images into a trained encoder-decoder neural network, the encoder-decoder neural network having been trained to compare the plurality if 2D input images with a universe of possible virtual 3D internal anatomical models representing the internal anatomy shown in the 2Dmedical images; finding (s22.3), using the encoder-decoder neural network, amaximum likelihood virtual 3D model from the universe of possible virtual 3D internal anatomical models that probabilistically most closely matches the internal anatomy of the subject as represented in the plurality of 2D medical images input P604279PC00to the encoder-decoder neural network; and outputting (s22.4) the maximumlikelihood virtual 3D model as representative of the internal anatomy of thesubject. The method 2200 can include various of the elements 1-4 of the TrainingPhase discussed above, or any other element described herein. Deployment Phase: 1. Image Acquisition: Multiple 2D X-ray images of a patient's coronary arteriesare captured from different angles using a clinical X-ray apparatus.2. 3D Reconstruction: The trained encoder-decoder network processes theseimages to generate a detailed 3D model of the coronary anatomy. 3. Output and Analysis: The 3D model is, for example, used to compute vFFRand can be visualised or analysed further for clinical decision-making. One advantage of the present disclosure relates to the reconstruction process and specifically to the fact that the output is a 3D model.Figure 23 shows a generalised method 2300 relating to the deployment phase, which is performed by certain embodiments of the present disclosure. In particular, Figure23 shows a method 2300 of training an encoder-decoder neural network, comprising:receiving (s23.1) a plurality of 2D medical images of the internal anatomy of a plurality of subjects, the 2D medical images each having been obtained from at least2 different imaging angles or planes; and training (s23.2) the encoder-decoderneural network, wherein training the encoder-decoder neural network comprises: training (s23.3) an encoder neural network to convert the 2D medical image data into latent space data; and training (s23.4) a decoder neural network to reconstruct a virtual 3D model of the subject’s internal anatomy from the latent space data, wherein the latent space data is of a lower dimensionality than the 2D medical image data and the virtual 3D model. The method 2300 can include various of the elements 1-3 of the Deployment Phase discussed above, or any other element described herein. Detailed Description of the Embodiments A more detailed and quantitative description of the embodiments will now be described below, with reference to Figures 1 to 12.We define angiographic reconstruction as the recovery of a maximum-likelihoodcoronary anatomy given the available observations (section 3.1). The anatomy is P604279PC00represented as an implicit field neural network (section 3.2) trained with a surfacefocused reconstruction loss function (section 3.3). An online data synthesis routinesupports training despite the lack of multimodal clinical data (section 3.4). Unlikeexisting methods, reconstruction can be performed in a single step from rawangiography images, without the need for preprocessing such as 2D segmentationor centreline detection prior to the 3D reconstruction and the output is a fullyvolumetric representation of the coronary arteries with arbitrary precision.3.1 Probabilistic reconstruction with encoder-decoder neural networksLet Y be the space of all possible 3D shapes and X be the space of all possible 2D X-ray image sequences. Also let a shape y ∼ Y be the unknown coronary anatomy of areal patient and the image sequence X ∼ X be acquired of that patient during clinicalcoronary angiography. We define 2D-to-3D reconstruction as the 3D geometry, ˆy where P(y|X) is the conditional probability of some shape y being the true geometrygiven the observations X.According to Bayes’ theorem, this is identical towhere P(X|y) can be considered to model the stochastic image forming process andP(y) models the (unconditional) prior distribution of feasible coronary anatomies.Lu and Lu

[0056] show that an arbitrarily large neural network with a suitable activationfunction can approximate any probability distribution to an arbitrary degree ofprecision. Therefore, we use a neural network to model the process of reconstructionas z = fe(X) (3) y = fd(z) (4) P604279PC00where fe is an encoder network which models the data-dependent distribution P(X|y),fd is a decoder network which models the unconditional distribution P(y) (discussedfurther in section 3.2), and z is a low-dimensional latent embedding of the observeddata X in order to make the computation tractable. 3.1.1 AdvantagesNon-robustness to realistic clinical conditions is a major unsolved challenge inautomatic angiographic reconstruction. Traditionally, angiographic reconstruction isframed as the recovery of depth for key 2D points in X-ray images by matching themacross different views and applying physical constraints based on epipolar geometry.This has two main limitations: first, the constraints used are fixed, meaning theycan’t adapt or improve during the learning process of the model; second, defining2D points in space becomes problematic in certain scenarios, such as when only oneimage is available or when matching points between two images is challenging dueto factors such as patient motion, image noise, vessel occlusion, or contrast agenttransport.Conversely, under our formulation in eq. (1), there exists a unique and well-definedsolution to reconstruction as long as P(X|y) and P(y) exist, which we call the data-dependent and unconditional components respectively. The challenging conditions ofthe realistic clinical environment can therefore be accounted for by simply trainingsuch a probabilistic model on realistic data without the need to increase thecomplexity in the design of the system.3.2 Representing coronary anatomy as an implicit fieldThe formulation in eq. (3) is agnostic of any particular choice of representation ofthe geometry y, which include pointclouds, voxel grid, mesh, depth map, etc.Equally, the decoder architecture can be of any form which does not violate theassumptions of Lu and Lu

[0056] .In this work, we use an implicit representation of y in which fd is used as a parametricrepresentation of a scalar field, such that ˆs = fd(t, z) : s ∈ R, t ∈ R3(5) P604279PC00The implicit representation is then used to recover the explicit geometry ˆy for a newobservation set x as ˆy = M({ fd(t, fe(x)) : t ∈ T}) (6)Where M is some pointset-to-mesh scheme such as the modified marching cubesalgorithm by Mescheder et al.

[0053] and T is a set of query points, for example auniform 3D grid across the domain. The resulting mesh can then be used for anydownstream clinical task, such as visualisation or haemodynamic modelling. Weinvestigate both the continuous signed-distance field, where d ∈ R specifies thedistance of any point t from the nearest surface of the geometry, and the binaryoccupancy field, where o ∈ {0, 1} specifies whether any point t is inside or outsideof the geometry.3.3 Surface-focused loss functionCoronary vascular anatomy is characterised by its thin, sparse structures, in which asubstantial majority of the bounding volume of the geometry is empty space andsmall errors in surface location result in large changes to the recovered shape (andhence, its haemodynamics). It is therefore important that training the implicitreconstruction network is focused on regions where t is closer to the surface byweighting reconstruction errors in these regions as being more important than thosefar from the surface.In the case of occupancy field reconstruction, focal loss

[0057] has this surface-focusing effect, with where o = f (t) is the true (or observed) value for the occupancy at position t, ˆo = f (t) is its predicted value according to eq. (5) and γ is a hyperparameter used to scale the focusing effect of the loss function. To achieve a similar effect in signed-distance field reconstruction, we introduce the novel loss function P604279PC00 where Lp is the p−norm, p is the order of the p−norm, and α is a hyperparameterused to scale the focusing effect of the loss function.3.4 Generating training dataSynthetic coronary angiography data is generated in three stages (fig. 2). First, adirected acyclic graph representation containing geometry and connectivityinformation for the coronary network is generated stochastically using a physiologicalangiogenesis method based on VascuSynth

[0058] . Next, a 3D surface mesh isgenerated by inflating the centreline graph using node radius data, providing theground-truth 3D representation for anatomy reconstruction (examples shown in fig.3). Finally, this mesh is rendered on-the-fly using a fast-approximation of X-rayformation which allows for train-time 3D data augmentation (examples shown in fig.4).The use of synthetically generated vascular anatomies allows for the creation of anarbitrarily large dataset with controllable complexity, disease status, and diversity.The addition of a mechanism which produces 2D rendered images further adds twovaluable properties to the solution. Firstly, a model for the realistic clinical effectssuch as motion or contrast agent transport can be applied to the shape betweenframes, which allows the model to learn to produce a reconstruction under suchcircumstances without modification of the reconstruction algorithm. Secondly, itmeans that clinical datasets containing only the 3D component can still be used fortraining, which greatly increases the number of clinical data sources to be used tofurther train the system. 3.5 Experimental procedureWe generate a single dataset of synthetic data as defined in table 1 to use for trainingand evaluation and randomly split the data to produce separate training, validation,and test sets with a ratio of 80:10:10. P604279PC00During training and testing, a pair of orthogonal, 256 × 256- pixel images aregenerated using our on-the-fly X-ray rendering procedure. The images are depth-stacked to produce an input tensor with dimension 2 × 256 × 256, which is then fedinto a ResNet34 convolutional encoder network. The only modification made to theencoder is to remove the final classification layer and replace it with a n-dimensionallinear layer and ReLU non-linearity to generate an n-dimensional latent embedding,z, of each imageset, with n = 256 for all experiments. The output z is then passedinto the decoder.To test the implicit field reconstruction as described in section 3.2 we use the ONetarchitecture presented in Mescheder et al.

[0053] for occupancy field reconstructionand the SDFNet architecture presented in Thai et al.

[0055] for signed-distance fieldreconstruction. Both architectures are almost identical: first, a 1D convolutional layeris used to upsample the number of features of input query points from 3 to 256,followed by a series of conditional batch normalisation blocks. We use the conditionalresidual batch normalisation block, which consists of conditional batch normalisationfollowed by a 1D convolution along the feature dimension with ReLU activation, whichis then repeated again on the output before adding the outputs from each of the twostages to yield a residual skip connection. Conditional batch normalisation itself isperformed by calculating the divisor and summand of the normalisation using twoseparate 1D convolutional layers, the weights of which are included in optimisationduring training. Following the conditional batch normalisation layers, a final 1Dconvolution layer is used to reduce the 256 features per point to a single outputfeature. Only at this point do the two methods diverge: for occupancy fieldreconstruction we use a sigmoid output layer to create a valid probability output,while for signed-distance field reconstruction we use a tanh output layer to create anormalised signed distance. Finally, we calculate loss according to eqs. (7) and (8)for the occupancy field and signed-distance field methods respectively, with γ = 2,α = 5, and p = 2. P604279PC00We use several evaluation metrics to assess reconstruction quality: Hausdorffdistance (HD) and Chamfer distance (CD), which should be minimised, and the F1score at three different thresholds (F1@τ) and Intersection-over-union (IoU), whichshould be maximised.In all experiments, we use a batch size of 512, an Adam optimiser with an initiallearning rate of 1 × 10−3, β1 of 0.9, and β2 of 0.999. All trainable networkparameters are randomly initialised. We schedule a one-half learning rate reductionwhen the validation loss plateaus with a patience of 10, a minimum value of 1 × 10−4 and no cool-down. We train each model until the validation loss does not increaseany further and select the model with the minimum validation loss for test andinference. 4 The performance of implicit field reconstructionThis work represents the first application of implicit field reconstruction to the taskcoronary angiography reconstruction. To assess the suitability of this approach, wecompare the reconstruction quality achieved when using the implicit field decoderswith that of two other state-of-the-art methods for reconstruction of voxel grids andpointclouds. In particular, we compare the implicit field decoders with the R2N2 voxelgrid decoder

[0048] and the PSGN pointcloud decoder

[0051] , using the sameconfigurations presented in the original publications. 4.1 ResultsResults are shown quantitatively in table 2 and example reconstructions for theimplicit field reconstruction methods are visualised in figs. 5 and 6. Implicit fielddecoders generate superior reconstructions in all metrics, with the only exceptionbeing an equal F1@0.05 between ONet and PSGN when reconstructing the simplerbifurcation category. This effect becomes greater as the vessel complexity increases,with the signed-distance reconstruction network achieving the best performanceoverall. A qualitative assessment of the occupancy field reconstruction model alsoshows a general ability for both implicit field reconstruction networks to capture mostof the vascular anatomy, despite the complexity of the vessel tree. In particular, theyoften successfully recover a good approximation of the high curvature at vesseljunctions, suggesting implicit fields can represent realistic coronary artery trees withhigh fidelity. Despite this, there are some regions where the reconstruction anatomy P604279PC00may be discontinuous or missing vessel segments, with such effects becominggreater as the vessel anatomy becomes increasingly complex. Table 2: Comparative results of the investigated reconstruction methods with twoorthogonal, X-ray rendered images as input when trained and evaluated on theVessel dataset. The splits are also provided for the bifurcation (bf.), doublebifurcation (db.), coronary (co.), and mean categories within the dataset. Results arepresented in terms of Hausdorff distance (HD), Chamfer distance (CD), F1 score atthree different thresholds (F1@τ), and Intersection-over-union (IoU). In eachcategory, the best performing method is highlighted in bold font. A down-arrow (↓)indicates a low value is better, while an up-arrow (↑) indicates a higher value is better 5 The effect of the surface-focused loss functionTo test the impact of the surface-focused loss function, we also separately trainimplicit field reconstruction models with alternative loss functions. In the case of theoccupancy field decoder, we use the standard binary cross-entropy loss function. Forthe signed-distance field decoder, we investigate both a standard L2-loss and a‘clamped’-L2 loss as P604279PC00which focuses training on regions near the surface using a different method than eq.(8). Table 3 below shows that the choice of loss function significantly affectsreconstruction performance, with both occupancy field and signed-distance fieldreconstruction methods showing much better reconstruction quality when such a lossfunction is used.The use of the hyperbolic p−norm (from eq. (8)) generally leads to marginallyimproved reconstruction performance. However, fig. 7 shows that the hyperbolic p−norm loss function is more robust to suboptimal choice of the α hyperparameterthan the clamped L2 loss function (eq. (9)).Table 3: Comparative reconstruction results between implicit decoders with differentloss functions when trained and evaluated on the Vessel dataset. The splits are alsoprovided for the bifurcation (bf.), double bifurcation (db.), coronary (co.), and mean categories within the dataset. Results are presented in terms of Hausdorff distance(HD), Chamfer distance (CD), F1 score at three different thresholds (F1@τ), andIntersection-over-union (IoU). For all clamped (Clamp) models α = 5 and for allhyperbolic p-norm (Hyp) model, a scaled hyperbolic L2 norm with α = 5 is used.6 Simulating camera pose uncertainty and patient motion P604279PC00Despite routine calibration of angiography equipment in clinics, there is uncertaintyin the system from extrinsic and intrinsic sources. Imperfect calibration of the C-armintroduces intrinsic uncertainty in the camera pose and state-of-the-art methods toaccount for this calibration uncertainty are currently limited to iterative optimisationschemes based on deterministic mechanisms

[0059] . Another major source of extrinsicuncertainty is patient motion, including body motion, respiratory motion, and cardiacmotion. This is particularly impactful in monoplane angiography, which requires atleast two acquisition phases separated by camera positioning and, consequently,time. While cardiac motion can be mitigated by electrocardiogram (ECG) gating

[0060] ,it is not always possible to ensure breath-hold nor immobilisation for cardiac patients,ostensibly for patient comfort and their ability to tolerate such protocols

[0061] . In thisexperiment, we measure the degradation in reconstruction quality subject to bothimperfect equipment calibration and simulated patient motion.6.1 MethodsTo model calibration uncertainty, we add Gaussian noise of scale σ to camera poseprior to rendering. This is applied during both training and testing and means thatthe view separation angle is different for every sample and is also recomputed eachepoch. To measure the effect of calibration uncertainty, a new model is trained andevaluated for each of a series of values of σ.We model patient motion in two different forms: first, as a rigid translation, whichwe perform in 3D by translating the anatomy along a random vector rt = δˆrt, whereˆrt ∈ R3 is a random unit vector and δ ∈ [0, Δ] is a random uniform variable. Tomeasure the effect of this rigid transformation model of motion, a new model istrained and evaluated for each of a series of values of Δ.A second rigid motion model is also investigated, modelled as a random rotationaround a random point within the bounding cuboid of the geometry. Again, this isperformed in 3D by translating the anatomy along a random unit vector ˆrt ∈ R3,before rotating by a uniformly random angle θ ∈ [0,Θ] around a second random unitvector with components ˆrx, ˆry, ˆrz, such that, for a point p the new position p′ isgiven by P604279PC00 where c = cos(θ) and s = sin(θ). Following the translation or rotation, the image isacquired using the same rendering method used before. In all cases, the first imagein the set is without uncertainty or motion, since it defines the canonical location andorientation of the target geometry. 6.2 ResultsThe effect of calibration uncertainty is shown in fig. 8. In most cases, the effect ofcalibration uncertainty is either insignificant or even positive, with improvedreconstruction in all metrics when σ = 1◦, and improved HD and F1@0.05 when σ= 10◦. However, when σ > 10◦ the reconstruction quality is significantly reduced.The effect of rigid translation is shown in fig. 9. Adding rigid translation of any degreeactually increases the reconstruction quality in all metrics apart from IoU.The effect of rigid rotation is shown in fig. 10. Across all metrics, there is a significanttrend of worsening performance for larger rotation ranges Θ. However, in terms ofHD, CD, F1@0.02, and F1@0.05, reconstruction quality is significantly improvedwhen Θ = 5, with a smaller but still significant improvement also shown in CD andF1@0.05 when Θ = 10. 7 Non-orthogonal input imagesThe earlier experiments in this work have so far used orthogonal separation, whichis essentially equivalent to biplane angiography. However, due to the lower dosageand wider availability of equipment, coronary angiography is also often performedwith monoplane angiography. In such cases, the separation angle between two (ormore) images acquired during a sequence may not be: 1. orthogonal nor; 2.consistent.In this experiment, we investigate the effects of non-orthogonal image separationand inconsistent image separation. P604279PC00 7.1 MethodsWe specify the view separation angle as a hyperparameter for a series of independentexperiments, training and testing a specific model until convergence for each choiceof angle separation. To test the effect of non-orthogonality, we sweep in incrementsof 20◦ starting at 10◦ until 170◦ in order to coarsely evaluate the selection of viewseparation angles, in addition to a finer sweep at 2◦ intervals between 2◦ and 10◦.To test the effect of inconsistent angle separation, we use a fixed viewpoint for thefirst image in the set, while the second image in the set is allowed to vary throughoutthe remainder of the angular domain. We randomly sample the two angularcoordinates for the second image using a uniform spherical sampling method suchthat where A1 ∼ [0, 2π] and A2 ∼ [−1, 1] are uniform variates.Commercial angiography imaging devices record the relative positioning of theapparatus for each frame of imaging, so camera pose angles for each frame in theangiography sequence are available in the clinic. These angles are incorporated intothe reconstruction model as a pair of features for each camera pose angle θi 7→(cos(θi), sin(θi)), which is continuous for all θi ∈ R and unique for all θi ∈ [0, 2π). Since the C-arm apparatus has two angular degrees of freedom (θ and ϕ), thereare four angle features per image. The angle features are then concatenated withthe latent embedding z (see eq. (3)) before being passed to the decoder.7.2 ResultsFigure 12 shows that there is generally a worse reconstruction performance whenthe view angle separation is not 90◦, though the difference in quality is generallysmall; for example, at 8◦ separation angle the F1@0.01 is reduced by 0.0342 andIoU by 0.0165. Even with view separation angles as low as 2◦ separation angle, thequality is still reasonably good, with HD of 0.1880, CD of 0.0266, and IoU of 0.2124.The effect of using inconsistent camera pose angle is shown in table 4 below. Withthe exception of F1@0.05, when the second image is taken from an arbitraryorientation there is a reduction in reconstruction quality. The addition of encoded P604279PC00angles does improve performance in terms of CD, F1@0.01, F1@0.02, and F1@0.05for the random mode. However, the impact is small (ΔF1@0.01 is +0.0094) andinsufficient to meaningfully offset the drop in performance when moving away fromthe consistent biplane configuration. Despite this, a visualisation of the reconstructedmeshes in fig. 11 shows that even in a random monoplane setting without the useof angle information, ONet can learn to reconstruct the relatively simple geometriesof the bifurcation and double bifurcation categories of the Vessel dataset. The qualityis negatively impacted on more complex geometries, with a noticeable smearing ofthe vessel in the reconstruction. Table 4: The difference in median reconstruction metrics of the ONet method whenadding inconsistency to view angles, with and without angle encoding. For example,for the item Random vs Orthogonal, each value is calculated by subtracting the valuefor the Random configuration from the corresponding value with the Orthogonalconfiguration 8 ConclusionAutomatic reconstruction of coronary arteries from 2D angiography sequencespresents the chance to radically improve the effectiveness of treatment for patientswith coronary artery disease. Current methods are limited due to their basis inepipolar backprojection across frames, which is prone to failure due to patientmotion, image noise, and vessel foreshortening. In this paper, we demonstrated analternative approach which reframes reconstruction as a probabilistic process andhence motivates a more generalisable reconstruction algorithm. Using an implicitfield reconstruction neural network, trained on synthetic data with a surface-focusedloss function, we are able to achieve high-quality reconstructions even under theclinically-realistic conditions of (simulated) patient motion and unknown camerapose. Various modifications whether by way of addition, deletion, or substitution of features may be made to above described embodiment to provide further embodiments, any and all of which are intended to be encompassed by the appended claims. P604279PC00 References[1] Sida Jia, Yue Liu, and Jinqing Yuan. Evidence in guidelines for treatment ofcoronary artery disease. Coronary Artery Disease: Therapeutics and Drug Discovery,pages 37–73, 2020.[2] Authors / Task Force members, Philippe Kolh, Stephan Windecker, FernandoAlfonso, Jean-Philippe Collet, Jochen Cremer, Volkmar Falk, Gerasimos Filippatos,Christian Hamm, Stuart J Head, et al. 2014 esc / eacts guidelines on myocardialrevascularization: the task force on myocardial revascularization of the europeansociety of cardiology (esc) and the european association for cardiothoracic surgery(eacts) developed with the special contribution of the european association ofpercutaneous cardiovascular interventions (eapci). European journal of cardio-thoracic surgery, 46(4):517–592, 2014.[3] Pim AL Tonino, Bernard De Bruyne, Nico HJ Pijls, Uwe Siebert, Fumiaki Ikeno,Marcel vant Veer, Volker Klauss, Ganesh Manoharan, Thomas Engstrøm, Keith GOldroyd, et al. Fractional flow reserve versus angiography for guiding percutaneouscoronary intervention. New England Journal of Medicine, 360(3):213–224, 2009.[4] Gabor G Toth, Balint Toth, Nils P Johnson, Frederic De Vroey, Luigi Di Serafino,Stylianos Pyxaras, Dan Rusinaru, Giuseppe Di Gioia, Mariano Pellicano, EmanueleBarbato, et al. Revascularization decisions in patients with stable angina andintermediate lesions: results of the international survey on interventional strategy.Circulation: Cardiovascular Interventions, 7(6):751–759, 2014.[5] BCIS Audit Report for 2019-20. British Cardiovascular Intervention Society, 2020.[6] Bernard De Bruyne, Nico HJ Pijls, Bindu Kalesan, Emanuele Barbato, Pim ALTonino, Zsolt Piroth, Nikola Jagic, Sven Möbius-Winkler, Gilles Rioufol, Nils Witt, etal. Fractional flow reserve–guided pci versus medical therapy in stable coronarydisease. New England Journal of Medicine, 367(11):991–1001, 2012.[7] Paul D Morris, Desmond Ryan, Allison C Morton, Richard Lycett, Patricia VLawford, D Rodney Hose, and Julian P Gunn. Virtual fractional flow reserve fromcoronary angiography: modeling the significance of coronary lesions: results fromthe virtu-1 (virtual fractional flow reserve from coronary angiography) study. JACC: P604279PC00Cardiovascular Interventions, 6(2):149–157, 2013. Direct CFD vFFR, >24 hours tocalculate.[8] Bjarne L Nørgaard, Jonathon Leipsic, Sara Gaur, Sujith Seneviratne, Brian S Ko,Hiroshi Ito, Jesper M Jensen, Laura Mauri, Bernard De Bruyne, Hiram Bezerra, et al.Diagnostic performance of noninvasive fractional flow reserve derived from coronarycomputed tomography angiography in suspected coronary artery disease: the nxttrial (analysis of coronary blood flow using ct angiography: Next steps). Journal ofthe American College of Cardiology, 63(12):1145–1155, 2014.[9] Frederik M Zimmermann, Angela Ferrara, Nils P Johnson, Lokien X Van Nunen,Javier Escaned, Per Albertsson, Raimund Erbel, Victor Legrand, Hyeong-Cheol Gwon,Wouter S Remkes, et al. Deferral vs. performance of percutaneous coronaryintervention of functionally nonsignificant coronary stenosis: 15-year follow-up of thedefer trial. European heart journal, 36(45):3182–3188, 2015.

[0010] Adriaan Coenen, Young-Hak Kim, Mariusz Kruk, Christian Tesche, Jakob DeGeer, Akira Kurata, Marisa L Lubbers, Joost Daemen, Lucian Itu, Saikiran Rapaka, etal. Diagnostic accuracy of a machine-learning approach to coronary computedtomographic angiography–based fractional flow reserve: result from the machineconsortium. Circulation: Cardiovascular Imaging, 11(6):e007217, 2018. Validationof @itu2016machine. CTA-based lesion classification. Clinical data (n=351) is semi-automatically segmented. 99.7% correlation between CFD and ML was achieved, and62% between ML and invasive FFR. Relatively weak results but authors focus on itsimprovement vs CTA alone.

[0011] Medical Technologies Guidance. Heartflow ffrct for estimating fractional flowreserve from coronary ct angiography. NICE UK.

[0012] Juhani Knuuti and Valeriu Revenco. 2019 esc guidelines for the diagnosis andmanagement of chronic coronary syndromes. European heart journal, 41(5):407–477, 2020.

[0013] Uwe Jandt, Dirk Schäfer, Michael Grass, and Volker Rasche. Automaticgeneration of 3d coronary artery centerlines using rotational x-ray angiography.Medical image analysis, 13(6):846–858, 2009.

[0014] Albert KW Law, Hui Zhu, and Francis HY Chan. 3d reconstruction of coronaryartery using biplane angiography. In Proceedings of the 25th Annual International P604279PC00Conference of the IEEE Engineering in Medicine and Biology Society (IEEE Cat. No.03CH37439), volume 1, pages 533–536. IEEE, 2003.

[0015] Jia Li and Laurent D Cohen. Reconstruction of 3d tubular structures from cone-beam projections. In 2011 IEEE International Symposium on Biomedical Imaging:From Nano to Macro, pages 1162–1166. IEEE, 2011.

[0016] Serkan Çimen, Ali Gooya,Michael Grass, and Alejandro F Frangi. Reconstruction of coronary arteries from x-ray angiography: A review. Medical image analysis, 32:46– 68, 2016.

[0017] Kenneth R Hoffmann, Anindya Sen, Li Lan, Kok-GeeChua, Jacqueline Esthappan, and Marco Mazzucco. A system for determination of 3dvessel tree centerlines from biplane images. The International Journal of CardiacImaging, 16(5):315–330, 2000.

[0018] Zeyu Fu, Zhuang Fu, Zening Gong, Xin Feng, Haoran Gu, Rongli Xie, Jun Zhang,and Jian Fei. Optimization for 3d reconstruction of coronary artery tree by two-stagelevenberg-marquardt algorithm. In 2021 27th International Conference onMechatronics and Machine Vision in Practice (M2VIP), pages 84–89. IEEE, 2021.

[0019] Panagiota I Tsompou, Ioannis O Andrikos, Georgia S Karanasiou, Antonis ISakellarios, Nikolaos Tsigkas, Vassiliki I Kigka, Savvas Kyriakidis, Lampros KMichalis, and Dimitrios I Fotiadis. Validation study of a novel method for the 3dreconstruction of coronary bifurcations. In 2020 42nd Annual InternationalConference of the IEEE Engineering in Medicine & Biology Society (EMBC), pages1576–1579. IEEE, 2020.

[0020] Francesca Galassi, Mohammad Alkhalil, Regent Lee, Philip Martindale, Rajesh KKharbanda, KeithMChannon, Vicente Grau, and Robin P Choudhury. 3dreconstruction of coronary arteries from 2d angiographic projections using non-uniform rational basis splines (nurbs) for accurate modelling of coronary stenoses.PloS one, 13(1):e0190650, 2018.

[0021] Mathias Unberath, Oliver Taubmann, Michaela Hell, Stephan Achenbach, andAndreas Maier. Symmetry, outliers, and geodesics in coronary artery centerlinereconstruction from rotational angiography. Medical Physics, 44(11):5672–5685,2017. P604279PC00

[0022] Christian Kromm and Karl Rohr. Inception capsule network for retinal bloodvessel segmentation and centerline extraction. In 2020 IEEE 17th internationalsymposium on biomedical imaging (ISBI), pages 1223–1226. IEEE, 2020.

[0023] Giles Tetteh, Velizar Efremov, Nils D Forkert, Matthias Schneider, Jan Kirschke,Bruno Weber, Claus Zimmer, Marie Piraud, and Björn H Menze. Deepvesselnet:Vessel segmentation, centerline prediction, and bifurcation detection in 3-dangiographic volumes. Frontiers in Neuroscience, page 1285, 2020.

[0024] Jiafa He, Chengwei Pan, Can Yang, Ming Zhang, Yang Wang, Xiaowei Zhou, andYizhou Yu. Learning hybrid representations for automatic 3d vessel centerlineextraction. In International Conference on Medical Image Computing and Computer-Assisted Intervention, pages 24–34. Springer, 2020.

[0025] Shanhui Sun, Stefan Kluckner, Ahmet Tuysuzoglu, Ankur Kapoor, GünterLauritsch, and Terrence Chen. 3-d vessel tree surface reconstruction method, April 32018. US Patent 9,934,566.

[0026] DM Bappy, Ayoung Hong, Eunpyo Choi, Jong-Oh Park, and Chang-Sei Kim.Automated three-dimensional vessel reconstruction based on deep segmentation andbi-plane angiographic projections. Computerized Medical Imaging and Graphics,92:101956, 2021.

[0027] Yaosong Jia, Deqiang Xiao, Qing Yan, and Mingwei Gao. A method forreconstructing 3d skeleton of coronary artery from 2d x-ray angiographic images. In2022 4th International Conference on Intelligent Medicine and Image Processing,pages 70–75, 2022.

[0028] Ali Iskurt, Yasar Becerikli, and Kamran Mahmutyazicioglu. A fast and automaticcalibration of the projectory images for 3d reconstruction of the branchy structures.In 201347th Annual Conference on Information Sciences and Systems (CISS), pages1–6. IEEE, 2013.

[0029] Abhirup Banerjee, Francesca Galassi, Ernesto Zacur, Giovanni Luigi De Maria,Robin P Choudhury, and Vicente Grau. Point-cloud method for automated 3dcoronary tree reconstruction from multiple non-simultaneous angiographicprojections. IEEE transactions on medical imaging, 39(4):1278–1290, 2019.

[0030] Michael Kass, Andrew Witkin, and Demetri Terzopoulos. Snakes: Active contourmodels. International journal of computer vision, 1(4):321–331, 1988. P604279PC00

[0031] Sun Zheng, Tian Meiying, and Sun Jian. Sequential reconstruction of vesselskeletons from x-ray coronary angiographic sequences. Computerized MedicalImaging and Graphics, 34(5):333–345, 2010.

[0032] Sara Gaur, Hiram G Bezerra, Evald H Christiansen, Kentaro Tanaka, Jesper MJensen, Anne K Kaltoft, Hans Erik Botker, Jens F Lassen, Christian J Terkelsen, andBjarne LNorgaard. Reproducibility of invasively measured and noninvasivelycomputed fractional flow reserve. Journal of the American College of Cardiology,63(12S):A999–A999, 2014.

[0033] Roshni Solanki, Rebecca Gosling, Vignesh Rammohan, Giulia Pederzani, PankajGarg, James Heppenstall, D Rodney Hose, Patricia V Lawford, Andrew J Narracott,John Fenner, et al. The importance of three dimensional coronary arteryreconstruction accuracy when computing virtual fractional flow reserve from invasiveangiography. Scientific reports, 11(1):1–12, 2021.

[0034] Kritika Iyer, Brahmajee K Nallamothu, C Alberto Figueroa, and Raj R Nadakuditi.A multi-stage neural network approach for coronary 3d reconstruction fromuncalibrated x-ray angiography images. 2023.

[0035] Lu Wang, Dong-xue Liang, Xiao-lei Yin, Jing Qiu, Zhiyun Yang, Jun-hui Xing, Jian-zeng Dong, and Zhao-yuan Ma.Weakly-supervised 3d coronary artery reconstruction from two-view angiographicimages. arXiv preprint arXiv:2003.11846, 2020.

[0036] Jingyi Zuo. 2d to 3d neurovascular reconstruction from biplane view via deeplearning. In 2021 2nd International Conference on Computing and Data Science(CDS), pages 383–387. IEEE, 2021.

[0037] Minki Hwang, Sa-Bin Hwang, Hyosang Yu, Jaehyeok Kim, Daehyun Kim, WonjaeHong, Ah-Jin Ryu, Han Yong Cho, Jinlong Zhang, Bon Kwon Koo, et al. A simplemethod for automatic 3d reconstruction of coronary arteries from x-ray angiography.Frontiers in Physiology, 12, 2021.

[0038] Athanasios Vlontzos, Samuel Budd, BenjaminHou, Daniel Rueckert, and Bernhard Kainz. 3d probabilistic segmentation andvolumetry from 2d projection images. In International Workshop on Thoracic ImageAnalysis, pages 48–57. Springer, 2020.

[0039] Carlo Biffi, Juan J Cerrolaza, Giacomo Tarroni, Antonio de Marvao, Stuart ACook, Declan P O’Regan, and Daniel Rueckert. 3d high-resolution cardiacsegmentation reconstruction from 2d views using conditional variational P604279PC00autoencoders. In 2019 IEEE 16th International Symposium on Biomedical Imaging(ISBI 2019), pages 1643–1646. IEEE, 2019.

[0040] David Stojanovski, Uxio Hermida, Marica Muffoletto, Pablo Lamata, Arian Beqiri,and Alberto Gomez. Efficient pix2vox++ for 3d cardiac reconstruction from 2d echoviews. arXiv preprint arXiv:2207.13424, 2022.

[0041] Alejandro F Frangi, Julia A Schnabel, Christos Davatzikos, Carlos Alberola-López,and Gabor Fichtinger. Medical Image Computing and Computer AssistedIntervention– MICCAI 2018: 21st International Conference, Granada, Spain,September 16-20, 2018, Proceedings, Part IV, volume 11073. Springer, 2018.

[0042] Pak-Hei Yeung, Linde Hesse, Moska Aliasi, Monique Haak, Weidi Xie, Ana ILNamburete, et al. Implicitvol: Sensorless 3d ultrasound reconstruction with deepimplicit representation. arXiv preprint arXiv:2109.12108, 2021.

[0043] Kristijan Bartol,David Bojani´c, Tomislav Petkovi´c, Nicola D’Apuzzo, and Tomislav Pribanic. Areview of 3d human pose estimation from 2d images. In Int. Conf. and Exhibition on3D Body Scanning and Processing Technologies, 2020.

[0044] Zhiping Zhang. Review of 3d reconstruction technology of uav aerial image. InJournal of Physics: Conference Series, volume 1865, page 042063. IOP Publishing,2021.

[0045] Zhiliang Ma and Shilong Liu. A review of 3d reconstruction techniques incivil engineering and their applications. Advanced Engineering Informatics, 37:163–174, 2018.

[0046] Zhizhong Kang, Juntao Yang, Zhou Yang, and Sai Cheng. A reviewof techniques for 3d reconstruction of indoor environments. ISPRS InternationalJournal of Geo- Information, 9(5):330, 2020.

[0047] Xian-Feng Han, Hamid Laga, and Mohammed Bennamoun. Image-based 3dobject reconstruction: Stateof- the-art and trends in the deep learning era. IEEEtransactions on pattern analysis and machine intelligence, 43 (5):1578–1604, 2019.

[0048] Christopher B Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and SilvioSavarese. 3d-r2n2: A unified approach for single and multi-view 3d objectreconstruction. In European conference on computer vision, pages 628–644.Springer, 2016.

[0049] Jakob Engel, Thomas Schöps, and Daniel Cremers. Lsdslam: Large-scale directmonocular slam. In European conference on computer vision, pages 834–849.Springer, 2014. P604279PC00

[0050] Seong-Woo Kim, Baoxing Qin, Zhuang Jie Chong, Xiaotong Shen, Wei Liu,Marcelo H Ang, Emilio Frazzoli, and Daniela Rus. Multivehicle cooperative drivingusing cooperative perception: Design and experimental validation. IEEE Transactionson Intelligent Transportation Systems, 16(2):663–680, 2014.

[0051] Haoqiang Fan, Hao Su, and Leonidas J Guibas. A pointset generation networkfor 3d object reconstruction from a single image. In Proceedings of the IEEEconference on computer vision and pattern recognition, pages 605–613, 2017.

[0052] Zhiqin Chen and Hao Zhang. Learning implicit fields for generative shapemodeling. In Proceedings of the IEEE / CVF Conference on Computer Vision andPattern Recognition, pages 5939–5948, 2019.

[0053] Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, andAndreas Geiger. Occupancy networks: Learning 3d reconstruction in function space.In Proceedings of the IEEE / CVF conference on computer vision and patternrecognition, pages 4460–4470, 2019.

[0054] Jeong Joon Park, Peter Florence, JulianStraub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuoussigned distance functions for shape representation. In Proceedings of the IEEE / CVFconference on computer vision and pattern recognition, pages 165–174, 2019.

[0055] Anh Thai, Stefan Stojanov, Vijay Upadhya, and James M Rehg. 3d reconstructionof novel object shapes from single images. In 2021 International Conference on 3DVision (3DV), pages 85–95. IEEE, 2021.

[0056] Yulong Lu and Jianfeng Lu. A universal approximation theorem of deep neuralnetworks for expressing probability distributions. Advances in neural informationprocessing systems, 33:3094–3105, 2020.

[0057] Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. Focal lossfor dense object detection. In Proceedings of the IEEE international conference oncomputer vision, pages 2980–2988, 2017.

[0058] Ghassan Hamarneh and Preet Jassi. Vascusynth: Simulating vascular trees forgenerating volumetric image data with ground-truth segmentation and tree analysis.Computerized medical imaging and graphics, 34(8):605–616, 2010.

[0059] Negar Chabi, Domenico Iuso, Oliver Beuing, Bernhard Preim, and SylviaSaalfeld. Self-calibration of c-arm imaging system using interventional instruments P604279PC00during an intracranial biplane angiography. International Journal of ComputerAssisted Radiology and Surgery, 17(7):1355– 1366, 2022.

[0060] V Rasche, A Buecker, M Grass, R Koppe, J Op de Beek, R Bertrams, R Suurmond,H Kuehl, and RW Guenther. Ecg-gated 3d-rotational coronary angiography (3drca).In CARS 2002 Computer Assisted Radiology and Surgery: Proceedings of the 16 thInternational Congress and Exhibition Paris, June 26–29, 2002, pages 827–831.Springer, 2002.

[0061] Merlin J Fair, Peter D Gatehouse, Eliana Reyes, Ganesh Adluru, Jason Mendes,Tina Khan, Ranil de Silva, Rick Wage, Edward VR DiBella, and David N Firmin. Initialinvestigation of free-breathing 3d whole-heart stress myocardial perfusion mri.Global Cardiology Science & Practice, 2020(3), 2020.

[0062] G. Guven, H. F. Ates and H. F. Ugurdag, "X2V: 3D Organ Volume Reconstruction From a Planar X-Ray Image With Neural Implicit Methods," in IEEE Access, vol. 12, pp. 50898-50910, 2024, doi: 10.1109 / ACCESS.2024.3385668.

[0063] Mahesh Shakya, Bishesh Khanal, Benchmarking Encoder-Decoder Architectures for Biplanar X-ray to 3D Shape Reconstruction, https: / / arxiv.org / pdf / 2309.13587

Claims

P604279PC00 Claims1. A method of generating a 3D virtual model of internal anatomy of a subject,comprising: a) receiving a plurality of 2D medical images of the internal anatomyof the subject that is to be modelled, the 2D medical images having been obtained from at least 2 different imaging angles or planes; b) inputting the plurality of 2D medical images into a trained encoder-decoder neural network, the encoder-decoder neural network having been trained to reconstruct a feasible estimate for theinternal anatomy shown in the plurality of 2D medical images froma universe of possible 3D anatomical models; c) finding, using the encoder-decoder neural network, a best estimate3D anatomy from the universe of possible 3D anatomical models that probabilistically most closely matches the internal anatomy of the subject as represented in the plurality of 2D medical images input to the encoder-decoder neural network; and d) outputting the best estimate 3D anatomy as representative of theinternal anatomy of the subject.

2. A method according to claim 1, wherein the finding comprises:i. encoding, using an encoder neural network, the 2D medical image datainto latent space data; and ii. decoding, using a decoder neural network, the latent space data intoa virtual 3D model of the subject’s internal anatomy; wherein the latent space data is of a lower dimensionality than the 2D medical image data and the virtual 3D model.

3. A method according to claim 2, wherein the encoder neural network has beentrained to encode the received 2D imagery into a maximum-information latent code z, being the latent space data, wherein the latent code z is a low-dimensional latent embedding of observed data in the input 2D image sequence, based on a stochastic image forming process of the 2D input images.

4. A method according to claims 2 or 3, wherein the latent space data is inputinto the decoder neural network to condition the decoder on the observedP604279PC00 data in the 2D input images, the decoder having been trained to learn the distribution of feasible internal anatomies, including healthy, diseased, andimpossible anatomies as the universe of possible 3D internal anatomical models, resulting in an output 3D reconstruction that is the best estimate 3Dreconstruction given the input 2D images.

5. A method according to claim 4, wherein the decoder neural network uses animplicit field reconstruction to represent the output best effort 3Dreconstruction.

6. A method according to claims 4 or 5, wherein the latent space data is suchthat it suppresses any 3D reconstruction by the decoder which is implausiblegiven the available information in the 2D input images.

7. A method according to any of the preceding claims, and further comprising:generating a synthetic 2-dimensional image of the internal anatomy of the subject at a desired imaging plane by projecting the best effort 3D model of the internal anatomy to the desired imaging plane.

8. A method according to any of the preceding claims, and further comprisingtransforming the best effort 3D model into a surface and / or volume meshmodel.

9. A method according to any of the preceding claims, wherein the virtual 3Dmodel is of the coronary artery, the method further comprising: using the 3Dmodel, computing one or more haemodynamic measurements that relating to the coronary artery of the subject. 10.A method according to claim 9, wherein the one or more haemodynamic measures is a virtual Fractional Flow Reserve (vFFR) measurement of the coronary artery of the subject.11.A method of training an encoder-decoder neural network, comprising:a. receiving a plurality of 2D medical images of the internal anatomy ofa plurality of subjects, the 2D medical images each having been obtained from at least 2 different imaging angles or planes; andP604279PC00 b. training the encoder-decoder neural network, wherein training theencoder-decoder neural network comprises: i. training an encoder neural network to convert the 2D medicalimage data into latent space data; and ii. training a decoder neural network to reconstruct a virtual 3Dmodel of the subject’s internal anatomy from the latent spacedata, wherein the latent space data is of a lower dimensionality than the 2D medical image data and the virtual 3D model. 12.A method according to claim 11, wherein training the encoder neural network comprises training the encoder to encode the received 2D imagery into a maximum-information latent code z, being the latent space data, wherein the latent code z is a low-dimensional latent embedding of observed data in theinput 2D image sequence, based on a stochastic image forming process of the 2D input images.13.A method according to claim 11 or claim 12, further comprising training thedecoder to learn the unconditional distribution of internal anatomies, includinghealthy, diseased, and impossible anatomies as the universe of possible virtual 3D internal anatomical models, resulting in an output 3D reconstruction that is a best effort 3D reconstruction given a set of input 2Dimages. 14.A method according to claim 13, wherein the decoder neural network uses an implicit field reconstruction to determine the output best effort 3Dreconstruction.15.A method according to any of claims 11 to 14, wherein the latent space datais such that it suppresses any 3D reconstruction by the decoder which isimplausible given the available information in the 2D input images.16.A method according to any of claims 11 to 15, further comprising increasingresolution of the virtual 3D model of the subject’s internal anatomy near to a surface of the anatomy by modifying a loss function to increase the weighting of points in the virtual 3D model close to the surface compared to points that are further away from the surface.P604279PC00 17.A method according to claim 16, wherein the increase in the weighting of points close to a surface effectively balances the set of points in terms ofpoints falling within and outside of the anatomy. 18.A method according to any of the preceding claims, and further comprising undertaking 3D data augmentations as part of the procedure of training of the encoder-decoder model. 19.A method according to claim 18, wherein the data augmentations include one or more selected from the group comprising: motion modelling, contrast agent transport, disease modelling, and arbitrary 3D affine transformations.20.A computer system, comprising:d. one or more processors;e. computer readable memory; andf. a computer readable medium storing computer readable instructionsthat when executed by the processor cause the processor to operate in accordance with the method of any of the preceding claims.21.A medical imaging system comprising:c. a medical imager capable of producing medical images of internalanatomy of a subject to be imaged; and d. a computer system according to claim 20, the computer system beingarranged such that in use it receives 2D medical images of the internal anatomy of the subject from the medical imager, and processes those 2D images according to the method of any of claims 1 to 19 to produce a virtual 3D model of the internal anatomy of the subject. 22.A medical imaging system according to claim 21, wherein the medical imager is an X ray medical imaging system, and more preferably a C-arm X ray medical imaging system.

Citation Information

Patent Citations

  • 3-D vessel tree surface reconstruction method

    US9934566B2

  • Three-Dimensional Shape Reconstruction from a Topogram in Medical Imaging

    US20220028129A1

  • Anatomical and functional assessment of coronary artery disease using machine learning

    US20230368398A1