Method and apparatus for controllable generation of virtual populations of anatomy
The conditional generative framework using latent flow transformations and graph convolutional neural networks addresses the limitations of traditional clinical trials by generating diverse and realistic virtual populations of anatomy, enhancing in-silico testing and design of medical devices and drugs.
Patent Information
- Application Number
- PCT/GB2024/051608
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-24
- Publication Date
- 2026-01-02
AI Technical Summary
Traditional clinical trials for medical devices are limited by patient variability, leading to insufficient validation of safety and efficacy, and costly late-phase failures, while existing in-silico models lack the flexibility to generate diverse and realistic virtual populations of anatomy.
A conditional generative framework using latent flow transformations to synthesize virtual populations of anatomy, constrained by relevant patient attributes, enabling controllable generation of diverse and realistic organ shapes that match target populations, using graph convolutional neural networks and normalizing flows.
Enables the creation of virtual populations that accurately capture the variability and correlations between patient attributes and organ shapes, facilitating more effective in-silico testing and design of medical devices and drugs, reducing the need for costly clinical trials.
Smart Images

Figure GB2024051608_02012026_PF_FP_ABST
Abstract
Description
[0001] METHOD AND APPARATUS FOR CONTROLLABLE GENERATION OF VIRTUAL POPULATIONS OF ANATOMY
[0002] FIELD OF THE DISCLOSURE
[0003] The disclosed embodiments generally relate to a computer-implemented method, computer program and apparatus for controllable generation of virtual populations of anatomy for use in in-silico studies, and the use thereof for medical device or drug design, manufacture, testing and regulatory approval.
[0004] BACKGROUND
[0005] The development costs of medical devices have grown steadily in recent years, driven by lengthy regulatory approval processes and costly clinical trials. The high rate of late-phase failures of medical devices in clinical trials can lead to significant financial losses for medical device manufacturers, and the financial losses incurred by such late-phase failures are often subsequently rolled over onto the much smaller number of successful medical devices that are eventually approved. The resulting financial burden on healthcare systems and the medical device industry worldwide is fast becoming unsustainable, and there is a clear need for a paradigm shift in medical device innovation and regulatory approval processes.
[0006] Traditional (i.e. in-vivo) clinical trials for medical devices are often restricted by the limited variability in operational regimes (e.g. patient variability) that medical devices are able to be subjected to when evaluating their safety and efficacy. This may be due to the difficulty of recruiting sufficient patients that cover the variability of target populations for the medical device under development (i.e. the target patient as intended for use with the developed medical devices in routine clinical care). Insufficient validation of medical device safety and efficacy can lead to serious adverse side effects only becoming apparent post-market approval, particularly when used in operational regimes not assessed adequately in pre-market approval clinical trials.
[0007] However, in-silico trials (ISTs) provide an alternative and complementary route to medical device innovation, and the regulatory approval process, which underpins the route to commercialisation of medical devices and their adoption in routine patient care. ISTs use computational modelling and simulation to evaluate the safety and efficacy of medical devices virtually, in virtual patient populations and provide digital (or in-silico) as opposed to real-world evidence in support of the regulatory approval of the medical devices under development. ISTs thus have the potential to explore medical device performance in a wider range of patient characteristics than would be feasible to recruit for in a real world clinical trial, and hence ISTs can help refine, reduce, and partially replace in-vivo clinical trials for medical device testing.
[0008] SUMMARY
[0009] The technology described herein provides a generative framework that enables simulation of animal, including human, anatomy and physiology, in a controlled manner, also referred to as conditional generation or synthesis, where, the conditional generation is controlled using conditioning variables that are relevant attributes or characteristics in some desired target population. The conditionally generated anatomical and / or physiological parts may then be used for in-silico support and enablement across medical products (i.e. medical devices, pharmaceuticals, etc.) lifecycle (from conception, design, investment, testing, trialling, regulation, health economics to post-market surveillance), which in general terms may be referred to as an in-silico assessment. In particular, there is provided a method and apparatus / system for conditional generation (i.e. synthesis) of plausible single- or multi-part organs (or multi-organ shape assemblies) using shapes of macroscopic anatomical structures represented as unstructured graphs and conditioning variable(s) such as patient demographic information (e.g. age, sex, ethnicity, etc.), physical measurements (e.g. height, weight, body mass index, etc.), clinical measurements (e.g. systolic and diastolic blood pressure, blood glucose and cholesterol, etc.), presence of disease or abnormal health conditions (e.g. diabetes, hypertension, congenital disorders, etc.) and lifestyle factors (e.g. smoking status, alcohol consumption, diet and exercise habits, etc.), as input data. According to examples, the conditioning variables may be collected at the time of collection of real imaging data that is used to train the disclosed methods and apparatuses. This allows the disclosed methods and apparatus to learn how the conditioning variable(s) affect the imaged anatomical and / or physiological parts, such as organs, such that the trained method or system may then generate plausible further examples of anatomical and / or physiological parts / organs based on subsequently provided one or more new conditioning variable(s). According to examples, the disclosed conditional generative framework (i.e. method and / or system) may encode single- or multi-part organ shapes (or multi-organ shape assemblies) into a latent representation (i.e. low-dimensional representation, for example a vector), conditioned on (i.e. subject to or constrained by) some relevant patient attributes (such as the aforementioned demographic information, clinical measurements, etc.). The conditionally encoded latent representation in turn may be decoded to synthesise instances single- or multi-part organs (or multiorgan shape assemblies). The organ shapes or multi-organ shape assemblies synthesised using the trained system are referred to as virtual populations and the process is referred to herein as ‘controllable generation’, or ‘controllable synthesis’, hence providing the overall controllable generation of virtual populations of anatomy, for use in known, existing, or future developed, medical in-silico methodologies.
[0010] Advantageously, given some relevant patient attributes in a target population of interest, the disclosed method, and associated apparatuses, allow controllable synthesis of virtual populations that are similar in terms of the variability in organ shapes observed in the target population, but additionally, are controllably synthesised virtual populations that capture the correlation (i.e. the relationships between patients’ attributes and their organs’ shapes) between patient attributes and organ shapes in the target population of interest. For example, in the specific area of Cardiac modelling (which is used as the specific example area of application of the disclosed technology, but the disclosure as a whole is not so limited), previous clinical studies have established a positive correlation between physical measurements such as height, weight and body mass index with the size or volume of the heart. Thus, using variables such as the height, weight, and / or body mass index as the conditioning variables during training, once trained, the disclosed method(s) and system(s) can controllably generate a virtual population of hearts that matches a target population in terms of the variability or distribution of heart volumes observed in the latter, given the target populations range of heights, weights, and / or body mass indices used as input conditioning variables. This disclosed process of conditional synthesis can, for example, emulate specific patient recruitment in real clinical trials used to assess medical device and / or drug performance in specific patient types. Or, put another way, controllable synthesis using the disclosed method(s) and system(s) enables virtual patient recruitment, to create a specified form of virtual population, suitable for use in in-silico design or trials of medical devices and / or drugs. Patient recruitment in real clinical trials adhere to strict predefined criteria, referred to as inclusion and exclusion criteria, which are characteristics of the target patient population considered appropriate or safe to evaluate the safety and efficacy of a medical device and / or drug. Therefore, virtual patient recruitment to an in-silico trial according to the present disclosure enables the ability to specify inclusion and exclusion criteria to ensure virtual or in-silico design or testing of a medical device and / or drug in a virtual population matches the desired target patient population forthat medical device or drug. As such, conditional synthesis using the disclosed method(s) and system(s) enables creation of virtual populations according to such inclusion and exclusion criteria, or in other words, using the inclusion and exclusion criteria as conditioning variables, enables to create virtual populations that match the characteristics of the desired target patient population.
[0011] Controllable synthesis of virtual populations of organ shapes or multi-organ shape assemblies is not possible with prior known technologies based on unconditional generative models, i.e. unconditional generative models that do not constrain the learning of latent representations of shapes of interest with relevant patient attributes (such as inclusion and exclusion criteria defined in real clinical trials). Existing conditional generative models for organ shapes or multi-organ shape assemblies impose strict constraints on the latent representations learned by the generative model, specifically, by requiring the latent representations to be multivariate Gaussian distributions. Constraining latent representations in a conditional generative model in this way limits the flexibility of the model as multivariate Gaussian distributions are unimodal. Conversely, the conditional latent distribution of organ shapes or multi-organ shape assemblies, as per the present disclosure, conditioned on aforementioned patient attributes and using latent flow transformations, may be multimodal in nature. As such, prior technologies for conditional generative models provide poor approximations to the true underlying conditional latent distribution of organ shapes or multi-organ shape assemblies and this in turn limits the diversity or variability observed in virtual populations synthesised using such models (and hence their usefulness in developing or testing medical devices and / or drugs using in-silico techniques). A key aspect of the proposed approach is that it applies ‘normalising flows’ (hereinafter referred to as ‘latent flow transformations’), which refers to a sequence of non-linear, invertible, transformations, that are applied during training of the conditional generative model to transform the conditional latent representation from a unimodal multivariate Gaussian distribution to a multi-modal distribution in a data-driven manner. Once trained, the disclosed method(s) and system(s) can therefore controllably synthesise virtual populations that are more diverse (i.e. have greater variability) than permitted by prior known technologies, and correspondingly, the disclosed controllable synthesis techniques can synthesise virtual populations that better match the variability observed in target populations. Whilst the disclosed controllable generative framework for conditional synthesis of virtual populations is described in the specific use-case of generating left ventricular heart shapes (where shapes are represented as triangular meshes or unstructured graphs), the disclosed approach(s) can be similarly used to synthesise any multi-part or multi-organ shape assemblies (for example, all four heart chambers and associated blood vessels, abdominal organs, lungs, bone structures, or any other sub-selection of a physical entity within an anatomy).
[0012] Advantages of the present disclosure may include: the ability to synthesise virtual populations of organ shapes or multi-organ shape assemblies in a controllable manner, i.e., conditioned on relevant patient attributes (e.g. demographic information, physical and clinical measurements, lifestyle factors, etc.), such that, the synthesised virtual populations are realistic, capture the non-linear relationship between desired patient attributes and the corresponding shape of organs of interest; and providing a flexible generative framework that captures the variability in organ shapes more accurately than prior known technologies, with respect to a target patient population.
[0013] BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments and, togetherwith the description, serve to explain the disclosed principles. In the drawings:
[0015] Figure 1 shows an example system for controllable synthesis of virtual organ shapes (or training thereof) according to an example of the disclosure;
[0016] Figure 2 shows an example conditional generative encoder neural network for use in the system of Figure 1 , according to an example of the disclosure;
[0017] Figure 3 shows an example latent flow neural network for use in the system of Figure 1 , according to an example of the disclosure;
[0018] Figure 4 shows an example conditional decoder neural network for use in the system of Figure 1 , according to an example of the disclosure;
[0019] Figure 5 shows an example residual graph convolution down-sampling block for use in the encoder of Figure 2, according to an example of the disclosure;
[0020] Figure 6 shows an example residual graph convolution up-sampling block for use in the decoder of Figure 4, according to an example of the disclosure;
[0021] Figure 7 shows an example patient attribute mapping network for use in the system of Figure 1 , according to an example of the disclosure;
[0022] Figure 8 shows a method of pre-training the system of Figure 1 , according to an example of the disclosure;
[0023] Figure 9 shows a method of training the system of Figure 1 , according to an example of the disclosure;
[0024] Figure 10 shows a method of inferencing new virtual populations using the trained system of Figure 1 , according to an example of the disclosure;
[0025] Figure 11 shows an example apparatus suitable for carrying out the method(s) according to examples of the disclosure.
[0026] DETAILED DESCRIPTION
[0027] The technical operation and advantages of the present disclosure shall now be provided by way of a plurality of examples that are merely illustrative of the novel and inventive features, and the disclosed examples are intended to be fully combinable in any reasonable combination, unless explicitly stated otherwise (for example, are defined as true alternatives), or breaks the laws of physics.
[0028] Exemplary embodiments are described with reference to the accompanying drawings. Wherever convenient, the same reference numbers are used throughout the drawings to refer to the same or like parts. Also, it will be appreciated that the drawings show processing of different data as it passes through the disclosed system (or method) portion, and as such the input to a later stage may be the output of an earlier stage. Thus, there the label terms inputs and outputs have been applied in parentheses to show this linkage. Items in the Figures that are optional, or show alternative embodiments, may be shown in dotted lines. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the disclosed example embodiments. However, it will be understood by those skilled in the art that the principles of the example embodiments may be practiced without every specific detail. Well-known methods, procedures, and components have not been described in detail so as not to obscure the principles of the example embodiments. Unless explicitly stated, the example methods and processes described herein may be carried out in any order or sequence, and is not constrained to a particular system configuration, unless specifically stated. Additionally, some of the described embodiments or elements thereof can occur or be performed (e.g., executed) simultaneously, at the same point in time, or concurrently.
[0029] This disclosure may be described in the general context of customized hardware capable of executing customized preloaded instructions such as, e.g., computer-executable instructions for performing program modules that provide the described neural networks and other modules of the disclosed method(s) and system(s). Program modules may include one or more of routines, programs, objects, variables, commands, scripts, functions, applications, components, data structures, and so forth, which may perform particular tasks on data (such as, but not limited to, the conditioning variables data, input meshes, other training data and subsequent data used to define inference outputs), or implement particular abstract data types. The disclosed embodiments may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network (for example, in the so called ‘Cloud’). In a distributed computing environment, program modules may be located in local and / or remote computer storage media including memory storage devices.
[0030] The embodiments discussed herein generally involve or relate to artificial intelligence (Al), including machine learning (ML), based (i.e. AI / ML-enhanced) medical device (or drug) design, manufacture, testing and / or regulatory approval. As such, the disclosed Al-based techniques may involve perceiving, training, synthesizing, inferring, predicting and / or generating information using computerized tools and techniques. For example, such Al systems may use a combination of hardware and software as a foundation for rapidly performing complex operation to perceive, train, synthesize, infer, predict, and / or generate information about real and virtual patients. Such Al systems may use one or more models, specific ones of which are disclosed herein, and so may have a particular configuration (e.g., model parameters and relationships between those parameters, as discussed below). While any of the disclosed models may have an initial configuration, their configuration may change overtime as the respective model(s) learns from input data (e.g., training input data), which allows the model(s) to improve function and / or abilities. For example, a dataset may be input to one of the disclosed model(s), which may produce an output based on the dataset and the configuration of the model itself. Then, based on additional information (e.g., an additional input dataset, validation data, reference data, feedback data), the model(s) may deduce and automatically electronically implement a change to its configuration that will lead to an improved output.
[0031] Figure 1 shows an example system 100 for controllable generation of virtual organ shapes (or assemblies) according to an example of the disclosure, in general terms. The disclosed system 100, in its most generic form (also referred to as a conditional generative shape modelling system or conditional generative neural network), comprises four main components - an encoder neural network 110, a latent flow neural network 120, a decoder neural network 130 and a patient attribute mapping neural network 140. The encoder neural network 110 comprises multiple (N) feature extraction blocks. According to the example, the encoder neural network 110 takes as input a batch of organ shapes or a batch of multi-organ shape assemblies, in the form of surface meshes 101 , and their associated sets of patient attributes 102 (e.g. demographic information, physical and clinical measurements, etc.) that act as conditioning variables, and provides as output a sampled learned conditional latent representation 111 (i.e. a lower dimensional representation) of those meshes. Here, the term ‘sampled’ means the actual output is sampled from a learned posterior latent distribution, where the posterior latent distribution represents a broader set of possible values for the latent representations. Each surface mesh 101 may be a computational mesh represented by a list of 3D points (i.e. spatial coordinates, which may be referred to as vertices or nodes) defining the boundary of the anatomical part / shape, and an adjacency matrix defining vertex / node connectivity. Note in the following, ‘representations’ are used interchangeably with ‘variables’. Prior to use by the encoder 110, the patient attributes / conditioning variables 102 are first transformed through a multi-layer perceptron or fully connected network, referred to as the patient attribute mapping network 140, to yield mapped patient attributes 141 which are the actual input to the encoder neural network 110 that carries the patient-based conditioning variable data. Note these mapped patient attributes are, in effect, a form of transformed patient attributes, where the transformation applied is of a particular type for use in a given use-case and patient data type combination, under control of the patient attribute mapping network 120. This initial mapping is done in order to normalise the patient parameters, for example, to put onto a similar scale, since, otherwise, the values can be of very different types, scales, etc. For example, blood pressure values are very different to BMI values, or weights, heights, etc.
[0032] The learned conditional latent representations 111 output from the encoder neural network 110 are input to the latent flow neural network 120, along with the same original conditioning variables 102, input into the patient attribute mapping network 140. The latent flow neural network 120 comprises multiple (F) flow layers that apply sequential, non-linear, invertible transformations (i.e. with a one to one mapping) to the inputted learned conditional latent representations 111 provided by the encoder neural network 110, and outputs transformed conditional latent flow representations 121 . The output transformed conditional latent flow representations 121 of the latent flow neural network 120, comprising transformed conditional latent representations, are input to the decoder neural network 130, which in turn, learns to output reconstructions 131 of (i.e. predicted approximations for) any inputted batch of organ shapes or multi-organ shape assemblies 101 , but adjusted for the patient attributes as conditioning variables. The decoder neural network 130 comprises multiple (N) feature extraction blocks.
[0033] The system may be trained on meshes derived from real-life medical imaging data of respective organs / parts thereof, which are captured together with the respective patient attributes of the patients being imaged. This allows the system to learn how the different patient attributes affect the respective organ / organ assembly shapes. The system may also be pre-trained using only meshes derived from real-life medical imaging data of respective organs / parts thereof, without the respective patient attributes, so that each of the parts of the system, e.g. encoder, latent flow, and decoder neural networks, are correctly initialised for use on respective organ / organ assembly types. The real- life medical imaging data (e.g. imaging data of organs) used fortraining, and subsequently determinant of the type of output from the trained system, may be any suitable type of medical imaging data, for example, but not limited to: magnetic resonance imaging, computed tomography, computed tomography angiography, ultrasound, positron emission tomography, single-photon emission tomography. Put another way, the disclosed system is medical imaging technology agnostic.
[0034] During training of the disclosed system 100, given a training set of organ shapes or multiorgan shape assemblies 101 as inputs to the encoder neural network 110, along with patient attributes / conditioning variables 102 as inputs to the patient attribute mapping network 140, the decoder neural network 130 outputs reconstructions 131 that approximate the input training set of organ shapes 101 , but conditioned on or subject to the input patient attributes 102. Or, put another way, during training, the conditional generative model in the disclosed system 100 learns the conditional latent representations (i.e. learns the conditional posterior distribution) of the input training set of shapes 101 subject to their associated patient attributes 102, which captures the relationships between the organ shapes 101 and the patient attributes 102, and the decoder neural network learns to ‘decode’ or recover the training set of organ shapes by predicting or reconstructing approximations 131 of the input shapes. During training, the outputs of the decoder neural network 131 are approximations of the input training set of organ shapes. Once the disclosed system 100 is trained, during inference, given a new set of patient attributes 102 as inputs to the patient attribute mapping network 140, the disclosed system samples from the conditional posterior latent distribution learned during training to generate new latent representations and the decoder neural network 130 uses these new latent representations to synthesise virtual organ shapes representative of the new set of patient attributes 102 input to the system 100. Note, the new set of patient attributes will be of the same type as used in the training, but may be of different values. In some examples, the system may be further trained on new patient attribute types, thereby allowing whole new types of patient data to be used in the future.
[0035] According to an example, both the encoder neural network 110 and the decoder neural network 130 are parameterised (i.e. defined by) graph convolutional neural networks. A conditional generative model, as used in the example implementation of the disclosed system 100 shown in Figure 1 , is an algorithm that is able to learn from real data and synthesise virtual instances that are similar to, but not the same as, the real data instances that it is trained with. The conditional generative model of the disclosed system 100 is used for synthesising one or more virtual instances of anatomical parts / organ shapes or multi-organ shapes or assemblies according to a desired set of patient attributes or conditioning variables 102. In the following example disclosure, the example used will be a human heart comprising one or more heart chambers. However, other undescribed examples may include any organ, and even including multiple abdominal organs or assemblies thereof (such as liver, kidneys, pancreas, etc.). The individual virtual organ shapes conditionally synthesised by the disclosed system 100 may be combined together resulting in a ‘virtual population’.
[0036] As noted above, Figure 1 depicts a schematic diagram that provides an overview of the conditional generative shape modelling system 100 for controllable synthesis of virtual populations from data sets 101 of anatomical / organ shapes and their associated patient attributes / conditioning variables 102 (e.g. demographic information, physical and clinical measurements, etc.). In this example, the four components of the conditional generative shape modelling system 100, comprising the encoder neural network 110, the latent flow neural network 120, the decoder neural network 130, and the patient attribute mapping network 140, are all trained together / simultaneously.
[0037] In an example based on cardiac modelling, the conditional generative shape modelling system 100 can be used to learn individual or shared conditional latent representations for anatomical parts that constitute a human heart, which in the present example comprises the left ventricle (LV), right ventricle (RV), left atrium (LA), right atrium (LA) and the aortic root (AR), using anatomical / organ shape data from example patient data covering different individuals (and / or coming from different data sources) and using data covering different patient attributes (such as, demographic information, physical and clinical measurements, lifestyle factors, etc.) as conditioning variables 102. The disclosed conditional generative shape modelling system 100 is not restricted to modelling the abovementioned cardiac structures but may be extended to include additional cardiac structures if available, such as, coronary arteries and their branches, pulmonary artery, pulmonary veins, cardiac valves, etc. Additionally, the disclosed system 100 is not restricted to cardiac modelling and can be used to model other organs (e.g. lungs, liver, spine) or even multiple organs as instances of a multiorgan shape assembly (e.g. thoracic organs and abdominal organs). This is to say, whilst the specific exemplary embodiments described herein focus on conditionally generating cardiac virtual cohorts, other embodiments may be used to conditionally synthesise virtual cohorts of any other organs or multi-organ shape assemblies.
[0038] Figure 2 provides an example implementation 200 of the encoder neural network 110 in the disclosed system of Figure 1 , with N=5 feature extraction blocks (210 - 250).
[0039] Referring now to Figure 2, each feature extraction block (210 - 250) in the example implementation of the encoder neural network 110, is defined by a sequential combination of Chebyshev graph convolution operations (a type of spectral graph convolution operation), batch normalisation operations and graph pooling (down-sampling) operations. These are shown in, and described later, in respect of Figure 5 (which refers to an input 501 and output 503 of an individual one of the feature extraction blocks - item 230 in this case). The feature extraction blocks in the encoder neural network 110 may be implemented using other types of convolution operations such as standard spatial convolutions used for image and point cloud processing, K-nearest neighbour spatial graph convolutions or other types of spectral graph convolution operations. An example of the architecture used for each feature extraction block 210-250 in the disclosed system is shown in Figure 5, discussed below. In the example implementation 200 of the encoder neural network 110 shown, the encoder comprises a plurality feature extraction blocks defined by residual graph convolution down-sampling (RGCDS) blocks (210-250). The first RGCDS block 210 takes as input 201 a mesh (or a batch of meshes) for a single part of the heart (e.g. left ventricle) or for the whole heart comprising multiple structures (e.g. left and right ventricle, left and right atria, aortic vessel, etc.), and the mapped patient attributes 141 output by the patient attribute mapping network 140. The input mesh 101 is down-sampled in the encoder neural network 110 through operation of the successive RGCDS blocks (210-250). The output of the last feature extraction block 250 is vectorised, resulting in a vector 111 of size VxNf (where V represents the number of vertices in the down-sampled mesh within the last feature extraction block 250 and Nf represents the number of feature maps used in the block 250) and taken as input to two different linear layers (i.e. two separate fully connected layers) covering Mean 260a and standard Deviation 260b, respectively. In the example shown, the outputs of the two linear layers 260a and 260b are each vectors (e.g. of size 20x1) and represent the mean 261 and standard deviation 262 of the conditional posterior latent distribution, defined as a multivariate Gaussian distribution. The vector sampled from this conditional posterior latent distribution based on the input mean 261 and standard deviation 262 vectors constitutes the output, i.e. the sampled conditional latent variables (e.g. another vector of size 20x1) 111 , from the encoder neural network 110.
[0040] Figure 3 shows an example implementation 300 of the latent flow neural network 120 in the disclosed system shown in Figure 1 , and in this example has F=6 flow layers (310 - 360). Other numbers of flow layers may also be used.
[0041] Referring now to Figure 3, the latent flow neural network 120 receives as input the output sampled conditional latent variables 111 from the encoder network 110, together with the original patient attributes 102. The latent flow layers in the example implementation 300 shown in Figure 3, comprise F=6 ‘planar’ flow layers (310 - 360) which apply a sequence of six sequential, invertible transformations based on ‘planar’ flow, to the input sampled conditional latent vectors 111 , and in turn, output transformed conditional latent vectors 121 . The sequence of planar flow transformations 310 - 360 applied by the latent flow neural network 120 transform the sampled conditional latent vectors 111 from a unimodal multivariate Gaussian distribution to a multi-modal distribution in a data- driven manner (i.e. where the data drives the estimation of the latent flow network’s parameters to approximate the unknown, multi-modal latent / hidden low-dimensional representation of the input data, as opposed to forcing the latent / hidden representation to conform to a prior belief / distribution, as done with conventional unconditional VAEs and conditional VAEs that may use a multi-variate Gaussian prior distribution for example)., which provides the overall conditional generative model its improved flexibility and ability to conditionally synthesise more diverse virtual populations, relative to prior known technologies for controllable synthesis of virtual populations of anatomical / organ shapes. The original patient attributes (i.e. conditioning variables) 102 input to the latent flow network 120 are processed through a fully connected network 370, in this case comprising two fully connected / linear layers 372, 376, and a sigmoid linear unit activation (SiLU) function layer 374 as shown at the top portion of Figure 3. The fully connected network 370 is a part of the latent flow network 120 which transforms / maps the original patient attributes / conditioning variables 102 to match the scale / range of the features learned within the latent flow neural network 120. In other words, the role of the fully connected network 370 is similar to that of the patient attribute mapping neural network 140, i.e. the patient attribute mapping neural network 140 learns the mapping from the original patient attribute values to new values that are compatible with the features in each RGCDS block in the encoder neural network, while the fully connected network 370 learns the mapping from the original patient attribute values 102 to new values that are compatible with the features learned by the latent flow neural network 120. In the figure, these are referred to as (output) transformed conditioning variables 378.
[0042] In other words, the fully connected network 370 carries out a similar process to the patient mapping neural network 140, but is done independently so that the patient attribute values 102 input to the fully connected network 370 are scaled and transformed to match the features learned within the latent flow neural network 120. The fully connected network 370 in the latent flow neural network 120 is described in more detail below.
[0043] Figure 4 provides an example implementation 400 of the decoder neural network 130 in the disclosed system 100, with N=5 feature extraction blocks (420 - 460). Other numbers of feature extraction blocks may be used.
[0044] Referring now to Figure 4, the decoder neural network 130 receives as input the transformed conditional latent flow variables 121 output from the latent flow network 120. In the example implementation 400 of the decoder neural network 130, the decoder comprises a linear layer (or fully connected layer) 410 followed by a plurality of feature extraction blocks 420-460 defined by residual graph convolution up-sampling (RGCUS) blocks. In other examples, these may be different types of convolution blocks, for example to match ones used in alternative examples of the down sampling blocks shown in Figure 2, as noted previously (e.g. K-nearest neighbour spatial graph convolutions or other types of spectral graph convolution operations). An example of the architecture used for RGCUS block in the disclosed system is shown in Figure 6, discussed below (which refers to an input 601 and output 602 of an individual one of the feature extraction blocks - item 440 in this case). The output of the decoder neural network 130 during training comprises predictions or reconstructions 402 which are approximations of the anatomical / organ mesh (or batch of meshes) that are input to encoder neural network 110, subject to the associated set of patient attributes / conditioning variables 102 input to the patient attribute mapping network 140.
[0045] Figure 5 shows an example of the architecture 500 used for each of the feature extraction blocks, specifically residual graph convolution down-sampling (RGCDS) blocks 210-250, in the encoder neural network Figure 2, as discussed above. Each block 210-250 comprises the following sequence of operations: a given input 501 (which may be the original input mesh, or the output from the previous RGCDS block) is passed through a first Chebyshev graph convolution layer 510, a first batch normalisation layer 520, a first exponential linear unit (ELU) activation function layer 530, a second Chebyshev graph convolution layer 540, a second batch normalisation layer 550, a second ELU layer 560, a residual connection / addition layer 570, and a mesh pooling (i.e. down-sampling) layer 580. The mesh pooling layer provides the output 503 (which may then be used by the next RGCDS block, or be used by the fully connected layers (Mean) 260a / (Standard deviation) 260b. The residual connection / addition layer 570 adds the output 502 of the first Chebyshev graph convolution layer and the output of the second ELU layer 560 together. Whilst the example shown uses two of each of the Chebyshev graph convolution layers, batch normalisation layers, and exponential linear unit (ELU) activation function layers, other numbers of each layer may also be used, dependent on the use-case. It is also to be noted that in this disclosure as a whole, each of the instances of the Chebyshev graph convolution layers, Batch normalisation layers, Exponential Linear unit layers, fully connected layers, residual connection / addition layer and SiLU layers are all distinct, trainable components of the system, i.e., none of these layers, instances of which are used in several parts of the component networks in the overall system, share parameters with one another.
[0046] Referring now to the right hand side of Figure 5, each RGCDS block 210-250 in the encoder neural network 1 10 also receives the mapped conditioning variables / mapped patient attribute values 141 output by the patient attribute mapping network 140. The mapped patient attribute values 141 input to each RGCDS block are passed through two separate fully connected networks 590-591 , which further transform the inputs 141 before combining them with the output 502 of the first Chebyshev graph convolution layer 510, in each RGCDS block 210-250. The fully connected networks 590-591 in each RGCDS block 210-250 are used to scale and translate the inputs 141 to match size and range of features output by the first Chebyshev graph convolution layer 510 in each RGCDS block 210-250. In the example shown, each fully connected network 590 / 591 comprises: a first fully connected layer 590a / 591 a, a SiLU activation layer 590b / 591 b, and a second fully connected layer 590c / 591 c). The fully connected networks 590-591 and the combination of their respective outputs, with the output 502 of the first Chebyshev graph convolution layer 510 in each RGCDS block 210-250 is described in more detail below.
[0047] Figure 6 shows an example of the architecture 600 used for each RGCUS block 420-460 in the decoder neural network of Figure 4, as discussed above. Each RGCUS block comprises the following sequence of operations: a mesh un-pooling (i.e. up-sampling) layer 610, a first Chebyshev graph convolution layer 620, a first batch normalisation layer 630, a first exponential linear unit (ELU) activation function layer 640, a second Chebyshev graph convolution layer 650, a second batch normalisation layer 660, a second ELU layer 670, and a residual connection between the output of the first Chebyshev graph convolution layer 620 and the output of the RGCUS block formulated as an addition layer 680. Whilst the example shown uses two of each of the Chebyshev graph convolution layers, batch normalisation layers, and exponential linear unit (ELU) activation function layers, other numbers of each layer may also be used, dependent on the use-case. The input 601 of each RGCUS block may be the transformed conditional latent flow variables (via the fully connected layer 410), or the output of the previous RGCUS block. Meanwhile, the output 602 may be the input to the next RGCUS block, or the Chebyshev layer 470). Once trained, the conditional generative model based system 100, also referred to as a conditional flow graph variational autoencoder (cFGVAE) in the example implementation described herein, can be sampled from given a new set of target patient attributes 102 as inputs, to conditionally generate new shapes of organs / multi-organ assemblies that are representative of the inputted new set of target patient attributes 102, and are similar to, but not exactly the same as, any input organ shape data set(s) 101 used in the training set - i.e. the conditional generative model of the conditional flow graph variational autoencoder can now generate synthetic examples of organs that are known to be realistic, and conformative to the new patient attributes / conditioning variables.
[0048] In more detail: the example implementation of the conditional generative model based system 100 described above represents a conditional flow variational autoencoder neural network, wherein, the constituent encoder neural network 110, latent flow neural network 120, decoder neural network 130 and patient attribute mapping network 140 are trained jointly / simultaneously to learn conditional latent representations for the input mesh or batch of meshes representing shapes of anatomical parts and / or organs of interest. Conditional flow variational autoencoders (cFVAEs) (and the graphconvolution variants, used in the disclosed system 100 herein) may be Bayesian conditional latent variable models that couple: 1) a recognition function represented by the encoder neural network 110, to conditionally infer a low dimensional hidden representation of the input shape / mesh data (i.e. to infer a conditional latent representation), given some relevant patient attributes (e.g. demographic information, physical and clinical measurements, lifestyle factors, etc.) as input conditioning variables 102, with 2) a generative function represented by the latent flow neural network 120 and the decoder neural network 130, which initially transforms the unimodal conditional latent representation inferred by the encoder neural network 110 to a multi-modal conditional latent representation as output by the latent flow neural network 120, and finally transforms the multi-modal conditional latent representation back to the original data space through the decoder neural network 140. Unconditional VAEs infer the latent representation of input data by approximating the posterior distribution of the latent variables given the observed input data. Meanwhile, conditional VAEs infer the conditional latent representation of input data, by approximating the conditional posterior distribution of the latent variables given some input data (e.g. mesh-based representations of anatomical / organ shapes) and appropriate conditioning variables (e.g. patient attributes such as demographic information, physical and clinical measurements, etc.). Whereas, conditional flow VAEs and their graph convolution variants as used in the example implementations provided for the disclosed system, work similar to but not exactly the same as conditional VAEs. Specifically, conditional VAEs developed in prior known technologies are limited to approximating the conditional latent distribution as a unimodal multivariate Gaussian distribution (although other unimodal probability distributions other than Gaussians have also been used). Whereas, conditional flow VAEs described herein relax this constraint on the conditional latent distribution that is inferred from the input data by applying sequential, invertible, transformations that transform the unimodal distribution to a multi-modal distribution in a data-driven manner (see earlier explanation of data driven manner). This provides conditional flow VAEs with greater flexibility than conditional VAEs, allowing the former to synthesise virtual populations with greater diversity and variability in anatomical / organ shapes than the latter. This is achieved by jointly optimising the encoder neural network 110, latent flow neural network 120, decoder neural network 130, and patient attribute mapping network 140, for example to maximise the associated evidence lower bound (ELBO) of the observed input data, conditioned on the input patient attributes (or conditioning variables) 102. Conditional flow VAEs are flexible and can be used to learn latent representations of any type of data (e.g. images, audio signals, mesh-based representations of shapes), given some appropriate conditional information or covariates. In the example conditional generative shape modelling system 100, conditional flow VAEs are combined with graph-convolutional neural networks to learn conditional latent representations of anatomical shapes represented by computational meshes (also referred to as unstructured graphs), subject to or conditioned on relevant patient attributes 102 (including patient demographic information, physical and clinical measurements, and lifestyle factors).
[0049] The encoder neural network 110 and decoder neural network 130 are graph-convolutional neural networks that utilise Chebyshev convolution operations to extract a hierarchy (i.e. across multiple scales) of shape features and learn a representative conditional latent representation describing variability in shape of the corresponding part(s), and / or assemblies, of organ(s) of interest across the training population and their relationships with the input conditioning variables 102, i.e. via the mapped patient attributes 141 input through the patient attribute mapping network 140. The encoder and decoder neural networks each have their own set of down- (210-250) and up-sampling (420-460) blocks, and in the examples shown, there are 5 levels used, but other numbers of levels may be used instead. The number of levels typically affects processing speed and accuracy, so a value that is the best compromise for a given use-case may be chosen for each implementation. The number of levels used in the disclosed system were determined empirically to be the best for the data used to train the developed system. Given different training data for the same or different applications, the chosen number of levels may vary. Spatial convolution operations used in convolution neural networks are well-defined for structured / gridded data in the Euclidean domain and cannot be applied directly to irregularly structured data such as graphs. Therefore, Chebyshev convolution operations are used to provide a generalisation of spatial convolutions, enabling the spatial convolutions to be applied to data defined by graphs (i.e. meshes). Chebyshev convolution operations are performed in the Fourier domain. According to the convolution theorem, a convolution operation performed in the spatial domain is equivalent to element-wise multiplication in the Fourier domain. Chebyshev convolutions are applied by first transforming the graph-based data to the Fourier domain. The Fourier transformed graph-based data is then multiplied elementwise with the Fourier transformed convolution filter, and the inverse Fourier transform is subsequently applied to convert the result of the multiplicative operation back to the spatial domain. The parameters / weights of the convolution filter in a Chebyshev convolution operation are parameterised by truncated Chebyshev polynomials, where the polynomial coefficients represent the learnable weights / parameters of the convolution filter. Truncated Chebyshev polynomials are used in the disclosed conditional generative shape modelling system as they reduce the computational complexity relative to other spectral convolution operations and relative to using general N-dimensional Chebyshev polynomials. However, other spectral convolution operations may also be used instead of truncated Chebyshev polynomial convolution operations, according to examples. Alternatively, spatial convolution operations designed for graphbased data, such as Feature-steered graph convolution or K-nearest neighbour / Dynamic Edge graph convolution operations may also be used instead of spectral / Chebyshev convolution operations.
[0050] To train the conditional flow graph convolution VAE (cFGVAE), which is another way to call the overall system 100, consider an input data set denoted X= {x}n=i...N where, each xndenotes the shape of an anatomical part (e.g. left ventricle) of an individual (i.e. patient) represented in the training data set 101 as a computational mesh (also known as a triangular surface mesh, or an unstructured graph), and a set of patient attributes / conditional variables for the same patients denoted C = {c}n=i...N. Each computational mesh is represented by a list of 3D points (i.e. spatial coordinates, which may be referred to as vertices or nodes) defining the boundary of the anatomical part / shape, and an adjacency matrix defining vertex / node connectivity. Each individual patient’s data in the input data set X is denoted by subscript n, where {n=1...N} and N individuals’ data is included in X. Each set of patient attributes / conditional variables, cn, 102 are vectors of the same size (Cx1) across all patients and the elements of these vectors are numerical values representing different types of attributes (e.g. age, gender, systolic and diastolic blood pressure, height, weight, etc.). Each mesh xn101 and associated conditioning variables cn102, in the input data set X and associated conditional variables data set C, respectively, is used to train the cFGVAE 100. The training data samples {xn}n=i...N in X, and their corresponding conditioning variables {cn}n=i...N in C, each come from different N individuals. In other words, the data X and C used to train the cFGVAE comprise anatomical / organ shape data and associated patient attributes / conditional variables from the same individuals. The cFGVAE however, can be trained with X and C wherein, some instances cn102 contain missing information, i.e. where some of the attributes are missing at random, for some individuals. The patient attribute mapping network 140, takes the conditioning variables cn102 for one or more individuals as input and outputs a mapped set of conditioning variables denoted c‘n141 . The conditioning variables cn102 are transformed in the patient attribute mapping network 140 using a fully connected network (or multilayer perceptron) with three linear layers, and sigmoid linear unit activation (SiLU) function layers (see Figure 7 described below). Other types of activation functions may also be used for the patient attribute mapping network 140, instead of SiLU.
[0051] Figure 7 shows an example architecture implementation 700 of the patient attribute mapping network 140 of Figure 1 , and it comprises the following sequence of operations (applied to each cn, to output a transformed, i.e. mapped, version of the conditioning variables c‘n141): a first fully connected / linear layer 710 that takes the original conditioning variables cn102 (vector of size Cx1 , e.g. 14x1 in the example shown) as input and outputs a vector 702 of size Mx1. Note, conditioning variables may be individual values, or collated into a vector as used here. In the example implementation of the disclosed system of Figure 7, M=32 was used as the output size for the first fully connected / linear layer 710 in the patient attribute mapping network 140, however, the output size can be defined to be higher or lower than this chosen value. The output 702 of the first fully connected / linear layer 710 is then transformed using a first SiLU activation function layer 720, and the output 703 of the first SiLU activation function layer is passed through the second fully connected / linear layer 730. The second fully connected / linear layer 730 takes the output 703 vector of size Mx1 (M=32 was used in this example) and outputs a vector 704 of the same size. The output 704 is once again transformed using a second SiLU activation function layer 740, and its output 705, is passed to the final (third) fully connected / linear layer 750 as input. The final fully connected / linear layer 750 takes a vector of size Mx1 as input (i.e. output 705 from the preceding second SiLU activation function layer 740) and outputs 706 a vector of size Cx1 , which denotes the mapped set of patient attributes / conditioning variables c‘n141 , where C is the same value as the input C noted above - e.g. 14 in this case) The encoder neural network 110 in the cFGVAE takes a mesh-based representation of the shape for a given anatomical part or organ xnas input along with the mapped set of conditioning variables, c‘n141 , output from the patient attribute mapping network 140, and maps this high-dimensional representation of shape to a low-dimensional vector or latent representation (denoted zn), conditioned on the mapped set of conditioning variables c‘n. The lowdimensional vector or conditional latent representation resulting from each sample, i.e. each pair of xnand c‘nbeing passed through the encoder neural network 110, is subsequently further transformed by the latent flow neural network 120 and finally mapped back or reconstructed to approximate the original mesh of the input shape, by the decoder neural network 130. By training the cFGVAE 100 on its corresponding data sets X and C, each conditional latent representation for each sample {xn, cn} is discovered. Or put another way, by training the cFGVAE 100 on its corresponding data sets X and C, the conditional posterior latent distribution of the training set of shapes and their conditioning variables is inferred / learned by the model.
[0052] The conditional generative model, cFGVAE, in the disclosed system 100 is trained using the training data set X and the associated set C of conditioning variables as inputs, to learn a conditional latent representation from which the observed data X is assumed to be generated (subject to the conditions / constraints imposed by C), by approximating the true but intractable conditional posterior distribution of the latent variables using variational inference. Variational inference is achieved by optimising the evidence lower bound (ELBO) of the observed data X, conditioned on the conditioning variables C, with respect to the parameters of the patient attribute mapping 140, encoder 110, latent flow 120 and decoder 130 networks in each cFGVAE. The encoder 110 approximates the conditional posterior distribution (denoted q(Z | X,C)) given the input data X and associated set of conditioning variables transformed by the patient attribute mapping network 140, denoted C, and the decoder 130 approximates the data likelihood given the latent variables (denoted p(X | Z, C)). Thus, during training, for a given input sample {xn, c‘n} the encoder 110 maps the input to a unique conditional latent representation znwhich in turn is used as input by the latent flow network 120, which transforms znto zn. The output of the latent flow network znin turn is used as input by the decoder 130 to reconstruct the data (denoted rn). Put another way, during training, the encoder 110 learns to reduce the inputted data to a low-dimensional or compressed form, conditional on the inputted conditioning variables transformed by the patient attribute mapping network 120, the latent flow network 120 transforms the low-dimensional or compressed form received from the encoder 110, and the decoder 130 learns to unravel or transform / map this compressed information received from the latent flow network back to an approximation of the data inputted to the encoder 110. Once the patient attribute mapping 140, encoder 110, latent flow 120 and decoder 120 networks are trained, in the generation phase (also referred to as inference phase or synthesis phase) new compressed forms of the data may be created given a new set of patient related conditional variables as input, and transformed by the trained decoder 130 to generate synthetic data that appear realistic (i.e. has certain properties or characteristics that are shared and / or similar to real data, but now also taking into account conditional patient attributes).
[0053] There now follows a more detailed description of the operation of Figure 5. Given an input mesh xn101 and a mapped set of conditioning variables c‘n141 output by the patient attribute mapping network 140 as inputs to the encoder neural network 110, c‘nis first passed as inputs to each RGCDS block (210 - 250). Within each RGCDS block, as shown in Figure 5, c‘n141 is first passed through two fully connected networks each with three independent layers, namely, a linear layer (590a, 591 a), a SiLU activation layer (590b, 591 b) and another, second, linear layer (590c, 591 c). The output of the second linear layer (590c, 591 c) from each fully connected network results in two vectors each of size (1xNf) where Nf denotes the number of feature maps used in the first Chebyshev graph convolution layer 510 in each RGCDS block (210 - 250). The output of the first Chebyshev graph convolution layer 510 in each RGCDS block is combined with the output of the two fully connected networks in each block as:
[0054] On= (On* ) + fn (1) where, Onrepresents the output 510a of the first Chebyshev graph convolution layer 510 in the nthRGCDS block, f„ and f„ represent the output vectors of the two fully connected networks in the nthRGCDS block and (*) represents the element-wise multiplication between the array Onand vector f„. Reverting to Figure 2, this operation is performed in each RGCDS block within the encoder 110 neural network to scale the feature maps learned in each RGCDS block using mapped set of conditioning variables c‘n141 and to combine or embed the conditioning information with the features learned within each RGCDS block in the encoder neural network 110. This type of operation is similar to, but not exactly the same as, adaptive instance normalisation used commonly in training different types of deep neural networks. The output of the final RGCDS block (250) is vectorised, resulting in a vector 203 of size VxNf (where the variables V and Nf are as described previously) input to two separate linear layers, one for the Mean 260a and one for the Standard Deviation 260b. The outputs of the two linear layers 260a and 260b are each vectors of size (e.g. 20x1 , in Figure 2) and represent the mean 261 and standard deviation 262 of the conditional posterior latent distribution, defined by a multivariate Gaussian distribution. The vector sampled from this conditional posterior latent distribution (e.g. of size 20x1), based on the input mean 261 and standard deviation 262 vectors, constitutes the output 111 , i.e. the conditional latent representations znfrom the encoder neural network 110, corresponding to the inputted sample {xn, c‘n} . Put another way, the encoder approximates the conditional posterior latent distribution q(Z | X, C) given input data X and C and samples drawn from this distribution are the conditional latent representations zncorresponding to the input sample {xn, c‘n}.
[0055] Note, it will be appreciated that the prior distribution is an assumed form for the unknown true distribution of the latent / hidden representation of the input data, wherein, the ‘unknown true distribution’ is also referred to as the ‘unknown true posterior distribution’ of the latent representations corresponding to some given input data. In the example implementation of the disclosed system, as shown in Figure 2, the prior distribution assumed for the unknown latent representations is a multivariate Gaussian distribution with the mean given by a zero vector of size 20x1 and the covariance given as an identity matrix (i.e. a diagonal matrix where all entries on the diagonal are equal to one and all other entries are equal to zero), respectively. The output of the encoder neural network 110 comprises a vector 261 and another vector 262, both of size 20x1 , where the first vector 261 represents the mean (referred to as mean vector in Figure 2) and the other vector 262 represents the diagonal entries of the covariance matrix (referred to as standard deviation vector in Figure 2), of a multi-variate Gaussian distribution used to define / represent the approximated posterior distribution of the latent representations of the input data. Or, put another way, the encoder neural network 110 approximates the posterior distribution of the latent representations given some input data, where the approximated posterior distribution is in the form of the assumed prior distribution, i.e. it is constrained to be a multi-variate Gaussian distribution whose parameters are defined by the mean 261 and standard deviation 262 vectors (note, standard deviation 262 vector refers to the diagonal entries of a covariance matrix, where all other entries in the matrix are equal to zero) output by the encoder neural network 110. The latent representation 111 output by the encoder neural network 110 for each instance in the input data is sampled from a multi-variate Gaussian distribution defined by the estimated mean 261 and standard deviation 262 vectors. The latent posterior distribution approximated by the encoder neural network 110 is constrained to be a multi-variate Gaussian distribution that is close to the prior distribution by minimising the Kullback-Leibler divergence between the approximate posterior and assumed prior distributions, which is described in more detail below.
[0056] Reverting to Figure 3, the latent flow neural network 120 receives as input the conditional latent representations zn111 output by the encoder neural network 110, which is transformed through a sequence of F normalising flow layers, which is to say, a sequence of F invertible and differentiable transformations of mapping functions. In the example implementation of the disclosed system in Figure 3, F=6 flow layers are used defined by ‘planar’ flow functions; however, other types of invertible and differentiable transformations or flow functions may be used in place of planar flow functions (such as Radial flows, Affine flows, etc.). More generally, given some initial latent representation zn111 as input, planar flow functions or planar flow layers, apply a sequence of contractions and expansions in the direction perpendicular to the hyperplane defined by wTzn+ b = 0 where, wTand b are learnable or trainable parameters of the planar flow layers. As these sequential invertible transformations are applied in the direction perpendicular to a hyperplane defined based on the input initial conditional latent representation znthey are referred to as planar flow functions. In the Example of Figure 3, given an input conditional latent representation zn, 111 a planar flow function f may be defined as:
[0057] / (zn) = zn+ ( u * h(wTzn+ b ) ( ) where, u, wTand b are learnable or trainable parameters in the planar flow layer, each comprising the same number of parameters as the dimension of the inputted conditional latent representation zn111 (in the present example implementation znis of size 20x1); h is a sequential, invertible element-wise non-linear function, and in the present example implementation, h is defined by a Hyperbolic Tangent (TanH) function; and as before, (*) represents the element-wise multiplication between the array it and the result of h(wTzn+ b). By applying a sequence of F=6 such planar flow transformations, the initial conditional latent representation denoted zno is transformed as,
[0058] Given an invertible and differentiable transformation function f, such as the planar flow function used in the present example, for which the inverse of the function may be expressed as -1= g exists, such that (f(zy) = z, a latent representation z with a distribution q(z) that is transformed using f results in a random variable or latent representation z’ = / (z) that has a distribution given by:
[0059] Q(z) = q(z) | det where, (det is the Jacobian determinant of / . Therefore, by applying a sequence of invertible transformation functions within the latent flow network 120 as shown in Equation 3 to an input conditional latent representation zno as in the present example, a complex multi-modal probability distribution can be constructed from the simple, unimodal conditional posterior latent distribution approximated by the encoder neural network 110. Specifically, the resulting multi-modal probability distribution obtained from applying F=6 planar flow layers as described in this example is given by: where, q(zn0) is the initial conditional posterior latent distribution approximated by the encoder neural network 110 from which the initial conditional latent representation zno is sampled, In is the natural logarithm operation and q6(zn6) is the multi-modal conditional posterior latent distribution approximated by the latent flow network 120 through a sequence of F=6 successive planar flow transformations resulting in a transformed conditional latent representation zne for the sample {xn, c‘n} originally inputted to the encoder neural network 110. Or put another way, the path traversed by the initial conditional latent representation zno through the sequential application of transformation functions is called the ‘flow’ and the path formed by the successive resulting distributions qF=1 6is referred to as a ‘normalising flow’ or a ‘planar normalising flow’ when planar flow transformation functions are used as in the present example. The transformed conditional latent representation zne output by the final planar flow layer 360 in the latent flow network 120 is combined with the original set of conditioning variables 102, denoted cn, corresponding to the inputted mesh xnby first passing cnthrough a fully connected network comprising a first fully connected / linear layer 372, a SiLU activation function layer 374 and a second fully connected / linear layer 376. The output 378 of this fully connected network is a vector of the same size as zne (i.e. vector of size 20x1 in the present example) and is denoted cn. The output cnof the fully connected network 376 is finally combined with zne by simple addition, represented as an addition layer 380, resulting in the final conditional latent representation znoutput 303 from the latent flow neural network 120 and inputted to the decoder neural network 130, shown in detail in Figure 4. Note the above mentioned equation (5) is for the specific case of F=6, but can be generalised to:
[0060] Reverting to Figure 4, the decoder network 130 receives as input the transformed conditional latent representations output 121 from the latent flow neural network 120. The input 121 is a vector (of size 20x1 in the present example) which is first passed through a Fully connected layer / linear layer 410 in the decoder 130. The output of the linear layer 410 is a vector of size VxNf where V represents the number of vertices in the down-sampled mesh in the final RGCDS block 250 in the encoder neural network 110 (see Figure 2) and Nf represents the number of feature maps used in the first RGCUS block 420 in the decoder neural network 130. The output of the fully connected layer 410 is a vector that is reshaped into an array or tensor and inputted to the first RGCUS block 420 and subsequently is successively up-sampled through all five RGCUS blocks (420 - 460). The output of the final RGCUS block 460 is passed through a Chebyshev graph convolution layer 470 which in turn outputs 402 the reconstructed mesh rnas an approximation to the original mesh xninputted to the encoder neural network 110 of Figure 2.
[0061] The conditional generative model, cFGVAE, in the disclosed system is trained by jointly optimising the parameters of all its constituent networks, namely, the patient attribute mapping 140, the encoder 110, the latent flow 120 and the decoder 130 networks, to maximise the modified evidence lower bound (ELBO) of the observed input data X and their corresponding conditioning variables C, where, the modified ELBO is expressed as: where, lnp(xn\ cn) is the marginal log-likelihood of the observed data xn(i.e. the mesh of an anatomical part / organ shape input to the cFGVAE) conditioned on the associated set of conditioning variables cn(i.e. patient attributes 102 such as demographic information, physical and clinical measurements also input to the cFGVAE); Inp(xn|zn6,cn) is the conditional likelihood of the input data parameterised or defined by the decoder neural network 130 (which outputs a reconstruction rnas an approximation to the original mesh xngiven the conditional latent variables znoutput by the latent flow neural network 120 and cn); subscript denotes the number of flow steps or the number of planar flow layers in the latent flow network (F=6 planar flow layers were used in the present example); and K£(q(z0| xn,cn) || p(z0)) represents the Kullback-Leibler divergence of the approximate conditional posterior distribution of the latent variables (z0| xn,cn) from the assumed prior distribution p(z0) over the latent variables. The complete loss function, denoted Ltotal, minimised to train the cFGVAE in the disclosed system is implemented as the negative of the modified ELBO described above, which is defined as:
[0062] Ltotalrecon ^logdet + ^O^KL (®) where, Lreconis the reconstruction loss used to represent the negative of the conditional likelihood of the data - Inp(xn|zn,cn) (defined in equation 6), Llogdetis the summation of the log Jacobian determinant, summed over all flow steps (i.e. summed over all i=1 ...6 planar flow steps as defined in equation 6) and LKIis the Kullback-Leibler divergence KL(c / (z0\ xn,cn) | | p(z0)) (also defined in equation 6).
[0063] The reconstruction loss Lreconis computed as the vertex-wise Li- norm evaluated between each reconstructed (rn) and original input shape / mesh (xn). Lreconis formulated as:
[0064] Brecon = I |x„ — rn11x(8)
[0065] The Kullback-Leibler divergence loss term LKLis a distance measure between probability distributions and here it is used to minimise the divergence between the approximated conditional posterior distribution over the latent variables (z0| xn,cn) and the true unknown (or target) posterior distribution over the latent variables p(z| xn,cn). However, as the target / true posterior distribution p(z| xn,cn) is intractable (cannot be estimated) in the present setting, it is factorised (expressed as) as a product of the conditional data likelihood p(xn|zn,cn) and the assumed prior distribution over the latent variables p(z0) using Bayes’ theorem. As a result, the Kullback-Leibler divergence loss term LKLreduces to minimising the divergence between the approximated conditional posterior distribution over the latent variables (z0| xn,cn) and the prior distribution over the latent variables p(z0). The prior distribution over the latent variables is assumed to be a centered isotropic multivariate Gaussian distribution, i.e. p(z0) = JV(0, I). In the complete loss function Ltotalexpressed in equation 7, w0denotes the weight that is multiplied with the LKLterm in order to balance its influence on the training process relative to the other terms in the overall loss function.
[0066] The conditional generative model, cFGVAE, in the disclosed system 100 is trained by minimising the overall loss function Ltotalwith respect to the parameters of the constituent networks, namely, the patient attribute mapping 140, the encoder 110, the latent flow 120 and the decoder 130 networks. The parameters of all four constituent networks in the cFGVAE are learned iteratively, for example via the error backpropagation algorithm.
[0067] There now follows an overview of an example specific network architecture that may be used for the cFGVAE system 100 as depicted in Figure 1. In the example, the encoder 110 and decoder 130 neural networks comprise five blocks, where each block (e.g. blocks 210, 220, 230, 240 and 250 for the down-sampling blocks in encoder architecture 200) learns shape features at a different spatial resolution, enabling the encoder and decoder networks to learn shape features with both global and local context, across multiple mesh resolutions. This is enabled by including mesh down-sampling (and up-sampling for the decoder network architecture 400) operations in each of the five blocks of the encoder 110 and decoder 130 neural networks respectively. Table 1 below summarises the architectural details of the encoder 110 and decoder 130 networks in the cFGVAE that may be used to model the shapes of organs / anatomical parts or of multi-organ shape assemblies. The architecture of residual graph convolution down-sampling blocks and the residual graph convolution up-sampling blocks used in the encoder and decoder networks, respectively, are structurally similar and are depicted in Figures 5 and 6, respectively. According to an example, the number of feature channels (i.e. the number of different learnable Chebyshev convolution filters / operations) used within each graph convolution layer in the residual graph convolution down-sampling blocks, from blocks 1 to 5 are: 16, 32, 32, 64 and 64, and the number of feature channels used within each graph convolution layer in the residual graph convolution up-sampling blocks, from blocks 1 to 5 are: 64, 64, 32, 32, and 16. Other numbers of feature channels may also be used. An example of the mesh down-sampling factors used across down-sampling blocks 1 to 5 in the encoder network architecture 200 is 4. Similarly, and example of the up-sampling factors used across up-sampling bocks 1 to 5 in the decoder network is 4. The number of feature channels and mesh resampling factors used in the disclosed encoder and decoder networks in the disclosed cFGVAE were determined empirically to be the best for the example application (i.e. cardiac modelling) and training data used. Their values may be changed and varied given different training data for the same or different applications in generative modelling of other organs or multi-organ shape assemblies. All residual graph convolution downsampling and up-sampling blocks used throughout the encoder and decoder networks in the disclosed system also contain batch normalisation layers, a residual connection and an exponential linear activation unit organised as shown in Figures 5 and 6.
[0068] There now follows some details of an example specific implementation as applicable to left ventricular (LV) shapes and associated patient attributes / conditioning variables, which may form part of application to a wider cardiac example. As will be appreciated, the general principles as discussed below may be applied to any other organ type or shape.
[0069] Table 1 : An example network architecture for the encoder 110 and decoder 120 neural networks in the cFGVAE applied to left ventricular (LV) shapes and associated patient attributes / conditioning variables cn. VLV represents the number of nodes / vertices present in each LV mesh / unstructured graph, for all samples in the training, validation and test sets used to train and evaluate the disclosed system.
[0070] The architecture of the patient attribute mapping network implementation 700 (depicted in
[0071] Figure 7), is implemented as summarised in Table 2 and is made up of a fully connected network (or multi-layer perceptron) with three linear layers, and sigmoid linear unit activation (SiLU) function layers. The patient attribute mapping network implementation 700 shown is used to transform the set of conditioning variables associated with the organ shapes input to the encoder neural network 110.
[0072] Table 2: An example network architecture for the patient attribute mapping network 140 in the cFGVAE applied to a set of patient attributes / conditioning variables associated with the left ventricular (LV) shapes used for training and testing the disclosed system. A summary of the patient attributes used and their data type / numerical encoding format, in this example implementation, is also provided.
[0073] The architecture of the latent flow network 300 (depicted in Figure 3), implemented in the disclosed cFGVAE is summarised in Table 3 and is made up of F=6 planar flow layers which receives the conditional latent representation zn(output from the encoder 200 network) as input along with the vector of conditioning variables cnand outputs a transformed conditional latent vector znwhich is input to the decoder network 400.
[0074] Table 3: An example network architecture for the latent flow network in the cFGVAE, applied to the conditional latent vector znoutput by the encoder 200 network. znare low-dimensional representations of the left ventricular (LV) shapes used for training and testing the disclosed system. Once the conditional generative model system 100 is trained, new synthetic organ shapes or multi-organ shape assemblies can be generated, conditioned on new sets of patient attributes / conditioning variables 102 inputted to the model system 100. This synthesis is done by the encoder neural network 110 sampling new conditional latent vectors given the input conditional variables, from the learned conditional latent distribution. The sampled conditional latent vectors are passed as inputs to the latent flow neural network 120 along with the new sets of patient attributes / conditioning variables 102, and the transformed conditional latent vectors output by the latent flow neural network 120 are in turn inputted to the decoder neural network 130 to reconstruct / synthesise new anatomical parts’ / organs’ shapes. This process of synthesising organ shapes using the trained conditional generative model 100, conditioned on new sets of patient attributes / conditioning variables inputted to the model system 100, is called controllable synthesis or controllable generation.
[0075] Data used for training the overall conditional generative model:
[0076] The disclosed conditional generative shape modelling system 100, also known as a conditional flow graph variational autoencoder (cFGVAE), for controllable synthesis of virtual cohorts of anatomical parts / organs (e.g. left ventricle or whole heart) may be trained and evaluated using data derived from real individuals’ data, for example as available in the UK Biobank imaging database. According to an example, the general data requirements and the specific data and pre-processing steps employed in the current example, fortraining the disclosed conditional generative shape modelling system 100 include -
[0077] All anatomical parts’ and / or organs’ shapes are represented as surface or volumetric meshes / unstructured graphs, but may also be represented as contours in the form of an unstructured graph. In essence this is so that the examples may use the described graph convolutional based methodologies. Other input formats, or formatting types, may be used in examples using different types of learnable features extraction operations. In the current example, left ventricular shapes were represented as triangular surface meshes, but other forms may equally be used (e.g. polygonal meshes).
[0078] The mesh / unstructured-graph based representations of all left ventricle shapes used in the training population are to have the same number of nodes and mesh connectivity / graph topology, prior to their use for training the disclosed generative framework. This can be achieved by coregistering all samples to establish spatial correspondences across the training population, i.e. for example all left ventricle meshes considered in the training population are co-registered to one another.
[0079] These forms of pre-processing may be avoided in some examples. For example, where the respective data is already in the appropriate data format.
[0080] According to an example, a high-resolution 3D cardiac atlas mesh (where atlas refers to an average representation of a population of cardiac meshes derived from multiple patients’ images) available from a previous study may be non-rigidly registered to contours defining the boundaries of the cardiac chambers annotated manually for each of the four cardiac chambers of interest, in cardiac cinematic magnetic resonance (cine-MR) images by expert cardiologists. These manual contours are available in the UK Biobank database. By registering the high-resolution 3D cardiac atlas mesh to the manual contours, the training populations of shapes for the cardiac structure of interest in this example, namely the left ventricle, may be generated (i.e. as the same atlas mesh was deformed / non- rigidly registered to fit the manual contours of all individuals considered from UK Biobank, spatial correspondence was automatically established across all left ventricle samples in the training population). The high-resolution cardiac atlas used to generate the training populations comprised the following parts / structures (each represented by distinct meshes) as part of a multi-part shape assembly - left and right ventricles (LV and RV), left and right atria (LA and RA) and the aortic vessel root (AR). Although the manual contours, and hence the generated training population of left ventricle meshes (based on the deformed / registered atlas mesh), were defined on cardiac cine-MR images of individuals in the UK Biobank, the disclosed approach is not dependent on the left ventricular meshes in the training population being extracted / inferred from cardiac cine-MR images only. Left ventricle meshes may be obtained from different imaging modalities and also from different patient populations and processed by co-registering the meshes to establish spatial correspondence, prior to being used to train the conditional generative model. For example, some left ventricle meshes may be extracted / inferred from cardiac cine-MR images of patients / individuals in population ‘A’, while other left ventricle meshes may be extracted / inferred from cardiac computed tomography angiography (CTA) images of patients / individuals in population ‘B’. The left ventricle meshes from both population ‘A’ and ‘B’ are then co-registered to establish spatial correspondence and form the training population.
[0081] According to an example, the overall cohort used to train and test the disclosed generative shape compositional framework comprised 2360 left ventricle shapes (represented as meshes) from 2360 individuals / participants in the UK Biobank population. Additionally, patient attributes comprising demographic information, physical and clinical measurements and lifestyle factors for the same 2360 individuals, were used as conditioning variables in this example. The complete list of patient attributes used as conditioning variables is summarised in Table 2. The subject-specific left ventricle meshes were created by registering the high-resolution cardiac atlas mesh to the manually annotated cardiac contours available for individuals in the UK Biobank database. These 2360 subjects were selected from 4000 subjects in the UK Biobank for whom the manually annotated cardiac contours are available. This selection was done based on visual assessment of the quality of the subject-specific left ventricle meshes resulting from the registration process (initially executed for all 4000 subjects). Only those meshes were retained where no obvious topological errors were introduced following registration of the atlas mesh to the manual contours, resulting in the chosen cohort of 2360 subjectspecific left ventricle meshes. The resulting cohort of 2360 subject-specific left ventricle meshes shared node / vertex-wise spatial correspondence across all samples, i.e. all meshes have the same number of vertices / nodes, the same mesh connectivity / graph topology, and each node / vertex across all meshes represent the same anatomical feature / position. Having the same mesh connectivity / unstructured graph topology means that the edges that connect each node to neighbouring nodes in the mesh / unstructured graph are identical across all subject-specific left ventricle meshes in the cohort. Another way to define having the same mesh connectivity / unstructured graph topology is to say that the adjacency matrices of all subject-specific left ventricles meshes / unstructured graphs were identical. Identical graph topology across all subjectspecific meshes enables the use of spectral (i.e. Chebyshev polynomial-based) convolutions in the graph-convolution layers used throughout the disclosed conditional generative shape modelling framework, as fixed / identical graph topology across all inputs is a pre-requisite for graph neural networks that utilise spectral convolution operations in their constituent graph convolution layers. The developed system may be adapted to work with meshes that are not identical in terms of their vertex connectivity / graph topology by utilising spatial graph convolution operations in place of spectral convolution operations. The cohort of 2360 subject-specific left ventricle meshes were randomly split into training, validation and test sets, each comprising 422, 59, and 1879 samples, respectively. An aim of the split may be to have a substantially larger test set than training set to ensure evaluation is on a more diverse population than used for training.
[0082] Examples of the present disclosure may be trained using data comprising organ meshes and linked patient attributes / conditioning variables as described above. These may be derived from existing real-life patient imaging databases, or captured afresh specifically for use in the present disclosure. In such specifically captured examples more patient attribute data may be acquired at the same time as the organ imaging, so as to provide better training data, and / or broader output capability of the so-trained system. The disclosed system may also be trained over two phases, comprising an optional pre-training phase (see Figure 8) and a main training phase (see Fig. 9).
[0083] The pre-training phase 800 may comprise initial training of the RGCDS 210-250 blocks in the encoder neural network 110, the planar flow layers 310-360 in the latent flow neural network 120, and the RGCUS 420-460 blocks in the decoder neural network 130 of the disclosed system in an unconditional manner, i.e. given pre-training data only comprising organ meshes / unstructured graphs as inputs. Or put another way, the disclosed system may be pre-trained in an unconditional manner using input organ meshes / unstructured graphs without imposing any conditioning variables / patient attributes on the learning of the weights / parameters of the encoder, latent flow and decoder neural networks. Once the pre-training phase is complete, the learned weights / parameters of the encoder 110, latent flow 120 and decoder 130 neural networks may be used to initialise the corresponding weights of the corresponding networks in the conditional flow graph variational autoencoder of the disclosed system. The unconditional pre-training phase may then be followed by a conditional training phase. In the conditional training phase, real-life acquired organ meshes / unstructured graphs and linked patient attributes / conditioning variables are input to the disclosed system, wherein, the patient attribute mapping network 140 is initialised with random weights / parameters and the RGCDS 210-250 blocks of the encoder 110 neural network, the planar flow layers 310-360 in the latent flow 120 neural network, and the RGCUS 420-460 blocks in the decoder 130 neural network are initialised with the weights / parameters learned from the pre-training phase. Note, the fully connected networks 590-591 in the encoder 110 network and the fully connected network 370 in the latent flow 120 network are not pre-trained and are initialised with random weights (similar to the patient attribute mapping 140 network) in the training phase. In some examples, the systems can be re-trained using new sets of data in the unconditional pre-training and conditional training portions. Referring to Fig. 8, in broad outline, the pre-training phase 800 may start 801 off by receiving suitable pre-training data 810 (e.g. organ only mesh data from real-life medical imaging data of organs or the like), These are encoded at step 820, and one or more latent flow neural networks are applied at step 830. The output from the one or more latent flow networks are then decoded at step 840, to thereby output reconstructed / decoded mesh(es) at step 850. Typically, the training is repeated over a number of epochs, n, 845, until a suitable level of training is reached.
[0084] Referring to Fig. 9, in broad outline, the main conditional training phase 900 may start 901 off by apply the afore-described unconditional pre-training 800, and applying the initialisation 905 of the encoder, latent flow and decoder neural networks with parameters derived from the pre-training 800. However, in some examples, the main training 900 can be done without pre-training 800. The main training 900 proceeds by receiving suitable training data 910 (e.g. this time organ data as well as related patient attribute data). The patient attribute data is mapped at step 920, for use in encoding mesh(es) as described in detail above at step 930. The latent flow neural networks are applied at step 940, the output of which is decoded at step 950. Again, typically, the training is repeated over a number of epochs, n, 955, until a suitable level of training is reached.
[0085] Once the method / system is suitably trained, the system may then proceed to carry out inference in order to controllably synthesise new instances of virtual patients according to predefined, or pre-chosen, target population parameters. Figure 10 shows, in broad outline, such an inference method. This starts off at 1001 , and receives the predefined new patient attributes / conditioning variables and a randomly sampled latent representation (i.e. starting point) from the learned posterior latent distribution at step 1010, to which are applied the latent flow layers (to transform the latent representations as learned in the training phase). At 1030, the transformed latent representations are decoded to form the new reconstructed meshes in alignment with (i.e. subject to) those predefined patient attributes, which can then be stored at 1040, or output at 1050, and then subsequently used in any form of existing, or yet to be devised, in-silico based medical device or drug development, testing or approval process.
[0086] Examples of the present disclosure may be implemented by suitably programmed computer hardware. Fig. 11 is a block diagram 1100 illustrating components, according to some example embodiments, able to read instructions from a machine-readable or computer-readable medium (e.g., a non-transitory machine-readable storage medium) and perform any one or more of the methodologies discussed herein, hence providing the apparatus or system to carry out the described conditional, or controllable virtual population synthesis. Specifically, Fig. 11 shows a diagrammatic representation of hardware resources 1105 including one or more processors (or processor core(s), tile(s), or any other sub-unit type of processing resource) 1110, one or more linked memory / storage devices 1120, and one or more communication resources 1130, each of which may be communicatively coupled via a bus 1140. The processors 1110 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), neural processing unit (NPU), a digital signal processor (DSP) such as a baseband processor, an application specific integrated circuit (ASIC), a cloud processing function (such as AWS instance), another processor, or any suitable combination thereof) may include, for example, a processor 1112 and a processor 1114.
[0087] The linked memory / storage devices 1120 may include main memory, disk storage, or any suitable combination thereof. The memory / storage devices 1120 may include, but are not limited to any type of volatile or non-volatile memory such as dynamic random access memory (DRAM), static random-access memory (SRAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), Flash memory, solid-state storage device (SSD), magnetic storage based hard disk drive (HDD) media, etc.
[0088] The communication resources 1130 may include interconnection or network interface components or other suitable devices to communicate with one or more peripheral devices 1104 or one or more databases 1106 via a network 1108, or via a direct communications link 1132. For example, the communication resources 1130 may include wired communication components (e.g., for coupling via Ethernet, a Universal Serial Bus (USB) or the like), cellular communication components, NFC components, Bluetooth® components (e.g., Bluetooth® Low Energy), Wi-Fi® components, and other communication components.
[0089] The instructions 1150, neural network model(s) 1152, conditioning variable(s) 1154, and / or mesh(es) data 1156 for use in the disclosed methods or systems may comprise software, a program, an application, an applet, an app, or other executable code for causing at least any of the processors 1110 to perform any one or more of the methodologies discussed herein. As shown, the instructions 1150, neural network model(s) 1152, conditioning variable(s) 1154, and / or mesh(es) data 1156 for use in the disclosed methods or systems may reside, completely or partially, within at least one of the processors 1110 (e.g., within the processor’s cache memory), the memory / storage devices 1120, or any suitable combination thereof. Furthermore, any portion of the instructions 1150, neural network model(s) 1152, conditioning variable(s) 1154, and / or mesh(es) data 1156 for use in the disclosed methods or systems may be transferred to the hardware resources 1105 from any combination of the peripheral devices 1104 or the databases 1106. Accordingly, the memory of processors 1110, the memory / storage devices 1120, the peripheral devices 1104, and the databases 1106 are examples of computer-readable and machine-readable media.
[0090] In some embodiments, the electronic device(s), network(s), system(s), chip(s) or component(s), or portions or implementations thereof, of Figures 11 , or some other figure herein may be configured to perform one or more processes, techniques, or methods as described herein, or portions thereof.
[0091] In a particular example the disclosed conditional generative shape compositional framework may be implemented using PyTorch, and PyTorch Geometric python libraries on a PC with an NVIDIA RTX 2080Ti graphics processing unit (GPU).
[0092] In an example, the conditional generative model in the disclosed system was trained using the AdamW optimizer with an initial learning rate of le-03and a batch size of 16, for a maximum of 1000 training epochs.
[0093] While training the conditional generative model system 100, a warm-up strategy may be adopted to improve stability and prevent model collapse in the learned conditional posterior distribution over the latent variables. This is achieved by initially training the constituent networks in the model system 100, with the weight of KL loss term (i.e. OJO in equation 7) set to zero for 200 epochs (i.e. they are trained as plain autoencoders). The learned weights initialise the subsequent training step for constituent networks, wherein, the weight of the KL loss term is initially set to a small value, i.e. 1e-6, and subsequently increased each epoch, for example by multiplying by a factor of 1 .25, up to a maximum value of 1e-4. All hyperparameters associated with the constituent networks’ architectures, training routines, and the optimiser used were tuned / set empirically based on the average loss values observed for the validation set in preliminary prototyping experiments.
[0094] Examples provide a computer-implemented method for training a machine learning system to generate at least one organ shape for use in an in-silico assessment, the method comprising: receiving at least one organ mesh, and at least one patient attribute data associated with the at least one organ mesh; encoding the at least one organ mesh using data derived from the at least one associated patient attribute to learn a conditional latent representation of the at least one organ mesh; applying at least one latent flow layer to the conditional latent representation of the at least one organ mesh to form a transformed conditional latent representation of the at least one organ mesh, wherein the at least one latent flow layer transforms a unimodal conditional latent representation output from the encoder into a multi-modal conditional latent representation; and decoding the transformed conditional latent representation of the at least one organ mesh to form a reconstructed organ mesh dependent on the at least one patient attribute, wherein the reconstructed organ mesh data comprises an instance of a virtual patient to be included in a virtual population for use in the in-silico assessment.
[0095] In some examples, the at least one latent flow layer comprises a set of sequential, invertible transformations.
[0096] In some examples, the sequential, invertible transformations are planar flow functions.
[0097] In some examples, the planar flow function f is defined as:
[0098] / (zn) = zn+ ( u * h(wTzn+ 6)) wherein: zncomprises an input initial conditional latent representation; it, wTand b are learnable or trainable parameters in the planar flow layer, each comprising the same number of parameters as the dimension of the inputted conditional latent representation zn; and h is a sequential, invertible element-wise non-linear function. In some examples, h is a Hyperbolic Tangent (TanH) function.
[0099] In some examples, the set of sequential, invertible transformations comprises F functions, wherein their application results in the multi-modal probability distribution given by: wherein, (zn0) is a initial conditional posterior latent distribution approximated by an encoder from which the initial conditional latent representation zno is sampled, In is the natural logarithm operation and qF(znF) is a multi-modal conditional posterior latent distribution approximated by applying a sequence of F successive planar flow functions. Some examples further comprise transforming the at least one patient attribute to form a transformed version of the at least one patient attribute, wherein the transforming comprises applying a first fully connected neural network layer to the at least one patient attribute.
[0100] In some examples, the applying the first fully connected neural network layer comprises: applying a first fully connected layer; applying a first sigmoid linear unit activation (SiLU) layer; and applying a second fully connected layer.
[0101] Some examples further comprise adding the transformed version of the at least one patient attribute to the transformed conditional latent representation of the at least one organ mesh.
[0102] In some examples, the method further comprises mapping the at least one patient attribute to a corresponding mapped patient attribute, wherein said mapping comprises: applying at least one fully connected layer; and applying at least one sigmoid linear unit activation (SiLU) layer.
[0103] In some examples, the applying at least one fully connected layer and applying at least one SiLU layer are carried out alternately.
[0104] In some examples, the encoding the at least one organ mesh using data derived from the at least one associated patient attribute further comprises applying at least one feature extraction block to form the conditional latent representation, wherein the at least one feature extraction block takes as an input a transformed version of the corresponding mapped patient attribute.
[0105] In some examples, the encoding the at least one organ mesh using data derived from the at least one associated patient attribute comprises applying at least one corresponding mapped patient attribute to at least one feature extraction block used in the encoding to learn a conditional latent representation of the at least one organ mesh.
[0106] Some examples further comprise transforming the corresponding mapped patient attribute by: applying a second fully connected neural network to the mapped patient attribute; and applying a third fully connected neural network to the mapped patient attribute; wherein the second and third connected neural networks each comprises at least one fully connected layer, and a SiLU layer.
[0107] In some examples the encoder comprises N feature extraction blocks.
[0108] In some examples the at least one feature extraction block comprises a sequential combination of convolution operations, normalisation operations and graph pooling operations arranged to form a down-sampling block.
[0109] In some examples the convolution operations comprises at least one spectral graph convolution operation.
[0110] In some examples the at least one spectral graph convolution operation comprises a residual graph convolution operation.
[0111] In some examples the residual graph convolution operation comprises a Chebyshev graph convolution.
[0112] In some examples decoding the transformed conditional latent representation of at least one organ mesh further comprises: applying at least one feature extraction block to form a sampled set of conditional latent representations. In some examples the feature extraction block comprises a sequential combination of convolution operations, normalisation operations, and mesh up-sampling operations arranged to form an up-sampling block.
[0113] In some examples the convolution operations comprises at least one spectral graph convolution operation.
[0114] In some examples the spectral graph convolution operation comprises a residual graph convolution operation.
[0115] In some examples the residual graph convolution operation a Chebyshev graph convolution.
[0116] Examples also provide a method of controllably synthesising at least one organ part for use in an in-silico assessment comprising: training a machine learning system according to any of the previous examples, then, receiving at least one new patient attribute and a random latent representation; and inferencing a new reconstructed organ shape using the trained system and the received at least one new patient attribute and a random latent representation.
[0117] Examples also provide a method of developing a new medical device, or medicinal drug, comprising carrying out in-silico testing of the new medical device, or medicinal drug, on at least one organ shape reconstructed using any of the disclosed example method.
[0118] Examples also provide a non-transitory computer readable medium comprising instructions which, when executed by one or more processors cause the one or more processors to carry out any of the disclosed methods.
[0119] Examples also provide a machine learning system for generating at least one conditionally synthesised organ mesh, comprising: a patient mapping neural network configured to receive one or more real life example patient attributes and output one or more associated mapped patient attributes; an encoder neural network configured to receive at least one example organ mesh and the mapped patient attributes, and output learned conditional latent representations of said received at least one organ mesh based on said mapped patient attributes; a latent flow neural network configured to apply a plurality of flow functions to transform the learned conditional latent representations of said received at least one organ mesh into transformed conditional latent representations; and a decoder neural network configured to receive the transformed conditional latent representations and output synthesised reconstructed organ meshes based on the transformed conditional latent representations; wherein the patient mapping neural network, encoder neural network, latent flow neural network and decoder neural network are trained collectively, in a training phase, on an initial set of real-life medical imaging data and associated patient attribute data, and then, during an inference phase, output realistic synthesised reconstructed organ meshes based on a new set of patient attributes data.
[0120] In some examples the at least one conditionally synthesised organ mesh comprises an organ mesh reconstructed dependent on a set of patient attributes data as applied by the trained machine learning system after training based on a plurality of real life organ imaging data and associated patient attributes.
[0121] In some examples the patient mapping neural network comprises: at least one fully connected layer; and at least one SiLU layer. In some examples the latent flow neural network is configured to apply a set of sequential, invertible transformations. In some examples, a set is just one.
[0122] In some examples the sequential, invertible transformations are planar flow functions. Or they can be affine flow, Radial flow, or any other similar flow functions.
[0123] In some examples the planar flow function f is defined as:
[0124] / (zn) = zn+ ( u * h(wTzn+ 6)) wherein: zncomprises an input initial conditional latent representation; it, wTand b are learnable or trainable parameters in the planar flow layer, each comprising the same number of parameters as the dimension of the inputted conditional latent representation zn; and h is a sequential, invertible element-wise non-linear function.
[0125] In some examples h is a Hyperbolic Tangent (TanH) function.
[0126] In some examples the set of sequential, invertible transformations comprises F functions, wherein their application results in the multi-modal probability distribution given by: where, (zn0) is a initial conditional posterior latent distribution approximated by an encoder from which the initial conditional latent representation zno is sampled, In is the natural logarithm operation and qF(znF) is a multi-modal conditional posterior latent distribution approximated by applying a sequence of F successive planar flow functions.
[0127] In some examples the latent flow network is further configured to transform the at least one patient attribute to form a transformed version of the at least one patient attribute.
[0128] In some examples the latent flow network is further configured to apply a first fully connected neural network layer to the at least one patient attribute.
[0129] In some examples the latent flow network is further configured to: apply a first fully connected layer; apply a first sigmoid linear unit activation (SiLU) layer; and apply a second fully connected layer.
[0130] In some examples the latent flow network is further configured to add a transformed version of the at least one patient attribute to the transformed conditional latent representation of the at least one organ mesh.
[0131] In some examples the patient mapping neural network is further configured to: map the at least one patient attribute to a corresponding mapped patient attribute.
[0132] In some examples the patient mapping neural network is further configured to: apply at least one fully connected layer; and apply at least one sigmoid linear unit activation (SiLU) layer.
[0133] In some examples the patient mapping neural network is further configured to apply the at least one fully connected layer and apply the at least one SiLU layer alternately.
[0134] In some examples the encoder neural network is further comprises at least one feature extraction block configured to form the conditional latent representation, wherein the at least one feature extraction block takes as an input a transformed version of the corresponding mapped patient attribute. In some examples the encoder neural network is further configured to: apply a second fully connected neural network to the mapped patient attribute; and apply a third fully connected neural network to the mapped patient attribute; wherein the second and third connected neural networks each comprises at least one fully connected layer, and a SiLU layer.
[0135] In some examples the at least one feature extraction block further comprises a sequential combination of convolution operations, normalisation operations and graph pooling operations arranged to form a down-sampling block.
[0136] In some examples the convolution operations comprises at least one spectral graph convolution operation.
[0137] In some examples the at least one spectral graph convolution operation comprises a residual graph convolution operation.
[0138] In some examples the residual graph convolution operation comprises a Chebyshev graph convolution.
[0139] In some examples the decoder is further configured to apply at least one feature extraction block to form a sampled set of conditional latent representations.
[0140] In some examples the feature extraction block comprises a sequential combination of convolution operations, normalisation operations, and mesh up-sampling operations arranged to form an up-sampling block.
[0141] In some examples the convolution operations comprises at least one spectral graph convolution operation.
[0142] In some examples the spectral graph convolution operation comprises a residual graph convolution operation.
[0143] In some examples the residual graph convolution operation a Chebyshev graph convolution. The foregoing examples may be combined in any reasonable combination,
[0144] In the foregoing, functions may be described as neural networks, modules or blocks, i.e. functional units that are operable to carry out the described function, algorithm, or the like. These terms may be interchangeable. Where networks, modules, blocks, or functional units have been described, they may be formed as processing circuitry, where the circuitry may be general purpose processor circuitry configured by program code to perform specified processing functions and / or Al specific processor circuitry (e.g. NPU, GPU, etc) configured by program code to perform specified processing functions. The circuitry may also be configured by modification to the processing hardware. Configuration of the circuitry to perform a specified functions may be entirely in hardware, entirely in software or using a combination of hardware modification and software execution. Program instructions may be used to configure logic gates of general purpose or special-purpose processor circuitry to perform a processing function.
[0145] Circuitry may be implemented, for example, as a hardware circuit comprising custom Very Large Scale Integrated, VLSI, circuits or gate arrays, off-the-shelf semiconductors such as logic chips, transistors, or other discrete components. Circuitry may also be implemented in programmable hardware devices such as field programmable gate arrays, FPGA, programmable array logic, programmable logic devices, A System on Chip, SoC, or the like. Machine readable program instructions may be provided on a transitory medium such as a transmission medium or on a non-transitory medium such as a storage medium. Such machine readable instructions (computer program code) may be implemented in a high level procedural or object oriented programming language. However, the program(s) may be implemented in assembly or machine language, if desired. In any case, the language may be a compiled or interpreted language, and combined with hardware implementations. Program instructions may be executed on a single processor or on two or more processors in a distributed manner.
Claims
Claims:1 . A computer-implemented method fortraining a machine learning system to generate at least one organ shape for use in an in-silico assessment, the method comprising: receiving at least one organ mesh, and at least one patient attribute data associated with the at least one organ mesh; encoding the at least one organ mesh using data derived from the at least one associated patient attribute to learn a conditional latent representation of the at least one organ mesh; applying at least one latent flow layer to the conditional latent representation of the at least one organ mesh to form a transformed conditional latent representation of the at least one organ mesh, wherein the at least one latent flow layer transforms a unimodal conditional latent representation output from the encoder into a multi-modal conditional latent representation; and decoding the transformed conditional latent representation of the at least one organ mesh to form a reconstructed organ mesh dependent on the at least one patient attribute, wherein the reconstructed organ mesh data comprises an instance of a virtual patient to be included in a virtual population for use in the in-silico assessment.
2. The method of claim 1 , wherein the at least one latent flow layer comprises a set of sequential, invertible transformations.
3. The method of claim 2, wherein the sequential, invertible transformations are planar flow functions.
4. The method of claim 3, wherein the planar flow function f is defined as: / (zn) = zn+ ( u * h(wTzn+ 6)) wherein: zncomprises an input initial conditional latent representation; u, wTand b are learnable or trainable parameters in the planar flow layer, each comprising the same number of parameters as the dimension of the inputted conditional latent representation zn; and h is a sequential, invertible element-wise non-linear function.
5. The method of claim 4, wherein h is a Hyperbolic Tangent (TanH) function.
6. The method of claim 2, wherein the set of sequential, invertible transformations comprises F functions, wherein their application results in the multi-modal probability distribution given by:where, (zn0) is a initial conditional posterior latent distribution approximated by an encoder from which the initial conditional latent representation zno is sampled, In is the natural logarithm operation and qF(znF) is a multi-modal conditional posterior latent distribution approximated by applying a sequence of F successive planar flow functions.
7. The method of claim 1 , further comprising transforming the at least one patient attribute to form a transformed version of the at least one patient attribute, wherein the transforming comprises applying a first fully connected neural network layer to the at least one patient attribute.
8. The method of claim 7, wherein applying the first fully connected neural network layer comprises: applying a first fully connected layer; applying a first sigmoid linear unit activation (SiLU) layer; and applying a second fully connected layer.
9. The method of claim 7, further comprising adding the transformed version of the at least one patient attribute to the transformed conditional latent representation of the at least one organ mesh.
10. The method of claim 1 , wherein the method further comprises mapping the at least one patient attribute to a corresponding mapped patient attribute, wherein said mapping comprises: applying at least one fully connected layer; and applying at least one sigmoid linear unit activation (SiLU) layer.11 . The method of claim 10, wherein applying at least one fully connected layer and applying at least one SiLU layer are carried out alternately.
12. The method of claim 10, wherein encoding the at least one organ mesh using data derived from the at least one associated patient attribute further comprises: applying at least one feature extraction block to form the conditional latent representation, wherein the at least one feature extraction block takes as an input a transformed version of the corresponding mapped patient attribute.
13. The method of claim 12, further comprising transforming the corresponding mapped patient attribute by: applying a second fully connected neural network to the mapped patient attribute; and applying a third fully connected neural network to the mapped patient attribute; wherein the second and third connected neural networks each comprises at least one fully connected layer, and a SiLU layer.
14. The method of claim 12, wherein the at least one feature extraction block comprises a sequential combination of convolution operations, normalisation operations and graph pooling operations arranged to form a down-sampling block.
15. The method of claim 14, wherein the convolution operations comprises at least one spectral graph convolution operation.
16. The method of claim 15, wherein the at least one spectral graph convolution operation comprises a residual graph convolution operation.
17. The method of claim 16, wherein the residual graph convolution operation comprises a Chebyshev graph convolution.
18. The method of claim 1 , wherein decoding the transformed conditional latent representation of at least one organ mesh further comprises: applying at least one feature extraction block to form a sampled set of conditional latent representations.
19. A non-transitory computer readable medium comprising instructions which, when executed by one or more processors cause the one or more processors to carry out the method of claim 1 .
20. A machine learning system for generating at least one conditionally synthesised organ mesh, comprising: a patient mapping neural network configured to receive one or more real life example patient attributes and output one or more associated mapped patient attributes; an encoder neural network configured to receive at least one example organ mesh and the mapped patient attributes, and output learned conditional latent representations of said received at least one organ mesh based on said mapped patient attributes; a latent flow neural network configured to apply a plurality of flow functions to transform the learned conditional latent representations of said received at least one organ mesh into transformed conditional latent representations; and a decoder neural network configured to receive the transformed conditional latent representations and output synthesised reconstructed organ meshes based on the transformed conditional latent representations; wherein the patient mapping neural network, encoder neural network, latent flow neural network and decoder neural network are trained collectively, in a training phase, on an initial set of real-life medical imaging data and associated patient attribute data, and then, during an inference phase, output realistic synthesised reconstructed organ meshes based on a new set of patient attributes data.
Citation Information
Patent Citations
Three-dimensional shape reconstruction from a topogram in medical imaging
CN113272869A
Modelling method using a conditional variational autoencoder
WO2021058710A1