Method and apparatus for generating a virtual population of anatomical structures

A generative framework using graph convolutional neural networks synthesizes virtual anatomical structures from non-redundant datasets, addressing the limitations of traditional clinical trials by creating diverse virtual populations for efficient in silico testing of medical devices.

JP2025535033APending Publication Date: 2025-10-22UNIVERSITY OF LEEDS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025519166
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-03
Filing Date
2023-10-03
Publication Date
2025-10-22

AI Technical Summary

Technical Problem

Traditional clinical trials for medical devices are limited by the range of operational populations that can be targeted, leading to inadequate verification of safety and effectiveness, and the high cost and inefficiency of regulatory approval processes.

Method used

A generative framework that synthesizes virtual populations of anatomical structures using non-redundant or partially overlapping datasets, utilizing graph convolutional neural networks to model and simulate multi-part organs, allowing for the generation of anatomically plausible virtual chimeras.

Benefits of technology

Enables the creation of diverse virtual patient populations for in silico testing, reducing the need for extensive clinical trials and enhancing the evaluation of medical devices across a broader range of patient characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025535033000001_ABST
    Figure 2025535033000001_ABST
Patent Text Reader

Abstract

A computer-implemented method for generating a virtual chimera population of multi-part organ shapes for use in in silico testing, where the organ shapes are represented as surface or volumetric meshes, is provided, the method including: using a part-aware generative model to learn a latent representation of each part of the multi-part organ shape to be included in the virtual population and outputting composite parts of the multi-part organ shape; using a spatial configuration model to align the output composite parts of the multi-part organ shape and output anatomically significant instances of the entire multi-part organ shape as virtual chimeras; and storing the virtual chimeras in the virtual population for use in in silico studies. Also provided are an apparatus and a computer-readable medium for practicing this method.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to computer-implemented methods, computer programs, and apparatus for generating virtual populations of anatomical structures for use in in silico research and their use in the design, manufacture, testing, and regulatory approval of medical devices or drugs. [Background technology]

[0002] The cost of developing medical devices has steadily increased in recent years due to lengthy regulatory approval processes and costly clinical trials. A high rate of late-stage failures of medical devices in clinical trials can result in significant financial losses for medical device manufacturers, yet the financial losses resulting from these late-stage failures are often sequentially deferred and imposed on the much smaller number of successful medical device cases that ultimately achieve approval. As a result, the economic burden on healthcare systems and the medical device industry worldwide is rapidly becoming unsustainable, clearly calling for a paradigm shift in medical device innovation and the regulatory approval process.

[0003] Traditional (i.e., in vivo) clinical trials of medical devices are often limited by the limited range of operational populations (e.g., patient variations) that can be targeted by a medical device when evaluating its respective safety and effectiveness. This is likely due to the difficulty of recruiting sufficient patients to encompass the range of target populations for a medical device under development (i.e., patients in whom the device is intended to be used in routine clinical care). If the safety and effectiveness of a medical device are inadequately verified, particularly in operational populations not adequately evaluated in pre-market approval clinical trials, serious adverse reactions may only become apparent after marketing approval.

[0004] In silico testing (IST) offers an alternative and complementary pathway for medical device innovation and regulatory approval processes, which underpin the path to commercialization and adoption in routine patient care. IST uses computational modeling and simulation to virtually evaluate the safety and effectiveness of medical devices in virtual patient populations and provide digital (or in silico) evidence, as opposed to the real-world evidence that supports regulatory approval of medical devices under development. In this way, IST potentially allows for the investigation of device performance across a broader range of patient characteristics than would be possible in actual clinical trials, potentially refining, reducing, and partially replacing in vivo clinical trials for medical device testing. Summary of the Invention

[0005] The technology described herein provides a generative framework that enables simulation of the anatomy and physiology of animals, including humans, for the design and testing of medical devices or drugs. In particular, the synthesis of plausible multi-part organs (or multi-organ shape assemblies) uses the shapes of macroscopic anatomical structures, represented as unstructured graphs, as input data. The disclosed generative framework allows for modeling each part / organ individually or simultaneously, and then encoding the single-part or multi-part anatomical shapes as a latent representation (i.e., a low-dimensional representation (e.g., a vector)). Decoding this latent representation then allows for the individual synthesis of each part instance or the simultaneous synthesis of multiple part instances. The synthesized parts can be spatially configured (i.e., jointly and spatially combined) to generate plausible multi-part / multi-organ shape assemblies, referred to herein as "virtual chimeras." The described methods and associated apparatus advantageously allow for training using non-redundant or partially overlapping information from various datasets, after which the trained system can be used to synthesize virtual chimera populations. For example, in the specific area of ​​cardiac modeling, a complete cardiac virtual chimeric population can be generated by combining aortic vascular geometries extracted from computed tomography angiography (CTA) data of a clinical trial cohort with cardiac ventricular geometries extracted from cine magnetic resonance images (cine MRI) acquired in the general population (such as those available from UK Biobank).

[0006] Generating a cardiac virtual population including aortic vascular geometry is not feasible using cardiac cine magnetic resonance images alone, as the field of view of such images does not capture the aorta. However, cardiac CTA images capture the aorta, allowing for the extraction and 3D characterization of its respective geometry. An advantage of the proposed method is that it allows for the combination of one or more organ regions extracted from one type of image (i.e., imaging modality) for a specific population (e.g., the aorta from a CTA of a transcatheter aortic valve implantation (TAVI) clinical trial cohort) with other regions of the same or other organs from one or more other imaging modalities that may be acquired for different patient populations (e.g., the ventricles from cine magnetic resonance imaging of population imaging initiatives such as UK Biobank).

[0007] The disclosed methods may also include a construction neural network that spatially organizes real or synthetic anatomical sites belonging to a multi-site or multi-organ assembly of interest. In such examples, the construction neural network can learn spatial relationships between adjacent anatomical sites and / or organs of interest and recover affine transformations (i.e., linear mappings) and non-rigid spatial transformations that can organize sites and / or organs into anatomically plausible combined (global) assemblies. Such construction neural networks are advantageous because they can be constructed, for example, in a self-supervised manner and trained on both non-overlapping and partially overlapping data, thereby making them independent of the availability of fully overlapping data for training.

[0008] Compared to previously known techniques that rely on the availability of fully redundant data in a training population, i.e., that require all site and / or organ structures to be available for each patient included in the training population, the disclosed techniques are capable of generating more diverse virtual patient populations and can utilize cross-modality and cross-population data sources. The disclosed generative framework for synthesizing virtual chimeric populations is illustrated by the schematic diagram shown in FIG. 1 for the specific use case of generating a multi-site whole heart (cardiac). However, the disclosed techniques can be similarly used to synthesize any multi-site or multi-organ geometric assembly (e.g., blood vessels, bone structures, or any other partial selection of physical entities within an anatomy).

[0009] An advantage of the present disclosure is that it allows heterogeneous medical datasets containing potentially non-overlapping or only partially overlapping anatomical structures to be utilized in training generative models of multi-part anatomical shape assemblies. Thus, examples of the present disclosure address the problem of generating models (i.e., so-called "virtual chimeras") by combining data (or datasets) from different subjects containing partially overlapping (or non-overlapping) anatomical structures in a generative shape modeling framework that can synthesize multi-part shape assemblies that are representative of realistic native anatomy.

[0010] The disclosed generative shape configuration learning techniques utilize configuration networks that predict rigid or affine transformations to compose individual parts that are combined into coherent multi-part assemblies, but these can be extended to "stitch" individual parts (anatomical structures) together by including a non-rigid registration component in the configuration neural network.

[0011] An advantage of the present disclosure is that heterogeneous patient datasets can be used to train various portions of the disclosed generative framework (e.g., portions including non-overlapping and partially overlapping anatomical structures / multi-part shape assemblies). This provides a generative framework that can effectively capture variations in the shape of individual parts in an organ or organ assembly, as well as accommodate different topologies at all parts within the organ or organ assembly. Ultimately, while being anatomically realistic, it does not rely on the availability of all parts of the shape assembly across all samples in the training set (i.e., it does not require fully overlapping input data).

[0012] Examples of the present invention will now be described in detail with reference to the accompanying drawings. [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 illustrates an exemplary system for generating (or training) a virtual organ assembly, according to one example of the present disclosure. [Figure 2] FIG. 2 illustrates an exemplary independent generator for use in the system of FIG. 1 according to an example of the present disclosure. [Figure 3] 2 illustrates an exemplary slave generator for use in the system of FIG. 1 according to an example of the present disclosure. [Figure 4] FIG. 2 illustrates an example architecture of a residual graph convolution downsampling block for use in all graph neural networks of the system of FIG. 1 according to an example of the present disclosure. [Figure 5] FIG. 2 illustrates an example architecture of a residual graph convolution upsampling block for use in all graph neural networks of the system of FIG. 1 according to an example of the present disclosure. [Figure 6] FIG. 2 illustrates an exemplary graph neural network-based architecture of an affine spatial construction network for use in the spatial construction module of FIG. 1 according to an example of the present disclosure. [Figure 7]FIG. 2 illustrates an exemplary graph neural network-based architecture of a non-rigid spatial construction network for use in the spatial construction module of FIG. 1 according to an example of the present disclosure. [Figure 8] FIG. 1 illustrates a method for training a part-aware model according to an example of the present disclosure. [Figure 9] FIG. 1 illustrates a method for training a spatial configuration model according to an example of the present disclosure. [Figure 10] FIG. 1 illustrates a method for generating a virtual population for use in in silico testing, according to one example of the disclosure. [Figure 11] FIG. 1 illustrates an apparatus for carrying out a method for generating a virtual population for use in in silico testing, according to one example of the disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0014] The technical effects and advantages of the present disclosure are presented below as several examples, which are merely examples of novel and original features, and it is intended that these disclosed examples can be fully combined in any reasonable combination (e.g., defined as true alternatives) unless otherwise expressly stated.

[0015] 1 illustrates, in general terms, an exemplary system 100 for generating virtual organ assemblies according to one example of the present disclosure. In its most general form (also referred to as a generative geometric modeling system), the disclosed system 100 comprises two main components: a part-aware generation module 110 and a spatial configuration module 120. The part-aware generation module 110 takes as input one or more sets of patient data 101 (e.g., partially overlapping or non-overlapping data in the form of meshes 101a-101e) and comprises a part-aware generative model block 102 that provides as output a trained latent representation 103 of these meshes. This output is used to define a set of composite parts (of organs) or organ-shaped (sub)structures (e.g., in a multi-organ assembly). The spatial configuration module 120 comprises a spatial configuration model block 122 that takes the composite parts / structures 115 as input and forms a realistic set of spatially configured parts (e.g., organ parts) or part assemblies 125 (e.g., of a multi-part organ assembly) of multi-shape parts by their respective spatial alignment.

[0016] According to one example, both modules 110 and 120 are parameterized (defined) by graph convolutional neural networks. For use in the part-aware generative model block 102, a generative model is an algorithm capable of learning from real data and synthesizing virtual instances similar to (but not identical to) the real data instances on which it trains. The part-aware generative model 102 is used to synthesize one or more virtual instances of anatomical parts / shapes of a multi-part organ (i.e., a single organ with a predefined set of parts or sub-parts) or a multi-part organ shape or assembly (i.e., multiple organs with a predefined set of parts / sub-parts). In the following exemplary disclosure, the human heart, which includes multiple parts / structures, is an exemplary application. However, other undescribed examples include multiple abdominal organs, such as the liver, kidneys, and pancreas. The spatial configuration model 122 is used to spatially organize (i.e., move / align or "configure") one or more synthetic virtual instances of parts (of an organ or organ assembly) into an anatomically meaningful multi-part / multi-organ shape assembly. Each spatially configured multi-site / multi-organ shape assembly output by the spatial configuration model 122 is referred to as a "virtual chimera" (considered to be the only complete model) for each organ, and a "virtual population" is generated by combining them together to generate multiple instances of such virtual chimeras, which may also be referred to as a "virtual chimera population."

[0017] As mentioned above, FIG. 1 is a schematic diagram providing an overview of a generative shape modeling system 100 for synthesizing a virtual chimeric population from non-overlapping and partially overlapping datasets of anatomical shapes. In this example, the partially overlapping / non-overlapping nature of the input data used is indicated by shading. The first step / component of the generative shape modeling system 100 is a part-aware generative model 102, which has two different implementations: an independent generator and a dependent generator. These two different implementations are not used simultaneously; either an independent generator is used or a dependent generator is used. That is, these are two ways of addressing the same problem. The independent generator implementation is illustrated by the schematic diagram in FIG. 2, which shows how an independent generator can be configured and trained to learn a latent representation (i.e., a hidden low-dimensional representation) for each anatomical part or organ of interest independently from the others. Meanwhile, FIG. 3 illustrates a dependent generator implementation, which is configured and trained to simultaneously learn a shared latent representation across the shapes of multiple parts and / or organs of interest to synthesize a desired multi-part / multi-organ shape assembly.

[0018] In one example based on cardiac modeling, the generative shape modeling system 100 can be used to learn individual or shared latent representations of the anatomical regions that make up a human heart, including in this example the left ventricle (LV), right ventricle (RV), left atrium (LA), right atrium (RA), and aortic root (AR), using data from example patient data covering a variety of individuals (and / or derived from a variety of data sources), particularly data that does not fully overlap the example patient data (i.e., not all anatomical regions are available for all patient data used to build the system). The disclosed generative shape modeling system 100 is not limited to modeling the aforementioned cardiac structures and can be extended to include additional cardiac structures, such as coronary arteries and their respective branches, pulmonary arteries, pulmonary veins, and heart valves, if available. Additionally, the disclosed system 100 is not limited to cardiac modeling and can be used to model other organs as multi-part shape assemblies (e.g., lungs, liver, spine) or to model multiple organs as instances of multi-part shape assemblies (e.g., thoracic and abdominal organs). That is, while certain exemplary embodiments described herein focus on generating virtual cohorts of the heart, other embodiments may be adapted to synthesize virtual cohorts of other multi-part organs or multi-organ shape assemblies.

[0019] Referring now to FIG. 2 , the independent generator 200 can include multiple (N) independent generator sub-models parameterized (i.e., defined) by graph convolutional variational autoencoders (gcVAEs) 202(A)-202(E). Each sub-model (i.e., gcVAE) corresponds to one of the N anatomical regions and / or organs that combine to generate the multi-region / multi-organ shape assembly of interest. In the example of FIG. 2 , N=5 because there is a gcVAE for each of the left ventricle, right ventricle, left atrium, right atrium, and aortic root that combine to form the multi-region shape assembly (in this case, a human heart). That is, sub-model (A) is a region-specific gcVAE network architecture used in the independent generator 200 to learn variations in left ventricular (LV) shape observed across a training population. Sub-model (B) is a region-specific gcVAE network architecture used to learn variations in right ventricle (RV) shape observed across a training population. Submodel (C) is a region-specific gcVAE network architecture used to learn variations in left atrium (LA) shape observed across the entire training population. Submodel (D) is a region-specific gcVAE network architecture used to learn variations in right atrium (RA) shape observed across the entire training population. Submodel (E) is a region-specific gcVAE network architecture used to learn variations in aortic root (AR) shape observed across the entire training population.

[0020] After training, sampling each region-specific gcVAE 202(A)-202(E) allows us to generate new regions (of organs / organ assemblies) that are not exactly the same as the input dataset used for training, but are representative of the input training dataset, i.e., synthetic examples that are known to be realistic.

[0021] Each independent generator 200 includes a set of graph convolutional variational autoencoders (gcVAEs) 202A-202E. Each gcVAE 202 includes a pair of neural networks, referred to as an encoder (e.g., encoder 206A) and a decoder (e.g., decoder 216A), where the encoder includes multiple residual graph convolution downsampling (RGCDS) blocks (e.g., 210A-214A), and the decoder includes multiple residual graph convolution upsampling (RGCUS) blocks (e.g., 218A-222A). RGCDS and RGCUS blocks are discussed in more detail with respect to FIGS. 4 and 5 below. As shown in FIG. 2, each of the five sub-models (denoted by letters A-E) receives as input a mesh for each region (e.g., LV mesh 201A for sub-model gcVAE 202(A)), which is downsampled in encoder 206 to form latent vectors 215 for the fully connected layer. A different vector size can be applied to each of the five different sub-models. In the illustrated example, 16×1 is applied to the LV and LA, 12×1 to the RV and RA, and 8×1 to the AR. The size of the latent vector for each of the five different sub-models is determined empirically and may be different for the application and the dataset used for training. The latent vectors are then upsampled in the decoder 216 to form the output reconstructed mesh 223 (e.g., reconstructed LV mesh 223A).

[0022] More specifically, each gcVAE learns a separate and independent latent representation for each of the N corresponding anatomical regions and / or organs of interest through separate training. A variational autoencoder (graph convolutional variational field, or gcVAE, used in the independent generator 200 disclosed herein) may be a Bayesian latent variable model that combines 1) a recognition function represented by the encoder 206 network to infer a hidden low-dimensional representation of the input data (i.e., infer a latent representation) and 2) a generation function represented by the decoder 216 network to transform the inferred latent representation back to the original data space. The VAE infers a latent representation of the input data by approximating the posterior distribution of the latent variables given the observed input data. This is achieved by simultaneously optimizing the encoder 206 and decoder 216 networks, for example, to maximize the associated evidence lower bound (ELBO) of the observed input data. The VAE is flexible and can be used to learn latent representations of any type of data (e.g., images, audio signals, mesh-based representations of shapes). In the independent generator 200 of the exemplary generative shape modeling system 100, a VAE is combined with a graph convolutional neural network to learn a latent representation of anatomical shapes represented by a computational mesh (also called an unstructured graph).

[0023] The encoder-decoder neural network pairs (206A / 216A-206E / 216E and 306A / 316A-306E / 306E) in both the independent generator 200 and the dependent generator 300 are graph convolutional neural networks that utilize Chebyshev convolution operations to extract a hierarchy of shape features (i.e., across multiple scales) and learn representative latent representations that describe the shape variations of the corresponding region and / or organ of interest across the training population. Each level has its own set of upsampling and downsampling blocks (e.g., 210A / 218A), so while five levels are used in the illustrated example, other numbers of levels may alternatively be used. Also, while the same number of levels is used across all exemplary figures, different numbers may be used for different regions. The number of levels typically impacts processing speed and accuracy, so a value that provides the best compromise for a given use case may be selected for each implementation. The number of levels used in the independent and dependent generators developed was empirically determined to be optimal for the data used to train the developed system. Given different training data for the same or different applications, the number of levels selected may vary. The spatial convolution operation used in convolutional neural networks is well-defined for structured / gridded data in the Euclidean domain and cannot be directly applied to irregularly structured data such as graphs. The Chebyshev convolution operation is a generalization of spatial convolution, allowing for the application of spatial convolution to data defined by graphs (i.e., meshes). Furthermore, the Chebyshev convolution operation is performed in the Fourier domain. According to the convolution theorem, a convolution operation performed in the spatial domain is equivalent to element-wise multiplication in the Fourier domain. Chebyshev convolution is applied by transforming the graph-based data into the Fourier domain. The Fourier-transformed graph-based data is then element-wise multiplied by a Fourier-transformed convolution filter, after which an inverse Fourier transform is applied to transform the result of the multiplication operation back into the spatial domain.The parameters / weights of the convolution filters in a Chebyshev convolution are parameterized by truncated Chebyshev polynomials, and the coefficients of the polynomials represent the learnable weights / parameters of the convolution filters. The truncated Chebyshev polynomials are used in the disclosed generative shape modeling system because of their reduced computational complexity relative to other spatial convolutions and the use of general N-dimensional Chebyshev polynomials. However, in some examples, other spatial convolutions may be used instead of the truncated Chebyshev polynomial convolution. Alternatively, spatial convolutions designed for graph-based data, such as feature-operated graph convolutions, may be used instead of the spatial / Chebyshev convolution.

[0024] 2, the exemplary independent generator 200 comprises five site-specific gcVAEs corresponding to the five cardiac structures of interest. To train each of the N=5 gcVAEs, X n=1...N ={x i} i=1...D Consider an input data set denoted by, where each x i denotes the shape of an anatomical region (e.g., left ventricle) of an individual (i.e., patient) represented in the training dataset as a computational mesh (also known as a triangular surface mesh or unstructured graph). Each computational mesh is represented by a list of 3D points (i.e., spatial coordinates, which can be referred to as vertices or nodes) that define the boundary of the anatomical region / shape, and an adjacency matrix that defines the vertex / node connectivity. The input dataset X n The data of each individual in is denoted by the subscript i, where {i=1...D}, and the data of D individuals are stored in X n Included in each X n is the input data used to train each of the {n=1...N} site-specific gcVAEs, so in this example, N=5. n The D individuals whose data make up X do not have to be identical. n The training dataset sample x i} i=1...D may come from the same individual or from different individuals. In other words, the dataset X used to train each gcVAE n may contain non-overlapping, partially overlapping, or fully overlapping data, where the non-overlapping data is n This means that all samples in each X are obtained from different individuals. n This means that a certain proportion of the samples in each X are obtained from the same individual. n This means that all samples in the dataset are taken from the same individual. Each encoder network in each gcVAE receives a mesh-based representation of the shape of a given anatomical region or organ as input, i.e., x i and convert this high-dimensional representation into a low-dimensional vector or latent representation (z i The low-dimensional vector or latent representation resulting from each sample passing through the encoder 206 network is then inversely mapped or reconstructed by the decoder 216 network to approximate the original mesh of the input shape. n For each dataset X, n Each latent representation of is found separately (i.e., independently of the other regions).

[0025] Each gcVAE 202A-202E is used to calculate the corresponding input training dataset X n By approximating the intractable true posterior distribution of the latent variables with variational inference, we can obtain the observed data X n Each gcVAE is trained independently to learn a latent representation that is expected to generate a given set of observed data X n This is achieved by optimizing the evidence lower bound (ELBO) of the input data X n Given the posterior distribution, we approximate (q(z n |X n)), the decoder approximates the data likelihood given the latent variables (p(X n |z n ) during training. Thus, for a given input sample x i For each, the encoder converts the input into a unique latent representation z n i which is used as input by the decoder to reconstruct the data (r i In other words, during training, the encoder learns to reduce the input data into a lower-dimensional or compressed form, and the decoder learns to decompose or inverse transform / map this compressed information into an approximation of the data input to the encoder. Once the encoder and decoder are trained, in the generation phase (also called the inference or synthesis phase), a new compressed form of the data is generated and transformed by the trained decoder to produce synthetic data that looks realistic (i.e., has certain characteristics or properties that are shared and / or similar to the real data).

[0026] Each gcVAE 202A-202E is trained independently by jointly optimizing its constituent encoder 206 and decoder 216 networks to maximize the associated evidence lower bound (ELBO) of the observed input data. The ELBO is the sum of the expected negative log-likelihoods of the data (L recon ), a hypothetical prior distribution p(z n ) from the approximate posterior distribution q(z n |X n ) Kullback-Leibler divergence (L KL ), and an additional regularization loss (L regular The total loss function (L total ) is given by: L total =L recon +w0L KL +L regular (1) where L recon is the reconstruction loss used to represent the negative log-likelihood of the data, and i ) and the original input shape / mesh (x i ) is computed as the vertex-wise L1-norm evaluated between L recon is formulated as follows: L recon =||x i -r i ||1(2)

[0027] Kullback-Leibler divergence loss term L KL is a distance measure between probability distributions, where the approximate posterior distribution q(z n |X n ) and the true unknown (or target) posterior distribution p(z n |X n ) is used to minimize the divergence between the target posterior distribution p(z n |X n ) is difficult to handle (cannot be estimated) in the current setting, so by using Bayes' theorem, we can calculate the data likelihood p(X n |z n ) and the assumed prior distribution of the latent variable p(z). As a result, the Kullback-Leibler divergence loss term L KL is the approximate posterior distribution q(z n |X n ) and a hypothetical prior distribution p(z n ) is reduced to minimize the divergence between them. The prior distribution on the latent variables is assumed to be a centrally isotropic multivariate Gaussian distribution, i.e., p(z) = N(0,I), for all n = {1...N} site-specific GcVAEs. w0 is chosen to balance the influence on the training process relative to the other terms in the overall loss function, L KL The term denotes the weight by which the weight is multiplied. And the additional regularization loss term L regular penalizes outliers and reduces the decoder (r i ) (i.e., the reconstructed shape on the anatomical region input to the encoder) is smooth. regularis formulated as a weighted sum of Laplacian smoothness loss and edge length loss. laplacian The Laplacian loss, denoted by L, encourages adjacent vertices to move together, i.e., it penalizes adjacent vertices in the reconstructed shape / mesh from moving away from each other. Similarly, L edge The edge length loss, denoted by L, penalizes spurious movements of a vertex relative to its neighbors. regular is given by: L regular =w1L laplacian +w2L edge (3)

[0028] where w1 and w2 are the total loss function (L total ) to balance their relative influence on the regularization loss term and the overall training process, respectively. laplacian and L edge Indicates the weight by which the term is multiplied. L laplacian is given by:

[0029]

number

[0030]

number

[0031] The N=5 gcVAEs 202A to 202E in the independent generator 200 corresponding to the five cardiac structures of interest (LV, RV, LA, RA, AR) each have a total loss function L totalThe parameters of the encoder-decoder network pair in each gcVAE 202 are iteratively learned, for example, by the backpropagation algorithm.

[0032] Next, we outline an exemplary specific network architecture that may be used for the region-specific gcVAE in the independent generator 200, as shown in FIG. 2. The encoder-decoder network pair in each gcVAE 202A-202E includes five blocks, each of which (e.g., blocks 210A, 211A, 212A, 213A, and 214A for the downsampling block in encoder 206A) learns shape features at a different spatial resolution, enabling the encoder-decoder network to learn shape features in both global and local contexts across multiple mesh resolutions. This is made possible by including mesh downsampling (and upsampling for the decoder 216 network) operations in each of the five blocks of the encoder 206 and decoder 216 networks. Tables (1-5) below summarize the architectural details of each encoder-decoder network pair in each of the five gcVAEs that may be used to model the LV, RV, LA, RA, and AR. The architectures of the residual graph convolution downsampling and upsampling blocks used in the encoder and decoder networks in each of the five gcVAEs may be structurally similar, e.g., only the downsampling and upsampling factors used differ between different gcVAEs, as shown in Figures 4 and 5, respectively. According to one example, the number of feature channels (i.e., the number of different Chebyshev convolution filters / operations that can be trained) used within each graph convolution layer in the residual graph convolution downsampling blocks of blocks 1 through 5 is 16, 32, 32, 64, and 64, and the number of feature channels used within each graph convolution layer in the residual graph convolution upsampling blocks of blocks 1 through 5 is 64, 64, 32, 32, and 16. Other different numbers of feature channels may also be used. The mesh downsampling factors used across downsampling blocks 1 through 5 vary in the encoder network in each site-specific gcVAE, as detailed below.The left ventricle is [4, 4, 4, 6, 6], the right ventricle is [4, 4, 4, 6, 6], the left atrium is [4, 4, 4, 5, 5], the right atrium is [4, 4, 4, 5, 5], and the aorta is [4, 4, 4, 4, 4]. Here, each number in parentheses corresponds to the downsampling factor used in the residual graph convolution downsampling blocks of blocks 1 through 5 for each region-specific gcVAE. Similarly, the upsampling factors used across upsampling blocks 1 through 5 differ in the decoder network in each region-specific gcVAE and are detailed as follows: the left ventricle is [6, 6, 4, 4, 4], the right ventricle is [6, 6, 4, 4, 4], the left atrium is [5, 5, 4, 4, 4], the right atrium is [5, 5, 4, 4, 4], and the aorta is [4, 4, 4, 4, 4]. Here, the numbers in parentheses correspond to the upsampling factors used in the residual graph convolution upsampling blocks of blocks 1 through 5 for each region-specific gcVAE. However, the resampling factors are the same at each corresponding resolution level between the encoder and decoder. That is, for the decoder, the order is reversed as we move from low to high resolution (while the encoder moves from high to low resolution (i.e., downsamples)). The number of feature channels and mesh resampling factors used in the disclosed region-specific gcVAE were empirically determined to be optimal for the example application (i.e., cardiac modeling) and the training data used. The respective values ​​may vary given different training data for the same or different applications in generative modeling of multi-region or multi-organ shape assemblies. Furthermore, all residual graph convolution downsampling and upsampling blocks used in all networks in the disclosed system include instance normalization layers, residual connections, and exponential linear activation units, organized as shown in Figures 4 and 5.

[0033] [Table 1]

[0034] [Table 2]

[0035] [Table 3]

[0036] [Table 4]

[0037] [Table 5]

[0038] As clearly shown in FIG. 2, the independent generators 200 are called independent to maintain separation between the processing of each gcVAE (202A-202E), and each gcVAE is trained independently with its own training loop.

[0039] FIG. 3 illustrates an exemplary subordinate generator 300 for use in the system of FIG. 1 according to one example of the present disclosure. This exemplary subordinate generator 300 is a single generative model formulated as a graph convolutional multi-channel variational autoencoder (gcmcVAE) that includes multiple encoder-decoder pairs, one for each region to be modeled (i.e., considered as a single gcmcVAE with multiple branches, as opposed to the multiple independent gcmcVAEs in the independent generator of FIG. 2 ), resulting in N=5 pairs of encoder-decoder networks corresponding to the five cardiac structures of interest (e.g., LA, RA, LV, RV, AR). The encoder 306 includes multiple residual graph convolutional downsampling (RGCDS) blocks (e.g., 310A-314A), and the decoder includes multiple residual graph convolutional upsampling (RGCUS) blocks (e.g., 318A-322A). Again, the RGCDS and RGCUS blocks are the same as those discussed in more detail with respect to FIGS. 4 and 5 below. As shown in FIG. 3, each of the five encoder-decoder pairs in the gcmcVAE (indicated by letters A through E) receives a mesh (e.g., LV mesh 301A) for each region as input, which is downsampled in encoder 306 to form latent vector 315A for the fully connected layer. In the example of FIG. 3, the same vector size (=12×1) is applied to each of the five different sub-models, although other sizes may alternatively be used. However, unlike the independent generator 200, the latent vectors 315A through 315E corresponding to each of the five different sub-models in the dependent generator 300 are realized with the same fixed size (e.g., =12×1). This is because in the dependent generator 300, the latent vectors are not independent as in the gcVAEs (202A through 202E) of the independent generator 200, but are shared, or "dependent," with each other. In other words, each encoder (306A to 306E) in the gcmcVAE of the dependent generator 300 maps the input mesh of each part to a shared latent vector (315A to 315E).The latent vectors 315A-315E are considered shared because, by design, the output of each encoder 306A-306E is formulated as a different approximation of the same shared latent vector (or as an approximation of the same shared hidden representation across all input channels of information). This formulation of the output of each encoder 306A-306E as an approximation of the same shared latent vector is achieved by imposing constraints on the loss function (described in the following paragraph) used to train the gcmcVAE. Each encoder's approximation of the shared latent vector is passed as input to all decoders 316A-316E (e.g., via multiplexing layer 317), which reconstructs all channels of information or meshes for all segments input to the encoders 306A-306E. In other words, the encoders (306A-306E) in the gcmcVAE are channel- or region-specific and output different approximations of the shared latent vectors (315A-315E). Meanwhile, the decoders (316A-316E) operate across channels or regions, simultaneously outputting / reconstructing information for all regions or all channels of the mesh. Because the outputs of all encoders (306A-306E) are used as inputs to all decoders (316A-316E), the input size of each decoder (316A-316E) must be the same. Therefore, in the case of the dependent generator 300, the latent vectors (315A-315E) output by each encoder (306A-306E) have the same size.

[0040] More specifically, the input to each encoder-decoder network pair in the gcmcVAE is a mesh / unstructured graph 301A-301E representing the shape of the associated anatomical part or organ, which together constitute the multi-part / multi-organ shape assembly synthesized by the disclosed generative shape modeling system.

[0041] All encoder-decoder network pairs 302A-302E in the gcmcVAE are trained simultaneously to learn a shared / common latent representation across all anatomical regions and / or organs of interest in the input data. The mcVAE is an extension of the VAE that simultaneously analyzes heterogeneous data by projecting observations from different input sources / channels to a shared / common latent representation. The training procedure by which the mcVAE learns a shared latent representation across multiple data input sources / channels is similar to that of the VAE in that it is achieved by approximating the posterior distribution of the shared latent representation given the observed heterogeneous input data. This is achieved by simultaneously optimizing all N constituent encoder-decoder network pairs 302A-302E (representing the N data input sources / channels or the N anatomical region and / or organ shapes in the case of the gcmcVAE in the present invention) to maximize the associated evidence lower bound (ELBO) of the observed heterogeneous input data.

[0042] Training the gcmcVAE requires X n=1...N ={x n i} i=1...D Consider a data set denoted by, where each x n i denotes the shape of an individual's anatomical region (e.g., the left ventricle) represented as a computational mesh. Each mesh / unstructured graph is represented by a list of 3D points / spatial coordinates (called vertices) that define the boundary of the anatomical region / shape, as well as an adjacency matrix that defines the vertex connectivity. n Individual patient data in is denoted by the subscript i, where {i=1...D} and D is the n is the number of personal data included in each X n is the data input to each of the {n=1...N} pairs of encoder-decoder networks in the gcmcVAE, where N=5 in this example. n The D individuals whose data make up X do not have to be identical. n The training data samples in Xn i} i=1...D may come from the same individual or from different individuals. In other words, the dataset X used to train each gcmcVAE n may contain partial or complete overlapping data, where the partial overlapping data is n This means that a certain proportion of the samples in each X are obtained from the same individual. n This means that all samples in the gcmcVAE are obtained from the same individual. The dependent generator 300, unlike the independent generator 200, cannot be trained on non-redundant data. By design, the gcmcVAE learns a shared / common latent representation across all input channels of information / anatomical region by exploiting the correlations that exist in the provided inputs. Thus, the dependent generators use inputs obtained from different patients / patient populations, making the information redundant (i.e., at least partially overlapping). Each encoder network in the gcmcVAE receives a mesh / unstructured graph-based representation of a given anatomical region or organ shape as input, i.e., x n i (see Figure 3), and all encoder networks in the gcmcVAE share a high-dimensional representation of these shapes as a common low-dimensional vector or latent representation (z i In other words, each encoder network simultaneously maps to its corresponding anatomical region or organ x n i is given as input, and different approximations are used to share the latent vector z iThe shared / common latent representation approximations resulting from passing each anatomical region or organ shape through its corresponding encoder network are then provided as input to all decoder networks in the gcmcVAE. Each decoder network then uses the latent vector approximations output by each encoder network as input to inversely map or reconstruct the original mesh of the anatomical region or organ input to its corresponding encoder. In other words, given that there are N regions in the shape assembly of each patient's data used to train the gcmcVAE, each of the N encoders in the gcmcVAE outputs N approximations of the latent representation shared by all N regions. Each of the N decoders in the gcmcVAE then takes as input all N approximations of the shared latent representation and outputs / reconstructs N versions / approximations of the shape of its corresponding region (i.e., the region input to the encoder corresponding to each of the N decoders). During training, the N different reconstruction results output by the corresponding decoders for each region result in N different losses, which are aggregated / summed together and used to guide the training of the gcmcVAE. n=1...N By training the gcmcVAE on the dataset X n=1...N A shared / common latent representation for is found.

[0043] Each of the N=5 encoder-decoder network pairs in gcmcVAE is a training dataset X n=1...N By using variational inference to approximate the intractable true posterior distribution of the latent variables, we can estimate the observed dataset X n=1...N are simultaneously trained to learn a shared / common latent representation that is assumed to generate a common latent representation for the observed dataset X n=1...NThis is achieved by optimizing the evidence lower bound (ELBO) of the posterior distribution (q n (z|X n ) or, in other words, the input data set X n Given a shared latent vector z, each decoder network outputs an approximation of the shared latent representation. Each corresponding decoder network calculates the data likelihood (p n (X n Since each of the N=5 input channels (i.e., anatomical regions / organ shapes) in the gcmcVAE is assumed to be conditionally independent from all others given the shared / common latent representation z (i.e., the input channels are considered independent from each other but depend on the shared latent variable z), the joint data likelihood across all input channels is formulated as p(X|z), given by:

[0044]

number

[0045] To train the gcmcVAE, we use input samples s i Given (where s i ={x i} n=1...N represents a set of shapes (i.e., where the shapes belong to one of the five cardiac structures of interest), s i Individual shapes in x n i is passed as input to its corresponding encoder network. i All x's in n i come from the same individual, denoted by the subscript “i”. Each sample s used to train the gcmcVAE iIt is not necessary for all five cardiac structures to be present in the gcmcVAE. That is, the gcmcVAE can be trained with partially overlapping data. All five encoder networks in the gcmcVAE are trained with their corresponding inputs x n i (x n i ga s i (if present) is a unique latent representation z i This is used as input by all decoder networks to obtain the shape of each corresponding anatomical part or organ (r n i(denoted by ). Thus, the gcmcVAE comprises a shape-specific encoder network that maps each shape to a shared / common latent representation, and a cross-shape decoder network that can simultaneously reconstruct all shapes given the shared / common latent representation as input. This allows the subordinate generator 300, once trained, to "complete / fill" missing regions / channels of information conditioned on the given input. Specifically, this completion or conditional generation of missing regions / channels of information is achieved by training each encoder in the gcmcVAE to reconstruct the corresponding region, given an uncompleted input region / channel of information, using an approximation to the shared latent space output by all encoders in the gcmcVAE (i.e., by training a decoder for cross-shape / cross-channel synthesis). Additionally, the design constraints of the gcmcVAE that facilitate shape-to-shape training and subsequent shape-to-shape imputation for missing regions / channels of information are that the size of the latent vectors output by all encoder networks is fixed, and correspondingly, the size of the input vectors expected by all decoder networks is fixed, leading to the formulation of the loss function used to train the gcmcVAE (described in the next paragraph). Because the learned latent representations are shared across all input regions / channels, after training the gcmcVAE, multiple regions / channels can be reconstructed using the latent representation obtained from a single input region / channel of information. Furthermore, by sampling the shared / common latent representation from a hypothetical prior distribution on the latent variables (i.e., p(z) = N(0,I)), novel / virtual shapes for regions (not observed in the training data) can be synthesized by passing the sampled latent representations through all decoder networks simultaneously.

[0046] The gcmcVAE is a network model that simultaneously optimizes the pair of encoder-decoder networks that compose it to generate the observed input dataset {X n} n=1...5The training is done by maximizing the weighted evidence balance (ELBO). The ELBO is the sum of the expected joint negative log-likelihoods of the data (L recon ), and the approximate posterior distribution q(z|X n ) Kullback-Leibler divergence (L KL ), and an additional regularization loss (L regular The gcVAE (202A-202E) in the independent generator 200 and the gcmcVAE (302A-302E) in the dependent generator 300 are primarily differentiated by the ELBO formulation used to train the encoder and decoder networks that comprise them. The gcVAE (202A-202E) in the independent generator 200 assume that each anatomical region or organ shape of interest is completely independent, and therefore each gcVAE learns an independent / unique latent representation for its corresponding region. On the other hand, the gcmcVAE (302A-302E) assume that all regions or organ shapes in the shape assembly are conditionally independent, i.e., each region depends on a shared latent representation, and therefore learns a shared latent representation across all regions. In this way, the formulation of the ELBO loss function used to train the gcmcVAEs (302A to 302E) in the dependent generator 300 is based on the shared latent variables q n (z|X n ) and the unknown target posterior distribution p(z|X 1 ,X 2 ,...,X n ) is derived by imposing a constraint to minimize the Kullback-Leibler divergence between the gcVAEs. Conversely, each gcVAE (202A-202E) in the independent generator 200 is trained by maximizing a separate ELBO loss function, independent of the other gcVAEs. Here, each ELBO loss function is derived by maximizing the latent variables q(z n |X n ) and the unknown target posterior distribution p(z n|X n ) is formulated to be minimized independently of other approximate posterior distributions and latent variables output by the encoders in other gcVAEs. As in the case of the gcVAE in the independent generator 200, for training the gcmcVAE in the dependent generator 300, each approximate posterior distribution q on the shared latent variables is n (z|X n ) and the true unknown (i.e., target) posterior distribution p(z|X 1 ,X 2 ,...,X n ) is intractable (cannot be computed / estimated). This is because the target posterior distribution p(z|X 1 ,X 2 ,...,X n ) is unknown and itself intractable in the current setting. Instead, the ELBO loss function used to train the gcmcVAE (302A-302E) is formulated by using Bayes' theorem to factorize the target posterior distribution as the product of the joint data likelihood p(X|z) across all sites (see Equation 6) and a hypothetical prior distribution on the shared latent variable p(z). As a result, the Kullback-Leibler divergence loss term L KL is the approximation q of the posterior distribution on the shared latent variables output by each encoder (306A to 306E) network in the gcmcVAE. n (z|X n ) and a tentative prior distribution on the shared latent variable p(z). Or, in other words, to learn a shared latent representation across all input parts, the gcmcVAE is trained by minimizing an approximation of the posterior distribution on the shared latent variable output by each encoder (306A-306E) network in the gcmcVAE from the target posterior distribution, under the underlying assumption that each input part carries useful information about the shared latent representation of itself and all other parts in the shape assembly. The ELBO total loss function (L total) is the L used to train each gcVAE in the independent generator 200. total It is formulated in the same way as L total The individual loss terms in are formulated differently for the dependent generator 300. The reconstruction loss L (which represents the joint negative log-likelihood of the data) recon is the reconstruction (r i n|n’ ) shape / mesh and the corresponding original shape / mesh (x n i ) where each of the N encoders (306A-306E) outputs an approximation of the shared latent vector (z), and each decoder (316A-316E) receives as input the N approximations of z and is trained to reconstruct the corresponding region / mesh (i.e., the region / mesh input to its corresponding encoder) using each of the N approximations of z output by each of the N encoder networks (306A-306E). Thus, r i n|n’ Given an approximation of the sample shared latent vector n' output by each of the n = {1...N} encoders (306A-306E) in the gcmcVAE, L denotes the region / mesh n' that one of the n = {1...N} decoders (316A-316E) reconstructs for sample i'. Since there are N encoders, n' = {1...N} different approximations of the shared latent space z are passed as input to each decoder (316A-316E), resulting in N different reconstructions of the nth region for the ith sample. As a result, N different reconstruction loss terms are evaluated for each region in the shape assembly, and by joint aggregation / summation, L recon is formulated as follows:

[0047]

number

[0048] Kullback-Leibler divergence loss term L KL is the posterior distribution on the shared / common latent variables that each encoder network approximates, i.e., q n (z|X n ) and a tentative prior distribution on the shared latent variables. The prior distribution on the latent variables is assumed to be a central isotropic multivariate Gaussian distribution, i.e., p(z)=N(0,I). While an isotropic multivariate Gaussian prior is assumed to simplify the computation, other types of prior distributions (e.g., Dirichlet distributions, Gaussian mixture models, Gaussian processes, Dirichlet processes, among others) may also be used in the disclosed system. As mentioned above, the gcmcVAE is based on the reconstruction loss term L as shown in Equation 7. recon and each approximate posterior distribution q on the shared latent variable corresponding to each input site / channel of information from the tentative prior distribution p(z) and its specific encoder (306A-306E). n (z|X n ) is trained to learn a shared latent representation across all parts / input channels of information by minimizing the divergence of L recon Training gcmcVAE using this formulation allows for shape-to-shape / channel-to-channel imputation of missing information (i.e., given partially overlapping data as input, the trained gcmcVAE can be used to imputation / conditional synthesis of missing information / channels). In other words, this L recon Training the gcmcVAE using the formulation allows for the reconstruction of multi-channel / multi-site data given a single channel / site input.

[0049] L of the dependent generator total The regularization loss term L regular is formulated as follows:

[0050]

number

[0051] The gcmcVAE is a network model that uses the total loss function L via the parameters of the encoder-decoder network pair that composes it. total The parameters of all encoder-decoder network pairs in the gcmcVAE are iteratively learned by the backpropagation algorithm.

[0052] We now outline an exemplary specific network architecture that may be used for the gcmcVAEs 306A-306E in the subordinate generator 300, as shown in FIG. 3 . In the current implementation of the gcmcVAE, five pairs of encoder-decoder networks are used, corresponding to the five cardiac structures of interest; however, the number of encoder-decoder network pairs may vary to suit different applications (i.e., when the number of structures comprising the multi-part shape assembly of interest is less than or greater than five). The encoders 306 and decoders 316 in all five pairs (A-E) each contain five blocks, and each block (e.g., blocks 310A, 311A, 312A, 313A, and 314A for downsampling blocks in encoder 306A) learns shape features at different spatial resolutions, enabling the encoder-decoder networks to learn shape features in both global and local contexts across multiple mesh resolutions. This is made possible by including mesh downsampling and upsampling operations in each of the five blocks of the encoder 306 and decoder 316 networks. Table 6 below summarizes the architectural details of each encoder-decoder network pair 302A-302E in the gcmcVAE used to model the five cardiac structures of interest, namely, the LV, RV, LA, RA, and AR. In all five pairs, the architectures of the residual graph convolution downsampling blocks (310-314) and residual graph convolution upsampling blocks (318-322) used in the encoder and decoder networks, respectively, are similar (i.e., only the downsampling and upsampling factors used differ between different encoder-decoder network pairs), as shown in Figure 4. The number of feature channels used in each graph convolution layer in the residual graph convolution downsampling blocks (Blocks 1-5) is 16, 32, 32, 64, and 64, respectively. Similarly, the number of feature channels used in each graph convolution layer in the residual graph convolution upsampling blocks (Blocks 1-5) is 64, 64, 32, 32, and 16, respectively.The mesh downsampling factors used across downsampling blocks 1 to 5 were different for each encoder network in each encoder-decoder network pair in the gcmcVAE. These are described as follows, according to the cardiac structure each encoder network accepts as input: [4, 4, 4, 6, 6] for the left ventricle, [4, 4, 4, 6, 6] for the right ventricle, [4, 4, 4, 5, 5] for the left atrium, [4, 4, 4, 5, 5] for the right atrium, and [4, 4, 4, 4, 4] for the aorta. Here, the numbers in parentheses correspond to the downsampling factors used in the residual graph convolution downsampling blocks in blocks 1 to 5 for each encoder-decoder network pair. Similarly, the mesh upsampling factors used across upsampling blocks 1 to 5 were different for each decoder network in each encoder-decoder network pair in the gcmcVAE. These are described as follows, according to the cardiac structure each decoder network reconstructs / outputs: That is, the left ventricle is [6, 6, 4, 4, 4], the right ventricle is [6, 6, 4, 4, 4], the left atrium is [5, 5, 4, 4, 4], the right atrium is [5, 5, 4, 4, 4], and the aorta is [4, 4, 4, 4, 4]. Here, the numbers in parentheses correspond to the upsampling factors used in the residual graph convolution upsampling blocks of blocks 1 through 5 in each encoder-decoder network pair.

[0053] [Table 6]

[0054] The dependent generator 300 is called dependent because a shared latent representation is assumed across all parts / input channels of information, and the output of each decoder (316A-316E) in the gcmcVAE depends on the shared latent representation across all parts. Therefore, the dependent generator 300 is trained in a single training loop by simultaneously optimizing all constituent encoder (306A-306E) and decoder (316E-316E) networks using a single loss function. This contrasts with the independent generator 200, in which each gcVAE (202A-202E) is trained independently for each part, and the latent representations learned for each part are independent of those learned for the other parts. Furthermore, the independent generator 200 is trained in separate / independent training loops for each constituent gcVAE (202A-202E) using separate / independent loss functions.

[0055] In the above description, we have mentioned the residual graph convolution downsampling (RGCDS) block and the residual graph convolution upsampling (RGCUS) block, which will be described below with reference to Figures 4 and 5. In our example, we use these residual graph convolution downsampling / upsampling blocks to enhance the efficiency of feature exchange between vertices (i.e., nodes in a mesh) and mitigate the problem of vanishing gradients when training a network.

[0056] FIG. 4 illustrates an exemplary architecture of a residual graph convolution downsampling block used in all graph neural networks of the system of FIG. 1 according to one example of the present disclosure. The RGCDS block takes a set of meshes (e.g., the output from a preceding RGCDS in the encoder network used in the disclosed independent and dependent generators) as input 401, applies a Chebyshev graph convolution 410, applies an instance normalization process 420, and then applies an exponential linear unit function 430. However, in some examples, the input 401 may be a single mesh. The result may then undergo another Chebyshev graph convolution 440 and an instance normalization process 450 before being summed together with separate paths 415 (i.e., residual branches) from the original Chebyshev graph convolution 410. The summed result may then undergo another exponential linear unit function 470 before being applied with a final mesh pooling layer function to output the result 491. To this end, each residual block includes at least one graph convolutional layer that extracts features using the input provided to that block, and a residual connection that combines the outputs of the first and second graph convolutional layers by addition. Chebyshev graph convolutional operations are used throughout this example because they are strictly localized filters that, when combined with mesh pooling operations, enable learning of multi-scale hierarchical patterns and reduce computational complexity. Each graph convolutional layer is followed by an instance normalization layer and an exponential linear unit (ELU) for activation. The instance normalization layer is also used in the residual branch to ensure similarity of feature statistics to the output of the second graph convolutional layer in the block.

[0057] FIG. 5 illustrates an exemplary architecture of a residual graph convolution upsampling block for use in all graph neural networks of the system of FIG. 1 according to one example of the present disclosure.

[0058] The RGCUS block is similar to the RGCDS block in FIG. 4 , but in reverse order. The RGCUS block is used in the decoder network within the disclosed part-aware generative model 102. As shown, the RGCUS may take as input 501 the latent vector output from the encoder network or the output from a previous RGCUS block in the same decoder network, apply a mesh upsampling function 510, a Chebyshev graph convolution 520, and a first instance normalization process 530, then apply an exponential linear unit function 540, followed by another Chebyshev graph convolution 550 and an instance normalization process 560. The result is then jointly summed with a separate path 515 (i.e., the residual branch) from the original Chebyshev graph convolution 520. Finally, the summed result may be applied with another exponential linear unit function 570 before outputting the result 591. To this end, each residual block first processes the given input by upsampling it using a mesh upsampling layer, then includes two graph convolutional layers that extract the feature output of the mesh upsampling layer, and a residual connection that combines the outputs of the first and second graph convolutional layers by summation. Chebyshev graph convolutional operations are used throughout this example because they are strictly localized filters that, when combined with mesh pooling operations, enable learning of multiscale hierarchical patterns and reduce computational complexity. Each graph convolutional layer is followed by an instance normalization layer and an exponential linear unit (ELU) for activation. The instance normalization layer is also used in the residual branch to ensure similarity of feature statistics to the output of the second graph convolutional layer in the block.

[0059] Once the part-aware generation module 110 is trained using the independent generator 200 or the dependent generator 300, all anatomical parts and / or organs in the shape assembly can be synthesized. However, there are two ways to train the part-aware generation model 102: using the independent generator 200 (e.g., of FIG. 2) or the dependent generator 300 (e.g., of FIG. 3). In the case of the independent generator 200, non-overlapping data, partial overlapping data, or full overlapping data can be used. In the case of the dependent generator 300, only partial overlapping data or full overlapping data can be used. This synthesis can be achieved by sampling new latent vectors from the learned latent distribution. That is, in the case of the independent generator 200, a new latent vector is sampled from each gcVAE to synthesize each corresponding anatomical part / organ. On the other hand, in the case of the dependent generator 300, the new latent vector may be sampled from a shared / common latent distribution learned by the gcmcVAE across all anatomical parts / organs of interest. The sampled latent vectors may be passed as inputs to corresponding decoder networks for the reconstruction / synthesis of new anatomical regions / organ shapes. That is, the sampled latent vectors from each gcVAE in the independent generator 200 may be passed to their corresponding decoder networks independently of each other. Alternatively, all sampled latent vectors from the gcmcVAEs may be passed to all constituent decoder networks simultaneously (i.e., in the dependent generator, all decoder networks may receive the same sampled latent vector as input). Using the independent generator 200 or the dependent generator 300, anatomical regions / organs may be synthesized separately through their corresponding decoder networks. As a result, the synthesized regions may not initially be spatially organized / aligned into a shape assembly that is realistic and representative of the original anatomy (e.g., the synthesized regions may significantly intersect / overlap each other, which does not represent naturally occurring variations in anatomy).1 may then apply a spatial construction network to learn the spatial relationships between the regions in the shape assembly and estimate the spatial transformations required to combine the combined individual regions into an anatomically meaningful (i.e., realistic) shape assembly (such as the whole heart shape assembly of interest in this example). One example of a spatial construction network of the present disclosure takes as input meshes of individual anatomical regions / organs combined using independent or dependent generators and spatially organizes the input regions into a multi-region / multi-organ shape assembly (i.e., a virtual chimera) that is representative of the native anatomy by estimating both affine and non-rigid (also referred to as deformable) transformations. According to this example, the input to the spatial construction network includes mesh / graph-based representations of the shapes of the left ventricle, right ventricle, left atrium, right atrium, and aortic root, and the spatially constructed output is a virtual chimera of the heart (also referred to as the whole heart shape assembly).

[0060] As shown in FIG. 1 , the disclosed spatial construction module 120 applies two networks: 1) an affine construction network (see FIG. 6 ) and 2) a non-rigid construction network (see FIG. 7 ). The affine construction network first recovers the affine transformation to coarsely construct the shapes of the input anatomical parts / organs. This coarsely constructed part is used as input to a non-rigid construction network that estimates local non-rigid / deformable transformations, outputting a finely constructed multi-part / multi-organ shape assembly, i.e., a virtual chimera. The affine construction network includes N branches / sub-networks, each of which takes one of the N anatomical part / organ shapes as input (N is determined by the structures of interest in the multi-part / multi-organ shape assembly to be synthesized). In this example, the affine construction network includes N=5 part-specific sub-networks corresponding to the five cardiac structures of interest. Meanwhile, the non-rigid construction network includes a single network that takes as input the coarsely constructed multi-part / multi-organ shape assembly output by the affine construction network. In this example, the affine and non-rigid construction networks are trained separately and sequentially (i.e., sequentially, with the affine network trained first, followed by the non-rigid network), although they could also be trained simultaneously. Both the affine and non-rigid construction networks could be trained with real or synthetic data, or a combination of the two. That is, they could be trained with input shapes derived from actual patient data or shapes synthesized by a generative network such as the disclosed independent generator 200 or dependent generator 300 (or any combination of real and synthetic data). Thus, the overall spatial construction network applied by the spatial construction module 120 is not limited to the use of real or synthetic shapes for training, nor is it limited to the use of synthetic data generated by the disclosed independent generator 200 or dependent generator 300.In this example, a combination of real and synthetic data was used to train both the affine and non-rigid construction networks of spatial construction module 120.

[0061] {x} n=1...N Given the meshes of all N regions in the shape assembly of interest, denoted by {T} (where the meshes of the N regions may be derived from actual patient data, synthesized using a generative shape model (e.g., independent generator 200 or dependent generator 300), or a combination of real and synthetic data), each of the N regions is passed as input to its own branch / sub-network in the affine construction network. In this example, the shapes of the left ventricle, right ventricle, left atrium, right atrium, and aortic root are passed as input to each of its own branch / sub-networks (A-E). Each sub-network in the affine construction network extracts information about the position, orientation, and shape of its corresponding input mesh and outputs a 3D affine transformation for each region in the shape assembly. As a result, {T} n=1...N In this example, the affine construction network estimates five affine transformations corresponding to the five input cardiac structures of interest (LV, RV, LA, RA, AR). Each of the five (N=5) estimated affine transformations is a translation (t=[t x ,t y ,t z ], where each component of the translation vector t represents the estimated displacement along a particular direction / axis in 3D Euclidean space), rotation (denoted by R=[Q1,Q2,Q3,Q4], where components Q1-Q4 represent the quaternion used to parameterize the rotation in 3D), and scaling (denoted by S, which is a scalar value that controls the global scaling estimated per region). The affine configuration network is based on a loss function L given by: affine The method is trained in a "self-supervised manner" to estimate the affine transformation of each of the five cardiac structures of interest by minimizing

[0062]

number

[0063] where T n is estimated at site x n where the subscripts {n=1...N} represent each region of interest in the shape assembly (in this example, N=5, corresponding to the five cardiac structures of interest). n is the nth site in the assembly (i.e., x n ) mesh, shared with all other (N-1) parts in the assembly. o transf represents the nodes / vertices in the meshes of all other (N-1) parts in the shape assembly that are shared with the mesh of the nth part. o transf is calculated by applying the affine transformation estimated for each of the (N-1) sites (site x n It is constructed by connecting the shared vertices from the meshes of all (N-1) parts in the assembly (given

number

[0064] FIG. 6 illustrates an exemplary architecture of an affine construction network according to the present disclosure. Briefly, an input mesh 601 (e.g., LV 601A) is input to each encoder sub-network 606, which operates to downsample the data and output a fully connected layer latent vector 615. The output latent vectors 615 are then concatenated as a shared latent vector 616, which continues to two fully connected layers 1 617 and 2 618. Fully connected layers, also known as dense layers, and multilayer perceptrons are individual stacked neurons / neural units in which each neuron in a fully connected layer is connected to all features of the vector input to that layer. The second fully connected layer 618 provides input to each of three sub-branches that process one of three transformations: scaling, rotation, and translation. These are each processed by two fully connected layers (619 / 620, 622 / 623, 625 / 626), which each output a set of parameters as a vector (scaling output vector 621, rotation output vector 624, translation output vector 627). Here, the fully connected layers, which output the translation, rotation, and scaling parameters for all five cardiac structures of interest, act as a regression model / network. Details of the five branches / sub-networks available in the example affine construction network are shown in Table 7. Each branch / sub-network is composed of five residual graph convolution downsampling blocks (as shown in Figure 4). The number of feature channels used in each graph convolution layer in the residual graph convolution downsampling blocks (blocks 1-5) is 16, 32, 32, 64, and 64. The mesh downsampling factors used across downsampling blocks 1-5 were different for each branch / sub-network in the affine construction network. The downsampling factors used for each site-specific branch / sub-network according to an example are as follows:The left ventricular branch / sub-network is [4, 4, 4, 6, 6], the right ventricular branch / sub-network is [4, 4, 4, 6, 6], the left atrial branch / sub-network is [4, 4, 4, 5, 5], the right atrial branch / sub-network is [4, 4, 4, 5, 5], and the aortic root branch / sub-network is [4, 4, 4, 4, 4]. Here, each number in parentheses corresponds to the downsampling factor used in the residual graph convolution downsampling block in blocks 1 through 5. The output of the final residual graph convolution downsampling block in each branch / sub-network is flattened into a vector. The flattened vectors from each branch / sub-network are then concatenated into a single long vector that serves as input to a series of two fully connected layers. These two fully connected layers extract features from the provided input and provide the output in the form of a shorter vector (in this example, 1 / 4) to three regression branches, each containing two fully connected layers. These three regression branches then estimate the translation, rotation, and scaling parameters (representing the desired 3D affine transformation) of all sites in the assembly.

[0065] [Table 7]

[0066] Following the affine spatial construction of real and / or synthetic parts by the affine construction network in Figure 6, a coarse construction / aligned multi-part shape assembly (X = {x n} n=(1...N)The non-rigid construction network may be configured to refine the N-part shape assembly (denoted by gcVAE 202A-202E) by estimating the deformations of each of the N parts in the multi-part shape assembly (i.e., N=5 in this example, corresponding to the five cardiac structures of interest). The non-rigid construction network improves the continuity or coherence at the boundaries between adjacent parts in the multi-part shape assembly by estimating the deformations for each local part and reducing the intersections / gaps between adjacent parts remaining after the initial affine construction step. The non-rigid construction network is a graph convolutional neural network with an encoder-decoder network pair (shown in FIG. 7), similar to the gcVAEs 202A-202E in the independent generator 200 of FIG. 2. The non-rigid construction network 700 extracts features from the coarsely constructed / registered multi-part shape assembly (i.e., the entire heart shape assembly in this example) through a series of graph convolution blocks (710-714) and estimates the displacements for local nodes / vertices in the mesh 701 of each part in the assembly. The non-rigid construction network 700 is computed using a loss function L given by: non-rigid It is trained in a self-supervised manner (similar to an affine registration network) by minimizing

[0067]

number

[0068] where T nr is a non-rigid transformation, i.e., the transformation of all parts {x n} n=(1...N) represents the node / vertex displacements estimated for all vertices in the mesh of T n nr is the mesh of the nth part in the assembly (v n represents the estimated displacements for nodes / vertices shared between the mesh of the (N-1) other parts in the assembly (denoted by v). onr represents the deformed / transformed nodes / vertices in the meshes of all other (N-1) parts in the geometry assembly that are shared with the mesh of the nth part. o nr is the displacement of the nodes / vertices estimated for each of the (N-1) sites (site x n It is constructed by connecting the shared vertices from the meshes of all (N-1) parts in the assembly (given

number

number

number

[0069] Table 8 details the encoder-decoder network pair that can be used in the nonrigid construction network. The encoder 706 and decoder 716 each contain five blocks, each of which learns shape features at different spatial resolutions, enabling the entire nonrigid construction network 700 to learn shape features in both global and local contexts across multiple mesh resolutions. This is made possible by including mesh downsampling and upsampling operations in each of the five blocks (710-714) of the encoder 706 network and the five blocks (718-722) of the decoder 716 network. The encoder 706 takes as input a coarsely constructed / registered multi-part shape assembly, in this case a whole-heart shape assembly including meshes for the left ventricle (LV), right ventricle (RV), left atrium (LA), right atrium (RA), and aortic root (AR). The features extracted through the residual graph convolution downsampling blocks (710-714) in the encoder network are output to a feature embedding array 715, and the residual graph convolution upsampling blocks (718-722) in the decoder network are used to decompose / expand the feature embedding array 715. The output of the final residual graph convolution upsampling block is condensed to the size of the input array / coarsely aligned whole heart mesh using a Chebyshev convolution layer 723. The Chebyshev convolution layer 723 also estimates the displacements of the nodes / vertices of all parts in the assembly, which becomes the decoder output 724. The residual graph convolution downsampling and upsampling blocks used in the encoder and decoder are shown in Figure 4. The number of feature channels used in each graph convolution layer in the residual graph convolution downsampling blocks (blocks 1-5) is 16, 32, 32, 64, and 64, respectively. Similarly, the number of feature channels used in each graph convolution layer in the residual graph convolution upsampling blocks (blocks 1 to 5) is 64, 64, 32, 32, and 16, respectively.Skip connections 717 are included between corresponding blocks in the encoder 706 and decoder 716 in the network to ensure better propagation of features and gradients during training. The mesh downsampling factor used across downsampling blocks 1-5 is [4, 4, 4, 6, 6], where each number in parentheses corresponds to the downsampling factor used in the residual graph convolution downsampling blocks of blocks 1-5. Similarly, the upsampling factor used across upsampling blocks 1-5 is [6, 6, 4, 4, 4], where each number in parentheses corresponds to the upsampling factor used in the residual graph convolution upsampling blocks of blocks 1-5. The output of the final residual graph convolution upsampling block in the decoder is passed to a single convolutional layer that estimates node / vertex displacements for all regions, ultimately constructing the entire heart shape assembly, i.e., the cardiac virtual chimera.

[0070] [Table 8]

[0071] The data used to train the entire generative shape construction framework is as follows:

[0072] The disclosed generative shape modeling system (also known as a generative shape construction framework) for synthesizing a cohort of organ virtual chimeras (e.g., in a specific example, cardiac virtual chimeras representing virtual whole-heart shape assemblies) may be trained and evaluated using data derived from data of real individuals available, for example, in the UK Biobank image database. General data requirements and specific data and pre-processing steps employed in this example to train the disclosed generative shape construction framework include the following:

[0073] The shapes of all anatomical parts and / or organs are represented as surface or volumetric meshes / unstructured graphs, but may also be represented as contours in the form of unstructured graphs. In this example, all cardiac structures of interest (i.e., left ventricle, right ventricle, left atrium, right atrium, and aortic root) are represented as triangular surface meshes, although other formats (e.g., polygonal meshes) could be used as well.

[0074] All mesh / unstructured graph-based representations of region and / or organ geometry used in a training population are assumed to have the same number of nodes and mesh connectivity / graph topology prior to use in training the disclosed generative framework. This can be achieved by jointly aligning all samples of a particular region and establishing spatial correspondence across the entire training population considered for that region. That is, for example, all left ventricular meshes considered in the training population are jointly aligned with each other, while all right ventricular meshes are jointly aligned with each other, and so on. The training populations for each region may vary in size, and the individuals / patients whose data are used in the region-specific training populations may differ in terms of the degree of overlap of individuals / patients between the training populations of the various regions.

[0075] According to one example, in cardiac cine magnetic resonance (cine MR) images by a cardiologist, a high-resolution 3D cardiac atlas mesh available from previous studies (where the atlas represents an average representation of a population of cardiac meshes derived from images of multiple patients) can be non-rigidly registered to manually annotated contours defining the boundaries of each of the four cardiac chambers of interest. These manual contours are available in the UK Biobank database. By registering the high-resolution 3D cardiac atlas mesh to the manual contours, a training population of shapes for each region / cardiac structure of interest can be generated (i.e., the same atlas mesh was deformed / non-rigidly registered to fit the manual contours of all possible individuals from UK Biobank, automatically establishing spatial correspondence across all samples in all region-specific training populations). The high-resolution cardiac atlas used to generate the training population included regions / structures (each represented by a different mesh) such as the left and right ventricles (LV and RV), the left and right atria (LA and RA), and the aortic root (AR) as part of a multi-region shape assembly. The atlas also includes vertices shared between meshes of adjacent structures (i.e., the boundaries separating any pair of adjacent structures in the entire cardiac shape assembly). Vertices in the atlas mesh are shared at boundaries between structure pairs such as LV-RV, LV-LA, LA-RA, RV-RA, LV-AR, LA-AR, and RA-AR. The availability of these shared vertices between adjacent regions / structures helps enable self-supervised learning schemes used to train the disclosed spatial construction networks (i.e., both affine and non-rigid spatial construction networks). Thus, the disclosed generative shape construction framework formulation uses information about shared vertices between adjacent regions in multi-region and / or multi-organ shape assemblies. Here, the shared vertex information may be generated manually, semi-automatically, or automatically, or may be available prior to training the spatial construction network in the framework.The generated training population of region-specific meshes (based on the manual contours and thus the deformed / registered atlas mesh) was defined on cardiac cine MR images of individuals in UK Biobank. The disclosed approach does not rely on region-specific training populations of meshes extracted / inferred solely from cardiac cine MR images. Any region and / or organ shape of interest may be obtained from various imaging modalities and various patient populations. For example, a left ventricular mesh may be extracted / inferred from cardiac cine MR images of patients / individuals in population "A," while an aortic root mesh may be extracted / inferred from cardiac computed tomography angiography (CTA) images of patients / individuals in population "B." Furthermore, the patients / individuals included in populations "A" and "B" may (a) not overlap, (b) partially overlap, or (c) completely overlap. Here, scenario (a) means that patients included in population "A" are not included in population "B." Scenario (b) means that some patients from population "A" are included in population "B." Scenario (c) means that all patients in population "A" are also included in population "B", and vice versa. However, it is important to note that for scenario (a), where training populations for individual regions / structures are extracted / inferred from images of completely different patients / patient populations, only the independent generator 200 within the disclosed generative shape construction framework is suitable for use, i.e., the independent generator is designed to allow training with non-overlapping data (as well as partially overlapping and fully overlapping data), while the dependent generator can only be trained with either partially overlapping or fully overlapping data.

[0076] According to one example, the entire cohort used to train and test the disclosed generative shape construction framework included 2,360 subject-specific whole-heart shape assemblies (represented as meshes). These subject-specific meshes were generated by aligning a high-resolution cardiac atlas mesh to manually annotated cardiac contours available for individuals in the UK Biobank database. These 2,360 subjects were selected from 4,000 subjects in the UK Biobank for whom manually annotated cardiac contours were available. This selection was based on a visual assessment of the quality of the subject-specific meshes resulting from the alignment process (initially performed on all 4,000 subjects). Only meshes that did not introduce obvious topology errors after aligning the atlas mesh to the manual contours were retained and selected as the cohort of 2,360 subject-specific meshes. The resulting cohort of 2,360 subject-specific meshes shared spatial correspondences for nodes / vertices across all regions / cardiac structures in the shape assembly. That is, all meshes have the same number of vertices / nodes, the same mesh connectivity / graph topology, and each node / vertex across all meshes represents the same anatomical feature / location. Having the same mesh connectivity / unstructured graph topology means that the edges connecting each node to its neighbors in the mesh / unstructured graph are identical across all subject-specific meshes of a cohort. Another way to specify having the same mesh connectivity / unstructured graph topology is to indicate that the adjacency matrices of all subject-specific meshes / unstructured graphs were identical. Identical graph topology across all subject-specific meshes enables the use of spectral (i.e., Chebyshev polynomial-based) convolution in the graph convolution layers used throughout the disclosed generative shape construction framework. A fixed / identical graph topology across all inputs is a prerequisite for graph neural networks that utilize spectral convolution operations in their respective constituent graph convolution layers.The developed system may be configured to handle meshes that are not identical in terms of their vertex connectivity / graph topology by utilizing spatial graph convolution instead of spectral convolution. The cohort of 2360 subject-specific mesh / whole-heart shape assemblies was randomly split into training, validation, and test sets containing 422, 59, and 1879 samples, respectively. The purpose of this split was to have a test set that was substantially larger than the training set, thereby allowing evaluation on a more diverse population than the training set.

[0077] The disclosed generative shape construction framework may be trained with non-overlapping data, partially overlapping data, or fully overlapping data, as defined above. In this example, the disclosed generative shape construction framework was trained using partially overlapping and fully overlapping data derived from cardiac cine MR images available in UK Biobank. To ensure a fair comparison between the disclosed independent and dependent generators, only the partial overlapping and fully overlapping scenarios were performed (because the dependent generator, unlike the independent generator, cannot be trained with non-overlapping data). For the fully overlapping scenario, the data of all subjects included in the training set included all cardiac structures / regions (i.e., LV, RV, LA, RA, and AR) in the entire cardiac shape assembly, and both the independent and dependent generators were trained using a complete dataset containing 422 training samples (each with five regions / cardiac structures). For the partial overlap scenario, 300 subjects in the training set were considered to have "missing data" by including only one region / cardiac structure from each subject in the training set (i.e., this was done to mimic a real-world scenario in which only partially overlapping data is available across subjects). For the partial overlap scenario, the training set included 60 regions / cardiac structures for each of the five cardiac structures of interest from different subjects (resulting in 300 training samples of different regions / cardiac structures from different subjects). The other 122 subjects in the training set were considered to have a complete full heart shape assembly (i.e., all five regions / cardiac structures were available). This partially overlapping data was also used to train both the independent and dependent generators.

[0078] The spatial configuration networks (including both affine and non-rigid spatial configuration networks) were trained using both real and synthetic data. First, the spatial configuration networks were pre-trained using real data. The data split for training, validation, and testing was maintained similar to that used in the region-aware generation module trained in the "perfect overlap" scenario (i.e., 422 samples for training, 59 samples for validation, and 1879 samples for testing). After the initial training using real data, the spatial configuration networks were fine-tuned (i.e., training continued) using synthetic data generated using the independent generator 200 (i.e., using regions / structures synthesized using each region-specific gcVAE in the independent generator 200). The synthetic data cohort included a total of 2000 synthetic samples per region / cardiac structure of interest. Of these, 1600 were used for training, 200 for validation, and 200 for testing.

[0079] FIG. 8 illustrates a method for training a body part-aware generative model 800 according to one example of the present disclosure. The method begins at 801 by using specific non-overlapping, partially overlapping, or fully overlapping anatomical shape / mesh data (from multiple subjects / patients) as input (i.e., training data) to the body part-aware generative model (810). The body part-aware generative model is then trained to learn hidden / latent representations of the input data by minimizing a suitable loss function through numerical optimization and learning / updating the body part-aware generative model parameters using an error backpropagation algorithm (810). The training process is iterative, and the error / loss generated by the body part-aware generative model in each iteration (also called an epoch) is used to update the body part-aware generative model parameters at the end of each iteration / epoch. Training continues until a suitable convergence criterion is reached or a user-specified maximum number of training iterations / epochs is reached (820), at which point the training loop is stopped / terminated (830) and the learned body part-aware generative model parameters are saved.

[0080] FIG. 9 illustrates a method for training a spatial construction model 900 according to one example of the present disclosure. The method begins at 901 with training data (each training sample including all the features in a geometric assembly of interest) as input. The features in each training sample used to begin training the spatial construction model (910) may belong to a real patient / subject or may be synthetic / virtual features synthesized using a suitable statistical / machine learning model (e.g., the independent generator 200 or the dependent generator 300 in the disclosed system). The spatial construction model is then trained (910) to learn the spatial relationships between the input features in the geometric assembly and to estimate the spatial transformations required for the construction / registration (i.e., integration) of the features into an anatomically plausible / realistic geometric assembly. The training process is iterative, minimizing a suitable loss function through numerical optimization and learning / updating the parameters of the spatial construction model using a back-propagation algorithm. In the iterative training process, the error / loss made by the spatial configuration model at each iteration (also called an epoch) is used to update the parameters of the spatial configuration model at the end of each iteration / epoch. Training continues 920 until a suitable convergence criterion is reached or a user-specified maximum number of training iterations / epochs is reached, at which point the training loop is stopped / terminated 930 and the learned parameters of the spatial configuration model are saved.

[0081] 8 and 9 show the training occurring independently, the training may also be performed sequentially, i.e., training the part recognition model followed by training the spatial configuration model in one cycle, repeated over multiple epochs.

[0082] 10 illustrates a method for generating a virtual chimera population 1000 for use in in silico testing according to one example of the disclosure. According to this example, which may be a computer-implemented method in which organ shapes are represented as surface or volumetric meshes, the method, starting in step 1001, includes generating organ parts (1010) using a part-recognition and generation model, e.g., trained on non-overlapping or partially overlapping (or fully overlapping) patient data (to learn latent representations of each part of the multi-part organ shape to be included in the virtual population) and configured to output composite parts of the multi-part organ shape; registering previously output composite parts (registering output composite parts of the multi-part organ shape) using a spatial configuration model similarly trained on non-overlapping and partially overlapping real patient data or synthetic / virtual patient data; and outputting anatomically significant examples of the entire multi-part organ shape as virtual chimeras (1020); and storing the resulting virtual chimeras as a set of combined chimeras, known as a virtual chimera population (1030). The resulting complete virtual chimeric population may then be output (1040) for multiple subsequent uses, such as in the design or testing of medical devices, particularly in silico testing.

[0083] Generating a virtual population of anatomies and physiology as disclosed herein is important for in silico testing because it provides a systematic and comprehensive framework for investigating and evaluating medical device performance in silico. A virtual population may be a collection of samples (e.g., referred to as virtual patients) that represent plausible instances of anatomies and / or physiology that may realistically be found in the target patient population of the medical device under test, but is not necessarily representative of real patient data. A virtual population can be synthesized as disclosed herein using statistical or machine learning techniques, where the machine learning technique first learns the underlying hidden (or latent) distribution of the data by training on real patient data. The trained statistical / machine learning model may then be used to generate new data, i.e., new virtual patients, by sampling from the learned latent distribution. The resulting virtual population generated by sampling from such a trained statistical / machine learning model may also be referred to as a virtual population of anatomies and / or physiology (depending on the type of real patient data used to train the statistical / machine learning model).

[0084] Existing methods for synthesizing virtual patient populations require complete overlap of the data used to build the models, limiting their use, particularly for large heterogeneous datasets with non-overlapping or partially overlapping information. The inability to utilize (and therefore maximize the value of) non-overlapping / partially overlapping data also limits the overall anatomical and physiological variation that can be captured in virtual patient populations synthesized using existing methods.

[0085] Because the virtual chimeras / models output by the disclosed system are complete and anatomically realistic models of the physical configuration of each organ or set of organs, they can be used for any purpose for which such models are utilized, particularly for in silico testing of medical devices for the purposes of medical device development, design, and manufacturing, as well as subsequent medical testing. For example, a heart-specific virtual chimera population generated / created using the disclosed system can be used to conduct in silico testing of all implantable cardiac devices, such as transcatheter aortic valve implants, mitral valve replacement devices, defibrillators, pacemakers, and coronary stents. Virtual populations generated using the disclosed system can also find use in in silico studies that investigate the effects of drugs, genetic, lifestyle, environmental, or demographic factors on cardiac biochemical and / or biophysical processes through computational simulation and modeling. The generated virtual populations can also be used to conduct virtual imaging studies and develop and calibrate medical imaging systems (e.g., intravascular ultrasound, computed tomography, computed tomography angiography, magnetic resonance imaging, positron emission tomography, single-photon emission computed tomography, cardiac ultrasound / echocardiography, and 3D ultrasound), interventional robots, surgical guide systems, catheters, guidewires, and other non-implantable medical devices. The generated virtual populations can also be used for 3D printing of realistic physical cardiac models / phantoms for use in experimental studies, development of surgical guide systems, surgical training, and educational purposes.

[0086] It should be appreciated that some methods for providing virtual chimeras use "strong labels" within a self-supervised or fully supervised learning framework to train the neural network used to generate the virtual chimera. "Strong labels" represent reference / ground truth spatial transformations and rely on available actual complete multi-part / shape assemblies. Thus, if the reference / ground truth spatial transformations are available, fully supervised training of the neural network can predict the spatial transformations required to compose individual parts into a complete multi-part shape assembly. Accessing the reference / ground truth spatial transformations requires prior availability / knowledge of the corresponding actual complete multi-part shape assemblies. However, training a neural network using strong labels precludes the use of partially overlapping (or non-overlapping) data. Therefore, the present disclosure provides a novel self-supervised approach for a composition network that can be trained with weak labels, where the weak labels represent information shared between adjacent parts in an actual incomplete shape assembly (i.e., including partially overlapping or non-overlapping data).

[0087] Examples of the present disclosure may be implemented by suitably programmed computer hardware. Figure 11 is a block diagram 1100 illustrating components capable of reading instructions from a machine-readable or computer-readable medium (e.g., a non-transitory machine-readable storage medium) and performing any one or more of the methods described herein, thus providing an apparatus or system for performing the virtual population modeling described, according to some exemplary embodiments. Specifically, Figure 11 shows a diagrammatic representation of hardware resources 1105 comprising one or more processors (or processor cores) 1110, one or more memory / storage devices 1120, and one or more communication resources 1130, each of which may be communicatively coupled via a bus 1140. Processor 1110 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a multiple instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP) such as a baseband processor, an application specific integrated circuit (ASIC), a cloud processing function (e.g., an AWS instance), another processor, or any suitable combination thereof) may include, for example, processor 1112 and processor 1114.

[0088] The memory / storage 1120 may include main memory, disk storage, or any suitable combination thereof, including, but not limited to, any type of volatile or non-volatile memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, solid state storage (SSD), magnetic storage-based hard disk drive (HDD) media, etc.

[0089] Communications resources 1130 may include interconnect or network interface components, or other suitable devices, that communicate with one or more peripherals 1104 or one or more databases 1106 over network 1108. For example, communications resources 1130 may include wired communications components (e.g., for coupling via Ethernet, Universal Serial Bus (USB), etc.), cellular communications components, NFC components, Bluetooth® components (e.g., Bluetooth® Low Energy), Wi-Fi® components, and other communications components.

[0090] The instructions 1150 may include software, programs, applications, applets, apps, or other executable code for causing at least one of the processors 1110 to perform any one or more of the methods discussed herein. The instructions 1150 may reside, in whole or in part, on at least one of the processors 1110 (e.g., in the processor's cache memory), the memory / storage 1120, or any suitable combination thereof. Furthermore, any portion of the instructions 1150 may be transferred to the hardware resources 1105 from any combination of the peripherals 1104 or the database 1106. Thus, the memory of the processor 1110, the memory / storage 1120, the peripherals 1140, and the database 1106 are examples of computer-readable and machine-readable media.

[0091] In some embodiments, an electronic device, network, system, chip, component, or portion or implementation thereof of FIG. 11 or any other figure herein may be configured to perform one or more processes, techniques, methods, or portions thereof described herein.

[0092] In a particular example, the disclosed generative shape construction framework may be implemented using the PyTorch, PyTorch Geometric, and PyTorch3D Python libraries on a PC with an NVIDIA RTX 2080Ti graphics processing unit (GPU).

[0093] In one example, the part-aware generation modules (i.e., both the independent and dependent generators) are trained using the Adam optimizer with an initial learning rate of 5e -04 The learning rate decay per training epoch was 0.9. All networks in the part recognition generation module were trained with a batch size of 16 for up to 100 training epochs. Both the affine and non-rigid spatial construction networks were trained with a batch size of 8 for up to 100 training epochs. The initial learning rate used for both spatial construction networks was 1e -04 The learning rate decay per epoch was 0.9. All hyperparameters associated with the network architecture, training routine, and optimizer used were empirically tuned / set based on the average loss values ​​observed on the validation set in preliminary prototyping experiments.

[0094] During training of both the independent generator 200 and the dependent generator 300, a warm-up strategy can be employed to improve stability and prevent mode collapse in the learned posterior distributions on the latent variables. This is achieved by first training the gcVAE for the regions in the independent generator and the gcmcVAE in the dependent generator (i.e., trained as a regular autoencoder) with the weight of the KL loss term (i.e., w0 in Equation 1) set to zero for 200 epochs. The learned weights initialize subsequent training steps for the independent and dependent generators, where the weight of the KL loss term is initially set to a small value, i.e., 1e -6After being set to , it is multiplied by, say, 1.25 times every epoch to a maximum value of 1e -4 The other weights in the regularization term of the composite loss function (see Equation 3) and the self-supervised alignment loss function, namely w1, w2, w3, and w4, are increased to 8, 1, 8, and 5e, respectively. -7 was set to.

[0095] The initial learning rate is 1e -4 The same network architecture and optimization parameter settings as those of the body part recognition generator can be retained for the constituent networks, except that the metric is set to and the batch size is set to 8. Body part shapes may be randomly sampled from data from various patients when training the constituent networks. The constituent networks may be pre-trained on original shapes in the actual population, for example, for 100 epochs, and then trained on previously generated shapes. The generated dataset may be randomly split (e.g., 1600 / 200 / 200 split) into training, validation, and test portions. In particular, the constituent networks are trained only on partially overlapping data, i.e., structures obtained from various patients.

[0096] The functions have been described above as modules or blocks, i.e., functional units operable to perform a described function, algorithm, etc. These terms are considered synonymous. Where modules, blocks, or functional units are described, they may be formed as processing circuitry, and the circuitry may be a general-purpose processor circuit configured with program code to perform a particular processing function. Alternatively, the circuitry may be configured with processing hardware modifications. The configuration of a circuit to perform a particular function may be entirely hardware, entirely software, or may use a combination of hardware modifications and software implementations. Program instructions may be used to configure logic gates in a general-purpose or special-purpose processor circuit to perform a processing function.

[0097] The circuitry may be implemented as a hardware circuit comprising, for example, custom very large scale integrated circuit (VLSI) circuits or gate arrays, off-the-shelf semiconductors such as logic chips, transistors, or other discrete components. The circuitry may also be implemented in programmable hardware devices such as field programmable gate arrays (FPGAs), programmable array logic, programmable logic devices, system-on-chip (SoCs), etc.

[0098] The machine-readable program instructions may be provided on a transitory medium such as a transmission medium, or a non-transitory medium such as a storage medium. Such machine-readable instructions (computer program code) may be implemented in a high-level procedural or object-oriented programming language, although the program may also be implemented in assembly or machine language if desired. In either case, the language may be a compiled or interpreted language, and combined with hardware implementations. The program instructions may be adapted for execution on a single processor, or for distributed execution across two or more processors.

[0099] An example provides a computer-implemented method for generating a virtual chimera population of multi-site organ shapes for use in in silico studies (e.g., including in silico testing or in silico observational studies), where the organ shapes are represented as surface or volumetric meshes, the method including: using a site-aware generative model to learn a latent representation for each site of the multi-site organ shape to be included in the virtual population and outputting a composite site of the multi-site organ shape; using a spatial configuration model to align the output composite sites of the multi-site organ shape and output an anatomically significant example of the entire multi-site organ shape as a virtual chimera; and storing the virtual chimera in the virtual population for use in in silico testing. The composite site may include one or more meshes. The site-aware generative model may include an independent generator that learns a latent representation for each site of the multi-site organ shape. The independent generator may include a graph convolutional variational autoencoder (gcVAE) network architecture operable to learn variations of each site of the multi-site organ shape observable across the training population. The independent generators may be configured to learn the independent variations of each part of the multi-part organ shape observable across the training population. Each of the independent generators, for example, the gcVAEs, may be configured with a total loss function L totalThe gcVAEs may be trained independently of each other by minimizing the difference (x, y, y). Parameters of the encoder-decoder network pairs in each gcVAE may be iteratively learned, for example, by a backpropagation algorithm. The part-aware generative model may include a dependent generator that learns a shared latent representation of all parts of the multi-part organ shape. The independent generators and dependent generators may each include at least one encoder and at least one decoder. In an example of a dependent generator using multiple encoder-decoder pairs, the dependent generator may further include at least one multiplexing layer between the encoder and decoder, the multiplexing layer operable to provide the output of each active encoder to each active decoder. The dependent generator may include a graph convolutional multi-channel autoencoder (gcmcVAE) network architecture operable to learn joint variations in the shapes of each part of the multi-part organ shape observable across the entire training population. The joint variations may be combinatorial variations observed across various organ parts or organs of the multi-part organ. The multi-part organ shape may include a multi-organ shape assembly. The disclosed methods may be nested in use to construct higher-order assemblies from sub-organs. That is, many unique multi-part organ shapes may be modeled and then combined into an overall multi-organ assembly, up to the entire virtual patient. For example, the heart, liver, and kidneys may all be modeled and combined into a heart-liver-kidney assembly. The spatial construction model may further comprise at least one of an affine construction network and a non-rigid construction network. By way of example, the gcVAE network, the gcmcVAE network, the affine construction network, and the non-rigid construction network may each further comprise at least one residual graph convolution downsampling function block and at least one residual graph convolution upsampling function block.At least one residual graph convolution downsampling function block and / or at least one residual graph convolution upsampling function block may include at least one Chebyshev graph convolution function, at least one exponential linear identity function, and at least one instance normalization function. At least one residual graph convolution downsampling function block may include a mesh pooling layer function, and / or at least one residual graph convolution upsampling function block may include a mesh upsampling layer function. The affine construction network may include at least one of a scaling branch, a rotation branch, and a translation branch. The affine construction network may further include a concatenation layer and / or at least one fully connected layer (also known as a dense layer or single-layer / multi-layer perceptron). The non-rigid construction network may further include a feature embedding array and / or a Chebyshev convolution layer. The disclosed exemplary method may further include providing non-overlapping or partially overlapping patient data during training. Non-overlapping or partially overlapping patient data may include shape data extracted from imaging data of the same or different patients, and the imaging data may be acquired using the same or different imaging modalities. That is, according to various examples, the nature of the data overlap may be defined with respect to the extent of a particular patient source or original data source (i.e., modality). In some examples, the anatomical extent may be a result of the modality used. Because different imaging modalities can detect different parameters of the patient or organ under study, a complete patient can be modeled by combining data from two or more modalities. For example, even if the modalities capture the same field of view, they may detect different aspects of the patient's anatomy within that same field of view. As a specific example, one modality may image (i.e., "view") muscle structures, while another modality may image vascular structures.In other words, by allowing a mix of patient imaging modalities, more complete and accurate input training data sets can be derived and used to train the disclosed generation system 100, resulting in a more complete and accurate composite variable virtual chimera (while still being variable within these more accurate training derivation boundaries). Various imaging modalities may include any one or more of magnetic resonance imaging (including cross-sectional, angiographic, or any variant of structural, functional, or molecular imaging), computed tomography (including cross-sectional, angiographic, or any variant of structural or functional imaging), ultrasound (including any variant of structural or functional imaging), positron emission tomography, and single-photon emission tomography. Each of these imaging techniques may be enhanced with appropriate endogenous or exogenous contrast agents to highlight anatomical or functional structures of interest.

[0100] An example provides a method of training a generative shape modeling system (e.g., including a site-aware generative model and a spatial configuration model) as disclosed herein, comprising iteratively providing non-overlapping patient data and / or partially overlapping patient data to the generative shape modeling system over a predefined number of training epochs. The method may further include utilizing the trained generative shape modeling system to generate (i.e., derive) a virtual chimera population. The disclosed exemplary method may further include performing in silico design or testing of a medical device using the generated virtual chimera population.

[0101] Examples also provide computer-readable media containing instructions that, when executed by one or more processors, cause the one or more processors to perform any of the described methods or portions thereof (e.g., training methods and / or methods for generating virtual chimeras using a trained system).

[0102] The examples also provide apparatus arranged or configured to carry out any of the described methods.

[0103] The examples also provide output synthetic datasets (referred to as virtual chimeras or chimera populations) generated (i.e., derived and synthesized) according to the disclosed examples for use in in silico studies.

[0104] Examples also provide a method for designing or testing a medical device, the method including: training a generative shape modeling system using non-overlapping or partially overlapping patient data; (optionally testing the trained system); generating new instances of virtual chimeras based on the trained generative shape modeling system; incorporating the new instances of virtual chimeras into a virtual chimera population; and designing or testing the medical device using the virtual chimera population. Thus, examples herein provide a method for generating virtual patient populations (e.g., for IST-based testing of medical devices) that capture sufficient anatomical and physiological variation representative of a wide variety of potential target patient populations to enable meaningful evaluation of the performance of each medical device.

[0105] The examples of the present disclosure may be used in any in silico research (including, but not limited to, in silico trials and observational studies) such as those used in drug development and the development of general medical devices (including, for example, digital health products, AI-based medical software, etc.), as well as in the development of medical devices such as implants intended for use inside or outside a patient. In other words, the examples may be used in any type of in silico research.

[0106] The independent generator 200 and the dependent generator 300 in the disclosed system have different designs and therefore have different characteristics and advantages. The dependent generator can capture nonlinear correlations between input regions in a shape assembly by learning a shared latent representation across all regions. Therefore, the dependent generator 300 can generate virtual regions (or synthetic regions) with higher specificity / anatomical plausibility than the independent generator 200. Because the dependent generator is trained to learn a shared latent representation that explains the shape variations observed across all regions in the shape assembly, the shared latent representation learned in the dependent generator 300 is more constrained than the independent latent representation learned in the independent generator 200. In other words, the shared latent representation learned in the dependent generator 300 is constrained to capture the correlations that exist between regions in the shape assembly, and therefore must be able to explain simultaneous or coordinated variations in the shapes of all regions in the shape assembly. In contrast, the latent representation learned in the independent generator 200 is not subject to such constraints. That is, the latent representations learned for each region are completely independent of each other and therefore focus only on explaining the shape variation of each specific region. Therefore, the independent generator 200 is more flexible than the dependent generator and can capture greater variation in the shape of the region of interest during shape assembly. In other words, the dependent generator 300 can generate more anatomically plausible / realistic virtual chimera cohorts than the independent generator 200. Meanwhile, the independent generator 200 can generate virtual chimera cohorts with more diverse shapes / capture greater variation in shapes (of individual regions and their combinations) than the dependent generator 300. The dependent generator 300 may be used instead of the independent generator 200 when the application of interest requires higher statistical fidelity. For example, in the context of conducting an in silico trial, the dependent generator 300 would be a better choice if the trial design includes very specific inclusion and exclusion criteria to determine the virtual cohort to be used in the in silico trial for evaluating the performance of a medical device or drug.If the application of interest requires a higher degree of virtual population diversity, then the independent generator 200 may be used instead of the dependent generator 300. For example, in the context of an in silico trial, if the trial aims to investigate the performance of a medical device or drug in a non-standard or niche patient population that is less commonly / frequently observed in the real world, then the independent generator 200 would be a better choice.

[0107] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art will recognize that numerous variations, modifications, and substitutions will occur to those skilled in the art without departing from the scope of the present disclosure. It is understood that various alternatives to the embodiments of the present invention described herein may be employed in any combination in practicing the present disclosure. The following claims define the scope of the invention, and it is intended to cover methods and structures within the scope of the claims and their respective equivalents.

Claims

1. 1. A computer-implemented method for generating a multi-site organ-shaped virtual chimeric population for use in in silico studies, comprising: using a part-aware generative model to learn a latent representation for each part of the multi-part organ shape to be included in the virtual population, and outputting a composite part of the multi-part organ shape; using a spatial configuration model to register the output composite portions of the multi-part organ shape and output an anatomically significant example of the entire multi-part organ shape as a virtual chimera; storing said virtual chimeras in said virtual population for use in said in silico testing; A method comprising:

2. The method of claim 1 , wherein the part-aware generative model comprises an independent generator that learns a latent representation for each part of the multi-part organ shape.

3. 3. The method of claim 2, wherein the independent generator comprises a graph convolutional variational autoencoder (gcVAE) network architecture operable to learn variations for each part of the multi-part organ shape observable across a training population.

4. The method of claim 1 , wherein the part-aware generative model further comprises a dependent generator that learns a shared latent representation of all parts of the multi-part organ shape.

5. 5. The method of claim 4, wherein the dependent generator comprises a graph convolutional multi-channel variational autoencoder (gcmcVAE) network architecture operable to learn simultaneous variation in the shape of each part of the multi-part organ shape observable across a training population.

6. The method of claim 1 , wherein the spatial configuration model further comprises at least one of an affine configuration network and a non-rigid configuration network.

7. 7. The method of claim 3, wherein the gcVAE network, the gcmcVAE network, the affine construction network, and the non-rigid construction network each further comprise at least one residual graph convolution downsampling function block and at least one residual graph convolution upsampling function block.

8. 8. The method of claim 7, wherein the at least one residual graph convolution downsampling function block or the at least one residual graph convolution upsampling function block comprises at least one Chebyshev graph convolution function, an exponential linear identity function, or an instance normalization function.

9. 9. The method of claim 7 or 8, wherein the at least one residual graph convolution downsampling function block comprises a mesh pooling layer function and / or the at least one residual graph convolution upsampling function block comprises a mesh upsampling layer function.

10. The method of any one of claims 6 to 9, wherein the affine construction network includes at least one of a scaling branch, a rotation branch, and a translation branch.

11. The method of claim 5 , wherein the non-rigid configuration network further comprises a feature embedding array or a Chebyshev convolutional layer.

12. 12. The method of any one of claims 1 to 11, comprising providing non-overlapping or partially overlapping patient data during training.

13. 13. The method of claim 12, wherein the non-overlapping patient data or the partially overlapping patient data comprises shape data extracted from imaging data of the same patient or different patients, the imaging data being acquired using the same imaging modality or different imaging modalities.

14. 14. The method of claim 13, wherein the different imaging modalities include magnetic resonance imaging, computed tomography, computed tomography angiography, ultrasound, positron emission tomography, and single photon emission tomography.

15. 15. The method of any one of claims 1 to 14, further comprising using the generated virtual population to perform in silico design, testing, or regulatory approval of a medical device (including, but not limited to, an implant, digital health, or artificial intelligence device) or drug.