Method and apparatus for generating virtual anatomical populations
By synthesizing multi-part organ-shaped components through the generation framework, the problem of patient variability limitation in clinical trials of medical devices is solved, the diversity of virtual patient populations and the utilization of cross-data sources is achieved, and the efficiency of medical device design and testing is improved.
Patent Information
- Application Number
- CN202380075239.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-03
- Filing Date
- 2023-10-03
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is limited by patient variability in clinical trials of medical devices, resulting in insufficient safety and efficacy verification, increasing the risk of late failure and development costs.
Using a generative framework, a reasonable multipart organ shape component is synthesized through the macroanatomy of the unstructured graph as input data and encoded into a potential representation to synthesize instances of multiple parts independently or simultaneously.
It realizes the generation of diverse virtual patient groups, leveraging cross-modal and cross-popular data sources, reducing dependence on fully overlapping data, and improving the efficiency and accuracy of medical device design and testing.
Smart Images

Figure CN120113010A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer-implemented methods, computer programs and apparatus for generating virtual anatomical populations for use in computer simulation studies and their use for medical device or drug design, manufacture, testing and regulatory approval. Background Art
[0002] The cost of developing medical devices has increased steadily in recent years, driven by lengthy regulatory approval processes and expensive clinical trials. The high rate of late-stage failures of medical devices in clinical trials results in significant financial losses for medical device manufacturers, and the financial losses incurred by such late-stage failures are often subsequently passed on to the far fewer successful medical devices that are ultimately approved. The resulting financial burden on the global healthcare system and medical device industry is rapidly becoming unsustainable, and a paradigm shift in medical device innovation and the regulatory approval process is clearly needed.
[0003] Traditional (i.e., in vivo) clinical trials of medical devices are often limited by the limited variability of the protocols (e.g., patient variability) that the medical device can withstand when evaluating its safety and efficacy. This may be due to the difficulty in recruiting enough patients to cover the variability of the target population of the medical device under development (i.e., the target patients for whom the developed medical device is intended to be used in routine clinical care). Inadequate validation of the safety and efficacy of a medical device can result in serious adverse side effects becoming apparent only after post-marketing approval, particularly when used in protocols that were not adequately evaluated in pre-marketing approval clinical trials.
[0004] However, in silico trials (ISTs) offer an alternative, complementary pathway to medical device innovation and the regulatory approval process that strengthens the path to medical device commercialization and its adoption in routine patient care. ISTs use computational modeling and simulation to virtually evaluate the safety and efficacy of medical devices in virtual patient populations, and provide digital (or in silico) rather than real-world evidence to support regulatory approval of medical devices under development. As a result, ISTs have the potential to explore medical device performance in a wider range of patient characteristics that would not be feasible when recruiting in real clinical trials, and thus ISTs can help improve, reduce, and partially replace in vivo clinical trials for medical device testing. Summary of the invention
[0005] The technology described herein provides a generation framework that enables the simulation of animals (including humans) in anatomy and physiology to design and test medical devices or drugs. In particular, the shape of the macroscopic anatomical structure represented as an unstructured graph is used as input data to synthesize a reasonable multi-part organ (or multi-organ shape component). The disclosed generation framework can model each part / organ independently or jointly, and then encode the single or multi-part anatomical shape into a potential representation (i.e., a low-dimensional representation, such as a vector), which can be decoded to independently synthesize an instance of each part or simultaneously synthesize an instance of multiple parts. The synthesized parts can be combined in space (i.e., combined in space) to generate reasonable multi-part / multi-organ shape components, which are referred to as "virtual chimeras" in this article. Advantageously, the described method and associated equipment allow the use of non-overlapping / partially overlapping information from different data sets for training, and then, the trained system is used to synthesize a virtual chimera population. For example, in the specific area of cardiac modeling, aortic vascular geometry extracted from computed tomography angiography (CTA) data from a clinical trial cohort can be combined with cardiac chamber geometry extracted from cine magnetic resonance images (cine MRI) acquired across a general population, such as that provided by the UK Biobank, to produce a complete cardiac virtual chimera population.
[0006] It is not feasible to generate a cardiac virtual population including the aortic vascular geometry using only cardiac cine MRI images, since the field of view of images of this kind does not capture the aorta. However, cardiac CTA images do capture the aorta and enable its geometry to be extracted and characterized in 3D. The benefit of the proposed method is that it enables parts / organs of an organ extracted from one type of images (i.e., imaging modality) of a specific population (e.g., the aorta from CTA of a transcatheter aortic valve implantation (TAVI) clinical trial cohort) to be combined with other parts / other organs of the same organ from one or several other imaging modalities, which may even be acquired for a different patient population (e.g., cardiac chambers from cine MRI of a population imaging initiative such as the UK Biobank).
[0007] The disclosed method may also include a combined neural network for spatially combining real or synthetic anatomical parts belonging to a multi-part or multi-organ assembly of interest. In such an example, the combined neural network can learn the spatial relationships between adjacent anatomical parts and / or organs of interest, and recover affine transformations (i.e., linear mappings) and non-rigid spatial transformations that can configure the parts and / or organs into an anatomically reasonable combined (holistic) assembly. Beneficially, such a combined neural network can also be independent of the availability of fully overlapping data for training, for example, because it is constructed through self-supervision, it can be trained with both non-overlapping data and partially overlapping data.
[0008] Compared to prior known techniques that depend on the availability of fully overlapping data in a training population (i.e., they require all parts and / or organ structures to be available for each patient included in the training population), the disclosed methods can generate a more diverse virtual patient population and can exploit cross-modality and cross-population data sources. In the specific use case of generating a multi-part whole heart (heart), Figure 1 The disclosed generation framework for synthesizing a virtual chimera population is described by schematic diagrams in FIG. However, the disclosed method can be similarly used to synthesize any multi-part or multi-organ shape components (e.g., blood vessels, bone structures, or any other sub-selection of physical entities within an anatomical structure).
[0009] An advantage of the present disclosure is that disparate medical datasets that may have only non-overlapping or only partially overlapping anatomical structures can be used to learn generative models of multi-part anatomical shape assemblies. Thus, examples of the present disclosure address the problem of creating models (i.e., so-called "virtual chimeras") by combining data from different subjects (or datasets) with partially overlapping (or non-overlapping) anatomical structures in a generative shape modeling framework that is capable of synthesizing multi-part shape assemblies that represent realistic natural anatomical structures.
[0010] The disclosed generative shape combination learning methods utilize combination networks that predict rigid or affine transformations to constitute independent parts that are synthesized into coherent multi-part assemblies, however, these can be extended to include non-rigid registration components in the combination neural networks to "stitch" the independent parts (anatomical structures) together.
[0011] Advantages of the present disclosure may include: the ability to use disparate patient datasets to train different parts of the disclosed generative framework (e.g., those with non-overlapping and partially overlapping anatomical structures / multi-part shape components); providing a generative framework that can effectively capture the variability of the shape of each independent part in an organ or organ component, and can adapt to the changing topology across all parts within an organ or organ component. While maintaining anatomical realism, it ultimately does not depend on the availability of all parts of the shape component across all samples in the training set (i.e., does not require fully overlapping input data). BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Examples of the present invention are further described below with reference to the accompanying drawings, in which:
[0013] Figure 1 An example system for generating a virtual organ assembly (or training thereof) according to an example of the present disclosure is shown;
[0014] Figure 2 An example according to the present disclosure is shown. Figure 1 An example standalone generator used in the system;
[0015] Figure 3 An example according to the present disclosure is shown. Figure 1 An example slave generator used in the system;
[0016] Figure 4 An example according to the present disclosure is shown. Figure 1 Example architecture of the residual graph convolution downsampling block used in all graph neural networks in our system.
[0017] Figure 5 An example according to the present disclosure is shown. Figure 1 Example architecture of the residual graph convolution upsampling block used in all graph neural networks in our system.
[0018] Figure 6 An example according to the present disclosure is shown. Figure 1 An example of a graph neural network-based architecture using an affine spatial combination network used in the spatial combination module of ;
[0019] Figure 7 An example according to the present disclosure is shown. Figure 1 An example of a graph neural network-based architecture with a non-rigid spatial combination network used in the spatial combination module of
[0020] Figure 8 A method of training a partial perception model according to a disclosed example is shown;
[0021] Fig. 9A method of training a spatial combination model according to a disclosed example is shown;
[0022] Fig.10 A method of generating a virtual population for use in a computer simulation experiment according to a disclosed example is shown;
[0023] Fig.11 An apparatus for performing a method of generating a virtual population for use in a computer simulation experiment according to a disclosed example is shown. DETAILED DESCRIPTION
[0024] The technical operations and advantages of the present disclosure should now be presented by way of a number of examples illustrating only the novel and inventive features, and the disclosed examples are intended to be fully combinable in any reasonable combination unless expressly stated otherwise (eg, defined as true alternatives).
[0025] Figure 1 An example system 100 for generating a virtual organ assembly according to an example of the present disclosure is generally shown. The disclosed system 100 includes two main components in its most general form (also referred to as a generative shape modeling system) - a part-aware generation module 110 and a spatial combination module 120. The part-aware generation module 110 includes a part-aware generation model block 102 that takes as input 101 a set of one or more patient data (e.g., partially overlapping data or non-overlapping data in the form of meshes 101a to 101e) and provides as output 103 a learned latent representation of those meshes, which in turn is used to define a set 115 of (sub)structures of composite parts (of an organ) or organ shapes (e.g., in a multi-organ assembly). The spatial combination module 120 includes a spatial combination model block 122 that takes as input the composite parts / structures 115 and spatially arranges them to form a realistic set 125 of spatial combinations of parts of parts (e.g., organ parts) or assemblies of parts (e.g., of a multi-part organ assembly).
[0026] According to the example, both modules 110 and 120 are parameterized (defined by) a graph convolutional neural network. A generative model such as that used in the part-aware generator model block 102 is an algorithm that is capable of learning from real data and synthesizing virtual instances that are similar but not identical to the real data instances used to train it. The part-aware generative model 102 is used to synthesize one or more virtual instances of anatomical parts / shapes of a multi-part organ (i.e., a single organ with a predefined set of parts or sub-parts) or a multi-organ shape or assembly (i.e., multiple organs including a predefined set of parts / sub-parts). In the example disclosure below, the example used will be a human heart including multiple parts / structures. However, other undescribed examples may include multiple abdominal organs (such as the liver, kidneys, pancreas, etc.). The spatial combination model 122 is used to spatially organize (i.e., move / arrange, or "combine") one or more synthesized virtual instances of parts (of an organ or organ assembly) into a multi-part / multi-organ shape assembly with anatomical significance. The independent spatially combined multi-part / multi-organ shape components output by the spatial combination model 122 are called "virtual chimeras" (which can be considered as a single complete model of the corresponding organ or organ component) and can be combined together to obtain a "virtual population" created by generating multiple such virtual chimera instances (which can be referred to as a "virtual chimera population").
[0027] As mentioned above, Figure 1 A schematic diagram is depicted that provides an overview of a generative shape modeling system 100 for synthesizing a population of virtual chimeras from non-overlapping and partially overlapping data sets of anatomical shapes. In this example, the partial / non-overlapping nature of the input data used is shown by shading. The first step / component of the generative shape modeling system 100 is a part-aware generative model 102, where the part-aware generative model has two different implementations - an independent generator implementation and a dependent generator implementation. These two different implementations are not used together - either an independent generator is used or a dependent generator is used, i.e., they are two ways of solving the same problem. By Figure 2 The schematic diagram shown in , depicts an independent generator implementation, which illustrates how an independent generator can be constructed and trained to learn independent latent representations (i.e., hidden low-dimensional representations) for each anatomical part or organ of interest independently of the others. Meanwhile, the dependent generator implementation in Figure 3 , and is constructed and trained to jointly learn a shared latent representation of shapes across multiple parts and / or multiple organs of interest in order to synthesize desired multi-part / multi-organ shape components.
[0028] In an example based on cardiac modeling, the generative shape modeling system 100 can be used to use data from example patient data covering different individuals (and / or originating from different data sources), and in particular those data that lack complete overlap of the example patient data (i.e., all anatomical parts are not available for all patient data used to build the system) to learn independent or shared latent representations of the anatomical parts that make up the human heart, which in this example includes the left ventricle (LV), right ventricle (RV), left atrium (LA), right atrium (LA), and aortic root (AR). The disclosed generative shape modeling system 100 is not limited to modeling the above-mentioned cardiac structures, but can be extended to include additional cardiac structures (if available), such as coronary arteries and their branches, pulmonary arteries, pulmonary veins, heart valves, etc. In addition, the disclosed system 100 is not limited to cardiac modeling, and can be used to model other organs as instances of multi-part shape components (e.g., lungs, liver, spine), or even multiple organs as instances of multi-part shape components (e.g., chest organs and abdominal organs). That is, while the particular exemplary embodiments described herein focus on generating virtual queues of hearts, other embodiments may be used to synthesize virtual queues of other multi-part organs or multi-organ shaped assemblies.
[0029] Now refer to Figure 2 , the independent generator 200 may include multiple (N) independent generator sub-models parameterized by (i.e., defined by) graph convolutional variational autoencoders (gcVAEs) 202(A) to 202(E), each sub-model (i.e., gcVAE) corresponding to each of the N anatomical parts and / or organs to be synthesized to create the multi-part / multi-organ shape assembly of interest. Figure 2 In the example of , N=5 because there is a gcVAE provided for each of the left ventricle, right ventricle, left atrium, right atrium, and aortic root that combine to make up a multi-part shape component (in this case, the human heart). That is, sub-model (A) is a partially specific gcVAE network architecture used in the independent generator 200 to learn the variability of the left ventricle (LV) shape observed across the training population; sub-model (B) is a partially specific gcVAE network architecture used to learn the variability of the right ventricle (RV) shape observed across the training population; sub-model (C) is a partially specific gcVAE network architecture used to learn the variability of the left atrium (LA) shape observed across the training population; sub-model (D) is a partially specific gcVAE network architecture used to learn the variability of the right atrium (LA) shape observed across the training population; and sub-model (E) is a partially specific gcVAE network architecture used to learn the variability of the aortic root (AR) shape observed across the training population.
[0030] Once trained, each part-specific gcVAE 202(A) to 202(E) can be sampled to generate new parts (of organs / organ components) that are representative of the input training dataset, but are not identical to any input dataset used in training, i.e., they can now generate synthetic examples that are known to be realistic.
[0031] Each independent generator 200 includes a set of graph convolutional variational autoencoders (gcVAE) 202A to 202E. Each gcVAE 202 includes a pair of neural networks called an encoder (e.g., encoder 206A) and a decoder (e.g., decoder 216A), where the encoder includes a plurality of residual graph convolutional downsampling (RGCDS) blocks (e.g., 210A to 214A) and the decoder includes a plurality of residual graph convolutional upsampling (RGCUS) blocks (e.g., 218A to 222A). Figure 4 and Figure 5 The RGCDS block and the RGCUS block are discussed in more detail. Figure 2 As shown in , each of the five sub-models (represented by letters A to E) takes as input the grid for the corresponding part (e.g., LV grid 201A for sub-model gcVAE 202(A), etc.) (which is downsampled in encoder 206) to form a fully connected layer latent vector 215. Each of the five different sub-models can apply a different vector size, and in the example shown, 16×1 is applied for LV and LA, 12×1 is applied for RV and RA, and 8×1 is applied for AR. The size of the latent vector for each of the five different sub-models is determined empirically and can be of different sizes for different applications and datasets used for training. The latent vector is then upsampled in decoder 216 to form an output reconstruction grid 223, such as reconstructed LV grid 223A.
[0032] In more detail: Each gcVAE is trained separately to learn an individual independent latent representation of each of the N corresponding anatomical parts and / or organs of interest. A variational autoencoder (VAE) (and the graph convolution variant used in the independent generator 200 disclosed herein, i.e., gcVAE) can be a Bayesian latent variable model that couples 1) a recognition function represented by an encoder 206 network for inferring a low-dimensional hidden representation of the input data (i.e., for inferring the latent representation) with 2) a generation function represented by a decoder 216 network that transforms the inferred latent representation back to the original data space. The VAE infers the latent representation of the input data by approximating the posterior distribution of the latent variables given the observed input data. This is achieved by jointly optimizing the encoder 206 network and the decoder 216 network, for example, to maximize the associated evidence lower bound (ELBO) of the observed input data. VAE is flexible and can be used to learn latent representations of any type of data (e.g., images, audio signals, grid-based shape representations). In the standalone generator 200 of the example generative shape modeling system 100, a VAE is combined with a graph convolutional neural network to learn a latent representation of anatomical shapes represented by a computational grid (also called an unstructured graph).
[0033] The encoder-decoder neural network pairs (i.e., 206A / 216A to 206E / 216E and 306A / 316A to 306E / 306E) in both the independent generator 200 and the slave generator 300 are graph convolutional neural networks that utilize Chebyshev convolution operations to extract hierarchies of shape features (i.e., across multiple scales) and learn representative latent representations that describe the variability of the shape of the corresponding parts and / or organs of interest across the training population. Each level has its own set of upsampling blocks and downsampling blocks (e.g., 210A / 218A), so in the example shown, 5 levels are used, but other numbers of levels may be used instead. In addition, the same number of levels is used across all example graphs, but different numbers of levels may be used in different parts. The number of levels generally affects processing speed and accuracy, so a value that is the best compromise for a given use case may be selected for each implementation. The number of levels used in the developed independent generator and slave generator is empirically determined to be optimal for the data used to train the developed system. Given different training data for the same or different applications, the number of levels selected can vary. The spatial convolution operation used in convolutional neural networks is well defined for structured / grid data in the Euclidean domain and cannot be directly applied to irregular structured data, such as graphs. The Chebyshev convolution operation provides a generalization of spatial convolution, so that spatial convolution can be applied to data defined by graphs (i.e., grids). The Chebyshev convolution operation is performed in the Fourier domain. According to the convolution theorem, the convolution operation performed in the spatial domain is equivalent to the element-by-element multiplication in the Fourier domain. The application of Chebyshev convolution is first to transform the graph-based data to the Fourier domain. Then, the Fourier-transformed graph-based data is multiplied element-by-element with the Fourier-transformed convolution filter, and then the inverse Fourier transform is applied to convert the result of the multiplication operation back to the spatial domain. The parameters / weights of the convolution filters in the Chebyshev convolution operation are parameterized by truncated Chebyshev polynomials, where the polynomial coefficients represent the learnable weights / parameters of the convolution filters. Truncated Chebyshev polynomials are used in the disclosed generative shape modeling system because they reduce computational complexity relative to other spectral convolution operations and relative to the use of general N-dimensional Chebyshev polynomials. However, according to the example, other spectral convolution operations may also be used in place of the truncated Chebyshev polynomial convolution operations. Alternatively, spatial convolution operations designed for graph-based data (such as feature-driven graph convolution operations) may also be used in place of spectral / Chebyshev convolution operations.
[0034] like Figure 2As shown in , the example independent generator 200 includes five part-specific gcVAEs corresponding to five cardiac structures of interest. To train each of the N=5 gcVAEs, consider n=1…N ={x i} i=1…D An input data set where each x i Represents the shape of an anatomical part (e.g., left ventricle) of an individual (i.e., patient) represented in the training dataset as a computational mesh (also called a triangular surface mesh, or unstructured graph). Each computational mesh is represented by a list of 3D points (i.e., spatial coordinates, which may be referred to as vertices or nodes) defining the boundaries of the anatomical part / shape and an adjacency matrix defining the vertex / node connectivity. Input dataset X n The data of each individual in X is represented by the subscript i, where {i=1...D} and the data of D individuals are included in X n In. Each X n is the input data used to train each of the {n=1...N} part-specific gcVAEs, so N=5 in this example. n The individuals of X are not necessarily the same, that is, each n The training dataset samples {x i} i=1…D can be from the same or different individuals. In other words, the dataset X used to train each gcVAE n It can include non-overlapping, partially overlapping, or completely overlapping data, where non-overlapping data means that each X n All samples in are obtained from different individuals; partially overlapping data means that each X n A certain proportion of samples are obtained from the same individuals; and completely overlapping data means that each X n All samples in are obtained from the same individual. Each encoder network in each gcVAE takes as input a grid-based shape representation of a given anatomical part or organ, i.e., x i , and maps this high-dimensional representation of the shape to a low-dimensional vector or latent representation (denoted as z i ). The low-dimensional vector or latent representation obtained by each sample passing through the encoder 206 network is then mapped back or reconstructed through the decoder 216 network to approximate the original grid of the input shape. n On training, find each data set X separately n Each potential representation of (i.e., independent of the findings for other parts).
[0035] Each gcVAE 202A to 202E is trained using its corresponding input training dataset X nis trained independently to learn the hypothesis used to generate the observed data X by approximating the posterior distribution of the true but intractable latent variable using variational inference n The latent representation of the observed data X is obtained by optimizing the parameters of the encoder-decoder network in each gcVAE. n The ELBO is used to implement variational inference, where the encoder approximates the posterior distribution (denoted as q(z)) given the input data Xn. n |X n )), and the decoder is given a latent variable (denoted as p(X n |z n )) approximates the data likelihood. Therefore, during training, for a given input sample x i , the encoder maps the input to a unique latent representation z n i , which is then used as input by the decoder to reconstruct the data (denoted as r i ). In other words, during training, the encoder learns to reduce the input data to a low-dimensional or compressed form, and the decoder learns to unpack or transform / map this compressed information back into an approximation of the data input to the encoder. Once the encoder and decoder are trained, in a generation phase (also called an inference phase or synthesis phase), new compressed forms of the data can be created and transformed by the trained decoder to generate synthetic data that looks realistic (i.e., has certain properties or characteristics shared and / or similar to real data).
[0036] Each gcVAE 202A to 202E is trained independently by jointly optimizing its constituent encoder 206 and decoder 216 networks to maximize the associated evidence lower bound (ELBO) of the observed input data. The ELBO is formulated as a loss function that includes the sum of the expected negative log-likelihood of the data (denoted as L recon ), the approximate posterior distribution of the latent variable q(z n |X n ) and the hypothesized prior distribution p(z n ) of the Kullback-Leibler divergence (denoted as L KL ) and an additional regularization loss (denoted as L regular ). The overall loss function (L) used to train each part-specific gcVAE 202A to E in the independent generator 200 total ) is given by:
[0037] L total =L recon +w 0 L KL +L regular(1)
[0038] Among them, L recon is the reconstruction loss for the negative log-likelihood of representing the data, which is calculated as i ) and the shape / grid of the original input (x i ) is evaluated vertex by vertex L 1 Norm. L recon is formulated as:
[0039] L recon =||x i -r j || 1 (2)
[0040] Kullback-Leibler divergence loss term L KL is a distance measure between probability distributions, and here, it is used to approximate the posterior distribution q(z) over the latent variables n |X n ) and the true unknown (or target) posterior distribution p(z n |X n ) is minimized. However, since the target posterior distribution p(z n |X n ) is intractable (cannot be estimated) in the current setting, so Bayes’ theorem is used to factorize it into (expressed as) the data likelihood p(X n |z n ) multiplied by the prior distribution of the hypothesis within the latent variable p(z). As a result, the Kullback-Leibler divergence loss term L KL Simplify so that the approximate posterior distribution q(z n |X n ) and the latent variable p(z n ) is minimized. The prior distribution over the latent variables is assumed to be a centered isotropic multidimensional Gaussian distribution, i.e., p(z) = N(0,I) for all n = {1...N} part-specific gcVAEs. 0 Represents the weight, which is related to L KL The term is multiplied by , in order to balance its impact on the training process relative to other terms in the overall loss function. And, the additional regularization loss term L regular Penalize outliers and encourage i )’s reconstructed output (i.e., the reconstructed shape of the anatomical part input to the encoder) is smooth. The regularization loss L regular It is formulated as a weighted sum of the Laplacian smoothness loss and the edge length loss. It is represented by Llaplacian The Laplace loss of encourages vertices in a neighborhood to move together, i.e., it penalizes neighboring vertices in the reconstructed shape / mesh to move away from each other. Similarly, denoted as L edge The edge length loss of L penalizes spurious motions of vertices relative to their neighbors. regular Given by:
[0041] L regular =w 1 L laplacian +w 2 L edge (3)
[0042] Here, w 1 and w 2 Respectively represent and L laplacian and L edge The term is multiplied so that it is relative to the overall loss function (L total ) balances their relative impact on the regularization loss term and the overall training process. laplacian Given by:
[0043]
[0044] Where v represents the shape / grid of each input (x i ) and v k ∈N(v) represents its neighboring vertices. And, L edge Given by:
[0045]
[0046] Each of the N=5 gcVAEs 202A-E in the independent generators 200 corresponding to the five cardiac structures of interest (LV, RV, LA, RA, AR) can be calculated by subjecting the overall loss function L to the parameters of their constituent encoder 206 and decoder 216 networks, respectively. total The parameters of the encoder-decoder network pair in each gcVAE 202 are learned iteratively, for example, via an error back-propagation algorithm.
[0047] Now the latter can be used as Figure 2An overview of an example specific network architecture for a partially specific gcVAE in a standalone generator 200 depicted in . The encoder-decoder network pair in each gcVAE 202A to 202E includes five blocks, where each block (e.g., blocks 210A, 211A, 212A, 213A, and 214A for the downsampling block in encoder 206A) learns shape features at a different spatial resolution, thereby enabling the encoder-decoder network to learn shape features with both global context and local context across multiple grid resolutions. This is enabled by including a grid downsampling (and upsampling for the decoder 216 network) operation in each of the five blocks of the encoder 206 and decoder 216 networks, respectively. The tables (1 to 5) below summarize the architectural details of each encoder-decoder network pair in each of the five gcVAEs that can be used to model LV, RV, LA, RA, and AR. The architectures of the residual graph convolution downsampling blocks and residual graph convolution upsampling blocks used in the encoder and decoder networks in each of the five gcVAEs, respectively, may be similar in structure, e.g., where only the downsampling factors and upsampling factors used across different gcVAEs vary, and are respectively Figure 4 and Figure 5According to the example, the number of feature channels (i.e., the number of different learnable Chebyshev convolution filters / operations) used in each graph convolution layer in the residual graph convolution downsampling blocks (from blocks 1 to 5) is: 16, 32, 32, 64, and 64, and the number of feature channels used in each graph convolution layer in the residual graph convolution upsampling blocks (from blocks 1 to 5) is: 64, 64, 32, 32, and 16. Other numbers of feature channels may also be used. For the encoder network in each part-specific gcVAE, the grid downsampling factors used across downsampling blocks 1 to 5 are different and are detailed as follows - Left ventricle: [4,4,4,6,6]; Right ventricle: [4,4,4,6,6]; Left atrium: [4,4,4,5,5]; Right atrium: [4,4,4,5,5]; and Aorta: [4,4,4,4,4], where each number in brackets corresponds to the downsampling factor used in the residual graph convolution downsampling block (from block 1 to 5) of each part-specific gcVAE. Similarly, for the decoder network in each part-specific gcVAE, the upsampling factors used across upsampling blocks 1 to 5 are different and are detailed as follows - Left Ventricle: [6,6,4,4,4]; Right Ventricle: [6,6,4,4,4]; Left Atrium: [5,5,4,4,4]; Right Atrium: [5,5,4,4,4]; and Aorta: [4,4,4,4,4], where each number in brackets corresponds to the upsampling factor used in the residual graph convolution upsampling block (from block 1 to 5) in each part-specific gcVAE. Note that the resampling factors are the same at each corresponding resolution level between the encoder and decoder, i.e., for the decoder, the order is reversed as it is from low resolution to high resolution (whereas the encoder is from high resolution to low resolution - i.e., downsampling). The number of feature channels and grid resampling factors used in the disclosed part-specific gcVAEs are empirically determined to be optimal for the example application (i.e., heart modeling) and the training data used. In generative modeling of multi-part or multi-organ shape components, their values can be changed and varied given different training data for the same or different applications. All residual graph convolution downsampling blocks and upsampling blocks used in all networks in the disclosed system also contain Figure 4 and Figure 5 Instance normalization layers, residual connections, and exponential linear activation units organized as shown in .
[0048] Table 1: Example network architecture of the gcVAE 202A applied to a left ventricle (LV) shape. LV represents the number of nodes / vertices present in each LV grid / unstructured graph for all samples in the training, validation, and test sets used to train and evaluate the disclosed system.
[0049]
[0050]
[0051] Table 2: Example network architecture of gcVAE 202B applied to right ventricle (RV) shape. RV represents the number of nodes / vertices present in each RV grid / unstructured graph for all samples in the training, validation, and test sets used to train and evaluate the disclosed system.
[0052]
[0053] Table 3: Example network architecture of the gcVAE 202C applied to the left atrium (LA) shape. LA represents the number of nodes / vertices present in each LA grid / unstructured graph for all samples in the training, validation, and test sets used to train and evaluate the disclosed system.
[0054]
[0055]
[0056] Table 4: Example network architecture of gcVAE 202D applied to the right atrium (RA) shape. RA represents the number of nodes / vertices present in each RA grid / unstructured graph for all samples in the training, validation, and test sets used to train and evaluate the disclosed system.
[0057]
[0058] Table 5: Example network architecture of gcVAE 202E applied to the aortic root (AR) shape. AR represents the number of nodes / vertices present in each AR mesh / unstructured graph for all samples in the training, validation, and test sets used to train and evaluate the disclosed system.
[0059]
[0060] As in Figure 2 As can be clearly seen in FIG, the independent generator 200 is called independent because it maintains separation between the processing of each gcVAE (202A to E), where each gcVAE is independently trained with its own training loop.
[0061] Figure 3 An example according to the present disclosure is shown. Figure 1The example slave generator 300 used in the system of . The example slave generator 300 is a single generative model formulated as a graph convolutional multi-channel variational autoencoder (gcmcVAE), which includes multiple encoder-decoder pairs, one for each part being modeled (i.e., it can be considered as a gcmcVAE with multiple branches, rather than Figure 2 The encoder 306 includes a plurality of slave gcVAEs in independent generators of the present invention, and thus includes N=5 pairs of encoder-decoder networks corresponding to the five cardiac structures of interest (e.g., LA, RA, LV, RV, AR). The encoder 306 includes a plurality of residual graph convolutional downsampling (RGCDS) blocks (e.g., 310A to 314A), and the decoder includes a plurality of residual graph convolutional upsampling (RGCUS) blocks (e.g., 318A to 322A). The RGCDS blocks and the RGCUS blocks are also similar to those described below with respect to Figure 4 and Figure 5 The same blocks are discussed in more detail. Figure 3 As shown in , each of the five encoder-decoder pairs (denoted by letters A to E) in the gcmcVAE takes as input the grid for the corresponding part (e.g., LV grid 301A) (which is downsampled in encoder 306) to form a fully connected layer latent vector 315A. Figure 3In the example of , each of the five different sub-models applies the same vector size = 12×1, but other sizes may be used instead. However, unlike the independent generator 200, each latent vector 315A to E corresponding to each of the five different sub-models in the slave generator 300 is implemented using the same fixed size (e.g., = 12×1). This is because, in the slave generator 300, each latent vector is no longer independent as in the gcVAE (202A to E) of the independent generator 200, but is shared or "slaved" to each other. In other words, each encoder (306A to E) in the gcmcVAE of the slave generator 300 maps the input grid for the corresponding part to a shared latent vector (315A to E), where the latent vector (315A to E) is considered to be shared because, by design, the output of each encoder (306A to E) is formulated as a different approximation of the same shared latent vector (or an approximation of the same shared hidden representation across all information input channels). This formulation is implemented by imposing a constraint on the loss function (described in a subsequent paragraph) used to train the gcmcVAE that the output of each encoder (306A to E) is an approximation of the same shared latent vector, where the approximation of the shared latent vector of each encoder (306A to E) is passed as input to all decoders (316A to E) (e.g., via a multiplexing layer 317), and each decoder (316A to E) is used to reconstruct the grid of all information channels or all parts input to the encoder (306A to E). In other words, the encoders (306A to E) in the gcmcVAE are channel or part specific, which output different approximations of the shared latent vector (315A to E), while the decoders (316A to E) operate across channels or parts and output / reconstruct the grid of all information channels or all parts simultaneously. Since the outputs of all encoders (306A to E) are used as inputs to all decoders (316A to E), the input size of each decoder (316A to E) must be the same. Therefore, for the slave generator 300, the latent vectors (315A to E) output by each encoder (306A to E) have the same size.
[0062] In more detail, the input to each encoder-decoder network pair in the gcmcVAE is a mesh / unstructured graph 301A to E representing the shape of the relevant anatomical part or organ, which together constitute the multi-part / multi-organ shape component to be synthesized by the disclosed generative shape modeling system.
[0063] All encoder-decoder network pairs 302A to E in the gcmcVAE are jointly trained to learn shared / common latent representations in the input data across all anatomical parts and / or organs of interest. mcVAE is an extension of VAE that jointly analyzes heterogeneous data by projecting observations from different input sources / channels to a shared / common latent representation. The training process of mcVAE learning shared latent representations across multiple data input sources / channels is similar to VAE in that it is achieved by approximating the posterior distribution of shared latent variables given the observed heterogeneous input data. This is achieved by jointly optimizing all N constituent encoder-decoder network pairs 302A to E (which represent the N data input sources / channels, or in the case of the gcmcVAE in the present invention, the shapes of the N anatomical parts and / or organs) to maximize the associated evidence lower bound (ELBO) of the observed heterogeneous input data.
[0064] To train gcmcVAE, consider n=1…N ={x n i} i=1…D A data set, each x n i represents the shape of an individual anatomical part (e.g., left ventricle) represented as a computational mesh. Each mesh / unstructured graph is represented by a list of 3D points / spatial coordinates (called vertices) defining the boundaries of the anatomical part / shape and an adjacency matrix defining the vertex connectivity. n The data of each individual patient in is represented by the subscript i, where {i=1...D} and D is X n The number of individuals included in the data. Each X n is the data input to the encoder-decoder network of {n=1...N} pairs in gcmcVAE, where N=5 in this example. Its data constitutes each X n The D individuals of X are not necessarily the same, that is, each X n The training data samples {x n i} i=1…D can be from the same or different individuals. In other words, the dataset X used to train each gcmcVAE n Can include partially overlapping or fully overlapping data, where - partially overlapping data means that each X n A certain proportion of samples are obtained from the same individuals; and completely overlapping data means that each X nAll samples in are obtained from the same individual. Unlike the independent generator 200, the dependent generator 300 cannot be trained with non-overlapping data. The gcmcVAE is designed to learn a shared / common latent representation across all information input channels / anatomical parts, and does so by exploiting the correlations present in the provided inputs. Therefore, the dependent generator uses inputs obtained from different patients / patient populations to make the information have some overlap (i.e., at least partially overlap). Each encoder network in the gcmcVAE takes as input a grid / unstructured graph-based representation of the shape of a given anatomical part or organ, i.e., x n i (Refer to Figure 3 ), and all encoder networks in the gcmcVAE jointly map these high-dimensional representations of the shape to a shared / common low-dimensional vector or latent representation (denoted as z i ), or in other words, given each encoder network corresponds to an anatomical part or organ x n i As input, the encoder network outputs a shared latent vector z i . Subsequently, the approximation of the shared / common latent representation obtained from each anatomical part or organ shape that has passed through its corresponding encoder network is provided as input to all decoder networks in the gcmcVAE. Using each of the latent vector approximations output by each of the encoder networks as input, each decoder network in turn maps back or reconstructs the original mesh of the anatomical part or organ input to its corresponding encoder. In other words, given that there are N parts in the shape component of the data for each individual patient used to train the gcmcVAE, each of the N encoders in the gcmcVAE outputs N approximations to a latent representation shared across all N parts. And, each of the N decoders in the gcmcVAE takes all N approximations of the shared latent representation as input and outputs / reconstructs N versions / approximations of the shape for its corresponding part (i.e., for the part input to the encoder corresponding to each of the N decoders). During training, the N different reconstructions output by the decoder corresponding to each part for that part result in N different losses, which are aggregated / added together and used to guide the training of the gcmcVAE. By n=1…N Train gcmcVAE on the dataset X n=1…N A shared / public latent representation of .
[0065] Each of the N=5 encoder-decoder network pairs in the gcmcVAE is trained using its corresponding input training dataset X n=1…Nis jointly trained to learn the hypothesis used to generate the observed dataset X by approximating the posterior distribution of the true but intractable latent variable using variational inference n=1…N This is the shared / common latent representation of the dataset X observed by optimizing the parameters of all encoder-decoder network pairs in the gcmcVAE n=1…N In this formulation, each encoder network in gcmcVAE approximates the posterior distribution (denoted as q n (z|X n ), or in other words, given its input data set X n , outputs an approximation of the shared latent representation, and each corresponding decoder network uses each of the N approximations of the shared latent vector z output by each of the N encoder networks in the gcmcVAE to approximate the data likelihood given z (denoted as p n (X n Since each of the N=5 input channels (i.e., shapes of anatomical parts / organs) in the gcmcVAE is assumed to be conditionally independent from all other channels given a shared / common latent representation z (i.e., these input channels are considered independent of each other but depend on the shared latent variable z), the joint data likelihood across all input channels is formulated as p(X|z) given by:
[0066]
[0067] To train gcmcVAE, given s i ={x i} n=1…N An input sample s represents a set of shapes (i.e., here, the shapes belong to one of the five cardiac structures of interest) i In the case of i Each independent shape x in n i Passed as input to its corresponding encoder network. Each s i All x in n i Each sample s used to train the gcmcVAE is derived from the same individual denoted by the subscript “i”. i All five cardiac structures do not need to be present in the gcmcVAE, that is, the gcmcVAE can be trained with partially overlapping data. All five encoder networks in the gcmcVAE convert their corresponding input x n i (If in s i There is x n i ) is mapped to a unique latent representation zi These approximations are then used as input by all decoder networks to reconstruct the shape of their corresponding anatomical parts or organs (denoted as r n i ). Therefore, the gcmcVAE includes a shape-specific encoder network that maps each shape to a shared / common latent representation and a cross-shape decoder network that can simultaneously reconstruct all shapes given the shared / common latent representation as input. This enables the subordinate generator 300, once trained, to "complete / impute" missing portions / channels of information conditioned on the provided input. Specifically, by training each decoder in the gcmcVAE to reconstruct its corresponding portion using an approximation of the shared latent space output by all encoders in the gcmcVAE (i.e., by training a decoder for cross-shape / cross-channel synthesis), this interpolation or conditional generation of missing portions / channels of information is enabled given some incomplete input portions / channels of information. In addition, design constraints in the gcmcVAE that facilitate this cross-shape training and subsequent cross-shape synthesis / interpolation of missing portions / channels of information are fixed-size latent vectors output by all encoder networks and correspondingly fixed-size input vectors expected by all decoder networks; and the formulation of the loss function for training the gcmcVAE (described in the next paragraph). Since the latent representations learned across all input parts / channels are shared, after training of the gcmcVAE, multiple parts / channels can be reconstructed using the latent representations derived from a single informative input part / channel. Additionally, a shared / common latent representation can be sampled from an assumed prior distribution over the latent variables (i.e., p(z) = N(0, I)) to synthesize new / virtual shapes of parts (not observed in the training data) by passing the sampled latent representations through all decoder networks simultaneously.
[0068] The gcmcVAE is trained by jointly optimizing all its constituent encoder-decoder network pairs so that the observed input dataset {X n} n=1…5 The ELBO is formulated as a loss function that consists of the sum of the expected joint negative log-likelihoods of the data (denoted as L recon ), shared latent variables q(z|X n ) and the Kullback-Leibler divergence (denoted as L KL ) and an additional regularization loss (denoted as L regular). The key difference between the gcVAE (202A to E) in the independent generator 200 and the gcmcVAE (302A to E) in the slave generator 300 lies in the formulation of the ELBO used to train its constituent encoder and decoder networks. The gcVAE (202A to E) in the independent generator 200 assumes that each anatomical part or organ shape of interest is completely independent, and therefore each gcVAE learns an independent / unique latent representation of its corresponding part. However, the gcmcVAE (302A to E) assumes that all parts or organ shapes in a shape component are conditionally independent, or in other words, each part depends on a shared latent representation, and therefore learns a shared latent representation across all parts. Therefore, by imposing constraints that the posterior distribution q within the shared latent variables output by each encoder (306A to E) network is n (z|X n ) is similar to the unknown target posterior distribution p(z|X 1 ,X 2 ,…,X n ) is derived by minimizing the Kullback-Leibler divergence between the gcVAEs (302A to E) in the slave generator 300. In contrast, each gcVAE (202A to E) in the independent generator 200 is trained by maximizing a separate ELBO loss function independently of the other gcVAEs. Here, each ELBO loss function is formulated so that the approximate posterior distribution q(z) over the latent variables output by each encoder (206A to E) network specific to each gcVAE is n |X n ) and the unknown target posterior distribution p(z n |X n ) is minimized independently of the other approximated posterior distributions and latent variables output by the encoders in the other gcVAEs. As in the case of the gcVAEs in the independent generators 200, it is straightforward to minimize each approximated posterior distribution q within the shared latent variable n (z|X n ) and the shared latent variable p(z|X 1 ,X 2 ,…,X n ) to train the gcmcVAE in the slave generator 300 is intractable (cannot be calculated / estimated). This is because the target posterior distribution p(z|X 1 ,X 2 ,…,X n) is unknown in the current setting and inherently intractable. Instead, the ELBO loss function used to train the gcmcVAE (302A to E) is formulated by factorizing (expressing) the target posterior distribution as the product of the joint data likelihood p(Xz) across all parts (see Equation 6) and the prior distribution of the hypothesis within the shared latent variable p(z) using Bayes' theorem. Therefore, the Kullback-Leibler divergence loss term L KL Simplified to make the posterior distribution q within the shared latent variable output by each encoder (306A to E) network in gcmcVAE n (z|X n ) and the prior distribution of the hypothesis within the shared latent variable p(z). Or, in other words, to learn a shared latent representation across all input parts, the gcmcVAE is trained by minimizing the approximation of the posterior distribution within the shared latent variable output by each encoder (306A to E) network in the gcmcVAE relative to the target posterior distribution, with the underlying assumption that each input part carries useful information about the shared latent representation of itself and all other parts in the shape component. The overall ELBO loss function (L total ) to be used to train each gcVAE in the independent generator 200 total The same way is formulated, however, for the slave generator 300, L total The independent loss terms in are formulated differently. The reconstruction loss L recon (denoting the joint negative log-likelihood of the data) is calculated as each reconstructed (r i n|n′ ) shape / grid and the original shape / grid (x) that is input to the corresponding encoder network n i ) evaluates each vertex L between 1 Here, when each of the N encoders (306A to E) outputs an approximation of the shared latent vector (z), each decoder (316A to E) receives as input the N approximations of z, and each decoder is further trained to reconstruct its corresponding part / grid (i.e., the part / grid input to its corresponding encoder) using each of the N approximations of z output by each of the N encoder networks (306A to E). Therefore, r i n|n′L represents the part / grid "n" reconstructed by one of the n = {1 ... N} decoders (316A to E) for sample "i" given an approximation of the shared latent vector n' output by each of the n = {1 ... N} encoders (306A to E) in the gcmcVAE. Since there are N encoders, n' = {1 ... N} different approximations of the shared latent space z are passed as input to each decoder (316A to E), resulting in N different reconstructions of the nth part for the i-th sample. Therefore, N different reconstruction loss terms are evaluated for each part in the shape component and are aggregated / added together to convert L recon Formulated as:
[0069]
[0070] Kullback-Leibler divergence loss term L KL is computed as the posterior distribution over the shared / common latent variable approximated by each encoder network (i.e., q n (z|X n )) and the sum of the divergences between the assumed prior distribution within the shared latent variables. The prior distribution within the latent variables is assumed to be a centered isotropic multidimensional Gaussian distribution, i.e., p(z)=N(0,I). To simplify the calculation, an isotropic multidimensional Gaussian prior is assumed, and other types of prior distributions may be used within the disclosed system, such as Dirichlet distribution, Gaussian Mixture Model, Gaussian process, Dirichlet process, and other distributions. As previously described, the gcmcVAE is trained to transform the reconstruction loss term L by as shown in Equation 7 recon Formulated and by making each approximated posterior distribution q within the shared latent variable corresponding to each input portion / channel of information and its specific encoder (306A to E) n (z|X n ) with respect to the divergence of the assumed prior distribution p(z) to learn a shared latent representation across all parts of the information / input channels. In addition, using L shown in Eq. 7 recon The formulation of L enables cross-shape / cross-channel interpolation of missing parts / channels of information (i.e., given partially overlapping data as input, gcmcVAE once trained can be used to interpolate / conditionally synthesize missing parts / channels of information). In other words, using L recon This formulation of training gcmcVAE allows reconstructing multi-channel / multi-part data given a single-channel / part input.
[0071] The L of the slave generator totalThe regularization loss term L in regular is formulated as:
[0072]
[0073] Among them, L laplacian n and L edge n is the estimated output of each decoder for each shape / grid, i.e. r n i Each of the {n=1...N} Laplace smoothness and edge length loss terms (in this example, N=5) is formulated in the same way as each gcVAE in the independent generator 200. As with each gcVAE in the independent generator 200, here, w 1 and w 2 Respectively represent and L laplacian n and L edge n The terms are multiplied so that relative to the overall loss (L total ) balances the weights of their impact on the regularization loss term and the overall training process.
[0074] gcmcVAE transforms the overall loss function L total The parameters of all encoder-decoder network pairs in the gcmcVAE are iteratively learned via the error back-propagation algorithm.
[0075] Now the latter can be used as Figure 3An overview of an example specific network architecture of a gcmcVAE 306A to E in a slave generator 300 depicted in FIG. Five pairs of encoder-decoder networks corresponding to the five cardiac structures of interest are used in the current implementation of the gcmcVAE, however, the number of encoder-decoder network pairs may vary to accommodate different applications (i.e., where the number of structures constituting the multi-part shape assembly of interest is less than / greater than five). The encoder 306 and decoder 316 networks in all five pairs (A to E) each include five blocks, where each block (e.g., blocks 310A, 311A, 312A, 313A, and 314A of the downsampling block in encoder 306A) learns shape features at a different spatial resolution, thereby enabling the encoder-decoder network to learn shape features with both global and local context across multiple grid resolutions. This is enabled by including grid downsampling and upsampling operations in each of the five blocks of the encoder 306 and decoder 316 networks, respectively. Table 6 below summarizes the architectural details of each encoder-decoder network pair 302A to E in the gcmcVAE used to model the five cardiac structures of interest (i.e., LV, RV, LA, RA, and AR). In all five pairs, the architectures of the residual graph convolution downsampling blocks (310 to 314) and residual graph convolution upsampling blocks (318 to 322) used in the encoder and decoder networks, respectively, are similar (i.e., only the downsampling factors and upsampling factors used vary across different encoder-decoder network pairs), and Figure 4The number of feature channels used in each graph convolution layer in the residual graph convolution downsampling blocks (from block 1 to block 5) are: 16, 32, 32, 64, and 64, respectively. Similarly, the number of feature channels used in each graph convolution layer in the residual graph convolution upsampling blocks (from block 1 to block 5) are: 64, 64, 32, 32, and 16, respectively. For the encoder network in each encoder-decoder network pair included in gcmcVAE, the grid downsampling factors used across downsampling blocks 1 to 5 are different. These are described as follows based on the cardiac structures accepted as input by each encoder network, - Left Ventricle: [4,4,4,6,6]; Right Ventricle: [4,4,4,6,6]; Left Atrium: [4,4,4,5,5]; Right Atrium: [4,4,4,5,5]; and Aorta: [4,4,4,4,4], where each number in brackets corresponds to the downsampling factor used in the residual graph convolution downsampling block (from block 1 to 5) in each encoder-decoder network pair. Similarly, the upsampling factors used across upsampling blocks 1 to 5 are different for the encoder network in each encoder-decoder network pair included in the gcmcVAE. These are described as follows based on the cardiac structures reconstructed / output by each encoder network, - Left ventricle: [6,6,4,4,4]; Right ventricle: [6,6,4,4,4]; Left atrium: [5,5,4,4,4]; Right atrium: [5,5,4,4,4]; and Aorta: [4,4,4,4,4], where each number in brackets corresponds to the upsampling factor used in the residual graph convolution upsampling block (from blocks 1 to 5) in each encoder-decoder network pair.
[0076] Table 6: Example network architecture of gcmcVAEs 302A to 302E that can be used with the slave generator 300 to learn shape variability for all five cardiac structures of interest.
[0077]
[0078]
[0079]
[0080] The dependent generator 300 is called dependent because a shared latent representation across all parts / input channels of information is assumed, so that the output of each decoder (316A to E) in the gcmcVAE depends on the shared latent representation across all parts. In this way, the dependent generator 300 is trained in a single training cycle by optimizing all constituent encoder (306A to E) and decoder (316A to E) networks simultaneously using a single loss function. This is in contrast to the independent generator 200, in which each gcVAE (202A to E) is trained independently for each part, and the latent representation learned for each part is independent of the latent representation learned for the other parts. In addition, the independent generator 200 is trained using a separate / independent training cycle using a separate / independent loss function for each constituent gcVAE (202A to E).
[0081] The previous description will now refer to Figure 4 and Figure 5 The example uses these residual graph convolution downsampling / upsampling blocks to enhance the efficiency of feature exchange between vertices (i.e., nodes in the grid) and alleviate the problem of gradient vanishing when training the network.
[0082] Figure 4 An example according to the present disclosure is shown. Figure 100147] An example architecture of a residual graph convolution downsampling block used in all graph neural networks in the system of . The RGCDS block takes as input a collection of grids (e.g., outputs from previous RGCDS in encoder networks used within the disclosed independent generators and slave generators) 401 and applies a Chebyshev graph convolution 410 and then an instance normalization process 420 before applying an exponential linear unit function 430. Note that in some examples, the input 401 can be a single grid. Next, another Chebyshev graph convolution 440 and an instance normalization process 450 can be applied before adding the result to a separate path 415 (i.e., the residual branch) from the original Chebyshev graph convolution 410. The summed result can then have another exponential linear unit function 470 applied before applying the final grid pooling layer function and outputting the result 491. Thus, each residual block includes at least one graph convolution layer and a residual connection that uses the input provided to the block to extract features, and the residual connection combines the outputs of the first and second graph convolution layers using summation. Chebyshev graph convolution operations are used throughout this example because its strictly localized filters enable learning of multi-scale hierarchical patterns when combined with grid pooling operations and have low computational complexity. Each graph convolution layer is followed by an instance normalization layer and an exponential linear unit (ELU) for activation. In addition, an instance normalization layer is used in the residual branch to ensure the similarity of feature statistics relative to the output of the second graph convolution layer in the block.
[0083] Figure 5 An example according to the present disclosure is shown in Figure 1 Example architecture of the residual graph convolution upsampling block used in all graph neural networks in our system.
[0084] The RGCUS block is similar to Figure 4, but in reverse. Within the disclosed partially perceptual generation model 102, an RGCUS block is used in the decoder network. As shown, the RGCUS takes as input 501 the latent vector output from the encoder network or the output from a previous RGCUS block in the same decoder network, and applies a grid upsampling function 510 before applying an exponential linear unit function 540, followed by a Chebyshev graph convolution 520 and a first instance normalization process 530, which may be followed by another Chebyshev graph convolution 550 and an instance normalization process 560. The result is then added to a separate path 515 (i.e., the residual branch) from the original Chebyshev graph convolution 520. Finally, another exponential linear unit function 570 may be applied to the summed result before outputting the result 591. As such, each residual block first processes the provided inputs by upsampling them using a grid upsampling layer, and includes two graph convolution layers that extract the feature outputs of the grid upsampling layer and a residual connection that combines the outputs of the first and second graph convolution layers using summation. Chebyshev graph convolution operations are used throughout this example because their strictly localized filters enable learning of multi-scale hierarchical patterns when combined with grid pooling operations and have low computational complexity. Each graph convolution layer is followed by an instance normalization layer and an exponential linear unit (ELU) for activation. In addition, an instance normalization layer is used in the residual branch to ensure the similarity of feature statistics relative to the output of the second graph convolution layer in the block.
[0085] Once the part-aware generation module 110 is trained using either the independent generator 200 or the dependent generator 300, all anatomical parts and / or organs in the shape component can be synthesized. Note that there are two ways to train the part-aware generation model 102. Using (e.g. Figure 2 ) independent generator 200 or (for example Figure 3) slave generator 300. For the independent generator 200, non-overlapping, partially overlapping and fully overlapping data can be used. For the slave generator 300, only partially overlapping or fully overlapping data can be used. The synthesis can be done by sampling new latent vectors from the learned latent distribution, that is, in the case of the independent generator 200, a new latent vector can be sampled from each gcVAE to synthesize each corresponding anatomical part / organ. For the slave generator 300, new latent vectors can be sampled from a shared / common latent distribution learned by the gcmcVAE across all anatomical parts / organs of interest. The sampled latent vectors can be passed as input to their corresponding decoder networks to reconstruct / synthesize the shape of the new anatomical part / organ, that is, the latent vectors sampled from each gcVAE in the independent generator 200 can be passed to their corresponding decoder networks independently of each other. At the same time, the latent vectors sampled from the gcmcVAE can all be passed to all constituent decoder networks at the same time (that is, in the slave generator, all decoder networks can receive the same sampled latent vector as input). Using either the independent generator 200 or the dependent generator 300, the anatomical parts / organs may be synthesized individually using their corresponding decoder networks. Therefore, the synthesized parts may not initially be spatially organized / arranged into shape components that represent realistic natural anatomical structures (e.g., the synthesized parts may have significant intersections / overlaps with each other, which does not account for naturally occurring variations in anatomical structures). Therefore, then, Figure 1 The spatial combination module 120 can apply a spatial combination network to learn the spatial relationship between the parts in the shape component, and estimate the spatial transformations necessary to combine (combination can also be referred to as organization or arrangement) the synthesized independent parts into an anatomically meaningful (i.e., realistic) shape component (such as in this example, the entire heart shape component of interest). An example of a spatial combination network of the present disclosure takes as input a mesh of individual anatomical parts / organs synthesized using an independent generator or a slave generator, and estimates both an affine transformation and a non-rigid (also referred to as a deformable) transformation to spatially organize the input parts into a multi-part / multi-organ shape component (i.e., a virtual chimera) that represents the natural anatomical structure. According to this example, the input to the spatial combination network includes a grid / graph-based representation of the shape of the left ventricle, right ventricle, left atrium, right atrium, and aortic root, and the spatially combined output is a virtual chimera of the heart (also referred to as the entire heart shape component).
[0086] like Figure 1 As shown in FIG. 1 , the disclosed spatial combination module 120 applies two networks, namely, 1) an affine combination network (see Figure 6 ) and 2) non-rigid combination networks (see Figure 7). The affine combination network first recovers an affine transformation to roughly combine the shapes of the input anatomical parts / organs. The roughly combined parts are then used as input to the non-rigid combination network that estimates the localized non-rigid / deformable transformation, and outputs a finely combined multi-part / multi-organ shape component, i.e., a virtual chimera. The affine combination network includes N branches / sub-networks, each of which takes one of the N anatomical parts / organ shapes as input (where N is determined by the structure of interest in the multi-part / multi-organ shape component to be synthesized). In this example, the affine combination network includes N=5 part-specific sub-networks corresponding to the five cardiac structures of interest mentioned above. At the same time, the non-rigid combination network includes a single network that takes as input the roughly combined multi-part / multi-organ shape component output by the affine combination network. The affine combination network and the non-rigid combination network are trained separately and one after another in this example (i.e., sequentially, so that the affine network is trained first, followed by the non-rigid network), but can also be trained jointly. Both the affine combination network and the non-rigid combination network can be trained with real data or synthetic data or a combination of both, i.e., they can be trained with input shapes derived from real patient data or using shapes synthesized by a generating network (such as the independent generator 200 or the slave generator 300 disclosed above) (or some combination of real data and synthetic data). In this way, the overall spatial combination network applied by the spatial combination module 120 is not limited to training with real or synthetic shapes, and in addition, is not limited to using synthetic data generated by the disclosed independent generator 200 and slave generator 300. In this example, a combination of real data and synthetic data is used to train both the affine combination and the non-rigid combination network of the spatial combination module 120.
[0087] Given a representation of {x} n=1…N In the case of meshes of all N parts in the shape component of interest (where the meshes of the N parts may be derived from real patient data, or may be synthesized using a generative shape model (e.g., an independent generator 200 or a slave generator 300), or may be a combination of real data and synthetic data), each of the N parts is passed as input to its own specific branch / sub-network in the affine combination network. In this example, the shapes of the left ventricle, right ventricle, left atrium, right atrium, and aortic root are passed as input to its specific branch / sub-network (A to E). Each sub-network in the affine combination network extracts information about the position, orientation, and shape of its corresponding input mesh, and outputs a 3D affine transformation for each part in the shape component, resulting in a 3D affine transformation denoted as {T} n=1…NIn this example, five affine transformations corresponding to the five input cardiac structures of interest (LV, RV, LA, RA, AR) are estimated by the affine combination network. Each of the five (N=5) estimated affine transformations includes eight parameters, including a translation (denoted as t=[t x ,t y ,t z ], where each component in the translation vector t represents an estimated displacement along a specific direction / axis in 3D Euclidean space), rotation (denoted as R = [Q1, Q2, Q3, Q4], where components Q1 to Q4 represent quaternions used to parameterize the rotation in 3D), and scaling (denoted as S, a scalar value that controls the global scaling estimated for each part). The affine combination network is trained in a "self-supervised manner" to obtain the estimated loss function L by making it affine The affine transformation of each of the five cardiac structures of interest is estimated by minimizing the loss function given by the following equation:
[0088]
[0089] Here, T n represents the estimated value applied to part x n , where the subscripts {n=1...N} represent each part of interest in the shape component (in this example, N=5, corresponding to five cardiac structures of interest); v n Indicates the nth part in the component (i.e. x n )'s mesh that are shared with all other (N-1) parts in the component; Represents the nodes / vertices in the meshes of all other (N-1) parts in the shape component that are shared with the mesh of the nth part. is constructed by: after applying the affine transformation estimated for each of the (N-1) parts, all (N-1) parts of the cascade component (at a given part x n In the case of ), the shared vertices in the mesh are Among them, v m represents the vertices shared between part "m" and part "n" in the component (m≠n), and T m Denotes the affine transformation estimated for the mth part in the component. Because the training of the affine combination network is driven by a loss function that depends only on the nodes / vertices shared between adjacent / neighboring parts in the component, the network is said to be trained in a "self-supervised" manner. This is in contrast to fully supervised learning, in which the ground truth affine transformation must be known a priori (i.e. in advance) to drive the training of the affine combination network.
[0090] Figure 6 An example architecture of an affine combination network according to the present disclosure is shown. In short, a corresponding input grid 601 (e.g., 601a for LV, etc.) is input to a corresponding encoder subnetwork 606, which operates to downsample the data and outputs a corresponding fully connected layer latent vector 615. The output latent vector 615 is then cascaded into a shared latent vector 616, followed by two fully connected layers - layer 1 617 and layer 2 618. Fully connected layers (also called dense layers) and multi-layer perceptrons are stacked individual neurons / neural units, where each neuron in a fully connected layer is connected to each feature of the vector input to the layer. The second fully connected layer 618 provides input to each of the three sub-branches, which handle one of the three transformations (i.e., scaling, rotation, and translation). Each of these is processed by two fully connected layers (619 / 620, 622 / 623, 625 / 626) which output the corresponding parameter sets as vectors - scaled output vector 621, rotation output vector 624, translation output vector 627. The fully connected layers that output the translation, rotation and scaling parameters of all five cardiac structures of interest here act as regression models / networks. Details of the five branches / sub-networks that can be used in the example of the affine combination network are presented in Table 7. Each branch / sub-network consists of five residual graph convolutional downsampling blocks (e.g. Figure 4 ). The number of feature channels used within each graph convolution layer in the residual graph convolution downsampling blocks (from block 1 to block 5) are: 16, 32, 32, 64, and 64. For each branch / subnetwork in the affine combination network, the grid downsampling factor used across downsampling blocks 1 to 5 is different. For each part-specific branch / subnetwork, the downsampling factors used according to the example are as follows - left ventricle branches / subnetwork: [4,4,4,6,6]; right ventricle branches / subnetwork: [4,4,4,6,6]; left atrium branches / subnetwork: [4,4,4,5,5]; right atrium branches / subnetwork: [4,4,4,5,5]; and aortic root branches / subnetwork: [4,4,4,4,4], where each number in the brackets corresponds to the downsampling factor used in the residual graph convolution downsampling block (from block 1 to 5). The output of the final residual graph convolution downsampling block in each branch / subnetwork is flattened into a vector. The resulting flattened vectors from each branch / subnetwork are then concatenated into a single long vector, which is passed as input to a sequence of two fully connected layers. These two fully connected layers extract features from the provided input and pass the output in the form of shorter vectors (in this example, vectors divided by 4) to three regression branches, each of which includes two fully connected layers. The three regression branches in turn estimate the translation, rotation, and scaling parameters (which represent the desired 3D affine transformation) of all parts in the component.
[0091] Table 7: Example network architecture of an affine combination network that can be used in the spatial combination module 120 (including five part-specific branches / sub-networks corresponding to five cardiac structures of interest).
[0092]
[0093]
[0094]
[0095] In use Figure 6 After the real and / or synthetic parts are combined in affine space by the affine combination network, the roughly combined / arranged multi-part shape components (denoted as X = {x n} n=(1…N) ) can be input to a non-rigid combination network for improvement. Here, the subscript "n" is used to represent each of the N parts in the multi-part shape component (i.e., in the current example, N=5, corresponding to the five cardiac structures of interest). The non-rigid combination network estimates localized part-by-part deformations to improve continuity or coherence at boundaries between adjacent parts by reducing any intersections / gaps between adjacent parts in the multi-part shape component that remain after the initial affine combination step. The non-rigid combination network is a network with an encoder-decoder network pair (in Figure 7 , similar to the graph convolutional neural network depicted in Figure 2 The non-rigid combination network 700 utilizes a series of graph convolution blocks (710 to 714) to extract features from a roughly combined / arranged multi-part shape component (i.e., in the current example, the entire heart shape component) and estimates the localized node-by-node / vertex displacements in the mesh 701 of each part in the component. The non-rigid combination network 700 is constructed by making the loss function L (non-rigid) Minimize the loss function trained in a self-supervised manner (as with the affine registration network) given by:
[0096]
[0097] Here, T nr represents a non-rigid transformation, i.e., a node-by-node / vertex displacement, which is applied to all parts {x n} n=(1…n) All vertices in the mesh are estimated; T n nrrepresents the nodes / vertices (denoted as v) shared between the mesh for the "nth" part of the component and the meshes of all other (N-1) parts of the component. n ) estimated displacement; v o nr Represents the deformed / transformed nodes / vertices in the meshes of all other (N-1) parts in the shape component that are shared with the mesh of the nth part. After applying the node-by-node / vertex displacements estimated for each of the (N-1) parts, the shape component is transformed by concatenating all (N-1) parts in the component (at a given part x n In the case of ) the shared vertices in the mesh are used to construct v o nr ,Right now Among them, v m represents the vertices shared between part "m" and part "n" in the component (m≠n), and T m nr denotes the non-rigid / deformable transformation estimated for the mth part in the assembly. Because the training of the non-rigid combination network is driven by a loss function that depends only on the nodes / vertices shared between adjacent / neighboring parts in the assembly, the network is said to be trained in a "self-supervised" manner. This is in contrast to fully supervised learning, where the ground truth non-rigid transformation must be known a priori (i.e. in advance) to drive the training of the affine combination network. L (non-rigid) The first item in ) is designed to iteratively refine the The estimated value of L minimizes the distance between nodes / vertices shared by each pair of adjacent structures across the component. (non-rigid) The other terms in (i.e., L laplacian and ‖T nr ‖ 1 ) are respectively designed to encourage the estimated non-rigid transformation to be smooth and sparse / localized (i.e., meaning that a large number of displacements / transformations are pushed to zero in cases where deformation is not expected, and deformations are localized only to the region of shared vertices - i.e., for displacements that are non-zero at shared vertices). laplacian The definition of L is exactly the same as before in the loss function used to train the independent generator and the dependent generator (see Equation 4 above). (non-rigid) The third term in represents the L applied to the displacements estimated for all nodes / vertices of all parts in the component. 1 norm, and is used to encourage sparsity so that the estimated deformation / non-rigid transformation is localized to the boundaries between neighboring structures in the component (i.e., to the nodes / vertices shared between parts in the component). The weights w 3 and w 4Used to balance L (non-rigid) The second and third loss terms in and are hyperparameters that are tuned empirically.
[0098] Details of the encoder-decoder network pair that can be used in the non-rigid combination network are presented in Table 8. The encoder 706 and decoder 716 networks each include five blocks, each of which learns shape features at a different spatial resolution, thereby enabling the overall non-rigid combination network 700 to learn shape features with both global context and local context across multiple grid resolutions. This is enabled by including grid downsampling and upsampling operations in each of the five blocks (710 to 714) of the encoder 706 network and the five blocks (718 to 722) of the decoder 716 network, respectively. The encoder 706 takes as input a roughly combined / arranged multi-part shape component, which in the case of the current example is an entire heart shape component including grids for the left ventricle (LV), right ventricle (RV), left atrium (LA), right atrium (RA), and aortic root (AR). The features extracted using the residual graph convolution downsampling blocks (710 to 714) in the encoder network are output to the feature embedding array 715, and the residual graph convolution upsampling blocks (718 to 722) in the decoder network are used to disentangle / expand the feature embedding array 715. The output of the final residual graph convolution upsampling block is compressed to the size of the input array / roughly arranged entire heart mesh using a Chebyshev convolution layer 723, and the node-by-node / vertex displacement of all parts in the component output 724 by the decoder is estimated using a Chebyshev convolution layer 723. Figure 4 The residual graph convolution downsampling blocks and upsampling blocks used in the encoder and decoder are depicted in . The number of feature channels used in each graph convolution layer in the residual graph convolution downsampling blocks (from block 1 to block 5) are: 16, 32, 32, 64, and 64, respectively. Similarly, the number of feature channels used in each graph convolution layer in the residual graph convolution upsampling blocks (from block 1 to block 5) are: 64, 64, 32, 32, and 16, respectively. Skip connections 717 are included between each corresponding encoder 706 and decoder 716 blocks in the network to ensure better propagation of features and gradients during training. The grid downsampling factors used across downsampling blocks 1 to 5 are [4, 4, 4, 6, 6], where each number in the brackets corresponds to the downsampling factor used in the residual graph convolution downsampling blocks (from block 1 to 5). Similarly, the upsampling factors used across upsampling blocks 1 to 5 are [6, 6, 4, 4, 4], where each number in brackets corresponds to the upsampling factor used in the residual graph convolution upsampling block (from block 1 to 5). The output of the final residual graph convolution upsampling block in the decoder is passed to a single convolutional layer that estimates the node-by-node / vertex displacements of all parts, resulting in the final combined whole heart shape component, i.e., the heart virtual chimera.
[0099] Table 8: Example network architecture for a non-rigid combination network used in the spatial combination module 120.
[0100]
[0101]
[0102] Data used to train the overall generative shape composition framework:
[0103] The disclosed generative shape modeling system (also referred to as a generative shape composition framework) for synthesizing a cohort of organ virtual chimeras (e.g., in a specific example, a cardiac virtual chimera describing a virtual whole heart shape assembly) can be trained and evaluated using data derived from data of real individuals (e.g., as available in the UK Biobank imaging database). The general data requirements and specific data and pre-processing steps used to train the disclosed generative shape composition framework employed in the current example include-
[0104] The shapes of all anatomical parts and / or organs are represented as surface or volume meshes / unstructured graphs, but may also be represented as outlines of unstructured graphs. In the current example, all cardiac structures of interest (i.e., left ventricle, right ventricle, left atrium, right atrium, and aortic root) are represented as triangular surface meshes, but other forms (e.g., polygonal meshes) may equally be used.
[0105] Before the grid / unstructured graph-based representations of the shapes of all parts and / or organs used in the training population are used to train the disclosed generative framework, these representations will have the same number of nodes and grid connectivity / graph topology. This can be achieved by co-registering all samples of a particular part to establish spatial correspondence across the training population considered for that part, i.e., for example, all left ventricle grids considered in the training population are co-registered with each other, while all right ventricle grids are co-registered with each other, and so on. The size of the training population for each part can vary, and the individuals / patients whose data are used in the part-specific training population can also vary in the degree of overlap between individuals / patients in training populations of different parts.
[0106] According to an example, a high-resolution 3D cardiac atlas grid available from a previous study (where the atlas refers to an average representation of a population of cardiac grids derived from multiple patient images) can be non-rigidly registered to contours defining cardiac chamber boundaries that were manually annotated by a cardiologist in cardiac cine-MR images for each of the four cardiac chambers of interest. These manual contours are available in the UK Biobank database. By registering the high-resolution 3D cardiac atlas grid to the manual contours, a training population of shapes for each part / cardiac structure of interest can be generated (i.e., since the same atlas grid is deformed / non-rigidly registered to fit the manual contours of all individuals considered from the UK Biobank, spatial correspondence is automatically established across all samples in the part-specific training population). The high-resolution cardiac atlas used to generate the training population includes the following parts / structures (each represented by a different grid) as parts of a multi-part shape assembly - left and right ventricles (LV and RV), left and right atria (LA and RA), and aortic vascular root (AR). In addition, the atlas includes vertices shared between meshes of adjacent structures (i.e., the boundaries separating any adjacent structure pairs in the entire heart shape component). The vertices in the atlas mesh are shared at the boundaries between the following structure pairs - LV-RV, LV-LA, LA-RA, RV-RA, LV-AR, LA-AR, and RA-AR. The availability of these vertices shared between adjacent parts / structures helps enable self-supervised learning schemes for training the disclosed spatial combination networks (i.e., both affine spatial combination networks and non-rigid spatial combination networks). Therefore, the disclosed formulation of the generative shape combination framework uses information about vertices shared between adjacent parts in multi-part and / or multi-organ shape components, where the shared vertex information can be created manually / semi-automatically / automatically, or is available before training the spatial combination network in the framework. Although the training population of the manual contours and thus generated part-specific meshes (based on the deformed / registered atlas meshes) is defined on cardiac movie MR images of individuals in the UK Biobank, the disclosed method does not depend on the part-specific mesh training population extracted / inferred only from cardiac movie MR images. The shape of any part and / or organ of interest may be obtained from different imaging modalities and also from different patient populations. For example, the left ventricle mesh may be extracted / inferred from cardiac cine MR images of patients / individuals in population "A", while the aortic root mesh may be extracted / inferred from cardiac computed tomography angiography (CTA) images of patients / individuals in population "B".In addition, the patients / individuals included in groups "A" and "B" can be - (a) non-overlapping; (b) partially overlapping; or (c) completely overlapping; wherein scenario (a) means that no patients in group "A" are included in group "B"; scenario (b) means that a certain proportion of patients in group "A" are included in group "B"; scenario (c) means that all patients included in group "A" are included in group "B", and vice versa. However, it is important to note that scenario (a) for training groups of each independent part / structure extracted / inferred from images of completely different patients / patient groups is only suitable for use with independent generators 200 within the disclosed generative shape combination framework, that is, while the independent generators are designed to allow training with non-overlapping data (as well as partially overlapping and fully overlapping data), the subordinate generators can only be trained with partially overlapping or fully overlapping data.
[0107] According to an example, the overall cohort for training and testing the disclosed generative shape combination framework includes 2360 subject-specific whole heart shape components (represented as meshes). These subject-specific meshes are created by registering high-resolution heart atlas meshes to individually available manually annotated heart contours in the UK Biobank database. These 2360 subjects are selected from 4000 subjects in the UK Biobank, and manually annotated heart contours are available from these subjects. The selection is based on a visual assessment of the quality of the subject-specific meshes obtained from the registration process (initial execution for all 4000 subjects). Only those meshes that do not introduce obvious topological errors after registering the atlas mesh to the manual contour are retained, thereby obtaining a selected cohort of 2360 subject-specific meshes. The resulting cohort of 2360 subject-specific meshes shares node-by-node / vertex spatial correspondence across all parts / heart structures in the shape component, that is, all meshes have the same number of vertices / nodes, the same mesh connectivity / graph topology, and each node / vertex across all meshes represents the same anatomical feature / position. Having the same mesh connectivity / unstructured graph topology means that the edges connecting each node to neighboring nodes in the mesh / unstructured graph are the same across all subject-specific meshes in the cohort. Another way to define having the same mesh connectivity / unstructured graph topology is to say that the adjacent matrices of all subject-specific meshes / unstructured graphs are the same. The same graph topology across all subject-specific meshes enables the use of spectral (i.e., based on Chebyshev polynomials) convolutions in the graph convolution layers used throughout the disclosed generative shape combination framework, because the fixed / same graph topology across all inputs is a prerequisite for graph neural networks that utilize spectral convolution operations in their constituent graph convolution layers. The developed system can be adapted to adapt to meshes that are not identical in terms of their vertex connectivity / graph topology by replacing spectral convolution operations with spatial graph convolution operations. A cohort of 2360 subject-specific meshes / entire heart shape components was randomly split into training, validation, and test sets that each included 422, 59, and 1879 samples, respectively. The goal of the split can be to have a test set that is significantly larger than the training set to ensure that the evaluation is done on a more diverse population than the one used for training.
[0108] The disclosed generative shape combination framework can be trained with non-overlapping, partially overlapping, or fully overlapping data as previously defined. However, in the current example, the training of the disclosed generative shape combination framework is performed using partially overlapping and fully overlapping data derived from cardiac movie MR images available in the UK Biobank. In order to fairly compare the disclosed independent generator with the slave generator, only partially overlapping scenarios and fully overlapping scenarios are implemented (because, unlike the independent generator, the slave generator cannot be trained with non-overlapping data). For fully overlapping scenarios, the data of all subjects included in the training set contain all cardiac structures / parts (i.e., LV, RV, LA, RA, and AR) in the entire heart shape assembly, and a complete data set including 422 training samples (each with five parts / heart structures) is used to train both independent generators and slave generators. For partially overlapping scenarios, 300 subjects in the training set are considered to have "missing data" because only one part / heart structure from each subject in the training set is included (i.e., this is done to simulate real scenarios where only partially overlapping data across subjects is available). The samples in the training set for the partially overlapping scenario included 60 parts / heart structures for each of the five cardiac structures of interest obtained from different subjects (thus resulting in 300 training samples, each of the different parts / heart structures obtained from different subjects). The remaining 122 subjects in the training set were considered to have complete whole heart shape components (i.e., all five parts / heart structures were available). This partially overlapping dataset was also used to train both the independent generator and the dependent generator.
[0109] The spatial combination network (including both affine spatial combination networks and non-rigid spatial combination networks) is trained using both real data and synthetic data. First, the spatial combination network is pre-trained using real data to maintain the data segmentation for training, verification and testing identical to the partial perception generation module trained in the "complete overlap" scene (i.e. 422 samples are used for training, 59 samples are used for verification, and 1879 samples are used for testing). After initial training using real data, the spatial combination network is fine-tuned (i.e., continuing training) using synthetic data generated using independent generator 200 (i.e., using the part / structure synthesized using the gcVAE specific to each part in the independent generator 200). A total of 2000 synthetic samples of each part / heart structure of interest are included in the synthetic data queue. Among these, 1600 are used for training, 200 are used for verification and 200 are used for testing.
[0110] Figure 8A method 800 of training a partially perceptual model according to a disclosed example is shown. The method begins at 801 by using certain non-overlapping or partially overlapping or fully overlapping anatomical shape / mesh data (from multiple subjects / patients) as input (i.e., training data) to a partially perceptual generative model 810. The partially perceptual generative model is trained 810 to learn hidden / latent representations of the input data by minimizing a suitable loss function using numerical optimization and learning / updating the parameters of the partially perceptual generative model using an error back-propagation algorithm. The training process is iterative, wherein the parameters of the partially perceptual generative model are updated at the end of each iteration / epoch using the error / loss incurred by the partially perceptual generative model at each iteration (also referred to as an epoch). Training continues until a suitable convergence criterion is reached, or a user-specified maximum number of training iterations / epochs is reached 820, at which point the training loop stops / terminates 830 and the learned parameters of the partially perceptual generative model are saved.
[0111] Fig. 9 A method 900 for training a spatial combination model according to a disclosed example is shown. The method starts at 901 using training data, wherein each training sample includes all parts of the shape component of interest as input. The components in each training sample used to start training 910 the spatial combination model can belong to a real patient / subject, or can be a synthetic / virtual part synthesized using a suitable statistical / machine learning model (e.g., an independent generator 200 or a slave generator 300 in the disclosed system). The spatial combination model is trained 910 to learn the spatial relationship between the input components of the shape component and estimate the spatial transformation required to combine / arrange (i.e., put together) the parts into an anatomically reasonable / realistic shape component. The training process is iterative and involves minimizing a suitable loss function using numerical optimization and using an error back propagation algorithm to learn / update the parameters of the spatial combination model. During the iterative training process, the parameters of the spatial combination model are updated at the end of each iteration / epoch using the error / loss incurred by the spatial combination model at each iteration (also referred to as an epoch). Training continues until a suitable convergence criterion is reached or a user-specified maximum number of training iterations / epochs is reached 920, at which point the training loop stops / terminates 930 and the parameters of the learned spatial combination model are saved.
[0112] Although Figure 8 and Fig. 9 The training is shown to be done independently, but the training can also be performed sequentially, i.e., the partial perception models are trained in one cycle, and then the spatial combination model is trained, and this is repeated over multiple epochs.
[0113] Fig.10A method 1000 of generating a virtual chimera population for use in a computer simulation experiment according to the disclosed example is shown. According to this example (which can be a computer-implemented method) that an organ shape can be represented as a surface or volume mesh, the method starts at step 1001 and then includes the following steps: 1010-generating organ parts using a part-aware generative model that has been trained, for example, on non-overlapping or partially overlapping (or fully overlapping) patient data (so as to learn a latent representation of each part of the multi-part organ shape to be included in the virtual population) and is arranged to output composite parts of the multi-part organ shape; 1020 using a spatial combination model that has been trained, for example, on non-overlapping and partially overlapping real patient data or synthetic / virtual patient data to arrange the previously output composite parts (to arrange the output composite parts of the multi-part organ shape), and outputting an anatomically meaningful overall multi-part organ shape example as a virtual chimera; and 1030-storing the resulting virtual chimera in a collection of combined chimeras, the so-called virtual chimera population. The resulting complete virtual chimera population may then be exported 1040 for a number of subsequent uses - for example, for designing or testing medical devices, particularly for use in computer simulation experiments.
[0114] Generating an anatomical and physiological virtual population as disclosed herein is important for computer simulation experiments because it provides a systematic and comprehensive framework for studying and evaluating the performance of medical devices in computer simulations. A virtual population can be a collection of samples (e.g., called virtual patients) representing reasonable instances of anatomy and / or physiology, which can be found in reality in a target patient population for a medical device under test, but does not necessarily represent data of any real patient. As disclosed herein, statistical or machine learning methods can be used to synthesize virtual populations, wherein machine learning methods are initially trained on data of real patients to learn the underlying hidden (or potential) distribution of the data. Then, a trained statistical / machine learning model can be used to create new data, i.e., new virtual patients, by sampling from the learned potential distribution. The resulting virtual population created by sampling from such a trained statistical / machine learning model can be referred to as an anatomical and / or physiological virtual population (depending on the type of real patient data used to train the statistical / machine learning model).
[0115] Existing methods for synthesizing virtual patient populations require that the data used to construct the models fully overlap, thus limiting their ability to exploit disparate large datasets (particularly, datasets with non-overlapping or partially overlapping information). The inability to exploit (and thereby maximize the value of) non-overlapping / partially overlapping data also limits the variability in overall anatomy and physiology captured in virtual patient populations synthesized using existing methods.
[0116] The virtual chimera / model output by the disclosed system is a complete anatomically realistic model of the physical structure of the corresponding organ or organ collection, and as such, can be used for any purpose that utilizes such a model, in particular for computer simulation testing of medical devices for the purpose of developing, designing and manufacturing medical devices and subsequent medical testing thereof. For example, a heart-specific virtual chimera population generated / created using the disclosed system can be used to conduct computer simulation testing of all implantable cardiac devices, such as transcatheter aortic valve implant devices, mitral valve replacement devices, cardioverter-defibrillators, pacemakers, and coronary stents. Computer simulation studies that utilize computational simulation and modeling to study the effects of drugs, genetics, lifestyle, environmental, or demographic factors on biochemical and / or biophysical processes in the heart can also be used for virtual populations generated using the disclosed system. The generated virtual population can also be used to conduct virtual imaging experiments and develop and calibrate non-implantable medical devices, such as: medical imaging systems (e.g., intravascular ultrasound, computed tomography, computed tomography angiography, magnetic resonance imaging, positron emission tomography, single photon emission computed tomography, cardiac ultrasound / echocardiography, and 3D ultrasound); interventional robotics; surgical guidance systems; catheters, guidewires, etc. The generated virtual population can also be used to 3D print realistic physical heart models / phantoms for experimental research, development of surgical guidance systems, surgical training, and educational purposes.
[0117] As can now be appreciated, some methods for providing virtual chimeras use "strong labels" within a self-supervised or fully supervised learning framework for training a neural network for creating the virtual chimera. "Strong labels" refer to reference / ground truth spatial transformations and rely on the availability of real and complete multi-part / shape components. In this way, using the availability of reference / ground truth spatial transformations, neural networks can be trained in a fully supervised manner to predict the spatial transformations necessary to combine independent parts into complete multi-part shape components. Access to the reference / ground truth spatial transformations requires a priori availability / knowledge of the corresponding real and complete multi-part shape components. However, training neural networks using strong labels excludes the use of partially overlapping (or even non-overlapping) data. Therefore, the present disclosure provides a novel self-supervised method for training a combination network with available weak labels, wherein weak labels refer to information shared between adjacent parts in real and incomplete shape components (i.e., including partially overlapping or non-overlapping data).
[0118] Examples of the present disclosure may be implemented by appropriately programmed computer hardware. Fig.111100 is a block diagram illustrating components according to some example embodiments that are capable of reading instructions from a machine-readable or computer-readable medium (e.g., a non-transitory machine-readable storage medium) and performing any one or more of the methods discussed herein, thereby providing an apparatus or system for performing the described virtual population modeling. Specifically, Fig.11 A diagrammatic representation of hardware resources 1105 is shown including one or more processors (or processor cores) 1110, one or more memory / storage devices 1120, and one or more communication resources 1130, each of which may be communicatively coupled via a bus 1140. Processor 1110 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP) such as a baseband processor, an application specific integrated circuit (ASIC), a cloud processing function such as an AWS instance, another processor, or any suitable combination thereof) may include, for example, processor 1112 and processor 1114.
[0119] The memory / storage device 1120 may include a main memory, a disk storage device, or any suitable combination thereof. The memory / storage device 1120 may include, but is not limited to, any type of volatile or non-volatile memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, solid-state storage device (SSD), magnetic storage device-based hard disk drive (HDD) media, etc.
[0120] The communication resources 1130 may include an interconnect or network interface component or other suitable means to communicate with one or more peripheral devices 1104 or one or more databases 1106 via the network 1108. For example, the communication resources 1130 may include a wired communication component (e.g., for coupling via Ethernet, a universal serial bus (USB) or the like), a cellular communication component, an NFC component, Components (e.g. Low power consumption), components and other communication components.
[0121] The instructions 1150 may include software, programs, applications, applets, applications, or other executable code for causing at least any one of the processors 1110 to perform any one or more of the methods discussed herein. The instructions 1150 may reside, in whole or in part, in at least one of the processors 1110 (e.g., within a cache memory of the processor), the memory / storage device 1120, or any suitable combination thereof. In addition, any portion of the instructions 1150 may be transmitted to the hardware resources 1105 from any combination of the peripheral device 1104 or the database 1106. Thus, the memory of the processor 1110, the memory / storage device 1120, the peripheral device 1104, and the database 1106 are examples of computer-readable and machine-readable media.
[0122] In some embodiments, Fig.11 Or an electronic device, network, system, chip or component of some other figure in this document, or a portion or implementation thereof may be configured to perform one or more processes, techniques or methods, or portions thereof, as described herein.
[0123] In a specific example, the disclosed generative shape combination framework can be implemented on a PC with an NVIDIA RTX 2080Ti graphics processing unit (GPU) using PyTorch, PyTorch Geometric, and PyTorch3D python libraries.
[0124] In the example, the Adam optimizer is used to train the part-aware generation module (i.e. both the independent generator and the slave generator) with an initial learning rate of 5e -04 And the learning rate decays to 0.9 for each training epoch. All networks in the part-aware generation module are trained with a batch size of 16 for a maximum of 100 training epochs. Both the affine spatial combination network and the non-rigid spatial combination network are trained with a batch size of 8 for a maximum of 100 training epochs. The initial learning rate used for both spatial combination networks is le -04 , where the learning rate decays to 0.9 per epoch. All hyperparameters associated with the network architecture, training routine, and the optimizer used were empirically tuned / set based on the average loss values on the validation set observed in preliminary prototype experiments.
[0125] When training both the independent generator 200 and the slave generator 300, a warm-up strategy can be employed to improve stability and prevent mode collapse in the learned posterior distribution over the latent variables. This is achieved by initially training a part-by-part gcVAE in the independent generator and a gcmcVAE in the slave generator, where the weights of the KL loss term (i.e., ω0 in Eq. 1) are set to zero for 200 epochs (they are trained as ordinary autoencoders). The learned weights initialize subsequent training steps for the independent and slave generators, where the weights of the KL loss term are initially set to a small value, i.e., 1e -6 , and then for example by multiplying by a factor of 1.25 up to a maximum value of 1e -4 The other weights (i.e., ω1, ω2, ω3, and ω4) in the regularization term of the composite loss function (see Equation 3) and the self-supervised registration loss function are set to 8, 1, 8, and 5e, respectively. -7 .
[0126] In addition to setting the initial learning rate to 1e -4 Except for setting the batch size to 8, the same network architecture and optimized parameter settings as the part-aware generator can be retained for the combined network. During the training of the combined network, the shapes of the parts can be randomly sampled from different patient data. The combined network can be pre-trained with the original shapes in the real population for 100 epochs, and then it can be trained on the previously generated shapes. The generated dataset can be randomly split into parts for each of training, validation and testing (e.g., 1600 / 200 / 200 splits). It is worth noting that the combined network is trained only with partially overlapping data, i.e., structures obtained from different patients.
[0127] Above, functions are described as modules or blocks, i.e., functional units that are operable to perform the described functions, algorithms, etc. These terms may be interchangeable. Where modules, blocks, or functional units have been described, they may be formed as processing circuit systems, wherein the circuit systems may be general-purpose processor circuit systems that are configured to perform the specified processing functions by program code. The circuit systems may also be configured by modifying the processing hardware. The configuration of the circuit systems for performing the specified functions may be entirely in hardware, entirely in software, or using a combination of hardware modification and software execution. Program instructions may be used to configure the logic gates of a general-purpose or special-purpose processor circuit system to perform the processing functions.
[0128] For example, the circuit system can be implemented as a hardware circuit that includes a custom very large scale integration, VLSI, circuit or gate array, off-the-shelf semiconductors such as logic chips, transistors or other discrete components. The circuit system can also be implemented in a programmable hardware device (such as a field programmable gate array, FPGA, programmable array logic, programmable logic device, system on chip, SoC, etc.).
[0129] Machine readable program instructions may be provided on a temporary medium (such as a transmission medium) or on a non-temporary medium (such as a storage medium). Such machine readable instructions (computer program code) may be implemented in a high-level procedural or object-oriented programming language. However, if desired, the program may be implemented in assembly language or machine language. In any case, the language may be a compiled or interpreted language and combined with a hardware implementation. The program instructions may be executed on a single processor or in a distributed manner on two or more processors.
[0130] An example provides a computer-implemented method for generating a population of virtual chimeras of multi-part organ shapes for use in a computer simulation study (e.g., including a computer simulation experiment or a computer simulation observation study), wherein the organ shape is represented as a surface or volume mesh, the method comprising: using a part-aware generative model to learn a latent representation of each part of the multi-part organ shape to be included in the virtual population, and outputting a composite part of the multi-part organ shape; using a spatial composition model to arrange the output composite parts of the multi-part organ shape, and outputting an anatomically meaningful example of the overall multi-part organ shape as a virtual chimera; and storing the virtual chimera in the virtual population for use in the computer simulation experiment. The composite part may include one or more meshes. The part-aware generative model may include independent generators to learn the latent representation of each part of the multi-part organ shape. The independent generators may include a graph convolutional variational autoencoder (gcVAE) network architecture that is operable to learn the variability of each part of the multi-part organ shape that can be observed across a training population. The independent generators can learn the variability of each part of the multi-part organ shape that can be observed across a training population independently of each other. For example, by making the overall loss function L relative to the parameters of its constituent encoder and decoder networks respectively totalMinimization, each of the gcVAEs of the independent generators can be trained independently of each other. The parameters of the encoder-decoder network pair in each gcVAE can be iteratively learned, for example, via an error back propagation algorithm. The part-aware generative model may include a slave generator to learn the shared potential representation of all parts of the multi-part organ shape. Each of the independent generator and the slave generator may include at least one encoder and at least one decoder. In the example of a slave generator using multiple encoder-decoder pairs, the slave generator may also include at least one multiplexing layer between the encoder and the decoder, wherein the multiplexing layer is operable to provide the output of each of the encoders in use to each of the decoders in use. The slave generator may include a graph convolutional multi-channel variational autoencoder (gcmcVAE) network architecture, which is operable to learn the joint variability of the shape of each part of the multi-part organ shape that can be observed across the training population. The joint variability can be the combined variability observed across different organ parts or organs in the multi-part organ. The multi-part organ shape may include a multi-organ shape component. The disclosed method can be nested to construct high-order components from low-order organs - that is, multiple single multi-part organ shapes can be modeled and then combined into an overall multi-organ component until the entire complete virtual patient is included. For example, the heart, liver, and kidneys can all be modeled and combined into a heart-liver-kidney component. The spatial combination model can also include at least one of an affine combination network and a non-rigid combination network. According to an example, each of the gcVAE network; the gcmcVAE network; the affine combination network; and the non-rigid combination network can also include at least one residual graph convolution downsampling function block and at least one residual graph convolution upsampling function block. The at least one residual graph convolution downsampling function block and / or the at least one residual graph convolution upsampling function block can include at least one Chebyshev graph convolution function, at least one exponential linear unit function, and at least one instance normalization function. The at least one residual graph convolution downsampling function block can include a grid pooling layer function, and / or the at least one residual graph convolution upsampling function block can include a grid upsampling layer function. The affine combination network can include at least one of the following: a scaling branch, a rotation branch, and a translation branch. The affine combination network may also include a cascade layer, and / or at least one fully connected layer (also referred to as a dense layer or a single / multilayer perceptron). The non-rigid combination network may also include a feature embedding array and / or a Chebyshev convolution layer. The disclosed example method may also include providing non-overlapping patient data or partially overlapping patient data during training. The non-overlapping patient data or partially overlapping patient data may include shape data extracted from imaging data of the same or different patients, where the imaging data may be acquired using the same or different imaging modalities.That is, according to different examples, the overlapping nature of the data can be defined relative to the coverage of a specific patient source or original data source (i.e., modality). In some examples, the anatomical coverage can be the result of using a modality. Different imaging modalities are capable of detecting different parameters of the patient or organ under study, so that a complete patient can be modeled by combining data from more than one modality. For example, even if the modalities are viewing the same field of view, they can detect different aspects of each part of the patient within the same field of view. By way of specific example, a first modality may be able to image (i.e., "see") muscle structures, and another modality may be able to image vascular structures. In other words, by enabling a mixture of patient imaging modalities, a more complete and accurate input training data set can be derived and used for training the disclosed generation system 100, so that the resulting synthetic variable virtual chimera is also more complete and accurate (while still being variable within the limits derived from those more accurate training). The different imaging modalities may include any one or more of the following: magnetic resonance imaging (including cross-sectional, angiographic, or any variant of structural, functional, or molecular imaging), computed tomography (including cross-sectional, angiographic, or any variant of structural, or functional imaging), ultrasound (including any variant of structural or functional imaging), positron emission tomography, single photon emission tomography. Each of these imaging techniques may be enhanced with appropriate endogenous or exogenous contrast agents to highlight the anatomical or functional structures of interest.
[0131] The example provides a method of training a generative shape modeling system as disclosed herein (e.g., including a partially perceptual generative model and a spatial combination model), the method comprising iteratively providing non-overlapping and / or partially overlapping patient data to the generative shape modeling system over a predefined number of training periods. The method may also include generating (i.e., deriving) a virtual chimera population using the trained generative shape modeling system. The disclosed example method may also include performing a computer simulation design or testing of a medical device using the generated virtual chimera population.
[0132] Examples also provide a computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform any of the described methods or portions thereof (e.g., training methods, and / or methods of generating virtual chimeras using a trained system).
[0133] Examples also provide an apparatus arranged or configured to perform any of the described methods.
[0134] The examples also provide generated output synthetic data sets, called virtual chimeras or chimera populations, for use in computer simulation studies generated (ie, derived and synthesized) according to the disclosed examples.
[0135] The example also provides a method for designing or testing a medical device, the method comprising: training a generative shape modeling system using non-overlapping or partially overlapping patient data (optionally testing the trained system); generating new instances of virtual chimeras based on the trained generative shape modeling system; combining the new instances of virtual chimeras into a virtual chimera population; designing a medical device using the virtual chimera population or testing a medical device using the virtual chimera population. Thus, the examples of the present application provide a way to generate a virtual patient population (e.g., for IST-based medical device testing) that captures sufficient anatomical and physiological variability to represent a myriad of different potential target patient populations, thereby enabling a meaningful evaluation of the performance of the corresponding medical device.
[0136] Examples according to the present disclosure may be used in any computer simulation study, including but not limited to computer simulation experiments and observational studies, for example, for developing medical devices intended for use in or on a patient (such as implants), but also for drug development and general medical devices (including, for example, digital health products, AI-based medical software, etc.). In other words, the examples may be used in any kind of computer simulation study.
[0137] The independent generator 200 and the slave generator 300 in the disclosed system are different in design and thus have different properties and advantages relative to each other. The slave generator can capture the nonlinear correlation between the input parts in the shape component by learning a shared potential representation across all parts. Therefore, the slave generator 300 can generate a virtual part (or synthetic part) with better specificity / anatomical rationality than the independent generator 200. Since the slave generator is trained to learn a shared potential representation that explains the shape variability observed across all parts in the shape component, the shared potential representation learned in the slave generator 300 is more restricted than the independent potential representation learned in the independent generator 200. In other words, the shared potential representation learned in the slave generator 300 is restricted to capturing the correlation between the parts in the shape combination and must be able to explain the joint changes or common changes in the shapes of all parts in the shape component. However, the potential representation learned in the independent generator 200 has no such limitation, that is, the potential representations learned for each part are completely independent of each other and only focus on explaining the shape changes of its specific part. Therefore, the independent generator 200 has greater flexibility than the slave generator, and in this way, a greater degree of shape variation of the part of interest in the shape combination can be captured. In other words, the slave generator 300 can generate a virtual chimera queue that is more anatomically reasonable / realistic than the independent generator 200, while the independent generator 200 can generate a virtual chimera queue containing more diverse shapes / capturing a greater degree of variability in the shape (of individual parts and their combinations) than provided by the slave generator 300. When the application of interest requires a higher level of statistical fidelity, the slave generator 300 can be used to replace the independent generator 200. For example, in the context of conducting a computer simulation experiment, if the experimental design includes very specific inclusion and exclusion criteria to determine the virtual queue to be used in the computer simulation experiment to evaluate the performance of a medical device or drug, the slave generator 300 will be a better choice. When the application of interest requires the virtual population to have more diversity, the independent generator 200 can be used to replace the slave generator 300. For example, in the context of an in silico trial, if the purpose of the trial is to explore the performance of a medical device or drug in a non-standard or niche patient population that is less commonly / frequently observed in the real world, the independent generator 200 would be a better choice.
[0138] Although preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Without departing from the scope of the present disclosure, those skilled in the art will now expect numerous changes, modifications and substitutions. It should be understood that the various alternatives of the embodiments of the present invention described herein can be used in any combination when practice is disclosed. The following claims are intended to define the scope of the present invention, and thus encompass methods and structures and their equivalents within the scope of these claims.
Claims
1. A computer-implemented method for generating a virtual chimera population of multi-part organ shapes for use in computer simulation studies, the method comprising: include: using a part-aware generative model to learn a latent representation for each part of a multi-part organ shape to be included in a virtual population, and outputting a synthesized part of the multi-part organ shape; arranging the composite parts of the output multi-part organ shape using a spatial combination model, and outputting an anatomically meaningful example of the overall multi-part organ shape as a virtual chimera; as well as The virtual chimeras are stored in the virtual population for use in computer simulation experiments.
2. The method according to claim 1, in, The part-aware generative model includes an independent generator to learn a latent representation for each part of the multi-part organ shape.
3. The method according to claim 2, in, The independent generator comprises a graph convolutional variational autoencoder (gcVAE) network architecture operable to learn the variability of each part of the multi-part organ shape observed across a training population.
4. The method according to claim 1, in, The part-aware generative model also includes a slave generator to learn a shared latent representation of all parts of the multi-part organ shape.
5. The method according to claim 4, in, The slave generator comprises a graph convolutional multi-channel variational autoencoder (gcmcVAE) network architecture operable to learn the joint variability of the shape of each part of the multi-part organ shape observable across a training population.
6. The method according to any one of claims 1 to 5, in, The spatial combination model also includes at least one of an affine combination network and a non-rigid combination network.
7. The method according to any one of claims 3 to 6, in, The gcVAE network, gcmcVAE network, affine combination network, and non-rigid combination network all include at least one residual graph convolution downsampling functional block and at least one residual graph convolution upsampling functional block.
8. The method according to claim 7, in, The at least one residual graph convolution downsampling functional block or the at least one residual graph convolution upsampling functional block includes at least one Chebyshev graph convolution function, an exponential linear unit function, or an instance normalization function.
9. The method according to claim 7 or 8, in, The at least one residual graph convolution downsampling functional block comprises a grid pooling layer function, and / or the at least one residual graph convolution upsampling functional block comprises a grid upsampling layer function.
10. The method according to any one of claims 6 to 9, in, The affine combination network includes at least one of the following: a scaling branch, a rotation branch, and a translation branch.
11. The method according to any one of claims 5 to 10, in, Non-rigid combination networks also include feature embedding arrays or Chebyshev convolutional layers.
12. A method according to any preceding claim, in, The method includes providing non-overlapping patient data or partially overlapping patient data during training.
13. The method according to claim 12, in, The non-overlapping patient data or the partially overlapping patient data includes shape data extracted from imaging data of the same or different patients, wherein the imaging data is acquired using the same or different imaging modalities.
14. The method according to claim 13, in, The different imaging modalities include magnetic resonance imaging, computed tomography, computed tomography angiography, ultrasound, positron emission tomography, single photon emission tomography.
15. The method according to any of the preceding claims further comprises using the generated virtual population to perform computer simulation design, testing or regulatory approval of a medical device (including but not limited to implants, digital health or artificial intelligence devices) or a drug.