Probabilistic variant interpretation
Patent Information
- Application Number
- JP2026513086
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-07
- Filing Date
- 2024-08-30
- Publication Date
- 2026-09-03
AI Technical Summary
)である場合、それらのそれぞれの予測間のミスマッチ(例えば、健康な集団コホートにおける高い対立遺伝子頻度対罹患コホートにおける疾患表現型との強い相関)は、病原性予測における不確実性を増加させるはずである。相関モデルは、上流変数と下流変数との間のこれらの種類の区別を行うことができない場合がある。更に、相関モデリングアプローチは、生物学的に無意味であり、より広いデータセットに一般化することができないモデルを生成する可能性がある。対照的に、因果的確率的グラフィカルモデルは、変数間の因果関係を定義するためにドメイン知識に依存する。
Smart Images

Figure 2026530012000001_ABST
Abstract
Description
[Technical Field]
[0001] [Cross-Reference to Related Applications] This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 579,939 filed on August 31, 2023 and U.S. Provisional Patent Application No. 63 / 562,696 filed on March 7, 2024, and each of the U.S. provisional patent applications is incorporated herein by reference in its entirety.
[0002] The technical field to which the present application relates is genetic testing. Another technical field to which the present application relates is machine learning-based systems for genetic variant classification or variant interpretation. [Background Art]
[0003] Genetic variants are differences in DNA sequences between individuals in a population. There are many different types of variants including, but not limited to, structural variations, single nucleotide polymorphisms, insertion and deletion mutations, copy number variations, and translocations and inversions.
[0004] Gene sequencing technology continues to evolve rapidly. High-throughput sequencing technology increasingly enables genetic testing covering genotyping of genetic diseases, single-gene testing, gene panel testing, exome testing, genome testing, transcriptome testing and epigenetic assays. The increasing complexity of analysis and interpretation of clinical genetic testing and the increasing volume of testing have brought new challenges in the interpretation of genetic variants. For example, clinical molecular laboratories are increasingly detecting novel variants in the process of testing patient specimens, where the number of disease-associated genes is rapidly increasing. While some phenotypes are associated with a single gene, many are associated with multiple genes.
[0005] The continued expansion of gene sequencing into more areas of clinical medicine, facilitated by cost reductions and the expansion of clinical guidelines, has resulted in an explosion in the number of rare variants observed through testing for genetic disorders. Despite a corresponding increase in the amount of available data that can potentially be used to interpret these variants, a high proportion of variants identified in genetic testing (e.g., about half) are classified as variants of unknown significance (VUS). [Brief explanation of the drawing]
[0006] This disclosure will be better understood from the following detailed description and the accompanying drawings of various embodiments of this disclosure. The drawings are for illustrative and illustrative purposes only and should not be construed as limiting this disclosure to the specific embodiments shown.
[0007] [Figure 1A] Examples of machine learning-based processes for variant interpretation are shown according to several embodiments of this disclosure.
[0008] [Figure 1B] Examples of causal models are shown in some embodiments of this disclosure.
[0009] [Figure 2A] Examples of probabilistic models for variant interpretation are shown according to several embodiments of this disclosure. [Figure 2B] Examples of probabilistic models for variant interpretation are shown according to several embodiments of this disclosure. [Figure 2C] Examples of probabilistic models for variant interpretation are shown according to several embodiments of this disclosure. [Figure 2D] Examples of probabilistic models for variant interpretation are shown according to several embodiments of this disclosure. [Figure 2E] Examples of probabilistic models for variant interpretation are shown according to several embodiments of this disclosure. [Figure 2F]Examples of probabilistic models for variant interpretation are shown according to several embodiments of this disclosure. [Figure 2G] Examples of probabilistic models for variant interpretation are shown according to several embodiments of this disclosure.
[0010] [Figure 3] Examples of component-based processes for creating causal models for variant interpretation using components of a modeling system are provided by some embodiments of this disclosure.
[0011] [Figure 4] This disclosure illustrates exemplary processes for training and validating probabilistic models for variant interpretation, according to several embodiments of this disclosure.
[0012] [Figure 5A] Examples of variant interpretation systems according to several embodiments of this disclosure are shown.
[0013] [Figure 5B] Examples of component-based processes for variant interpretation systems are shown according to several embodiments of this disclosure.
[0014] [Figure 6] Examples of evidence-based variant interpretation platforms are shown in several embodiments of this disclosure.
[0015] [Figure 7] This table lists examples of models that may be included in an evidence-based variant interpretation platform according to some embodiments of the present disclosure.
[0016] [Figure 8A] Examples of experimental results for probabilistic variant interpretation models according to several embodiments of this disclosure are shown. [Figure 8B] Examples of experimental results for probabilistic variant interpretation models according to several embodiments of this disclosure are shown. [Figure 8C] An example of experimental results of a probabilistic variant interpretation model according to some embodiments of the present disclosure is shown. [Figure 8D] An example of experimental results of a probabilistic variant interpretation model according to some embodiments of the present disclosure is shown.
[0017] [Figure 9A] A method, system, apparatus, and / or non-transitory computer-readable medium configured to generate a probabilistic model and use the probabilistic model for variant interpretation according to some embodiments of the present disclosure is shown. [Figure 9B] A method, system, apparatus, and / or non-transitory computer-readable medium configured to generate a probabilistic model and use the probabilistic model for variant interpretation according to some embodiments of the present disclosure is shown. [Figure 9C] A method, system, apparatus, and / or non-transitory computer-readable medium configured to generate a probabilistic model and use the probabilistic model for variant interpretation according to some embodiments of the present disclosure is shown. [Figure 9D] A method, system, apparatus, and / or non-transitory computer-readable medium configured to generate a probabilistic model and use the probabilistic model for variant interpretation according to some embodiments of the present disclosure is shown. [Figure 9E] A method, system, apparatus, and / or non-transitory computer-readable medium configured to generate a probabilistic model and use the probabilistic model for variant interpretation according to some embodiments of the present disclosure is shown. [Figure 9F] A method, system, apparatus, and / or non-transitory computer-readable medium configured to generate a probabilistic model and use the probabilistic model for variant interpretation according to some embodiments of the present disclosure is shown. [Figure 9G] A method, system, apparatus, and / or non-transitory computer-readable medium configured to generate a probabilistic model and use the probabilistic model for variant interpretation according to some embodiments of the present disclosure is shown.
[0018] [Figure 10] This disclosure presents exemplary computing systems, including probabilistic variant interpretation models, according to several embodiments of this disclosure.
[0019] [Figure 11] This aspect of the disclosure is a block diagram of an exemplary computer system in which it can operate. [Modes for carrying out the invention]
[0020] In some fields, variant classification or variant interpretation may interchangeably refer to the process of interpreting information about a genetic variant based on evidence supporting or rejecting a causal relationship with a given disease. In some fields, variant classification may refer to the process of classifying variants based on a combination of evidence, and variant interpretation may refer to the use of variant classification information to make a diagnostic decision. In some fields, variant interpretation may refer to the process of evaluating variant information to apply classifications and notify one or more clinicians to make a diagnostic decision. The amount and significance of any given evidence can vary depending on the variant, the gene, the disease, and / or its relationship to other evidence. Currently, variant interpretation does not provide a diagnosis on its own, but may be used by clinicians to make a diagnostic decision.
[0021] The primary objective of clinical genetic testing is to identify gene variants in an individual and, for each variant, determine whether it is pathogenic (potentially capable of causing disease) or benign (potentially not capable of causing disease). Currently, the problem is that many variants cannot be definitively answered due to insufficient data or inadequate tools for evaluating available data. Certain guidelines attempt to estimate the likelihood of a variant being pathogenic or benign using rule-based systems such as the American College of Medical Genetics and Genomics (ACMG) guidelines and the Sherloc framework (see, e.g., Nykamp K, Anderson M, Powers M, et al. Sherloc: a comprehensive refinement of the ACMG AMP variant classification criteria. Genet Med. 2017;19(10):1105-1117). Given a set of data and information as input, these other systems combine using heuristics to transform the data into qualitative evidence segments, and then output a qualitative classification. For example, these other systems may assign each part of the data to a qualitative category such as very strong, strong, moderate, or supporting evidence based on expert intuition, which introduces the possibility of bias, subjectivity, and inconsistency. Furthermore, in such systems, data is lost in the qualitative classification process because the qualitative evidence categories are coarse. For example, if a portion of the data is more compelling than "moderate" but does not reach the "strong" threshold, one of the evidence categories may be selected for the data depending on the judgment of the domain expert or historical precedents, thus reducing precision. The other systems then classify variants as pathogenic (P), likely pathogenic (LP), variant of unknown significance (VUS), likely benign (LB), or benign (B) based on a specific combination of evidence available for that variant.For example, a variant with one strong piece of evidence and two moderate pieces of evidence may be identified as likely to be pathogenic. Similar to data binning, other systems perform these categorization assignments based on intuition and thus introduce subjectivity. Furthermore, evidence of interactions can be more complex than simply summing the evidence. For example, there may be synergistic interactions between different types of evidence or partially overlapping or redundant evidence that should not be simply added together. While the ACMG guidelines aim for 90%–99% certainty for the two “likely” categories (likely pathogenic and likely benign) and >99% certainty for pathogenicity and benign, this certainty-based thresholding is inherently ideal, and the guidelines do not produce quantitative classifications. Other systems, being inherently qualitative, do not measure the accuracy or precision of the resulting variant classification or interpretation. Consequently, these systems are likely to introduce subjectivity and inconsistency, resulting in suboptimal outcomes for patients and the clinicians managing them. Additionally, these other systems limit the clinical utility of genetic test results by generating only the minimum set of outcomes that can be used as decision-making boundaries for medical management. Current management recommendations establish a single boundary where LP and P variants are considered medically intervenable, while VUS, LB, and B are considered not, and this single boundary is used regardless of the gene, disease, and medical intervention.
[0022] The described approach, which can be implemented using methods, systems, apparatus, non-temporary computer-readable media, etc., is designed to achieve a transition from existing suboptimal qualitative systems to a quantitative variant interpretation platform capable of consuming all available relevant data for any given variant, any gene in real time, and generating accurate and precise quantitative calculations of the probability of pathogenicity (PoP). The described approach includes methods for evaluating data associated with a variant to quantify the usefulness of that data as evidence that the variant is pathogenic or benign. An example uses a machine learning model to calculate the PoP of a variant, considering the complex constellation of phenotypes observed in the set of patients observed with this variant. Furthermore, the described approach includes methods for evaluating the PoP of a variant, considering a set of evidence (which may include quantitatively measured evidence and / or qualitatively binned evidence). For example, the set of evidence may include model prediction scores (i.e., quantitative evidence) and / or criteria (i.e., qualitatively binned evidence) based on in vitro functional assay results, population allele frequency data, mRNA splicing changes, and / or predictions of nonsense mutation-dependent degradation for a given variant.
[0023] Traditional machine learning models are correlated and non-causal in that they do not necessarily consider causal relationships between variables. This aspect of traditional models has statistical divergences because the model cannot leverage known statistical relationships between variables. In correlated models, it can be difficult to know how to interpret the model output when two predictor variables correlate with an outcome such as pathogenicity. For example, if two variables represent two independent biological mechanisms (causes) of pathogenicity (e.g., protein abundance and protein stability), only one of those variables may be sufficient to predict pathogenicity, while the other is irrelevant. On the other hand, if two variables are causally downstream (effects) of pathogenicity, mismatches between their respective predictions (e.g., high allele frequency in a healthy population cohort versus a strong correlation with disease phenotype in an affected cohort) should increase uncertainty in pathogenicity prediction. Correlated models may not be able to make these kinds of distinctions between upstream and downstream variables. Furthermore, correlated modeling approaches may produce models that are biologically meaningless and cannot be generalized to broader datasets. In contrast, causal probabilistic graphical models rely on domain knowledge to define causal relationships between variables.
[0024] In a probabilistic model, for example, if three variables are continuously connected in a graph, for instance, A predicts B and B predicts C, then given an estimate of the value of A, a probabilistic model can be used to predict the most likely value of C. Traditional models take point values as estimates and produce point values as outputs. On the other hand, probabilistic graphical models can track not only point values but also the entire probability distribution. In a traditional model, the input of A could be a point value like "0.5", but in a probabilistic graphical model, the input of A could be a distribution such as a bell curve centered at 0.5.
[0025] The ability to consider the distribution of values that could be considered as possible evidence and predictions can be very useful in general situations such as when measurements have some degree of uncertainty and / or when the modeling task has true randomness in the process. For example, having a BRCA1 mutation does not guarantee that a person will develop cancer, but it does increase a person's risk of cancer. In the case of incomplete penetrating diseases, seeing that a particular variant causes the disease in 50% of individuals may indicate that the variant is causal, but it never guarantees that the variant is causal. Probabilistic modeling can measure the likelihood that a variant is the root cause of the disease and can account for the uncertainty in predicting that a variant is the root cause of the disease.
[0026] In some examples, the described approaches include scalable models that utilize natural language processing (NLP) alone or in combination with other evidence modeling techniques, and Bayesian estimation (e.g., from clinical phenotypic data reported by requesting providers, such as test application forms) to predict variant pathogenicity. In some examples, an NLP classifier is used to learn features of indication and family history fields to predict the molecular diagnosis of a given condition. The classifier combines these features and demographic information for a given patient to determine a patient score, e.g., the probability that a given patient has the condition. The patient score is thus calculated for a population of patients. A generative hierarchical Bayesian inference model is fitted to these patient scores and labeled variant pathogenicity data, where available. By sampling the posterior predictive distribution from the inference model, a variant score, e.g., the probability that a given variant is pathogenic for the relevant condition, is obtained. In this way, the described approaches can account for both the uncertainty of disease status in each patient observation and the uncertainty due to the number of observations for each variant.
[0027] In some examples, the approaches described include evidence modeling techniques, a broad set of biological variables representing the molecular and / or cellular causes of pathogenicity, a broad set of variables representing the observable effects or impacts of pathogenicity in the human organism, and scalable models that utilize Bayesian estimation to predict variant pathogenicity.
[0028] In some embodiments or examples, probabilistic graphical modeling (PGM) uses directed acyclic graphs (DAGs) to represent conditional dependency structures between random variables. Nodes in a DAG represent conditionally independent variables. Edges in a DAG, indicated as arrows, signify dependencies between corresponding nodes, and the direction of the arrows signifies the direction of causality. In some examples, PGM implementations involve generating DAGs that represent causal relationships between a broad set of biological variables related to pathogenicity assessment, including nodes, edges, and edge directions.
[0029] The described DAG modeling approach includes a central node representing pathogenicity, with upstream (cause of pathogenicity) variables and / or downstream (effect of pathogenicity) variables pointing from this central node. Upstream nodes above the central node include variables representing data on the biological characteristics that cause pathogenicity, such as variables relating to the effect of variants on protein function, the effect of variants on protein stability, the effect of variants on mRNA processing such as mRNA splicing, the effect of variants on nonsense mutation-dependent degradation, the effect of variants on mRNA and / or protein expression, and variables relating to the biological importance of disrupted DNA and / or RNA nucleotides and / or protein amino acid residues, such as those confirmed from evolutionarily conserved data.
[0030] Downstream nodes below the central node include variables representing data on traits that are effects of pathogenicity, such as the effect of pathogenicity on individual patient phenotypes, the effect of pathogenicity on general population allele frequencies, the effect of pathogenicity on selected cohort allele frequencies, and the effect of pathogenicity on familial genotype-phenotypic cosegregation patterns. Upstream or downstream nodes may have additional nodes and edges representing one or more other relevant variables. For example, a first node representing the effect of variants on mRNA processing may be coupled to a second node representing data from in vitro mRNA splicing assays, and / or a third node representing data from in silico mRNA splicing prediction algorithms, and / or a fourth node representing positional information for consensus splice sequences.
[0031] The relationships visually represented by a DAG are encoded within the model by specifying the probability of available observations for the variables and parameters within the model. For example, the variable "pathogenicity" is only partially observed. It is represented with respect to all of its parent variables that have a function, and either a linear model or a deep learning model may be valid, for example. Using observations of the variable (e.g., known labels of existing variants) and observations of all upstream variables, the model is fitted by taking the data into account and maximizing the likelihood of all those parameters. The parameters learned from the fitting process can be used to sample the distribution of the variable "pathogenicity" or any other variable if the variable is not observed.
[0032] In some examples, evidence-based variant modeling platforms include many different evidence models, each or any of which can be used alone or in combination for variant interpretation. For example, models can be grouped according to categories of evidence criteria used by the Sherloc framework (e.g., assay data, in silico data, population data, etc.), and each grouping can include multiple alternative types of models for predicting pathogenicity, given the relevant type of evidence. For each grouping of models, the best-performing model based on various performance metrics can be selected for a particular gene. The selected model can then be used in existing variant interpretation systems such as the ACMG guidelines and the Sherloc framework, or alternatively, as a node within a PGM. Alternatively, or in addition, the Bayesian approaches described herein can be used to develop one or more models in any of the groupings on the platform, which can also be used in existing variant interpretation systems or as nodes within a PGM.
[0033] The described modeling strategies are applied to proprietary genotype / phenotype databases consisting of data collected with permission from thousands (e.g., 1,000, 10,000, 100,000, 500,000) or millions (e.g., 1 million, 2 million, 5 million, 10 million, 20 million, 50 million, 100 million) patients queried for clinical genetic testing of their genetic status, and thousands (e.g., 1,000, 10,000, 100,000, 500,000) or millions (e.g., 1 million, 2 million, 5 million, 10 million, 20 million, 50 million, 100 million) classified gene variants. In one embodiment, the described approach yielded high-performance models (AUROC > 0.8, evaluated with a holdout set of labeled variants for each gene) for over 1,100 genes associated with a wide range of clinical areas (e.g., oncology, cardiology, neurology, metabolism). In this embodiment, more than 20,000 VUSs across these genes and conditions received highly reliable stochastic predictions (PoP greater than 0.95 or less than 0.05) that could provide new evidence for interpretation.
[0034] In some examples, combined machine learning-based models are generated to compute PoPs from various individual machine learning-based models that are considered together. Each machine learning-based model is built and trained using large datasets containing thousands to millions of data records and automated and / or semi-automated processes. The outputs of multiple machine learning-based models (e.g., an in silico model of a particular gene + an assay model of that particular gene + a population model) can be used as individual nodes in a PGM. Thus, the PGM itself is a synthesis of multiple different machine learning-based models, each constructed and trained using large datasets and automated and / or semi-automated processes as described above, and the relationships between these models are machine-learned using automated and / or semi-automated tools. The approach described demonstrates the usefulness of Bayesian models for effectively integrating predictive data generated by many different models, with significant potential to reduce VUS, accelerate genetic diagnostics, and improve therapeutic personalization for patients and providers.
[0035] This disclosure will be better understood from the following detailed description with reference to the accompanying drawings. The detailed description of the drawings is for illustrative and understanding purposes only and should not be construed as limiting this disclosure to the specific embodiments described.
[0036] In the drawings and the following description, components that have the same name but different reference numbers may be referenced in different drawings. The use of different reference numbers in different drawings indicates that components with the same name may represent the same or different embodiments of the same component. For example, components with the same name but different reference numbers in different drawings may have the same or similar function in some embodiments, such that the description of one of those components in one drawing may apply to other components with the same name in other drawings.
[0037] Furthermore, components illustrated and described in the drawings and the following description in relation to some embodiments may be used in conjunction with other embodiments or incorporated into other embodiments. For example, components shown in a particular drawing may not be limited to use in relation to the embodiment to which that drawing relates, but may be used in conjunction with other embodiments, including embodiments shown in other drawings, or incorporated into other embodiments.
[0038] Figure 1A shows examples of machine learning-based processes for variant interpretation according to several embodiments of the present disclosure. In the examples in Figure 1A, a causal machine learning model is constructed using a Bayesian inference approach to model the relationships between different types of pathogenicity evidence for a given combination of genetic variants and health status, using phenotypic data, genotypic data, and domain knowledge. Bayesian inference is a method of statistical inference in which Bayes' theorem is used to update the probability of a hypothesis as more evidence or information becomes available. Bayesian inference uses prior knowledge in the form of a prior distribution to estimate the posterior probability. The prior probability distribution of an uncertainty is often simply called the prior and is an assumed probability distribution before any evidence is considered. The posterior probability is a kind of conditional probability that arises from updating the prior probability with information summarized by a likelihood function through the application of Bayes' theorem. The likelihood function (often simply called the likelihood) measures how well a statistical model explains observed data by calculating the probability of seeing that data under different parameter values of the model. The likelihood function is constructed from the joint probability distribution of the random variables that (presumably) generated the observation.
[0039] In Figure 1A, Person 1 and Person 2 represent two examples of subjects within population N, but any population N can contain any number M subjects (N and M are positive integers, and the value of M may differ for different populations). Population N contains M human subjects who participated in genetic testing, where N represents one or more defining characteristics of the population, such as demographic criteria, and M represents the number of subjects in that population. For example, M may be 10 human subjects, 100 human subjects, 1,000 human subjects, 10,000 human subjects, 100,000 human subjects, 1,000,000 human subjects, 2,000,000 human subjects, 5,000,000 human subjects, 10,000,000 human subjects, 100,000,000 human subjects, or 1,000,000 human subjects. Each subject in population N has associated phenotypic data 110 and genotypic data 112. The phenotypic data 110 includes observable features or traits of the subject, such as morphology, developmental processes, biochemical and physiological characteristics, behavior, and the consequences of behavior. For a given subject, phenotypic data 110 may include demographic information, family history of diseases and other health conditions, and indications. Indications include signs, symptoms, or medical conditions that suggest treatment, testing, or procedures may be appropriate or necessary. Indications may also refer to reasons why a drug or other treatment may be used. For example, aspirin is indicated for adults with cardiovascular risk factors such as hypertension or diabetes. Phenotypic data 110 are collected and used with the subject's permission. For example, phenotypic data may be collected through a test request form completed by the subject or clinician.
[0040] Genotype data 112 contains information about the genetic structure of each individual subject in population N. For example, a subject may submit a DNA (deoxyribonucleic acid) sample for genetic testing using DNA sequencing, and the results of the genetic testing provide information about the subject's genetic structure. A DNA sample contains one or more genes 102. Each gene contains one or more regions 104, for example, a contiguous group or sequence of nucleotides or proteins (e.g., AGACGCT, where A refers to adenine, G to guanine, C to cytosine, and T to thymine). Each nucleotide within a region has one or more positions where a genetic variant may be located. For example, in Figure 1A, person 1 of population N has cytosine at position 106 of region 104 of their copy of gene 102, and person 2 of population N has thymine at position 106 of region 104 of their copy of gene 102. In the example in Figure 1A, thymine is considered genomic variant 108. The source of the genotype data 112 may be a publicly available population database (e.g., gnomAD), a private population database, or any database that contains population data.
[0041] For each population in human populations 1 to N (each such population includes M subjects), the genotype data 112 may include gene-level data corresponding to each gene 102, region-level data corresponding to one or more regions 104 of the gene, location-level data corresponding to one or more positions 106, and variant-level data corresponding to one or more genomic variants 108. Gene-level data refers to any characteristics, attributes, or metrics determined for a defined sequence of DNA identified as a gene. For example, gene-level data may include information specific to a particular gene, such as the length of the gene. Region-level data refers to any characteristics, attributes, or metrics determined for a region of a gene, i.e., a defined part of a gene that is typically larger than a position but smaller than the entire gene. Region-level data may include information specific to a particular region of the gene. For example, region-level data may include the identification of a region containing a variant. A region may refer to a part of a gene containing a variant and one or more adjacent or neighboring nucleotides. Location-level data refers to a position within a gene, i.e., any characteristics, attributes, or metrics determined for a single nucleotide or amino acid position in the gene. Location-level data contains information specific to a particular location in a gene, regardless of whether a variant is present at that location. Typically, a gene can be several thousand nucleotides long, while a variant typically resides at a single location within those several thousand nucleotides. Variant-level data contains information about a specific variant located at a particular location in a gene. A variant may contain one nucleotide or more. If a variant contains more than one nucleotide, the variant's location refers to a region, and the terms location and region may be synonymous in that context.
[0042] Domain knowledge 114 can include learning information such as quantitative and / or qualitative compilations of clinical expertise and / or other evidence for one or more features (including, but not limited to, pathogenicity) of one or more gene variants. Domain knowledge 114 can be stored in one or more datasets and / or encoded in one or more machine learning models. For example, domain knowledge 114 can include statistical correlations between different types or categories of evidence for variant pathogenicity for a given health condition. One example of domain knowledge is a population frequency model that models the correlation between allele frequencies and variant pathogenicity. Another example of domain knowledge is historical variant interpretation information, e.g., historical outputs of a variant classification framework that can score different pathogenicity evidence differently for different populations, genes, variants, and / or health conditions, and the history of these scores can be used to develop probabilistic relationships between evidence types and pathogenicity. Additional examples of domain knowledge can include causal relationships between different types or categories of evidence for pathogenicity.
[0043] For example, domain knowledge 114 may be used by a causal model builder 118 to construct, train, and / or validate causal models (e.g., probabilistic graphical models such as directed acyclic graphs representing hierarchical Bayesian networks). Domain knowledge 114 may include one or more outputs, or combinations of outputs, generated by an evidence modeling platform (EMP), such as scores, distributions, or labels, examples of which are described herein. Examples of additional domain knowledge 114 include one or more of the following, or any combination thereof: CVM (patient clinical data assessed through EMP), PFM (population frequency data assessed through EMP), FIM (evolutionary conservation, protein structure, and protein stability data assessed through EMP), DMA (publicly available in vitro experimental study data assessed through EMP), DML (in-house performed in vitro experimental study data assessed through EMP), splice site loss predictions obtained via open-source models such as PANGOLIN, splice site gain predictions (e.g., obtained via PANGOLIN), and variant classification framework outputs (e.g., predictions of nonsense mutation-dependent degradation, human-curated data based on the application of the Sherloc framework).
[0044] Data preprocessing 116 includes one or more components for selecting data from sources of phenotypic data 110, genotypic data 112, and domain knowledge 114 to be used to construct a causal model. For example, given variants and health status, data preprocessing 116 can filter one or more of the phenotypic data 110, genotypic data 112, and domain knowledge 114 to exclude less predictive data and include more predictive data. Where available, data preprocessing 116 can combine known pathogenicity labels with relevant variant information. Data preprocessing 116 can convert one or more of the phenotypic data 110, genotypic data 112, and domain knowledge 114 into a format suitable or required for input to the causal model. Data preprocessing 116 is optional in some examples. For example, the phenotypic data 110, genotypic data 112, and / or domain knowledge 114 may be obtained from sources where they are already pre-formatted so that preprocessing is not required. The output of data preprocessing 116 may include training and / or validation datasets, and / or one or more structured representations of phenotypic data 110 such as feature sets, embeddings, tensors, rules, weights, and parameters, genotypic data 112, and / or domain knowledge 114.
[0045] The causal model builder 118 applies the output of domain knowledge 114 and / or data preprocessing 116 to construct a graphical representation of the causal model, or selects a model structure for the causal model from a library of available model architectures. The graphical representation of the causal model includes nodes and edges, each node representing a type or category of evidence, and each edge representing a type of relationship between two nodes. In a causal model or probabilistic graphical model (PGM), edges can be acyclic directed edges. Using the output of domain knowledge 114 and / or data preprocessing 116, the causal model builder 118 creates nodes to map to evidence categories, creates and weights edges between nodes, and determines the overall topography of the resulting graphical model.
[0046] Training / validation 120 applies the causal model constructed or selected by the causal model builder 118 to training and validation datasets prepared using inputs that include one or more of the phenotypic data 110, genotypic data 112, and domain knowledge 114. Although not specifically shown in Figure 1A, the training and validation processes involving each dataset are shown and described in more detail in subsequent figures. Once the causal model meets or exceeds the performance criteria applicable throughout training and validation 120, the trained and validated causal model is made available via the model providing / inference component 122 to respond to queries, such as sampling. The model providing / inference component 122 provides access to the causal model via, for example, one or more application programming interfaces (APIs), user interfaces, batch processes, query mechanisms, etc. Variant interpretation data 124 can be sampled and output via the model providing / inference component 122, which can be used, for example, by engineers, researchers, or clinicians to assist in diagnosis, adjust the causal model, update the domain knowledge 114, etc. The variant interpretation data 124 may include one or more values representing the probability of one or more pathogenicity (PoP), part or all of a causal model, updated domain knowledge 114, a graphical representation of a causal model, etc. Other examples of variant interpretation data 124 include the probability that a patient exhibits a particular phenotype (in addition to the probability that the variant is pathogenic, the disease risk for a given patient, the risk of expressing a particular phenotype), and the predicted functional / cellular assay output before the assay is performed, which may be useful for prioritizing experiments. An exemplary use of variant interpretation data 124 is to simulate the impact of missing data, e.g., the predicted impact on model outputs (PoP, etc.), if more data (e.g., more patient information, more family history information, functional assay results, etc.) were available. Another exemplary use of variant interpretation data 124 is to identify and highlight specific fragments of information that are missing but are needed to make a proper classification, and thus inform patient care (e.g., a specific type of blood test is needed for a proper classification).
[0047] The components and processes shown in Figure 1A are described in more detail with reference to subsequent figures. The examples shown in Figure 1A and the accompanying descriptions above are provided for illustrative purposes only. This disclosure is not limited to the examples described. Additional or alternative details and implementations are described herein.
[0048] Figure 1B shows examples of causal models according to some embodiments of the present disclosure. The examples in Figure 1B are illustrative and non-limiting.
[0049] In the example in Figure 1B, the causal model 150 is represented graphically as a probabilistic graphical model (PGM), or more specifically as a directed acyclic graph (DAG), or even more specifically as a Bayesian network represented as a DAG. A PGM is a modeling framework that uses graphs to represent conditional dependencies between random variables. A DAG is a specific type of graph that can be used within a PGM. In a DAG, each edge represents the direction of causality by an arrow, where the arrowhead indicates the direction of the causality, the arrowtail connects to a node representing the causal variable, and the arrowhead connects to a node affected by the causal variable. Also, DAGs are cycleless, meaning that a path through the graph cannot return to the same node (there are no loops). A Bayesian network is a type of PGM that can be used for causal and probabilistic inference, and its graph structure encodes the factorization of the joint probability distribution of variables. As illustrated with the examples shown in Figures 1B, 2A, 2B, 2C, 2D, 2E, 2F, and 2G, PGM can encode complex distributions across multiple different variables.
[0050] Causal model 150 is a probabilistic graphical model (PGM) represented as a directed acyclic graph and used to represent conditional dependency structures between random variables. Nodes in causal model 150 (e.g., Node 1, Node 2, Node 3, Node 4) represent independent variables. Edges (e.g., E12, E13, E14, E24, E34) represent dependencies between connected nodes, and the direction of the arrows on the edges indicates the direction of the causal relationship from one node to another. For example, in causal model 150, edge E12 represents a conditional dependency between the independent variable represented by Node 1 and the independent variable represented by Node 2. Node 2 is conditionally independent from Node 3, given the observation of Node 1. Edge E13 represents a conditional dependency between the independent variable represented by Node 1 and the independent variable represented by Node 3. Edge E14 represents a conditional dependency between the independent variable of Node 1 and the independent variable of Node 4. Edge E24 represents the conditional dependency between the independent variables of node 2 and node 4, and implicitly encodes the relationship between node 1 and node 2. Edge E34 represents the conditional dependency between the independent variables of node 3 and node 4, and implicitly encodes the relationship between node 1 and node 3.
[0051] As applicable to variant classification or variant interpretation problems, each node in the causal model 150 represents a type of evidence, which may relate to the pathogenicity of one or more variants associated with one or more health conditions, and / or any other characteristics of one or more variants or health conditions. The nodes are based on biological characteristics or any other relevant information, and the number of nodes in the causal model 150 is determined using domain knowledge about the types of evidence variables, such as useful, available, and reliable.
[0052] Each edge computationally represents the relationship between the types of evidence represented by the nodes connected by the edge. For example, different types of evidence variables can influence other evidence variables, either substituting for or in combination with each other. Each node can represent the output of a particular type of evidence model, and each relationship between nodes can be further represented by another machine learning model, mathematical model, function, heuristic, weighted linear model, deep learning model, or any other mathematical or logical function.
[0053] In some examples, each node in the causal model 150 represents a type of evidence modeled by an evidence modeling platform, such as platform 600, as described with reference to Figure 6. The configuration and topology of the causal model created by the causal model builder 118 can vary depending on the types of evidence available, the reliability of each type of evidence (e.g., weight values assigned by the EMP or variant classification framework), and / or other factors.
[0054] In some examples of the PGM implementation forms described, the DAG is generated using domain knowledge to accurately represent the assumed causal relationships between biological variables related to the assessment of variant pathogenicity, including nodes, edges, and edge orientation. For example, as applied to variant interpretation, the DAG is described as having a single node representing pathogenicity at its center (e.g., node 4 in causal model 150), with upstream variables above and pointing to this central node (e.g., nodes 2 and 3), or downstream variables below and pointing from this central node (e.g., nodes 5 and 6).
[0055] Upstream or parent nodes above the central node include variables representing data on pathogenicity-causing biological characteristics, such as variables relating to the effect of variants on protein function, the effect of variants on protein stability, the effect of variants on mRNA processing such as mRNA splicing, the effect of variants on nonsense mutation-dependent degradation, the effect of variants on mRNA and / or protein expression, and variables relating to the biological importance of disrupted DNA and / or RNA nucleotides and / or protein amino acid residues, such as those confirmed from evolutionarily conserved data.
[0056] Downstream or child nodes below the central node are variables representing data on biological properties that are effects of pathogenicity: variables relating to the impact of pathogenicity on individual patient phenotypes, the impact of pathogenicity on general population allele frequencies, the impact of pathogenicity on selected cohort allele frequencies, and the impact of pathogenicity on familial genotype-phenotypic cosegregation patterns. For nodes both above and below the central pathogenicity node, additional nodes and edges may exist representing other relevant variables (e.g., node 1). For example, a node relating to the impact of variants on mRNA processing may be linked to nodes representing data from in vitro mRNA splicing assays, nodes representing data from in silico mRNA splicing prediction algorithms, and nodes representing positional information relative to consensus splice sequences.
[0057] The relationships visually represented by a DAG are encoded within the model by specifying the likelihood, e.g., probability, of available observations for the variables and parameters within the model. For example, a variable such as "pathogenicity" may only be partially observed and is therefore represented using a function for all of its parent variables. For example, a linear model in the context of a GLM (Generalized Linear Model), or a deep learning model, can both be valid. The model is fitted by using observations of the variable (e.g., known labels of existing variants) and observations of all upstream variables to maximize the likelihood of all those parameters, taking the data into account. The learned parameters can then be used to sample the distribution of the variable "pathogenicity" when it is not observed.
[0058] As shown in Figures 1B, 2A, 2B, 2C, 2D, 2E, 2F, and 2G, the structure of causal model 150 is flexible and adaptable, and can be restructured and redesigned when new evidence is obtained, when new information about the supporting model is obtained (e.g., changes in the weighting of the evidence model provided by EMP), etc. In some examples, causal model 150 is a unified model having the same topology, structure, variables, parameters, etc. for all gene variants, which may or may not correspond to the example shown in Figure 1B, or to the examples shown in Figures 1B, 2A, 2B, 2C, 2D, 2E, 2F, and 2G. In other examples, as shown in Figures 2C and 2D, for example, causal model 150 can be modified or adapted for different variants, groups of variants, genes, or groups of genes, so that different configurations of causal model 150 are constructed, trained, maintained, and served for different variants, groups of variants, genes, or groups of genes.
[0059] New variables can be added to the causal model 150 using domain knowledge by deciding to encode the representations and relationships to the causal model 150 using a modeling tool (e.g., software), and to retrain the model. This includes determining the appropriate representation of the new variable (e.g., which type of model best represents the variable, such as a distribution or point value), the relative position of the variable to other variables represented by nodes in the graph (e.g., independence or conditional independence from other variables), the relationship between the causal variable and the target variable, and the relationship between downstream variables and the target variable (e.g., by determining the causal function to which different variables are connected). Variables and / or relationships between variables, and / or other aspects of the model can be updated in a similar manner.
[0060] The components and processes shown in Figure 1B are described in more detail with reference to subsequent figures. The examples shown in Figure 1B and the accompanying descriptions above are provided for illustrative purposes only. This disclosure is not limited to the examples described. Additional or alternative details and implementations are described herein.
[0061] Figures 2A, 2B, 2C, 2D, 2E, 2F, and 2G illustrate examples of probabilistic models for variant interpretation according to some embodiments of the present disclosure. In the illustrated examples, the probabilistic model is described using a graphical representation, for example, a directed acyclic graph containing nodes and acyclic directed edges between nodes.
[0062] One or more of the examples shown in Figures 2A, 2B, 2C, 2D, 2E, 2F, and 2G can be pre-built and stored, for example, in a model library. These examples and / or variations thereof are built using domain knowledge (e.g., domain knowledge 114 in Figure 1A). For example, domain knowledge is used to identify evidence categories to include in a probabilistic model and to exclude evidence categories that are not relevant to variant classification or variant interpretation. The evidence categories included in the probabilistic model are represented by nodes in a directed acyclic graph (DAG). Domain knowledge is also used to specify causal relationships between categories of evidence (e.g., A causes B, and B causes C). These causal relationships are represented as acyclic directed edges between nodes in the DAG.
[0063] Throughout the training process, weights can be assigned to nodes and / or edges, and these weight values reflect varying degrees of availability or reliability of different types of evidence. For example, nodes represent variables corresponding to evidence categories. Variables can be, for example, independent or conditionally independent. Variables can also be represented as distributions or point values. Variables can be parameterized using domain knowledge and / or other methods that enable the learning and estimation of parameters. For example, a reasonable range of prior probabilities of pathogenicity of a variant in a given gene can be determined based on domain knowledge such as the observed pathogenicity rate in that gene. However, if the distribution of inputs is not well understood through domain knowledge, other tools can be used to estimate the distribution of inputs.
[0064] Additionally or alternatively, domain knowledge and / or other methods or tools can be used to determine how different variables (e.g., evidence categories) relate to each other, and these relationships can be encoded in the model's linking function. For example, the knowledge that two variables are linearly inversely correlated is an example of domain knowledge that can be encoded in the model.
[0065] Figure 2A shows a probabilistic graphical model 200 having a central node 202, an upstream node 204, and a downstream node 206. The central node 202 represents the pathogenicity variable 202. For example, sampling of the pathogenicity node 202 provides the probability that a variant is pathogenic with respect to a disease (or health condition), and this probability is influenced by the upstream node 204. The upstream node 204 represents a causal variable, for example, a variable that has a likelihood of causing a disease given the functional effect of a variant. In the example in Figure 2A, node 204 represents the modified protein function variable. The relationship between the modified protein function variable and the pathogenicity variable is represented by a directed edge 203. The direction of edge 203 indicates the direction of the causal relationship. Thus, edge 203 indicates that the modified protein function variable is the cause of pathogenicity.
[0066] The downstream node 206 represents an effect variable, such as a variable that has the likelihood that a variant was pathogenic, taking into account how common the variant is in the general population. In the example in Figure 2A, node 206 represents an observation in the population database. The relationship between the observation represented by node 206 and the variant pathogenicity variable is represented by the directed edge 205. The direction of edge 205 indicates the direction of causality. Therefore, edge 205 indicates that the observation in the population database represented by node 206 is an effect arising from variant pathogenicity.
[0067] Figure 2B shows another example of a probabilistic graphical model 210 having a central node 212, an upstream node 214, and two downstream nodes 216 and 218. The central node 212 represents the variant pathogenicity variable. The upstream node 214 represents the EMP model output from a causal variable, e.g., an MSE (Molecular Stability Engine) that predicts the structure and stability of a protein with a given gene variant. The relationship between the upstream node 214 and the central node 212 is represented by a directed edge 213. The direction of the edge 213 indicates the direction of the causal relationship. Thus, edge 213 indicates that the MSE (Molecular Stability Engine) variable is causally related to pathogenicity. In some examples, the MSE is a model supported by EMP, and node 214 represents the output of the MSE, e.g., a distribution representing molecular stability.
[0068] Downstream nodes 216 and 218 each represent effect variables, for example, variables that have the likelihood that a variant was pathogenic, taking into account how common the variant is in the general population. In the example in Figure 2B, node 216 represents an effect variable, a population frequency model, for example, a model of the correlation between population allele frequencies and pathogenicity. Node 218 represents an effect variable, for example, an EMP model based on evolutionarily conserved data. The relationship between nodes 212 and 216 is represented by a directed edge 215, and the relationship between nodes 212 and 218 is represented by a directed edge 217. The direction of edges 215 and 217 indicates the direction of causality. Thus, edges 215 and 217 indicate that the variables represented by nodes 216 and 218 are effects arising from variant pathogenicity.
[0069] Figure 2C shows another example of a probabilistic graphical model 220 having a central node 224, an upstream node 234, and downstream nodes 226, 228, and 230. In the example in Figure 2C, the variant-level model, which includes variant-level variables represented by nodes 224, 226, 228, and 230, is extended to include gene-level variables at nodes 234 and 236. The directed edges 225, 227, 229, 235, and 237 each represent causal relationships between nodes connected by edges.
[0070] In the example in Figure 2C, node 228 represents the allele count (AN) in a population database such as gnomAD, which reflects the number of individual chromosomes sampled (e.g., representing the sampling size). Node 230 represents the allele count (AC) in a population database such as gnomAD, which reflects the number of times a particular variant has been detected. Along with the allele count (node 228), these variables are used to calculate the confidence of estimates through an understanding of allele frequency (node 226) and sample size. Node 234 represents the label ratio (lp) variable, which is the ratio of known pathogenic labels to known benign labels in a dataset such as gnomAD. The label ratio is used to determine the prior probability. Node 236 represents the beta distribution (β_i) variable, which is a parameter for the AF distribution defined at the gene level (i.e., specific to each gene).
[0071] Figure 2D shows another example of a probabilistic graphical model 240 similar to the example in Figure 2C, except that model 240 imposes constraints 246 on the label ratio variable 249 and the beta hyperparameter 248, and includes both variant-level variables 244 and gene-level variables 242. Constraints 246 can be used for genes that have variables satisfying constraints 246. Thus, model 240 can be used for groups of genes where all satisfy constraints 246 (rather than, for example, having to build a separate model for each gene in the group).
[0072] Figure 2E shows another example of a probabilistic graphical model 250 having a central node 252, upstream nodes 254, 256, 262, and downstream nodes 264, 266, 270. In the example in Figure 2E, each of the upstream nodes 254, 256, 262 and the downstream nodes 264, 266, 270 represents a probabilistic graphical model specific to the variable represented by each node 254, 256, 262, 264, 266, 270. Each variable-specific submodel is a probabilistic graphical model in the illustrated example, but could be of other types.
[0073] In the example in Figure 2E, nodes 254, 256, and 262 represent variables that express biological concepts related to the cause of pathogenicity (e.g., the effect of the variant on protein, mRNA processing, or disruption of cell / tissue function may cause pathogenicity). Nodes 264, 266, and 270 represent variables that express observable effects of pathogenicity (e.g., the number of patients affected by the disease, the impact on allele frequencies in a population, the impact on the degree to which a reference allele is conserved over evolution). Node 252 represents the predicted pathogenicity of the variant.
[0074] Figure 2F shows another example of a probabilistic graphical model 280, similar to model 250 but extended to include a second central node 288 in addition to the central node 294. The second central node 288 represents the drug responsiveness of the variant in a somatic setting to the anticancer drug vemurafenib. Extending the PGM to include multiple central nodes improves the usefulness of somatic / tumor tests for drug responsiveness, for example, by enabling the utilization of data / knowledge from germline tests. Currently, germline variant interpretation (assessing the pathogenicity of a variant) and somatic (tumor) variant interpretation (disease management information such as diagnosis, prognosis, and expected drug responsiveness) have little overlap. That is, their assessments are often performed independently, and data related to one variant is not considered related to another variant and is therefore not used at all. By creating linked DAGs as shown and described, it becomes possible to use data associated with one variant (e.g., simultaneous isolation of variants and disease in germline settings) to inform other variants (e.g., drug response in tumors).
[0075] Nodes 284, 286, 290, and 294 are upstream of node 288, and node 292 is downstream of node 288. In the example in Figure 2F, the model of protein effect variables is extended by node 288 because the node of submodel 286 has a causal relationship with the variable represented by node 288.
[0076] Figure 2G shows another example of a probabilistic graphical model 281. Model 281 includes a central pathogenicity node 291, several upstream (causal) nodes 283, several downstream (effect) nodes 285, an upstream acyclic directed edge 289, and a downstream acyclic directed edge 287. Each of the upstream nodes 283 is connected to the central pathogenicity node 291 by one of the acyclic directed edges 289, and each of the downstream nodes 285 is connected to the central pathogenicity node 291 by one of the acyclic directed edges 287. For each acyclic directed edge, the direction of the arrow indicates the direction of causality.
[0077] In the example in Figure 2G, each of the upstream nodes 283 and each of the downstream nodes 285 corresponds to a different machine learning model that generates predictive outputs related to a specific type or category of pathogenicity evidence (e.g., distribution or point value). For example, each of the upstream nodes 283 and each of the downstream nodes 285 could correspond to a machine learning model on an evidence modeling platform such as platform 600, as described with reference to Figure 6. In Figure 2G, CVM refers to the clinical variant model (including patient clinical data modeled and evaluated through EMP), PFM refers to the population frequency model (population frequency data modeled and / or evaluated through EMP), FIM refers to the evolutionary conservation, protein structure, and protein stability data modeled through EMP, DMA refers to the publicly available in vitro experimental study data modeled and / or evaluated through EMP, DML refers to the internally performed in vitro experimental study data modeled and / or evaluated through EMP, PANGOLIN splice loss refers to the splice site loss prediction model obtained from an external source such as the PANGOLIN model, and splice gain refers to the splice site gain prediction model obtained from an external source such as the PANGOLIN model.
[0078] In summary, the examples shown in Figures 2A, 2B, 2C, 2D, 2E, 2F, and 2G illustrate several different configurations of causal models for variant interpretation. In any of the models, it is possible to sample any of the nodes to obtain predictive information associated with the sampled node.
[0079] The examples shown in Figures 2A, 2B, 2C, 2D, 2E, 2F, and 2G, and the accompanying descriptions above, are provided for illustrative purposes only. This disclosure is not limited to the examples described. Additional or alternative details and implementations are described herein.
[0080] Figure 3 illustrates an exemplary component-based process for creating a probabilistic model for variant interpretation, according to several embodiments of the present disclosure. Figure 3 shows an example of communication between various components of a computing system involved in executing a component-based process 300, which includes machine learning techniques used by various components of the computing system to create a causal model for variant interpretation. The component-based process 300 is executed by processing logic embodied in various components of the computing system, including hardware (e.g., processing devices, circuits, proprietary logic, programmable logic, microcode, device hardware, integrated circuits, etc.), software (e.g., instructions that are run on or executed on processing devices), or a combination thereof. In some embodiments, the component-based process 300 is executed by one or more components of the computing system, including components or flows shown in Figure 3, which in some embodiments may not be specifically shown in other figures, and / or components or flows shown in other figures, which in some embodiments may not be specifically shown in Figure 3. Although shown in a specific sequence or order, the order of processes can be changed unless otherwise specified. Thus, the illustrated embodiments should be understood as examples only, the illustrated processes can be executed in different orders, and some processes can be executed in parallel. Additionally, in various embodiments, at least one process can be omitted. Therefore, not all processes are required in all embodiments. Other process flows are possible.
[0081] The component-based process 300 uses one or more machine learning techniques to prepare one or more datasets, e.g., a training dataset 338 and a validation dataset 342, to input into a causal model 328 via a model training component 326 and a model validation subsystem 330, respectively. Part of the process 300 is performed by components of a modeling computing system (e.g., a modeling system 1050 as described with reference to Figure 10), which includes a data preprocessing subsystem 308, a modeling subsystem 320, a model validation subsystem 330, and a model providing / inference subsystem 332. Data sources from which data is received between the various parts of the process 300 include unlabeled data 302, labeled data 304, domain knowledge 306, feature manifests 318, model performance criteria 336, a training dataset 338, model validation criteria 340, and a validation dataset 342.
[0082] Unlabeled data 302 includes unlabeled phenotypic and / or genotype data. For example, unlabeled data 302 includes population data extracted from gnomAD or similar databases, and / or features extracted from test application forms using natural language processing, such as natural language descriptions of family history, indications, etc.
[0083] Unlabeled data 302 does not include associated pathogenicity labels or scores. Pathogenicity labels or scores (e.g., benign or pathogenic) can be obtained from labeled data 304. Labeled data 304 is a reference database such as ClinVar, or an internally developed database curated by genetic scientists and / or other genetic experts, linking variants to associated ground truth pathogenicity labels or scores. Examples of labeled data 304 include gene-specific labels, labels pooled from collections of genes with similar biological characteristics (e.g., gene families, genes in the same pathway), labels obtained from proprietary databases (e.g., compiled based on historical data), labels generated and manipulated from proprietary databases (e.g., simulated recalculations of labels with added or deleted data), labels obtained from (e.g., external) data sources such as ClinVars with various sorting or filtering parameters (e.g., all ClinVars, only data from trusted ClinVar submitters, only specific classification categories, etc.), and labels for one or more subsets of variant types (e.g., missense, splice sites, etc.). Other examples of unlabeled data 302 and / or labeled data 304 include other models, empirical measurements, predicted variant effects, and outputs from public datasets.
[0084] Labeled data 304 can be combined or merged with corresponding unlabeled data 302 (for example, using a common identifier or key value) to form a training dataset 338, or separate training datasets can be formed for each of the unlabeled data 302 and labeled data 304. In other words, the model training component 326 of the modeling subsystem 320 can train a causal model 328 with a first training dataset 338 containing only unlabeled data 302, and separately train a causal model 328 with a second training dataset containing only labeled data 304, or the model training component 326 can combine the unlabeled data 302 and labeled data 302 to train a causal model 328 with a combined training dataset. Additional details and examples regarding training and validation are provided by referring to the modeling subsystem 320 and the model validation subsystem 330.
[0085] The data preprocessing subsystem 308 includes a fetch component 310, a natural language processing (NLP) component 312, a filtering component 314, and a feature generation component 316. For example, the data preprocessing subsystem 308 converts variant identifiers to a common format, identifies and fetches the correct data when available, filters the fetched data to a given scope (e.g., missense variants), and creates an appropriate data structure (e.g., annotated tensor). In some embodiments, the data input to process 300 is already in a format that can be used directly by the modeling subsystem 320. In these embodiments, the data preprocessing subsystem 308 may be omitted.
[0086] The fetch component 310 performs one or more processes in which raw data is retrieved from one or more data sources (e.g., unlabeled data 302, labeled data 304, domain knowledge 306) using, for example, a query mechanism or retrieval engine. The fetch component 310 evaluates the unlabeled data 302 and labeled data 304 and selects portions of the unlabeled data 302 and labeled data 304 to be used to generate datasets that can be input to create training and / or validation datasets for the causal model 328.
[0087] NLP component 312 is used for natural language processing of raw input data as needed. NLP component 312 may be employed when the fetched data (e.g., unlabeled data 302) is in the form of unstructured content such as natural language text. If the fetched data does not contain unstructured content, NLP component 312 may be omitted. NLP component 312 applies one or more NLP techniques to unstructured data such as natural language text extracted from documents such as examination application forms. For example, NLP component 312 may use entity recognition techniques to find and identify, for example, family history and / or indication information in a document, and then extract the identified information.
[0088] The filtering component 314 filters the input data (e.g., unlabeled data 302, labeled data 304, domain knowledge 306) to create one or more datasets for processing by, for example, the feature generation component 314, the modeling subsystem 320, and / or the model validation subsystem 330. The filtering component 314 can selectively filter the raw data to obtain datasets that should be used to build, train, or validate the causal model 328. The filtering component 314 applies one or more filters to the data 302, 304. The filters include criteria for determining whether some of the data 302, 304 should be included in or excluded from the model training and / or validation process. For example, data associated with variants that have no known association with any disease may be filtered by the filtering component 314 and thereby excluded from the training dataset 338 and / or validation dataset 342. As another example, raw data obtained from an external source, such as gnomAD data, can be filtered by the filtering component 314 to extract only the data related to specific variables, such as allele frequencies. In some embodiments, the input data may already be filtered, and therefore the filtering component 314 may be omitted.
[0089] The feature generation component 316 converts the input data into a format that can be input to the causal model 328. For example, the feature generation component 316 uses the feature manifest 318 to convert the output resulting from the application of the fetch component 310, the NLP component 312, and / or the filtering component 314, or a combination thereof, to the data 302, 304 into a format that can be input to, read by, and processed by the causal model 328, such as a tensor, vector, or embedding. The feature manifest 318 includes specifications for the types and configurations of features that can be input to the causal model 328. In some embodiments, the input data may already be in a suitable format, and therefore the feature generation component 316 may be omitted.
[0090] The modeling subsystem 320 constructs and trains a causal model 328. In some embodiments, the modeling subsystem 320 receives a feature set designed and output by the feature generation component 316 as input. In other embodiments, the modeling subsystem 320 receives formatted data from another system or component as input. The modeling subsystem 320 includes a dataset creation component 322, a construction component 324, and a model training component 326.
[0091] The dataset creation component 322 creates datasets for model training and validation. For example, the dataset creation component 322 divides the feature set designed and output by the feature generation component 316 into training and validation (e.g., holdout) datasets, such as the training dataset 338 and the validation dataset 342. For example, in some embodiments, the dataset creation component 322 creates training and validation datasets specific to a gene or variant.
[0092] The model building component 324 uses domain knowledge 306 to select or construct a causal model 328. For example, the model building component 324 may query a library of causal models (e.g., model library 334) to identify a causal model from the library that corresponds to the domain knowledge 306. As another example, the model building component 324 uses domain knowledge 306 to create a graphical representation of the causal model 328, including identifying nodes, determining the arrangement or topology of nodes, and creating edges between nodes (e.g., acyclic directed edges). For example, the model building component 324 may use domain knowledge 306 to identify the type of pathogenicity evidence that is a key determinant of pathogenicity, determine whether the evidence is upstream (e.g., causal) or downstream (e.g., effect) of pathogenicity, and establish relationships between nodes. Examples of causal models that can be constructed by the model building component 324 are described with reference to Figures 2A to 2F.
[0093] The model training component 326 executes a model training process that applies the causal model 328, selected or constructed by the model building component 324, to one or more of the datasets created by the dataset creation component 322. For example, the model training component 326 iteratively applies the causal model to the training dataset 338 and adjusts the weights of one or more model parameters and / or features until the comparison between the predicted model output generated by the causal model 328 by sampling the pathogenicity nodes of the causal model and the predicted model output proven by ground truth labels obtained via labeled data 304 satisfies (e.g., satisfies or exceeds) the model performance criterion 336. Once the model performance criterion 336 is met, the modeling subsystem 320 terminates the model training process and generates the trained causal model 328.
[0094] The model validation subsystem 330 applies the model validation process to the trained causal model 328 generated by the modeling subsystem 320. The model validation subsystem 330 applies the trained causal model 328 to the validation dataset 342 to determine whether the model validation criteria 340 are met (e.g., met or exceeded). If the trained causal model 328 is successfully validated by the model validation subsystem 330, the validated causal model 328 is provided to the model providing / inference subsystem 332 for inference, for example, to generate pathogenicity predictions, pathogenicity probabilities, or estimates of novel (i.e., previously unknown) variants. Through the model providing / inference subsystem 332, predictive data can be obtained by sampling pathogenicity nodes or any other nodes of the causal model 328. Alternatively or additionally, the predictive data output by the validated causal model 328 can be stored for future use (e.g., for access or lookup by one or more downstream processes, systems, or services). In some embodiments, predictive data can be used to update domain knowledge 306.
[0095] The example shown in Figure 3 and the accompanying explanation above are provided for illustrative purposes only. This disclosure is not limited to the example described. Additional or alternative details and implementations are described herein.
[0096] Figure 4 shows exemplary processes for training and validating probabilistic models for variant interpretation according to several embodiments of the present disclosure.
[0097] In Figure 4, the model training and / or validation process 400 applies a causal machine learning model 404 to the input dataset 402 using machine learning. The input dataset 402 can be the training dataset 338 or the validation dataset 342, as described with reference to Figure 3.
[0098] In response to the input dataset 402, the causal machine learning model 404 generates a model output 406, for example, by sampling the pathogenicity nodes of the causal machine learning model 404. Based on the evaluation of the model output 406 in subprocess 412 using model performance or validation criteria, one or more model parameters 410, node weights, and / or edge weights 408 of the causal machine learning model 404 may be adjusted in subprocess 414, and another iteration of process 400 may be initiated. In each iteration, the decision subprocess 414 determines, based on the evaluation performed by subprocess 412, whether to continue the iteration of process 400 or proceed to model provision / inference in subprocess 416.
[0099] In some examples, subprocess 412 includes evaluating model performance using separation performance techniques (e.g., area under receiver operating curve or AUROC) to determine how well variants known to be pathogenic or benign can be separated (e.g., accurately predicted by the model). In some examples, subprocess 412 uses calibration metrics to determine the accuracy of PoP calculations output by the model over different prediction ranges. In some examples, subprocess 412 evaluates model performance using one or more of the following: test set reclassification rate, VUS reclassification rate, agreement with other classifications (e.g., Sherloc classification), model precision, or error analysis to identify the stage of the model where errors occur and the reasons for the errors. In some examples, based on the output of subprocess 412, subprocess 414 modifies the input data to a set of data considered reliable or ground truth labeled data by adding or removing input features or by changing the method used to label the input data. In some examples, subprocess 414 adjusts one or more model parameters 410 based on the output of subprocess 412.
[0100] Other examples of criteria that can be used to evaluate a causal machine learning model 404 include how well the model handles missing data, how easy it is to explain how the model arrived at its conclusions for a particular variant, how easy it is to modify or update the input data by simulating it (e.g., adding splice predictions for a particular variant and seeing how that modifies the predictions), how well the model incorporates domain knowledge such as scientific or clinical expertise, how well the model's parameters correspond to scientific concepts, whether the model can represent differences between pooled groups of variants, the proportion of variants for which the model can provide usable predictions, the difference between the model's predictions for labels in a control set (e.g., pathogenic / benign) and a random set of VUS variants, the agreement with a variant classification or variant interpretation framework, and the model's ability to update in real time and make on-demand predictions.
[0101] The examples shown in Figure 4 and the accompanying description above are provided for illustrative purposes only. This disclosure is not limited to the examples described. Additional or alternative details and implementations are described herein.
[0102] Figure 5A shows examples of variant interpretation or variant classification systems according to some embodiments of the present disclosure.
[0103] In Figure 5A, system 500 includes a data layer 502, a model layer 512, an application layer 514, and an infrastructure layer 516. Each or any of layers 502, 512, 514, and 516 may be implemented as one or more software-based components on one or more computing devices.
[0104] The data layer 502 includes several different datasets 536, 538, 540, and 542. Each dataset 536, 538, 540, and 542, which may include, for example, model training data and / or model validation data, phenotypic data, genotypic data, domain knowledge, and model outputs, is supported by one or more data ingestion components, including an online data ingestion component 528, an offline data ingestion component 530, a data ingestion utility component 532, and a labeling method 534. The data ingestion components 528, 530, and the data ingestion utility component 532 are configured according to the requirements of their respective datasets and / or system requirements. For example, system 500 connects to a source of updates for datasets 536, 538, 540, and 542 via one or more components of the data layer 502. For example, system 500 may connect to one or more online and / or offline data stores.
[0105] In some examples, data ingestion components 528, 530, and 532 allow the model layer 512 to systematically access various types of data used when creating, training, validating, and using the model. The online data ingestion component 528 can be used to ingestion data into a dataset when the data is frequently updated, such as new patient phenotypic data that is updated when a new patient is referred for examination. The online data ingestion component 528 can connect the system 500 to updated data (e.g., an online data store such as an online database). The offline data ingestion component 530 can be used to ingestion static data types that can be stored in an offline or nearline data store (e.g., gnomAD population allele frequency data or Pangolin splice effect predictions).
[0106] The data acquisition utility component 532 may include the timing, frequency, and method for selecting, activating, or calling one or more tools, e.g., one or both of the online and / or offline data acquisition components 528, 530, which can be used to schedule or coordinate the data acquisition process. The labeling method 534 may include, for example, a method (e.g., a computer program) that can apply labels to items in a dataset based on rules or classification algorithms. The labeling method 534 may include one or more tools that can be used to acquire label data from various data sources (e.g., proprietary variant databases, external data sources such as ClinVar, etc.). The labeling method 534 may include one or more processes for selectively assigning labels to data, e.g., for determining which labels to apply to different types of data. For example, the labeling method 534 may decide not to use unverified label data, or to use only label data from certain trusted sources.
[0107] System 500 is flexible in that the labels that can be obtained and used by the labeling method 534 can be adapted to the requirements of a particular design or implementation. In some examples, pathogenic (P) and benign (B) labels are used (e.g., for binary classification via supervised machine learning). In other implementations, pathogenic, meaningless variant (VUS), and benign labels are used (e.g., in multi-class machine learning-based classification). In some embodiments, P and LP (likely pathogenic) data are grouped under the pathogenic label, and / or LB (likely benign) and B data are grouped together under the benign label.
[0108] The model layer 512 includes a model abstraction layer 510 and a model implementation layer 504. The model abstraction layer 510 is configured to delegate model interactions to the underlying model in the model implementation layer 504. The model abstraction layer 510 can determine the types of interactions available to a given model type. For example, in response to a request (e.g., a sampling request), a Bayesian model may deliver a distribution, while another model type may deliver only point estimates. The model abstraction layer 510 may include a set of common methods that can be used across or by any of the models in the model implementation layer 504, as well as logic for unique interactions.
[0109] The model implementation layer 504 contains N different models 506, 508, where N is a positive integer (two models are shown for illustrative purposes, but are not limited to two). For example, model 1 506 could be a Bayesian inference model, and model 2 508 could be a deep learning model or a regression model. As another example, model 1 506 could be a first configuration of a Bayesian inference model, and model 2 508 could be a second configuration of a Bayesian inference model different from model 1 506. Models 506 and 508 may have the same architecture but be trained on different training data, or they may have different model architectures. For example, the model implementation layer 504 may contain one or more of the models described with reference to Figures 1B, 2A-2F, or 6. The models included in the model implementation layer 504 are models trained and validated using, for example, the platform 600 described with reference to Figure 6.
[0110] The application layer 514 includes interaction tools (APIs, UIs) and tools for querying and monitoring models within the model implementation layer 504. For example, the application layer 514 includes a model providing component 518, one or more application programming interfaces (APIs) 520, an interaction component 522, one or more cloud platform components 524, and one or more monitoring components 526. The model providing component 518 enables access to the model within the model implementation layer 504, for example, to retrieve samples or predictions from the model. The APIs 520, interaction component 522, cloud platform component 524, and monitoring component 526 support the model providing component 518 depending on the type of interaction (e.g., connection, query, communication, etc.) being received by the application layer 514.
[0111] In some examples, the model-providing component 518 provides model-providing functionality that enables online access and querying capabilities. For example, the model-providing component 518 can enable a clinical reporting software application to access pathogenicity prediction outputs from one or more models in the model layer 512, and the pathogenicity prediction outputs can be imported into the clinical reporting software, for example, to provide an interpretation of variants observed during genetic testing.
[0112] It should be noted that the model-providing component 518 can provide predictive outputs of one or more models in the model layer 512, and such predictive outputs of models (including causal models as described herein) may include pathogenicity predictions and / or other types of predictions. For example, a causal model can be sampled at a molecular-cellular function node, and the predictive model outputs provided at that node can be used to determine whether molecular-cellular function is disrupted by a variant. As another example, model outputs can be separated or grouped into specific families or groups of genetically related individuals, which may be useful for patients in determining recommended next steps. In yet another example, model outputs can be sampled at a molecular function node to obtain predictive data on molecular function that may be useful in designing drugs.
[0113] In an example of a causal model constructed such that the pathogenicity node is connected to both the first and second parent nodes, with molecular function (first parent node) and transcript / protein expression node (second parent node) being causal inputs to the pathogenicity node, if a variant is known to be pathogenic and its molecular function is also known to be unimpaired, then in response to inputs of these parameters, the model output can be used to infer that the variant is pathogenic through means other than molecular function, which, based on the model architecture, can indicate that the possible cause of pathogenicity is transcript / protein expression and not molecular function. In this way, by sampling data from other nodes in the causal model, potential alternative sources of causality can be identified. These and other examples demonstrate that the potential applications of the described causal models are not limited to clinical situations but may also be useful, for example, in pharmaceutical research, drug discovery, and / or other applications.
[0114] The topologies and architectures of causal models (e.g., DAGs) described herein have no known limitations, as their outputs reflect the level of uncertainty in the model design. Causal models described herein may also inform of any evidence, data sources, or experiments that may be useful in pursuing ways to incorporate uncertainty and help address it.
[0115] One or more APIs 520 provide one or more application programming interfaces that can be used to enable a calling program to access the model input and / or output data of the model layer 512. For example, data can be retrieved from the model layer 512 via one or more APIs 520 and used for data science experiments and analyses.
[0116] The interactive component 522 can provide a web interface that allows, for example, scientists and / or clinicians to interact with one or more models in the model layer 512. For example, the interactive component 522 can generate and present simulations in response to requests received through the interactive component 522 (e.g., from a user device), and the simulations can simulate the impact of new data on model predictions (e.g., "What will happen to the predictions if we perform another patient observation?").
[0117] The monitoring component 526 provides a monitoring service that monitors the performance of one or more models in the model layer 512 and can detect changes in model performance over time, for example, for quality control purposes.
[0118] The cloud platform component 524 can provide access control and security services, such as protected cloud access to model output, for example, for research purposes or clinical reporting applications.
[0119] The infrastructure layer 516 supports the other layers 502, 512, and 514 by providing / with tools and utilities for computation (e.g., computation components 544), data storage (e.g., storage 546), event logging (e.g., logging 548), communication (e.g., Docker 550, CI / CD 552), and security (e.g., security 554).
[0120] For example, the infrastructure layer 516 may include backend computing components and associated distribution tools to support the application layer 514. As another example, the infrastructure layer 516 may support an Evidence Modeling Platform (EMP) in addition to the causal model-based variant classification or variant interpretation system described. For example, the infrastructure layer 516 can facilitate, but is not required, the supply of information from the EMP for input to the causal modeling system. The example shown in Figure 5A and the accompanying description above are provided for illustrative purposes only. This disclosure is not limited to the examples described. Additional or alternative details and implementations are described herein.
[0121] Figure 5B shows examples of component-based processes for variant classification or variant interpretation according to some embodiments of the present disclosure.
[0122] In Figure 5B, the computing system includes several datasets 562, a data preparation subsystem 564, a feature store 572, a label store 574, a model experiment 576, a model training component 578, a model registry 582, a model providing component 584, an online prediction 586, a batch prediction / inference component 588, and a batch prediction store 590.
[0123] Dataset 562 contains evidence data that may be relevant to variant classification or interpretation, such as various types of pathogenicity evidence. For a given variant, the evidence data may or may not include labels (e.g., pathogenic, benign, or pathogenicity labels such as VUS).
[0124] Dataset 562 may include a data store and / or a searchable database that stores raw data and / or the outputs of machine learning-based models, such as the outputs of one or more machine learning models. For example, one or more machine learning models may be included in an evidence modeling platform. Dataset 562 may include internally generated data and / or data obtained from one or more external sources. For example, dataset 562 may include one or more of the following: CVM dataset (including patient clinical data modeled and evaluated via EMP), PFM dataset (population frequency data modeled and / or evaluated via EMP), FIM dataset (evolutionary conservation, protein structure, and protein stability data evaluated via EMP), DMA dataset (published in vitro experimental study data modeled and / or evaluated via EMP), DML (in-house performed in vitro experimental study data modeled and / or evaluated via EMP), splice site loss predictions from external sources such as PANGOLIN models, splice site gain predictions from external sources such as PANGOLIN models, allele frequency (AF) data, or predictions of nonsense mutation-dependent degradation, or human-managed data based on the application of a variant classification framework (e.g., the Sherloc framework).
[0125] The data preparation subsystem 564 includes a feature ingestion and preprocessing component 566, a label pipeline 568, and a data versioning component 570. The feature ingestion and preprocessing component 566 ingestion and preprocesses data obtained individually or in combination from one or more of the datasets 562, which may or may not include labels. For example, the data can be fetched and preprocessed as described with reference to Figures 1A, 3 and / or 5A. The feature ingestion and preprocessing component 566 can create a combined or linked set of features extracted from different types of evidence obtained from multiple different datasets 562. Features extracted from the datasets 562 can be annotated with a common identifier, such as a variant identifier. Alternatively or additionally, features extracted from the datasets 562 can be combined, for example, using a common identifier to create a combined set of features that can be input into a machine learning model (e.g., a Bayesian causal model as described herein). The feature acquisition and preprocessing component 566 outputs the feature set to the feature store 572 for use by the model training component 578.
[0126] The label pipeline 568 determines whether labeled data exists for a given type of evidence data for a variant, and, if available, to what extent the label is reliable (e.g., obtained from a reliable source or independently verified). The label pipeline 568 can assign VUS labels to variants that otherwise do not have labels. The label pipeline 568 aligns the labeled data with their respective variant identifiers and outputs the labeled data to the label store 574 for use by the model training component 578.
[0127] The data versioning component 570 applies timestamp data to the feature set output by the feature acquisition and preprocessing component 566 and the label data output by the label pipeline 568 in order to synchronize and manage different versions of the feature set and the corresponding label data, if any. The label data and feature set are stored separately in the example in Figure 5B (for example, the feature set is stored in the feature store 572 and the label data is stored in the label store 574). In this example, the separate storage of these data facilitates unsupervised training of one or more parts of the machine learning model by the model training component 578, while also enabling supervised training of other parts of the machine learning model. This example is illustrative, and the label data and feature set may be stored together without departing from the scope of this disclosure.
[0128] Once prepared, the labels and feature sets are used for training / inference. The model training component 578 retrieves training data from one or more of the feature store 572 or the label store 574 and applies the machine learning model to the training data. In the context of the Bayesian causal model described herein, inference and / or training may be used to refer to the process of fitting the Bayesian model to the training data before the Bayesian model is used for prediction. As described above, a Bayesian model can be developed using various combinations of labeled and / or unlabeled training data. In this regard, the model training component 578 treats labels as simply another variable that may or may not be observed for a variant.
[0129] Fitting a Bayesian model to training data involves iteratively applying model experiments 576 to the Bayesian model (e.g., in a closed loop) until the model training component 578 meets applicable performance criteria in the decision block 580, to examine how the model has responded to the training data (e.g., updates to weight values, parameter values, etc.). The model experiments 576 and applicable performance criteria may include one or more of the experiments described, for example, with reference to Figure 4. The decision block 580 represents a function included in the model training component 578, but is shown separately in Figure 5B for illustrative purposes. When new training data becomes available, the Bayesian model is fitted to the new data, thereby creating a new version of the model. In this way, the Bayesian model naturally accommodates sparse or missing data and adapts as more complete data becomes available.
[0130] The model training component 578 allows the Bayesian model to be fitted to the training set of variant data independently of labels, so that any or all of the variant-related data can be used for inference, with or without labels (for example, even if labels are not observed). Furthermore, the feature set can be arbitrary in the sense that, in a given version of the feature set, some features may be observed for a given variant, while others may not be observed. The Bayesian model can handle such arbitrary feature sets in that it can use observed features for inference, although it must be noted that unobserved features remain unobserved.
[0131] For example, if a request or query identifies a variant and only one feature, the Bayesian model can still provide some information about the distribution of that feature, but with a level of uncertainty that reflects the fact that only one feature is known. Another example is when the Bayesian model does not have labels coupled to the variant's feature set; the model can still provide information about the feature distribution of the known features. As mentioned above, any node in the graphical representation of the Bayesian model can be sampled, and the output from the sampled node can provide predictive information about the corresponding variable.
[0132] The model training component 578 continues to apply the training data and model experiments 576 to the model until the determination block 580 determines that the applicable performance criteria have been met. In response to the determination block 580 determining that the applicable performance criteria have been met, the trained model is registered and stored in the model registry 582 and made available to the batch prediction / inference component 588. The model registry 582 can store one or more different versions of the Bayesian model. The batch prediction / inference component 588 generates batch predictions using the trained model and stores the batch predictions in the batch prediction store 590. The batch prediction store 590 can be queried for prediction data generated by the trained model. A batch prediction is a prediction provided by a model trained across a large number of variants with known labels.
[0133] The model registry 582 makes the trained models accessible to the model provider 584, for example, via an API. The model provider 584 samples or queries the model registry 582 in response to requests or queries received, for example, via one or more user devices over a network. In response to a request or query, the model provider 584 samples the trained Bayesian model and uses the sampled model output to generate an online prediction 586, for example, for presentation via one or more user devices. An example scenario involving the online prediction 586 is a clinical application where an individual's test results identify a new variant that has not been observed previously. In this scenario, the model provider 584 can query the Bayesian model in real time to retrieve predictions and return predictions in response to queries.
[0134] The examples shown in Figure 5B and the accompanying description above are provided for illustrative purposes only. This disclosure is not limited to the examples described. Additional or alternative details and implementations are described herein.
[0135] Figure 6 shows examples of evidence-based variant classification or variant interpretation platforms according to several embodiments of the present disclosure.
[0136] In Figure 6, the evidence-based modeling platform 600 provides a processing pipeline 602. Pipeline 602 is used to generate, evaluate, and integrate various sources of evidence for clinical variant classification or interpretation using different sources of input data. A wide range of machine learning (ML) models / algorithms are used in conjunction with domain knowledge (e.g., clinical expertise) to generate and output interpretations of gene variants. Different models are queried to obtain information about different attributes of gene variants that may or may not contribute to their pathogenicity. For example, attributes such as allele frequencies, functional data, and the resulting protein structure stability can be modeled using different modeling approaches, and pipeline 602 can select and retrieve evidence from these different models. Platform 600 provides a framework for evaluating and using diverse evidence to resolve variants that are otherwise classified as having uncertain significance (VUS) with respect to health status. For example, the configuration of pipeline 602 allows platform 600 to improve variant classification or interpretation for a wide range of genetic conditions, resulting in a reduction of VUS, particularly in ancestral groups that are not well represented in genome studies and databases.
[0137] One study measured the performance of Platform 600 against established approaches to variant classification and evaluated its impact on resolving variant unresolved syndromes (VUS) across various clinical areas and diverse ancestral populations. The hypothesis was that utilizing Platform 600 to assess and integrate novel evidence on a scale could accelerate the rate of VUS resolution, particularly for individuals from ancestral populations that are not well represented in genomic studies, and thus lead to improved fairness in variant classification or interpretation. Anonymized patient data were approved for analysis under Western Independent Review Board protocol number without requiring further individual informed consent. This study followed the Strengthening of Reporting in Observational Studies in Epidemiology for Cohort Studies (STROBE) reporting guidelines. Next-generation sequencing data and clinical and demographic information were obtained for individuals referred to laboratories where clinician-directed germline testing of single-gene or multi-gene panels was performed. Individuals who requested data deletion or opted out of data use were excluded from the study.
[0138] The variants were classified as benign, likely benign, VUS, likely pathogenic, or pathogenic using the Sherloc framework, a validated system based on guidelines from the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Sherloc classifies variants using semi-quantitative, point-based rules to match the distinct contributions of heterogeneous evidence types to the classification decision.
[0139] In this study, 27,477 out of 496,209 novel variants were used to validate Platform 600 because they received evidence from at least one model and obtained a non-VUS classification regardless of that evidence. Across 8,517 unique pathogenic variants and 14,960 unique benign variants, the model combinations provided in Pipeline 602 achieved a 98% negative predictive value and a 90% positive predictive value. Overall, 22.3% of individuals participating in the study had at least one variant classified based on input from one or more models in Platform 600 (e.g., 21.6% from VUS to B / LB, and 1.4% to P / LP). The impact varied by clinical area, being greatest among individuals tested for immunology (60.5%) and neurology (48.5%). The impact also varied by ancestry, with the greatest impact observed among individuals with Asian (33.7%) and Black (28.7%) ancestry compared to Caucasian individuals (18.5%, two-tailed t-test p<1e-10). Results from this study demonstrate that a systematic approach to integrating machine learning models into variant classification can yield highly accurate interpretations, enabling a significant proportion of individuals with VUS to receive more definitive results. The results reflect the fact that models in Platform 600 are not biased by ancestral origins of evidence, are primarily informed by biologically sound hypotheses, and yield greater equivalence in variant classification or interpretation.
[0140] Platform 600 provides an analytical infrastructure for evaluating and weighting diverse evidence for use in variant interpretation at scale. Based on machine learning from labeled examples (e.g., known pathogenic and benign variants), Platform 600 can be used for both gene-specific and genome-wide analyses and is designed to incorporate new evidence through iterative updates over time. The function of Platform 600 is to evaluate whether the type of evidence meets a quality threshold for inclusion in variant interpretation. In one embodiment, if the evidence meets or exceeds the quality threshold, Platform 600 assigns a standardized weight to that value for use in ACMG-based clinical variant classification or interpretation. In contrast to other approaches that can combine heterogeneous data into a metamodel, Platform 600 consists of multiple different evidence models, each independently evaluating distinct categories of evidence and determining whether that evidence can reliably inform variant classification or interpretation.
[0141] In the illustrated example of Platform 600, Platform 600 provides several functions. Firstly, Platform 600 can build and train evidence models (however, this capability is not necessarily required for all models in Platform 600). Models developed using Platform 600 include models generated by supervised machine learning and trained using gene-specific labeled data (i.e., known benign and pathogenic variants) that can be based on a diverse range of evidence types. Examples of models that can be included in Platform 600 include, but are not limited to, CVM (patient clinical data assessed through EMP), PFM (population frequency data assessed through EMP), DMA (publicly available in vitro experimental study data assessed through EMP), DML (in-house performed in vitro experimental study data assessed through EMP), splice site loss prediction (from PANGOLIN), splice site gain prediction (from PANGOLIN), Sherloc EV0016 (nonsense mutation-dependent degradation prediction, human-curated data based on Sherloc application), and / or FIM (evolutionary conservation, protein structure, and protein stability data assessed through EMP).
[0142] Secondly, Platform 600 evaluates models (this functionality applies to all models included in Platform 600). Being a flexible platform, Platform 600 can evaluate both models built using Platform 600, as well as numerous external unsupervised and / or genome-wide algorithms and models. For example, Platform 600 can evaluate whether a particular dataset or model can reliably distinguish benign variants from pathogenic ones. To perform such an evaluation, Platform 600 uses a holdout set of labeled data (e.g., a portion of the training data not used to train the model) to test how well a given model separates known benign and pathogenic variants. Platform 600 includes evidence from models that meet or exceed a defined quality threshold in variant interpretation.
[0143] Thirdly, Platform 600 can be used to predict the pathogenicity of novel variants and / or previously classified variants. The platform uses model evidence that meets or exceeds quality thresholds to predict the pathogenicity of variants that otherwise have uncertain clinical significance.
[0144] Fourth, Platform 600 standardizes pathogenicity predictions across various evidence-based models. Platform 600 calibrates predictions for each variant using empirical data (e.g., adjusting weight values) so that the strength (or confidence) of the prediction can be represented by attribute points (e.g., 2P, 1B). This score is determined by how clearly the model separates known benign and pathogenic variants. In other words, the evidence is weighted according to the model's predictive performance (i.e., PPV and NPV) for known variants. This evidence-based weighting enables Platform 600 to standardize diverse data and models (e.g., population allele frequencies, evolutionary conservation, physicochemical molecular properties, etc.) into a common framework that can be used for variant classification or interpretation. Platform 600 evaluates and standardizes diverse models and data types. In one embodiment, the output generated by Platform 600 (e.g., certified and weighted evidence) can be input into an ACMG-based clinical variant interpretation process that can yield pathogenicity / benign scores. In another embodiment, the output generated by platform 600 (e.g., certified and weighted evidence) can be input into a probabilistic or causal (e.g., Bayesian) clinical variant interpretation process.
[0145] Platform 600 is a flexible platform designed to address the specific challenges of evidence-driven variant interpretation in the age of big data. By using an adaptable framework based on supervised learning from labeled data, Platform 600 can estimate predictive values for many different types of data, thereby bringing heterogeneous forms of evidence into a common quantitative system for use in variant interpretation.
[0146] Before model construction or evaluation, labels can be calculated at the genomic, premRNA, mRNA, or protein level, depending on the nature of the evidence. Similarly, when building models, Platform 600 can use gene-specific (most common) or genome-wide data, as long as it is biologically appropriate. Because gene products have unique features and properties, many models do not generalize well across all genes, leading to inaccurate predictions. Rather, models often perform better (with respect to PPV and NPV) when trained on gene-specific data. For example, a model of protein tertiary structure may be a better predictor of pathogenicity for proteins with structural function in cells (e.g., ion channels) than for cytoplasmic enzymes where only a small portion of the protein (e.g., the active site) may be required for function. Gene-specific model training and evaluation can enable Platform 600 to determine which type of data produces the best predictive model for each gene. The platform also allows training across genes, if necessary, for example, when an insufficient number of labels are available for a given gene.
[0147] In contrast, other biological phenomena are not specific to individual genes (e.g., RNA splicing, allele frequencies) and are therefore modeled using genome-wide data. Importantly, model evaluation and standardization steps (using labels) can be applied to both the data generated by Platform 600 and external algorithms and models such as the Evolutionary Model of Variant Effects (EVE).
[0148] Furthermore, model evaluation and standardization offer greater flexibility by allowing the use of different types of labeled data beyond known pathogenic and benign variants. For example, Platform 600 uses empirically measured splicing results to evaluate the accuracy of in silico SpliceAI predictions. The flexible design of Platform 600 allows for the application of new evidence and new tools as they emerge.
[0149] Platform 600 is designed to update its predictive models using machine learning iterations, rather than statically, whenever new information becomes available. Therefore, the models included in Platform 600 are regularly updated and continuously refined based on new empirical evidence. This ensures that the models included in Platform 600 leverage the latest clinically validated information.
[0150] Platform 600 includes a circular filter 604. The circular filter 604 filters the model input data for circularity before using it in future iterations. For example, previous predictions or variant classifications / interpretations generated based on the evidence provided by Platform 600 may be filtered by the circular filter 604 so that they are not used as training data in future iterations.
[0151] In contrast to some other tools, Platform 600 can make predictions rooted in biological processes, and it maintains a transparent connection to those primary data. Independent evidence types are modeled, evaluated, and standardized separately, rather than being mixed together. Pathogenicity predictions derived from each piece of evidence are each passed to a separate ACMG-based variant interpretation process. This allows variant classification scientists using predictions based on Platform 600 to identify the types of data that are key determinants of the output score and to evaluate the evidence provided by Platform 600 in the context of other available information. This is in contrast to ensemble or meta-predictors, which typically combine different evidence types, thereby obscuring information about which types of data drive its output score and creating a risk of double-dipping during variant interpretation.
[0152] Furthermore, because many of the models in Platform 600 are based on fundamental biology (e.g., physicochemical properties), they are less susceptible to biases associated with uneven representation of ancestral groups in genome databases, resulting in improved equity in variant interpretation.
[0153] Referring to Figure 6, pipeline 602 includes software-based components for model training, model evaluation, prediction using the trained model, and standardization of the model output.
[0154] The model training component trains the machine learning algorithms used by the model using data of variants with known pathogenicity. For each model, the model training component applies the applicable machine learning algorithms to one or more training datasets until the model output meets or exceeds one or more performance criteria (for example, the model converges in that the difference between the model output and the expected output consistently falls below an error tolerance threshold level). As part of the model training process, the model training component includes a feature generation process and a label generation process.
[0155] For feature generation, the model training component can convert input data from an assay, dataset, or method into a format that the model can accept as input (e.g., tensor, embedding, etc.), and / or the model can combine data from different sources that have a common unique identifier (ID), such as a variant ID. For label generation, the model training component can obtain variant pathogenicity label data (e.g., benign, pathogenic, etc.) from one or more data sources, such as the ClinVar database, EVE, and / or others. The model training component can timestamp the label data and associate it with an applicable variant ID. Label data may be stored separately from model feature data. For example, some models may be trained only on unlabeled feature data, other models may be trained only on labeled data (in this case, the feature data and label data may be combined by variant ID), and / or yet another model may be trained on both unlabeled and labeled data (e.g., the training dataset may include data on variants whose pathogenicity has not been determined, as well as data on variants with known pathogenicity labels).
[0156] The model evaluation component evaluates the quality of the output of the model trained by the model training component and, based on the evaluation, determines whether to include the model in Platform 600. For example, the model evaluation component applies a holdout dataset (e.g., a portion of the training dataset not used to train the model, containing data on variants with known pathogenicity values) to the trained model and compares one or more model performance metrics (e.g., AUROC, area under the receiver operating curve, or others) to the metric threshold. The performance metric could, for example, measure the ratio of true positives to false positives in the model output. If the ratio falls below the threshold (e.g., there are more false positives than true positives in the model output), the model may be rejected. If the ratio meets or exceeds the performance threshold, the model may be accepted for inclusion in Platform 600.
[0157] The prediction component of pipeline 602 uses only the best-performing model among those evaluated by the model evaluation component for prediction. For example, upon encountering a novel variant (e.g., a previously unlabeled variant), the prediction component applies one or more models accepted into platform 600 by the model evaluation component to the data associated with the novel variant and uses those models to generate a prediction output, which may include the predicted effect of the novel variant on the predicted pathogenicity, e.g., the likelihood of a particular health condition occurring. The prediction component may select the model to use for prediction based on the data available for the novel variant. For example, the prediction component may select one model if the available data contains data for one biological characteristic, and select different models if the available data contains data for different biological characteristics.
[0158] In some examples, the prediction component employs multiple different models on the platform and, for each model, bins the model outputs using a prediction threshold that determines the relative weights assigned to the outputs of each model.
[0159] The standardization component converts the model output generated by the prediction component into a standardized score that can be used for variant interpretation. The standardized score reflects the weight or reliability of relevant types of evidence with respect to predicting the pathogenicity of the variant. These standardized scores are output by platform 600 and can be stored, for example, in an evidence list, which can be queried by a clinician, for example, for clinical variant interpretation. Once a variant is newly interpreted using the model output provided by platform 600, this new data can be used to iteratively update one or more models within platform 600. For example, the new data can be used to improve previously approved models, or to improve previously rejected models so that they may be approved for use within the platform, or to create new models for platform 600. A circular filter 604 can be used to determine whether to update platform 600 with a given set of new data.
[0160] Platform 600 is adaptable to a wide range of training and validation data. For example, artificial intelligence (AI)-based in silico splicing predictions can be validated using empirical splicing data. For instance, if gene-specific EVE is not possible due to insufficient gene-level labeling data, genome-wide gwEVE, which relies on genome-wide (gene-independent) labeling data, can be used as an alternative.
[0161] In several studies, information on patients' age, sex, and genetic ancestry was provided by the clinician who requested the test. For genetic ancestry, clinicians selected from predefined genetically similar groups (referred to as racial / ethnic groups on the test request form): Ashkenazi Jews, Asians, Black, French Canadians, Hispanics, Native Americans, Pacific Islanders, Sephardic Jews, White (non-Jewish and non-French Canadians for the purposes of this study), or Other. If "Other" was selected, a free-text response could be added. If the free-text response matched a predefined group, the individual was included in that group. Individuals with more than one reported genetic ancestry were grouped as multiple.
[0162] In several studies, the clinical domains included cardiology, oncology (hereditary cancer), immunology, hereditary metabolic disorders, and neurology. Both clinical domains were considered when a gene was included in a test associated with two clinical domains (e.g., NF1 in both hereditary cancer and neurology). Because some genes are involved in multiple hereditary disorders, patient attributes were analyzed based on the ordered panel's clinical domains (i.e., test-clinical domains) and variant attributes based on primary gene-disease relationships (i.e., gene-clinical domains).
[0163] To examine the performance of the models in Platform 600, we evaluated / compared the agreement between the classification of variants made without evidence provided by Platform 600 and the predictions output by Platform 600 for those variants. Novel variants observed internally between May 1, 2022 and May 1, 2023 (N=496,204 unique variants) were selected for examination. The analysis was limited to a subset of variants that received evidence from at least one model in Platform 600, and we estimated the non-performing variable value (NPV) and performing variable value (PPV) of variants that reached classification without evidence provided by Platform 600 (8,517 unique pathogenic variants and 14,960 unique benign variants) compared to the evidence provided by Platform 600.
[0164] Additionally, the accuracy of predictions output by Platform 600 was prospectively evaluated against externally validated variants. More specifically, for variants classified as (likely) pathogenic or benign using predictions generated by Platform 600, the number and frequency of occurrences of those same variants later classified as pathogenic or benign by external laboratories using independent new data were evaluated. In some evaluations, the analysis was limited to novel ClinVar variants that reached classification between May 1, 2022 and May 1, 2023. Only variants that reached classification and were submitted by third parties were considered.
[0165] The impact of the machine learning model integrated into variant classification was also evaluated using an internal proprietary system. To do so, we sampled 10 subsets of 5,000 patients observed internally between May 1, 2022, and May 1, 2023. Across those patients, all observed variants that received any type of evidence from the modeling platform were removed from the list of evidence for those variants.
[0166] The calculated impact rate corresponds to the proportion of patients with at least one variant that would have been VUS without prediction generated by Platform 600, which subsequently reached a more definitive classification (e.g., likely pathogenic, pathogenic, likely benign, and benign). This impact was calculated across major clinical areas (oncology, metabolism, immunology, neurology, and cardiology) and across patient ethnicity (e.g., for these groups: Caucasian, Latino, Black / African American, Asian, and Ashkenazi Jewish, using self-reported data from laboratory application forms).
[0167] Predictive modeling can help standardize and quantify the interpretation of assay data. By learning from known pathogenic variants (P) or benign variants (B) (effectively positive and negative controls), predictive values for each data type can be determined and mapped to a single standard of "probability of pathogenicity" that can be used for variant classification or interpretation. For example, if the input data (e.g., high-throughput assay data) does not correlate well with known P / B variants, the model cannot be expected to be useful for variant classification or interpretation. Therefore, a performance evaluation step is used to ensure high-quality predictions. This may limit the number of genes that Platform 600 can predict, but it ensures higher reliability in the results.
[0168] The output of Platform 600 may be just one component of a variant classification or variant interpretation system. The results of the predictive model in Platform 600 can be evaluated, for example, by a clinical genomics scientist, a variant classification or variant interpretation system, along with all other data related to variant classification or interpretation, such as patient phenotypes and isolation data, in order to arrive at a final variant classification or otherwise interpret the variant.
[0169] Platform 600 utilizes diverse models to address challenges specific to different genes and gene variants. Platform 600 implements complex interactions of various models, continuously updated with the addition of newly validated models within the context of clinical expertise. Different models are used to examine various attributes of gene variants that may or may not contribute to pathogenicity, such as allele frequencies, functional data, and the stability of the resulting protein structures. Examples of such models that may be included in Platform 600 are shown in Figure 7. Each such model, leveraging diverse data types, enables the elucidation of critical insights into the molecular mechanisms of disease, even while using data that may be biased from highly represented cohorts within the genetic testing landscape.
[0170] Platform 600 can generate variant pathogenicity scores that can be used in many different variant classification or variant interpretation systems. As mentioned above, Platform 600 includes a suite of evidence models that use different data sources and different modeling techniques to generate predictions for different types of variants (i.e., missense, splicing, etc.). For consistent use of these models in variant classification or variant interpretation systems, the models may be required to pass performance criteria, and if a semi-quantitative interpretation system such as Sherloc is used, the models are thresholded based on performance to provide discrete evidence (i.e., attribute points). Similarly, for use of the outputs of these models as inputs in a probabilistic variant classification model, the models may be required to pass performance criteria.
[0171] Within the Sherloc interpretation system, the model's performance across a set of variants is used to determine the number of points a variant can receive at the gene-based level. During the evidence assignment step, a metric threshold (e.g., a variant with PPV >= 0.95 can receive 2 points) is used to assign points to variants. Filters are applied across groups of the model's functional elements to ensure a minimum quality across relevant functional elements.
[0172] In the first step, validation predictions are used to find segments of the training data that satisfy a specified prediction threshold. In some examples, the predictions generated by Platform 600 are converted into discrete point values. To achieve this, variants are assigned to performance segments using PPV and NPV. Different performance segments have different point value associations based on a specific model type and performance threshold. The transcripts are then filtered by performance. To do this, a performance metric is calculated over the validation data predictions and returns a list of transcripts that pass one or more performance filters (e.g., AUROC≧0.7, AUROC≧0.8, AUROC≧0.9, AUROC≧0.95, AUROC≧0.99). If a transcript does not pass a filter, the point values of all variants within that transcript are set to 0.
[0173] When releasing a model for use in a variant interpretation system, the best-performing model for a given transcript is selected from among competing models (e.g., models from the same category) using points assigned to each variant within that transcript.
[0174] In several examples, SpliceAI validation was performed. SpliceAI is a deep residual neural network that uses only the genomic sequence of an mRNA precursor transcript as input to predict whether each position in the mRNA precursor transcript is a splice donor, a splice acceptor, or neither. SpliceAI was validated using variants in regions of genes potentially affected by splicing. To evaluate SpliceAI's ability to accurately predict changes in splicing, internal data from oncology RNA assays were used to select variants that were associated with splicing changes (true positive set) or not (true negative set) in at least one of our patients. In total, a set of 1236 variants was considered a positive control and 81 variants were considered a negative control. SpliceAI predictions were considered positive if all scores for a given variant were greater than or equal to 0.2. Overall, SpliceAI yielded a sensitivity of 91.34% and a specificity of 79% across the variants tested and was incorporated into Platform 600.
[0175] The accuracy of pathogenicity predictions generated using Platform 600 was evaluated using a set of known pathogenic and benign variants, and high agreement with independent variant classification methods was found. A retrospective cohort study was used to evaluate the platform's VUS reclassification performance against external independent laboratories, again finding high agreement.
[0176] The overall performance of a given model may vary from gene to gene. A model type may function well for a particular gene or group of genes, but not well for others, based on the characteristics of the gene. Models that can adequately distinguish between pathogenic and benign variants receive higher weight for their predictions in pipeline 602 for variant classification or interpretation. Conversely, models that show insufficient distinction between pathogenic and benign variants in a particular gene are not used in pipeline 602. This gene-specific validation of each model is useful when considering the integration of predictions from the model into platform 600.
[0177] The example shown in Figure 6 and the accompanying explanation above are provided for illustrative purposes only. This disclosure is not limited to the examples described. Additional or alternative details and implementations are described herein.
[0178] Figure 7 is a table listing examples of models that may be included in an evidence-based variant interpretation platform according to several embodiments of the present disclosure. One or more of the models described in Figure 7 may be incorporated into the platform 600 described with reference to Figure 6, in accordance with the model evaluation and other processes described with reference to Figure 6.
[0179] The example shown in Figure 7 and the accompanying description above are provided for illustrative purposes only. This disclosure is not limited to the examples described. Additional or alternative details and implementations are described herein.
[0180] Figures 8A, 8B, 8C, and 8D show examples of experimental results for probabilistic variant interpretation models according to several embodiments of the present disclosure.
[0181] Figure 8A shows an example histogram, which illustrates the distribution of pathogenicity predictions generated by the Bayesian model described herein across variants in the test set (not used for training). Figure 8B shows an example histogram illustrating the distribution of pathogenicity predictions generated by the same Bayesian model across variants in the VUS set (variants of unknown significance). For both examples in Figures 8A and 8B, the Bayesian model was trained on the MMR gene. Figure 8A shows that for the test set, known pathogenic variants tend to receive a high pathogenicity probability, while known benign variants tend to receive a low pathogenicity probability. These results indicate that the model tends to separate known pathogenic variants from known benign variants. Figure 8B shows that for the VUS set, variants tend to be split between low and high pathogenicity probabilities, suggesting that some variants may be considered either benign or pathogenic, respectively.
[0182] Figure 8C shows an example of a calibration plot of a Bayesian model configured as described herein, for example, the performance of various implementations of the Bayesian model. In Figure 8C, the x-axis is the predicted probability of pathogenicity generated by the various models, and the y-axis is the proportion of variants in the predicted probability of pathogenicity known to be pathogenic. A fully calibrated system would show that 20% of variants with a score of 0.2 are pathogenic, and 80% of variants with a score of 0.8 are actually pathogenic, and so on. Each line in the plot represents a different model implementation. The line labeled "pyro_02.csv" shows the results of a probabilistic graphical model-based implementation.
[0183] Figure 8D illustrates pairwise analysis for an exemplary probabilistic graphical model-based implementation. Figure 8D includes a series of heatmaps showing learned pairwise interactions between various features used in the probabilistic graphical model-based implementation. Regions of the heatmap corresponding to higher values on the x and y axes (typically the top and / or right of the map) indicate a higher probability of pathogenicity. Regions of the heatmap corresponding to lower values on the x and / or y axes (typically the lower and / or left of the map) indicate a lower probability of pathogenicity. For example, the DMA vs. FIM heatmap shows that the model has learned that the probability of pathogenicity is primarily driven by DMA (e.g., a high DMA score indicates a high probability of pathogenicity, regardless of the FIM score). In another example, the DMA vs. CVM heatmap shows that the model has learned that both DMA and CVM scores contribute approximately equally to the final probability of pathogenicity (e.g., a variant with a high CVM score can still have a low probability of pathogenicity if the DMA score is low, and vice versa).
[0184] The examples shown in Figures 8A, 8B, 8C, and 8D, and the accompanying descriptions above, are provided for illustrative purposes only. This disclosure is not limited to the examples described. Additional or alternative details and implementations are described herein.
[0185] Figures 9A, 9B, 9C, 9D, 9E, 9F, and 9G illustrate methods, systems, apparatus, and / or non-temporal computer-readable media configured to create and use probabilistic models for variant interpretation, according to some embodiments of the present disclosure.
[0186] Figure 9A is a flowchart of an exemplary method 900. Parts of method 900 may be embodied in one or more non-temporary computer-readable media, executed by one or more processors, and / or implemented in one or more components of a computing system. Exemplary examples of parts of method 900 are further described throughout this disclosure with reference, for example, to Figures 1A, 1B, 2A–2F, 3, 4, 5A, and 5B.
[0187] In operation 902, the processing device uses pathogenicity evidence data associated with gene variants and health status to create input data for a causal machine learning model. Pathogenicity evidence data is obtained through multiple evidence sources. In some examples, for a category of pathogenicity evidence data, the multiple evidence sources include alternative types of models for predicting pathogenicity given the category of pathogenicity evidence data, and the processing device uses at least one evaluation criterion to create input data using a model selected from among the alternative types of models. In some examples, part of operation 902 is performed by a data preprocessing component of the computing system.
[0188] In operation 904, the processing device applies a causal machine learning model to the input data to generate a trained causal model. The graphical representation of the trained causal model includes nodes connected via acyclic directed edges. The first of the nodes represents a pathogenicity evidence variable among the pathogenicity evidence variables related to health status. At least one second node among the nodes represents the cause of the pathogenicity evidence variable. At least one third node among the nodes represents the effect of the pathogenicity evidence variable. The acyclic directed edges represent the relationship between two of the nodes.
[0189] In some examples, in the graphical representation of the trained causal model, at least one second node is upstream of the first node, and at least one third node is downstream of the first node and at least one second node. In some examples, in the graphical representation of the trained causal model, at least one second node includes at least one variable representing data about at least one biological trait that causes pathogenicity of a gene variant with respect to health status. In some examples, part of operation 904 is performed by the model training component of the computing system. In some examples, in the graphical representation of the trained causal model, at least one third node includes at least one variable representing data about at least one biological trait that is the effect of pathogenicity of a gene variant with respect to health status. In some examples, the graphical representation of the trained causal model includes at least one of gene-level variables, gene-level constraints, or variant-level constraints.
[0190] In some examples, at least one variable includes at least one of the following: a variable relating to the effect of a gene variant on protein function; a variable relating to the effect of a gene variant on protein stability; a variable relating to the effect of a gene variant on mRNA processing; a variable relating to the effect of a gene variant on nonsense mutation-dependent degradation; a variable relating to the effect of a gene variant on mRNA; a variable relating to the effect of a gene variant on protein expression; a variable relating to the biological importance of disrupted DNA; a variable relating to the biological importance of disrupted RNA nucleotides; or a variable relating to the biological importance of disrupted protein amino acid residues. In some examples, at least one variable includes at least one of the following: a variable relating to the effect of pathogenicity on individual patient phenotypes; a variable relating to the effect of pathogenicity on population allele frequencies; a variable relating to the effect of pathogenicity on selected cohort allele frequencies; or a variable relating to the effect of pathogenicity on familial genotype-phenotypic cosegregation patterns.
[0191] In operation 906, the processing device outputs predictive data sampled from a trained causal model in response to a request. In some examples, the processing device uses the trained causal model to generate and output predictions about whether a gene variant is benign or pathogenic in terms of health status. In some examples, the processing device provides the predictions about whether a gene variant is benign or pathogenic to a clinician for use in formulating a patient's diagnosis. In some examples, the processing device stores the predictions for retrieval via at least one query.
[0192] In some examples, the processing device applies at least one performance criterion to the trained causal model, for example, through the model evaluation component of the computing system. In some examples, the processing device samples data from any node of the trained causal model, and the nodes sampled by the prediction component correspond to unobserved variables with respect to gene variants and health status. In some examples, part of operation 906 is performed by the prediction component of the computing system.
[0193] Figure 9B is a flowchart of an exemplary method 910. Method 910 can be used to create a training dataset for a causal machine learning model. Parts of Method 910 may be embodied in one or more non-temporal computer-readable media, executed by one or more processors, and / or implemented in one or more components of a computing system. Exemplary examples of parts of Method 900 are further described throughout this disclosure with reference, for example, to Figures 1A and 3.
[0194] In operation 912, the processing device applies a natural language processing (NLP) model to clinical phenotypic data of gene variants. In some examples, the processing device uses the NLP model to extract clinical phenotypic data from at least one field of a test application form, where at least one field includes at least one of the indication field or family history field. In some examples, the processing device obtains data related to gene variants through multiple sources, filters the data using at least one filtering criterion, converts the filtered data into a standardized format, and includes the standardized formatted data in a training dataset. In some examples, the data includes at least one of the outputs of at least one machine learning model, empirical measurements, predicted variant effects, or a public dataset. In some examples, the standardized format includes a data structure that can be used as input to a causal machine learning model. In some examples, the processing device splits the data into at least one training dataset for training a causal machine learning model and at least one holdout dataset for validating the trained causal machine learning model.
[0195] In operation 914, the processing device uses an NLP model to identify features in clinical phenotypic data that predict the molecular diagnosis of a health condition related to a gene variant. In operation 916, for each subject in the target population, the processing device uses the identified features and demographic information to calculate a first score, which includes the probability that the subject has a certain health condition. In operation 918, for the target population, the processing device includes the first score in the training dataset. In some examples, the processing device includes labeled pathogenicity data in the training dataset.
[0196] Figure 9C is a flowchart of an exemplary method 920. Method 920 may be used to train a causal machine learning model for variant interpretation. Parts of Method 920 may be embodied in one or more non-temporal computer-readable media, executed by one or more processors, and / or implemented in one or more components of a computing system. Exemplary examples of parts of Method 900 are further described throughout this disclosure with reference, for example, to Figures 1A, 1B, 2A–2F, 3, 4, 5A, and 5B.
[0197] In operation 922, the processing device creates model input data using variant pathogenicity evidence data. In operation 924, the processing device applies a causal machine learning model to the model input data, the graphical representation of the causal machine learning model including nodes connected via acyclic directed edges, where first nodes represent pathogenicity variables, at least one second node represents the causes of the pathogenicity variables, at least one third node represents the effects of the pathogenicity variables, and at least one acyclic directed edge represents the relationship between two of the nodes. In operation 926, the processing device iteratively evaluates the data sampled from the causal machine learning model until at least one performance metric exceeds at least one threshold performance criterion.
[0198] In some examples, the processing device is used to build a causal machine learning model by, for example, linking at least one second node to a first node via at least one upstream directed acyclic edge, and linking at least one third node to the first node via at least one downstream directed acyclic edge. In some examples, the processing device is used to build a causal machine learning model using domain knowledge to identify at least one cause of a pathogenicity variable and assign at least one cause of the pathogenicity variable to at least one second node, identify at least one effect of the pathogenicity variable and assign at least one effect of the pathogenicity variable to at least one third node, link at least one second node to a first node via at least one upstream directed acyclic edge, or link at least one third node to a first node via at least one downstream directed acyclic edge. In some examples, Method 920 involves constructing a causal machine learning model by fitting a hierarchical Bayesian inference model to model input data, which includes phenotypic data and variant pathogenicity data.
[0199] In some examples, Method 920 includes obtaining model input data through multiple sources and filtering the model input data using domain knowledge, where the domain knowledge includes the types of relationships between variables, and the variables include probability distributions, and the Method further includes building a causal machine learning model by assigning weights to the probability distributions using the domain knowledge.
[0200] Figure 9D is a component-based flowchart of an example of communication between components of a system or apparatus 930. Exemplary examples of parts of the system or apparatus are further described throughout this disclosure with reference to, for example, Figures 1A, 1B, 2A-2F, 3, 4, 5A, 5B, 6, 10, and 11.
[0201] The system or device 930 includes at least one processor 932 and at least one memory 934. At least one memory 934 includes a causal machine learning model. The graphical representation of the causal machine learning model includes nodes connected via acyclic directed edges, where a first node represents a pathogenicity variable, at least one second node represents the cause of the pathogenicity variable, at least one third node represents the effect of the pathogenicity variable, and at least one acyclic directed edge represents the relationship between two of the nodes. In some examples, at least one memory further includes submodels, each submodel capable of providing an output to a corresponding node in the graphical representation of the causal machine learning model. In some examples, at least one memory further includes model abstraction components that can determine the output types of the submodels and specific interactions between the submodels. In some examples, the system or apparatus 930 further includes at least one first device that can receive requests and, in response to those requests, provide outputs generated by a causal machine learning model to at least one of a data store, a model, a network, a user interface, or at least one other device.
[0202] Figure 9E is a flowchart of an exemplary method 940. Method 940 may be used to interpret gene variants. Parts of Method 940 may be embodied in one or more non-temporal computer-readable media, executed by one or more processors, and / or implemented in one or more components of a computing system. Exemplary examples of parts of a system or apparatus are further described throughout this disclosure with reference, for example, to Figures 1A, 1B, 2A–2F, 3, 4, 5A, 5B, 6, and 7.
[0203] In operation 942, the processing device samples nodes of a causal machine learning model to obtain sampled data, where the nodes are associated with gene variants, and the graphical representation of the causal machine learning model includes nodes connected via acyclic directed edges, where the first of the nodes represents a pathogenicity variable, at least one second of the nodes represents the cause of the pathogenicity variable, at least one third of the nodes represents the effect of the pathogenicity variable, and at least one acyclic directed edge represents the relationship between two of the nodes. In some examples, sampling a node includes sampling the posterior predictive distribution associated with the node.
[0204] In operation 944, the processing device uses the sampled data to determine and output an interpretation of the gene variant. In some examples, the interpretation includes the probability that the gene variant is pathogenic to a health condition. In some examples, the processing device receives a request via the device and provides the interpretation of the gene variant to at least one of the data store, network, model, or device.
[0205] Figure 9F is a flowchart of an exemplary method 950. Parts of method 950 may be embodied in one or more non-temporary computer-readable media, executed by one or more processors, and / or implemented in one or more components of a computing system. Exemplary examples of parts of a system or device are further described throughout this disclosure with reference, for example, to Figures 1A, 1B, 2A-2F, 3, 4, 5A, 5B, 6, and 7.
[0206] In operation 952, the processing device uses pathogenicity evidence data associated with gene variants and health status to create input data for a causal machine learning model, the pathogenicity evidence data being obtained through multiple evidence sources, the multiple evidence sources including at least one of patient clinical data, population frequency data, published in vitro experimental study data, internally performed in vitro experimental study data, splice site loss prediction, splice site gain prediction, or prediction of nonsense mutation-dependent degradation curated based on the application of a variant classification or variant interpretation framework.
[0207] In some examples, multiple evidence sources include patient clinical data assessed through the Evidence Modeling Platform (EMP), population frequency data assessed through the EMP, published in vitro experimental study data assessed through the EMP, internally performed in vitro experimental study data assessed through the EMP, splice site loss predictions, splice site gain predictions, and predictions of nonsense mutation-dependent decomposition curated based on the application of variant classification or variant interpretation frameworks.
[0208] In operation 954, the processing device applies a causal machine learning model to the input data to generate a trained causal model, the graphical representation of which includes nodes connected via acyclic directed edges, where first nodes represent pathogenicity variables related to health status, at least one second node represents the causes of the pathogenicity variables, at least one third node represents the effects of the pathogenicity variables, and the acyclic directed edges represent the relationships between any two of the nodes. In operation 956, in response to a request, the processing device outputs predictive data sampled from the trained causal model.
[0209] Figure 9G is a flowchart of an exemplary method 960. Parts of method 960 may be embodied in one or more non-temporary computer-readable media, executed by one or more processors, and / or implemented in one or more components of a computing system. Exemplary examples of parts of a system or device are further described throughout this disclosure with reference, for example, to Figures 1A, 1B, 2A-2F, 3, 4, 5A, 5B, 6, and 7.
[0210] In operation 962, the processing device applies at least one evaluation criterion to one or more outputs of a plurality of machine learning-based evidence models, the plurality of machine learning-based evidence models including at least one of a first model of patient clinical data, a second model of population frequency data, a third model of published in vitro experimental study data, a fourth model of in vitro experimental study data, a fifth model of splice site loss prediction, a sixth model of splice site gain prediction, or a seventh model of nonsense mutation-dependent decomposition prediction. In some examples, the plurality of machine learning-based evidence models include a first model of patient clinical data, a second model of population frequency data, a third model of published in vitro experimental study data, a fourth model of in vitro experimental study data, a fifth model of splice site loss prediction, a sixth model of splice site gain prediction, and a seventh model of nonsense mutation-dependent decomposition prediction.
[0211] In operation 964, the processing device determines one or more weight values for the outputs of multiple machine learning-based evidence models. In operation 966, the processing device applies one or more weight values to the outputs of multiple machine learning-based evidence models. In operation 968, in response to a request, the processing device outputs predictive data based on the applied one or more weight values.
[0212] In some examples, Method 960 includes training at least one model from a group of machine learning-based evidence models using supervised machine learning. In some examples, the training further includes training at least one model from a group of machine learning-based evidence models using gene-specific labeled data. In some examples, Method 960 includes at least one of evaluating at least one first model from a group of machine learning-based evidence models individually, or evaluating at least two second models from a group of machine learning-based evidence models in combination.
[0213] The examples shown in Figures 9A, 9B, 9C, 9D, 9E, 9F, and 9G, and the accompanying descriptions above, are provided for illustrative purposes only. This disclosure is not limited to the examples described. Additional or alternative details and implementations are described herein.
[0214] Figure 10 shows an exemplary computing system 1000, including a modeling system, according to some embodiments of the present disclosure.
[0215] In the embodiment of Figure 10, the computing system 1000 includes one or more user systems 1010, a network 1020, an application software system 1030, a modeling system 1050, and a data storage system 1080. The modeling system 1050 includes a model 1052, a model creator component 1054, and a dataset 1060. The modeling system 1050 may include one or more of the components described with reference to any of the figures described above. For example, the modeling system 1050 may include one or more components of the platform 600 described with reference to Figure 6, the system 500 described with reference to Figure 5A, the method 560 described with reference to Figure 5B, a part of the process 400 described with reference to Figure 4, the component-based process 300 described with reference to Figure 3, any of the models described with reference to Figure 1B or any of Figures 2A to 2F, and / or any part of the method 100 described with reference to Figure 1A.
[0216] Model 1052 includes one or more models configured to determine a probabilistic or statistical relationship between inputs and outputs using one or more predictive algorithms. For example, given one or more inputs, Model 1052 outputs labels that can be used to classify the inputs into different categories, or scores that can be used to sort or rank the inputs into groups or ranking lists. An example of Model 1052 includes one or more of the models described above.
[0217] In some embodiments, the model creator component 1054 generates or constructs one or more models of model 1052 by, for example, applying model 1052 to one or more datasets 1060. One or more of the datasets 1060 may include input data and training examples of known or ground truth labels or scores. In some implementations, the predicted outputs generated by one or more models 1052 in response to one or more datasets 1060 are observed iteratively until a set of model validation criteria is met. For example, the difference between the predicted output and the expected output is quantified using a loss function. The model validation criteria are used to determine when one or more models have converged to provide reliable outputs with some degree of certainty. The required level of certainty and validation criteria are determined based on the requirements or design of the particular implementation of one or more models.
[0218] One or more datasets 1060, in some implementations, contain data used to construct one or more models 1052. A dataset 1060 may include, for example, a set of model inputs, such as labeled or unlabeled training data. In some embodiments, one or more of the datasets 1060 may include, for example, a database of historical population data, or be derived therefrom. One or more of the datasets 1060 may include, for example, variant-level data and / or gene-level data.
[0219] The user system 1010 includes at least one computing device, such as a personal computing device, a server, a mobile computing device, or a smart appliance. The user system 1010 includes at least one software application, including a user interface 1012, which is installed on the computing device or accessible via a network. For example, an embodiment of the user interface 1012 includes a graphical display screen that displays controls and graphical elements for operating and / or manipulating one or more of the model 1052, the model creator component 1054, and the dataset 1060.
[0220] The user interface 1012 can be used to input data, initiate user interface events, and view or otherwise perceive output containing data generated by the modeling system 1050. Examples of the user interface 1012 include a web browser, a command-line interface, and a mobile application frontend. The user interface 1012 used herein may include an application programming interface (API).
[0221] The application software system 1030 is any type of application software system that provides or enables the generation, display, or manipulation of the output generated by the modeling system 1050. Examples of the application software system 1030 include, but are not limited to, DNA (deoxyribonucleic acid) analysis software, genetic testing software, medical testing software, health management software, or any combination of any of the above.
[0222] The data storage system 1080 includes a data store and / or data service for storing data received, used, manipulated, and generated by the application software system 1030 and / or the modeling system 1050, such as training data, model parameters, validation criteria, and model outputs. In some embodiments, the data storage system 1080 includes several different types of data storage and / or distributed data services. As used herein, a data service may refer to a physical, geographical grouping of machines, a logical grouping of machines, or a single machine. For example, a data service may be a data center, a cluster, a group of clusters, or a machine.
[0223] The data storage system 1080 resides on at least one persistent and / or volatile storage device that may reside on the same local network as at least one other device of the computing system 1000, and / or on a network that is remote to at least one other device of the computing system 1000. Thus, although shown as being included in the computing system 1000, parts of the data storage system 1080 may also be parts of the computing system 1000, or may be accessed by the computing system 1000 via a network such as network 1020.
[0224] Although not specifically shown, it should be understood that each of the user system 1010, application software system 1030, modeling system 1050, and data storage system 1080, when executed, includes an interface embodied as computer programming code stored in computer memory that enables bidirectional communication between the computing device and any other of the user system 1010, application software system 1030, modeling system 1050, and data storage system 1080 using communication coupling mechanisms. Examples of communication coupling mechanisms include network interfaces, inter-process communication (IPC) interfaces, and application program interfaces (APIs).
[0225] Each of the user system 1010, application software system 1030, modeling system 1050, and data storage system 1080 is implemented using at least one computing device communicatively coupled to the electronic communication network 1020. Any of the user system 1010, application software system 1030, modeling system 1050, and data storage system 1080 may be bidirectionally coupled to the network 1020. The user system 1010 and other different user systems (not shown) may be bidirectionally coupled to the application software system 1030 and / or the modeling system 1050.
[0226] A typical user of user system 1010 may be an administrator or end user of application software system 1030 and / or modeling system 1050. User system 1010 is configured to communicate bidirectionally with application software system 1030 and / or modeling system 1050 via network 1020.
[0227] The features and functions of the user system 1010, application software system 1030, modeling system 1050, and data storage system 1080 may include a combination of automated functions, data structures, and digital data implemented using computer software, hardware, or software and hardware, which are schematically represented in the figure. While the user system 1010, application software system 1030, modeling system 1050, and data storage system 1080 are shown as separate elements in Figure 4 for ease of explanation, this figure does not imply that separation of these elements is necessary unless otherwise stated. Each illustrated system, service, and data store (or their functions) of the user system 1010, application software system 1030, modeling system 1050, and data storage system 1080 can be divided into any number of physical systems, including a single physical computer system, and can communicate with each other in any appropriate manner.
[0228] Network 1020 can be implemented on any medium or mechanism that provides for the exchange of data, signals, and / or instructions between various components of the computing system 1000. Examples of network 1020 include, but are not limited to, a local area network (LAN), a wide area network (WAN), an Ethernet® network or the Internet, or at least one terrestrial link, satellite link or wireless link, or any combination of any number of different networks and / or communication links.
[0229] The example shown in Figure 10 and the accompanying description above are provided for illustrative purposes only. This disclosure is not limited to the example described. Additional or alternative details and implementations are described herein.
[0230] For the sake of clarity, in Figure 11, an embodiment of modeling system 1050 is represented as modeling system 1150.
[0231] Figure 11 shows an exemplary machine of a computer system 1100 capable of executing a set of instructions to cause a machine to perform any of the methods described herein. In some embodiments, the computer system 1100 may correspond to a component of a networked computer system (e.g., the computing system 1000 in Figure 10) that includes, is coupled to, or utilizes a machine for running an operating system to perform the operations described above, corresponding to an aspect of the modeling system 1050 in Figure 10.
[0232] The machine is connected to other machines in a local area network (LAN), intranet, extranet, and / or the internet (e.g., it is networked). The machine can operate as a server or client machine in a client-server network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or client machine in a cloud computing infrastructure or environment.
[0233] A machine is a personal computer (PC), smartphone, tablet PC, set-top box (STB), personal digital assistant (PDA®), mobile phone, web appliance, server, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be performed by such machine. Furthermore, although a single machine is shown, the term “machine” should also be interpreted to include any set of machines that individually or collectively execute a set (or set) of instructions to perform any of the methods described herein.
[0234] An exemplary computer system 1100 includes processing devices 1102, main memory 1104 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM) or rhombus DRAM (RDRAM), etc.), memory 1105 (e.g., flash memory, static random access memory (SRAM), etc.), input / output system 1110, and data storage system 1140, all communicating with each other via a bus 1130.
[0235] The processing device 1102 represents at least one general-purpose processing device, such as a microprocessor or a central processing unit. More specifically, the processing device may be a composite instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor implementing another instruction set, or a processor implementing a combination of instruction sets. The processing device 1102 may also be at least one dedicated processing device, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), or a network processor. The processing device 1102 is configured to execute instructions 1112 for performing the operations and steps discussed herein.
[0236] Instruction 1112 includes parts of the modeling system 1150 when the processing device is executing those parts of the modeling system 1150. Therefore, the modeling system is sometimes indicated by a dashed line as part of instruction 1112 to show that parts of the modeling system are executed by the processing device 1102. For example, when at least some parts of the modeling system are embodied in instructions causing the processing device 1102 to execute in the manner described above, some of those instructions can be loaded into the processing device 1102 from the main memory 1104 and / or the data storage system 1140 (for example, into the internal cache or other memory). However, the entire modeling system does not need to be included in instruction 1112 at the same time; parts of the modeling system may be stored in at least one other component of the computer system 1100 at other times, for example, when at least some parts of the modeling system are not being executed by the processing device 1102.
[0237] The computer system 1100 further includes a network interface device 1108 for communication over the network 1120. The network interface device 1108 provides bidirectional data communication to connect to the network. For example, the network interface device 1108 may be an Integrated Services Digital Network (ISDN®) card, a cable modem, a satellite modem, or a modem for providing data communication connectivity to a corresponding type of telephone line. As another example, the network interface device 1108 may be a local area network (LAN) card for providing data communication connectivity to a compatible LAN. A wireless link may also be implemented. In any such implementation, the network interface device 1108 can transmit and receive electrical, electromagnetic, or optical signals carrying digital data streams representing various types of information.
[0238] A network link can provide data communication to other data devices via at least one network. For example, a network link can provide a connection to a global packet data communication network, commonly referred to as the "Internet," to a host computer via a local network, or to data devices operated by an Internet Service Provider (ISP). The local network and the Internet use electrical, electromagnetic, or optical signals to carry digital data to and from the computer system 1100.
[0239] The computer system 1100 can send messages and receive data, including program code, through the network and network interface device 1108. In the example of the internet, a server can send requested code for an application program through the internet and network interface device 1108. The received code can be executed by the processing device 1102 when it is received and / or stored in the data storage system 1140 or other non-volatile storage for later execution.
[0240] The input / output system 1110 includes a display for displaying information to the computer user, such as a liquid crystal display (LCD) or a touchscreen display, or an output device such as a speaker, a haptic device, or another form of output device. The input / output system 1110 may also include an input device configured to communicate information and command selections to the processing device 1102, such as alphanumeric keys and other keys. The input device may also include, or alternatively, a cursor control such as a mouse, trackball, or cursor directional keys for communicating directional information and command selections to the processing device 1102 and for controlling cursor movement on the display. The input device may also include, or alternatively, a microphone, sensor, or array of sensors for communicating sensed information to the processing device 1102. The sensed information may include, for example, voice commands, audio signals, geographic location information, and / or digital images.
[0241] The data storage system 1140 includes a machine-readable storage medium 1142 (also known as a computer-readable medium) storing at least one instruction set 1144 or software that embodies any of the methods or functions described herein. The instructions 1144 may also be entirely or at least partially present in the main memory 1104 and / or the processing device 1102 during their execution by the computer system 1100, and the main memory 1104 and the processing device 1102 also constitute the machine-readable storage medium.
[0242] In one embodiment, instruction 1144 includes instructions for implementing a function corresponding to a modeling system (e.g., modeling system 1050 in Figure 10).
[0243] The dashed lines in Figure 11 are used to indicate that the modeling system does not need to be fully implemented simultaneously in instructions 1112, 1114, and 1144. In one example, a portion of the modeling system is implemented in instruction 1144, which is loaded into main memory 1104 as instruction 1114, and a portion of instruction 1114 is loaded into processing device 1102 as instruction 1112 for execution. In another example, some portions of the modeling system are implemented in instruction 1144, others in instruction 1114, and yet another portion in instruction 1112.
[0244] Although the machine-readable storage medium 1142 is shown as a single medium in exemplary embodiments, the term “machine-readable storage medium” should be interpreted to include a single or more mediums that store at least one set of instructions. The term “machine-readable storage medium” should also be interpreted to include any medium that can store or encode a set of instructions for machine execution, causing a machine to perform any of the methods of the present disclosure. Accordingly, the term “machine-readable storage medium” should be interpreted to include, but not be limited to, solid-state memory, optical media, and magnetic media.
[0245] Some parts of the detailed description above are presented with respect to algorithms and symbolic representations of operations on data bits in computer memory. These algorithmic descriptions and representations are the methods used by those skilled in the data processing art to most effectively communicate the content of their work to others skilled in the art. An algorithm is considered herein, and also generally, to be a self-consistent set of operations that produce a desired result. An operation is one that requires the physical manipulation of physical quantities. Usually, but not always, these quantities take the form of electrical or magnetic signals that can be stored, combined, compared, and otherwise manipulated. Referring to these signals as bits, values, elements, symbols, characters, terms, numbers, etc., has sometimes proven convenient, mainly for reasons of general use.
[0246] However, it should be noted that all these and similar terms should be associated with appropriate physical quantities and are merely convenient labels applied to those quantities. This disclosure may refer to actions and processes of a computer system or similar electronic computing device that manipulate data represented as physical (electronic) quantities in the registers and memory of a computer system and convert them into other data similarly represented as physical quantities in the computer system memory or registers or other such information storage systems.
[0247] This disclosure also relates to an apparatus for performing the operations described herein. This apparatus may include a general-purpose computer that can be specifically constructed for an intended purpose or that can be selectively invoked or reconfigured by a computer program stored in the computer. For example, a computer system such as computing system 1000 or other data processing system can perform the techniques described above in response to its processor executing a computer program (e.g., a sequence of instructions) contained in memory or other non-temporary machine-readable storage medium. Such computer programs may be stored in computer-readable storage mediums, each coupled to a computer system bus, including, but not limited to, any type of disk including floppy disks, optical disks, CD-ROMs, and magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, or any type of medium suitable for storing electronic instructions.
[0248] The algorithms and representations presented herein are not inherently related to any particular computer or other device. Various general-purpose systems can be used with the programs in accordance with the teachings herein, or it may be convenient to construct more specialized devices for performing the methods. Various structures for these systems will appear as described below. Furthermore, this disclosure is not described with reference to any particular programming language. It will be understood that various programming languages can be used to implement the teachings of this disclosure as described herein.
[0249] This disclosure may be provided as a computer program product or software that includes a machine-readable medium storing instructions that can be used to program a computer system (or other electronic device) to perform the processes described herein. The machine-readable medium includes a mechanism for storing information in a machine-readable form. In some embodiments, the machine-readable (e.g., computer-readable) medium includes machine-readable (e.g., computer) storage media such as read-only memory ("ROM"), random-access memory ("RAM"), magnetic disk storage media, optical storage media, and flash memory components.
[0250] The example shown in Figure 11 and the accompanying explanation above are provided for illustrative purposes only. This disclosure is not limited to the example described. Additional or alternative details and implementations are described herein.
[0251] Exemplary embodiments of the technology disclosed herein are provided below. One embodiment of the technology may include any of the embodiments described herein, any combination of any of the embodiments described herein, or any combination of any parts of the embodiments described herein. One embodiment may include a system, method (e.g., a computer implementation method), apparatus, or non-temporary computer-readable medium configured according to one or more of the embodiments described herein.
[0252] In some embodiments, the technology described herein is a system for modeling causal relationships between pathogenicity evidence variables of gene variants and health status, comprising: at least one processor; at least one memory coupled to the at least one processor, wherein the at least one memory is a data preprocessing component that uses pathogenicity evidence data associated with gene variants and health status to create input data for a causal machine learning model, the pathogenicity evidence data being obtained through multiple evidence sources; and a data preprocessing component that applies the causal machine learning model to the input data to produce a trained causal model. The present invention relates to a system comprising a model training component, the graphical representation of the trained causal model comprising nodes connected via acyclic directed edges, where a first node represents a pathogenicity evidence variable among pathogenicity evidence variables related to a health state, at least one second node represents the cause of the pathogenicity evidence variable, at least one third node represents the effect of the pathogenicity evidence variable, and the acyclic directed edges represent the relationship between two of the nodes; and a prediction component for outputting predictive data sampled from the trained causal model in response to a request.
[0253] In some embodiments, the techniques described herein relate to a system in which a predictive component uses a trained causal model to generate and output predictions about whether a gene variant is benign or pathogenic in relation to health status.
[0254] In some embodiments, the technology described herein relates to a system in which predictive components provide clinicians with predictions about whether a gene variant is benign or pathogenic for use in formulating a patient's diagnosis.
[0255] In some embodiments, the technology described herein relates to a system in which a prediction component stores predictions for retrieval via at least one query.
[0256] In some embodiments, the techniques described herein relate to a system in which, in a graphical representation of a trained causal model, at least one second node is upstream of a first node, and at least one third node is downstream of the first node and at least one second node.
[0257] In some embodiments, the techniques described herein relate to a system in which, in a graphical representation of a trained causal model, at least one second node includes at least one variable representing data relating to at least one biological characteristic that causes pathogenicity of a gene variant with respect to a health state.
[0258] In some embodiments, the techniques described herein relate to a system in which at least one variable includes at least one of the following: a variable relating to the effect of a gene variant on protein function; a variable relating to the effect of a gene variant on protein stability; a variable relating to the effect of a gene variant on mRNA processing; a variable relating to the effect of a gene variant on nonsense mutation-dependent degradation; a variable relating to the effect of a gene variant on mRNA; a variable relating to the effect of a gene variant on protein expression; a variable relating to the biological importance of disrupted DNA; a variable relating to the biological importance of disrupted RNA nucleotides; or a variable relating to the biological importance of disrupted protein amino acid residues.
[0259] In some embodiments, the techniques described herein relate to a system in which, in a graphical representation of a trained causal model, at least one third node includes at least one variable representing data on at least one biological characteristic which is the effect of a gene variant on the pathogenicity of a health state.
[0260] In some embodiments, the techniques described herein relate to a system in which at least one variable is one of the following: a variable relating to the pathogenicity effect on individual patient phenotypes, a variable relating to the pathogenicity effect on population allele frequencies, a variable relating to the pathogenicity effect on selected cohort allele frequencies, or a variable relating to the pathogenicity effect on familial genotype-phenotypic cosegregation patterns.
[0261] In some embodiments, the technology described herein relates to a system in which at least one memory further includes a model evaluation component for applying at least one performance criterion to a trained causal model.
[0262] In some embodiments, the techniques described herein relate to a system in which, with respect to a category of pathogenicity evidence data, multiple evidence sources include alternative types of models for predicting pathogenicity given a category of pathogenicity evidence data, and a data preprocessing component creates input data using a model selected from among the alternative types of models, with respect to a category of pathogenicity evidence data, using at least one evaluation criterion.
[0263] In some embodiments, the techniques described herein relate to a system in which a graphical representation of a trained causal model includes at least one of a gene-level variable, a gene-level constraint, or a variant-level constraint.
[0264] In some embodiments, the techniques described herein relate to a system in which a predictive component samples data from any node of a trained causal model, and the nodes sampled by the predictive component correspond to unobserved variables relating to gene variants and health status.
[0265] In some embodiments, the techniques described herein are methods for creating training datasets for causal machine learning models, comprising: applying a natural language processing (NLP) model to clinical phenotypic data of gene variants; using the NLP model to identify features in the clinical phenotypic data that predict a molecular diagnosis of a health condition relating to the gene variants; for each subject in a target population, calculating a first score, which includes the probability that the subject has a health condition, using the identified features and demographic information; and including the first score in a training dataset for the target population.
[0266] In some embodiments, the techniques described herein further include the step of incorporating labeled pathogenicity data into a training dataset.
[0267] In some embodiments, the techniques described herein further include a step of using an NLP model to extract clinical phenotypic data from at least one field of a test application form, wherein the at least one field includes at least one of an indication field or a family history field.
[0268] In some embodiments, the techniques described herein further include the steps of: obtaining data on gene variants through multiple sources; filtering the data using at least one filtering criterion; converting the filtered data into a standardized format; and including the data in the standardized format into a training dataset.
[0269] In some embodiments, the techniques described herein relate to methods, where the data include at least one of the outputs of at least one machine learning model, empirical measurements, predicted variant effects, or public datasets.
[0270] In some embodiments, the techniques described herein relate to methods, and the standardized format includes a data structure that can be used as input to a causal machine learning model.
[0271] In some embodiments, the techniques described herein further include the step of splitting data into at least one training dataset for training a causal machine learning model and at least one holdout dataset for validating the trained causal machine learning model.
[0272] In some embodiments, the techniques described herein relate to a method for training a causal machine learning model for variant interpretation, the method comprising: creating model input data using pathogenicity evidence data of variants; applying a causal machine learning model to the model input data, wherein the graphical representation of the causal machine learning model includes nodes connected via acyclic directed edges, where a first node represents a pathogenicity variable, at least one second node represents a cause of the pathogenicity variable, at least one third node represents an effect of the pathogenicity variable, and at least one acyclic directed edge represents a relationship between two of the nodes; and iteratively evaluating data sampled from the causal machine learning model until at least one performance metric exceeds at least one threshold performance criterion.
[0273] In some embodiments, the technique described herein further includes the step of constructing a causal machine learning model by linking at least one second node to a first node via at least one upstream directed acyclic edge and linking at least one third node to the first node via at least one downstream directed acyclic edge.
[0274] In some embodiments, the techniques described herein further include the step of constructing a causal machine learning model using domain knowledge for at least one of the following: identifying at least one cause of a pathogenicity variable and assigning at least one cause of the pathogenicity variable to at least one second node; identifying at least one effect of the pathogenicity variable and assigning at least one effect of the pathogenicity variable to at least one third node; linking at least one second node to a first node via at least one upstream acyclic directed edge; or linking at least one third node to a first node via at least one downstream acyclic directed edge.
[0275] In some embodiments, the techniques described herein further include the steps of acquiring model input data through multiple sources, filtering the model input data using domain knowledge, wherein the domain knowledge includes types of relationships between variables, and the variables include probability distributions, and further includes the step of constructing a causal machine learning model by assigning weights to the probability distributions using the domain knowledge.
[0276] In some embodiments, the techniques described herein further include a step of constructing a causal machine learning model by fitting a hierarchical Bayesian inference model to model input data, wherein the model input data includes phenotypic data and variant pathogenicity data.
[0277] In some aspects, the technology described herein relates to an apparatus for variant interpretation, the apparatus comprising at least one processor and at least one memory, the at least one memory comprising a causal machine learning model, wherein a graphical representation of the causal machine learning model comprises nodes connected via acyclic directed edges, a first one of the nodes represents a pathogenicity variable, at least one second one of the nodes represents a cause of the pathogenicity variable, at least one third one of the nodes represents an effect of the pathogenicity variable, and at least one acyclic directed edge represents a relationship between two of the nodes.
[0278] In some aspects, the technology described herein relates to an apparatus, wherein the at least one memory further comprises submodels, and each submodel is capable of providing an output to a corresponding node in the graphical representation of the causal machine learning model.
[0279] In some aspects, the technology described herein relates to an apparatus, wherein the at least one memory further comprises a model abstraction component capable of determining output types of the submodels and specific interactions between the submodels.
[0280] In some aspects, the technology described herein relates to an apparatus, the apparatus further comprising at least one first device capable of receiving a request and providing an output generated by the causal machine learning model in response to the request to at least one of a data store, a network, a user interface, or at least one device.
[0281] In some embodiments, the techniques described herein are methods for interpreting gene variants, comprising the steps of: acquiring data by sampling nodes of a causal machine learning model, wherein the nodes are associated with gene variants, and the graphical representation of the causal machine learning model includes nodes connected via acyclic directed edges, where a first node represents a pathogenicity variable, at least one second node represents a cause of the pathogenicity variable, at least one third node represents an effect of the pathogenicity variable, and at least one acyclic directed edge represents a relationship between two of the nodes; and determining and outputting an interpretation of the gene variant using the sampled data.
[0282] In some embodiments, the techniques described herein relate to methods, and sampling a node further includes sampling a posterior predictive distribution associated with the node.
[0283] In some embodiments, the techniques described herein relate to methods, and the interpretation includes the probability that a gene variant is pathogenic to a health condition.
[0284] In some embodiments, the techniques described herein further include the steps of receiving a request via a device and providing an interpretation of a gene variant to at least one of a data store, a network, or a device.
[0285] In some embodiments, the technology described herein is a system for modeling causal relationships between pathogenicity evidence variables of gene variants and health status, comprising at least one processor and at least one memory coupled to the at least one processor, wherein the at least one memory is a data preprocessing component that uses pathogenicity evidence data associated with gene variants and health status to create input data for a causal machine learning model, the pathogenicity evidence data being obtained through multiple evidence sources, the multiple evidence sources being patient clinical data, population frequency data, published in vitro experimental study data, internally performed in vitro experimental study data, splice site loss prediction, splice site gain prediction, or curation based on the application of variant classification or variant interpretation frameworks The present invention relates to a system comprising: a data preprocessing component that includes at least one of the predictions of a nonsense mutation-dependent decomposition; a model training component that applies a causal machine learning model to input data to generate a trained causal model, wherein the graphical representation of the trained causal model includes nodes connected via acyclic directed edges, where a first node represents a pathogenicity evidence variable among pathogenicity evidence variables related to a health state, at least one second node represents the cause of the pathogenicity evidence variable, at least one third node represents the effect of the pathogenicity evidence variable, and the acyclic directed edges represent the relationship between two of the nodes; and a prediction component that, in response to a request, outputs predicted data sampled from the trained causal model.
[0286] In some embodiments, the techniques described herein relate to a system in which multiple evidence sources include patient clinical data evaluated through an Evidence Modeling Platform (EMP), population frequency data evaluated through an EMP, published in vitro experimental study data evaluated through an EMP, internally performed in vitro experimental study data evaluated through an EMP, splice site loss prediction, splice site gain prediction, and prediction of nonsense mutation-dependent decomposition curated based on the application of a variant classification or variant interpretation framework.
[0287] In some embodiments, the technique described herein is a method comprising the step of creating input data for a causal machine learning model using pathogenicity evidence data associated with gene variants and health status, wherein the pathogenicity evidence data is obtained through multiple evidence sources, and the multiple evidence sources include at least one of patient clinical data, population frequency data, published in vitro experimental study data, internally performed in vitro experimental study data, splice site loss prediction, splice site gain prediction, or prediction of nonsense mutation-dependent degradation curated based on the application of a variant classification or variant interpretation framework. The present invention relates to a method comprising the steps of: generating a causal machine learning model by applying a causal machine learning model to input data to generate a trained causal model, wherein the graphical representation of the trained causal model includes nodes connected via acyclic directed edges, where a first node represents a pathogenicity evidence variable associated with a health condition, at least one second node represents the cause of the pathogenicity evidence variable, at least one third node represents the effect of the pathogenicity evidence variable, and the acyclic directed edges represent the relationship between two of the nodes; and outputting predictive data sampled from the trained causal model in response to a request.
[0288] In some embodiments, the techniques described herein relate to methods, and the multiple evidence sources include patient clinical data evaluated through an Evidence Modeling Platform (EMP), population frequency data evaluated through an EMP, published in vitro experimental study data evaluated through an EMP, internally performed in vitro experimental study data evaluated through an EMP, splice site loss prediction, splice site gain prediction, and prediction of nonsense mutation-dependent degradation curated based on the application of a variant classification or variant interpretation framework.
[0289] In some embodiments, the technology described herein relates to an apparatus comprising at least one processor and at least one memory, wherein the at least one memory comprises a plurality of machine learning-based evidence models, including at least one of a first model of patient clinical data, a second model of population frequency data, a third model of published in vitro experimental study data, a fourth model of in vitro experimental study data, a fifth model of splice site loss prediction, a sixth model of splice site gain prediction, or a seventh model of nonsense mutation-dependent decomposition prediction; a model evaluation component for applying at least one evaluation criterion to the output of one or more of the plurality of machine learning-based evidence models; a model calibration component for determining weight values using the output of the model evaluation component and applying the weight values to the output of one or more of the plurality of machine learning-based evidence models; and a prediction component for outputting the output of the model calibration component to at least one device, system, process, model, component, or application.
[0290] In some embodiments, the technology described herein relates to an apparatus wherein at least one memory further comprises a model training component for training at least one of a plurality of machine learning-based evidence models using supervised machine learning.
[0291] In some embodiments, the techniques described herein relate to an apparatus in which a model training component further trains at least one of several machine learning-based evidence models using gene-specific labeled data.
[0292] In some embodiments, the techniques described herein relate to an apparatus, and the multiple machine learning-based evidence models include a first model for patient clinical data, a second model for population frequency data, a third model for published in vitro experimental study data, a fourth model for in vitro experimental study data, a fifth model for predicting splice site loss, a sixth model for predicting splice site gain, and a seventh model for predicting nonsense mutation-dependent degradation.
[0293] In some embodiments, the techniques described herein relate to an apparatus, wherein the model evaluation component is for performing at least one of the following: individually evaluating at least one first model from a plurality of machine learning-based evidence models, or evaluating at least two second models from a plurality of machine learning-based evidence models in combination.
[0294] In some embodiments, the techniques described herein are methods comprising the steps of: applying at least one evaluation criterion to the outputs of one or more of a plurality of machine learning-based evidence models, wherein the plurality of machine learning-based evidence models include at least one of a first model of patient clinical data, a second model of population frequency data, a third model of published in vitro experimental study data, a fourth model of in vitro experimental study data, a fifth model of splice site loss prediction, a sixth model of splice site gain prediction, or a seventh model of nonsense mutation-dependent decomposition prediction; determining one or more weight values for the outputs of the plurality of machine learning-based evidence models; applying one or more weight values to the outputs of the plurality of machine learning-based evidence models; and outputting predictive data based on the applied one or more weight values in response to a request.
[0295] In some embodiments, the techniques described herein further include a step of training at least one of several machine learning-based evidence models using supervised machine learning.
[0296] In some embodiments, the techniques described herein further include a step of training at least one model from a group of machine learning-based evidence models using gene-specific labeled data.
[0297] In some embodiments, the techniques described herein, with respect to methods, include multiple machine learning-based evidence models, including a first model for patient clinical data, a second model for population frequency data, a third model for published in vitro experimental study data, a fourth model for in vitro experimental study data, a fifth model for splice site loss prediction, a sixth model for splice site gain prediction, and a seventh model for nonsense mutation-dependent degradation prediction.
[0298] In some embodiments, the techniques described herein further include at least one of the steps of individually evaluating at least one first model from a plurality of machine learning-based evidence models, or evaluating a combination of at least two second models from a plurality of machine learning-based evidence models.
[0299] In some embodiments, the techniques described herein relate to any one or more embodiments, steps, components, elements, processes, or limitations that are described in the accompanying description and / or shown in the accompanying drawings.
[0300] Clause 1. A system for modeling a causal relationship between pathogenicity evidence variables of a genetic variant and a health condition, comprising: at least one processor; and at least one memory coupled to the at least one processor, wherein the at least one memory stores: a data preprocessing component that generates input data for a causal machine learning model using pathogenicity evidence data associated with the genetic variant and the health condition, wherein the pathogenicity evidence data is obtained via a plurality of evidence sources; a model training component that applies the causal machine learning model to the input data to generate a trained causal model, wherein a graphical representation of the trained causal model comprises nodes connected via acyclic directed edges, a first one of the nodes represents a pathogenicity evidence variable among the pathogenicity evidence variables associated with the health condition, at least one second node among the nodes represents a cause of the pathogenicity evidence variable, at least one third node among the nodes represents an effect of the pathogenicity evidence variable, and the acyclic directed edge represents a relationship between two of the nodes; and a prediction component for outputting prediction data sampled from the trained causal model in response to a request.
[0301] Clause 2. The system according to Clause 1, wherein the prediction component uses the trained causal model to generate and output a prediction regarding whether the genetic variant is benign or pathogenic with respect to the health condition.
[0302] Clause 3. The system according to Clause 2, wherein the prediction component provides a prediction regarding whether the genetic variant is benign or pathogenic to a clinician for use in formulating a patient diagnosis.
[0303] Clause 4. The system according to Clause 2 or Clause 3, wherein the prediction component stores the prediction for retrieval via at least one query.
[0304] Clause 5. The system described in any one of Clauses 1 to 4, wherein in the graphical representation of the trained causal model, at least one second node is upstream of a first node, and at least one third node is downstream of the first node and at least one second node.
[0305] Clause 6. A system according to any one of Clauses 1 to 5, wherein in a graphical representation of a trained causal model, at least one second node includes at least one variable representing data relating to at least one biological characteristic that causes pathogenicity of a gene variant with respect to health status.
[0306] Clause 7. The system described in Clause 6, wherein at least one variable includes at least one of the following: a variable relating to the effect of a gene variant on protein function; a variable relating to the effect of a gene variant on protein stability; a variable relating to the effect of a gene variant on mRNA processing; a variable relating to the effect of a gene variant on nonsense mutation-dependent degradation; a variable relating to the effect of a gene variant on mRNA; a variable relating to the effect of a gene variant on protein expression; a variable relating to the biological importance of disrupted DNA; a variable relating to the biological importance of disrupted RNA nucleotides; or a variable relating to the biological importance of disrupted protein amino acid residues.
[0307] Clause 8. A system according to any one of Clauses 1 to 7, wherein in a graphical representation of a trained causal model, at least one third node includes at least one variable representing data on at least one biological characteristic which is the effect of a gene variant on the pathogenicity of a health status.
[0308] Clause 9. The system described in Clause 8, wherein at least one variable is a variable relating to the pathogenicity impact on individual patient phenotypes, a variable relating to the pathogenicity impact on population allele frequencies, a variable relating to the pathogenicity impact on selected cohort allele frequencies, or a variable relating to the pathogenicity impact on familial genotype-phenotypic cosegregation patterns.
[0309] Clause 10. The system described in any one of Clauses 1 to 9, wherein at least one memory further comprises a model evaluation component for applying at least one performance criterion to a trained causal model.
[0310] Clause 11. With respect to categories of pathogenicity evidence data, the multiple evidence sources include alternative types of models for predicting pathogenicity given categories of pathogenicity evidence data, and the data preprocessing component creates input data using a model selected from among the alternative types of models, using at least one evaluation criterion for categories of pathogenicity evidence data, as described in any one of Clauses 1 to 10.
[0311] Clause 12. A system described in any one of Clauses 1 to 11, wherein the graphical representation of the trained causal model includes at least one of the following: a gene-level variable, a gene-level constraint, or a variant-level constraint.
[0312] Clause 13. The predictive component is a system described in any one of Clauses 1 to 12, which samples data from any node of the trained causal model, and the nodes sampled by the predictive component correspond to unobserved variables with respect to gene variants and health status.
[0313] Clause 14. A method for creating a training dataset for a causal machine learning model, comprising: applying a natural language processing (NLP) model to clinical phenotypic data of gene variants; using the NLP model to identify features in the clinical phenotypic data that predict a molecular diagnosis of a health condition relating to a gene variant; for each subject in a population of interest, calculating a first score, which includes the probability that the subject has a health condition, using the identified features and demographic information; and for the population of interest, including the first score in a training dataset.
[0314] Clause 15. The method according to Clause 14, further comprising the step of including labeled pathogenicity data in the training dataset.
[0315] The method described in Clause 16, further comprising the step of extracting clinical phenotypic data from at least one field of the test application form using an NLP model, wherein at least one field includes at least one of the indication field or the family history field.
[0316] Clause 17. The method described in any one of Clauses 14 to 16, further comprising the steps of: obtaining data on gene variants through multiple sources; filtering the data using at least one filtering criterion; converting the filtered data into a standardized format; and including the data in the standardized format into a training dataset.
[0317] Clause 18. The method described in Clause 17, wherein the data includes at least one of the outputs of at least one machine learning model, empirical measurements, predicted variant effects, or public datasets.
[0318] Clause 19. A standardized format is a data structure that can be used as input to a causal machine learning model, as described in Clause 17 or Clause 18.
[0319] Clause 20. The method according to Clauses 17-19, further comprising the step of splitting the data into at least one training dataset for training a causal machine learning model and at least one holdout dataset for validating the trained causal machine learning model.
[0320] Clause 21. A method for training a causal machine learning model for variant interpretation, comprising the steps of: creating model input data using pathogenicity evidence data of a variant; applying the causal machine learning model to the model input data, wherein the graphical representation of the causal machine learning model includes nodes connected via acyclic directed edges, where a first node represents a pathogenicity variable, at least one second node represents a cause of the pathogenicity variable, at least one third node represents an effect of the pathogenicity variable, and at least one acyclic directed edge represents a relationship between two of the nodes; and iteratively evaluating data sampled from the causal machine learning model until at least one performance metric exceeds at least one threshold performance criterion.
[0321] Clause 22. The method according to Clause 21, further comprising the step of building a causal machine learning model by linking at least one second node to the first node via at least one upstream directed acyclic edge and at least one third node to the first node via at least one downstream directed acyclic edge.
[0322] The method according to Clause 21 or Clause 22, further comprising the step of constructing a causal machine learning model using domain knowledge for at least one of the following: identifying at least one cause of a pathogenicity variable and assigning at least one cause of the pathogenicity variable to at least one second node; identifying at least one effect of the pathogenicity variable and assigning at least one effect of the pathogenicity variable to at least one third node; linking at least one second node to a first node via at least one upstream acyclic directed edge; or linking at least one third node to a first node via at least one downstream acyclic directed edge.
[0323] The method of the
[0324] Clause 25. The method described in any one of Clauses 21 to 24, further comprising the step of constructing a causal machine learning model by fitting a hierarchical Bayesian inference model to model input data, wherein the model input data includes phenotypic data and variant pathogenicity data.
[0325] Clause 26. Apparatus for variant interpretation, comprising at least one processor and at least one memory, wherein at least one memory includes a causal machine learning model, the graphical representation of the causal machine learning model includes nodes connected via acyclic directed edges, the first of the nodes representing a pathogenicity variable, at least one second of the nodes representing a cause of the pathogenicity variable, at least one third of the nodes representing an effect of the pathogenicity variable, and at least one acyclic directed edge representing a relationship between two of the nodes.
[0326] Clause 27. The apparatus described in Clause 26, wherein at least one memory further includes submodels, each submodel capable of providing output to corresponding nodes of a graphical representation of a causal machine learning model.
[0327] Clause 28. The apparatus according to Clause 27, further comprising at least one memory which can determine the output type of a submodel and specific interactions between submodels.
[0328] Clause 29. The apparatus described in any one of Clauses 26 to 28, further comprising at least one first device capable of receiving a request and providing, in response to the request, an output generated by a causal machine learning model to at least one of a data store, a network, a user interface, or at least one of the at least one device.
[0329] Clause 30. A method for interpreting a gene variant, comprising the steps of: sampling nodes of a causal machine learning model and obtaining sampled data, wherein the nodes are associated with a gene variant, and the graphical representation of the causal machine learning model includes nodes connected via acyclic directed edges, where a first node of the nodes represents a pathogenicity variable, at least one second node of the nodes represents the cause of the pathogenicity variable, at least one third node of the nodes represents the effect of the pathogenicity variable, and at least one acyclic directed edge represents the relationship between two of the nodes; and determining and outputting an interpretation of the gene variant using the sampled data.
[0330] Clause 31. The method according to Clause 30, wherein sampling a node further comprises sampling a posterior predictive distribution associated with the node.
[0331] Clause 32. Interpretations are as described in Clause 30 or Clause 31, including the probability that a gene variant is pathogenic to a health condition.
[0332] The method described in any one of the clauses 30 to 32, further comprising the steps of receiving a request via a device and providing an interpretation of a gene variant to at least one of a data store, network, or device.
[0333] Clause 34. A system for modeling causal relationships between pathogenicity evidence variables of gene variants and health status, comprising at least one processor and at least one memory coupled to the at least one processor, wherein the at least one memory is a data preprocessing component that uses pathogenicity evidence data associated with gene variants and health status to create input data for a causal machine learning model, wherein the pathogenicity evidence data is obtained through multiple evidence sources, the multiple evidence sources being patient clinical data, population frequency data, published in vitro experimental study data, internally performed in vitro experimental study data, splice site loss prediction, splice site gain prediction, or a nonsense based on the application of variant classification or variant interpretation frameworks. A system comprising: a data preprocessing component that includes at least one of the predictions of mutation-dependent decomposition; a model training component that applies a causal machine learning model to input data to generate a trained causal model, wherein the graphical representation of the trained causal model includes nodes connected via acyclic directed edges, where a first node of the nodes represents a pathogenicity evidence variable among pathogenicity evidence variables related to a health status, at least one second node of the nodes represents the cause of the pathogenicity evidence variable, at least one third node of the nodes represents the effect of the pathogenicity evidence variable, and the acyclic directed edges represent the relationship between two of the nodes; and a prediction component for outputting predicted data sampled from the trained causal model in response to a request.
[0334] Clause 35. Multiple evidence sources include patient clinical data evaluated through the Evidence Modeling Platform (EMP), population frequency data evaluated through the EMP, published in vitro experimental study data evaluated through the EMP, internally performed in vitro experimental study data evaluated through the EMP, splice site loss predictions, splice site gain predictions, and predictions of nonsense mutation-dependent degradation curated based on the application of a variant classification or variant interpretation framework, as described in Clause 34.
[0335] Clause 36. A method comprising the step of creating input data for a causal machine learning model using pathogenicity evidence data associated with gene variants and health status, wherein the pathogenicity evidence data is obtained through multiple evidence sources, and the multiple evidence sources include at least one of patient clinical data, population frequency data, published in vitro experimental study data, internally performed in vitro experimental study data, splice site loss prediction, splice site gain prediction, or prediction of nonsense mutation-dependent degradation curated based on the application of a variant classification or variant interpretation framework, and causal machine learning A method comprising: a generating step of applying a trained model to input data to generate a trained causal model, wherein the graphical representation of the trained causal model includes nodes connected via acyclic directed edges, where a first node represents a pathogenicity evidence variable related to a health status, at least one second node represents the cause of the pathogenicity evidence variable, at least one third node represents the effect of the pathogenicity evidence variable, and the acyclic directed edges represent the relationship between two of the nodes; and a generating step of outputting predictive data sampled from the trained causal model in response to a request.
[0336] Clause 37. The method described in Clause 36, including multiple evidence sources, such as patient clinical data evaluated through the Evidence Modeling Platform (EMP), population frequency data evaluated through the EMP, published in vitro experimental study data evaluated through the EMP, internally performed in vitro experimental study data evaluated through the EMP, splice site loss predictions, splice site gain predictions, and predictions of nonsense mutation-dependent decomposition curated based on the application of a variant classification or variant interpretation framework.
[0337] Clause 38. Apparatus comprising at least one processor and at least one memory, wherein at least one memory contains a plurality of machine learning-based evidence models, each including at least one of a first model of patient clinical data, a second model of population frequency data, a third model of published in vitro experimental study data, a fourth model of in vitro experimental study data, a fifth model of splice site loss prediction, a sixth model of splice site gain prediction, or a seventh model of nonsense mutation-dependent decomposition prediction; a model evaluation component for applying at least one evaluation criterion to the output of one or more of the plurality of machine learning-based evidence models; a model calibration component for determining weight values using the output of the model evaluation component and applying the weight values to the output of one or more of the plurality of machine learning-based evidence models; and a prediction component for outputting the output of the model calibration component to at least one device, system, process, model, component, or application.
[0338] Clause 39. The apparatus described in Clause 38, further comprising at least one memory for training at least one of a plurality of machine learning-based evidence models using supervised machine learning.
[0339] Clause 40. The model training component is the apparatus described in Clause 39, which further trains at least one of several machine learning-based evidence models using gene-specific labeled data.
[0340] Clause 41. Multiple machine learning-based evidence models include, but are not limited to, a first model for patient clinical data, a second model for population frequency data, a third model for published in vitro experimental study data, a fourth model for in vitro experimental study data, a fifth model for predicting splice site loss, a sixth model for predicting splice site gain, and a seventh model for predicting nonsense mutation-dependent degradation, as described in any one of Clauses 38-40.
[0341] Clause 42. The apparatus described in any one of Clauses 38 to 41, wherein the model evaluation component is for performing at least one of the following: evaluating at least one first model from a plurality of machine learning-based evidence models individually, or evaluating at least two second models from a plurality of machine learning-based evidence models in combination.
[0342] Clause 43. A method comprising the steps of: applying at least one evaluation criterion to the outputs of one or more of a plurality of machine learning-based evidence models, wherein the plurality of machine learning-based evidence models include at least one of a first model of patient clinical data, a second model of population frequency data, a third model of published in vitro experimental study data, a fourth model of in vitro experimental study data, a fifth model of splice site loss prediction, a sixth model of splice site gain prediction, or a seventh model of nonsense mutation-dependent decomposition prediction; determining one or more weight values for the outputs of the plurality of machine learning-based evidence models; applying one or more weight values to the outputs of the plurality of machine learning-based evidence models; and outputting predictive data based on the applied one or more weight values in response to a request.
[0343] Clause 44. The method of Clause 43, further comprising a step of training at least one of several machine learning-based evidence models using supervised machine learning.
[0344] Clause 45. The method according to Clause 44, further comprising a stage of training at least one of several machine learning-based evidence models using gene-specific labeled data.
[0345] Clause 46. Multiple machine learning-based evidence models, including a first model for patient clinical data, a second model for population frequency data, a third model for published in vitro experimental study data, a fourth model for in vitro experimental study data, a fifth model for predicting splice site loss, a sixth model for predicting splice site gain, and a seventh model for predicting nonsense mutation-dependent degradation, as described in any one of Clauses 43 to 45.
[0346] Clause 47. The method described in any one of Clauses 38 to 46, further comprising at least one of the steps of individually evaluating at least one first model from a plurality of machine learning-based evidence models, or evaluating a combination of at least two second models from a plurality of machine learning-based evidence models.
[0347] In the above-mentioned specification, embodiments of the disclosure are described with reference to specific exemplary embodiments. It will be apparent that various modifications can be made without departing from the broader spirit and scope of the embodiments of the disclosure described in the following claims. The specification and drawings are therefore considered illustrative, not restrictive.
Claims
1. A system for modeling the causal relationship between pathogenicity-evidence variables of gene variants and health status, At least one processor; The system comprises at least one memory coupled to the at least one processor, and the at least one memory is A data preprocessing component that creates input data for a causal machine learning model using the gene variant and pathogenicity evidence data associated with the health status, wherein the pathogenicity evidence data is obtained through multiple evidence sources; A model training component that applies the causal machine learning model to the input data to generate a trained causal model, wherein the graphical representation of the trained causal model includes nodes connected via acyclic directed edges, where a first node represents a pathogenicity evidence variable among the pathogenicity evidence variables related to the health state, at least one second node represents the cause of the pathogenicity evidence variable, at least one third node represents the effect of the pathogenicity evidence variable, and the acyclic directed edges represent the relationship between two of the nodes; A system including a predictive component for outputting predictive data sampled from the trained causal model in response to a request.
2. The system according to claim 1, wherein the predictive component uses the trained causal model to generate and output a prediction as to whether the gene variant is benign or pathogenic with respect to the health status.
3. The system according to claim 2, wherein the predictive component provides a clinician with the prediction regarding whether the gene variant is benign or pathogenic for use in formulating a patient's diagnosis.
4. The system according to claim 2, wherein the prediction component stores the prediction for retrieval via at least one query.
5. The system according to claim 1, wherein in the graphical representation of the trained causal model, the at least one second node is upstream of the first node, and the at least one third node is downstream of the first node and the at least one second node.
6. The system according to claim 1, wherein in the graphical representation of the trained causal model, the at least one second node includes at least one variable representing data relating to at least one biological characteristic that causes pathogenicity of the gene variant with respect to the health state.
7. The system according to claim 6, wherein the at least one variable includes at least one of the following: a variable relating to the effect of the gene variant on protein function; a variable relating to the effect of the gene variant on protein stability; a variable relating to the effect of the gene variant on mRNA processing; a variable relating to the effect of the gene variant on nonsense mutation-dependent degradation; a variable relating to the effect of the gene variant on mRNA; a variable relating to the effect of the gene variant on protein expression; a variable relating to the biological importance of disrupted DNA; a variable relating to the biological importance of disrupted RNA nucleotides; or a variable relating to the biological importance of disrupted protein amino acid residues.
8. The system according to claim 1, wherein in the graphical representation of the trained causal model, the at least one third node includes at least one variable representing data relating to at least one biological characteristic which is the effect of the pathogenicity of the gene variant with respect to the health state.
9. The system according to claim 8, wherein the at least one variable includes at least one of the following: a variable relating to the effect of pathogenicity on individual patient phenotypes, a variable relating to the effect of pathogenicity on population allele frequencies, a variable relating to the effect of pathogenicity on selected cohort allele frequencies, or a variable relating to the effect of pathogenicity on familial genotype-phenotypic cosegregation patterns.
10. The system according to claim 1, wherein the at least one memory further comprises a model evaluation component for applying at least one performance criterion to the trained causal model.
11. The system according to claim 1, wherein, with respect to the category of pathogenicity evidence data, the plurality of evidence sources include alternative type models for predicting pathogenicity given the category of pathogenicity evidence data, and the data preprocessing component creates the input data using a model selected from the alternative type models using at least one evaluation criterion for the category of pathogenicity evidence data.
12. The system according to claim 1, wherein the graphical representation of the trained causal model includes at least one of a gene-level variable, a gene-level constraint, or a variant-level constraint.
13. The system according to claim 1, wherein the predictive component samples data from any node of the trained causal model, and the nodes sampled by the predictive component correspond to the gene variant and the unobserved variables relating to the health status.
14. A method for creating a training dataset for a causal machine learning model, The stage of applying natural language processing (NLP) models to clinical phenotypic data of gene variants; The steps include: using the NLP model to identify features of the clinical phenotypic data that predict the molecular diagnosis of health status related to the gene variant; For each subject within the target population, a first score is calculated, which includes the probability that the subject suffers from the health condition, using the identified characteristics and demographic information; A method comprising the step of including the first score in the training dataset for the target population.
15. The method according to claim 14, further comprising the step of including labeled pathogenicity data in the training dataset.
16. The method according to claim 14, further comprising the step of using the NLP model to extract the clinical phenotypic data from at least one field of the test application form, wherein the at least one field includes at least one of an indication field or a family history field.
17. The steps include obtaining data related to gene variants through multiple sources; A step of filtering the data using at least one filtering criterion; The steps include: converting the filtered data into a standardized format; The method according to claim 14, further comprising the step of including the data in the standardized format in the training dataset.
18. The method according to claim 17, wherein the data includes at least one of the outputs of at least one machine learning model, empirical measurements, predicted variant effects, or a public dataset.
19. The method according to claim 17, wherein the standardized format includes a data structure that can be used as input to the causal machine learning model.
20. The method according to claim 17, further comprising the step of splitting the data into at least one training dataset for training the causal machine learning model and at least one holdout dataset for validating the trained causal machine learning model.
21. A method for training causal machine learning models, The steps include creating model input data using variant pathogenicity evidence data; The step of applying the causal machine learning model to the model input data, wherein the graphical representation of the causal machine learning model includes nodes connected via acyclic directed edges, where a first node represents a pathogenicity variable, at least one second node represents the cause of the pathogenicity variable, at least one third node represents the influence of the pathogenicity variable, and at least one acyclic directed edge represents the relationship between two of the nodes; A method comprising the step of iteratively evaluating data sampled from the causal machine learning model until at least one performance metric exceeds at least one threshold performance criterion.
22. Linking the at least one second node to the first node via at least one upstream directed acyclic edge; By linking the at least one third node to the first node via at least one downstream directed acyclic edge, The method according to claim 21, further comprising the step of constructing the causal machine learning model.
23. Identifying at least one cause of the pathogenicity variable and assigning the at least one cause of the pathogenicity variable to the at least one second node; Identifying at least one effect of the pathogenicity variable and assigning the at least one effect of the pathogenicity variable to the at least one third node; Linking the at least one second node to the first node via at least one upstream acyclic directed edge; or Linking the at least one third node to the first node via at least one downstream acyclic directed edge, The method according to claim 21, further comprising the step of constructing the causal machine learning model using domain knowledge of at least one of the following.
24. The steps include: acquiring the model input data through multiple sources; The process further includes the step of filtering the model input data using domain knowledge, The method according to claim 23, wherein the domain knowledge includes types of relationships between variables, the variables include probability distributions, and the method further includes the step of constructing the causal machine learning model by assigning weights to the probability distributions using the domain knowledge.
25. The method according to claim 21, further comprising the step of constructing the causal machine learning model by fitting a hierarchical Bayesian inference model to the model input data, wherein the model input data includes phenotypic data and variant pathogenicity data.
26. A device for variant interpretation or variant classification, At least one processor; It has at least one memory, The device wherein the at least one memory includes a causal machine learning model, the graphical representation of the causal machine learning model includes nodes connected via acyclic directed edges, the first of the nodes representing a pathogenicity variable, at least one second of the nodes representing the cause of the pathogenicity variable, at least one third of the nodes representing the influence of the pathogenicity variable, and at least one acyclic directed edge representing the relationship between two of the nodes.
27. The apparatus according to claim 26, wherein the at least one memory further includes a submodel, each submodel capable of providing an output to a corresponding node of the graphical representation of the causal machine learning model.
28. The apparatus according to claim 27, wherein the at least one memory further comprises a model abstraction component capable of determining the output type of the submodels and specific interactions between the submodels.
29. The apparatus according to claim 26, further comprising at least one first device capable of receiving a request and providing the output generated by the causal machine learning model in response to the request to at least one of a data store, a network, a user interface, or at least one of the at least one device.
30. A method for interpreting or classifying gene variants, A step of sampling nodes of a causal machine learning model and obtaining sampled data, wherein the nodes are associated with the gene variants, and the graphical representation of the causal machine learning model includes nodes connected via acyclic directed edges, where a first node represents a pathogenicity variable, at least one second node represents the cause of the pathogenicity variable, at least one third node represents the influence of the pathogenicity variable, and at least one acyclic directed edge represents the relationship between two of the nodes; A method comprising the steps of determining and outputting an interpretation of the gene variant using the sampled data.
31. The method according to claim 30, wherein sampling the node further comprises sampling the posterior predictive distribution associated with the node.
32. The method according to claim 30, wherein the interpretation includes the probability that the gene variant is pathogenic to a health condition.
33. The stage of receiving a request via a device; The method according to claim 30, further comprising the step of providing the interpretation of the gene variant to at least one of a data store, a network, or the device.
34. A system for modeling the causal relationship between pathogenicity-evidence variables of gene variants and health status, At least one processor; The system comprises at least one memory coupled to the at least one processor, and the at least one memory is A data preprocessing component that creates input data for a causal machine learning model using pathogenicity evidence data associated with the gene variant and the health status, wherein the pathogenicity evidence data is obtained through a plurality of evidence sources, the plurality of evidence sources including at least one of patient clinical data, population frequency data, published in vitro experimental study data, internally performed in vitro experimental study data, splice site loss prediction, splice site gain prediction, or prediction of nonsense mutation-dependent decomposition curated based on the application of a variant classification or variant interpretation framework; A model training component that applies the causal machine learning model to the input data to generate a trained causal model, wherein the graphical representation of the trained causal model includes nodes connected via acyclic directed edges, where a first node represents a pathogenicity evidence variable among the pathogenicity evidence variables related to the health state, at least one second node represents the cause of the pathogenicity evidence variable, at least one third node represents the effect of the pathogenicity evidence variable, and the acyclic directed edges represent the relationship between two of the nodes; A system including a predictive component for outputting predictive data sampled from the trained causal model in response to a request.
35. The system according to claim 34, wherein the plurality of evidence sources include patient clinical data evaluated through an evidence modeling platform (EMP), population frequency data evaluated through the EMP, published in vitro experimental study data evaluated through the EMP, internally performed in vitro experimental study data evaluated through the EMP, splice site loss predictions, splice site gain predictions, and predictions of nonsense mutation-dependent decomposition curated based on the application of the variant classification or variant interpretation framework.
36. It is a method, A step of creating input data for a causal machine learning model using pathogenicity evidence data associated with gene variants and health status, wherein the pathogenicity evidence data is obtained through multiple evidence sources, the multiple evidence sources including at least one of patient clinical data, population frequency data, published in vitro experimental study data, internally performed in vitro experimental study data, splice site loss prediction, splice site gain prediction, or prediction of nonsense mutation-dependent degradation curated based on the application of a variant classification or variant interpretation framework; A step of generating a trained causal model by applying the causal machine learning model to the input data, wherein the graphical representation of the trained causal model includes nodes connected via acyclic directed edges, where a first node represents a pathogenicity evidence variable related to the health status, at least one second node represents the cause of the pathogenicity evidence variable, at least one third node represents the effect of the pathogenicity evidence variable, and the acyclic directed edges represent the relationship between two of the nodes; A method comprising the step of outputting predictive data sampled from the trained causal model in response to a request.
37. The method according to claim 36, wherein the plurality of evidence sources include patient clinical data evaluated through an evidence modeling platform (EMP), population frequency data evaluated through the EMP, published in vitro experimental study data evaluated through the EMP, internally performed in vitro experimental study data evaluated through the EMP, splice site loss predictions, splice site gain predictions, and predictions of nonsense mutation-dependent decomposition curated based on the application of the variant classification or variant interpretation framework.
38. It is a device, At least one processor; It comprises at least one memory, and the at least one memory is Multiple machine learning-based evidence models, including at least one of the following: a first model of patient clinical data, a second model of population frequency data, a third model of published in vitro experimental study data, a fourth model of in vitro experimental study data, a fifth model of splice site loss prediction, a sixth model of splice site acquisition prediction, or a seventh model of nonsense mutation-dependent degradation prediction; A model evaluation component for applying at least one evaluation criterion to one or more outputs of the aforementioned plurality of machine learning-based evidence models; A model calibration component for determining weight values using the output of the model evaluation component and applying the weight values to one or more outputs of the plurality of machine learning-based evidence models; An apparatus comprising a prediction component for outputting the output of the model calibration component to at least one device, system, process, model, component, or application.
39. The apparatus according to claim 38, wherein the at least one memory further comprises a model training component for training at least one of the plurality of machine learning-based evidence models using supervised machine learning.
40. The apparatus according to claim 39, wherein the model training component further trains at least one of the plurality of machine learning-based evidence models using gene-specific labeled data.
41. The apparatus according to claim 38, wherein the plurality of machine learning-based evidence models include a first model of patient clinical data, a second model of population frequency data, a third model of published in vitro experimental study data, a fourth model of in vitro experimental study data, a fifth model for predicting splice site loss, a sixth model for predicting splice site acquisition, and a seventh model for predicting nonsense mutation-dependent degradation.
42. The apparatus according to claim 38, wherein the model evaluation component is for performing at least one of the following: individually evaluating at least one first model from the plurality of machine learning-based evidence models, or evaluating at least two second models from the plurality of machine learning-based evidence models in combination.
43. It is a method, The step of applying at least one evaluation criterion to one or more outputs of a plurality of machine learning-based evidence models, wherein the plurality of machine learning-based evidence models include at least one of a first model of patient clinical data, a second model of population frequency data, a third model of published in vitro experimental study data, a fourth model of in vitro experimental study data, a fifth model of splice site loss prediction, a sixth model of splice site gain prediction, or a seventh model of nonsense mutation-dependent decomposition prediction; The steps include determining one or more weight values for the outputs of the plurality of machine learning-based evidence models; The step of applying the one or more weight values to the output of the plurality of machine learning-based evidence models; A method comprising the step of outputting predictive data based on one or more applied weight values in response to a request.
44. The method according to claim 43, further comprising the step of training at least one of the plurality of machine learning-based evidence models using supervised machine learning.
45. The method according to claim 44, wherein the training step further comprises training at least one of the plurality of machine learning-based evidence models using gene-specific labeled data.
46. The method according to claim 43, wherein the plurality of machine learning-based evidence models include a first model for patient clinical data, a second model for population frequency data, a third model for published in vitro experimental study data, a fourth model for in vitro experimental study data, a fifth model for predicting splice site loss, a sixth model for predicting splice site acquisition, and a seventh model for predicting nonsense mutation-dependent degradation.
47. The method according to claim 43, further comprising at least one of the steps of individually evaluating at least one first model from the plurality of machine learning-based evidence models, or evaluating a combination of at least two second models from the plurality of machine learning-based evidence models.