Interpretation of genetic and genomic variants via integrated computational and experimental deep mutational learning framework

An integrated computational and experimental deep mutation learning framework addresses the challenge of interpreting genotypic variants by calculating molecular and phenotypic scores and applying statistical learning, enhancing the accuracy and scalability of variant impact assessment.

JP2025114687APending Publication Date: 2025-08-05INVITAE CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025076385
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2018-03-08
Filing Date
2025-05-01
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Current methods struggle to accurately and scalably interpret the phenotypic impact of genotypic variants, particularly due to the high number of uncharacterized variants in disease-associated genes, leading to challenges in diagnostic and screening tests, and existing computational predictors exhibit low accuracy.

Method used

A computer-implemented method using an integrated computational and experimental deep mutation learning framework that determines phenotypic impact by receiving molecular variants, calculating molecular and phenotypic scores, and applying statistical learning to derive functional and evidence scores, incorporating machine learning techniques and high-throughput DNA sequencing platforms.

Benefits of technology

Enables robust and scalable interpretation of molecular variant impacts across diverse types of variants, biophysical processes, and functional elements, improving accuracy and applicability to less well-understood genes and pathways.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025114687000001_ABST
    Figure 2025114687000001_ABST
Patent Text Reader

Abstract

To provide a system, method and computer program product for determining phenotypic impacts of molecular variants identified within a biological sample.SOLUTION: The method comprises: receiving molecular variants associated with functional elements within a model system; determining molecular scores associated with the model system; determining molecular signals and population signals associated with the molecular variants based on the molecular scores; determining functional scores for the molecular variants based on statistical learning; deriving evidence scores of the molecular variants based on the functional scores; and determining phenotypic impacts of the molecular variants based on the functional scores or evidence scores.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Understanding the impact of genotypic (e.g., sequence) variants within functional genomic elements, such as protein-coding genes, non-coding genes, and regulatory elements, is critical for a wide variety of life science applications. Today, nearly half of disease-associated genes contain a greater number of uncharacterized variants in the population than variants with known clinical significance. This poses a significant challenge for both diagnostic and screening tests that evaluate genes and genomic sequences (Landrum et al. 2015; Lek et al. 2016). Large numbers of novel variants with unknown clinical significance are characteristic of nearly every gene (e.g., germline and somatic variants in the population), affecting even the most frequently tested genes. For example, tests evaluating gene panels for cancer predisposition mutations report the discovery of as many as 95 uncharacterized variants for every known disease-causing variant (Maxwell et al. 2016). Thus, predicting the phenotypic (e.g., cellular, biological, clinical, or other) consequences of genotypic variants poses a challenge to leveraging genetic and genomic information in a wide variety of clinical settings.

[0002] Genotypic (e.g., sequence) variants within gene-encoded functional elements can affect diverse biophysical processes and alter distinct molecular functions within each element, resulting in altered clinical and non-clinical phenotypes. For example, in the established tumor suppressor protein-encoding gene, phosphatase and tensin homolog (PTEN), genotypic variants affecting transcription (fg-903G>A, -975G>C, and -1026C>A), protein stability (fgC136R), phosphatase catalytic activity (fgC124S, H93R), and substrate recognition (fgG129E) all confer risk for breast, thyroid, endometrial, kidney, colorectal, and melanoma cancers and are associated with Cowden syndrome (CS) (Heikkinen et al. 2011; He et al. 2013; Myers et al. 1997; Myers et al. 1998). Variants affecting the same biophysical processes and molecular functions can lead to comorbidity between different disorders, as exemplified by PTEN variants affecting phosphatase activity (e.g., H93R), which are also found in autism spectrum disorders (ASD) (Johnston and Raines 2015), and the frequent comorbidity between ASD and cancer (Markkanen et al. 2016). Furthermore, variants affecting different biophysical processes and molecular mechanisms within functional elements can exhibit stereotypical and differentiated clinical and non-clinical phenotypes. Mutations in the lamina A / C gene (LMNA) cause a collection of over 15 disorders collectively referred to as "laminopathies," including A-EDMD (autosomal Emery-Dreifuss muscular dystrophy), DCM (dilated cardiomyopathy), LGMD1B (limb-girdle muscular dystrophy 1B), L-CMD (LMNA-related congenital muscular dystrophy), FPLD2 (familial partial lipodystrophy 2), HGPS (Hutchinson-Gilford progeria syndrome), atypical WRN (Werner syndrome), MAD (mandibular dysplasia), and CMT2B (Charcot-Marie-Tooth disorder type 2B) (Scharner et al. 2010).In LMNA, genotypic (e.g., sequence) variants leading to HGPS generate a cryptic splice site donor in lamin A-specific exon 11, resulting in a truncated form of lamin A, whereas variants leading to FPLD2 alter the surface charge of the Ig-like region and do not alter the crystal structure of the mutant protein (Scharner et al. 2010). Thus, reducing the complexity of genotype-phenotype relationships and cellular effects across a wide variety of variant types, functional elements, and molecular systems remains a challenge for robust and scalable interpretation of the phenotypic consequences of variants discovered in clinical and nonclinical genetic and genomic testing.

[0003] Indeed, assessing the significance of genotypic (e.g., sequence) variants can be a complex and challenging task. As recently as 2015, studies of variant classifications showed that as many as 17% (e.g., 2,229 / 12,895) of variant classifications were discordant between classification submitters (Rehm et al. 2015). Interpretation agreement as low as 34% has been measured between clinical testing laboratories, but with specific suggestions, interlaboratory agreement can increase to 71% (Amendola et al. 2016).

[0004] With over 5,300 genes evaluated by genetic tests in the market (e.g., by the NCBI Genetic Test Registry), scalable solutions for the interpretation (e.g., classification) of genotype (e.g., sequence) variants across a wide range of genes, diseases, and contexts (e.g., clinical and non-clinical) are important to the precision medicine and life sciences industries. With over 14,000,000 potential (e.g., unique) molecular variants in the clinical testing market, within the subset of molecular variants corresponding to single nucleotide variants (SNVs), within the subset of coding sequences, and within the subset of protein-coding genes, an effective solution for molecular variant classification needs to be robust and scalable.

[0005] While multiple strategies exist for identifying the phenotypic impact of molecular variants, including but not limited to family classification, functional measurements, and case-control studies, currently only computational variant effect predictors are capable of providing supporting evidence at the necessary scale. Indeed, analysis of clinical variant classifications from practitioners following the Joint Guidelines for Clinical Variant Interpretation from the American College of Clinical Genetics and Genomics (ACMG) and the Association for Molecular Pathology (AMP) indicates that up to 50% of clinical variant classifications rely on the use of computational variant effect predictors. However, despite their widespread use, benchmark studies show that the performance of computational variant effect prediction algorithms, such as SIFT, PolyPhen (v2), GERP++, Condel, CADD, REVEL, and others, is significantly lower, with accuracy rates (AUCs) ranging from 0.52 to 0.75 (Mahmood et al. 2017).

[0006] Direct measurement of molecular function may provide a basis for accurate interpretation of the clinical and nonclinical impact of genotypic (e.g., sequence) variants (Shendure and Fields 2016; Araya and Fowler 2011). To date, a diverse spectrum of measurements has been devised to directly assess the impact of variants on a wide variety of molecular functions. However, existing methods require a priori knowledge or assumptions about the mechanism of action of the variant associated with the clinical (and nonclinical) phenotype being investigated in order to define and measure molecular function (Shendure and Fields 2016). These methods are often limited to capturing and reporting only the effects of variants affecting the specific molecular function being measured, imposing limitations on the types of variants, molecular functions, and functional elements that can be measured on a large scale, as well as the genes. Thus, for example, phosphatase measurements may indicate (e.g., include) potential disease relevance for variants affecting the catalytic activity of the PTEN tumor suppressor, but such measurements may be unable to rule out (e.g., exclude) potential disease relevance for variants affecting protein stability because variants affecting protein stability may increase disease risk without an observable defect in catalytic activity. Conversely, for example, protein stability measurements may indicate (e.g., include) potential disease relevance for variants leading to stability defects in the PTEN tumor suppressor, but such measurements may be unable to rule out (e.g., exclude) potential disease relevance for variants affecting catalytic activity. The potential need for a priori knowledge or assumptions about the mechanism of action (and thus the relevant molecular function to measure) may limit the application of these methods to well-characterized functional elements (e.g., genes) and phenotypes, thereby preventing their application to less well-understood disease-associated genes.

[0007] Building on the technological foundation of high-throughput DNA sequencing platforms, recently developed large-scale functional assays such as deep mutation scanning (DMS), HITS-KIN, RNA map, and others, enable comprehensive or near-comprehensive coverage of potential sequence variants of different sequence classes, including single nucleotide variants (SNVs) and nonsynonymous variants (NSVs, missense variants) in coding, noncoding, and regulatory elements (Fowler et al. 2010; Araya et al. 2012; Guenther et al. 2013; Buenrostro et al. 2014; Kelsic et al. 2016; Patwardhan et al. 2009). Such methods may provide a basis for robust, statistically validated interpretation of the impact of molecular variants, such as genotypic (e.g., sequence) variants, on patient phenotypes, including clinical phenotypes such as lipodystrophy and increased risk of type 2 diabetes (T2D) in patients with mutations in PPARG or increased risk of breast and ovarian cancer in patients with mutations in BRCA1 (Starita et al. 2015; Majithia et al. 2016). While such methods may provide robust variant interpretation in clinical and nonclinical testing settings, these methods may require significant development and customization to measure each molecular function and each functional element. This may limit their utility as a generalizable, scalable solution for systematically assessing the clinical and nonclinical consequences of molecular variants, such as genotypic (e.g., sequence) variants, across diverse types of variants, biophysical processes, molecular functions, functional elements, genes, and ultimately, pathways. Thus, multifunctional platforms and methods for variant impact assessment are needed.

[0008] The accompanying drawings are incorporated into and form a part of this specification. The present invention provides, for example, the following items. (Item 1) 1. A computer-implemented method for determining the phenotypic impact of molecular variants identified in a biological sample, comprising: receiving molecular variants associated with one or more functional elements in a model system, wherein the model system comprises a single cell, a cellular compartment, a subcellular compartment, or a synthetic compartment; determining a molecular or phenotypic score for the single cell, the cellular compartment, the subcellular compartment, or the composite compartment; determining a molecular or phenotypic signal associated with a particular molecular variant based on the molecular or phenotypic score of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment containing the particular molecular variant; determining a population signal associated with a particular molecular variant based on the molecular or phenotypic score of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment containing the molecular variant; determining a functional score or functional classification for the molecular variant based on statistical learning, wherein the statistical learning relates the molecular signal, the phenotypic signal, or the population signal of the molecular variant to a phenotypic effect of the molecular variant; deriving an evidence score or classification for the molecular variant based on the functional score or classification, modeling of the functional score or classification, predictor score or classification, or hotspot score or classification; determining the phenotypic impact of the molecular variant based on the functional score, the functional classification, the evidence score, or the evidence classification. (Item 2) Item 10. The method of item 1, wherein the evidence score or the evidence classification is determined based on the molecular signal, the phenotypic signal, or the population signal from the molecular variants in one or more functional elements. (Item 3) Item 10. The method of item 1, wherein the evidence score or evidence classification is derived from the function score or functional classification, the predictor score or predictor classification, or the hotspot score or hotspot classification. (Item 4) 2. The method of claim 1, wherein the evidence score or evidence classification is derived by applying statistical learning that utilizes regression or classification to relate the evidence score and evidence classification to the phenotypic effect of the molecular variant. (Item 5) 2. The method of claim 1, wherein the functional score or functional classification of the molecular variant is derived by applying statistical learning that utilizes regression or classification to relate molecular signals to the phenotypic impact of the molecular variant. (Item 6) 5. The method of claim 4, wherein the phenotypic effect of the molecular variant is derived based on a clinical database, a phenotypic database, a population database, a molecular annotation database, or a functional database of variants, subjects, or populations. (Item 7) 5. The method of claim 4, wherein the phenotypic impact of the molecular variant is derived based on molecular signals such as mutational dose, mutation rate, and mutational signature. (Item 8) 2. The method of claim 1, wherein the functional score or functional classification of the molecular variant is derived from multiple statistical models generated using independent or disjoint estimates of the molecular, phenotypic, or population signals. (Item 9) 2. The method of claim 1, wherein the functional score or functional classification of the molecular variant is derived from a functional modeling engine (FME), and the FME is generated by applying machine learning techniques to associate non-measured features of the molecular variant with the functional score or functional classification, and the non-measured features include evolutionary, population, functional, structural, dynamic, and physicochemical features. (Item 10) 2. The method of claim 1, wherein the predictor scores or predictor classifications of the molecular variants are derived from a variant interpretation engine (VIE), and the VIE is generated by applying machine learning techniques to relate the functional scores or functional classifications and non-measured features to the phenotypic impact of the molecular variants. (Item 11) 2. The method of claim 1, wherein the predictor scores or predictor classifications are derived from lower-level variant interpretation engines (VIEs), and the lower-level VIEs are functional element, functional type, or condition-specific. (Item 12) 2. The method of claim 1, wherein the predictor score or predictor classification is derived from a higher-level variant interpretation engine (VIE), and the higher-level VIE is pathway, homolog family, enzyme family, or condition specific. (Item 13) 2. The method of claim 1, wherein the predictor scores or predictor classifications are derived from a higher-level variant interpretation engine (VIE), the VIE representing multiple pathways, homolog families, enzyme families, or conditions. (Item 14) 2. The method according to item 1, wherein the hotspot scores or hotspot classifications of the molecular variants are derived from a spatial clustering technique that applies significantly mutated regions and networks (SMR / SMN) calculation, and residual regions and networks with high density of molecular variants with high or low functional scores or specific functional classifications are detected. (Item 15) 16. The method of claim 1, wherein the molecular signal comprises a molecular signal below the molecular variant derived as a summary statistic, summary statistic, descriptive statistic, inferential statistic, or Bayesian inference model of the molecular score measured in the single cell, the cellular compartment, the subcellular compartment, or the composite compartment containing the molecular variant. 2. The method of claim 1, wherein the molecular signal comprises a higher-level molecular signal of the molecular variant derived by applying an existing model relating lower-level molecular signals to regulation, signal transduction, pathway, processing, cell cycle activity, alteration, defect, or state. (Item 17) 2. The method of claim 1, wherein the molecular signal comprises a higher-level molecular signal of the molecular variant that is derived from a lower-level molecular signal via unsupervised learning, feature representation learning, or dimensionality reduction techniques. (Item 18) 2. The method of claim 1, wherein the molecular signal comprises a lower molecular score corresponding to a molecular measurement, process, or feature from the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment. (Item 19) 2. The method of claim 1, wherein the molecular signal comprises a higher molecular score for the single cell, the cellular compartment, the subcellular compartment, or the composite compartment that is derived by applying an existing model that relates lower molecular scores to regulation, signaling, pathways, processes, cell cycle activity, alterations, defects, or states. (Item 20) 2. The method of claim 1, wherein the molecular signal comprises a higher-level molecular score for the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment derived through a lower-level molecular score from unsupervised learning, feature representation learning, or dimensionality reduction techniques. (Item 21) 21. The method of claim 20, wherein an autoencoder neural network is trained to learn condensed representations of lower-level molecular scores, and the autoencoder is utilized to encode lower-level molecular signals into condensed representations of higher-level molecular scores. (Item 22) 22. The method of claim 21, wherein the autoencoder is trained as a denoising autoencoder (DAE), or the autoencoder is constructed as a neural network with fully connected layers, or the autoencoder is constructed as a neural network with a symmetric number of neurons, or the autoencoder is constructed with rectified linear units (ReLu) for activation, or the autoencoder is trained using an Adam optimizer, or the autoencoder is cell type, gene, pathway, or disorder specific. (Item 23) 20. The method of claim 18, wherein the molecular measurement corresponds to a locus-specific measurement of gene expression, protein expression, chromatin accessibility, epigenetic modification, regulatory activity, post-transcriptional processing, post-translational modification, mutational status, mutational burden, or mutation rate of a molecule within the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment. (Item 24) 20. The method of claim 18, wherein the molecular treatment corresponds to a multi-locus measurement of gene expression, protein expression, chromatin accessibility, epigenetic modification, regulatory activity, transcriptional activity, translational activity, signaling activity, pathway activity, mutational status, mutational burden, or mutation rate derived from molecular measurements within the single cell, the cellular compartment, the subcellular compartment, or a synthetic compartment. (Item 25) 19. The method of claim 18, wherein the molecular signature corresponds to a global measure of gene expression, protein expression, chromatin accessibility, epigenetic modification, regulatory activity, transcriptional activity, translational activity, signaling activity, pathway activity, mutational status, mutational burden, or mutation rate derived from molecular measurements or processes within the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment. (Item 26) 20. The method of claim 18, wherein the molecular measurements are derived by applying single cell barcoding and nucleic acid sequencing techniques to the single cell, the cellular compartment, the subcellular compartment, or the population of synthetic compartments. (Item 27) 19. The method of claim 18, wherein the molecular measurements can include sequence read quality control, cell barcode identification or quality control, molecular barcode identification or quality control, alignment of sequence reads to a reference genome, sequence read alignment filtering or quality control, mapping of filtered and quality controlled sequence reads to functional elements, mapping of filtered and quality controlled molecular barcodes to functional elements, and mapping of filtered and quality controlled sequence reads or molecular barcodes to functional elements for specific cell barcodes. (Item 28) 2. The method of claim 1, wherein the molecular signal, the phenotypic signal, or the population signal is molecular state specific and is derived from a population of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment from a specific molecular state, enabling learning in a state-specific learning layer. (Item 29) 2. The method of claim 1, wherein the molecular signal, the phenotypic signal, or the population signal is molecular state agnostic and is derived from the single cell, the population of the cellular compartment, the subcellular compartment, or the synthetic compartment from multiple molecular states, enabling learning in a state-agnostic learning layer. (Item 30) 2. The method of claim 1, wherein the molecular signals, the phenotypic signals, or the population signals are ordered by molecular state and are derived from the single cell, the population of the cellular compartment, the subcellular compartment, or the synthetic compartment from multiple molecular states to enable learning in a multi-state learning layer. (Item 31) 2. The method of claim 1, wherein the molecular state of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment is derived by applying an existing model relating molecular or phenotypic scores to the molecular state, and the model assigns the single cell to a cell cycle phase based on a pre-characterized gene expression signature. (Item 32) 2. The method of claim 1, wherein the molecular state of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment is derived via unsupervised learning, feature representation learning, or dimensionality reduction techniques of molecular or phenotypic scores across the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment. (Item 33) 2. The method of claim 1, wherein the molecular signal, the phenotypic signal, or the population signal is calculated from independent or disjoint populations of single cells, cellular compartments, subcellular compartments, or synthetic compartments selected from the single cells, cellular compartments, subcellular compartments, or synthetic compartments containing the same molecular variant via random sampling. (Item 34) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within functional elements, genes and pathways associated with Mendelian diseases. (Item 35) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within functional elements, genes, and pathways associated with known cancer drivers. (Item 36) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within functional elements, genes and pathways associated with altered drug response. (Item 37) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified mutational hotspots of functional elements, genes and pathways associated with other clinically valuable genes. (Item 38) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified mutational hotspots of functional elements, genes and pathways associated with Mendelian diseases. (Item 39) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified mutational hotspots of functional elements, genes and pathways associated with known cancer drivers. (Item 40) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified mutational hotspots of functional elements, genes and pathways associated with altered drug response. (Item 41) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified mutational hotspots of functional elements, genes and pathways associated with other clinically valuable genes. (Item 42) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 10 bp pre-identified mutational hotspots of functional elements, genes and pathways associated with Mendelian diseases. (Item 43) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 10 bp pre-identified mutational hotspots of functional elements, genes and pathways associated with known cancer drivers. (Item 44) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 10 bp pre-identified mutational hotspots of functional elements, genes and pathways associated with altered drug response. (Item 45) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 10 bp pre-identified mutational hotspots of functional elements, genes and pathways associated with other clinically valuable genes. (Item 46) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 50 bp of pre-identified mutational hotspots of functional elements, genes and pathways associated with Mendelian diseases. (Item 47) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 50 bp of pre-identified mutational hotspots of functional elements, genes, and pathways associated with known cancer drivers. (Item 48) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 50 bp of pre-identified mutational hotspots of functional elements, genes and pathways associated with altered drug response. (Item 49) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 50 bp of pre-identified mutational hotspots of functional elements, genes and pathways associated with other clinically valuable genes. (Item 50) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 100 bp of pre-identified mutational hotspots of functional elements, genes and pathways associated with Mendelian diseases. (Item 51) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 100 bp of pre-identified mutational hotspots of functional elements, genes, and pathways associated with known cancer drivers. (Item 52) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 100 bp of pre-identified mutational hotspots of functional elements, genes and pathways associated with altered drug response. (Item 53) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 100 bp of pre-identified mutational hotspots of functional elements, genes and pathways associated with other clinically valuable genes. (Item 54) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 500 bp of pre-identified mutational hotspots of functional elements, genes and pathways associated with Mendelian diseases. (Item 55) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 500 bp of pre-identified mutational hotspots of functional elements, genes, and pathways associated with known cancer drivers. (Item 56) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 500 bp of pre-identified mutational hotspots of functional elements, genes and pathways associated with altered drug response. (Item 57) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 500 bp of pre-identified mutational hotspots of functional elements, genes and pathways associated with other clinically valuable genes. (Item 58) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 1,000 bp of pre-identified mutational hotspots of functional elements, genes and pathways associated with Mendelian diseases. (Item 59) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 1,000 bp of pre-identified mutational hotspots of functional elements, genes, and pathways associated with known cancer drivers. (Item 60) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 1,000 bp of pre-identified mutational hotspots of functional elements, genes, and pathways associated with altered drug response. (Item 61) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 1,000 bp of pre-identified mutational hotspots of functional elements, genes, and pathways associated with other clinically valuable genes. (Item 62) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified constraint regions of functional elements, genes and pathways associated with Mendelian diseases. (Item 63) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified constrained regions of functional elements, genes and pathways associated with known cancer drivers. (Item 64) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified constraint regions of functional elements, genes and pathways associated with altered drug response. (Item 65) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified constraint regions of functional elements, genes and pathways associated with other clinically valuable genes. (Item 66) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified 10 bp constraint regions of functional elements, genes and pathways associated with Mendelian diseases. (Item 67) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified 10 bp constraint regions of functional elements, genes and pathways associated with known cancer drivers. (Item 68) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified 10 bp constraint regions of functional elements, genes and pathways associated with altered drug response. (Item 69) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified 10 bp constraint regions of functional elements, genes and pathways associated with other clinically valuable genes. (Item 70) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified 50 bp constraint regions of functional elements, genes and pathways associated with Mendelian diseases. (Item 71) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified 50 bp constraint regions of functional elements, genes and pathways associated with known cancer drivers. (Item 72) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified 50 bp constraint regions of functional elements, genes and pathways associated with altered drug response. (Item 73) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified 50 bp constraint regions of functional elements, genes and pathways associated with other clinically valuable genes. (Item 74) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified 100 bp constraint regions of functional elements, genes and pathways associated with Mendelian diseases. (Item 75) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified 100 bp constraint regions of functional elements, genes and pathways associated with known cancer drivers. (Item 76) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified 100 bp constraint regions of functional elements, genes and pathways associated with altered drug response. (Item 77) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified 100 bp constraint regions of functional elements, genes and pathways associated with other clinically valuable genes. (Item 78) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified 500 bp constraint regions of functional elements, genes and pathways associated with Mendelian diseases. (Item 79) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified 500 bp constraint regions of functional elements, genes and pathways associated with known cancer drivers. (Item 80) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within 500 bp pre-identified constraint regions of functional elements, genes and pathways associated with altered drug response. (Item 81) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified 500 bp constraint regions of functional elements, genes and pathways associated with other clinically valuable genes. (Item 82) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified 1,000 bp constraint regions of functional elements, genes and pathways associated with Mendelian diseases. (Item 83) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified constraint regions of 1,000 bp of functional elements, genes, and pathways associated with known cancer drivers. (Item 84) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified constraint regions of 1,000 bp of functional elements, genes and pathways associated with altered drug response. (Item 85) 2. The method of claim 1, wherein the molecular variants correspond to coding or non-coding variants within pre-identified constraint regions of 1,000 bp of functional elements, genes, and pathways associated with other clinically valuable genes. (Item 86) 2. The method of claim 1, wherein the phenotypic score of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment represents the phenotypic relevance of the molecular variants identified within the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment. (Item 87) 2. The method of claim 1, wherein the phenotypic scores of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment comprise lower phenotypic scores, and the lower phenotypic scores correspond to scores or classifications generated by a phenotypic model through the use of statistical learning techniques that relate molecular scores and molecular states of model systems to the phenotypic impact of molecular variants within each model system. (Item 88) 88. The method of claim 87, wherein the phenotypic model is generated using a neural network architecture for single-task or multi-task statistical learning that relates molecular scores from one or more functional components to one or more phenotypic effects of molecular variants in the one or more functional components. (Item 89) 2. The method of claim 1, wherein the phenotypic score of the single cell, the cellular compartment, the subcellular compartment, or the composite compartment comprises a higher phenotypic score, and wherein the higher phenotypic score is derived by applying an existing model relating lower phenotypic scores to regulation, signaling, pathways, processes, cell cycle activity, alterations, defects, or states. (Item 90) 2. The method of claim 1, wherein the phenotypic scores of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment comprise a higher-level phenotypic score, and the higher-level phenotypic score is derived from lower-level phenotypic scores via unsupervised learning, feature representation learning, or dimensionality reduction techniques. (Item 91) 2. The method of claim 1, wherein the phenotypic signal associated with the molecular variant comprises lower-level phenotypic signals associated with the molecular variant, and the lower-level phenotypic signals associated with the molecular variant are derived as summary statistics, descriptive statistics, inferential statistics, or Bayesian inference models of the phenotypic scores measured in the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment containing the molecular variant. (Item 92) 2. The method of claim 1, wherein the phenotypic signal associated with the molecular variant comprises a higher-level phenotypic signal associated with the molecular variant, and the higher-level phenotypic signal associated with the molecular variant is derived by applying an existing model relating lower-level phenotypic signals to regulation, signal transduction, pathways, processes, cell cycle activity, alterations, defects, or states. (Item 93) 2. The method of claim 1, wherein the phenotypic signal associated with the molecular variant comprises a higher-order phenotypic signal associated with the molecular variant, and the higher-order phenotypic signal associated with the molecular variant is derived from the lower-order phenotypic signals via unsupervised learning, feature representation learning, or dimensionality reduction techniques. (Item 94) Accessing a collection of molecular variants with putative or known phenotypic effects from existing sources; and utilizing a predictive model to expand the collection of molecular variants with putative or known phenotypic impact; utilizing a sampling model to select a first set of genotypes with putative or known phenotypic effects; utilizing a sampling model to select a second set of genotypes having unknown, putative, or known phenotypic effect; utilizing the sampling model to select a third set of genotypes having unknown, putative, or known phenotypic effect; generating a functional model by applying statistical learning techniques that relate molecular, phenotypic, or population signals of said first set of genotypes to putative or known phenotypic effects; generating predicted phenotypic effects for the second set of genotypes by applying the functional model to make predictions based on molecular, phenotypic, or population signals of the second set of genotypes; generating an inferential model by applying statistical learning techniques, the inferential model relating non-measured features to the phenotypic effects of molecular variants; 2. The method of claim 1, further comprising applying the inference model to make a prediction based on non-measured features of the third set of genotypes, thereby generating a predicted phenotypic impact of the third set of genotypes. (Item 95) 95. The method of item 94, wherein the predictive model is a gene-specific, region-specific, homolog-specific, or genome-wide calculated predictor or functional measure. (Item 96) Item 95. The method of item 94, wherein the predictive model provides a performance or confidence estimate for each prediction of the predictive model. (Item 97) Item 95. The method of item 94, wherein the positive predictive value (PPV) of the predictive model comprises a function of the predictive performance or confidence estimate of the predictive model. (Item 98) Item 95. The method of item 94, wherein the negative predictive value (NPV) of the predictive model comprises a function of the predictive performance or confidence estimate of the predictive model. (Item 99) 95. The method of claim 94, wherein the predictive model is a molecular impact predictor. (Item 100) 95. The method of item 94, wherein the predictive model predicts that premature termination, nonsense, or truncation molecular variants in protein-encoding functional elements are loss-of-function variants. (Item 101) 95. The method of item 94, wherein the predictive model predicts that synonymous or silent molecular variants in protein-encoding functional elements are neural variants. (Item 102) 2. The method of claim 1, further comprising generating a functional model by applying statistical learning techniques that combine the molecular signals, the phenotypic signals, or the population signals with the phenotypic effects of the molecular variants of the functional element. (Item 103) generating the functional model, 103. The method of claim 102, further comprising generating the functional model using a neural network architecture for single-task or multi-task learning that relates the molecular signals, the phenotypic signals, or the population signals from the functional elements to the one or more phenotypic effects of the molecular variants of the functional elements. (Item 104) 2. The method of claim 1, further comprising generating a phenotypic model by applying statistical learning techniques that combine the molecular scores with the phenotypic effects of the molecular variants of the functional elements. (Item 105) generating the phenotypic model, 105. The method of claim 104, further comprising generating a phenotypic model using a neural network architecture for single-task or multi-task learning that relates the molecular scores from the functional elements to the one or more phenotypic effects of the molecular variants of the functional elements. (Item 106) directing the molecular variant to the functional element within the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment; identifying the molecular variant within the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment; determining the phenotypic effect of the molecular variant in the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment; 2. The method of claim 1, further comprising determining a molecular measurement, molecular signature, or molecular process within the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment. (Item 107) 2. The method of claim 1, wherein the population signal associated with the molecular variant describes the distribution of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment associated with the molecular variant across subpopulations of single cells, cellular compartments, subcellular compartments, or synthetic compartments from different molecular states. (Item 108) 2. The method of claim 1, wherein the population signal associated with a molecular variant describes the dynamics of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment associated with the molecular variant across different molecular states and subpopulations of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment. (Item 109) 2. The method of claim 1, wherein the population signal associated with the molecular variant describes a change in the distribution of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment from different molecular states associated with the molecular variant across subpopulations of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment. (Item 110) 2. The method of claim 1, wherein the population signal associated with the molecular variant describes changes in the dynamics of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment from different molecular states associated with the molecular variant to subpopulations of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment. (Item 111) 108. The method of claim 107, wherein a clustering technique is applied to cluster and assign the single cells, the cellular compartments, the subcellular compartments, or the synthetic compartments based on the molecular score or the phenotypic score. (Item 112) 112. The method of claim 111, wherein a Gaussian mixture model (GMM) is applied to cluster and assign the single cells, the cellular compartments, the subcellular compartments, or the composite compartments to a defined number of molecular states. (Item 113) 112. The method of claim 111, wherein a variational Gaussian mixture model (VGMM) is applied to cluster and assign the single cell, the cellular compartment, the subcellular compartment, or the composite compartment to an inferred number of molecular states using a Dirichlet process. (Item 114) 108. The method of claim 107, wherein the population signal associated with the molecular variant is determined as a fraction of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment associated with the molecular variant corresponding to a particular molecular state. (Item 115) 2. The method of claim 1, wherein the molecular score or the phenotypic score of the molecular variant comprises an adjusted molecular score or phenotypic score calculated as the difference between the molecular score or the phenotypic score of the molecular variant and the molecular score or the phenotypic score of a reference molecular variant or a reference single cell, cellular compartment, subcellular compartment, or synthetic compartment. (Item 116) 2. The method of claim 1, wherein the molecular score or phenotypic score of the molecular variant comprises an adjusted molecular score or phenotypic score calculated by normalizing the molecular score or phenotypic score of the molecular variant to a reference molecular variant or a reference single-cell, cellular compartment, subcellular compartment, or synthetic compartment molecular score or phenotypic score. (Item 117) 2. The method of claim 1, wherein the molecular signal, phenotypic signal, or population signal of a molecular variant comprises an adjusted molecular signal, phenotypic signal, or population signal, respectively, calculated as the difference between the molecular signal, phenotypic signal, or population signal of a molecular variant and the molecular signal, phenotypic signal, or population signal of a reference molecular variant. (Item 118) 2. The method of claim 1, wherein the molecular signal, phenotypic signal, or population signal associated with the molecular variant comprises an adjusted molecular signal, phenotypic signal, or population signal, respectively, calculated by normalizing the molecular signal, phenotypic signal, or population signal associated with the molecular variant by the molecular signal, phenotypic signal, or population signal of a reference molecular variant. (Item 119) 2. The method of claim 1, wherein the molecular signal, phenotypic signal, or population signal associated with the molecular variant comprises an adjusted molecular signal, phenotypic signal, or population signal, respectively, calculated as a quantile of the molecular signal, phenotypic signal, or population signal associated with the molecular variant among molecular signals, phenotypic signals, or population signals of reference molecular variants. (Item 120) selecting a first set of genotypes having a phenotypic effect; selecting a second set of genotypes having a phenotypic effect; applying single cell capture or barcoding techniques to obtain molecules from single cells, cellular compartments, subcellular compartments, or synthetic compartments of a first number of cells associated with said first set of genotypes; obtaining a first number of molecular reads per model system by performing sequencing, sequence read quality control, cellular barcode identification or quality control, molecular barcode identification or quality control, alignment of sequence reads to a reference genome, or read alignment filtering or quality control using model systems associated with the first set of genotypes; applying single cell capture or barcoding techniques to obtain molecules from the single cells, the cellular compartments, the subcellular compartments, or the synthetic compartments of a second number of cells associated with the first set of genotypes; obtaining a second number of molecular reads per model by performing sequencing, sequence read quality control, cellular barcode identification or quality control, molecular barcode identification or quality control, sequence read alignment to a reference genome, or read alignment filtering or quality control using the model system associated with the first set of genotypes; Deriving total molecular reads or total molecular measurements from the molecular reads of the total number of reads per model system from single cells, cellular compartments, subcellular compartments, or synthetic compartments of the total number of cells per genotype; generating a total dimensionality reduced model by applying statistical learning techniques for feature selection or dimensionality reduction to determine a molecular score, phenotypic score, molecular signal, phenotypic signal, or population signal for the first set of genotypes using the total molecular reads and the total molecular measurements; generating a summation functional model by applying statistical learning techniques utilizing the summation molecular reads and the summation molecular measurements to relate molecular signals, phenotypic signals, or population signals from the summation dimensionality reduced model to phenotypic effects for the first set of genotypes; determining a threshold performance of a functional score or functional classification using the total cell count, the total read count, the total dimensionality reduction model, or the total functional model for predicting the phenotypic impact of the first set of genotypes; deriving an optimal molecular read or optimal molecular measurement from the molecular reads of an optimal number of reads per model system from single cells, cellular compartments, subcellular compartments, or synthetic compartments of an optimal number of cells per genotype, wherein the optimal molecular read and the optimal molecular measurement are obtained by subsampling the total molecular reads or the total molecular measurements; generating an optimal dimensionality reduced model by applying statistical learning techniques for feature selection or dimensionality reduction to determine a molecular score, phenotypic score, molecular signal, phenotypic signal, or population signal for the first set of genotypes using the optimal molecular reads and the optimal molecular measurements; generating an optimal functional model by applying statistical learning techniques utilizing the optimal molecular reads and the optimal molecular measurements to relate molecular, phenotypic, or population signals from the optimal dimensionality reduced model to phenotypic effects for the first set of genotypes; validating the threshold performance of the functional score or functional classification based on the optimal cell number, the optimal read number, the optimal dimensionality reduction model, or the optimal functional model for predicting the phenotypic impact of the first set of genotypes; applying single cell capture or barcoding techniques to obtain molecules from the single cells, cellular compartments, subcellular compartments, or synthetic compartments of the optimal cell number associated with the second set of genotypes; obtaining the optimal number of molecular reads for each model system by performing sequencing, sequence read quality control, cell barcode identification or quality control, molecular barcode identification or quality control, sequence read alignment to a reference genome, or read alignment filtering or quality control using model systems associated with the second set of genotypes; generating a functional score or functional classification for the second set of genotypes based on the optimal cell number, the optimal number of reads, the optimal dimensionality reduction model, or the optimal functional model. (Item 121) 1. A computer-implemented method for scoring the phenotypic impact of molecular variants, comprising: evaluating the evidence dataset based on the accuracy of the evidence dataset; validating the evidence dataset based on the accuracy of the evidence dataset; optimizing the evidence dataset based on the accuracy of the evidence dataset; determining the phenotypic impact of the molecular variants based on evaluating, validating, and optimizing the evidence dataset. (Item 122) 122. The method of claim 121, wherein the evidence dataset comprises a functional score or functional classification of a molecular variant based on a machine learning model relating a molecular signal, a phenotypic signal, or a population signal of the molecular variant to the phenotypic effect of the molecular variant. (Item 123) 122. The method of claim 121, wherein the evidence dataset comprises predictor scores or predictor classifications from genome-wide, homolog-specific, enzyme class-specific, region-specific, or gene-specific calculated predictors. (Item 124) 122. The method of claim 121, wherein the evidence dataset comprises hotspot scores or hotspot classifications from mutational hotspots. (Item 125) 122. The method of claim 121, wherein the evidence dataset comprises population scores or population classifications from variant classifications derived based on population genomics indices. (Item 126) 122. The method of claim 121, further comprising a calculated metric for assessing agreement between the evidence dataset and a functional score or functional classification. (Item 127) Item 122. The method of item 121, wherein the evaluation index comprises Pearson's correlation coefficient, Spearman's rank correlation, Kendall's correlation, Matthew's correlation coefficient, Cohen's kappa coefficient, Youden's index, F value, true positive rate, true negative rate, positive predictive value, negative predictive value, positive likelihood ratio, negative likelihood ratio, or diagnostic odds ratio. (Item 128) Item 122. The method of item 121, wherein validating the evidence dataset includes validating the evidence dataset based on the evaluation index. (Item 129) Item 122. The method of item 121, wherein optimizing the evidence dataset includes selecting or removing data in the evidence dataset based on the evaluation index. (Item 130) 1. A computer-implemented method for scoring the phenotypic impact of molecular variants, comprising: evaluating the evidence dataset based on inherent biases of the evidence dataset; validating the evidence dataset based on the inherent bias of the evidence dataset; optimizing the evidence dataset based on the inherent bias of the evidence dataset; determining a score of the phenotypic impact of the molecular variant based on evaluating, validating, and optimizing an evidence dataset. (Item 131) 131. The method of claim 130, wherein the bias of the evidence dataset is measured as the statistical distance between the observed evidence score or evidence classification of a variant in the evidence dataset relative to the expected evidence score or evidence classification of the variant in a reference dataset. (Item 132) 131. The method of claim 130, wherein the validation bias of the evidence dataset is measured as the statistical distance between observed features and properties of variants in the evidence dataset relative to expected features and properties of variants in a reference dataset, defined based on a matching quantile or classification. (Item 133) 131. The method of claim 130, wherein the validation bias of the evidence dataset is measured as the statistical distance between observed features and properties of the variants in the evidence dataset relative to expected features and properties of the variants in a reference dataset, defined based on a matching distribution of evidence scores or evidence classifications. (Item 134) Item 131. The method of item 130, wherein validating the evidence dataset includes validating the evidence dataset based on a target evaluation bias index. (Item 135) Item 131. The method of item 130, wherein optimizing the evidence dataset includes selecting or removing data within the evidence dataset based on target validation criteria. (Item 136) Memory and at least one processor coupled to the memory, receiving molecular variants associated with one or more functional elements in a model system, the model system comprising a single cell, a cellular compartment, a subcellular compartment, or a synthetic compartment; determining a molecular or phenotypic score for the single cell, the cellular compartment, the subcellular compartment, or the composite compartment; determining a molecular or phenotypic signal associated with a particular molecular variant based on the respective molecular or phenotypic scores of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment containing the molecular variant; determining a population signal associated with a particular molecular variant based on the molecular or phenotypic score of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment containing the molecular variant; determining a functional score or functional classification for the molecular variant based on statistical learning relating the molecular signal, the phenotypic signal, or the population signal of the molecular variant to the phenotypic effect of the molecular variant; deriving an evidence score or evidence classification for the molecular variant based on the functional score or functional classification, modeling of the functional score or functional classification, modeling of the predictor score or predictor classification, or modeling of the hotspot score or hotspot classification; and the at least one processor configured to determine the phenotypic impact of the molecular variant based on the functional score, the functional classification, the evidence score, or the evidence classification. (Item 137) A tangible computer-readable apparatus that, when executed by at least one computing device, causes the at least one computing device to: receiving molecular variants associated with one or more functional elements in a model system, wherein the model system comprises a single cell, a cellular compartment, a subcellular compartment, or a synthetic compartment; determining a molecular or phenotypic score for the single cell, the cellular compartment, the subcellular compartment, or the composite compartment; determining a molecular or phenotypic signal associated with a particular molecular variant based on the molecular or phenotypic score of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment containing the particular molecular variant; determining a population signal associated with a particular molecular variant based on the molecular or phenotypic score of the single cell, the cellular compartment, the subcellular compartment, or the synthetic compartment containing the molecular variant; determining a functional score or functional classification for the molecular variant based on statistical learning, wherein the statistical learning relates the molecular signal, the phenotypic signal, or the population signal of the molecular variant to a phenotypic impact of the molecular variant; deriving an evidence score or evidence classification for the molecular variant based on the functional score or functional classification, modeling of the functional score or functional classification, predictor score or modeling of predictor classification, or hotspot score or modeling of hotspot classification; determining the phenotypic impact of the molecular variant based on the functional score, the functional classification, the evidence score, or the evidence classification. [Brief explanation of the drawings]

[0009] [Figure 1A] 1 shows an integrated functional measurement and computational deep mutation learning (DML) process and system for determining the phenotypic impact of molecular variants, according to some embodiments, and exemplary (e.g., intermediate) data generated from application of the process and system in two genes of the RAS / MAPK family of disorders. [Figure 1B] Same as above. [Figure 1C] Same as above. [Figure 2A] 1 illustrates the performance of deep mutation learning (DML) processes and systems, according to some embodiments, in discriminating (e.g., binary classification) disease-causing (e.g., pathogenic) and neutral (e.g., benign) molecular variants for germline (e.g., inherited) and somatic disorders in three genes of the RAS / MAPK pathway, HRAS, PTPN11, and MAP2K2. [Figure 2B] Same as above. [Figure 3A] 1 shows the performance of deep mutation learning (DML) processes and systems in identifying (e.g., binary classification) cells containing germline disease-causing (e.g., pathogenic) or neutral (e.g., benign) molecular variants in MAP2K2, according to some embodiments. [Figure 3B] Same as above. [Figure 4] 1 illustrates the architecture of a neural network-based denoising autoencoder trained and applied to generate a robust, condensed representation of molecular scores, according to some embodiments. [Figure 5] 1 shows normalized ERK pathway activation measured as the fraction of total ERK protein phosphorylated via enzyme immunoassay of cell extracts from H293 cells containing control, wild-type, and mutant versions of MAP2K2 and PTPN11, according to some embodiments. [Figure 6]

[0023] Figure 1 shows an example of a method for reducing the cost of deploying deep mutation learning (DML) to identify the phenotypic impact of molecular variants through stepwise optimization and deployment of measurements using various cell numbers, read depths, dimensionality reduction models (mDR), and functional models (mF), according to some embodiments, where optimization is first performed on a (reduced) truth set of molecular variants and deployment includes a target set of molecular variants. [Figure 7] 1 shows an example of how a phenotype score is calculated, according to some embodiments. [Figure 8] 1 illustrates an example of how a molecular score is calculated, according to some embodiments. [Figure 9] 1 illustrates a method for calculating molecular signals associated with individual molecular variants, according to some embodiments. [Figure 10] 1 illustrates a method for calculating molecular state-specific, independent, or disjoint estimates of molecular signals, according to some embodiments. [Figure 11] 1 illustrates a method for characterizing the distribution of cells with specific molecular variants across molecular states or phenotypic scores and deriving a population signal, according to some embodiments. [Figure 12] 1 shows an example of how unsupervised learning techniques can be utilized to distinguish higher-order from lower-order molecular signals associated with individual molecular variants, according to some embodiments. [Figure 13] 1 shows an example of how functional scores and classifications are derived via machine learning to relate molecular, phenotypic, or population signals to the phenotypic impact of molecular variants via regression and classification techniques, according to some embodiments. [Figure 14A] 1 shows examples of the performance of methods and systems for binary classification of molecular variants with two different phenotypic effects, as trained using various numbers of cells, according to some embodiments. [Figure 14B] Same as above. [Figure 15]

[0023] Figure 1 shows an example of how, according to some embodiments, functional scores and functional classifications from a subset of potential nonsynonymous variants can be utilized to enable the inference of a sequence-function map that describes the functional score or functional classification for all potential nonsynonymous variants in a protein-coding gene. [Figure 16] 1 illustrates an example system and method for determining the phenotypic impact of molecular variants through a series of modeling layers, according to some embodiments, that reduces the cost and increases the scope of DML processing. [Figure 17] 1 shows an example of how machine learning techniques are utilized to generate a lower level variant interpretation engine (VIE), which can be gene and condition specific, according to some embodiments. [Figure 18] 1 shows an example of a method for the identification of significantly mutated regions (SMRs) and networks (SMNs), according to some embodiments. [Figure 19] 1 is an exemplary computer system useful for implementing various embodiments.

[0010] In the drawings, like reference numbers indicate identical or similar elements. Additionally, the leftmost digit(s) of a reference number generally identifies the drawing in which the reference number first appears. DETAILED DESCRIPTION OF THE INVENTION

[0011] Provided herein are embodiments of systems, instruments, devices, methods and / or computer program products, and / or combinations and subcombinations thereof, to enable multi-functional, multi-component and multi-genic (e.g., pathway-scale) assessment of the phenotypic impact of variants across a wide variety of variant types, biophysical processes, molecular functions, and phenotypes.

[0012] The present disclosure provides embodiments of systems, instruments, devices, methods and / or computer program products that may leverage high-throughput molecular measurements (e.g., next-generation sequencing), single-cell manipulation, molecular biology, computational modeling, and statistical learning techniques, and may enable multi-functional, multi-component, and multi-gene (pathway-scale) assessment of the phenotypic impact of variants across a wide variety of variant types, biophysical processes, molecular functions, and phenotypes.

[0013] The present disclosure provides embodiments of systems, apparatus, devices, methods and / or computer program products for systematically determining and statistically validating one or more phenotypic (e.g., clinical or non-clinical) impacts (e.g., pathogenicity, functionality, or comparative effect) of identified molecular variants, such as genotypic (e.g., sequence) variants, in one or more (e.g., coding or non-coding) functional elements (e.g., molecular regions such as protein-coding genes, non-coding genes, protein or RNA regions, promoters, enhancers, silencers, regulatory binding sites, origins of replication, etc.) in the (e.g., nuclear, mitochondrial, etc.) genome(s), or derivable molecules thereof, within a subject's biological sample or record thereof.

[0014] The present disclosure provides embodiments of systems, devices, apparatus, methods and / or computer program products for classification (or regression) of putative phenotypic effects in a subject based on one or more molecular, phenotypic, or population signals measured in an in vivo or in vitro functional model system. The derived regression or classification may be referred to as a functional score or functional classification.

[0015] The embodiments herein represent a departure from existing computational or functional evidence-supported systems for molecular variant classification, such as those utilized in clinical genetic and genomic diagnostics.

[0016] First, existing computational methods and systems for variant classification rely on a wide variety of population, evolutionary, physicochemical, structural, and / or molecular annotations and properties for variant classification, but they do not utilize information about the impact of molecular variants on cell biology. As a result, such computational methods are unable to capture phenotypic effects that act through changes in molecular properties within cells or through changes in cell population and cellular heterogeneity.

[0017] Second, existing large-scale functional assays and solutions capable of measuring the activity of thousands of molecular variants provide activity measurements along a single dimension per molecular variant and often require a priori knowledge or assumption of the mechanism of action through which the molecular variant exerts its phenotypic effect.

[0018] Due to these limitations, although conventional computational methods and systems for variant classification can access data across a large number of annotations and parameters, these conventional approaches perform significantly poorly in classification (and regression) tasks regarding the phenotypic effects of molecular variants. Similarly, these conventional approaches require a priori knowledge or assumptions about the mechanism of action (and therefore the associated molecular function to be measured), thereby limiting their application to well-characterized functional elements (e.g., genes). This further precludes their application to poorly understood disease-related genes. Finally, these conventional approaches require significant development and customization to measure each molecular function and each functional element.

[0019] In embodiments herein, a technical solution to overcome these technical problems involves a data structure that provides a multidimensional assessment of cells and cell populations containing specific genotypes (e.g., molecular variants) in one or more functional elements (e.g., genes) and in one or more contexts (e.g., cell type, drug treatment, genotype background). Such a data structure enables systems and methods for statistical learning to achieve improved accuracy in classification tasks regarding the phenotypic impact of genotypes (e.g., molecular variants or combinations thereof).

[0020] Embodiments herein provide for the generation of hundreds to tens of thousands (~10 2 -10 4 ) molecular measurements, tens to thousands (~10 1 -10 3 ) molecular image construction, thousands (~10 3 ) and enables robust, scalable, multidimensional classification of molecular variants (and combinations thereof) across a wide variety of functional elements and phenotypes, through single or multiple functional elements in parallel.

[0021] As shown in FIG. 1A, embodiments of the present disclosure integrate mutant library generation 102 and cell library generation 104 methods for high-throughput mutagenesis and cell engineering techniques for generating profiles of model systems (e.g., cells) containing distinct molecular variants in target functional elements (e.g., genes). The present embodiments provide treatment, single-cell capture, library preparation, and sequencing 106 methods that utilize cellular, molecular biology, and genomics technologies and techniques for treatment and capture of model systems, preparation of libraries of molecular entities, and measurement of diverse molecular entities (e.g., transcripts) within the model systems. The present embodiments provide mapping, normalization 108 bioinformatics, computational biology, and statistical techniques for mapping, quantification, and normalization of relationships between molecular variants, model systems, and molecular entities within each model system. The present embodiments provide feature selection, dimensionality reduction 110 and contextualization, training, and classification 112 statistical (e.g., machine) learning, distributed high-performance computation, systems biology, population, and clinical genomics techniques for label generation, feature selection, dimensionality reduction, training, and classification of molecular variants.

[0022] In some embodiments, the present disclosure describes the use of these methods and techniques in FIG. 1A to determine the phenotypic impact of molecular variants identified in a biological sample. In some embodiments, the present disclosure describes the induction of molecular variants into one or more functional elements in a model system. The model system may include a single cell, a cellular compartment, a subcellular compartment, or a synthetic compartment. In some embodiments, the present disclosure describes the determination of a molecular or phenotypic score of a single cell, a cellular compartment, a subcellular compartment, or a synthetic compartment. In some embodiments, the present disclosure describes the identification of molecular variants within a single cell, a cellular compartment, a subcellular compartment, or a synthetic compartment. As will be appreciated by those skilled in the art, various methods can be utilized to identify molecular variants within a single cell, a cellular compartment, a subcellular compartment, or a synthetic compartment. This may be based on molecular measurements of the single cell, a cellular compartment, a subcellular compartment, or a synthetic compartment. In some embodiments, the present disclosure describes determining molecular or phenotypic signals associated with individual molecular variants based on molecular or phenotypic scores, respectively, from single cells, cellular compartments, subcellular compartments, or synthetic compartments associated with particular molecular variants. In some embodiments, the present disclosure describes determining population signals associated with molecular variants based on molecular or phenotypic scores from single cells, cellular compartments, subcellular compartments, or synthetic compartments associated with particular molecular variants.

[0023] In some embodiments, the present disclosure describes determining a functional score or functional classification of a molecular variant by applying a statistical (e.g., machine) learning approach that relates molecular, phenotypic, or population signals to the phenotypic impact of the molecular variant. In some embodiments, the present disclosure describes determining an evidence score or evidence classification of a molecular variant based on a functional score, functional classification, predictor score, predictor classification, hotspot score, or hotspot classification. In some embodiments, the present disclosure describes determining the phenotypic impact of a molecular variant identified in a biological sample based on a functional score, functional classification, evidence score, or evidence classification of the identified molecular variant.

[0024] Embodiments herein integrate methods, techniques, and science from many disciplines. 2 Statistical and machine learning techniques leveraging single-cell molecular measurements have been developed and applied for classification of model systems (e.g., cells) derived from different tissues or developmental stages (< 3 × 10), but not within the same cell line, tissue, or developmental stage. 9 The need to achieve accurate genotype-specific (e.g., molecular variant-specific) classification among thousands of cells with subtle differences such as single-base differences within a genomic background defined by larger nucleotides can present a significant challenge.

[0025] The present disclosure provides embodiments of deep mutation learning (DML) systems, instruments, devices, methods and / or computer program products, and / or combinations and subcombinations thereof, for overcoming challenges in identifying (e.g., classifying) the phenotypic impact of molecular variants identified in a subject based on biological signals measured in single or population model systems (e.g., cells).

[0026] The present disclosure provides system, apparatus, device, method and / or computer program product embodiments, and / or combinations and sub-combinations thereof, that improve cost-efficiency in molecular variant classification through (i) directed deployment of DML processes and systems with lower cost predictive models (see FIG. 16 ), and (ii) layered deployment of DML processes and systems that enable robust reconstruction of molecular signals at reduced cost (see FIG. 6 ).

[0027] The present disclosure provides system, apparatus, device, method and / or computer program product embodiments, and / or combinations and subcombinations thereof, that improve scalability and performance systems across functional elements (e.g., genes) through DML processing and systems that leverage information between functional elements (see Figures 3A and 3B).

[0028] The present disclosure provides embodiments of systems, apparatus, devices, methods, and / or computer program products, and / or combinations and subcombinations thereof, for assessing the phenotypic impact (e.g., pathogenicity, functionality, or comparative effect) of one or more molecular (e.g., genotypic) variants in one or more (e.g., coding or non-coding) functional elements (e.g., molecular regions such as protein-coding genes, non-coding genes, protein or RNA regions, promoters, enhancers, silencers, regulatory binding sites, origins of replication, etc.) in a (e.g., nuclear, mitochondrial, etc.) genome(s), or their derivable molecules. As will be appreciated by those skilled in the art, molecular variants may be genotypic (e.g., sequence) variants such as single nucleotide variants (SNVs), copy number variants (CNVs), or insertions or deletions affecting coding or non-coding sequences (or both) in nuclear, mitochondrial, or natural or synthetic episomal genomes. As will be appreciated by those skilled in the art, a molecular variant may also be a single amino acid substitution in a protein molecule, a single base substitution in an RNA molecule, a single base substitution in a DNA molecule, or any other molecular alteration to the cognate sequence of a polymeric biological molecule.

[0029] In some embodiments, classification (or regression) may relate to (e.g., putative) disease-causing (e.g., pathogenic) and neutral (e.g., benign) variants for predicting disorders having a genetic component or their severity based on molecular variants identified within a subject's biological sample or record. In some other embodiments, classification (or regression) may relate to molecular impact (e.g., loss-of-function, gain-of-function, or neutral) based on putative molecular outcome (e.g., nonsense or insertion and deletion mutations) and putative neutral (e.g., synonymous) molecular variants. In some other embodiments, classification (or regression) may relate to changes in response to therapeutic treatment (e.g., chemical, biochemical, physical, behavioral, digital, or other) based on molecular variants identified within a subject's biological sample or record. In some embodiments, phenotypic impact may refer to phenotypic class (e.g., probability of neutral, pathogenic, benign, high-risk, low-risk, positive-response variant, negative-response variant) and phenotypic score (e.g., probability of developing specific clinical and non-clinical phenotypes, levels of metabolites in the blood, and the rate at which specific compounds are absorbed or metabolized).

[0030] In some embodiments, the present disclosure provides systems and methods for modeling the diversity and prevalence of phenotypic properties within a population based on the diversity and prevalence of molecular variants in a representative population. In some embodiments, the present disclosure provides systems and methods for modeling the diversity and prevalence of phenotypic properties within a population based on the phenotypic effects of molecular variants with known or expected diversity and prevalence, where the phenotypic effects may be modeled from one or more molecular, phenotypic, or population signals previously associated with variants in an in vivo or in vitro functional model system. In some embodiments, such modeling may be utilized to inform the diversity and prevalence of mechanisms of drug resistance in a population.

[0031] In some embodiments, the present disclosure describes the use of models of the diversity and prevalence of phenotypic properties within a population of individuals (e.g., as informed by the phenotypic effects of molecular variants modeled from one or more molecular, phenotypic, or population signals in a functional model system) to construct cohorts of subjects (e.g., patients) and investigate the efficacy of therapeutic and non-therapeutic interventions.

[0032] In some embodiments, the present disclosure provides systems and methods for classification (or regression) of the phenotypic impact of molecular variants based on a functional score or functional classification derived from one or more molecular, phenotypic, or population signals associated with the variants as measured in a functional model system. In some embodiments, molecular variants may be functionally modeled within cells, cellular compartments, or synthetic compartments, such as in in vivo or in vitro model systems.

[0033] In some embodiments, modeled molecular variants (e.g., in vivo or in vitro) may be identified directly within the nucleic acid sequences of the modeled functional elements via library preparation, sequencing, and evaluation of nucleic acids or nucleic acid fragments within a single cell, cellular compartment, subcellular compartment, or synthetic compartment (e.g., collectively referred to as a model system). In some other embodiments, modeled molecular variants (e.g., in vivo or in vitro) may be inferred from barcode sequences associated with individual variants in the functional elements via library preparation, sequencing, and evaluation of nucleic acids or nucleic acid fragments within a model system (e.g., a single cell, cellular compartment, subcellular compartment, or synthetic compartment) utilizing a pre-assembled database of associated barcodes and variants. As will be appreciated by those skilled in the art, molecular variants may be produced through a variety of techniques, such as direct (e.g., chemical) synthesis, error-prone PCR, oligonucleotide-directed mutagenesis, nicking mutagenesis, or saturation genome editing (SGE), among others (Firnberg et al. 2012; Kitzman et al. 2014; Wrenbeck et al. 2016; and Findlay et al. 2014). As will be appreciated by those skilled in the art, the variant library can then be introduced (e.g., added) into a model system (e.g., a cell, a cellular compartment, a subcellular compartment, or a synthetic compartment) using a variety of approaches, including, but not limited to, homologous recombination (e.g., Cas9-mediated or adenovirus-mediated), site-specific recombination (e.g., Flp-mediated), or viral transduction (e.g., lentivirus-mediated) (Findlay et al. 2018; Wissink et al. 2016; and Macosko et al. 2015).

[0034] In some embodiments, in vivo or in vivo analyses of DNA, RNA, and protein molecules or modifications thereof containing mutations in functional elements, including but not limited to: Functional scores and functional classifications associated with individual molecular variants may be derived from measurements of molecules and / or chemical modifications present in an in vitro model system. For example, in some embodiments, measurements or models of molecular, cellular, or population signals may be created and utilized to learn functional scores and / or functional classifications. In some embodiments, functional scores and functional classifications may be derived from molecular measurements obtained through nucleic acid barcoding, isolation, enriched library preparation, sequencing, and evaluation of multiple nucleic acids or nucleic acid fragments within a single cell, cellular compartment, subcellular compartment, or synthetic compartment, including, but not limited to, RNA molecules, genomic DNA, chromatin-bound DNA, protein-bound DNA, accessible DNA fragments, or chemically modified nucleic acids. In some embodiments, these procedures may utilize molecular barcoding techniques to uniquely identify or associate nucleic acids, nucleic acid fragments, or nucleic acid sequences originating from individual single cells, cellular compartments, subcellular compartments, or synthetic compartments (Macosko et al. 2015; Buenrostro et al. 2015; Cusanovich et al. 2015; Dixit et al. 2016; Adamson et al. 2016; Jaitin et al. 2016; Datlinger et al. 2017; Zheng et al. 2017; Cao et al. 2017). These methods may build on developments from the field of single-cell genomics (Schwartzman and Tanay 2015; Tanay and Regev 2017; Gawad et al. 2016). In some embodiments, the systems and methods of the present disclosure may apply single-cell RNA-seq methods to derive molecular measurements from single cells, cellular compartments, subcellular compartments, or synthetic compartments.These methods include, but are not limited to, single-cell sequence library generation, high-throughput nucleic acid sequencing, sequence read quality control, barcode identification and quality control (e.g., of single cells, cellular compartments, subcellular compartments, or synthetic compartments), sequence read-specific molecular barcode identification and quality control, sequence read alignment, and read alignment filtering and quality control. In some embodiments, molecular measurements may correspond to locus-specific measurements of gene expression (e.g., RNA transcript abundance), protein abundance or modification (e.g., phosphoprotein abundance), chromatin accessibility (e.g., nucleosome occupancy), epigenetic modifications (e.g., DNA methylation), regulatory activity (e.g., transcription factor binding), post-transcriptional processing (e.g., splicing), post-translational modifications (e.g., ubiquitination), mutation abundance (e.g., number), mutation rate (e.g., frequency), mutation signatures (e.g., number or frequency of each type of mutation), or various other types of molecular measurements within single cells, cellular compartments, subcellular compartments, or synthetic compartments, as will be understood by those skilled in the art. In some embodiments, the present disclosure describes systems and methods for increasing the quality of molecular measurements for specific target genes and functional elements through the use of targeted enrichment or targeted capture techniques, via hybridization or amplicon-based techniques and interrogation, before, during, or after single-cell RNA library processing.

[0035] In some embodiments, molecular measurements from single cells, cellular (or subcellular) compartments, or synthetic compartments may be utilized to derive multilocus measures of molecular processes. For example, these measures of molecular processes may include multilocus measures of gene expression, chromatin accessibility, epigenetic modifications, regulatory activity, transcriptional activity, translational activity, signal transduction activity, pathway activity, mutational burden, mutation rate, mutational signatures, and a variety of other measures, as will be appreciated by those skilled in the art.

[0036] In some embodiments, molecular measurements and molecular processes from single cells, cellular (or subcellular) compartments, or synthetic compartments may be utilized to derive global (e.g., pan-locus or locus-independent) measures of molecular features. For example, these measures of molecular features may include global measures of gene expression, chromatin accessibility, epigenetic modifications, regulatory activity, transcriptional activity, translational activity, signal transduction activity, pathway activity, mutational burden, mutation rate, mutational signatures, and a variety of other measurements, as will be appreciated by those skilled in the art.

[0037] In some embodiments, molecular measurements, processes, or features of a single cell, cellular compartment, subcellular compartment, or composite compartment may directly contribute to a (e.g., lower-level) molecular score. In some embodiments, a (e.g., higher-level) molecular score may be derived by applying an existing model relating multiple lower-level (e.g., lower-level) molecular scores (e.g., molecular measurements, processes, or features) to a regulation, signaling pathway, process, cell cycle activity, alteration, defect, or state. In some embodiments, such methods may apply gene set enrichment analysis or other derivable methods, as will be appreciated by those skilled in the art. In some embodiments, as shown in Figure 8, molecular measurements, molecular processes, molecular features, or (e.g., lower-level) molecular scores 806 from single cells, cellular compartments, subcellular compartments, or synthetic compartments containing the same molecular variant 802 may be fed through a series of artificial neuron layers (e.g., convolutional or perceptron layers) in an artificial neural network 804 (ANN) to derive increasingly complex (e.g., higher-level) molecular scores 806 and generate an autoencoder with learned features. In some embodiments, molecular score calculation methods such as pathway-level analysis may be utilized to preserve information of biological function while allowing for dimensionality reduction.

[0038] In some embodiments, a database of molecular scores may be constructed via a cell scoring layer 902 from multiple individual single cells, cellular compartments, subcellular compartments, or synthetic compartments, as shown in Figure 9. In some embodiments, molecular scores from multiple single cells, cellular compartments, subcellular compartments, or synthetic compartments containing the same molecular variants 906 (e.g., v1, v2, and v3) may be accessed using a variant sampling layer 908 and analyzed in a variant scoring layer 910 to derive (e.g., direct measurements or models) summary statistics regarding the trend (e.g., mean, median, mode), variance (e.g., variation, standard deviation), shape (e.g., skewness, kurtosis), probability (e.g., quantiles), range (e.g., confidence interval, minimum, maximum), error (e.g., standard error), or covariation (e.g., covariance) of the molecular scores associated with the individual molecular variants. In some embodiments, summary statistics regarding the trend, variance, shape, range, or error of molecular scores may be utilized to generate a database of (e.g., quality-controlled) molecular signals 912 associated with individual molecular variants 906, as shown in Figure 9. In some embodiments, molecular measurements, molecular processes, molecular features, and molecular scores 904 may be properties of individual single cells, cellular compartments, subcellular compartments, or composite compartments. In some embodiments, molecular signals may be properties of molecular variants.

[0039] As will be appreciated by those skilled in the art, molecular measurements, processes, features, and scores from a model system (e.g., a single cell, a cellular compartment, a subcellular compartment, or a synthetic compartment) may define or correspond to different molecular states or specific subpopulations of the model system (e.g., a single cell, a cellular compartment, a subcellular compartment, or a synthetic compartment) having similar molecular properties. As will be appreciated by those skilled in the art, and as shown in Figure 10, a cell scoring layer 1002 can be applied to determine the molecular state, phenotypic score 1006 (e.g., s1, s2, s3) of the model system based on a variety of methods.

[0040] For example, molecular states of a model system can be identified based on cell cycle signatures derived from gene expression molecular scores (Macosko et al. 2015). As will be appreciated by those skilled in the art, molecular states can be derived through scoring using pre-derived models, such as scoring gene expression signatures of pre-characterized molecular states, such as gene expression signatures reflecting different phases of the cell cycle pre-characterized in chemically synchronized cells (Whitfield et al. 2002). As will be appreciated by those skilled in the art, molecular states can also be derived through scoring using internally derived models from partitions of the model system in which feature correlations between molecular signals (e.g., gene expression changes across different phases of the cell cycle) can be detected or predicted. As will be appreciated by those skilled in the art, internally derived models can be generated using various statistical techniques (e.g., machine learning techniques).

[0041] In some embodiments, as shown in FIG. 7 , the present disclosure provides a phenotypic model (m) for derivation of a phenotypic score through the use of statistical techniques (e.g., machine learning techniques) that relate molecular scores and the molecular state of a model system (e.g., single cell, cellular compartment, subcellular compartment, or synthetic compartment) to the phenotypic impact of molecular variants within each model system. PThe present invention provides systems and methods for generating molecular scores. While molecular scores may relate directly to molecular, biological, or physical properties within a particular model system, phenotypic scores may describe the (e.g., putative) phenotypic relevance of molecular variants. In some embodiments, phenotypic scores are derived by applying supervised learning techniques to relate the phenotypic impact (e.g., labels) of molecular variants within a model system to the molecular scores or molecular states (e.g., features) of the model system.

[0042] In some embodiments, a phenotypic model (m) of the phenotypic score (or phenotypic classification) is generated by accessing a database of features describing the (e.g., lower or higher) molecular score and molecular state 704 of the single cell 702, as well as input labels 708 (e.g., a database) describing the phenotypic impact 706 identified within the molecular variant single cell 702. P ) and database are generated. In some embodiments, the training / validation layer 710 generates a phenotypic model (m) that can predict the phenotypic effect 706 of an individual single cell 702. P In some embodiments, a database of features describing the molecular scores and molecular states 716 of a single cell (test) 714 is used to generate and quality control a generated phenotypic model (m) to calculate and generate a database of phenotypic scores 720 describing the predicted phenotypic impact 718 of molecular variants in the single cell (test) 714. P ) are provided to the training, validation, or testing stratum. As will be appreciated by those skilled in the art, the performance (e.g., accuracy) of the predicted phenotypic impact 718 in each cell (e.g., phenotypic score 720) can be determined relative to the phenotypic impact of known molecular variants in single cells (test) 714 in the test stratum 712. As will be appreciated by those skilled in the art, a phenotypic model (m P) may be applied. In some embodiments, such scoring and evaluation may occur in a phenotype scoring and classification layer 722. The phenotype scoring and classification layer 722 may consider the accuracy of classification of possible phenotype effects based on the phenotype scores 720.

[0043] In some embodiments, summary statistics regarding the trend, variance, shape, range, or error of phenotypic scores may be utilized to generate a database of (e.g., quality-controlled) phenotypic signals associated with individual molecular variants.

[0044] In some embodiments, and as shown in Figure 10, the present disclosure describes the use of molecular state-specific molecular signals for subsequent rounds of unsupervised and supervised learning in the generation of molecular state-specific or multi-state models. In some embodiments, and as shown in Figure 10, the present disclosure describes the use of a molecular state, variant-specific sampling layer 1008 to access molecular measurements, treatments, features, and scores 1004 and molecular state, phenotypic scores 1006 for model systems that have specific molecular variants 1010 (e.g., v1, v2, v3) and are in specific molecular states, with feature-phenotypic scores, or combinations thereof. In some embodiments, the molecular measurements, treatments, features, and scores 1004 or molecular state, phenotypic scores 1006 may be pre-computed or calculated on demand by the cell scoring layer 1002. In some embodiments, data, summary statistics, descriptive statistics (e.g., univariate, bivariate, or multivariate analysis), inferential statistics, Bayesian estimation models (e.g., variational Bayesian estimation models), Dirichlet processes, or other models are utilized for the data accessed by the molecular state, variant-specific sampling layer 1008 to construct a molecular, phenotypic signal matrix 1012 that describes the molecular and phenotypic signals at each molecular state for each molecular variant.

[0045] In some embodiments, the molecular, phenotypic signal matrix 1012 may be pre-computed or calculated on demand. In some embodiments, the molecular, phenotypic signal matrix 1012 may be pre-computed or calculated on demand by the molecular state, variant-specific scoring layer 1016, producing matrices that are molecular state specific. In some embodiments, the molecular, phenotypic signal matrix 1012 may be pre-computed or calculated on demand by the multi-state, variant-specific scoring layer 1014, producing matrices that include data from multiple molecular states.

[0046] In some embodiments, as shown in FIG. 11 , the present disclosure provides methods for characterizing the distribution of cells with specific molecular variants across molecular states (e.g., subpopulations) or phenotypic scores 1106, as produced by cell scoring layer 1102 utilizing molecular measurements, processes, features, and scores 1104 as inputs. These molecular states (e.g., subpopulations) or phenotypic scores may be associated with subpopulations of cells defined by, but not limited to, (a) feature levels of or correlations between molecular signals (e.g., cyclin-dependent kinases during cell cycle stages) determined by application of pre-existing or internally derived models, (b) feature levels of or correlations between phenotypic scores, or (c) dimensionality reduction techniques, including, but not limited to, principal component analysis (PCA), independent component analysis (ICA), and t-SNE (tSNE), as examples. 11, for each individual molecular variant 1110, the population sampling layer 1108 may generate a measure of the comparative representation (e.g., distribution, probability, etc.) of cells across molecular states (e.g., proportion or probability of variant-containing cells in a molecular state) or phenotypic scores (e.g., proportion or probability of variant-containing cells having a particular score), and may also serve to provide a population signal matrix 1112 that describes how the molecular variants affect cells at the population level. The population signal matrix 1112 may include multiple population signals for multiple molecular variants.

[0047] In some embodiments, subsampling of model systems (e.g., single cells, cellular compartments, subcellular compartments, or synthetic compartments) from molecular measurements, molecular treatments, molecular features, molecular scores, or phenotypic scores containing the same molecular variants may be applied to generate independent or disjoint estimates of summary statistics regarding the trend, variance, shape, probability, range, covariation, or error of molecular measurements, molecular treatments, molecular features, or molecular scores or phenotypic scores associated with individual molecular variants.

[0048] In some embodiments, independent or disjoint estimates of molecular measurements, molecular processes, molecular features, molecular scores, or summary statistics regarding trend, variance, shape, probability, range, covariation, or error of the molecular measurements, molecular processes, molecular features, molecular scores, or phenotypic scores may be utilized to generate a database of (quality-controlled) independent or disjoint estimates of molecular or phenotypic signals associated with individual molecular variants. As will be appreciated by those skilled in the art, independent or disjoint estimates of molecular or phenotypic signals may be utilized to generate a database of (quality-controlled) independent or disjoint estimates of molecular or phenotypic signals associated with individual molecular variants.

[0049] In some embodiments, the present disclosure describes systems and methods for deriving independent or disjoint estimates of summary statistics regarding the trend, variance, shape, probability, range, covariation, or error of molecular measurements, molecular treatments, molecular features, or molecular or phenotypic scores associated with individual molecular variants within a subpopulation of a model system (e.g., a single cell, a cellular compartment, a subcellular compartment, or a synthetic compartment) from a particular molecular state. As will be appreciated by those skilled in the art, these methods may utilize multiple statistical techniques (e.g., machine learning techniques).

[0050] In some embodiments, molecular state-specific, independent, or disjoint estimates of summary statistics regarding the trend, variance, shape, probability, range, covariation, or error of molecular measurements, molecular treatments, molecular features, molecular scores, or phenotypic scores may be utilized to generate a database of (e.g., quality-controlled) molecular state-specific, independent, and disjoint estimates of molecular and phenotypic signals associated with individual molecular variants in a particular molecular state.

[0051] In some embodiments, independent or disjoint estimates of summary statistics regarding the trend, variance, shape, probability, range, covariation, or error of population signals associated with individual molecular variants may be utilized to generate a database of (e.g., quality-controlled) population signals associated with individual molecular variants.

[0052] 12 , the present disclosure provides systems and methods that utilize a feature extraction layer 1208 (e.g., unsupervised learning techniques) for the identification of higher-order molecular, phenotypic, or population signals from lower-order molecular, phenotypic, or population signals 1204 associated with individual molecular variants 1202, including, but not limited to, feature representation learning (or representation learning) techniques that deploy an artificial neural network (ANN) 1210 to generate an autoencoder that can exploit lower-order associations to produce higher-order representations of lower-order molecular, phenotypic, or population signals. In some embodiments, these methods enable the construction of a database of lower-order and higher-order molecular, phenotypic, and population signals 1214. In some embodiments, the feature extraction layer 1208 may access or accept data from annotation features 1206 in addition to the lower-order molecular, phenotypic, or population signals 1204. In some embodiments, annotation features 1206 may include multiple independent (e.g., non-measured) features (e.g., evolutionary, population, functional (e.g., annotation-based), structural, dynamic, and physicochemical features associated with variants, genomic coordinates, transcriptional (e.g., RNA) coordinates, translated (e.g., protein) coordinates, amino acids, and various others, as will be understood by those skilled in the art) that describe changes associated with changes in genotype (e.g., sequence, molecular variants, etc.).

[0053] In some embodiments, the present disclosure describes the use of molecular state-specific, lower-level molecular or phenotypic signals to derive molecular state-specific, higher-level molecular or phenotypic signals. In some embodiments, the present disclosure describes the use of multi-state matrices of lower-level molecular, phenotypic, or population signals to derive multi-state, higher-level molecular, phenotypic, or population signals that exploit structured relationships between molecular signals across molecular states, such as structured gene expression patterns (e.g., molecular signals) across cell cycle stages (e.g., molecular states). In some embodiments, the present disclosure describes the use of convolutional neural networks (CNNs) to learn patterned associations in molecular, phenotypic, or population signals (and annotation features) across molecular states.

[0054] In some embodiments, and as shown in FIG. 13, the present disclosure provides functional models (m) that relate molecular, phenotypic, or population signals (e.g., features), i.e., single or multiple molecular measurements, molecular treatments, molecular features, and molecular scores, to the phenotypic impacts (e.g., labels) of molecular variants via regression and classification techniques, respectively. F The present invention provides a system and method for deriving feature scores and feature classifications via statistical (e.g., machine) learning to generate a feature score.

[0055] In some embodiments, a functional model (m) of functional scores (or functional classifications) is generated by accessing a database of features describing the molecular (e.g., lower or higher), phenotypic, or population signals 1304 of the molecular variants 1302 for training / validation, and a collection (e.g., database) of input labels 1310 describing the phenotypic impact 1308 of the molecular variants 1302. F ) and a database is generated. Generation is further accomplished by applying statistical (e.g., machine) learning techniques to relate molecular, phenotypic, or population signals 1304 (e.g., features) to phenotypic effects (e.g., labels).

[0056] In some embodiments, the training / validation layer 1312 generates a quality control function model (m) that can predict the phenotypic impact 1308 of a molecular variant 1302. F In some embodiments, the training / validation layer 1312 may deploy cross-validation techniques such as, but not limited to, K-fold or leave-one-out cross-validation (LOOCV). In some embodiments, the generated functional models (m) are used to compute and generate a database of functional scores 1324 describing the predicted phenotypic impact 1322 of the molecular variants (tests) 1316. F ) may be provided with a database of features describing the molecular, phenotypic, or population signal 1318 of the molecular variant (test) 1316. As will be appreciated by those skilled in the art, the performance (e.g., accuracy) of the predicted phenotypic effect 1322 (e.g., functional score 1324) of the molecular variant may be determined relative to the phenotypic effect of known molecular variants, such as the test molecular variant 1316. As will be appreciated by those skilled in the art, functional models (m F ) may be applied. In some embodiments, such scoring and evaluation may occur in a functional scoring and classification layer 1326, for example, to consider the accuracy of classification of possible phenotypic effects based on functional scores 1324.

[0057] In some embodiments, the functional model (m FDuring training and testing (prediction generation) of the annotation features 1306, 1320, additional annotation features 1306, 1320 may be provided. In some embodiments, annotation features 1306 and 1320 may include multiple independent (e.g., non-measured) features (e.g., evolutionary, population, functional (e.g., annotation-based), structural, dynamic, and physicochemical features associated with variants, genomic coordinates, transcriptional (e.g., RNA) coordinates, translated (e.g., protein) coordinates, amino acids, and various others, as will be understood by those skilled in the art) that describe changes associated with changes in genotype (e.g., sequence, molecular variants).

[0058] As will be appreciated by those of skill in the art, a wide variety of sources regarding the phenotypic impact (e.g., labels) of molecular variants can be utilized to define the truth set, including (e.g., public and / or private) clinical and non-clinical variant databases (e.g., ClinVar, HumVar, VariBench, SwissVar, PhenCode, PharmGKB, or locus-specific databases), and outcome databases.

[0059] In some other embodiments, the present disclosure provides functional models (m) that relate molecular, phenotypic, or population signals (e.g., features) derived from one or more molecular measurements, molecular treatments, molecular features, and / or molecular scores to phenotypic effects (e.g., labels) of molecular variants calculated directly from distinct molecular, phenotypic, or population signals via regression and classification techniques. F The present invention provides systems and methods for deriving functional scores and functional classifications via statistical (e.g., machine) learning to generate a set of functional scores and classifications. In some embodiments, this approach may enable the derivation of functional scores and classifications that predict, for example, the comparative mutation burden, mutation rate, or mutation signature of samples from subjects containing specific molecular variants. In some embodiments, the functional scores or classifications from such measurements may enable the notification of a test subject's lifetime risk of developing cancer.

[0060] As will be appreciated by those skilled in the art, the functional model (m F Regression and classification to generate the 's may rely on a variety of statistical (e.g., machine) learning techniques for semi-supervised or supervised learning, including, but not limited to, random forests (RF), gradient boosting trees (GBT), zero rule (ZR), naive Bayes (NB), naive logistic regression (LR), support vector machines (SVM), k-nearest neighbors (kNN), and approaches deploying a wide variety of artificial neural network (ANN) architectures and techniques. In some embodiments, the present disclosure describes the use of molecular state-specific, molecular signals for the derivation of molecular state-specific functional scores or functional classifications. In some other embodiments, the present disclosure describes the use of multi-state matrices of molecular signals for the derivation of molecular state-detected functional scores or functional classifications. In some embodiments, the present disclosure describes the use of convolutional neural networks (CNNs) to learn patterned associations between functional scores or functional classifications and molecular signals distributed across molecular states.

[0061] Figure 1A shows the application of DML treatment and systems to genes in the RAS / MAPK pathway, according to some embodiments. The RAS / mitogen-activated protein kinase (MAPK) pathway can play a role in cell proliferation, differentiation, survival, and death, and somatic mutations in RAS / MAPK genes can play a role in the development, progression, and therapeutic response of various cancer types through activation and dysregulation of MAPK / ERK signaling. In addition, inherited (e.g., germline) mutations in RAS / MAPK genes have been associated with several autosomal dominant congenital syndromes, including, but not limited to, Nunan syndrome (NS), Costello syndrome (CS), cardio-facio-cutaneous (CFC) syndrome, and Leopard syndrome (LS), which are found in patients with distinctive facial features, cardiac defects, musculocutaneous abnormalities, and mental retardation, as well as abnormalities of the skin, inner ear, and genitalia (Aoki et al. 2008). For example, mutations in the protein tyrosine phosphatase, non-receptor type 11 (PTPN11) and dual specificity mitogen-activated protein kinase kinase 1 / 2 genes (MAP2K1, MAP2K2) have been recurrently observed in Nu-Nan and CFC patients, with as many as 50% of Nu-Nan patients harboring PTPN11 mutations (Aoki et al. 2008).

[0062] Embodiments may utilize wild-type, somatic, and germline molecular variants of key RAS / MAPK pathway components, such as HRAS (e.g., G12V), PTPN11 (e.g., E76K and N308D), and MAP2K2 (e.g., F57C and P128Q), constructed and overexpressed in HEK293 cells. Embodiments may select cells with 1 mg / ml puromycin to ensure expression of exogenously induced functional elements (e.g., genes), and may verify RAS / MAPK pathway activation using enzyme-linked immunosorbent assays (ELISAs) for phospho-ERK protein and total ERK protein levels (see Figure 5). To generate single-cell RNA-seq data, embodiments may utilize a 10X Genomics Chromium system, targeting capture of 500 cells for each molecular variant. Capture and subsequent single-cell library generation may be performed according to the manufacturer's recommendations. The resulting libraries for each functional element (e.g., gene) can be pooled and sequenced on an Illumina MiniSeq sequencer until the average reads per cell for each genotype exceeds 30,000 reads / cell. Single-cell RNA-seq processing (e.g., single-cell quality control, normalization, transcriptome count, etc.) can be performed using the 10X Genomics Cell Ranger 2.1.0 pipeline and default settings.

[0063] 1B and 1C show projections of mammalian cells (e.g., HEK293) containing wild-type and mutant PTPN11 and MAP2K2 for molecular variants associated with germline disorders (F57C, P128Q, and N308D) and somatic disorders (E76K), according to some embodiments. Cells can be projected onto a two-dimensional plane derived by t-distributed stochastic neighbor embedding (tSNE) based on (e.g., lower) molecular scores determined from scaled, normalized unique molecular identifier (UMI) counts of single-cell gene expression. For each gene, tSNE projections are shown based on higher molecular scores derived through the application of a wide range of generalized algorithmic criteria (e.g., principal component analysis, PCA) and custom-developed solutions, including cell-type, gene-, or pathway-specific autoencoders (AEs) trained for robust, compressed representation of lower molecular scores. In some embodiments, the autoencoder may be constructed as a neural network with a symmetric number of neurons around the hidden layers (e.g., across layers) and fully connected layers with rectified linear units (ReLu) for activation. In some embodiments, the autoencoder may be trained and optimized for a mean squared error (MSE) loss function using the Adam optimizer.

[0064] As shown in Figures 1B and 1C, compared to generalized dimensionality reduction algorithms, cell projections from a customized, cell-type and pathway-specific autoencoder (AE) can improve hyperdimensional separation between model systems (e.g., cells) containing neutral (e.g., wild-type) and disease-associated molecular variants (e.g., N308D, E76K). The denoising autoencoder (AE) was trained on 8.3 million lower molecular scores from over 18,800 genes detected in 3,495 single HEK293 cells containing wild-type and mutant versions of RAS / MAPK genes. Training was performed for 30 epochs with a mini-batch size of 10, using noise simulation followed by a randomized 5% reduction in the sampling of UMI counts between epochs. The architecture of the fully connected, symmetric autoencoder utilized is shown in Figure 4. While traditional approaches in the area of scaling, normalization, and dimensionality reduction of lower molecular scores may fail to separate tSNE projections of cells containing Nunan syndrome (NS; N308D) molecular variants and wild-type PTPN11, customized cell-type and pathway-specific autoencoders can show robust separation of cells containing somatic (E76K) and germline (N308D) disorder molecular variants in PTPN11 from wild-type cells.

[0065] According to some embodiments, Figures 14A and 14B show the performance of systems and methods for binary classification of molecular variants with two distinct phenotypic effects, as determined in mammalian cells containing either a disease-associated (e.g., pathogenic) genotypic (e.g., sequence) variant (e.g., G12V) and a wild-type (e.g., benign) genotypic (e.g., sequence) version of the human HRAS gene, or a third member of the RAS / MAPK pathway encoding the oncoprotein h-Ras (also known as the transforming protein p21). Small GTPases, small G proteins in the Ras subfamily of the Ras superfamily, h-Ras, once bound to guanosine triphosphate, can activate RAF-family kinases (e.g., c-Raf), thereby leading to cellular activation of the MAPK / ERK pathway.

[0066] Figure 14A shows a projection 1402 of wild-type and mutant mammalian cells (HEK293) onto a two-dimensional plane, derived by cellular t-distributed stochastic neighbor embedding (tSNE) based on normalized, single-cell gene expression measurements of the cells. As shown in Figure 14A, on average, ~3,500 molecular measurements are made per cell, and lower-level molecular scores can be derived from the molecular measurements of over 33,500 genes. Principal component analysis (PCA) can be applied to derive higher-level molecular scores that reduce the dimensionality of the lower-level molecular scores. To assign the projected cells to molecular states 1404, a Gaussian mixture model (GMM) can be applied to define, for example, subpopulations of N = 6 cells based on their lower-level molecular scores derived from their normalized, single-cell gene expression measurements (e.g., UMI counts). Mutant and wild-type cells can be grouped, for example, by k P =15 disease-related and k B Pseudo disease-associated and benign genotypes can be generated by randomly assigning each of the 15 genotypes to a benign pseudo population. A machine learning functional model (m) that is capable of distinguishing between disease-associated and benign genotypes is then generated. F ) to train and test the pseudo-population (k P 1-15, k B1-15) are divided into training and testing sets, for example, applying an 80 / 20 cross-validation scheme, and k sets of each class label (e.g., disease-related and benign), collectively referred to as the truth set. TRAIN = 12 training and k TEST This procedure can be repeated, for example, with i=25 replicates in each of f=5 folds, and within each fold, a pseudopopulation (e.g., k P 1-15, k B 1-15) can be sampled by recovery to retain, for example, 20%, 40%, 60%, 80%, or 100% of the cells. At each iteration, folding, and sampling, lower molecular signals and higher molecular signals for disease-associated and benign genotypes can be calculated as the average of the lower molecular scores and higher scores, respectively. At each iteration, folding, and sampling, population signals for disease-associated and benign genotypes can be determined, for example, as the fraction of cells corresponding to each of the N=6 subpopulations. At each iteration, folding, and sampling, a machine learning functional model (m F ) is k TRAIN This functional model (m) can distinguish disease-associated and benign genotypes from the truth set based on lower-order molecular signals, higher-order molecular signals, or population signals observed in the data. F ) can be trained using a 10x cross-validation strategy and a random forest estimator to distinguish between variants. At each iteration, fold, and sample, the trained functional model (m F ) is k TEST The class label (e.g., disease-related or benign) of the pseudo-population can be predicted based on their lower molecular signals, higher molecular signals, or population signals. As shown in Figure 14B, this approach can result in a robust distinction between disease-related and benign genotypes based on the lower molecular signals, higher molecular signals, and population signals determined within the population of mutant and wild-type cells.

[0067] To evaluate the performance of the DML processing and system as a scalable solution for accurate identification of disease-associated (e.g., pathogenic) molecular variants across multiple genes and disorders, a uniform, distributed DML processing pipeline can be deployed for pre-processing, scaling, normalization, dimensionality reduction, and computation of molecular and population signals on, for example, three genes of the RAS / MAPK pathway, HRAS, PTPN11, and MAP2K2. Applying a similar training / testing scheme for assessment of classification accuracy as described above, the DML process can achieve (e.g., median) raw data classification accuracies 202 of ∼99.9% and ∼100% in the analysis of somatic cancer-causing molecular variants in HRAS (e.g., G12V) and PTPN11 (e.g., E76K), respectively, and ∼98.5% and ∼96.1% (e.g., median) raw data classification accuracies 204 of ∼98.5% and ∼96.1% in the analysis of molecular variant germline (e.g., inherited) disorders in PTPN11 (e.g., N308D) and MAP2K2 (e.g., F57C, P128Q), respectively, as shown in Figure 2A. As shown in Figure 2B, the mean accuracies (e.g., Matthew's correlation coefficient, MCC) in classifying molecular variants known to cause somatic disorders in HRAS, somatic disorders in PTPN11, germline disorders in PTPN11, and germline disorders in MAP2K2 can be ∼99.4%, ∼100%, ∼95.2%, and ∼90.1%, respectively. The raw data classification accuracy (e.g., ACC) and mean classification accuracy (e.g., MCC) in analyzing disease-associated (e.g., somatic and germline, combined) molecular variants can be ∼98.4% and ∼95.6%, respectively, based on the molecular and population signals described herein.

[0068] In some embodiments, the present disclosure provides systems and methods for the derivation of model system-level (e.g., cellular-level) phenotypic scores through the application of statistical machine learning models to relate lower and higher molecular scores to the known phenotypic effects of variants contained within the model system (e.g., cells). Figures 3A and 3B show the classification accuracy of raw cellular-level data for machine learning models trained to derive phenotypic scores in cells containing wild-type and mutant versions of MAP2K2, according to some embodiments.

[0069] In Figure 3A, the germline and extension bars may indicate the average classification accuracy of test cells containing MAP2K2 germline disorder molecular variants that were excluded from training based on the cell phenotype score, but training was based only on included data from MAP2K2 neural germline disorder molecular variants (e.g., germline 302) or PTPN11 germline disorder molecular variants (e.g., extension 304). In Figure 3B, the germline 302 and extension 304 bars indicate the average classification accuracy of test MAP2K2 germline disorder molecular variants that were excluded from training, as determined based on the main cell phenotype score for populations of cells with various numbers of cells. As in Figure 3A, the germline and extension bars may correspond to the raw data accuracy in classifying the test molecular variants, but training was based only on included data from MAP2K2 neural and germline disorder molecular variants (e.g., germline) or PTPN11 germline disorder molecular variants (e.g., extension).

[0070] Figures 3A and 3B show data obtained using a logistic regression (LR) classifier trained for binary classification of cells containing disease-associated molecular variants and cells containing wild-type MAP2K2 based on calculation of higher-order molecular scores as the top 100 principal components from lower-order molecular scores (e.g., scaled and / or normalized). Partitioning of molecular variants into training and testing bins and corresponding training and testing sets of cells on molecular variant genotypes can be generated for training and testing, such that certain populations of cells with specific disease-associated molecular variants are excluded from training. In this way, classification performance can be calculated on the complete population of cells containing variants excluded from training. As shown in Figures 3A and 3B, the average cell-by-cell classification accuracy across molecular variants associated with germline (e.g., inherited) disorders in MAP2K2 can be ∼80.3%.

[0071] In some embodiments, the present disclosure describes learning and predicting the phenotypic consequences of molecular variants based on molecular, phenotypic, or population signals measured in multiple genes or molecular entities within the same, related, or interacting pathway. As shown in Figures 3A and 3B, the inclusion of data from PTPN11 molecular variants associated with germline (e.g., inherited) disorders improves the average cell-by-cell classification accuracy across germline disorder molecular variants in MAP2K2 from ~80.3% (e.g., germline 302) to ~92.8% (e.g., extension 304), thereby demonstrating the ability of the disclosed DML processes and systems to identify and exploit coherent cellular properties for accurate classification of the phenotypic impact of molecular variants across multiple functional entities. As shown in Figures 3A and 3B, improved performance in cell-by-cell classification can result in increased classification of molecular variants based on majority-type classification from a population of cells containing the molecular variant.

[0072] In some embodiments, the present disclosure provides systems and methods for deriving functional scores and functional classifications for individual functional elements (e.g., individual genes). In some embodiments, the present disclosure provides methods for deriving functional scores and functional classifications across multiple functional elements that leverage coordinated molecular signals across molecular variants within multiple functional elements. In some embodiments, the present disclosure describes systems and methods that combine the use of mutagenists, molecular barcoding, molecular cloning, and cell pooling techniques to generate populations of cells in which molecular variants in different functional elements are uniquely generated, barcoded, or both.

[0073] In some embodiments, independent or disjoint estimates of molecular, phenotypic, or population signals (e.g., features) may be utilized to derive independent or disjoint functional scores and functional classifications via statistical (e.g., machine) learning to relate molecular signals (e.g., features) to the phenotypic effects (e.g., labels) of molecular variants via regression and classification techniques, respectively.

[0074] In some embodiments, as will be appreciated by those skilled in the art, feature weights from statistical (e.g., machine) learning models generated using independent or disjoint estimates of each molecular, phenotypic, or population signal are calculated, collected, and utilized for robust feature selection using techniques. In some embodiments, the present disclosure provides methods for deriving functional scores and functional classifications via statistical (e.g., machine) learning to relate robust molecular, phenotypic, or population signals (e.g., robust features) identified via regression and classification techniques, respectively, to the phenotypic impact (e.g., labels) of molecular variants.

[0075] In some embodiments, the present disclosure describes systems and methods for deriving feature scores and feature classifications from multiple statistical (e.g., machine) learning models generated using independent or disjoint estimates of molecular signals and applying either model selection or model combination (e.g., blending) techniques (Pan et al. 2006).

[0076] In some embodiments, model selection techniques may be applied to compare models using model selection criteria that measure the predictive performance or probability of a model being the true model, and selection may be applied to maximize the estimate of the selection criteria. As will be appreciated by those skilled in the art, various model selection criteria may be applied, including (but not limited to) Akaike Information Criterion (AIC), Bayesian Information Criterion (BIC), cross-validation (CV), bootstrap (Efron 1983; Efron 1986; Efron and Tibshirani 1997), or adaptive model selection criteria (George and Foster 2000; Shen and Ye 2002; Shen et al. 2004), calculated on training data or input test data, as exemplified by the test input dependent weight (IDW). The IDW for a candidate model may be defined as the probability of the model providing an accurate prediction for a given input or a reasonable measure to quantify the predictive performance of the model with respect to the input test data (Pan et al. 2006).

[0077] In some other embodiments, model combination techniques can be applied to generate a combined model by applying ensemble methods, taking an equal or unequal weighted average of the outputs from individual models (Ripley 2008; Hastie et al. 2001). For example, ensemble methods include Bayesian model averaging, stacking, bagging, random forests, boosting, ARM, and other methods computed on training data (Burnham and Anderson 2003; Hastie et al. 2001). Examples of methods for deriving feature scores and classifications include, but are not limited to, the use of performance metrics (e.g., AIC and BIC) as weights calculated on input test data (Pan et al. 2006) or calculated on input test data (Pan et al. 2001). In some other embodiments, model combination techniques may be applied to generate combined models using artificial neural network (ANN) architectures. In some embodiments, the present disclosure describes systems and methods for deriving feature scores and classifications from multiple statistical (e.g., machine) learning models generated using independent or disjoint estimates of molecular signals, involving the application of various noise control techniques (e.g., bootstrap ensemble with noise algorithms (Yuval Raviv 1996)).

[0078] In some embodiments, the present disclosure provides inferential models (m) that model the relationship between (e.g., measurement endpoint) functional scores or functional classifications and multiple dependent (e.g., measured) features (e.g., molecular, phenotypic, or population signals) or independent (e.g., non-measured) features (e.g., evolutionary, population, functional (e.g., annotation-based), structural, dynamic, and physicochemical features associated with variants, genomic coordinates, transcriptional (e.g., RNA) coordinates, translated (e.g., protein) coordinates, amino acids, and various others, as will be understood by those skilled in the art). I Systems and methods are described for estimating functional scores and functional classifications for molecular variants that apply statistical (e.g., machine) learning techniques to generate inference models (m). As will be appreciated by those skilled in the art, such inference models (m I) may enable the estimation of functional scores and functional classifications for molecular variants, with or without explicit use of molecular, phenotypic, or population signals, molecular measurements, molecular treatments, molecular features, or molecular scores. In some embodiments, such methods may enable the inference of sequence-function maps that describe functional scores and functional classifications for molecular variants other than those for which the functional scores and functional classifications were directly measured. In some embodiments, as shown in FIG. 15 , such systems and methods may enable the inference of sequence-function maps 1514 that describe functional scores or functional classifications for all potential nonsynonymous variants in a protein-coding gene, utilizing functional scores and functional classifications from sequence-function maps 1502 that represent a subset of potential nonsynonymous variants. In some embodiments, this inference may utilize a score regression layer 1504 that accesses as input an annotation matrix 1506 consisting of annotation features 1508, labels 1510, and functional scores 1512. As will be appreciated by those skilled in the art, numerous statistical validation and cross-validation techniques may be applied to monitor or ensure the accuracy of the estimated functional scores and functional classifications.

[0079] In some embodiments, and as shown in FIG. 16 , the present disclosure describes systems and methods for determining the phenotypic impact (e.g., pathogenic, functional, or comparative effect) of molecular variants through a series of modeling layers that (a) collect or generate existing knowledge or reliable predictions of the phenotypic impact of molecular variants, (b) grow the set of molecular variants with known or predicted phenotypic impact through functional modeling (e.g., via a functional modeling engine (FME)) of sampled molecular variants of known, highly-confidently predicted, or unknown phenotypic impact, and (c) further complete the set of molecular variants with known or predicted phenotypic impact through inferential modeling. These layers combine to produce a functional model (m F ) 1607 widens (or optimizes) the range of truth sets available for generation, and inferential models (m I)1609 Functional model for generation (m F ) 1607 may reduce (or optimize) the required scope of generation support. In some embodiments, these systems and methods may overcome limitations on training, validation, and testing for functional elements (e.g., genes) and contexts where molecular variants of known phenotypic impact (e.g., pathogenicity, functionality, or comparative effect) are limited. Such systems and methods may thus enable elucidation of the phenotypic impact of molecular variants for otherwise limited functional elements (e.g., genes) with data for model generation, reducing overall costs.

[0080] In some embodiments, and as shown in FIG. 16, such systems and methods may accomplish this by combining one or more of the following layers of modeling: (1) predictive models (m P ) 1603, (2) Sampling model (m S ) 1605, (3) Functional model (m F ) 1607, and (4) inference model (m I ) 1609. In some embodiments, the present disclosure describes systems and methods for accessing molecular variants with known phenotypic impact (e.g., pathogenic or benign) from existing sources to populate a sequence-function map 1602 describing the phenotypic impact of molecular variants in genes / functional elements. In some embodiments, well-characterized predictive models (m) are used to combine the phenotypic impact of molecular variants with high-confidence predictions to generate an extended sequence-function map 1604. P ) 1603 may be utilized. In some embodiments, a sampling model (m) is used to generate a set of genotypes (e.g., molecular variants) 1606 that includes (a) a truth set by selecting or subsampling molecular variants with known or highly confidently predicted phenotypic impact, and (b) a target set of molecular variants of unknown phenotypic impact. S )1605 is applied.

[0081] In some embodiments, the present disclosure provides a functional model (m) that associates molecular, phenotypic, or population signals and functional scores and classifications as learned from molecular variants in a truth set (e.g., from genotypes 1606) and predicts functional scores and classifications of functional molecular variants in a target set (e.g., from genotypes 1606), thereby producing a sequence-function map of functional scores 1608. F ) 1607 is described.

[0082] In some embodiments, as shown in FIG. 16, a functional model (m F ) 1607 accesses extended truth sets 1611 and 1612 that include molecular and population signals from multiple functional elements (e.g., genes) in the same, related, or interacting pathways. This capability may enable the system to generate functional models (mF) 1607 for functional elements (e.g., genes) with limited or missing availability of molecular variants with known or high-confidence predicted phenotypic impact, based on molecular, phenotypic, or population signals from functional elements (e.g., genes) with coherent mechanisms of action. Figures 3A and 3B show an example of this.

[0083] In some embodiments, the phenotypic impact of known molecular variants, high-confidence predicted molecular variants, and functionally modeled molecular variants can be leveraged by an inferential model (mI) 1609 that models the relationship between phenotypic impact and multiple dependent (e.g., measured) features (e.g., molecular, phenotypic, or population signals) or independent (e.g., non-measured) features (e.g., evolutionary, population, functional (e.g., annotation-based), structural, dynamic, and physicochemical features associated with variants, genomic coordinates, transcriptional (e.g., RNA) coordinates, translated (e.g., protein) coordinates, amino acids, and various others, as will be understood by those skilled in the art) to produce an augmented sequence function of functional score 1610. As will be understood by those skilled in the art, such an inferential model (mI) 1609 can be used to generate an augmented sequence function of functional score 1610. I) 1609 may allow estimation of the phenotypic effects of molecular variants, with or without explicit use of molecular, phenotypic, or population signals.

[0084] In some embodiments, the present disclosure describes systems and methods for cost-effective optimization of molecular variant classification through a staged deployment of deep mutation learning (DML) processes and systems on truth and target (query) sets of molecular variants. Some embodiments include a Stage I optimization 610 step, for example, as shown in FIG. 6, in which an autoencoder (m AE ) and other dimension reduction models (m DR )614 and functional model (m F To generate high-quality data for optimization 616, model systems (e.g., cells) containing truth set variants are measured at high model system (e.g., cell) numbers and read depths in cell number and read depth optimization 612. In this first phase, dimensionality reduction and classification accuracy for the phenotypic impact of target molecular variants may be optimized to identify combinations of dimensionality reduction model (614), functional model (616), and cell number and read depth (612) that ensure robust target performance. In some embodiments, subsampling and noise simulation may be utilized to train and model the performance of the dimensionality reduction model and functional model. As shown in FIG. 6 , some embodiments include a Phase II fabrication 620 step in which model systems (e.g., cells) containing target set variants and, optionally, truth set variants may be measured in deployments where a particular dimensionality reduction model 624 and functional model 626 are deployed at (e.g., optimal or minimal) cell number and / or read depth 622 is identified as robust.

[0085] In some embodiments, the present disclosure describes systems and methods for determining the phenotypic impact (e.g., pathogenicity, functionality, or comparative effect) of molecular variants identified in a subject's biological sample or record based on the functional scores and functional classifications determined as described above. In some embodiments, a time-stamped record of functional score and functional classification combinations for a set of (e.g., multiple unique) molecular variants may be generated, evaluated, validated, selected, and applied to determine the phenotypic impact of molecular variants identified in a subject's biological sample or record.

[0086] In some embodiments, the present disclosure describes systems and methods for determining the phenotypic impact (e.g., pathogenic, functional, or comparative effect) of molecular variants identified within a subject's biological sample or record based on predictor scores or predictor classifications from calculated predictors generated by applying statistical (e.g., machine) learning methods to utilize functional scores and classifications.

[0087] 17, the present disclosure describes a method for generating (e.g., lower-level) variant interpretation engines (VIEs), which may be gene- and condition-specific, through statistical (e.g., machine) learning techniques that model the phenotypic impact 1712 of molecular variants based on input labels 1714 and annotation matrices 1706, including their functional scores 1702, 1708 (or functional classifications) and other annotation features 1710, including features commonly used in generating computed predictors, including, but not limited to, evolutionary, population, functional (e.g., annotation-based), structural, dynamic, and physicochemical features associated with variants and residuals of functional elements. In some embodiments, a training and validation layer 1704 may utilize cross-validation techniques 1716 (e.g., K-fold or LOOCV) to train and quality control the VIE, which is subsequently evaluated by an examination layer 1718 to derive predictor scores 1720 used in molecular variant classification.

[0088] In some embodiments, the present disclosure further describes systems and methods for generating pathway and condition-specific (top-level) variant interpretation engines (VIEs) that apply model combination techniques that integrate (lower-level) gene and condition-specific variant interpretation engines (VIEs) from multiple genes in a target pathway of interest. In other embodiments, the present disclosure further describes systems and methods for generating pathway and condition-specific (top-level) variant interpretation engines (VIEs) through statistical (e.g., machine) learning techniques that model the phenotypic impact of molecular variants based on their functional scores, functional classifications, and other features commonly utilized in generating computational predictors, including, but not limited to, evolutionary, population, functional (annotation-based), structural, dynamic, and physicochemical features associated with variants and residuals of functional elements.

[0089] In some embodiments, the present disclosure describes systems and methods for determining the phenotypic impact (e.g., pathogenicity, functionality, or comparative effect) of molecular variants identified within a subject's biological sample or record thereof based on hotspot scores and hotspot classifications from mutational hotspots calculated by applying spatial clustering techniques described herein associated with the molecular variants and residuals and to identify networks of residuals with specific phenotypic impacts that utilize validated functional scores, functional classifications, and molecular signals.

[0090] In some embodiments, the present disclosure describes a system and method for deriving a matrix of functional distances between molecular variants or their corresponding residuals by calculating distance metrics between molecular variants projected in an N-dimensional space (1≦N≦M) defined by a set of functional scores, functional classifications, and molecular signals (as described above), where N<M when dimensionality reduction techniques are applied to reduce the feature space of the molecular variants. As will be understood by those skilled in the art, various dimensionality reduction techniques may be applied, including but not limited to techniques that rely on linear transformations such as in principal component analysis (PCA) or non-linear transformations such as in manifold learning techniques (e.g., t-distributed stochastic neighbor embedding (tSNE) and kernel principal component analysis (kPCA)). As will be understood by those skilled in the art, various distance metrics may be utilized, including but not limited to Euclidean distance, Manhattan distance (e.g., city block), Mahalanobis distance, or Chebyshev distance, and various others.

[0091] In some embodiments, the present disclosure describes a system and method for identifying significantly mutated regions (SMRs) and networks (SMNs) by measuring and scoring the phenotypic-related mutation density (e.g., the number of observed phenotypic-related variants per residual) within spatially proximal residuals of functional elements (e.g., protein-coding genes) through the application of spatial clustering techniques across multiple spatial distance metrics, including but not limited to the effective functional distances, sequence distances, structural distances, (co-)evolutionary distances, and combinations thereof described herein.

[0092] In some embodiments, also as shown in FIG. 18, the identification of SMR / SMN may apply a training / validation layer 1804 for identifying spatial clustering between phenotypic-related or functionally related molecular variants 1806 such that the determination is based on the commonality in the functional scores of the molecular variants. In some embodiments, these commonalities may be identified from the functional scores of the molecular variants in the sequence-function map of the protein-coding gene 1802.

[0093] In some embodiments, and as shown in FIG. 18 , identification of SMRs / SMNs in the training / validation layer 1804 may involve a series of steps, including but not limited to: (1) SMR / SMN detection techniques 1805 for identification of single residuals or networks of residuals enriched in molecular variants with specific phenotypic relevance, as previously described (Araya et al. 2016, US Patent Application 20160378915A1), and (2) SMR / SMN selection techniques 1815.

[0094] The SMR / SMN detection technique 1805 may involve a series of steps including, but not limited to: (1.1) projection 1810 of phenotype-associated molecular variants 1806 in functional, sequence, structural, or (co)evolutionary dimensions (or a combination thereof); (1.2) application of spatial clustering techniques (e.g., DBSCAN) 1812 to detect clusters of spatially close phenotype-associated variants; and (1.3) measurement of mutation density, scoring the number of phenotype-associated variants per residual in the cluster.

[0095] The SMN detection technique 1805 may further include steps represented at 1814, including, but not limited to: (1.4) scoring mutation density probabilities, for example, by calculating the (e.g., binomial) probability of obtaining k or more (e.g., greater than or equal to k) observed phenotype-associated variants per cluster given the residual-wise mutation rate within each functional element (e.g., protein-coding gene); (1.5) applying a multiple hypothesis correction (MHC) across the mutation density probabilities of the discovered clusters; and (1.6) calculating a false discovery rate (FDR) for the observed (e.g., raw or corrected) mutation density probabilities using a background model of mutation density probabilities derived by randomizing the location of the observed phenotype-associated variants within each functional element.

[0096] The training / validation layer 1804 may further perform an SMR / SMN selection technique 1815. The SMR / SMN selection technique may include the following steps: (2.1) defining the (e.g., raw or corrected) mutation density probability and / or false discovery rate (FDR) as a hotspot score and applying a cutoff to statistically define hotspot classifications, thereby specifying residuals in candidate clusters (e.g., sequences 1816, features 1818, and sequences 1820); (2.2) finding residuals in candidate clusters from multiple, different projections / spaces; (2.3) applying assignment heuristics to assign residuals to distinct clusters (e.g., selecting the cluster with the largest size (e.g., the cluster with the largest number of residuals)); and (2.4) identifying the final set of clusters that meet these SMR / SMN criteria. The final set of SMR / SMNs may be derived from multiple, different projections (eg, sequence 1820, function 1818, or sequence, function (combination) 1822).

[0097] In some embodiments, the present disclosure describes systems and methods for identifying SMRs / SMNs by measuring and scoring the density of phenotype-associated mutations (e.g., the number of observed phenotype-associated variants per residual) within spatially adjacent residuals of functional elements (e.g., protein-coding genes) through the application of spatial clustering techniques across multiple spatial distance metrics, where phenotype-associated variants may be defined based on the functional scores and functional classifications described herein. As will be appreciated by those skilled in the art, these methods may enable the determination of clusters of residuals that result in specifically defined phenotypic effects.

[0098] In some embodiments, the present disclosure describes systems and methods for assessing the accuracy, performance, or robustness of independent evidence datasets for the interpretation of molecular variants, such as quantitative (e.g., score) or qualitative (classification) evidence from computational predictors (e.g., M-CAP, REVEL, SIFT, and PolyPhen2), and gene-specific predictors (e.g., PON-P2), mutational hotspots, and population genomics indices (e.g., allele frequency-based variant classification), (Amendola et al. 2016), relative to the functional scores and classifications described herein.

[0099] In some embodiments, the present disclosure describes systems and methods for calculating evaluation metrics to assess the concordance between an evidence dataset and the functional scores and functional classifications described herein, based on which the best evidence dataset is selected for use in variant interpretation and prioritization. As will be understood by those skilled in the art, various evaluation metrics can be used to assess the concordance of an evidence dataset to the functional scores or functional classifications described herein. With regard to quantitative evidence (e.g., scores), these may include Pearson's correlation coefficient, Spearman's rank correlation, Kendall's correlation, and various others, as will be understood by those skilled in the art. With regard to qualitative evidence (e.g., classification), these may include accuracy, Matthew's correlation coefficient, Cohen's kappa coefficient, Youden's index (e.g., comprehension), F-measure (e.g., F1 score), true positive rate (e.g., sensitivity or recall), true negative rate (e.g., specificity), positive predictive value (e.g., precision), negative predictive value, positive likelihood ratio, negative likelihood ratio, and diagnostic odds ratio, as will be understood by those skilled in the art.

[0100] In some embodiments, the present disclosure describes systems and methods that may continuously evaluate, validate, and optimize (e.g., select, remove, or modify) diverse evidence datasets based on the above-mentioned evaluation metrics, and distribute the best (e.g., independent) evidence datasets to client systems via application of program interfaces (APIs) for use in variant interpretation and prioritization practices that determine the phenotypic impact (e.g., pathogenicity, functionality, or comparative effect) of molecular variants identified within a subject's biological sample or record thereof.

[0101] In some embodiments, the present disclosure describes systems and methods for determining the degree of validation bias, reporting bias, or outcome bias present in a variant dataset, including clinical datasets (e.g., ClinVar, HumVar, VariBench, SwissVar, PhenCode, or locus-specific databases), population datasets (e.g., ExAC, GnomAD, and 1000 Genomes), or independent evidence datasets for the interpretation of molecular variants, such as, but not limited to, calculated predictors (e.g., M-CAP, REVEL, SIFT, PolyPhen2, and PON-P2). In some embodiments, the present disclosure describes systems and methods for determining bias based on the expected distribution of functional scores, functional classifications, and molecular signals described herein associated with molecular variants and residuals.

[0102] In some embodiments, the present disclosure describes systems and methods for evaluating a target variant dataset by measuring and scoring the differences between the distribution of functional scores, functional classifications, and molecular signals of molecular variants and residuals in the target dataset relative to the expected distribution of functional scores, functional classifications, and molecular signals of molecular variants from a reference dataset. In some embodiments, measuring the inherent bias in the target variant dataset may include a series of steps, including, but not limited to: (1) collecting functional scores, functional classifications, and molecular signals associated with molecular variants in the target and reference datasets; (2) estimating probability density functions for functional scores, functional classifications, or molecular signals associated with molecular variants in the reference dataset; (3) estimating probability density functions for functional scores, functional classifications, or molecular signals associated with molecular variants in the target dataset; and (4) measuring the statistical distance between the probability density functions derived from the target dataset and the reference dataset for functional scores, functional classifications, or molecular signals. In some embodiments, measuring the inherent bias in the target variant dataset includes a series of steps, including: (5) sampling variants from a reference dataset (e.g., to match the sample population size of the target dataset); (6) estimating a probability density function for the functional scores, functional classes, or molecular signals of the reference dataset sampled in step 5; (7) measuring the statistical distance between the probability density function for the functional scores, functional classes, or molecular signals derived from the target dataset and the probability density function derived from the sampled reference dataset; and (8) repeating steps 5-8 to obtain a robust estimate and confidence interval of the statistical distance between the probability density functions for the functional scores, functional classes, or molecular signals of the target and reference datasets. In some embodiments, the above-described systems and methods for bias detection and statistical assessment enable the identification of clinical, population, or evidence datasets in which variants contained therein have functional scores, functional classes, or molecular signals that differ from those expected in the reference dataset.

[0103] In some other embodiments, the present disclosure describes systems and methods for assessing inherent bias in evidence datasets by a series of steps including, but not limited to: (1) partitioning the evidence and reference datasets into quantile-matching sets (e.g., for quantitative evidence scores) or classes (e.g., qualitative evidence classification), (2) scoring variants within each set (e.g., evidence vs. reference) across multiple properties (e.g., evolutionary, population, functional (e.g., annotation-based), structural, dynamic, and physicochemical features associated with the variants), (3) estimating a probability density function for each property score within each set (e.g., evidence vs. reference), (4) measuring the statistical distance between the evidence set-derived probability density function and the reference set-derived probability density function for each property score, and (5) identifying properties with statistically significant differences in scores between the reference and evidence sets.

[0104] In some embodiments, the present disclosure describes systems and methods that continuously evaluate and select diverse evidence datasets based on the above-mentioned bias indicators, and may distribute the most unbiased (e.g., independent) evidence datasets to client systems through the application of program interfaces (APIs) for use in variant interpretation and prioritization practices that determine the phenotypic impact (e.g., pathogenicity, functionality, or comparative effect) of molecular variants identified within a subject's biological sample or record.

[0105] In some embodiments, the present disclosure describes systems and methods for determining the phenotypic impact (e.g., pathogenic, functional, or comparative effect) of molecular variants identified within a subject's biological sample or record based on the functional scores, functional classifications, predictor scores, predictor classifications, hotspot scores, and hotspot classifications described herein in functional elements (e.g., genes) and pathways associated with Mendelian diseases (e.g., Table 1), known cancer drivers (e.g., Table 2), pharmacogenomic genes in which genotypic (e.g., sequence) variations are associated with altered drug response (Table 3), or other clinically valuable genes (e.g., Table 4).

[0106] In some embodiments, the present disclosure describes systems and methods for evaluating, selecting, distributing, and utilizing the best and most unbiased independent evidence, based on the functional scores and classifications described herein, for interpreting and prioritizing functional elements (e.g., genes) and pathways associated with Mendelian diseases (e.g., Table 1), known cancer drivers (e.g., Table 2), variants in pharmacogenomic genes where genotypic (e.g., sequence) variations are associated with altered drug response (e.g., Table 3), or other clinically valuable genes (e.g., Table 4).

[0107] As noted above, Table 1 is an exemplary table of functional elements and pathways associated with Mendelian diseases, according to some embodiments. Table 2 is an exemplary table of functional elements and pathways that are known cancer drivers, according to some embodiments. Table 3 is an exemplary table of pharmacogenomic genes whose genotypic (e.g., sequence) variations are associated with altered drug response, according to some embodiments. Table 4 is an exemplary table of other clinically valuable genes, according to some embodiments. Tables 1-4 may be found on page 47 herein.

[0108] In some embodiments, the present disclosure describes systems and methods for determining the phenotypic impact (e.g., pathogenicity, functionality, or comparative effect) of molecular variants identified within a subject's biological sample or record based on validated functional scores, functional classifications, predictor scores, and predictor classifications described herein for variants within known targets of pathogenic variation, including (but not limited to) mutational hotspots, or for variants within such hotspots, e.g., 50, 100, 500, and 1,000 base pairs (bp). In some embodiments, the present disclosure describes systems and methods for determining the phenotypic impact (e.g., pathogenicity, functionality, or comparative effect) of molecular variants identified within a subject's biological sample or record based on functional scores, functional classifications, predictor scores, and predictor classifications for variants within regions of constrained variation in a population, or for variants within such regions, e.g., 50, 100, 500, and 1,000 bp. As will be appreciated by those skilled in the art, a variety of methods for determining mutational hotspots and regions of constrained variation can be applied.

[0109] Various embodiments may be implemented using one or more computer systems, such as, for example, computer system 1900 shown in Figure 19. For example, computer system 1900 may be used to perform the methods of Figures 1A, 6-13, and 15-18. Computer system 1900 may be any computer capable of performing the functions described herein.

[0110] The computer system 1900 can be any well-known computer capable of performing the functions described herein.

[0111] Computer system 1900 includes one or more processors (also referred to as central processing units, or CPUs), such as processor 1904. Processor 1904 is connected to a communication infrastructure or bus 1906.

[0112] Each of the one or more processors 1904 may be a graphics processing unit (GPU). In one embodiment, a GPU is a processor that is a special-purpose electronic circuit designed to process mathematically intensive applications. A GPU may have a parallel structure that is useful for parallel processing of large blocks of data, such as mathematically intensive data common in computer graphics applications, images, video, etc.

[0113] The computer system 1900 also includes user input / output device(s) 1903 such as a monitor, keyboard, pointing device, etc. that communicate with a communications infrastructure 1906 through user input / output interface(s) 1902 .

[0114] The computer system 1900 also includes a main or primary memory 1908, such as random access memory (RAM). The main memory 1908 may include one or more levels of cache. The main memory 1908 stores control logic (e.g., computer software) and / or data.

[0115] Computer system 1900 may also include one or more secondary storage devices or memory 1910. Secondary memory 1910 may include, for example, a local, network, or cloud-accessible hard disk drive 1912 and / or a removable storage or drive 1914. Removable storage drive 1914 may be a floppy disk drive, a magnetic tape drive, a compact disk drive, optical storage, a tape backup device, and / or any other storage device / drive.

[0116] The removable storage drive 1914 may interact with a removable storage unit 1918 . The removable storage unit 1918 includes computer-usable or readable storage for storing computer software (control logic) and / or data. The removable storage unit 1918 may be a floppy disk, magnetic tape, compact disk, DVD, optical storage disk, and / or any other computer data storage. The removable storage drive 1914 reads from and / or writes to the removable storage unit 1918 in well-known fashion.

[0117] Secondary memory 1910, according to an exemplary embodiment, may include other means, techniques, or approaches that allow computer programs and / or other instructions and / or data to be accessed by computer system 1900. Such means, techniques, or approaches may include, for example, removable storage unit 1922 and interface 1920. Examples of removable storage unit 1922 and interface 1920 may include a program cartridge and cartridge interface (such as found in a video game device), a removable memory chip (such as an EPROM or PROM) and associated socket, a memory stick and USB port, a memory card and associated memory card slot, and / or any other removable storage unit and associated interface.

[0118] Computer system 1900 may further include a communications or network interface 1924. Communications interface 1924 enables computer system 1900 to communicate and interact with any combination of remote devices, remote networks, remote entities, etc. (individually and collectively referred to by reference numeral 1928). For example, communications interface 1924 may enable computer system 1900 to communicate with remote devices 1928 over communications path 1926, which may be wired and / or wireless and may include any combination of the Internet, etc., and may include any combination of a LAN, a WAN, the Internet, etc. Control logic and / or data may be transmitted to and from computer system 1900 over communications path 1926.

[0119] In one embodiment, a tangible apparatus or article of manufacture that includes a tangible computer-usable or readable medium that stores control logic (software) is also referred to herein as a computer program product or program storage device. This includes, but is not limited to, computer system 1900, main memory 1908, secondary memory 1910, and removable storage units 1918 and 1922, as well as tangible articles of manufacture that embody any combination of the above. Such control logic, when executed by one or more data processing devices (such as computer system 1900), causes such data processing devices to perform the operations described above.

[0120] Based on the teachings contained herein, it will be apparent to one skilled in the art(s) how to make and use embodiments of the present disclosure using data processing devices, computer systems and / or computer architectures other than those shown in Figure 12. In particular, embodiments may operate with software, hardware, and / or operating system implementations other than those described herein.

[0121] It will be understood that the Detailed Description section is intended to be utilized for interpreting the claims, and that none of the other sections are. The other sections may describe one or more, but not all, example embodiments as contemplated by the inventor(s), but are not intended to limit the disclosure or the appended claims in any way.

[0122] While this disclosure describes exemplary embodiments for exemplary fields and applications, it should be understood that the disclosure is not limited thereto. Other embodiments and modifications thereof are possible and fall within the scope and spirit of the present disclosure. For example, without limiting the generality of this paragraph, embodiments are not limited to the software, hardware, firmware, and / or entities shown in the drawings and / or described herein. Moreover, embodiments (whether or not explicitly described herein) have significant utility in fields and applications beyond the examples described herein.

[0123] In this specification, embodiments are described by functional building blocks that illustrate the implementation of certain functions and relationships thereof. The boundaries of these functional building blocks are arbitrarily defined in this specification for the convenience of description. Alternative boundaries may be defined so long as the specified functions and relationships (or equivalents) are appropriately performed. Alternative embodiments may also implement the functional blocks, steps, operations, and methods using an order different from that described in this specification.

[0124] References herein to “one embodiment,” “one embodiment,” “an exemplary embodiment,” or similar phrases indicate that an embodiment described herein may include a particular feature, structure, or characteristic, but not all embodiments necessarily include the particular feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in connection with one embodiment, it is within the knowledge of one skilled in the art(s) to incorporate such feature, structure, or characteristic into other embodiments, whether or not explicitly described herein. Furthermore, the terms “coupled” and “connected,” along with their derivatives, may be used to describe some embodiments. These terms are not necessarily intended as synonyms for each other. For example, some embodiments may be described using the terms “coupled” and / or “connected” to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term “coupled” may also mean that two or more elements are not in direct contact with each other, but still cooperate or interact with each other.

[0125] The breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined in accordance with the following claims and their equivalents. [Table 1-1] [Table 1-2] [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4]

Table 2-5

Table 2-6

Table 2-7

Table 2-8

Table 2-9

Table 2-10

Table 2-11

Table 2-12

Table 2-13

Table 2-14

Table 3-1

Table 3-2

Table 3-3

Table 3-4

Table 3-5

Table 3-6

Table 3-7

Table 3-8

Table 3-9

Table 3-10

Table 3-11

Table 3-12

Table 3-13

Table 3-14

Table 3-15

Table 3-16

Table 3-17

Table 3-18

Table 3-19

Table 3-20

Table 3-21

Table 3-22

Table 3-23

Table 3-24

Table 3-25

Table 3-26

Table 3-27

Table 3-28

Table 4-1

Table 4-2

Table 4-3

Table 4-4

Table 4-5

Table 4-6

Table 4-7

Table 4-8

Table 4-9

Table 4-10

Table 4-11

Table 4-12

Table 4-13

Table 4-14

Table 4-15

Table 4-16

Table 4-17

Table 4-18

Table 4-19

Table 4-20

Table 4-21

Table 4-22

Table 4-23

Table 4-24

Table 4-25

Table 4-26

Table 4-27

Table 4-28

Table 4-29

Table 4-30

Table 4-31

Table 4-32

Table 4-33

Table 4-34

Table 4-35

Table 4-36

Table 4-37

Claims

[Claim 1] The invention described in the specification.

Citation Information

Patent Citations

  • Methods and systems for predicting drug resistance and determining the genetic basis of drug resistance using neural networks

    JP2004523725A

  • Variant annotation, analysis and selection tool

    US20130332081A1

  • Methods of predicting pathogenicity of genetic sequence variants

    US20160371431A1