Population frequency modeling for quantitative variant pathogenicity estimation
A logistic regression-based population frequency model addresses the challenge of variant classification by quantifying pathogenicity using allele frequency data, improving accuracy and reducing VUS classification, thereby enhancing genetic testing precision.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2026-03-04
AI Technical Summary
Current genetic testing methodologies struggle to accurately and systematically define the relationship between allele frequency and pathogenicity for a given population and disease, leading to challenges in variant classification, particularly due to the complexity of gene-disease attributes and the limitations of large population databases like gnomAD, resulting in many variants being classified as variants of uncertain significance (VUS).
A population frequency model using logistic regression is developed to quantify the probability of variant pathogenicity based on allele frequency data, excluding gene-disease attributes and utilizing engineered feature combinations to provide a continuous, quantitative scoring approach, which can be integrated into frameworks like Sherlock for improved variant classification.
The model provides highly accurate and scalable pathogenicity estimates, reducing the number of variants classified as VUS by 2.5% compared to existing methods, and enhances the ability to confidently classify variants by incorporating subtle biological differences between genes.
Smart Images

Figure 2026507386000001_ABST
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 421,430, filed November 1, 2022, which is incorporated herein by reference in its entirety.
[0002] The technical field to which this application relates is genetic testing. Another technical field to which this application relates is machine learning-based variant classification systems. [Background technology]
[0003] Genetic variants are differences in DNA sequence between individuals within a population. There are many different types of variants, including structural variations, single nucleotide polymorphisms, insertion and deletion mutations, copy number variations, and translocations and inversions.
[0004] Gene sequencing technologies continue to evolve rapidly. High-throughput sequencing technologies are increasingly enabling genetic testing ranging from genotyping for inherited diseases, single genes, gene panels, exome, genome, transcriptome, and epigenetic assays. The increasing complexity of clinical genetic testing analysis and interpretation, and the increasing volume of testing, are accompanied by new challenges in interpreting sequence variants.
[0005] For example, clinical molecular laboratories are increasingly detecting novel sequence variants in the course of testing patient samples for a rapidly growing number of genes associated with genetic diseases. While some phenotypes are associated with a single gene, many are associated with multiple genes.
[0006] Variant classification refers to the process of classifying genetic variants based on evidence that supports or refutes a causal relationship with disease. The clinical significance of any given sequence variant decreases along a gradient, ranging from the variant being almost certainly pathogenic to almost certainly benign for a given disease.
[0007] Variant classifications are not themselves diagnostic, but can be used by clinicians to make a diagnosis. [Brief explanation of the drawings]
[0008] The present disclosure will be more fully understood from the detailed description given below and from the accompanying drawings of various embodiments of the present disclosure, which are for purposes of illustration and understanding only and should not be construed as limiting the disclosure to the particular embodiments illustrated.
[0009] [Figure 1] 1 illustrates an example of a feature generation process for machine learning-based population frequency modeling, according to some embodiments of the present disclosure.
[0010] [Figure 2] 1 illustrates an example process for constructing a population frequency model using a logistic regression model, according to some embodiments of the present disclosure.
[0011] [Figure 3] 1 illustrates an exemplary use of a population frequency model constructed using a logistic regression model, according to some embodiments of the present disclosure.
[0012] [Figure 4] 1 illustrates an exemplary process for constructing a population frequency model using a logistic regression model, according to some embodiments of the present disclosure.
[0013] [Figure 5]5A and 5B show an example of a pre-calibration curve for a population frequency model constructed using a logistic regression model, according to some embodiments of the present disclosure.
[0014]
[0015] [Figure 6] 1 illustrates an example process for validating a population frequency model constructed using a logistic regression model, according to some embodiments of the present disclosure.
[0016] [Figure 7-1] 7A and 7B show example decision boundaries and gene-specific response curves for a population frequency model constructed using a logistic regression model, according to some embodiments of the present disclosure.
[0017]
[0018] [Figure 7-2] FIG. 7C shows an example of variant classification based on a population frequency model constructed using a logistic regression model, according to some embodiments of the present disclosure.
[0019] [Figure 8-1] 8A and 8B show an example of a population frequency modeling result of a population frequency model constructed using a logistic regression model, according to some embodiments of the present disclosure.
[0020]
[0021] [Figure 8-2]FIG. 8C shows another example of population frequency modeling results of a population frequency model constructed using a logistic regression model, according to some embodiments of the present disclosure.
[0022] [Figure 8-3] 8D and 8E show another example of population frequency modeling results for a population frequency model configured using a logistic regression model, according to some embodiments of the present disclosure.
[0023]
[0024] [Figure 8-4] FIG. 8F shows another example of population frequency modeling results for a population frequency model constructed using a logistic regression model integrated with a variant classification framework according to some embodiments of the present disclosure.
[0025] [Figure 9] 1 illustrates a method for population frequency modeling using a population frequency model constructed using a logistic regression model, according to some embodiments of the present disclosure.
[0026] [Figure 10] 1 illustrates an exemplary computing system including a population frequency modeling system according to some embodiments of the present disclosure.
[0027] [Figure 11] FIG. 1 is a block diagram of an exemplary computer system in which aspects of the present disclosure can operate. DETAILED DESCRIPTION OF THE INVENTION
[0028] Given a gene, a particular genetic disorder, and a population, there are characteristics of the gene-disease relationship associated with the frequency of pathogenic variants in that population. These characteristics include penetrance, age of onset, severity, inheritance pattern, and disease prevalence. Prevalence can refer to the allele frequency of a variant potentially associated with a genetic disorder in a generally healthy population, or the frequency with which patients in a given population are affected by the disease.
[0029] Given a disease, a pattern in which a greater-than-expected prevalence of a variant in a population has been empirically observed to constitute evidence of a benign classification for that disease, or conversely, a pattern in which a lower-than-expected prevalence should be considered evidence of a pathogenic classification for that disease. According to the American College for Medical Genetics (ACMG) guidelines for the interpretation of sequence variants, given a disorder, a variant that is present in the population more frequently than the prevalence of that disorder should constitute strong evidence of a benign classification.
[0030] However, the prevalence of pathogenic variants in a large population is the result of a complex interaction of various gene-disease attributes that reflect the genetic makeup of individuals for whom sequence data are available. As a result, accurately and systematically defining the relationship between allele frequency and pathogenicity for a given population and disease is currently a challenge across all gene-disease relationships. One challenge is how to define and calculate the expected excess bounds that should be used to classify a specific variant as benign for a given disease based on population frequency data. Another challenge is improving traditional binary classification approaches.
[0031] The Genome Aggregation Database (gnomAD) is a large population database developed by an international consortium with the goal of aggregating and harmonizing exome and genome sequence data from diverse large-scale sequencing projects and making summarized data available to the broader scientific community. For example, the gnomAD v2.1.1 dataset spans 125,748 exome sequences and 15,708 whole-genome sequences from unrelated individuals sequenced as part of various disease-specific and population genetic studies. Subsequent versions of gnomAD will include more genomes from multiple lineages and even greater diversity.
[0032] A major industry-wide challenge with classifying variants using general population frequency data obtained from large general population databases such as gnomAD is that many of these attributes traditionally considered important predictive features are difficult to obtain and not easy to interpret. For example, penetrance is difficult to measure accurately, is often unavailable, and can vary from variant to variant. Age at onset is rarely a single value but more typically a range of values. Severity is typically a qualitative, not quantitative, measure. Prevalence is often imprecise or unclear and can vary from one ancestry group to another.
[0033] Furthermore, gene-level measurements only provide information about the cumulative frequency of all pathogenic variants within a particular gene, not about specific individual variants within that gene. For example, a 0.1% prevalence could mean that there is one pathogenic variant in that gene with an allele frequency of 0.1%, or that there are 1,000 different pathogenic variants, each with an allele frequency of 0.0001%, or any other combination of pathogenic variants totaling 0.1%. Differences in the distribution of pathogenic variants have very different implications for how population allele frequency data should be used in variant classification, raising the need for further refinement.
[0034] Another limitation of conventional approaches that is overcome by the described population frequency modeling approach is that large population databases such as gnomAD contain only data that are estimates of the true population frequencies of various variants; as a result, those estimates are subject to sampling uncertainty. Such sampling uncertainty is accounted for by the described feature generation approach applied to large population data for variant classification as described herein.
[0035] Despite these and other limitations, previous methodologies for evaluating this type of allele frequency data have relied heavily on the aforementioned gene-disease attribution. As a result, genetic testing laboratories have had to conservatively set high, discrete allele frequency thresholds for classifying variants and then broadly apply these conservative thresholds to large groups of genes based on simple parameters (e.g., thresholds for genes associated with dominant diseases versus thresholds for genes associated with recessive diseases). Consequently, laboratories have not been able to fully utilize the allele frequency data available within these large population databases for variant classification. This has resulted in too many variants being classified as variants of uncertain significance (VUS), even when the available allele frequency data suggest that the variant is too common to be expected to cause disease. In these cases, patients may be left uncertain about the extent to which genetic test results reflect a patient's disease risk or diagnosis and the extent to which those results should affect their treatment.
[0036] Complexities and limitations in both the current state of research on a given gene-disease pair and the relationships between gene-disease attributes have diminished the use of quantitative algorithms for variant classification. The disclosed approach addresses these and other challenges in variant classification. Embodiments of the disclosed approach apply machine learning techniques for developing computational algorithms to machine learning models for population frequency modeling. The described population frequency model uses logistic regression to quantitatively estimate the probability of variant pathogenicity based on allele frequency data obtained from large population databases such as gnomAD.
[0037] Certain embodiments of the described population frequency models examine allele frequencies in the context of a relatively small set of attributes that are not among the gene-disease attributes mentioned above. In embodiments, we avoid using those aforementioned gene-disease attributes as features and instead engineered particularly useful combinations of other variant-, gene-, and / or location-level characteristics of a given variant. For example, embodiments of the described population frequency models have been shown to generate reliable pathogenicity estimates based on combinations of allele frequencies and fewer than 30 other input features in total. For example, in some embodiments, the number of input features ranges from about 20 features to about 30 features. The feature combinations engineered as described herein enable embodiments to exclude the aforementioned gene-disease attributes from the feature set used to construct the population frequency model.
[0038] Rather than relying on gene-disease attribution, the population frequency model resulting from application of the described approach utilizes the expected frequency distribution of known benign and known pathogenic variants in a given gene, allowing for a quantitative assessment of how much the allele frequency of a particular variant deviates from what would be expected if the variant were pathogenic. Using the described approach, the model calculates and outputs the probability of pathogenicity for each variant.
[0039] Unlike previous approaches that are not quantitative, a population frequency model configured as described maps a continuous measure (e.g., allele frequency) to a continuous outcome (e.g., probability of pathogenicity). For example, embodiments of the population frequency model can provide a variant-specific quantitative measure of population frequency as a continuous score. Furthermore, embodiments of the population frequency model enable a continuous, quantitative scoring approach instead of, or in addition to, a binary approach.
[0040] The output of the population frequency model may be used directly for variant classification or may be provided as input to a variant classification framework. For example, the model output may be incorporated into the population data portion of the Sherlock framework. The Sherlock framework is a semi-quantitative method of variant interpretation (i.e., a process that considers and evaluates evidence of pathogenicity, classifies the variant, and communicates the pathogenicity information to the patient in an understandable manner; also used to classify variants). Experimental results show that when model output generated according to the disclosed approach is provided as input to the Sherlock framework, the number of variants that would otherwise be classified as VUS due to lack of other evidence is reduced.
[0041] Without any information about complex gene-disease attributes, the population frequency model constructed as described was able to reproduce the expected relationships between genes with different gene-disease attributes, even though those attributes were not included in the set of features provided as model inputs to the model. Using gene-disease attributes of inheritance patterns as an example, genes associated with recessive conditions are generally expected to have a higher prevalence of pathogenic variants than genes associated with dominant conditions. This is because heterozygous carriers of autosomal recessive conditions are expected to be part of the population database cohort, while individuals with a single pathogenic variant in an autosomal dominant condition may be affected by the disease and therefore may be excluded. Consequently, even if two variants have the same allele frequency, the one observed in the recessive gene may be more pathogenic than the one observed in the dominant gene.
[0042] As explained in more detail below, when one embodiment of the described population frequency model was implemented on two such genes, without any prior information about the inheritance pattern, the model reported a higher probability of pathogenicity for variants in genes associated with autosomal recessive conditions than for those in genes associated with autosomal dominant conditions (e.g., LAMA2 vs. TSC2). For example, Figure 8B shows results demonstrating that an embodiment of the described population frequency model can distinguish the effects of different inheritance patterns on allele frequencies even in the absence of any prior information about the inheritance pattern (i.e., information about the inheritance pattern is not provided as an input to the model).
[0043] Furthermore, for a given allele frequency, variants within genes associated with diseases with higher severity, higher penetrance, and earlier disease onset are traditionally assumed to have a lower calculated pathogenicity probability than variants within genes associated with diseases with lower severity, lower penetrance, and later onset. For example, Figure 8A shows the results of population frequency modeling using the described approach for two genes (i.e., MLH1, KMT2D) associated with different severity modes, penetrance, and ages of onset. As shown, use of the described approach reveals that the pathogenicity probability at a given allele frequency can vary significantly between genes with the same inheritance pattern but different severity, penetrance, and onset.
[0044] Even when comparing multiple genes with complex combinations of gene-disease attributes, one embodiment of a population frequency model constructed as described herein provided probabilities of variant pathogenicity that matched known gene-disease attributes, even when information about those gene-disease attributes was not provided to the model. For example, Figure 8C shows the results of population frequency modeling for five genes (i.e., EYS, TGM1, CACNA1C, FBN1, and DNAH11) associated with different modes of inheritance, severity, penetrance, and age of onset. As shown, use of the described approach reveals that the probability of pathogenicity at a given allele frequency can vary significantly between any two genes, even when both genes are associated with, for example, the same inheritance pattern (e.g., DNAH11 and EYS) or similar penetrance and age of onset (e.g., CACNA1C and FBN1).
[0045] Given the experimental results realized to date, the population frequency model constructed as described represents a highly accurate, highly scalable, and fully quantitative solution for determining predicted population allele frequencies in the context of individual genes and across multiple genes.
[0046] Additionally or alternatively, embodiments of the population frequency model can be used to identify subtle but potentially important biological differences between genes, such as the differences observed between DNAH11, DNAI1 and DNAH5.
[0047] Incorporating the described tools into variant classification frameworks, such as the Sherlock interpretation system, substantially increases the ability to accurately and confidently classify variants. Within Sherlock, for example, one embodiment of the described population frequency model was applied four times to as many variants as previous methods that utilized allele frequencies from gnomAD in an alternative manner, leading to a resolution of approximately 15,000 unique VUS when implemented. Through simulation experiments on historical VUS in a local database, using a conservative thresholding process of the model output, the described tool is estimated to provide approximately 2.5% VUS reduction compared to other existing methodologies for evaluating population allele frequency data. The VUS reduction provided by the described model can be further improved in other variant classification frameworks that do not utilize such a conservative thresholding approach.
[0048] The present disclosure will be more fully understood from the detailed description given below, which refers to the accompanying drawings, which are for purposes of explanation and understanding and should not be construed as limiting the disclosure to the particular embodiments described.
[0049] In these drawings and the following description, reference may be made to components that have the same name but different reference numbers in different figures. The use of different reference numbers in different figures indicates that components with the same name may represent the same or different embodiments of the same component. For example, in some embodiments, components with the same name but different reference numbers in different figures may have the same or similar functionality, such that a description of one of those components with respect to one drawing may apply to other components with the same name in other drawings.
[0050] Additionally, in these drawings and the following description, components illustrated and described in connection with some embodiments may be used with or incorporated into other embodiments. For example, components shown in a particular drawing are not limited to use in connection with the embodiment to which that drawing relates, but may be used with or incorporated into other embodiments, including embodiments shown in other drawings.
[0051] FIG. 1 illustrates an example of a feature generation process for machine learning-based population frequency modeling, according to some embodiments of the present disclosure.
[0052] In FIG. 1 , a dataset of sequence data 110 includes DNA sequence data for one to N human populations, where N is a positive integer and a population may be defined by any one or more demographic criteria. The DNA sample includes one or more genes 102. Each gene includes one or more regions 104, e.g., a contiguous group or sequence of nucleotides or proteins (e.g., AGACGCT, where A refers to adenine, G refers to guanine, C refers to cytosine, and T refers to taurine). Each nucleotide within a region has one or more positions where a variant may be located. For example, in FIG. 1 , person 1 from population N has a cytosine at position 106 in region 104 of their copy of gene 102, while person 2 from population N has a taurine at position 106 in region 104 of their copy of gene 102, where the taurine is considered a genomic variant 108.
[0053] For each of the 1 to N human populations, the sequence data 110 includes gene-level data corresponding to each gene 102, region-level data corresponding to one or more regions 104, location-level data corresponding to one or more positions 106, and variant-level data corresponding to one or more genomic variants 108. Gene-level data refers to any property, attribute, or metric determined for a defined sequence of DNA identified as a gene. Region-level data refers to any property, attribute, or metric determined for a region of a gene, i.e., a defined portion of a gene that is typically larger than a position but smaller than the entire gene. Location-level data refers to any property, attribute, or metric determined for a position within a gene, i.e., a single nucleotide or amino acid position within a gene. In contrast, a gene may typically span thousands of nucleotides, and a variant typically occurs at a single position among those thousands of nucleotides. A region may refer to a portion of a gene that includes a variant and one or more adjacent or neighboring nucleotides.
[0054] Gene-level data includes information specific to a particular gene, such as gene length. Region-level data includes information specific to a particular region of the gene. For example, region-level data may include characteristics of a region that contains a variant. Position-level data includes information specific to a particular position of the gene, regardless of whether a variant is present at that position. Variant-level data includes information about a particular variant located at a particular position of the gene. A variant may contain more than one nucleotide. When a variant contains more than one nucleotide, the location of the variant refers to a region, such that the terms position and region may be synonymous in that context. The source of sequence data 110 may be a public population database, such as gnomAD.
[0055] 1 , a feature generation process 100 generates a feature set 112 for input to a machine learning model. The feature generation process 100 includes a feature extraction and computation subprocess 114 and a feature transformation subprocess 116. The feature extraction and computation subprocess 114 extracts and / or computes gene-level features 118, variant-level features 120, region-level features 122, and location-level features 124. For example, the feature extraction and computation subprocess 114 extracts features from a population database and / or computes features based on data extracted from the population database. One example of a computed feature is a confidence interval associated with another feature, such as a confidence interval computed for an allele frequency.
[0056] Feature transformation subprocess 116 applies one or more feature transformations to one or more combinations of gene-level features 118, variant-level features 120, region-level features 122, and / or location-level features 124 to generate population frequency metafeatures 126. Feature transformation may refer to an operation performed on a computed feature. For example, a computational operation may be applied to raw features (e.g., raw observations or measured data) to generate a computed feature, and another computational operation may be applied to the computed feature to perform a feature transformation on a computed future, e.g., to arrive at a population frequency metafeature 126.
[0057] Examples of feature transformations that may be applied to one or more combinations of gene-level features 118, variant-level features 120, region-level features 122, and / or location-level features 124 to generate population frequency metafeatures 126 include mathematical operations such as logarithm, exponential, square root, sum, multiplication, average, mean, median, and / or other mathematical functions. Examples of population frequency metafeatures 126 include features that are mathematical combinations of gene-level features and variant-level features or location-level features. For example, a population frequency metafeature 126 quantifies the predicted value of the allele frequency for a particular variant, region, or location within a gene in a given population. For example, a population frequency metafeature 126 quantifies how reliable the population allele frequency is as an indicator of pathogenicity for a particular variant, region, or location within a gene in a population. Examples of population frequency metafeatures 126 include the expected frequency distribution of known benign variants and known pathogenic variants within a particular gene in a particular population.
[0058] Examples of features that may be included in Feature Set 112 are shown below in Table 1. In one embodiment, data from gnomAD (v2.1.1), including lineage-specific allele counts, allele numbers, and gene-level constraint estimates (e.g., LOEUF), were collected for single nucleotide variants within genes with at least one strong disease association according to a reference genetic knowledge base (e.g., a local, internally developed knowledge base). [Table 1] [Table 1] [Feature examples]
[0059] In one embodiment, features (also referred to as characteristics) and group-specific confidence intervals were generated, resulting in a feature set for a total of 33,402 variants. Some of these features, such as allele frequencies, are obtained from population databases such as gnomAD. Other features, such as position-level constraints, are calculated based on raw data obtained from population databases. Constraints are metrics calculated on external datasets, for example, published through gnomAD. Constraints indicate the amount of variation allowed for a gene in a human population. As such, constraints can be used as a signal to identify genes that are more or less constrained; i.e., more or less essential for normal, healthy human function. Constraints are gene-level features.
[0060] Another example of gene-level characteristics is the synonymous to non-synonymous ratio. Synonymous means that the variant does not cause a change in the protein, and non-synonymous means that the variant will cause a change in the protein. The synonymous to non-synonymous ratio indicates the degree to which non-synonymous variants have been selected throughout evolution, which indicates the intolerance of protein mutations for essential protein functions.
[0061] The fixation index (FST) is a calculated future that indicates the differentiation of allele frequencies between different subpopulations. For example, allele frequencies may be obtained from a population database for several different subpopulations (e.g., Southeast Asians, Africans, etc.), and a calculation may be performed based on the subpopulation allele frequency data to indicate how differentiated the allele frequencies are across multiple populations.
[0062] Another example of a location-level feature is the identity (ID) of the largest subpopulation, i.e., the phylogenetic group identified as having the highest allele frequency for a particular variant. The ID of the largest subpopulation can be calculated by first calculating the allele frequency of the variant for each subpopulation, and then ranking or sorting the subpopulations based on allele frequency. Yet another example of a location-level feature is the actual frequency in the largest population, e.g., an x% (e.g., 95%) upper bound of the allele frequency in the identified largest subpopulation.
[0063] An example of a region-level feature is the moving window missense. The moving window missense is the number of missense mutations (nucleotide mutations that result in amino acid changes, i.e., significant changes in the protein) within a region of a gene. The moving window missense can indicate the degree to which missense mutations are tolerated within a particular region of a gene.
[0064] Feature transformations can be applied to gene-level constraints to apply constraint information at the variant level. For example, a mathematical combination of allele frequencies and constraints (e.g., a logarithmic transformation) is an example of a population frequency metafeature. Another example of a population frequency metafeature is a position-level constraint. A position-level constraint is created by applying a feature transformation to a gene-level constraint. For example, a position-level constraint can be calculated by taking the average constraint per position across the gene. A position-level constraint indicates the degree to which a particular position or region within a gene is constrained (as opposed to the gene as a whole). A position-level constraint can be calculated by using the variants observed within a region as a normalization metric for the gene-level constraint. Another example of a population frequency metafeature is an allele frequency confidence interval to which a feature transformation is applied, such as the binomial ratio of the allele frequency confidence interval.
[0065] Another example of a population frequency meta-feature is AF_DIV_LOEUF, which is calculated by dividing the allele frequency by the constraint and then taking the exponent of the resulting quotient. Another example of a population frequency meta-feature is SIN2_MIS_EXP, which is an index of the ratio of synonymous to non-synonymous missense variants in a gene.
[0066] Other examples of population frequency metafeatures can be created by applying one or more feature transformations to a combination of one or more gene-level features and allele frequencies, or to a combination of gene-level constraints and one or more variant-level features and / or position-level features.
[0067] FIG. 2 illustrates an example process for constructing a population frequency model using a logistic regression model and machine learning techniques, according to some embodiments of the present disclosure. The process is performed by processing logic, including hardware (e.g., a processing device, circuitry, dedicated logic, programmable logic, microcode, device hardware, integrated circuit, etc.), software (e.g., instructions running or executed on a processing device), or a combination thereof. In some embodiments, the method is performed by components of a computing system, which may include components or flows shown in FIG. 2 that, in some embodiments, may not be specifically shown in other figures, and / or may include components or flows shown in other figures that, in some embodiments, may not be specifically shown in FIG. 2. Although shown in a particular order or sequence, unless otherwise specified, the order of the processes may be modified. Therefore, the illustrated embodiments should be understood as merely examples, and the illustrated processes may be performed in a different order, and some processes may be performed in parallel. Additionally, at least one process may be omitted in various embodiments. Thus, not all processes are required in all embodiments. Other process flows are possible.
[0068] 2 , model development process 200 uses one or more machine learning techniques to prepare one or more datasets, e.g., training dataset 234, validation dataset 238, for input to a machine learning model, e.g., logistic regression model 228. Portions of process 200 are performed by components of a population frequency modeling computing system (e.g., population frequency modeling system 1050 of FIG. 10 , described below), including genetic data selection subsystem 206, feature generation subsystem 212, modeling and calibration subsystem 220, and model validation subsystem 230. Data sources from which data is received during various portions of process 200 include unlabeled population data 202, labeled population data 204, logistic regression model 228, model performance criteria 232, training dataset 234, model validation criteria 236, and validation dataset 238.
[0069] The unlabeled population data 202 includes DNA sequence data for one or more human populations. For example, the unlabeled population data 202 includes population data extracted from gnomAD or a similar database. The unlabeled population data 202 does not include an associated pathogenicity label or score. The pathogenicity label or score (e.g., benign or pathogenic) is obtained from the labeled population data 204. The labeled population data 204 is a reference database that links variants with associated ground truth pathogenicity labels or scores, such as ClinVar or an internally developed database curated by genetic scientists and / or other genetic experts. The ground truth labels or scores can be concatenated or merged with the corresponding unlabeled population data 202 (e.g., using common key values) to generate labeled population data to which a logistic regression model 228 can be applied using a supervised machine learning approach.
[0070] The labeled population data 204 is concatenated or merged with the unlabeled population data 202 either before or after one or more operations of the genetic data selection subsystem 206. For example, the labeled population data 204 may be concatenated with the unlabeled population data 202 after the genetic data selection subsystem 206 to create one or more training and / or validation data sets.
[0071] The genetic data selection subsystem 206 evaluates gene-specific datasets of the unlabeled population data 202 and selects one or more gene-specific datasets to be used to generate a feature set that can be input into the logistic regression model 228. The genetic data selection subsystem 206 applies one or more filters to the population data that include criteria to determine whether data for particular genes will be included or excluded from the model development process. For example, genes that do not have a strong association with any disease may not be eligible for inclusion in the population frequency model.
[0072] For example, the genetic data selection subsystem 206 applies a gene-disease plausibility filter 208 to the unlabeled population data 202 to create a gene-disease plausibility filtered subset of the unlabeled population data 202, concatenates or merges the gene-disease plausibility filtered subset with the labeled population data 204 such that each item in the gene-disease plausibility filtered subset matches a corresponding ground truth label, and then applies a label minimum filter 210 to the labeled gene-disease plausibility filtered subset of the unlabeled population data 202 to generate and output a labeled gene-disease plausibility and label minimum filtered dataset.
[0073] The gene-disease plausibility filter 208 evaluates the unlabeled population data 202 to filter out data that is not strongly correlated with the disease. For example, the gene-disease plausibility filter 208 retains unlabeled population data 202 for variants that have a high probability (e.g., greater than or equal to 90%) of being either benign or pathogenic for a particular disease and filters out unlabeled population data 202 for variants that have an uncertain probability (e.g., less than 90%) of being either benign or pathogenic for a particular disease. Gene-disease association data can be obtained from public sources, such as the Geneticus database.
[0074] The label minimum filter 210 evaluates the labeled gene-disease plausibility filtered subset of the unlabeled population data 202 against a label minimum threshold. The label minimum filter 210 retains items of the labeled gene-disease plausibility filtered subset of the unlabeled population data 202 that have both multiple pathogenic labels and multiple benign labels that meet or exceed the applicable label minimum threshold, and excludes items of the labeled gene-disease plausibility filtered subset of the unlabeled population data 202 with multiple labels that do not meet or exceed the label minimum threshold, i.e., do not have enough benign and pathogenic labels to be useful for model training. For example, gene datasets that do not have at least a minimum number of benign labels are excluded from the dataset.
[0075] The feature generation subsystem 212 generates features for inclusion in a feature set for input to the logistic regression model. For example, embodiments of the feature generation subsystem 212 generate various features using the feature extraction, feature computation, and feature transformation techniques described above with reference to FIG. 1.
[0076] The feature generation subsystem 212 receives the labeled gene-disease relevance and label minimum filter dataset as input and generates a feature set (e.g., feature set 112) based on the received labeled gene-disease relevance and label minimum filter dataset. The feature generation subsystem 212 includes a feature extraction component 214, a feature computation component 216, and a feature transformation component 218. The feature extraction component 214 extracts features from the received labeled gene-disease relevance and label minimum filter dataset. The extracted features may include raw features and / or computed features. Examples of raw features and computed features are described above with reference to FIG. 1.
[0077] The feature computation component 216 applies one or more computational operations to one or more raw features produced by the feature extraction component 214. Examples of computational operations that may be applied to one or more raw features are described above with reference to FIG.
[0078] The feature transformation component 218 applies one or more computational operations to one or more computed features produced by the feature extraction component 214 or the feature computation component 216. Examples of feature transformations that may be applied to one or more computed features are described above with reference to FIG.
[0079] The feature generation subsystem 212 generates and outputs a feature set that includes raw features, computed features, and / or feature transforms. Examples of features and feature sets that can be generated and output by the feature generation subsystem 212 are described above with reference to FIG. 1.
[0080] The modeling and calibration subsystem 220 receives as input the feature set engineered and output by the feature generation subsystem 212. The modeling and calibration subsystem 220 includes a dataset creation component 222, a model training component 224, and a model calibration component 226.
[0081] The dataset creation component 222 divides the feature sets engineered and output by the feature generation subsystem 212 into training and validation datasets, e.g., training dataset 234 and validation dataset 238. For example, in some embodiments, the dataset creation component 222 creates gene-specific training and validation datasets for each of up to 600 or more different genes, where each gene-specific dataset includes a set of input features (e.g., feature set 112) associated with a particular variant within a particular gene. As an example, for a given variant, each feature set created by the dataset creation component 222 includes approximately 24 different variant characteristics, including allele frequency data extracted from a population database such as gnomAD, associated confidence intervals, gene length, and one or more population frequency meta-features.
[0082] The model training component 224 and the model calibration component 226 perform a model training process that causes the logistic regression model 228 to develop mathematical representations of the relationships between input features and ground truth labels, so that the resulting model can be used to predict the pathogenicity of novel variants (i.e., variants not previously seen by the model) based on the input features associated with those novel variants. The mathematical representations of these relationships can be shown, for example, as a two-dimensional plot of feature values (e.g., variant characteristics) (x-axis) versus pathogenicity labels (y-axis).
[0083] The model training component 224 iteratively applies the logistic regression model 228 to the training dataset 234, adjusting one or more model parameters and / or feature coefficients until the difference between the predicted model output generated by the logistic regression model 228 and the expected model output supported by ground truth labels obtained via the labeled population data 204 satisfies (e.g., meets or exceeds) the model performance criteria 232. If the model performance criteria 232 are met, the modeling and calibration subsystem 220 terminates the model training process and generates the trained logistic regression model 228. More detailed examples of model training and calibration processes that may be used to create the trained logistic regression model 228 are described below with reference to FIGS. 4, 5A, and 5B.
[0084] The model validation subsystem 230 applies a model validation process to the trained logistic regression model 228 generated by the modeling and calibration subsystem 220. The model validation subsystem 230 applies the trained logistic regression model 228 to a validation dataset 238 to determine whether the model validation criteria 236 are met (e.g., met or exceeded). A more detailed example of the model validation process is described below with reference to Figures 6, 7A, and 7B.
[0085] If the trained logistic regression model 228 is successfully validated by the model validation subsystem 230, the validated logistic regression model 228 may be used for inference, e.g., to generate pathogenicity predictions or estimates for novel (i.e., previously unseen) variants. Alternatively, or in addition, the predictions output by the validated logistic regression model 228 may be stored for future use (e.g., for access or lookup by one or more downstream processes, systems, or services). A more detailed example of inference-time uses of a logistic regression model and a logistic regression model configured for variant classification using the feature sets and techniques described herein is described below with reference to FIG. 3.
[0086] FIG. 3 illustrates an exemplary use of a population frequency model constructed using a logistic regression model, according to some embodiments of the present disclosure.
[0087] The logistic regression model 306 is a statistical machine learning model that models the relationship between X and Y using a logistic function, where the probability of Y is a linear combination of the independent variables in the input X. Mathematically, a simplified form of the logistic function is:
number
[0088] 3, a logistic regression model 306 is constructed via the supervised machine learning training, calibration, and validation process described. The logistic regression model 306 includes a logistic function 308. The logistic function 308 includes feature coefficients 310. The feature coefficients 310 include a regression coefficient β for each feature input x (e.g.,
number
[0089] The logistic regression model 306 also includes model hyperparameters 312, which are selected or tuned at a global level and are generally not modified based on specific instances of training data. The model hyperparameters 312 include, for the logistic regression model 306, a penalty or regularization parameter (e.g., L1 or L2) and a C or regularization strength parameter. The penalty or regularization parameter is adjustable to adjust the model generalization error and limit overfitting. Typical values for the penalty or regularization parameter are L1 and L2. In some embodiments of the logistic regression model 306, the penalty or regularization parameter is set to L1. The C or regularization strength parameter, in conjunction with the penalty, limits overfitting. A smaller value of C specifies stronger regularization. Regularization imposes a cost on model complexity by penalizing large coefficient values. Regularization limits the number of different ways the model can fit the data. The value of C is subject to a hyperparameter optimization process, such as a grid search. In some embodiments of the logistic regression model 306, the value of C is set to a value in the range of about 1 to about 10.
[0090] The model hyperparameters 312 may be tuned, for example, using a grid search cross-validation procedure. In some embodiments, for example, an automated hyperparameter tuning tool, e.g., GridSearchCV, is used for hyperparameter tuning. In other embodiments, other hyperparameter optimization methods, such as random search and Bayesian search, are used. In the illustrated embodiment, hyperparameter tuning was performed on the training dataset to compare the performance of various combinations of model type and hyperparameters. The simplest model with the strongest regularization (e.g., an L1 regularized logistic regression model, C=10) that resulted in high validation performance (e.g., AUROC=0.92) was selected for inference. Model predictions for the training and validation datasets were generated separately by averaging calibrated predictions from 10-fold cross-validation runs, such that the validation dataset was not used for training set predictions. Inference on VUS was performed by first refitting the model to the entire labeled set before making predictions for new variants.
[0091] The logistic regression model 306 can be configured as either a binary classifier or a scoring model. In binary classification mode, the output of the logistic regression model 306 indicates whether the predicted outcome is pathogenic or benign as a binary value, e.g., 0 indicates benign and 1 indicates pathogenic, for a given set of input features. In scoring mode, the output of the logistic regression model 306 includes a score (e.g., a number between 0 and 1, inclusive) corresponding to the probability that the predicted outcome is pathogenic or benign.
[0092] The logistic regression model 306 may be configured and implemented as a network service. For example, the logistic regression model 306 may be configured using a machine learning library such as scikit-learn. For example, the logistic regression model 306 may be configured via an application programming interface (API), e.g., via an API call such as ML_library.model.logistic_regression(p1, p2, ... pn), where p indicates a parameter or argument of the call, such as a model hyperparameter or an input feature set identifier. Once configured, the logistic regression model 306 and / or its output may be hosted on one or more servers and / or data storage devices for accessibility to one or more requesting processes, systems, devices, frameworks, or services.
[0093] 3, a Feature Set 302 includes an item or instance of a feature 304 for each gene-disease variant combination. For example, a Feature Set 302 may include items or instances of feature 304 for N genes, N diseases, and N variants, where the value of N may be the same or different in each instance. Each item or instance 304 may include one or more gene-level features x g1 ...x gN , one or more location-level features x p1 ...x pN , one or more variant-level features x v1 ...x vN and one or more population frequency metafeatures x mf1 ...x mfN Contains gene-level features x g1 ...x gN , location-level features x p1 ...x pN , variant-level feature x v1 ...x vN and population frequency meta feature x mf1 ...x mfNExamples include those described with reference to FIG. 1. Alternatively, or in addition, feature set 302 may include region-level features. In other embodiments, feature set 302 may include gene-level features, variant-level features, and population frequency meta-features, but may not include location-level features or region-level features. Features 304 may be grouped into feature sets 302 using, for example, a linkage function.
[0094] Embodiments of the feature set 302 are limited to quantitative, numerical, and categorical features, and do not include qualitative features or gene-disease attributes. Prior to input to the logistic regression model 306, items or instances of features 304 may be converted into vector representations. In some embodiments, these feature vector representations are converted into a compressed form, such as an embedding, prior to input to the logistic regression model 306.
[0095] In response to each instance of a feature in the feature set 302, the logistic regression model 306 estimates an outcome, P GDV The logistic regression model 306 calculates and outputs (Y|X) 314. The predicted outcome generated by the logistic regression model 306 based on the instances of features in the feature set 302 is in the form of a binary output (e.g., 0 for benign or 1 for pathogenic) or a score (e.g., a value between 0 and 1). The output may be stored in data storage for subsequent lookup or provided to one or more downstream systems, processes, devices, frameworks, and / or services.
[0096] FIG. 4 illustrates an exemplary process for constructing a population frequency model using a logistic regression model, according to some embodiments of the present disclosure. In FIG. 4, model training process 400 applies a logistic function to a training dataset 401 using regression-based supervised machine learning. Training dataset 401 includes feature set 402 and labels 404. Feature set 402 is similar to feature set 302, and each corresponding item or instance of the described feature set is accompanied by a ground truth label. In one embodiment, a feature set totaling 33,402 variants (i.e., labels) was used for training across 823 genes (n=18,148 benign variants and 15,234 pathogenic variants). All training variants from 25% of the genes were reserved as a test (or validation) set and remained unused until after model training was completed. Elimination of labeled training data based on gene selection is not required; i.e., not all embodiments are required to provide a training set in the described manner.
[0097] In subprocess 406, in the first training iteration, feature coefficients are initialized or assigned to each feature input in feature set 402. The feature coefficients are initialized, for example, by randomly setting coefficient values. The feature coefficient values assigned in subprocess 406 act as weights applied to each feature to generate weighted features 410. A logistic function is applied to the weighted features in subprocess 412 to generate predicted outputs 414. The ground truth labels 404 provide expected outputs 408 for supervised machine learning. In subprocess 416, the predicted outputs 414 are evaluated by calculating a loss (or error) based on the expected outputs 408 and the predicted outputs 414. The loss is calculated using a loss function such as a gradient descent algorithm.
[0098] At each iteration, a decision subprocess 418 evaluates the difference between the predicted output and the expected output, for example, using a comparison of the loss function output to a stopping condition that depends on the change in the loss function output. The change in the loss function output is compared to an error tolerance threshold (which may be referred to herein as a model performance criterion or a model convergence criterion). If the model performance criterion, e.g., an error tolerance threshold, is not met (e.g., below or equal to a threshold performance level, or exceeds a maximum allowed error value, or the loss has not stopped improving more than allowed, or the error has not stopped decreasing more than allowed threshold), model training continues for another iteration. If the loss has stopped improving more than allowed threshold, the model has converged and training ends.
[0099] In subsequent training iterations, the values of one or more of the feature coefficients are adjusted, the logistic function is applied to additional instances of the feature set 402, and the output of the logistic function is evaluated using the loss function and error tolerance described above. The training process 400 ends when the model performance criteria, i.e., error tolerance threshold, are met and the model has converged (e.g., the output of the comparison, e.g., the loss function, is greater than or equal to the threshold performance level, or is within or less than the maximum allowed error value, or the loss has stopped improving by more than the tolerance threshold).
[0100] FIG. 5A shows an example of a pre-calibration curve for a population frequency model constructed using a logistic regression model, according to some embodiments of the present disclosure.
[0101] In operation, for a given gene, variant, and disease, one embodiment of a population frequency model typically outputs many scores close to 0 (e.g., benign) and many scores close to 1 (e.g., pathogenic). Figure 5A shows an example of a pre-calibration curve compared to a desired or full calibration. Before calibration, the model tends to underestimate variant pathogenic probabilities in lower score regions (e.g., the portion of the pre-calibration curve below the full calibration line). The difference between the full calibration and pre-calibration curves may be referred to as error. The calibration process adjusts the values on the pre-calibration curve so that they more closely align with the full calibration diagonal. Depending on the error in a given region of the model output, the scores are shifted (e.g., by adjusting one or more model coefficients) so that the expected probabilities match the predicted probabilities. The pre-calibration curve in Figure 5A is an aggregate curve of the scores output by the model for multiple different genes. A similar calibration process can be performed using similar curves for individual genes.
[0102] Figure 5B shows an example of a post-calibration curve for a population frequency model constructed using a logistic regression model, according to some embodiments of the present disclosure. Figure 5B shows an example of a post-calibration curve resulting from a calibration applied to the curve of Figure 5A in comparison to a desired or perfect calibration. After calibration, the model output tends to align more closely with the perfect calibration line. The post-calibration curve in Figure 5B is an aggregate curve of the scores output by the model for multiple different genes. A similar calibration process can be performed using similar curves for individual genes.
[0103] FIG. 6 illustrates an example process for validating a population frequency model constructed using a logistic regression model, according to some embodiments of the present disclosure. In FIG. 6, the validation process 600 applies a trained logistic regression model 604 to a validation dataset 602. An example of the validation dataset 602 is a portion of the feature set 402 reserved for validation and not used for training. For example, referring to FIG. 2, the dataset creation component 222 creates several gene-specific partitions of the data to provide specific datasets for entire genes that are not used for model training but that can be used to create the validation dataset. The trained logistic regression model 604 is, for example, a logistic regression model trained and constructed as described herein.
[0104] In response to processing the validation data set 602, the trained logistic regression model 604 provides validation parameters, such as model outputs 606, feature weights 608, and decision boundaries 610, to subprocess 612 for evaluation. Subprocess 612 uses validation criteria to evaluate the performance of the trained logistic regression model 604, for example, by testing samples of the actual model outputs 606, feature weights 608, and / or decision boundaries 610 and comparing the tested samples to expected or reference values or by calculating one or more validation metrics.
[0105] Model validation may be performed over multiple iterations, each using a slightly different validation data set not used in training. The trained logistic regression model 604 generates virulence predictions across the validation set, and based on the average performance of the model across those splits, the performance of the model parameters may be validated and adjusted. Examples of metrics that may be used to evaluate model performance include mean squared error or area under the receiver operating characteristic curve.
[0106] Approaches that may be used to validate the trained logistic regression model 604 include evaluation of dataset performance metrics, decision boundary testing, evaluation of model-level performance metrics, gene and variant-level testing, feature weight testing, and benchmark comparisons with existing frameworks (e.g., Sherlock).
[0107] Decision boundary testing refers to the process of evaluating the actual boundaries of the model, the thresholds at which the model distinguishes between benign and pathogenic. An example of evaluation of a decision boundary, described below, is shown in Figure 7A. An example of evaluation of a dataset performance metric, described below, is shown in Figure 7B.
[0108] In response to the evaluation of model performance performed by sub-process 612, a particular gene-specific dataset and corresponding model output may be selected for inclusion in a database or downstream system, device, process, or service if the dataset meets the respective validation criteria. Alternatively, if a particular gene-specific dataset does not meet the respective validation criteria, the dataset may be rejected or marked as unavailable for subsequent use, e.g., not stored in a database or provided to any downstream system, device, process, or service.
[0109] For technical validation, model outputs were validated at the gene level for performance and calibration. Genes that showed poor calibration (i.e., Brier score >0.15) or poor discriminant performance (i.e., AUROC <0.80) were flagged for removal from the final model output. In total, 591 of the original 823 genes achieved sufficiently high calibration and discriminant performance requirements. Notably, the 232 genes that were flagged for removal were nevertheless used for subsequent stages when we calculated performance to avoid artificially biasing the overall performance assessment.
[0110] To gain further confidence in the predictive output of the population frequency model for clinical validation, a concordance analysis was performed on three independent, internally developed in silico algorithms that rely on orthogonal data types. As shown in Table 2 below, concordance was high for all three comparisons. [Table 2] [Table 2] [Example of match rate]
[0111] Additionally, model output was examined for a particularly challenging class of variants: high-allelic frequency pathogenic variants. These are variants that resemble benign variants based on population allele frequency data alone, but are classified as pathogenic or likely pathogenic due to other evidence supporting pathogenicity. These include variants such as lineage-specific founder variants and hypomorphic alleles associated with milder phenotypes or reduced penetrance relative to typical pathogenic variants in the same gene.
[0112] Among the 591 genes included in the analysis, 30 high-allele frequency pathogenic variants in 26 genes were identified that would have been classified as benign or likely benign based on allele frequency alone but had conflicting evidence supporting pathogenicity. The population frequency model still predicted 18 of these 30 variants to be strongly or moderately benign (NPV ≥ 95%), whereas the remaining 12 variants were predicted with strong or moderate support for benignity (95% > NPV ≥ 80%), support for pathogenicity (PPV ≥ 80%), or with insufficient certainty (NPV < 80% and PPV < 80%). These results demonstrate that the population frequency model has greater specificity for identifying benign variants than traditional approaches, despite the fact that important consideration of all conflicting evidence remains a necessary part of variant classification. An example of how the NPV / PPV values obtained from the population frequency model described below can be converted into discrete scores, e.g., Sherlock scores (such as 5 benign, 3 benign, 1 benign, and 1 pathogenic point), is shown in Figure 7C.
[0113] FIG. 7A shows an example of a decision boundary for a population frequency model constructed using a logistic regression model, according to some embodiments of the present disclosure. In FIG. 7A, variants are projected onto a two-dimensional plane. The plot in FIG. 7A is a two-dimensional representation of variants represented as circles and x's. To perform decision boundary testing, one or more locations within the plot in FIG. 7A can be sampled and tested to determine whether the prediction for that variant is pathogenic or benign. Decision boundary 706 separates locations within region 702 (pathogenic) from locations within region 704 (benign). Locations closer to decision boundary 706 have lower certainty for the associated prediction, while locations farther from decision boundary 706 are more confident classifications.
[0114] Figure 7B shows an example of a gene-specific response curve for a population frequency model constructed using a logistic regression model, according to some embodiments of the present disclosure. Using Figure 7B, the predictive value of the POPMAX feature can be evaluated for individual genes. The feature on the x-axis is the square root of the calculated allele frequency. An increase in allele frequency correlates with an increase in benign prediction. Figure 7B shows that the relationship between allele frequency and benign prediction differs for different genes.
[0115] FIG. 7C shows an example of variant classification based on a population frequency model constructed using a logistic regression model, according to some embodiments of the present disclosure. In FIG. 7C, counts of the number of variants with each score value are represented as a histogram. For example, approximately 2,900 variants were assigned a score of 0 by the model, and more than 3,000 variants were assigned a score greater than 0.8 by the model. The histogram is divided into classification groupings or bins and point values associated with each score grouping or bin. For example, the portion of the histogram associated with a score of 0 is assigned 5 points (e.g., benign), the portion of the histogram adjacent to the 0 score region is assigned 3 points (e.g., likely benign), while the portion of the histogram associated with a score greater than 0.6 is associated with a point value of 1. Each portion in the center of the histogram is assigned a point value of 0 because the associated score value has a lower confidence level. Point values assigned based on the model output may be combined with other evidence to which the classification framework is applied.
[0116] Generally, the number of points associated with a particular score correlates with the confidence or accuracy of the prediction. The number of bins need not be four; any number of bins (including no bins or an infinite number of bins) can be used. The points mapped to the scores output by the model can be incorporated into another variant classification framework, such as Sherlock. In this way, the output of the described population frequency model can be applied to the Sherlock framework or any other variant labeling or classification framework.
[0117] To integrate model predictions into the Sherlock variant classification framework, five tiers of predictive value were established based on predictive performance thresholds measured in negative predictive value (NPV) and positive predictive value (PPV). Four of these tiers reflect the existing Sherlock framework for evaluating population frequency data and were defined as: (1) sufficient confidence to classify as benign in the absence of convincing contradictory evidence [strong benign]; (2) sufficient confidence to classify as possibly benign in the absence of convincing contradictory evidence [moderate benign]; (3) consistent with benign classification but insufficient to reach a likely benign classification without orthogonal evidence supporting benign [supported benign]; and (4) consistent with pathogenic classification but insufficient to reach a likely pathogenic classification without orthogonal evidence supporting pathogenicity [supported pathogenic]. The predictive performance thresholds for these four tiers were defined as (1) strongly benign >99% NPV, (2) moderately benign >95% to 99% NPV, (3) favorable to benign >80% to 95% NPV, and (4) favorable to pathogenic >80% PPV, respectively. The fifth and final tier corresponded to predictions below 80% PPV and 80% NPV, which were considered to have insufficient certainty and assigned weights within the Sherlock scoring system. Overall model performance and tier-level classification performance were evaluated across 25% of the genes reserved for the test set. Overall model performance achieved an AUROC test of 0.92, confirming that the model generated from the training set generalized to the test set genes. For the four PPV and NPV tiers, predictions met or exceeded target performance across the test set. In the current iteration of the model, model output was limited to missense and single-nucleotide context nonsense variants. Future development of the model may extend its range of predictions to new variant types.
[0118] Figures 8A, 8B, 8C, and 8D show predicted outputs, i.e., population frequency modeling results, generated by an embodiment of a population frequency model configured as described. Figure 8A shows a comparison of variants found in the genes MLH1 (associated with Lynch syndrome) and KMT2D (associated with Kabuki syndrome). Figure 8B shows a comparison of variants found in the genes TSC2 (associated with tuberous sclerosis) and LAMA2 (associated with LAMA2-associated muscular dystrophy). Figure 8C shows a comparison of the genes EYS (associated with retinitis pigmentosa), TGM1 (associated with ichthyosis), CACNA1C (associated with infantile spasms, etc.), FBN1 (associated with Marfan syndrome, etc.), and DNAH11 (associated with primary ciliary dyskinesia).
[0119] All variants are plotted based on the output generated by an embodiment of the described population frequency model. Each variant (represented as a dot in each plot) is arranged along the x-axis according to its reported allele frequency in the gnomAD population database. The y-axis represents the probability that the variant is pathogenic based on the population frequency model configured as described. The gene-disease attributes for each gene are summarized in the table, where AD indicates autosomal dominant and AR indicates autosomal recessive.
[0120] These model outputs highlight the limitations of traditional approaches to variant classification, which involve extracting discrete allele frequency thresholds and then broadly applying them to many genes. When such a traditional approach is used, variants may cross the threshold with very different probabilities of pathogenicity, depending on the gene. Notably, variants in DNAH11 associated with autosomal recessive primary ciliary dyskinesia generally have lower pathogenicity probabilities than variants with similar allele frequencies in CACNA1C and FNB1 (both of which are associated with autosomal dominant conditions). While this is in some cases due to DNAH11 being associated with a highly penetrant, early-onset, severe, and rare condition, it could also be due to the large size of this gene (4,516 codons), which results in many different, but individually rare, pathogenic variants. This further highlights the challenges associated with applying a common allele frequency threshold to many genes.
[0121] While the plots in Figures 8A, 8B, 8C, and 8D show machine-learned relationships between variant pathogenicity probability and allele frequency for specific genes and variants generated by population frequency models, it should be understood that the model inputs included other features in addition to allele frequency.
[0122] Figure 8A shows an example of population frequency modeling results for a population frequency model configured using a logistic regression model, according to some embodiments of the present disclosure. Figure 8A shows an example of the pathogenicity probability estimated and output for each variant in two different genes, MLH1 and KMT2D, by the described population frequency model. In Figure 8A, each dot in the plot represents a different specific variant in a particular gene. The majority of variants in gene MLH1 are plotted within region 802, while the majority of variants in gene KMT2D are plotted within region 804. The y-axis shows the model output, i.e., the probability of pathogenicity of the variant, and the x-axis shows the allele frequency of the variant in a given human population. As the frequency of the variant in the population increases, the probability of pathogenicity decreases.
[0123] More specifically, Figure 8A shows an example of a population frequency model configured as described that outputs variant pathogenicity probabilities that account for differences in genes with the same inheritance pattern. For example, as shown in the table included in Figure 8A, both MLH1 and KMT2D are dominant for a particular disease, but have different severity, penetrance, and age-of-onset characteristics.
[0124] While it might be expected that population frequency information should be used differently for genes with different inheritance patterns, in the example of Figure 8A, these two genes are both dominant but have different other important characteristics (aggression, penetrance, age of onset). In conventional variant classification systems, it is difficult to determine how much to adjust the variant pathogenicity probabilities of these two genes to account for these differences.
[0125] In contrast, the model output, using a population frequency model constructed as described, can automatically determine the amount to adjust variant pathogenicity probabilities to account for the similar inheritance pattern and different attack, penetrance, and age-of-onset characteristics of these two genes, without having access to any information about these characteristics. Even though the model input does not include any information indicating that the two genes have the same inheritance pattern or different attack, penetrance, and age-of-onset, the model learned through machine learning-based training and calibration using the described feature set that variants in these two genes should be treated differently with respect to population frequency thresholds, and how much these probabilities need to be adjusted.
[0126] Figure 8B shows another example of population frequency modeling results for a population frequency model configured using a logistic regression model, according to some embodiments of the present disclosure. Figure 8B shows an example of the pathogenicity probability estimated and output for each variant in two different genes, TSC2 and LAMA2, by the described population frequency model. In Figure 8B, each dot in the plot represents a different specific variant in a particular gene. The majority of variants in gene LAMA2 are plotted within region 822, while the majority of variants in gene TSC2 are plotted within region 824. The y-axis shows the model output, i.e., the probability of pathogenicity of the variant, and the x-axis shows the allele frequency of the variant in a given human population. As the frequency of the variant in the population increases, the probability of pathogenicity decreases.
[0127] More specifically, Figure 8B shows an example of a population frequency model configured as described that outputs different variant pathogenicity probabilities for two genes with similar onset, penetrance, and disease characteristics but different inheritance patterns. As shown in the table included in Figure 8B, TSC2 is dominant for a particular disease, while LAMA2 is recessive. The model output supports the intuition that variants in genes with different inheritance patterns should be held to different population frequency standards. While these differences can be taken into account in traditional variant classification processes by setting different thresholds for recessive and dominant genes, traditionally used thresholds are still categorical (e.g., recessive, dominant) and often need to be manually adjusted.
[0128] In contrast, the model output confirms conventional intuition using a population frequency model constructed as described, but without access to any information about inheritance patterns: even though the model input does not contain any information indicating that these two genes have different inheritance patterns, the model learned through machine learning-based training and calibration using the described feature set that variants in these two genes should be treated differently with respect to population frequency thresholds, and how much these probabilities need to be adjusted.
[0129] Figure 8C shows another example of population frequency modeling results for a population frequency model configured using a logistic regression model, according to some embodiments of the present disclosure. Figure 8C shows an example of the pathogenicity probability estimated and output for each variant in several different genes, EYS, TGM1, CACNA1C, FBN1, and DNAH11, by the described population frequency model. In Figure 8C, each dot in the plot represents a different specific variant in a particular gene. The majority of variants in gene EYS are plotted within region 830, the majority of variants in gene TGM1 are plotted within region 832, the majority of variants in gene CACNA1C are plotted within region 834, the majority of variants in gene FBN1 are plotted within region 836, and the majority of variants in gene DNAH11 are plotted within region 838. The y-axis shows the model output, i.e., the probability of pathogenicity of the variant, and the x-axis shows the allele frequency of the variant in a given human population. As the frequency of a variant in a population increases, the probability of it being pathogenic decreases.
[0130] More specifically, Figure 8C shows an example of a population frequency model configured as described that outputs different variant pathogenicity probabilities for multiple different genes with differences that are counterintuitive to variant classification scientists. As shown in the table included in Figure 8C, these genes have a combination of similarities and differences in traditional gene-disease attributes. While the differences in pathogenicity probabilities between these genes are not apparent to human consciousness, the population frequency model configured as described takes these gene-specific differences into account through a machine learning-based training and calibration process using the described feature set. As a result, the pathogenicity curves generated based on the model output indicate that variants in these genes should be treated differently with respect to the population frequency threshold.
[0131] Figure 8D shows another example of population frequency modeling results for a population frequency model configured using a logistic regression model, according to some embodiments of the present disclosure. Figure 8D shows an example of pathogenicity probabilities estimated and output for each variant in a given gene, MLH1, by the described population frequency model. In Figure 8D, each dot in the plot represents a different specific variant. The y-axis indicates the probability of the variant being pathogenic, and the x-axis indicates the allele frequency of the variant in a given human population. As the frequency of the variant in the population increases, the probability of pathogenicity decreases. For a specific variant 870, the model output indicates a relatively high probability that the variant is pathogenic (i.e., the model output may indicate that the variant is in the pathogenic range, even though the actual probability may still be less than 50%). For a specific variant 872, the model output indicates a very low probability of pathogenicity. Using a variant classification framework such as Sherlock, in which points are assigned based on evidence of pathogenicity, benign points would not be awarded to variant 870 but would be awarded to variant 872.
[0132] Figure 8E shows another example of population frequency modeling results for a population frequency model configured using a logistic regression model, according to some embodiments of the present disclosure. Figure 8E shows an example of the pathogenicity probability estimated and output for each variant in a given gene, MLH1, by the described population frequency model. In Figure 8E, each dot in the plot represents a different specific variant. The y-axis shows the probability of the variant being pathogenic, and the x-axis shows the allele frequency of the variant in a given human population. As the frequency of the variant in the population increases, the probability of pathogenicity decreases.
[0133] In Figure 8E, the plot is the same as that in Figure 8D, but different variants are highlighted for illustrative purposes. For particular variant 880, the model output indicates that the variant has a high probability of being pathogenic. For particular variant 882, the model output indicates that variant 882 has a relatively low probability of being pathogenic, even though the allele frequencies indicate that variant 882 is relatively rare.
[0134] Additionally, variants 880, 882 in Figure 8E show the vertical spread for variants within a single gene at a given frequency, resulting from the model inputs, region-level and location-level features (e.g., explained gene-level features, region-level features and location-level features) provided to the population frequency model in addition to gene-level features.
[0135] Figure 8E shows that variant 882 at a given allele frequency is unlikely to be pathogenic, while variant 880 at the same allele frequency is much more likely to be pathogenic. Traditionally, it can be understood that different portions of the same gene may be more or less tolerant to mutations, for example, due to protein domains and functional sequence motifs, but this low-probability variant may affect a region of the gene that is less important than a high-probability variant. However, these differences are difficult to address using traditional variant classification techniques. In contrast, even if this domain and motif information is not provided to the population frequency model, the model constructed as described takes these differences into account through a machine learning-based training and calibration process using the described feature set. As a result, the pathogenicity plot generated based on the model output shows that different variants within the same gene should be treated differently with respect to the population frequency threshold.
[0136] Thus, embodiments of the population frequency model can not only model the expected frequency of a disease for a given gene, but can also determine that different variants within the same gene with the same frequency are associated with different probabilities of pathogenicity. Furthermore, embodiments of the model configured as described can identify different allele frequency thresholds for different regions within a gene because the model is trained and calibrated using region-level characteristics within the gene.
[0137] Figure 8F shows an example of population frequency modeling results for a population frequency model configured using a logistic regression model integrated with a variant classification framework according to some embodiments of the present disclosure. In particular, Figure 8F shows how the accuracy of variant pathogenicity predictions (measured as negative and positive predictive values) generated and output by embodiments of the described population frequency model can be mapped to locations within a variant classification system such as Sherlock. While Figure 8F shows an exemplary implementation using the Sherlock framework, the model output can be similarly adapted for incorporation into other variant classification frameworks.
[0138] As shown in Figure 8F, four new lines of evidence were created for population modeling data for use in Sherlock to correspond to the output of the population frequency model, where benign and pathogenic points were assigned to each new line of evidence based on the associated level of accuracy.
[0139] For example, a variant with a very highly predicted benign population modeling score with an NPV (negative predictive value) greater than 99% would be assigned, for example, 5 benign points and classified as a benign variant using the Sherlock framework in the absence of conflicting information from other evidence categories.
[0140] Conversely, a variant with a moderate or highly predicted pathogenic population modeling score with a PPV (positive predictive value) greater than 80% would be assigned, for example, 1 point toward pathogenicity, and would require, for example, 4 additional pathogenicity points from other independent evidence categories to reach the 5-point threshold within the Sherlock framework to be classified as pathogenic.
[0141] Population modeling using the modeling approach described herein can have a significant impact on variant classification. For example, gnomAD population frequency data can now be applied to four times as many variants as could be handled using previous approaches. This expansion of the use of population frequency data allows 15,000 variants from VUS to be reclassified into benign and likely benign variants. For example, 3,146 variants were reclassified from VUS to benign, and 11,640 variants were reclassified from VUS to likely benign.
[0142] Furthermore, population frequency models constructed as described allow for much more efficient use of population frequency data. As a result, many uncertainly significant variants that were not reclassified based on this data alone are now closer to receiving a final classification in the future. The reclassified variants affected 50,000 patients. It is estimated that the described population modeling approach could result in a 2-2.5% reduction in VUS rates for future patients.
[0143] Furthermore, the population modeling approach described was applied to variants from diverse populations and across multiple clinical areas, which at an aggregate level allowed for the identification of subpopulations with the highest frequency of a particular variant and the grouping of reclassified variants by major clinical area.
[0144] FIG. 9 illustrates a method 900 of population frequency modeling using a population frequency model constructed using a logistic regression model, according to some embodiments of the present disclosure.
[0145] Method 900 is performed by processing logic including hardware (e.g., a processing device, circuitry, dedicated logic, programmable logic, microcode, device hardware, integrated circuit, etc.), software (e.g., instructions running or executed on a processing device), or a combination thereof. In some embodiments, the method is performed by components of a computing system, which may include components or flows shown in FIG. 9 that, in some embodiments, may not be specifically shown in other figures, and / or may include components or flows shown in other figures that, in some embodiments, may not be specifically shown in FIG. 9. Although shown in a particular order or sequence, unless otherwise specified, the order of the processes may be modified. Thus, the illustrated embodiments should be understood as examples only, and the illustrated processes may be performed in a different order, and some processes may be performed in parallel. Additionally, at least one process may be omitted in various embodiments. Thus, not all processes are required in all embodiments. Other process flows are possible.
[0146] At operation 902, the processing device applies a logistic regression model to a first set of population data for a first set of genes. The items of the first set of population data include a set of features for variants located at intragenic positions in the first set of genes. The set of features includes at least one gene-level feature, at least one variant-level feature, and at least one population frequency meta-feature. The items of the first set of population data also include a reference label indicating whether the variant is benign or pathogenic.
[0147] In operation 904, the processing device compares, for each item in the first set of population data, the variant classification prediction output by the logistic regression model to the expected variant classification indicated by the reference label.
[0148] At operation 906, the processing device iteratively adjusts the value of at least one parameter or coefficient of the logistic regression model until one or more performance criteria are met, e.g., until the output of a loss function calculated based on the variant classification predictions output by the logistic regression model satisfies at least one first performance criterion (e.g., an error tolerance threshold), to generate a trained logistic regression model (e.g., the model converges). The trained logistic regression model can output variant pathogenicity estimates that satisfy at least one second performance criterion (e.g., one or more validation criteria).
[0149] In some implementations, the processing device uses the trained logistic regression model to generate a prediction regarding whether the variant is benign or pathogenic, hi some implementations, the processing device provides the prediction regarding whether the variant is benign or pathogenic to a clinician's computing device for use by the clinician in determining a diagnosis for the patient.
[0150] In some implementations, the processing device applies the trained logistic regression model to a second set of population data for a plurality of variants of a second plurality of genes. For each variant of the second plurality of genes, the processing device receives a variant classification prediction output by the trained logistic regression model. In response to the variant classification prediction satisfying at least a second performance criterion, the processing device stores the variant classification prediction in association with the variant for retrieval via at least one query.
[0151] In some implementations, the processing device calculates gene-level constraints and includes the gene-level constraints in at least one gene-level feature, calculates allele frequencies and includes the allele frequencies in at least one location-level feature, includes a mathematical combination of the gene-level constraints and the allele frequencies in at least one population frequency meta-feature, and applies a logistic regression model to the set of features including the mathematical combination of the gene-level constraints and the allele frequencies. The trained logistic regression model can output a variant pathogenicity estimate that meets at least one second performance criterion based on the set of features including the mathematical combination of the gene-level constraints and the allele frequencies.
[0152] In some implementations, the processing device calculates the allele frequencies by calculating a binomial ratio of the confidence values associated with the allele frequencies. In some implementations, the processing device calculates a mathematical combination of the gene-level constraints and the allele frequencies by dividing the allele frequencies by the gene-level constraints to generate a quotient and obtain an exponent of the quotient.
[0153] In some implementations, the processing device calculates a synonymous to non-synonymous missense variant ratio index for a gene. The processing device includes the synonymous to non-synonymous missense variant ratio index in at least one population frequency meta-feature. The processing device applies a logistic regression model to the set of features including the synonymous to non-synonymous missense variant ratio index. The trained logistic regression model can output a variant pathogenicity estimate that meets at least one second performance criterion based on the set of features including the synonymous to non-synonymous missense variant ratio index.
[0154] In some implementations, the processing device selects 30 or fewer features as the set of features. The processing device applies a logistic regression model to the selected set of 30 or fewer features. The trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion based on the selected set of 30 or fewer features.
[0155] In some implementations, the processing device calculates a fixation index. The fixation index includes subpopulation frequency data. The processing device includes the fixation index in a set of features. The processing device applies a logistic regression model to the set of features including the fixation index. The trained logistic regression model can output a variant pathogenicity estimate based on the set of features including the fixation index that meets at least one second performance criterion.
[0156] In some implementations, the processing device calculates a mathematical combination of subpopulation frequency data and population frequency data for the variant. The processing device includes the mathematical combination of the subpopulation frequency data and the population frequency data in at least one location-level feature. The processing device applies a logistic regression model to the set of features including the mathematical combination of the subpopulation frequency data and the population frequency data. The trained logistic regression model can output a variant pathogenicity estimate that meets at least one second performance criterion based on the set of features including the mathematical combination of the subpopulation frequency data and the population frequency data.
[0157] In some implementations, iteratively adjusting the value of at least one parameter of the logistic regression model includes adjusting a C value of the logistic regression model until a loss satisfies at least one first performance criterion to generate a trained logistic regression model. In some implementations, the processing device configures the logistic regression model using L1 regularization. In some implementations, the processing device estimates the first performance criterion using a mean squared error or an area based on a receiver operating characteristic curve. In some implementations, the processing device determines the second performance criterion using at least one of a decision boundary test, a feature weight test, or a comparison to a benchmark variant classification.
[0158] In some implementations, the processing device uses the variant classification estimates output by the trained logistic regression model as input to the variant classification framework. In some implementations, the at least one population frequency meta-feature comprises an expected frequency distribution of known benign variants and known pathogenic variants within the gene.
[0159] 10 illustrates an exemplary computing system including a population frequency modeling system according to some embodiments of the present disclosure. In the embodiment of FIG. 10, computing system 1000 includes one or more user systems 1010, a network 1020, an application system 1030, a population frequency modeling system 1050, and a data storage system 1080.
[0160] Population frequency modeling system 1050 includes a logistic regression model 1052, a model development subsystem 1054, model performance criteria 1056, model validation criteria 1058, a training dataset 1060, and a validation dataset 1062. Components of population frequency modeling system 1050 may correspond to similarly described components shown in other figures. For example, logistic regression model 1052 may correspond to logistic regression model 228, logistic regression model 306, or trained logistic regression model 604. Model development subsystem 1054 may correspond to a combination of components including genetic data selection subsystem 206, feature generation subsystem 212, modeling and calibration subsystem 220, and model validation subsystem 230. Model development subsystem 1054 may be configured to perform process 300 and / or process 400 and / or process 600. Model performance criteria 1056 may correspond to model performance criteria 232. The model validation criteria 1058 may correspond to the model validation criteria 236. The training data set 1060 may correspond to the training data set 234. The validation data set 1062 may correspond to the validation data set 238.
[0161] The logistic regression model 1052 includes one or more machine learning models trained to use machine learning algorithms to determine probabilistic or statistical relationships between inputs and outputs. For example, given one or more inputs, the logistic regression model 1052 outputs labels that can be used to classify the inputs into different categories or scores that can be used to sort or rank the inputs into groups or ranked lists. One example of a logistic regression model 1052 is the population frequency model described above.
[0162] The model development subsystem 1054 trains one or more machine learning models, e.g., the logistic regression model 1052, by applying supervised machine learning techniques to training data including training examples and ground truth labels of input data. The predicted outputs of the one or more machine learning models are iteratively observed until a set of model performance criteria is met. For example, the difference between the predicted output and the expected output is quantified using a loss function. The model performance criteria are used to determine when the one or more machine learning models have converged to provide an output that can be relied upon with a degree of accuracy. The required level of accuracy and performance criteria is determined based on the requirements or design of a particular implementation of the one or more machine learning models. An example of the model development subsystem 1054 includes data preparation, feature engineering, and model selection components, as described above.
[0163] In some implementations, the training dataset 1060 includes training data used to train the logistic regression model 1052. The training dataset 1060 includes, for example, a set of input features and corresponding ground truth labels. In some implementations, the training dataset 1060 includes or is derived from a database of historical population data. Examples of the training dataset 1060 include the variant-level dataset, gene-level dataset, and location-level dataset described above.
[0164] The user system 1010 includes at least one computing device, such as a personal computing device, a server, a mobile computing device, or a smart appliance. The user system 1010 includes at least one software application installed on the computing device or accessible to the computing device over a network, the software application including a user interface 1012. For example, an embodiment of the user interface 1012 includes a graphical display screen that displays controls and graphical elements for operating and / or manipulating one or more of the logistic regression model 1052, the model development subsystem 1054, and the training dataset 1060.
[0165] The user interface 1012 can be used to input data, initiate user interface events, and view or otherwise sense output, including pathogenicity predictions and / or other data, generated by the population frequency modeling system 1050. Examples of the user interface 1012 include a web browser, a command line interface, and a mobile app front end. As used herein, the user interface 1012 can include an application programming interface (API). The user interface 1012 can include a front end of an application system 1030 used by a clinician. For example, the output of the population frequency modeling system 1050 can be transmitted to and displayed by the user interface 1012 of a computing device used by the clinician. Alternatively, or in addition, another version of the user interface 1012 can include a front end of an application software system 1030 used by variant scientists and / or other individuals working in the field of genetic testing. As such, the output of population frequency modeling system 1050 may be transmitted to and displayed by user interface 1012 of a computing device used by any of these individuals.
[0166] Application system 1030 is any type of application software system that provides for or enables the generation, display, or manipulation of output generated by population frequency modeling system 1050. Examples of application system 1030 include, but are not limited to, a variant classification system, DNA (deoxyribonucleic acid) analysis software, genetic testing software, medical testing software, healthcare management software, or any combination of any of the foregoing.
[0167] Data storage system 1080 includes data stores and / or data services that store data received, used, manipulated, and generated by application system 1030 and / or population frequency modeling system 1050, such as training data, validation data, machine learning model parameters and coefficients, performance metrics, validation metrics, machine learning model outputs, etc. In FIG. 10 , data storage system 1080 includes one or more data stores that store unlabeled population data 1082, labeled population data 1084, and logistic regression model output 1086. The unlabeled population data may correspond to unlabeled population data 202. The labeled population data 1084 may correspond to labeled population data 204. The logistic regression model output 1086 may correspond to, for example, model output 314, predicted output 414, or model output 606. In some embodiments, data storage system 1080 includes multiple different types of data storage and / or distributed data services. As used herein, a data service may refer to a physical, geographical grouping of machines, a logical grouping of machines, or a single machine. For example, a data service may be a data center, a cluster, a group of clusters, or a machine.
[0168] Data storage system 1080 resides in at least one persistent and / or volatile storage device that may reside within the same local network as at least one other device of computing system 1000 and / or within a remote network relative to at least one other device of computing system 1000. Thus, portions of data storage system 1080, while shown as contained within computing system 1000, may be part of computing system 1000 or may be accessed by computing system 1000 over a network, such as network 1020.
[0169] Although not specifically shown, it should be understood that any of the user system 1010, application system 1030, population frequency modeling system 1050, and data storage system 1080 include interfaces embodied as computer programming code stored in computer memory that, when executed, enable a computing device to communicate bidirectionally using a communication coupling mechanism with any other of the user system 1010, application system 1030, population frequency modeling system 1050, and data storage system 1080. Examples of communication coupling mechanisms include network interfaces, inter-process communication (IPC) interfaces, and application program interfaces (APIs).
[0170] Each of user system 1010, application system 1030, population frequency modeling system 1050, and data storage system 1080 is implemented using at least one computing device communicatively coupled to an electronic communications network 1020. Any of user system 1010, application system 1030, population frequency modeling system 1050, and data storage system 1080 may be bidirectionally coupled by network 1020. User system 1010 and other different user systems (not shown) may be bidirectionally coupled to application system 1030 and / or population frequency modeling system 1050.
[0171] A typical user of user system 1010 may be an administrator or end user of application system 1030 and / or population frequency modeling system 1050. User system 1010 is configured to communicate bidirectionally with application system 1030 and / or population frequency modeling system 1050 via network 1020.
[0172] The features and functionality of user system 1010, application system 1030, population frequency modeling system 1050, and data storage system 1080 may be implemented using computer software, hardware, or software and hardware, and may include a combination of automated functions, data structures, and digital data, as represented schematically in the figures. While user system 1010, application system 1030, population frequency modeling system 1050, and data storage system 1080 are shown as separate elements in FIG. 10 for ease of illustration, this illustration is not intended to suggest that separation of these elements is required, unless otherwise stated. The depicted systems, services, and data stores (or their functionality) of each of user system 1010, application system 1030, population frequency modeling system 1050, and data storage system 1080 may be split across any number of physical systems, including a single physical computer system, and may communicate with each other in any suitable manner.
[0173] Network 1020 may be implemented on any medium or mechanism that provides for the exchange of data, signals, and / or instructions between various components of computing system 1000. Examples of network 1020 include, without limitation, a local area network (LAN), a wide area network (WAN), an Ethernet network, or the Internet, or at least one terrestrial, satellite, or wireless link, or a combination of any number of different networks and / or communication links.
[0174] For ease of explanation, an embodiment of population frequency modeling system 1050 is represented in FIG. 11 as population frequency modeling system 1150.
[0175] 11 is a block diagram of an exemplary computer system on which aspects of the present disclosure may operate. Figure 11 illustrates an exemplary machine, computer system 1100, within which a set of instructions for causing the machine to perform any of the methodologies described herein may be executed. In some embodiments, computer system 1100 may correspond to a component of a networked computer system (e.g., computing system 1000 of FIG. 10 ) coupled to or utilizing a machine running an operating system to perform the operations described above corresponding to aspects of population frequency modeling system 1050 of FIG. 10 .
[0176] The machine may be connected (e.g., networked) to other machines within a local area network (LAN), an intranet, an extranet, and / or the Internet. The machine may operate in the capacity of a server or a client machine in a client-server network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or client machine in a cloud computing infrastructure or environment.
[0177] The machine may be a personal computer (PC), smartphone, tablet PC, set-top box (STB), personal digital assistant (PDA), mobile phone, web appliance, server, or any machine capable of executing a set of instructions (sequentially or otherwise) that specify actions to be taken by the machine. Additionally, although a single machine is shown, the term "machine" shall also be taken to include any collection of machines that individually or jointly execute a set (or sets) of instructions to perform any of the methodologies described herein.
[0178] The exemplary computer system 1100 includes a processing device 1102, a main memory 1104 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM), such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM)), memory 1105 (e.g., flash memory, static random access memory (SRAM), etc.), an input / output system 1110, and a data storage system 1140, which communicate with each other via a bus 1130.
[0179] Processing device 1102 represents at least one general-purpose processing device, such as a microprocessor or central processing unit. More specifically, the processing device may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor implementing another instruction set or a combination of instruction sets. Processing device 1102 may also be at least one special-purpose processing device, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), or a network processor. Processing device 1102 is configured to execute instructions 1112 to perform the operations and steps described herein.
[0180] Instructions 1112 include portions of population frequency modeling system 1150 when those portions of the population frequency modeling system are being executed by processing device 1102. Accordingly, the population frequency modeling system is shown with dashed lines as part of instructions 1112 to illustrate that portions of the population frequency modeling system are potentially executed by processing device 1102. For example, when at least some portions of the population frequency modeling system are embodied in instructions for causing processing device 1102 to perform the methods described above, some of those instructions may be read into processing device 1102 (e.g., to an internal cache or other memory) from main memory 1104 and / or data storage system 1140. However, it is not required that all of the population frequency modeling system be included in instructions 1112 at the same time, and portions of the population frequency modeling system may be stored in at least one other component of computer system 1100 at other times, e.g., when at least one portion of the population frequency modeling system is not being executed by processing device 1102.
[0181] Computer system 1100 further includes a network interface device 1108 for communicating over network 1120. The network interface device 1108 provides a two-way data communication coupling to the network. For example, network interface device 1108 may be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem for providing a data communication connection to a corresponding type of telephone line. As another example, network interface device 1108 may be a local area network (LAN) card that provides a data communication connection to a compatible LAN. A wireless link may also be implemented. In any such implementation, network interface device 1108 can send and receive electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.
[0182] The network link may provide data communication through at least one network to other data devices. For example, the network link may provide a connection to the world wide packet data communication network commonly referred to as the “Internet,” for example, through a local network to a host computer or to data equipment operated by an Internet Service Provider (ISP). Local networks and the Internet use electrical, electromagnetic or optical signals that carry digital data to and from the computer system 1100.
[0183] Computer system 1100 can send messages containing program code and receive data containing program code through the network and network interface device 1108. In the Internet example, a server can transmit requested code for an application program through the Internet and network interface device 1108. The received code can be executed by processing device 1102 as it is received and / or stored in data storage system 1140 or other non-volatile storage for later execution.
[0184] The input / output system 1110 includes output devices, such as a display, e.g., a liquid crystal display (LCD) or touchscreen display, for displaying information to a computer user, or speakers, tactile devices, or another form of output device. The input / output system 1110 may include input devices, e.g., alphanumeric and other keys configured to communicate information and command selections to the processing device 1102. The input devices may alternatively or additionally include cursor controls, such as a mouse, trackball, or cursor direction keys, for communicating directional information and command selections to the processing device 1102 and for controlling cursor movement on the display. The input devices may alternatively or additionally include a microphone, sensor, or sensor array for communicating sensed information to the processing device 1102. The sensed information may include, for example, voice commands, audio signals, geographic location information, and / or digital images.
[0185] Data storage system 1140 includes machine-readable storage medium 1142 (also known as computer-readable medium) on which at least one set of instructions 1144 or software embodying any of the methodologies or functions described herein is stored. Additionally, instructions 1144 may reside, completely or at least partially, within main memory 1104 and / or within processing device 1102 during their execution by computer system 1100, main memory 1104 and processing device 1102 also constituting the machine-readable storage medium.
[0186] In one embodiment, instructions 1144 include instructions for implementing functionality corresponding to a population frequency modeling system (e.g., population frequency modeling system 1050 of FIG. 10).
[0187] 11 is used to indicate that the population frequency modeling system need not be entirely embodied in instructions 1112, 1114, and 1144 simultaneously. In one example, portions of the population frequency modeling system are embodied in instructions 1144, which are read into main memory 1104 as instructions 1114, and portions of instructions 1114 are read into processing device 1102 as instructions 1112 for execution. In another example, some portions of the population frequency modeling system are embodied in instructions 1144, while other portions are embodied in instructions 1114, and still other portions are embodied in instructions 1112.
[0188] While the exemplary embodiment shows machine-readable storage medium 1142 to be a single medium, the term "machine-readable storage medium" should be interpreted to include a single medium or multiple media that store at least one set of instructions. The term "machine-readable storage medium" should also be interpreted to include any medium that can store or encode a set of instructions for execution by the machine and cause the machine to perform any of the methodologies of the present disclosure. Accordingly, the term "machine-readable storage medium" should be interpreted to include, but is not limited to, solid-state memory, optical media, and magnetic media.
[0189] Some portions of the foregoing detailed descriptions have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the form used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of operations leading to a desired result. These operations require physical manipulations of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.
[0190] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. This disclosure may refer to the actions and processes of a computer system or similar electronic computing device that manipulate and transform data represented as physical (electronic) quantities in the computer system's registers and memory into other data that is also represented as physical quantities in the computer system's memory or registers or other such information storage systems.
[0191] The present disclosure also relates to apparatus for performing the operations herein. This apparatus may be specially constructed for an intended purpose, or may include a general-purpose computer selectively activated or reconfigured by a computer program stored on the computer. A computer system or other data processing system, such as computing system 1100, can perform the techniques described above in response to its processor executing a computer program (e.g., a sequence of instructions) contained in a memory or other non-transitory machine-readable storage medium. Such a computer program may be stored on a computer-readable storage medium, such as, but not limited to, any type of disk, including floppy disks, optical disks, CD-ROMs, and magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.
[0192] The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus to perform the methods. The structure of a variety of these systems appears as set forth in the description that follows. In addition, the present disclosure is described without reference to any particular programming language. It will be understood that a variety of programming languages can be used to implement the teachings of the present disclosure as described herein.
[0193] The present disclosure may be provided as a computer program product or software, which may include a machine-readable medium having stored thereon instructions that can be used to program a computer system (or other electronic device) to perform a process according to the present disclosure. A machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). In some embodiments, the machine-readable (e.g., computer-readable) medium includes a machine (e.g., computer) readable storage medium, such as, for example, read-only memory ("ROM"), random access memory ("RAM"), magnetic disk storage media, optical storage media, flash memory components, etc.
[0194] Exemplary aspects of the technology disclosed herein are provided below: An embodiment of the technology may include any of the aspects described herein, or any combination of any of the aspects described herein, or any combination of portions of any of the aspects described herein.
[0195] In some aspects, the techniques described herein include a method of configuring a machine learning model to model population frequencies for variant classification, the method comprising: applying a logistic regression model to a first set of population data for a first set of genes, wherein entries of the first set of population data include, for variants located at intragenic locations of the first set of genes, a set of features including at least one gene-level feature, at least one variant-level feature, and at least one population frequency meta-feature, and a reference label indicating whether the variant is benign or pathogenic, wherein the at least one population frequency meta-feature is a predictor of an allele frequency within the genes. evaluating, for each item in the first set of population data, a variant classification prediction output by the logistic regression model based on an expected variant classification indicated by the reference label; and iteratively adjusting the value of at least one parameter or coefficient of the logistic regression model until an output of a loss function calculated based on the variant classification prediction output by the logistic regression model satisfies at least one first performance criterion to generate a trained logistic regression model, wherein the trained logistic regression model is capable of outputting variant pathogenicity estimates that satisfy at least one second performance criterion.
[0196] In some aspects, the techniques described herein relate to methods further comprising using the trained logistic regression model to generate a prediction regarding whether a variant is benign or pathogenic.
[0197] In some aspects, the techniques described herein relate to methods further comprising providing said prediction as to whether said variant is benign or pathogenic to a clinician for use in the clinician's determination of a patient's diagnosis.
[0198] In some aspects, the techniques described herein relate to methods further comprising applying the trained logistic regression model to a second set of population data for a plurality of variants of a second plurality of genes; and receiving a variant classification prediction output by the trained logistic regression model for each variant of the second plurality of genes, and storing the variant classification prediction in association with the variant for retrieval via at least one query in response to the variant classification prediction satisfying at least the second performance criterion.
[0199] In some aspects, the techniques described herein relate to methods further comprising: calculating gene-level constraints; including the gene-level constraints in the at least one gene-level feature; calculating allele frequencies; including the allele frequencies in the at least one variant-level feature; including a mathematical combination of the gene-level constraints and allele frequencies in the at least one population frequency meta-feature; and applying the logistic regression model to the set of features comprising the mathematical combination of the gene-level constraints and the allele frequencies, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the mathematical combination of the gene-level constraints and the allele frequencies.
[0200] In some aspects, the techniques described herein relate to methods further comprising calculating the allele frequencies by calculating a binomial ratio of confidence values associated with the allele frequencies.
[0201] In some aspects, the techniques described herein relate to methods further comprising calculating the mathematical combination of the gene-level constraints and the allele frequencies by dividing the allele frequencies by the gene-level constraints to generate a quotient, and obtaining an exponent of the quotient.
[0202] In some aspects, the techniques described herein relate to methods further comprising: calculating an index of the ratio of synonymous to non-synonymous missense variants for the gene; including the index of the ratio of synonymous to non-synonymous missense variants in the at least one population frequency meta-feature; and applying the logistic regression model to the set of features including the index of the ratio of synonymous to non-synonymous missense variants, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features including the index of the ratio of synonymous to non-synonymous missense variants.
[0203] In some aspects, the techniques described herein relate to methods further comprising: selecting 30 or fewer features as the set of features; and applying the logistic regression model to the selected set of 30 or fewer features, wherein the trained logistic regression model is capable of outputting variant pathogenicity estimates that meet the at least one second performance criterion based on the selected set of 30 or fewer features.
[0204] In some aspects, the techniques described herein relate to methods further comprising: calculating a fixation index, wherein the fixation index comprises subpopulation frequency data; including the fixation index in the set of features; and applying the logistic regression model to the set of features comprising the fixation index, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the fixation index.
[0205] In some aspects, the techniques described herein relate to methods further comprising: calculating a mathematical combination of subpopulation and population frequency data for a variant; including the mathematical combination of the subpopulation and population frequency data in the at least one variant-level feature; and applying the logistic regression model to the set of features comprising the mathematical combination of the subpopulation and population frequency data, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the mathematical combination of the subpopulation and population frequency data.
[0206] In some aspects, the techniques described herein relate to methods in which iteratively adjusting the value of at least one parameter of the logistic regression model comprises adjusting a C-value of the logistic regression model until the output of the loss function satisfies the at least one first performance criterion to generate the trained logistic regression model.
[0207] In some embodiments, the techniques described herein include configuring the logistic regression model to model population frequencies for variant classification using L1 regularization. The present invention relates to a method further comprising:
[0208] In some aspects, the techniques described herein relate to methods further comprising estimating the first performance criterion using a mean squared error or an area based on the receiver operating characteristic curve.
[0209] In some aspects, the techniques described herein relate to methods further comprising determining the second performance criterion using at least one of a decision boundary test, a feature weight test, or a comparison to a benchmark variant classification.
[0210] In some aspects, the techniques described herein relate to methods further comprising using the variant classification estimates output by the trained logistic regression model as input to a variant classification framework.
[0211] In some aspects, the technology described herein relates to a method, wherein the at least one population frequency meta-feature comprises an expected frequency distribution of known benign variants and known pathogenic variants in the gene.
[0212] In some aspects, the techniques described herein comprise at least one processor; and at least one memory coupled to the at least one processor, wherein the at least one memory includes a procedure that, when executed by the at least one processor, applies a logistic regression model to a first set of population data for a first set of genes, wherein entries of the first set of population data include, for variants located at intragenic positions of the first set of genes, a set of features including at least one gene-level feature, at least one variant-level feature, and at least one population frequency meta-feature, and a reference label indicating whether the variant is benign or pathogenic, wherein the at least one wherein the population frequency metafeature quantifies the predictive value of the allele frequency within the gene; for each item in the first set of population data, evaluating a variant classification prediction output by the logistic regression model based on the expected variant classification indicated by the reference label; and adjusting the value of at least one parameter or coefficient of the logistic regression model until at least one first performance criterion is met to generate a trained logistic regression model, wherein the trained logistic regression model is capable of outputting variant pathogenicity estimates that satisfy at least one second performance criterion.
[0213] In some aspects, the technology described herein relates to a system wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including using the trained logistic regression model to generate a prediction regarding whether the variant is benign or pathogenic.
[0214] In some aspects, the technology described herein relates to a system wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including providing the prediction as to whether the variant is benign or pathogenic to a clinician for use in the clinician's determination of a patient's diagnosis.
[0215] In some aspects, the technology described herein relates to a system wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including: applying the trained logistic regression model to a second set of population data for a plurality of variants of a second plurality of genes; and receiving, for each variant of the second plurality of genes, a variant classification prediction output by the trained logistic regression model, and, in response to the variant classification prediction satisfying at least the second performance criterion, storing the variant classification prediction in association with the variant for retrieval via at least one query.
[0216] and applying the logistic regression model to the set of features comprising the mathematical combination of the gene-level constraints and the allele frequencies, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the mathematical combination of the gene-level constraints and the allele frequencies.
[0217] In some aspects, the technology described herein relates to a system wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including calculating the allele frequencies by calculating a binomial ratio of confidence values associated with the allele frequencies.
[0218] In some aspects, the technology described herein relates to a system, wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including calculating the mathematical combination of the gene-level constraints and the allele frequencies by dividing the allele frequencies by the gene-level constraints to generate a quotient and obtaining an exponent of the quotient.
[0219] In some aspects, the technology described herein relates to a system wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including: calculating an index of the ratio of synonymous to non-synonymous missense variants for the gene; including the index of the ratio of synonymous to non-synonymous missense variants in the at least one population frequency metafeature; and applying the logistic regression model to the set of features including the index of the ratio of synonymous to non-synonymous missense variants, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features including the index of the ratio of synonymous to non-synonymous missense variants.
[0220] In some aspects, the technology described herein relates to a system wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including: selecting 30 or fewer features as the set of features; and applying the logistic regression model to the selected set of 30 or fewer features, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the selected set of 30 or fewer features.
[0221] In some aspects, the techniques described herein relate to systems wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including: calculating a fixation index, wherein the fixation index comprises subpopulation frequency data; including the fixation index in the set of features; and applying the logistic regression model to the set of features comprising the fixation index, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the fixation index.
[0222] In some aspects, the technology described herein relates to a system wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including: calculating a mathematical combination of subpopulation frequency data and population frequency data for a variant; including the mathematical combination of the subpopulation frequency data and population frequency data in the at least one variant-level feature; and applying the logistic regression model to the set of features comprising the mathematical combination of the subpopulation frequency data and population frequency data, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the mathematical combination of the subpopulation frequency data and population frequency data.
[0223] In some aspects, the technology described herein relates to a system, wherein iteratively adjusting the value of at least one parameter of the logistic regression model comprises adjusting a C-value of the logistic regression model until the output of the loss function satisfies the at least one first performance criterion to generate the trained logistic regression model.
[0224] In some aspects, the technology described herein relates to a system, wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including configuring the logistic regression model to model population frequencies for variant classification using L1 regularization.
[0225] In some aspects, the techniques described herein relate to a system wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including estimating the first performance criterion using a mean squared error or an area based on the receiver operating characteristic curve.
[0226] In some aspects, the techniques described herein relate to a system wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including determining the second performance criterion using at least one of decision boundary checking, feature weight checking, or comparison to a benchmark variant classification.
[0227] In some aspects, the techniques described herein relate to a system wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including using a variant classification estimate output by the trained logistic regression model as input to a variant classification framework.
[0228] In some aspects, the technology described herein relates to a system, wherein the at least one population frequency meta-feature comprises an expected frequency distribution of known benign variants and known pathogenic variants in the gene.
[0229] In some aspects, the techniques described herein include a procedure that, when executed by at least one processor, applies a logistic regression model to a first set of population data for a first set of genes, wherein items of the first set of population data include a set of features, including at least one population frequency metafeature, for variants located at positions within genes of the first set of genes, and a reference label indicating whether the variant is benign or pathogenic, wherein the at least one population frequency metafeature quantifies a predictive value of an allele frequency within the gene; evaluating a variant classification prediction output by the logistic regression model based on an expected variant classification indicated by the reference label; and adjusting the value of at least one parameter or coefficient of the logistic regression model until at least one first performance criterion is met to generate a trained logistic regression model, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that satisfies at least one second performance criterion.
[0230] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including using the trained logistic regression model to generate a prediction regarding whether the variant is benign or pathogenic.
[0231] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including providing the prediction as to whether the variant is benign or pathogenic to a clinician for use in the clinician's determination of a patient's diagnosis.
[0232] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including: applying the trained logistic regression model to a second set of population data for a plurality of variants of a second plurality of genes; and receiving, for each variant of the second plurality of genes, a variant classification prediction output by the trained logistic regression model, and, in response to the variant classification prediction satisfying at least the second performance criterion, storing the variant classification prediction in association with the variant for retrieval via at least one query.
[0233] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including: calculating gene-level constraints; calculating allele frequencies; including a mathematical combination of the gene-level constraints and the allele frequencies in the at least one population frequency meta-feature; and applying the logistic regression model to the set of features comprising the mathematical combination of the gene-level constraints and the allele frequencies, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the mathematical combination of the gene-level constraints and the allele frequencies.
[0234] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including calculating the allele frequencies by calculating a binomial ratio of confidence values associated with the allele frequencies.
[0235] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including calculating the mathematical combination of the gene-level constraints and the allele frequencies by dividing the allele frequencies by the gene-level constraints to generate a quotient, and obtaining an exponent of the quotient.
[0236] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including: calculating an index of the ratio of synonymous to non-synonymous missense variants for the gene; including the index of the ratio of synonymous to non-synonymous missense variants in the at least one population frequency metafeature; and applying the logistic regression model to the set of features including the index of the ratio of synonymous to non-synonymous missense variants, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features including the index of the ratio of synonymous to non-synonymous missense variants.
[0237] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including: selecting 30 or fewer features as the set of features; and applying the logistic regression model to the selected set of 30 or fewer features, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the selected set of 30 or fewer features.
[0238] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including: calculating a fixation index, wherein the fixation index comprises subpopulation frequency data; including the fixation index in the set of features; and applying the logistic regression model to the set of features comprising the fixation index, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the fixation index.
[0239] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including: calculating a mathematical combination of subpopulation frequency data and population frequency data for a variant; and applying the logistic regression model to the set of features comprising the mathematical combination of the subpopulation frequency data and population frequency data, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the mathematical combination of the subpopulation frequency data and population frequency data.
[0240] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium, wherein adjusting the value of at least one parameter of the logistic regression model comprises adjusting a C-value of the logistic regression model until an output of a loss function satisfies the at least one first performance criterion to generate the trained logistic regression model.
[0241] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including configuring the logistic regression model to model population frequencies for variant classification using L1 regularization.
[0242] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including estimating the at least one first performance criterion using a mean squared error or an area based on the receiver operating characteristic curve.
[0243] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including determining the second performance criterion using at least one of decision boundary checking, feature weight checking, or comparison to a benchmark variant classification.
[0244] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation including using a variant classification estimate output by the trained logistic regression model as input to a variant classification framework.
[0245] In some aspects, the technology described herein relates to at least one non-transitory machine-readable medium, wherein the at least one population frequency metafeature comprises an expected frequency distribution of known benign variants and known pathogenic variants in the gene.
[0246] In some aspects, the technology described herein relates to any one or more of the aspects, steps, components, elements, processes, or limitations described in the accompanying description or illustrated in the accompanying drawings.
[0247] In the foregoing specification, embodiments of the present disclosure have been described with reference to specific exemplary embodiments thereof. It will be apparent that various modifications may be made thereto without departing from the broader spirit and scope of the embodiments of the present disclosure as set forth in the following claims. The specification and drawings are, therefore, to be regarded in an illustrative rather than a restrictive sense. (Other possible items) (Item 1) 1. A method of configuring a machine learning model to model population frequencies for variant classification, comprising: applying a logistic regression model to a first set of population data for a first set of genes, wherein entries of the first set of population data include a set of features for variants located at intragenic locations of the first set of genes, the set including at least one gene-level feature, at least one variant-level feature, and at least one population frequency meta-feature, and a reference label indicating whether the variant is benign or pathogenic, wherein the at least one population frequency meta-feature quantifies a predictive value of an allele frequency within the genes; evaluating, for each item in the first set of population data, a variant classification prediction output by the logistic regression model based on an expected variant classification indicated by the reference label; and iteratively adjusting the value of at least one parameter or coefficient of the logistic regression model until an output of a loss function calculated based on the variant classification predictions output by the logistic regression model satisfies at least one first performance criterion to generate a trained logistic regression model, wherein the trained logistic regression model is capable of outputting variant pathogenicity estimates that satisfy at least one second performance criterion. A method comprising: (Item 2) using the trained logistic regression model to generate a prediction as to whether the variant is benign or pathogenic. The method of claim 1, further comprising: (Item 3) providing said prediction as to whether said variant is benign or pathogenic to a clinician for use in the clinician's determination of a patient's diagnosis. The method of item 2 further comprises: (Item 4) applying the trained logistic regression model to a second set of population data for a plurality of variants of a second plurality of genes; and receiving a variant classification prediction output by the trained logistic regression model for each variant of the second plurality of genes, and storing the variant classification prediction in association with the variant for retrieval via at least one query in response to the variant classification prediction satisfying at least the second performance criterion; The method of claim 1, further comprising: (Item 5) A stage for computing gene-level constraints; including said gene-level constraint in said at least one gene-level feature; calculating allele frequencies; including said allele frequencies in said at least one variant-level signature; including a mathematical combination of the gene-level constraints and allele frequencies in the at least one population frequency metafeature; and applying the logistic regression model to the set of features comprising the mathematical combination of the gene-level constraints and the allele frequencies, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the mathematical combination of the gene-level constraints and the allele frequencies. The method of claim 1, further comprising: (Item 6) calculating the allele frequencies by calculating a binomial ratio of confidence values associated with the allele frequencies; Item 6. The method of item 5, further comprising: (Item 7) calculating the mathematical combination of the gene-level constraints and the allele frequencies by dividing the allele frequencies by the gene-level constraints to generate a quotient, and obtaining an exponent of the quotient; Item 6. The method of item 5, further comprising: (Item 8) calculating an index of the ratio of synonymous to non-synonymous missense variants for said gene; including the index of the ratio of synonymous to non-synonymous missense variants in the at least one population frequency metacharacteristic; and applying the logistic regression model to the set of features including the index of the ratio of synonymous to non-synonymous missense variants, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features including the index of the ratio of synonymous to non-synonymous missense variants. The method of claim 1, further comprising: (Item 9) selecting 30 or fewer features as the set of features; and applying the logistic regression model to the selected set of 30 or fewer features, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the selected set of 30 or fewer features. The method of claim 1, further comprising: (Item 10) calculating a fixation index, wherein said fixation index comprises subpopulation frequency data; including the fixed index in the set of features; and applying the logistic regression model to the set of features including the fixation index, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features including the fixation index. The method of claim 1, further comprising: (Item 11) calculating a mathematical combination of the subpopulation frequency data and the population frequency data for the variants; including a mathematical combination of the subpopulation frequency data and the population frequency data in the at least one variant-level feature; and applying the logistic regression model to the set of features comprising a mathematical combination of the subpopulation frequency data and the population frequency data, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising a mathematical combination of the subpopulation frequency data and the population frequency data. The method of claim 1, further comprising: (Item 12) 2. The method of claim 1, wherein iteratively adjusting the value of at least one parameter of the logistic regression model comprises adjusting a C value of the logistic regression model until the output of the loss function satisfies the at least one first performance criterion to generate the trained logistic regression model. (Item 13) configuring the logistic regression model to model population frequencies for variant classification using L1 regularization; The method of claim 1, further comprising: (Item 14) estimating the first performance criterion using a mean squared error or an area based on the receiver operating characteristic curve. The method of claim 1, further comprising: (Item 15) determining the second performance criterion using at least one of a decision boundary test, a feature weight test, or a comparison to a benchmark variant classification; The method of claim 1, further comprising: (Item 16) using the variant classification estimates output by the trained logistic regression model as input to a variant classification framework. The method of claim 1, further comprising: (Item 17) 2. The method of claim 1, wherein the at least one population frequency meta-feature comprises an expected frequency distribution of known benign variants and known pathogenic variants in the gene. (Item 18) at least one processor; and at least one memory coupled to the at least one processor; Equipped with wherein the at least one memory, when executed by the at least one processor, applying a logistic regression model to a first set of population data for a first set of genes, wherein items of the first set of population data include a set of features for variants located at intragenic locations of the first set of genes, the set including at least one gene-level feature, at least one variant-level feature, and at least one population frequency meta-feature, and a reference label indicating whether the variant is benign or pathogenic, wherein the at least one population frequency meta-feature quantifies a predictive value of an allele frequency within the genes; evaluating, for each item in the first set of population data, a variant classification prediction output by the logistic regression model based on an expected variant classification indicated by the reference label; and adjusting the value of at least one parameter or coefficient of the logistic regression model until at least one first performance criterion is met to generate a trained logistic regression model, wherein the trained logistic regression model is capable of outputting variant pathogenicity estimates that satisfy at least one second performance criterion. at least one instruction that causes the at least one processor to perform at least one operation including system. (Item 19) The at least one memory When executed by the at least one processor, using the trained logistic regression model to generate a prediction as to whether the variant is benign or pathogenic. at least one instruction that causes the at least one processor to perform at least one operation including further comprising Item 19. The system according to item 18. (Item 20) The at least one memory When executed by the at least one processor, providing said prediction as to whether said variant is benign or pathogenic to a clinician for use in the clinician's determination of a patient's diagnosis. at least one instruction that causes the at least one processor to perform at least one operation including further comprising Item 19. The system of item 19. (Item 21) The at least one memory When executed by the at least one processor, applying the trained logistic regression model to a second set of population data for a plurality of variants of a second plurality of genes; and receiving a variant classification prediction output by the trained logistic regression model for each variant of the second plurality of genes, and storing the variant classification prediction in association with the variant for retrieval via at least one query in response to the variant classification prediction satisfying at least the second performance criterion. at least one instruction that causes the at least one processor to perform at least one operation including further comprising Item 19. The system according to item 18. (Item 22) The at least one memory When executed by the at least one processor, Procedures for computing gene-level constraints; including said gene-level constraint on said at least one gene-level feature; Procedures for calculating allele frequencies; including said allele frequencies in said at least one variant-level feature; including a mathematical combination of the gene-level constraints and allele frequencies in the at least one population frequency metafeature; and applying the logistic regression model to the set of features comprising the mathematical combination of the gene-level constraints and the allele frequencies, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the mathematical combination of the gene-level constraints and the allele frequencies. at least one instruction that causes the at least one processor to perform at least one operation including further comprising Item 19. The system according to item 18. (Item 23) The at least one memory When executed by the at least one processor, calculating said allele frequencies by calculating a binomial ratio of confidence values associated with said allele frequencies; at least one instruction that causes the at least one processor to perform at least one operation including further comprising Item 23. The system according to item 22. (Item 24) The at least one memory When executed by the at least one processor, calculating the mathematical combination of the gene-level constraints and the allele frequencies by dividing the allele frequencies by the gene-level constraints to generate a quotient and taking an exponent of the quotient; at least one instruction that causes the at least one processor to perform at least one operation including further comprising Item 23. The system according to item 22. (Item 25) The at least one memory When executed by the at least one processor, calculating an index of the ratio of synonymous to non-synonymous missense variants for said gene; including the index of the ratio of synonymous to non-synonymous missense variants in the at least one population frequency metacharacteristic; and applying the logistic regression model to the set of features including the index of the ratio of synonymous to non-synonymous missense variants, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features including the index of the ratio of synonymous to non-synonymous missense variants. at least one instruction that causes the at least one processor to perform at least one operation including further comprising Item 19. The system according to item 18. (Item 26) The at least one memory When executed by the at least one processor, selecting 30 or fewer features as the set of features; and applying the logistic regression model to the selected set of 30 or fewer features, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the selected set of 30 or fewer features. at least one instruction that causes the at least one processor to perform at least one operation including further comprising Item 19. The system according to item 18. (Item 27) The at least one memory When executed by the at least one processor, calculating a fixation index, wherein said fixation index comprises subpopulation frequency data; including said fixed index in said set of features; and applying the logistic regression model to the set of features including the fixation index, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features including the fixation index. at least one instruction that causes the at least one processor to perform at least one operation including further comprising Item 19. The system according to item 18. (Item 28) The at least one memory When executed by the at least one processor, A procedure for calculating the mathematical combination of subpopulation and population frequency data for variants; including a mathematical combination of the subpopulation frequency data and the population frequency data in the at least one variant-level feature; and applying the logistic regression model to the set of features comprising the mathematical combination of the subpopulation frequency data and the population frequency data, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the mathematical combination of the subpopulation frequency data and the population frequency data. at least one instruction that causes the at least one processor to perform at least one operation including further comprising Item 19. The system according to item 18. (Item 29) Item 19. The system of item 18, wherein iteratively adjusting the value of at least one parameter of the logistic regression model comprises adjusting a C value of the logistic regression model until the output of the loss function satisfies the at least one first performance criterion to generate the trained logistic regression model. (Item 30) The at least one memory When executed by the at least one processor, Configuring the logistic regression model to model population frequencies for variant classification using L1 regularization at least one instruction that causes the at least one processor to perform at least one operation including further comprising Item 19. The system according to item 18. (Item 31) The at least one memory When executed by the at least one processor, estimating the first performance criterion using a mean squared error or an area based on the receiver operating characteristic curve; at least one instruction that causes the at least one processor to perform at least one operation including further comprising Item 19. The system according to item 18. (Item 32) The at least one memory When executed by the at least one processor, determining the second performance criterion using at least one of a decision boundary test, a feature weight test, or a comparison to a benchmark variant classification; at least one instruction that causes the at least one processor to perform at least one operation including further comprising Item 19. The system according to item 18. (Item 33) The at least one memory When executed by the at least one processor, using the variant classification estimates output by the trained logistic regression model as input to a variant classification framework. at least one instruction that causes the at least one processor to perform at least one operation including further comprising Item 19. The system according to item 18. (Item 34) 19. The system of claim 18, wherein the at least one population frequency meta-feature comprises an expected frequency distribution of known benign variants and known pathogenic variants in the gene. (Item 35) When executed by at least one processor, applying a logistic regression model to a first set of population data for a first set of genes, wherein entries of the first set of population data include a set of features for variants located at intragenic locations of the first set of genes, the set of features including at least one population frequency metafeature, and a reference label indicating whether the variant is benign or pathogenic, wherein the at least one population frequency metafeature quantifies a predictive value of an allele frequency within the genes; evaluating, for each item in the first set of population data, a variant classification prediction output by the logistic regression model based on an expected variant classification indicated by the reference label; and adjusting the value of at least one parameter or coefficient of the logistic regression model until at least one first performance criterion is met to generate a trained logistic regression model, wherein the trained logistic regression model is capable of outputting variant pathogenicity estimates that satisfy at least one second performance criterion. and at least one non-transitory machine-readable medium comprising at least one instruction that causes the at least one processor to perform operations that include: (Item 36) When executed by the at least one processor, using the trained logistic regression model to generate a prediction as to whether the variant is benign or pathogenic. Item 36. The at least one non-transitory machine-readable medium of item 35, further comprising at least one instruction that causes the at least one processor to perform at least one operation including: (Item 37) When executed by the at least one processor, providing said prediction as to whether said variant is benign or pathogenic to a clinician for use in the clinician's determination of a patient's diagnosis. 37. The at least one non-transitory machine-readable medium of claim 36, further comprising at least one instruction that causes the at least one processor to perform at least one operation including: (Item 38) When executed by the at least one processor, applying the trained logistic regression model to a second set of population data for a plurality of variants of a second plurality of genes; and receiving a variant classification prediction output by the trained logistic regression model for each variant of the second plurality of genes, and storing the variant classification prediction in association with the variant for retrieval via at least one query in response to the variant classification prediction satisfying at least the second performance criterion. Item 36. The at least one non-transitory machine-readable medium of item 35, further comprising at least one instruction that causes the at least one processor to perform at least one operation including: (Item 39) When executed by the at least one processor, Procedures for computing gene-level constraints; Procedures for calculating allele frequencies; including a mathematical combination of the gene-level constraints and the allele frequencies in the at least one population frequency metafeature; and applying the logistic regression model to the set of features comprising the mathematical combination of the gene-level constraints and the allele frequencies, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the mathematical combination of the gene-level constraints and the allele frequencies. Item 36. The at least one non-transitory machine-readable medium of item 35, further comprising at least one instruction that causes the at least one processor to perform at least one operation including: (Item 40) When executed by the at least one processor, calculating said allele frequencies by calculating a binomial ratio of confidence values associated with said allele frequencies; 40. The at least one non-transitory machine-readable medium of claim 39, further comprising at least one instruction that causes the at least one processor to perform at least one operation including: (Item 41) When executed by the at least one processor, calculating the mathematical combination of the gene-level constraints and the allele frequencies by dividing the allele frequencies by the gene-level constraints to generate a quotient, and obtaining an exponent of the quotient; 40. The at least one non-transitory machine-readable medium of claim 39, further comprising at least one instruction that causes the at least one processor to perform at least one operation including: (Item 42) When executed by the at least one processor, calculating an index of the ratio of synonymous to non-synonymous missense variants for said gene; including the index of the ratio of synonymous to non-synonymous missense variants in the at least one population frequency metacharacteristic; and applying the logistic regression model to the set of features including the index of the ratio of synonymous to non-synonymous missense variants, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features including the index of the ratio of synonymous to non-synonymous missense variants. Item 36. The at least one non-transitory machine-readable medium of item 35, further comprising at least one instruction that causes the at least one processor to perform at least one operation including: (Item 43) When executed by the at least one processor, selecting 30 or fewer features as the set of features; and applying the logistic regression model to the selected set of 30 or fewer features, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the selected set of 30 or fewer features. Item 36. The at least one non-transitory machine-readable medium of item 35, further comprising at least one instruction that causes the at least one processor to perform at least one operation including: (Item 44) When executed by the at least one processor, calculating a fixation index, wherein said fixation index comprises subpopulation frequency data; including said fixed index in said set of features; and applying the logistic regression model to the set of features including the fixation index, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features including the fixation index. Item 36. The at least one non-transitory machine-readable medium of item 35, further comprising at least one instruction that causes the at least one processor to perform at least one operation including: (Item 45) When executed by the at least one processor, A procedure for calculating a mathematical combination of subpopulation frequency data and population frequency data for variants; and applying the logistic regression model to the set of features comprising the mathematical combination of the subpopulation frequency data and the population frequency data, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the mathematical combination of the subpopulation frequency data and the population frequency data. Item 36. The at least one non-transitory machine-readable medium of item 35, further comprising at least one instruction that causes the at least one processor to perform at least one operation including: (Item 46) Item 36. The at least one non-transitory machine-readable medium of Item 35, wherein adjusting the value of at least one parameter of the logistic regression model comprises adjusting a C value of the logistic regression model until an output of a loss function satisfies the at least one first performance criterion to generate the trained logistic regression model. (Item 47) When executed by the at least one processor, Configuring the logistic regression model to model population frequencies for variant classification using L1 regularization Item 36. The at least one non-transitory machine-readable medium of item 35, further comprising at least one instruction that causes the at least one processor to perform at least one operation including: (Item 48) When executed by the at least one processor, estimating said at least one first performance criterion using a mean squared error or an area based on said receiver operating characteristic curve. Item 36. The at least one non-transitory machine-readable medium of item 35, further comprising at least one instruction that causes the at least one processor to perform at least one operation including: (Item 49) When executed by the at least one processor, determining the second performance criterion using at least one of a decision boundary test, a feature weight test, or a comparison to a benchmark variant classification; Item 36. The at least one non-transitory machine-readable medium of item 35, further comprising at least one instruction that causes the at least one processor to perform at least one operation including: (Item 50) When executed by the at least one processor, using the variant classification estimates output by the trained logistic regression model as input to a variant classification framework. Item 36. The at least one non-transitory machine-readable medium of item 35, further comprising at least one instruction that causes the at least one processor to perform at least one operation including: (Item 51) 36. The at least one non-transitory machine-readable medium of claim 35, wherein the at least one population frequency meta-feature comprises an expected frequency distribution of known benign variants and known pathogenic variants in the gene.
Claims
1. 1. A method of configuring a machine learning model to model population frequencies for variant classification, comprising: applying a logistic regression model to a first set of population data for a first set of genes, wherein the items of the first set of population data include a set of features for variants located at intragenic locations of the first set of genes, the set including at least one gene-level feature, at least one variant-level feature, and at least one population frequency meta-feature, and a reference label indicating whether the variant is benign or pathogenic, wherein the at least one population frequency meta-feature quantifies a predictive value of an allele frequency within the genes; evaluating, for each item in the first set of population data, a variant classification prediction output by the logistic regression model based on an expected variant classification indicated by the reference label; and iteratively adjusting the value of at least one parameter or coefficient of the logistic regression model until an output of a loss function calculated based on the variant classification predictions output by the logistic regression model satisfies at least one first performance criterion to generate a trained logistic regression model, wherein the trained logistic regression model is capable of outputting variant pathogenicity estimates that satisfy at least one second performance criterion. A method comprising:
2. using the trained logistic regression model to generate a prediction as to whether the variant is benign or pathogenic, and providing the prediction as to whether the variant is benign or pathogenic to a clinician for use in the clinician's determination of a diagnosis for the patient; or applying the trained logistic regression model to a second set of population data for a plurality of variants of a second plurality of genes, receiving a variant classification prediction output by the trained logistic regression model for each variant of the second plurality of genes, and storing the variant classification prediction in association with the variant for retrieval via at least one query in response to the variant classification prediction satisfying at least the second performance criterion. The method of claim 1 further comprising:
3. Computing gene-level constraints; including said gene-level constraint on said at least one gene-level feature; calculating allele frequencies; including said allele frequencies in said at least one variant-level feature; including a mathematical combination of the gene-level constraints and allele frequencies in the at least one population frequency metafeature; and applying the logistic regression model to the set of features comprising the mathematical combination of the gene-level constraints and the allele frequencies, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the mathematical combination of the gene-level constraints and the allele frequencies. The method of claim 1 further comprising:
4. calculating the allele frequencies by calculating a binomial ratio of confidence values associated with the allele frequencies; or calculating the mathematical combination of the gene-level constraints and the allele frequencies by dividing the allele frequencies by the gene-level constraints to generate a quotient, and obtaining an exponent of the quotient; The method of claim 3 further comprising:
5. calculating an index of the ratio of synonymous to non-synonymous missense variants for said gene; including the index of the ratio of synonymous to non-synonymous missense variants in the at least one population frequency metafeature; and applying the logistic regression model to the set of features including the index of the ratio of synonymous to non-synonymous missense variants, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features including the index of the ratio of synonymous to non-synonymous missense variants. The method of claim 1 further comprising:
6. calculating a mathematical combination of the subpopulation frequency data and the population frequency data for the variants; including a mathematical combination of the subpopulation frequency data and the population frequency data in the at least one variant-level feature; and applying the logistic regression model to the set of features comprising a mathematical combination of the subpopulation frequency data and the population frequency data, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising a mathematical combination of the subpopulation frequency data and the population frequency data. The method of claim 1 , further comprising:
7. (i)(a) selecting 30 or fewer features as the set of features; and (b) applying the logistic regression model to the selected set of 30 or fewer features, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the selected set of 30 or fewer features; or (ii)(a) calculating a fixation index, wherein said fixation index comprises subpopulation frequency data; (b) including the fixed index in the set of features; and (c) applying the logistic regression model to the set of features including the fixation index, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate based on the set of features including the fixation index that meets the at least one second performance criterion; or (iii) iteratively adjusting the value of at least one parameter of the logistic regression model comprises adjusting a C value of the logistic regression model until the output of the loss function satisfies the at least one first performance criterion to generate the trained logistic regression model; or (iv) configuring the logistic regression model to model population frequencies for variant classification using L1 regularization; or (v) estimating the first performance criterion using a mean squared error or area based receiver operating characteristic curve; or (vi) determining the second performance criterion using at least one of a decision boundary test, a feature weight test, or a comparison to a benchmark variant classification; or (vii) using the variant classification estimates output by the trained logistic regression model as input to a variant classification framework; or (viii) the at least one population frequency meta-feature comprises an expected frequency distribution of known benign variants and known pathogenic variants in the gene. The method of claim 1 , further comprising:
8. at least one processor; and at least one memory coupled to the at least one processor; Equipped with wherein the at least one memory, when executed by the at least one processor, applying a logistic regression model to a first set of population data for a first set of genes, wherein the items of the first set of population data include a set of features for variants located at intragenic locations in the first set of genes, the set including at least one gene-level feature, at least one variant-level feature, and at least one population frequency meta-feature, and a reference label indicating whether the variant is benign or pathogenic, wherein the at least one population frequency meta-feature quantifies a predictive value of an allele frequency within the genes; evaluating, for each item in the first set of population data, a variant classification prediction output by the logistic regression model based on an expected variant classification indicated by the reference label; and adjusting the value of at least one parameter or coefficient of the logistic regression model until at least one first performance criterion is met to generate a trained logistic regression model, wherein the trained logistic regression model is capable of outputting variant pathogenicity estimates that satisfy at least one second performance criterion. at least one instruction that causes the at least one processor to perform at least one operation including system.
9. The at least one memory When executed by the at least one processor, (i) using the trained logistic regression model to generate a prediction as to whether the variant is benign or pathogenic, and providing the prediction as to whether the variant is benign or pathogenic to a clinician for use in the clinician's determination of a diagnosis for the patient; or (ii) (a) applying the trained logistic regression model to a second set of population data for a plurality of variants of a second plurality of genes; and (b) receiving a variant classification prediction output by the trained logistic regression model for each variant of the second plurality of genes, and, in response to the variant classification prediction satisfying at least the second performance criterion, storing the variant classification prediction in association with the variant for retrieval via at least one query. at least one instruction that causes the at least one processor to perform at least one operation including: further comprising The system of claim 8.
10. The at least one memory When executed by the at least one processor, Procedures for computing gene-level constraints; including said gene-level constraint on said at least one gene-level feature; Procedures for calculating allele frequencies; including said allele frequencies in said at least one variant-level feature; including a mathematical combination of the gene-level constraints and allele frequencies in the at least one population frequency metafeature; and applying the logistic regression model to the set of features comprising the mathematical combination of the gene-level constraints and the allele frequencies, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the mathematical combination of the gene-level constraints and the allele frequencies. at least one instruction that causes the at least one processor to perform at least one operation including: further comprising The system of claim 8.
11. The at least one memory When executed by the at least one processor, (i) calculating the allele frequencies by calculating a binomial ratio of confidence values associated with the allele frequencies; or (ii) calculating the mathematical combination of the gene-level constraints and the allele frequencies by dividing the allele frequencies by the gene-level constraints to generate a quotient and taking an exponent of the quotient; at least one instruction that causes the at least one processor to perform at least one operation including: further comprising The system of claim 10.
12. The at least one memory When executed by the at least one processor, calculating an index of the ratio of synonymous to non-synonymous missense variants for said gene; including the index of the ratio of synonymous to non-synonymous missense variants in the at least one population frequency metafeature; and applying the logistic regression model to the set of features including the index of the ratio of synonymous to non-synonymous missense variants, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features including the index of the ratio of synonymous to non-synonymous missense variants. at least one instruction that causes the at least one processor to perform at least one operation including: further comprising The system of claim 8.
13. The at least one memory When executed by the at least one processor, a procedure for calculating a mathematical combination of subpopulation frequency data and population frequency data for variants; including a mathematical combination of the subpopulation frequency data and the population frequency data in the at least one variant-level feature; and applying the logistic regression model to the set of features comprising the mathematical combination of the subpopulation frequency data and the population frequency data, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the mathematical combination of the subpopulation frequency data and the population frequency data. at least one instruction that causes the at least one processor to perform at least one operation including: further comprising The system of claim 8.
14. The at least one memory When executed by the at least one processor, (i) (a) selecting 30 or fewer features as the set of features; and (b) applying the logistic regression model to the selected set of 30 or fewer features, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the selected set of 30 or fewer features; or (ii)(a) calculating a fixation index, wherein the fixation index comprises subpopulation frequency data; (b) including said fixed index in said set of features; and (c) applying the logistic regression model to the set of features including the fixation index, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate based on the set of features including the fixation index that meets the at least one second performance criterion; or (iii) iteratively adjusting the value of at least one parameter of the logistic regression model comprises adjusting a C value of the logistic regression model until the output of a loss function satisfies the at least one first performance criterion to generate the trained logistic regression model; or (iv) configuring the logistic regression model to model population frequencies for variant classification using L1 regularization; or (v) estimating the first performance criterion using a mean squared error or area based on a receiver operating characteristic curve; or (vi) determining the second performance criterion using at least one of a decision boundary test, a feature weight test, or a comparison to a benchmark variant classification; or (vii) using the variant classification estimates output by the trained logistic regression model as input to a variant classification framework; or (viii) the at least one population frequency meta-feature comprises an expected frequency distribution of known benign variants and known pathogenic variants in the gene. at least one instruction that causes the at least one processor to perform at least one operation including: further comprising 14. A system according to any one of claims 8 to 13.
15. When executed by at least one processor, applying a logistic regression model to a first set of population data for a first set of genes, wherein the items of the first set of population data include a set of features for variants located at intragenic positions in the first set of genes, the set of features including at least one population frequency metafeature, and a reference label indicating whether the variant is benign or pathogenic, wherein the at least one population frequency metafeature quantifies a predictive value of an allele frequency within the genes; evaluating, for each item in the first set of population data, a variant classification prediction output by the logistic regression model based on an expected variant classification indicated by the reference label; and adjusting the value of at least one parameter or coefficient of the logistic regression model until at least one first performance criterion is met to generate a trained logistic regression model, wherein the trained logistic regression model is capable of outputting variant pathogenicity estimates that satisfy at least one second performance criterion.
10. A computer program product comprising at least one instruction causing said at least one processor to perform operations including:
16. When executed by the at least one processor, Procedures for computing gene-level constraints; Procedures for calculating allele frequencies; including a mathematical combination of the gene-level constraints and the allele frequencies in the at least one population frequency metafeature; and applying the logistic regression model to the set of features comprising the mathematical combination of the gene-level constraints and the allele frequencies, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the mathematical combination of the gene-level constraints and the allele frequencies.
16. The computer program product of claim 15, further comprising at least one instruction that causes the at least one processor to perform at least one operation comprising:
17. When executed by the at least one processor, (i) calculating the allele frequencies by calculating a binomial ratio of confidence values associated with the allele frequencies; or (ii) calculating the mathematical combination of the gene-level constraints and the allele frequencies by dividing the allele frequencies by the gene-level constraints to generate a quotient and taking an exponent of the quotient; at least one instruction that causes the at least one processor to perform at least one operation including:
17. The computer program of claim 16, further comprising:
18. When executed by the at least one processor, calculating an index of the ratio of synonymous to non-synonymous missense variants for said gene; including the index of the ratio of synonymous to non-synonymous missense variants in the at least one population frequency metafeature; and applying the logistic regression model to the set of features including the index of the ratio of synonymous to non-synonymous missense variants, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features including the index of the ratio of synonymous to non-synonymous missense variants.
20. The computer program product of claim 17, further comprising at least one instruction that causes the at least one processor to perform at least one operation comprising:
19. When executed by the at least one processor, A procedure for calculating a mathematical combination of subpopulation frequency data and population frequency data for variants; and applying the logistic regression model to the set of features comprising the mathematical combination of the subpopulation frequency data and the population frequency data, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the set of features comprising the mathematical combination of the subpopulation frequency data and the population frequency data.
16. The computer program product of claim 15, further comprising at least one instruction that causes the at least one processor to perform at least one operation comprising:
20. When executed by the at least one processor, (i) (a) selecting 30 or fewer features as the set of features; and (b) applying the logistic regression model to the selected set of 30 or fewer features, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion based on the selected set of 30 or fewer features; or (ii)(a) calculating a fixation index, wherein the fixation index comprises subpopulation frequency data; (b) including said fixed index in said set of features; and (c) applying the logistic regression model to the set of features including the fixation index, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate based on the set of features including the fixation index that meets the at least one second performance criterion; or (iii) using the trained logistic regression model to generate a prediction as to whether the variant is benign or pathogenic, and providing the prediction as to whether the variant is benign or pathogenic to a clinician for use in the clinician's determination of a diagnosis for the patient; or (iv) (a) applying the trained logistic regression model to a second set of population data for a plurality of variants of a second plurality of genes; and (b) receiving a variant classification prediction output by the trained logistic regression model for each variant of the second plurality of genes, and, in response to the variant classification prediction satisfying at least the second performance criterion, storing the variant classification prediction in association with the variant for retrieval via at least one query; or (v) adjusting the value of at least one parameter of the logistic regression model includes adjusting a C value of the logistic regression model until an output of a loss function satisfies the at least one first performance criterion to generate the trained logistic regression model; or (vi) configuring the logistic regression model to model population frequencies for variant classification using L1 regularization; or (vii) estimating the at least one first performance criterion using a mean squared error or an area based receiver operating characteristic curve; or (viii) determining the second performance criterion using at least one of a decision boundary test, a feature weight test, or a comparison to a benchmark variant classification; or (ix) using the variant classification estimates output by the trained logistic regression model as input to a variant classification framework; or (x) the at least one population frequency meta-feature comprises an expected frequency distribution of known benign variants and known pathogenic variants in the gene.
20. The computer program product of claim 15, further comprising at least one instruction to cause the at least one processor to perform at least one operation comprising: