Population frequency modeling for quantitative variant pathogenicity estimation
The population frequency model is constructed through logistic regression model, which solves the problem of accurate calculation of the pathogenicity probability of variants, and achieves more accurate variant classification and genetic test results, reduces the classification of unknown variants, and improves the accuracy of genetic tests.
Patent Information
- Application Number
- CN202380086028.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-01
- Filing Date
- 2023-10-31
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art is difficult to accurately define and calculate the pathogenicity probability of variants, resulting in genetic testing laboratories being unable to make full use of allelic frequency data in large population databases, resulting in inaccurate classification of variants, affecting the patient's genetic test results and diagnosis.
Using machine learning technology, especially logistic regression modeling, based on population frequency modeling methods, the population frequency model is constructed using allele frequency and other characteristics combinations, quantitatively estimate the pathogenicity probability of variants, avoiding the use of gene-disease attributes as characteristics, and providing a mapping from continuous measurement values to continuous results.
It improves the accuracy and reliability of variant classification, reduces the number of unknown variants, enhances the quantitative assessment of variant pathogenicity, and improves the accuracy of genetic testing.
Smart Images

Figure CN120359571A_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 421,430, filed on November 1, 2022, which is incorporated herein by reference in its entirety. Technical field
[0003] The technical fields related to this application are genetic testing. Another technical field related to this application is a machine - learning - based variant classification system. Background art
[0004] Genetic variants are differences in DNA sequences between individuals in a population. There are many different types of variants, including structural variants, single - nucleotide polymorphisms, insertions and deletions, copy - number variants, and translocations and inversions.
[0005] Genetic sequencing technologies continue to develop rapidly. High - throughput sequencing technologies are increasingly enabling genetic tests that include the following: genotyping, single genes, gene panels, exomes, genomes, transcriptomes, and epigenetic analysis for genetic diseases. The increasing complexity of the analysis and interpretation of clinical genetic tests, along with the increasing volume of tests, is accompanied by new challenges in the interpretation of sequence variants.
[0006] For example, during the testing of patient specimens for a rapidly growing number of genes associated with genetic diseases, clinical molecular laboratories are increasingly detecting new sequence variants. Although some phenotypes are associated with single genes, many phenotypes are associated with multiple genes.
[0007] Variant classification refers to the process of classifying genetic variants based on evidence that supports or excludes a causal relationship with a disease. The clinical significance of any given sequence variant falls within a gradient range: from variants that are almost certainly pathogenic for a certain disease to variants that are almost certainly benign.
[0008] Variant classification itself is not a diagnosis, but it can be used by clinicians to make diagnostic decisions. Brief description of the drawings
[0009] The present disclosure will be more fully understood from the following detailed description and from the accompanying drawings of various embodiments of the present disclosure. The drawings are for purposes of explanation and understanding only and should not be considered as limiting the present disclosure to the specific embodiments shown.
[0010] Figure 1 Illustrates an example of a feature generation process for machine - learning - based population frequency modeling according to some embodiments of the present disclosure.
[0011] Figure 2 Illustrates an example of a process for configuring a population frequency model using a logistic regression model according to some embodiments of the present disclosure.
[0012] Figure 3 Illustrates an example use of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure.
[0013] Figure 4 Illustrates an example process for configuring a population frequency model using a logistic regression model according to some embodiments of the present disclosure.
[0014] Figure 5A Illustrates an example of a pre - calibration curve of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure.
[0015] Figure 5B Illustrates an example of a post - calibration curve of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure.
[0016] Figure 6 Illustrates an example of a process for validating a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure.
[0017] Figure 7A Illustrates an example of a decision boundary of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure.
[0018] Figure 7B Illustrates an example of a gene - specific response curve of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure.
[0019] Figure 7C Illustrates an example of variant classification based on a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure.
[0020] Figure 8A Illustrates an example of a population frequency modeling result of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure.
[0021] Figure 8B Illustrates another example of a population frequency modeling result of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure.
[0022] Figure 8C Illustrates another example of a population frequency modeling result of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure.
[0023] Figure 8D Illustrates another example of the population frequency modeling results of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure.
[0024] Figure 8E Illustrates another example of the population frequency modeling results of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure.
[0025] Figure 8F Illustrates another example of the population frequency modeling results of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure, the population frequency model being integrated with a variant classification framework.
[0026] Figure 9 Illustrates a method for population frequency modeling using a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure.
[0027] Figure 10 Illustrates an example computing system including a population frequency modeling system according to some embodiments of the present disclosure.
[0028] Figure 11 Is a block diagram of an example computer system in which aspects of the present disclosure may operate. Detailed Description
[0029] Given a gene, a specific genetic disease, and a population, there are characteristics of the gene-disease relationship associated with the frequency of pathogenic variants in the population. These characteristics include penetrance, age of onset, severity, mode of inheritance, and disease prevalence. Prevalence may refer to the allele frequency of variants potentially associated with a genetic disorder in the general healthy population, or may refer to the frequency of patients affected by the disease in a given population.
[0030] For a given disease, a higher-than-expected prevalence of variants in the population is a pattern that has been empirically observed and that serves as evidence for a benign classification relative to the disease, or conversely, a lower-than-expected prevalence should be considered evidence for a pathogenic classification relative to the disease. According to the American College of Medical Genetics (ACMG) guidelines for the interpretation of sequence variants, for a given disorder, the presence of a variant in the population at a frequency higher than the prevalence of the disorder should constitute strong evidence for a benign classification.
[0031] However, the prevalence of pathogenic variants in large populations is the result of complex interactions of various gene-disease attributes that reflect the genetic makeup of individuals for whom sequencing data are available. Thus, accurately and systematically defining the relationship between allele frequency and pathogenicity relative to a given population and disease is an ongoing challenge across all gene-disease relationships. One challenge is how to define and calculate the upper boundary beyond which a particular variant should be classified as benign relative to a given disease based on population frequency data. Another challenge is to improve on conventional binary classification methods.
[0032] The Genome Aggregation Database (gnomAD) is a large population database developed by an international consortium with the goal of aggregating and harmonizing exome and genome sequencing data from various large-scale sequencing projects and making the total data available to the broader scientific community. For example, the gnomAD v2.1.1 dataset encompasses 125,748 exome sequences and 15,708 genome sequences from unrelated individuals as part of various disease-specific studies and population genetics studies. Subsequent versions of gnomAD include even more genomes and even greater ancestral diversity.
[0033] For classifying variants, the major industry-wide challenge in using general population frequency data obtained from large general population databases such as gnomAD is that many of the attributes conventionally considered key predictive features are difficult to obtain and not easily interpretable. For example, penetrance is difficult to measure accurately, is often not available, and may vary among variants. Age of onset is rarely a single value but is more commonly a range of values. Severity is usually a qualitative rather than a quantitative measure. Prevalence is often inaccurate or imprecise and may vary among ancestral groups.
[0034] In addition, gene-level measurements only provide information about the cumulative frequency of all pathogenic variants in a particular gene and do not provide information about specific individual variants in that gene. For example, a prevalence of 0.1% could mean that there is one pathogenic variant in the gene (with an allele frequency of 0.1%), or there are 1,000 different pathogenic variants (each with an allele frequency of 0.0001%), or any other combination of pathogenic variants that add up to 0.1%. Differences in the distribution of pathogenic variants have very different implications for how population allele frequency data should be used in variant classification, leading to a need for further refinement.
[0035] Another limitation of conventional methods (which is overcome by the described population frequency modeling method) is that large population databases such as gnomAD only include data that are estimates of the true population frequencies of various variants; thus, there is sampling uncertainty around these estimates. As described herein, through the described feature generation method (which is applied to large population data for variant classification), this sampling uncertainty is taken into account.
[0036] Despite these and other limitations, conventional methods for evaluating this type of allele frequency data rely heavily on the aforementioned gene-disease attributes. Thus, genetic testing laboratories have had to conservatively set high discrete allele frequency thresholds for variant classification and then apply these conservative thresholds broadly to large gene groups based on simple parameters (e.g., a threshold for a gene associated with a dominant disease versus a threshold for a gene associated with a recessive disease). As a result, laboratories have been unable to fully utilize the allele frequency data available in these large population databases for variant classification. This has led to too many variants being classified as variants of unknown significance (VUS) even when the available allele frequency data indicate that the variant is too common to be expected to cause disease. In these cases, it may leave patients unsure of how their genetic test results reflect their disease risk or diagnosis and how these results will affect their treatment.
[0037] The state of the art of research on a given gene-disease pair and the complexity and limitations of the relationships between gene-disease attributes limit the use of quantitative algorithms for variant classification. The disclosed methods address these and other challenges in variant classification. Embodiments of the disclosed methods apply machine learning techniques to develop a computational algorithm into a machine learning model for population frequency modeling. The described population frequency model is based on allele frequency data obtained from a large population database such as gnomAD and uses logistic regression to quantitatively estimate the pathogenicity probability of a variant.
[0038] Certain embodiments of the described population frequency model examine allele frequencies in the context of a relatively small set of attributes that do not belong to the aforementioned gene-disease attributes. The embodiments avoid using the aforementioned gene-disease attributes as features and instead construct particularly useful combinations of other variant-level, gene-level, and / or position-level properties of a given variant. For example, embodiments of the described population frequency model have shown that reliable pathogenicity estimates can be generated based on a combination of allele frequency and a total of fewer than thirty other input features. For example, in some embodiments, the number of input features ranges from about twenty features to about thirty features. The feature combinations constructed as described herein enable the embodiments to exclude the aforementioned gene-disease attributes from the set of features used to configure the population frequency model.
[0039] The population frequency model obtained by applying the described method utilizes the expected frequency distributions of known benign variants and known pathogenic variants within a given gene, rather than relying on gene-disease attributes, thereby allowing for a quantitative assessment of the deviation between the allele frequency of a specific variant and the expected frequency in the case where the variant is pathogenic. Using the described method, the model calculates and outputs the pathogenicity probability for each variant.
[0040] Unlike non-quantitative existing methods, the population frequency model configured as described maps continuous measurements (e.g., allele frequencies) to continuous outcomes (e.g., pathogenicity probabilities). For example, an embodiment of the population frequency model can provide a variant-specific quantitative measurement of the population frequency as a continuous score. Additionally, embodiments of the population frequency model allow for a continuous quantitative scoring method in place of or in addition to a binary method.
[0041] The output of the population frequency model can be directly used for variant classification or supplied as input to a variant classification framework. For example, the model output can be incorporated into the population data section of the Sherloc framework. The Sherloc framework is a semi-quantitative method for variant interpretation (i.e., a process for reviewing and evaluating evidence of pathogenicity, classifying variants, and communicating pathogenicity information to patients in an understandable manner, which is also used for variant classification). Experimental results show that when the model output generated according to the disclosed method is provided as input to the Sherloc framework, the number of variants that would otherwise be routinely classified as VUS due to the lack of other evidence is reduced.
[0042] In the absence of any information about complex gene-disease attributes, the population frequency model configured as described is able to reproduce the expected relationships between genes with different gene-disease attributes, even though those attributes are not included in the feature set provided as model input. Using the gene-disease attributes of inheritance patterns as an example, it is generally expected that genes associated with recessive conditions have a higher prevalence of pathogenic variants than genes associated with dominant conditions. This is because heterozygous carriers of autosomal recessive conditions are expected to be part of the population database cohort, while individuals with a pathogenic variant in an autosomal dominant condition are likely to be affected by the disease, leading to their exclusion. Thus, even if two variants have the same allele frequency, the variant observed in a recessive gene is more likely to be pathogenic than the variant observed in a dominant gene.
[0043] As described in more detail below, when the described embodiments of the population frequency model are implemented in two such genes, the model reports a higher pathogenicity probability for variants in genes associated with autosomal recessive conditions than for variants in genes associated with autosomal dominant conditions (e.g., LAMA2 and TSC2) in the absence of any prior information about the mode of inheritance. For example, Figure 8B shows such results that demonstrate that, in the absence of any prior information about the mode of inheritance (i.e., information about the mode of inheritance is not provided as input to the model), the described embodiments of the population frequency model can distinguish the effects of different modes of inheritance on allele frequencies.
[0044] Additionally, for a given allele frequency, variants in genes associated with diseases (having higher severity, higher penetrance, and earlier onset) are routinely assumed to have a lower calculated pathogenicity probability than variants in genes associated with diseases (having lower disease severity, lower penetrance, and later onset). For example, Figure 8A shows the results of population frequency modeling using the described method for two genes (i.e., MLH1, KMT2D) associated with different modes of severity, penetrance, and age of onset. As shown, the use of the described method reveals that the pathogenicity probability at a given allele frequency can vary widely between genes having the same mode of inheritance but different severity, penetrance, and age of onset.
[0045] Even when comparing multiple genes with complex combinations of gene-disease attributes, embodiments of the population frequency model configured as described herein provide probabilities of variant pathogenicity that are consistent with known gene-disease attributes (even though information about those gene-disease attributes is not provided to the model). For example, Figure 8C shows the results of population frequency modeling for five genes (i.e., EYS, TGM1, CACNA1C, FBN1, DNAH11) associated with different modes of inheritance, severity, penetrance, and age of onset. As shown, the use of the described method reveals that even when two genes are associated with, for example, the same mode of inheritance (e.g., DNAH11 and EYS) or similar penetrance and age of onset (e.g., CACNA1C and FBN1), the pathogenicity probability at a given allele frequency can vary widely between any two genes.
[0046] Given the experimental results achieved to date, the population frequency model configured as described represents a highly accurate, highly scalable, and fully quantitative solution for determining predicted values of population allele frequencies in the context of individual genes and across genes.
[0047] Additionally or alternatively, embodiments of the population frequency model can be used to identify subtle but potentially important biological differences between genes, such as those observed between DNAH11, DNAI1, and DNAH5.
[0048] Incorporating the described tools into a variant classification framework, such as the Sherloc interpretation system, significantly increases the ability to accurately and confidently classify variants. Within Sherloc, for example, embodiments of the described population frequency model have been applied to four times as many variants as an earlier method that alternatively utilized allele frequencies from gnomAD, resulting in the resolution of approximately 15,000 unique VUSs when implemented. Through simulation experiments on historical VUSs in a local database and using a process of applying a conservative threshold to the model output, it is estimated that the described tools will result in a reduction of approximately 2.5% in VUSs compared to other existing methods for evaluating population allele frequency data. In other variant classification frameworks that do not utilize this conservative thresholding method, the reduction in VUSs provided by the described model can be further improved.
[0049] The present disclosure will be more fully understood from the detailed description given below, which refers to the accompanying drawings. The detailed description of the drawings is for explanation and understanding and should not be considered as limiting the present disclosure to the specific embodiments described.
[0050] In the drawings and the following description, reference will be made to components that have the same name but different reference numerals in different drawings. Using different reference numerals in different drawings indicates that components with the same name may represent the same or different embodiments of the same component. For example, in some embodiments, components with the same name but different reference numerals in different drawings may have the same or similar functions, such that the description of one of those components with respect to one drawing can be applied to other components with the same name in other drawings.
[0051] Furthermore, in the drawings and the following description, components shown and described in connection with some embodiments can be used with or incorporated into other embodiments. For example, a component illustrated in a particular drawing is not limited to being used in conjunction with the embodiment to which that drawing pertains, but can be used with or incorporated into other embodiments, including those shown in other drawings.
[0052] Figure 1 An example of a feature generation process for machine learning-based population frequency modeling according to some embodiments of the present disclosure is illustrated.
[0053] In Figure 1Among them, the dataset of sequence data 110 includes DNA sequence data of one to N human populations, where N is a positive integer and the populations can be defined by any one or more demographic criteria. The DNA sample includes one or more genes 102. Each gene includes one or more regions 104, for example, a continuous group or sequence of nucleotides or proteins (e.g., A G A C G C T, where A represents adenine, G represents guanine, C represents cytosine, and T represents thymine). Each nucleotide in the region has one or more positions where variants may be located. For example, in Figure 1 Person 1 in population N has cytosine at position 106 of region 104 of a copy of their gene 102, while person 2 in population N has thymine at position 106 of region 104 of a copy of their gene 102, where thymine is considered a genomic variant 108.
[0054] For each of the one to N human populations, the sequence data 110 includes gene-level data corresponding to each gene 102, region-level data corresponding to one or more regions 104, position-level data corresponding to one or more positions 106, and variant-level data corresponding to one or more genomic variants 108. Gene-level data refers to any property, attribute, or metric determined for the defined sequence of DNA identified as a gene. Region-level data refers to any property, attribute, or metric determined for a region of a gene (i.e., a defined part of the gene that is typically greater than one position but less than the entire gene). Position-level data refers to any property, attribute, or metric determined for a position within a gene (i.e., a single nucleotide or amino acid position within the gene). And typically, a gene may be thousands of nucleotides long, and variants usually occur at a single position among these thousands of nucleotides. A region can refer to a part of a gene that includes a variant and one or more adjacent or neighboring nucleotides.
[0055] Gene-level data includes information specific to a particular gene, such as gene length. Region-level data includes information specific to a particular region of a gene. For example, region-level data can include the properties of the region containing the variant. Position-level data includes information specific to a particular position in a gene, regardless of whether a variant is present at that position. Variant-level data includes information about a particular variant located at a particular position in a gene. A variant can include more than one nucleotide. When a variant includes more than one nucleotide, the position of the variant refers to a region, such that the terms "position" and "region" can be synonymous in this context. The source of the sequence data 110 can be a publicly available population database, such as gnomAD.
[0056] In Figure 1In this case, the feature generation process 100 generates a feature set 112 for input into a machine learning model. The feature generation process 100 includes a feature extraction and calculation sub-process 114 and a feature transformation sub-process 116. The feature extraction and calculation sub-process 114 extracts and / or calculates gene-level features 118, variant-level features 120, region-level features 122, and position-level features 124. For example, the feature extraction and calculation sub-process 114 extracts features from a population database and / or calculates features based on data extracted from the population database. An example of a calculated feature is a confidence interval associated with another feature, e.g., a confidence interval calculated for allele frequency.
[0057] The feature transformation sub-process 116 applies one or more feature transformations to one or more combinations of the gene-level features 118, variant-level features 120, region-level features 122, and / or position-level features 124 to produce population frequency meta-features 126. Feature transformation can refer to an operation performed on a calculated feature. For example, a calculation operation can be applied to an original feature (e.g., original observed or measured data) to generate a calculated feature, and another calculation operation can be applied to the calculated feature to perform a feature transformation on the calculated feature, e.g., to obtain the population frequency meta-feature 126.
[0058] Examples of feature transformations that can be applied to one or more combinations of the gene-level features 118, variant-level features 120, region-level features 122, and / or position-level features 124 to produce population frequency meta-features 126 include mathematical operations such as logarithm, exponentiation, square root, summation, product, mean, average, median, and / or other mathematical functions. Examples of population frequency meta-features 126 include features that are a mathematical combination of gene-level features and variant-level or position-level features. For example, a population frequency meta-feature quantifies a predicted value of the allele frequency of a specific variant, region, or position within a gene in a given population. For example, a population frequency meta-feature 126 quantifies the reliability of a population allele frequency as a pathogenicity marker for a specific variant, region, or position within a gene in a population. Examples of population frequency meta-features 126 include the expected frequency distribution of known benign variants and known pathogenic variants within a specific gene in a specific population.
[0059] Examples of features that can be included in the feature set 112 are shown in Table 1 below. In an embodiment, data including allele counts specific to an ancestry, allele numbers, and gene-level constraint estimates (e.g., LOEUF) are collected from gnomAD (v2.1.1) for single nucleotide variants in genes strongly associated with at least one disease, according to a reference genetics knowledge base (e.g., a local, in-house developed knowledge base).
[0060]
[0061] Table 1. Examples of features
[0062] In the examples, features (also referred to as properties) and confidence intervals for each group were generated, resulting in a set of features for a total of 33,402 variants. Some features, such as allele frequencies, were obtained from population databases such as gnomAD. Other features, such as position-level constraints, were calculated based on raw data obtained from population databases. A constraint is a metric calculated in an external dataset released to the public, e.g., via gnomAD. The constraint indicates the amount of tolerated variation of a gene in the human population. Similarly, the constraint can be used as a signal to identify more or less constrained genes; i.e., genes that are more or less essential for normal healthy human function. A constraint is a gene-level feature.
[0063] Another example of a gene-level feature is the synonymous-to-nonsynonymous ratio. Synonymous means that the variant does not result in a change in the protein, while nonsynonymous means that the variant does result in a change in the protein. The synonymous-to-nonsynonymous ratio indicates the degree to which variants have been spared from selection for nonsynonymous variants throughout evolution, thus indicating intolerance to protein variation for essential protein function.
[0064] The fixation index (FST) is a calculated feature that indicates the difference in allele frequencies between different subpopulations. For example, for several different subpopulations (e.g., Southeast Asians, Africans, etc.), allele frequencies can be obtained from a population database, and based on the subpopulation allele frequency data, a calculation can be performed that indicates the degree of difference in allele frequencies across populations.
[0065] Another example of a position-level feature is the identity (ID) of the largest subpopulation, i.e., the ancestral group that has been identified as having the highest allele frequency for a particular variant. The ID of the largest subpopulation can be calculated by first calculating the allele frequencies of the variant for each subpopulation and then ranking or classifying the subpopulations based on the allele frequencies. Another example of a position-level feature is the actual frequency in the largest population, e.g., the upper bound of x% (e.g., 95%) of the allele frequency in the identified largest subpopulation.
[0066] An example of a region-level feature is moving window missense. Moving window missense is the number of missense mutations (nucleotide mutations that result in an amino acid change, i.e., a meaningful change in the protein) in a region of a gene. Moving window missense can indicate the degree of tolerance to missense mutations in a particular region of a gene.
[0067] Feature transformations can be applied to gene-level constraints to apply constraint information at the variant level. For example, a mathematical combination of allele frequency and a constraint (e.g., a log transformation) is an example of a population frequency meta-feature. Another example of a population frequency meta-feature is a position-level constraint. A position-level constraint is created by applying a feature transformation to a gene-level constraint. For example, a position-level constraint can be calculated by taking the average constraint at each position on a gene. A position-level constraint indicates the degree to which a specific position or region within a gene (as opposed to the overall gene) is constrained. A position-level constraint can be calculated by using the variants observed in a region as a normalization metric for the gene-level constraint. Another example of a population frequency meta-feature is an allele frequency confidence interval to which a feature transformation has been applied, such as a binomial proportion transformation of the allele frequency confidence interval.
[0068] Another example of a population frequency meta-feature is AF_DIV_LOEUF, which is calculated by dividing the allele frequency by the constraint and then taking the exponent of the resulting quotient. Another example of a population frequency meta-feature is SIN2_MIS_EXP, which is the exponent of the ratio of synonymous missense variants to non-synonymous missense variants in a gene.
[0069] Other examples of population frequency meta-features can be created by applying one or more feature transformations to a combination of one or more gene-level features and allele frequencies, or to a combination of gene-level constraints and one or more variant-level features and / or position-level features.
[0070] Figure 2 An example of a process for configuring a population frequency model using a logistic regression model and machine learning techniques in accordance with some embodiments of the present disclosure is illustrated. The process is executed by processing logic that includes hardware (e.g., a processing device, circuitry, dedicated logic, programmable logic, microcode, device hardware, integrated circuits, etc.), software (e.g., instructions running or executing on a processing device), or a combination thereof. In some embodiments, the method is executed by components of a computing system, and in some embodiments, the components of the computing system include components or processes that may not be specifically shown in other figures Figure 2 shown in, and / or in some embodiments, the components of the computing system include components or processes that may not be specifically shown in Figure 2 other figures shown in that are not specifically shown in
[0071] InFigure 2 In this process, the model development process 200 uses one or more machine learning techniques to prepare one or more data sets, such as the training data set 234 and the validation data set 238, for input into a machine learning model (e.g., the logistic regression model 228). Parts of the process 200 are performed by components of a population frequency modeling computing system (e.g., the population frequency modeling system 1050 described below), which includes a gene data selection subsystem 206, a feature generation subsystem 212, a modeling and calibration subsystem 220, and a model validation subsystem 230. The data sources from which data is received during various parts of the process 200 include unannotated population data 202, annotated population data 204, the logistic regression model 228, model performance criteria 232, the training data set 234, model validation criteria 236, and the validation data set 238. Figure 10 The unannotated population data 202 includes DNA sequence data of one or more human populations. For example, the unannotated population data 202 includes population data extracted from gnomAD or a similar database. The unannotated population data 202 does not include associated pathogenicity labels or scores. The pathogenicity labels or scores (e.g., benign or pathogenic) are obtained from the annotated population data 204. The annotated population data 204 is a reference database that associates variants with associated ground truth pathogenicity labels or scores, such as ClinVar or an in-house developed database curated by genetic scientists and / or other genetic experts. The ground truth labels or scores can be combined or merged with the corresponding unannotated population data 202 (e.g., using a common key value) to produce the annotated population data, to which the logistic regression model 228 can be applied using supervised machine learning methods.
[0072] Before or after one or more operations of the gene data selection subsystem 206, the annotated population data 204 is combined or merged with the unannotated population data 202. For example, after the gene data selection subsystem 206, the annotated population data 204 can be combined with the unannotated population data 202 to create one or more training data sets and / or validation data sets.
[0073]
[0074] The gene data selection subsystem 206 evaluates gene-specific data sets of the unlabeled population data 202 and selects one or more gene-specific data sets for generating a feature set that can be input into the logistic regression model 228. The gene data selection subsystem 206 applies one or more filters to the population data, which include criteria for determining whether data for a particular gene will be included in or excluded from the model development process. For example, genes that have no strong association with any disease may not be eligible to be included in the population frequency model.
[0075] For example, the gene data selection subsystem 206 applies the gene-disease validity filter 208 to the unlabeled population data 202 to create a gene-disease validity-filtered subset of the unlabeled population data 202, combines or merges the gene-disease validity-filtered subset with the labeled population data 204 such that each item in the gene-disease validity-filtered subset matches a corresponding ground truth label, and then applies the label minimum filter 210 to the labeled gene-disease validity-filtered subset of the unlabeled population data 202 to generate and output a labeled gene-disease validity-filtered and label minimum-filtered data set.
[0076] The gene-disease validity filter 208 evaluates the unlabeled population data 202 to filter out data that is not strongly associated with a disease. For example, the gene-disease validity filter 208 retains the unlabeled population data 202 belonging to such variants that have a high probability (e.g., greater than or equal to 90%) of being benign or pathogenic with respect to a particular disease, and filters out the unlabeled population data 202 belonging to such variants that have an uncertain probability (e.g., less than 90%) of being benign or pathogenic with respect to a particular disease. Gene-disease association data can be obtained from publicly available sources such as the Geneticus database.
[0077] The label minimum filter 210 evaluates the labeled gene-disease validity-filtered subset of the unlabeled population data 202 against a label minimum threshold. The label minimum filter 210 retains those items in the labeled gene-disease validity-filtered subset of the unlabeled population data 202 that have a number of pathogenic and benign labels that reach or exceed the applicable label minimum threshold, and filters out those items in the labeled gene-disease validity-filtered subset of the unlabeled population data 202 that have a number of labels that do not reach or exceed the label minimum threshold, i.e., that do not have a sufficient number of benign and pathogenic labels to contribute to model training. For example, gene data sets that do not have at least a minimum number of benign labels are filtered out of the data set.
[0078] The feature generation subsystem 212 generates features to be included in a feature set for input into a logistic regression model. For example, embodiments of the feature generation subsystem 212 use the feature extraction, feature calculation, and feature transformation techniques described above with reference to Figure 1 to generate various features.
[0079] The feature generation subsystem 212 receives as input a labeled, gene-disease validity-filtered, and label minimum-filtered data set, and generates a feature set (e.g., feature set 112) based on the received labeled, gene-disease validity-filtered, and label minimum-filtered data set. The feature generation subsystem 212 includes a feature extraction component 214, a feature calculation component 216, and a feature transformation component 218. The feature extraction component 214 extracts features from the received labeled, gene-disease validity-filtered, and label minimum-filtered data set. The extracted features can include raw features and / or calculated features. Examples of raw features and calculated features were described above with reference to Figure 1 and
[0080] The feature calculation component 216 applies one or more calculation operations to one or more raw features created by the feature extraction component 214. Examples of calculation operations that can be applied to one or more raw features were described above with reference to Figure 1 and
[0081] The feature transformation component 218 applies one or more calculation operations to one or more calculated features created by the feature extraction component 214 or the feature calculation component 216. Examples of feature transformations that can be applied to one or more calculated features were described above with reference to Figure 1 and
[0082] The feature generation subsystem 212 generates and outputs a feature set that includes raw features, calculated features, and / or feature transformations. Examples of features and feature sets that can be generated and output by the feature generation subsystem 212 were described above with reference to Figure 1 and
[0083] The modeling and calibration subsystem 220 receives as input the feature set constructed and output by the feature generation subsystem 212. The modeling and calibration subsystem 220 includes a data set creation component 222, a model training component 224, and a model calibration component 226.
[0084] The dataset creation component 222 divides the feature set constructed and output by the feature generation subsystem 212 into a training dataset and a validation dataset, e.g., the training dataset 234 and the validation dataset 238. For example, in some embodiments, the dataset creation component 222 creates gene-specific training datasets and validation datasets for each of up to or more than six hundred different genes, where each gene-specific dataset includes an input feature set (e.g., the feature set 112) associated with a particular variant within a particular gene. As an example, for a given variant, each feature set created by the dataset creation component 222 includes approximately 24 different variant attributes, including allele frequency data, associated confidence intervals, and gene lengths extracted from a population database such as gnomAD, as well as one or more population frequency meta-features.
[0085] The model training component 224 and the model calibration component 226 perform a model training process that enables the logistic regression model 228 to derive a mathematical representation of the relationship between the input features and the ground truth labels, such that the resulting model can be used to predict the pathogenicity of new variants based on the input features associated with those new variants (i.e., variants not previously seen by the model). For example, the mathematical representation of these relationships can be presented as a two-dimensional plot of feature values (e.g., variant properties) (x-axis) versus pathogenicity labels (y-axis).
[0086] The model training component 224 iteratively applies the logistic regression model 228 to the training dataset 234 and adjusts one or more model parameters and / or feature coefficients until the difference between the predicted model output generated by the logistic regression model 228 and the expected model output indicated by the ground truth labels (obtained via the annotated population data 204) satisfies (e.g., reaches or exceeds) the model performance criterion 232. When the model performance criterion 232 is satisfied, the modeling and calibration subsystem 220 ends the model training process and produces the trained logistic regression model 228. A more detailed example of the model training and calibration process that can be used to create the trained logistic regression model 228 is described below with reference to Figure 4 、 Figure 5A and Figure 5B An example of a more detailed model training and calibration process that can be used to create the trained logistic regression model 228 is described below with reference to
[0087] The model validation subsystem 230 applies a model validation process to the trained logistic regression model 228 produced by the modeling and calibration subsystem 220. The model validation subsystem 230 applies the trained logistic regression model 228 to the validation dataset 238 to determine whether the model validation criterion 236 is satisfied (e.g., reaches or exceeds). A more detailed example of the model validation process is described below with reference to Figure 6 、 Figure 7A and Figure 7B A more detailed example of the model validation process is described below with reference to
[0088] If the trained logistic regression model 228 is successfully verified by the model verification subsystem 230, the verified logistic regression model 228 can be used for inference, e.g., to generate pathogenicity prediction values or estimates for new (i.e., previously unseen) variants. Alternatively or additionally, the prediction values output by the verified logistic regression model 228 can be stored for future use (e.g., for access or lookup by one or more downstream processes, systems, or services). A more detailed example of a logistic regression model and its use during inference (configured to use the feature sets and techniques described herein for variant classification) is described below with reference to Figure 3 more detailed examples of logistic regression models and the use of logistic regression models (configured to use the feature sets and techniques described herein for variant classification) during inference.
[0089] Figure 3 Illustrated is an example use of a population frequency model configured with a logistic regression model in accordance with some embodiments of the present disclosure.
[0090] The logistic regression model 306 is a statistical machine learning model that uses a logistic function to model the relationship between X and Y, where the probability of Y is a linear combination of the independent variables in the input X. Mathematically, the simplified form of the logistic function can be expressed as where e is the natural constant and β0 and β1 are feature coefficients. During the training of the logistic regression model 306, logistic regression estimates the coefficient values in the linear combination based on the feature values in the training dataset.
[0091] In Figure 3 , the logistic regression model 306 has been configured via the described supervised machine learning training, calibration, and verification processes. The logistic regression model 306 includes a logistic function 308. The logistic function 308 includes feature coefficients 310. The feature coefficients 310 include regression coefficients β for each feature input x (e.g., f(i) = β0 + β1x 1,i + … β m x m,i ), where i is a specific entry in the feature set (e.g., entry or row 304), and m is the number of feature inputs x in the feature set 302. The regression coefficients, based on the values of the feature inputs x in the feature set 302, indicate the relative impact of a specific feature input x in the feature set 302 on the prediction result P(Y|X) (e.g., the predicted pathogenicity label or score). The values of the feature coefficients are initialized and adjusted during model training and calibration.
[0092] The logistic regression model 306 also includes model hyperparameters 312, which are selected or adjusted at a global level and are generally not modified based on specific instances of the training data. For the logistic regression model 306, the model hyperparameters 312 include penalty or regularization parameters (e.g., L1 or L2) and C or regularization strength parameters. The penalty or regularization parameters are adjustable to adjust the model generalization error and to constrain overfitting. Typical values for the penalty or regularization parameters are L1 and L2. In some embodiments of the logistic regression model 306, the penalty or regularization parameter is set to L1. The C or regularization strength parameter, in combination with the penalty, constrains overfitting. A smaller C value specifies stronger regularization. Regularization imposes a cost on the model complexity by penalizing large coefficient values. Regularization limits the number of different ways the model can fit the data. The C value is subject to hyperparameter optimization processes such as grid search. In some embodiments of the logistic regression model 306, the C value is set to a value in the range of approximately 1 to approximately 10.
[0093] The model hyperparameters 312 can be adjusted using, for example, a grid search cross-validation procedure. For example, in some embodiments, an automated hyperparameter tuning tool (e.g., GridSearchCV) is used for hyperparameter tuning. In other embodiments, other hyperparameter optimization methods such as random search and Bayesian search are used. In the illustrated embodiment, hyperparameter tuning is performed on the training dataset to compare the performance of various combinations of model types and hyperparameters. The simplest model with the strongest regularization (e.g., an L1-regularized logistic regression model, C = 10) is selected for inference, which yields high validation performance (e.g., AUROC validation value = 0.92). The model prediction values for the training dataset and the validation dataset are generated separately by averaging the calibrated prediction values from ten-fold cross-validation such that the validation dataset is not used for training set predictions. Inference on VUS is performed by first refitting the model to the entire labeled set before making predictions on new variants.
[0094] The logistic regression model 306 can be configured as a binary classifier or a scoring model. In binary classification mode, for a given set of input features, the output of the logistic regression model 306 indicates the prediction result as pathogenic or benign with a binary value (e.g., 0 indicates benign and 1 indicates pathogenic). In scoring mode, the output of the logistic regression model 306 includes a score that corresponds to the probability that the prediction result is pathogenic or benign (e.g., a numerical value between 0 and 1, inclusive).
[0095] The logistic regression model 306 can be configured and implemented as a web service. For example, a machine learning library such as scikit-learn can be used to configure the logistic regression model 306. For example, the logistic regression model 306 can be configured via an application programming interface (API), e.g., via an API call such as ML_library.model.logistic_regression(p1,p2,…pn), where p indicates the parameter or argument of the call, such as a model hyperparameter or an input feature set identifier. Once configured, the logistic regression model 306 and / or its output can be hosted on one or more servers and / or data storage devices for access by one or more requesting processes, systems, devices, frameworks, or services.
[0096] In Figure 3 , the feature set 302 includes items or instances of features 304 for each gene-disease-variant combination. For example, the feature set 302 includes items or instances of features 304 for N genes, N diseases, and N variants, where the value of N can be the same or different in each case. Each item or instance 304 includes one or more gene-level features x g1 ...x gN , one or more position-level features x p1 ...x pN , one or more variant-level features x v1 ...x vN , and one or more population frequency meta-features x mf1 ...x mfN . Examples of the gene-level features x g1 ...x gN , the position-level features x p1 ...x pN , the variant-level features x v1 ...x vN , and the population frequency meta-features x mf1 ...x mfN include those described in Figure 1 . Alternatively or additionally, the feature set 302 can include region-level features. In other embodiments, the feature set 302 can include gene-level features, variant-level features, and population frequency meta-features, but not position-level features or region-level features. The features 304 can be grouped into the feature set 302 using, for example, a concatenation function.
[0097] The embodiments of the feature set 302 are limited to quantitative numerical features and categorical features and do not include qualitative features or gene-disease attributes. Prior to input into the logistic regression model 306, the entries or instances of the features 304 may be transformed into a vector representation. In some embodiments, prior to input into the logistic regression model 306, the vector representation of the features is transformed into a compressed form such as an embedding.
[0098] In response to each instance of the features in the feature set 302, the logistic regression model 306 computes and outputs an estimated result P GDV (Y|X)314. The estimated result generated by the logistic regression model 306 based on the instances of the features in the feature set 302 is in the form of a binary output (e.g., 0 for benign and 1 for pathogenic) or a score (e.g., a value between 0 and 1). This output may be stored in a data storage device for subsequent lookup or provided to one or more downstream systems, processes, devices, frameworks, and / or services.
[0099] Figure 4 Illustrated is an example process for configuring a population frequency model using a logistic regression model according to some embodiments of the present disclosure. In Figure 4 this, the model training process 400 applies a logistic function to a training data set 401 using regression-based supervised machine learning. The training data set 401 includes a feature set 402 and labels 404. The feature set 402 is similar to the feature set 302, where ground truth labels are appended to each corresponding entry or instance in the feature set as described. In an embodiment, a feature set for a total of 33,402 variants (i.e., labels) is used for training across 823 genes (n = 18,148 benign and 15,234 pathogenic variants). All training variants from 25% of the genes are held out as a test (or validation) set and remain unused until after model training is complete. Gene-based selection is not required to exclude labeled training data; that is, not all embodiments require a training set to be held out in the manner described.
[0100] In the first training iteration, at sub-process 406, feature coefficients are initialized or assigned to each feature input in the feature set 402. For example, the feature coefficients are initialized by randomly setting the coefficient values. The feature coefficient values assigned at sub-process 406 are used as weights applied to the corresponding features to produce weighted features 410. At sub-process 412, the logistic function is applied to the weighted features to produce a predicted output 414. The ground truth label 404 provides the expected output 408 for supervised machine learning. At sub-process 416, a loss (or error) is computed to evaluate the predicted output 414 by based on the expected output 408 and the predicted output 414. A loss function (such as a gradient descent algorithm) is used to compute the loss.
[0101] In each iteration, the decision sub - process 418 evaluates the difference between the predicted output and the expected output, for example, by comparing the output of a loss function with a stopping condition that depends on changes in the output of the loss function. The change in the output of the loss function is compared with an error tolerance threshold (which may be referred to herein as a model performance criterion or a model convergence criterion). If the model performance criterion (e.g., the error tolerance threshold) is not met (e.g., not greater than or equal to the threshold performance level, or exceeding the maximum allowable error value, or the loss has not stopped improving by an amount beyond the tolerance, or the error has not stopped decreasing by an amount beyond the tolerance threshold), then the model training continues for another iteration. If the loss has stopped improving by an amount beyond the tolerance threshold, then the model has converged and the training ends.
[0102] In subsequent training iterations, the values of one or more feature coefficients are adjusted, the logistic function is applied to additional instances in the feature set 402, and the output of the logistic function is evaluated using the loss function and error tolerance as described above. When the model performance criterion (i.e., the error tolerance threshold) is met and the model has converged (e.g., the output of the comparison (e.g., the loss function) is greater than or equal to the threshold performance level, or within or below the maximum allowable error value, or the loss has stopped improving by an amount beyond the tolerance threshold), the training process 400 ends.
[0103] Figure 5A An example of a pre - calibration curve of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure is illustrated.
[0104] In operation, for a given gene, variant, and disease, embodiments of the population frequency model typically output a large number of scores close to 0 (e.g., benign) and a large number of scores close to 1 (e.g., pathogenic). Figure 5A An example of a pre - calibration curve compared to an expected or ideal calibration is shown. Before calibration, the model tends to underestimate the variant pathogenicity probability in the lower score region (e.g., the part of the pre - calibration curve below the ideal calibration line). The difference between the ideal calibration and the pre - calibration curve can be referred to as an error. The calibration process adjusts the values on the pre - calibration curve so that they align more closely with the ideal calibration diagonal. Depending on the error in a given region of the model output, the scores are transformed (e.g., by adjusting one or more model coefficients) so that the expected probability matches the predicted probability. Figure 5A The pre - calibration curve in is an aggregated curve of scores output by models for multiple different genes. For individual genes, a similar curve can be used to perform a similar calibration process.
[0105] Figure 5B An example of a post - calibration curve of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure is illustrated.Figure 5B shows an example of a calibrated curve compared to an expected or ideal calibration, which is generated from the calibration applied to the curve in Figure 5A . After calibration, the model output tends to align more closely with the ideal calibration line. Figure 5B The calibrated curve in is an aggregated curve of the scores of the model outputs for multiple different genes. For individual genes, a similar curve can be used to perform a similar calibration process.
[0106] Figure 6 illustrates an example of a process for validating a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure. In Figure 6 , the validation process 600 applies the trained logistic regression model 604 to the validation dataset 602. An example of the validation dataset 602 is a portion of the feature set 402 that has been reserved for validation and not for training. For example, referring to Figure 2 , the dataset creation component 222 creates several gene-specific partitions of the data and sets aside a specific dataset for an entire gene that is not used for model training but can be used to create the validation dataset. The trained logistic regression model 604 is, for example, a logistic regression model trained and configured as described herein.
[0107] In response to the processing of the validation dataset 602, the trained logistic regression model 604 provides validation parameters for evaluation to the sub-process 612, where the validation parameters are the model output 606, the feature weights 608, and the decision boundary 610. The sub-process 612 uses validation criteria to evaluate the performance of the trained logistic regression model 604, for example, by examining samples of the actual model output 606, the feature weights 608, and / or the decision boundary 610 and comparing the examined samples to expected or reference values, or by calculating one or more validation metrics.
[0108] Model validation can be performed through multiple iterations, where each iteration uses a slightly different validation dataset that was not used in training. The trained logistic regression model 604 generates pathogenicity prediction values on the validation set, and based on the average performance of the model on these partitions, the performance of the model parameters can be verified and adjusted. Examples of metrics that can be used to evaluate model performance include mean squared error or area under the receiver operating characteristic curve.
[0109] Methods that can be used to validate the trained logistic regression model 604 include: evaluation of dataset performance metrics, decision boundary testing, evaluation of model-level performance metrics, gene and variant-level testing, feature weight testing, and benchmark comparison with existing frameworks (e.g., Sherloc).
[0110] Decision boundary testing refers to the process of evaluating the actual boundary of the model (i.e., the threshold at which the model distinguishes between benign and pathogenic). Examples of decision boundary evaluation are given in Figure 7A (described below). Examples of evaluation of performance metrics for datasets are shown in Figure 7B . (described below).
[0111] In response to the evaluation of model performance performed by subprocess 612, certain gene-specific data sets and corresponding model outputs may be selected to be included in a database or a downstream system, device, process, or service if the data sets meet corresponding validation criteria. Alternatively, if a certain gene-specific data set does not meet corresponding validation criteria, the data set may be excluded or marked as unavailable for subsequent use, e.g., not stored in a database and not provided to any downstream system, device, process, or service.
[0112] For technical validation, the model output was verified at the gene level for performance and calibration. Genes that showed poor calibration (i.e., Brier score>0.15) or poor discrimination performance (i.e., AUROC<0.80) were marked to be excluded from the final model output. Of the initial 823 genes, a total of 591 reached sufficiently high calibration and discrimination performance requirements. It is noteworthy that when calculating performance, the 232 genes marked for exclusion were still used in subsequent steps to avoid artificially causing bias to the overall performance evaluation.
[0113] For clinical validation, to further enhance confidence in the predictive output of the population frequency model, a concordance analysis was performed on three independent in-silico algorithms developed internally that rely on orthogonal data types. As shown in Table 2 below, the concordance rates were high in all three comparisons.
[0114]
[0115]
[0116] Table 2. Examples of agreement rates
[0117] In addition, model outputs were reviewed for a particularly challenging class of variants: high allele frequency pathogenic variants. These variants resemble benign variants based on population allele frequency data alone, but are classified as pathogenic or likely pathogenic due to other evidence supporting pathogenicity. These include variants such as founder variants, which are ancestry-specific, and hypomorphic alleles that are associated with milder phenotypes or reduced penetrance relative to typical pathogenic variants in the same gene.
[0118] Among the 591 genes included in the analysis, 30 high allele frequency pathogenic variants in 26 genes were identified that would have been identified as benign or likely benign based solely on allele frequency but had conflicting evidence supporting pathogenicity. Although the population frequency model still predicted 18 of these 30 variants as strongly or moderately benign (NPV ≥ 95%), the remaining 12 variants were predicted to support benign (95% ≥ NPV ≥ 80%), support pathogenicity (PPV ≥ 80%), or be insufficiently determined (NPV < 80% and PPV < 80%). These results indicate that the population frequency model has greater specificity for identifying benign variants than traditional methods, although a critical review of all conflicting evidence remains a necessary part of variant classification. An example of how NPV / PPV values obtained from the described population frequency model can be converted into discrete scores (e.g., Sherloc scores such as 5-point value benign, 3-point value benign, 1-point value benign, and 1-point value pathogenic) is shown in Figure 7C what follows.
[0119] Figure 7A FIG. illustrates an example of a decision boundary of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure. In Figure 7A this, the variants are projected onto a two-dimensional plane. Figure 7A The figures in Figure 7A are two-dimensional representations of the variants represented by circles and x's. To perform the decision boundary test, one or more points in the figure in
[0120] Figure 7B can be sampled and tested to determine whether the prediction for the variant is pathogenic or benign. The decision boundary 706 separates the points (pathogenic) in region 702 from the points (benign) in region 704. Points closer to the decision boundary 70 have lower certainty for the associated predicted value, while points further away from the decision boundary 706 are more reliable classifications. Figure 7B this, the predicted value of the POPMAX feature can be evaluated for an individual gene. The feature on the x-axis is the square root of the calculated allele frequency. Increasing allele frequency is associated with an increasing predicted value of benign. Figure 7B FIG. illustrates that the relationship between allele frequency and predicted value of benign is different for different genes.
[0121] Figure 7C FIG. illustrates an example of variant classification based on a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure. In Figure 7CIn it, the count of variants each having a fractional value is represented as a histogram. For example, the model assigns a score of 0 to approximately 2900 variants, and the model assigns a score greater than 0.8 to over 3000 variants. The histogram is partitioned into categorical groups or bins, and point values are associated with each group or bin of scores. For example, the portion of the histogram associated with a score of 0 is assigned 5 point values (e.g., benign), and the portion of the histogram adjacent to the zero-score region is assigned 3 point values (e.g., likely benign), while the portion of the histogram associated with scores greater than 0.6 is associated with a point value of 1. The middle portion of the histogram is assigned a point value of 0 because the associated score values have lower confidence. The point values assigned based on the model output can be combined with other evidence to which the classification framework is applied.
[0122] Generally, the number of point values associated with a particular score is related to the confidence or certainty of the predicted value. The number of bins need not be four; any number of bins can be used (including no bins or an infinite number of bins). The point values mapped to the scores output by the model can be incorporated into another variant classification framework, such as Sherloc. In this way, the output of the described population frequency model can be applied to the Sherloc framework or any other variant labeling or classification framework.
[0123] To integrate the model prediction values into the Sherloc variant classification framework, five prediction value levels were established based on prediction performance thresholds measured by the negative predictive value (NPV) and positive predictive value (PPV). Four of these levels reflect the existing Sherloc framework used to evaluate population frequency data and are defined as: (1) [Strongly Benign] Sufficient confidence to classify as benign in the absence of strong contradictory evidence; (2) [Moderately Benign] Sufficient confidence to classify as likely benign in the absence of strong contradictory evidence; (3) [Benign-Supporting] Consistent with a benign classification but insufficient to reach a likely benign classification in the absence of orthogonal evidence also supporting benignity; and (4) [Pathogenicity-Supporting] Consistent with a pathogenic classification but insufficient to reach a likely pathogenic classification in the absence of orthogonal evidence also supporting pathogenicity. The prediction performance thresholds for these four levels are defined as: (1) [Strongly Benign] > 99% NPV, (2) [Moderately Benign] > 95% to 99% NPV, (3) [Benign-Supporting] > 80% to 95% NPV, and (4) [Pathogenicity-Supporting] > 80% PPV. The fifth and final level corresponds to prediction values below 80% PPV and below 80% NPV, which are considered to have insufficient certainty to be weighted within the Sherloc scoring system. The overall model performance and level classification performance were evaluated on 25% of the genes reserved as a test set. The overall model performance reached an AUROC test value = 0.92, thus confirming that the model generated from the training set generalizes to the test set genes. For the four PPV and NPV levels, the prediction values matched or exceeded the target performance on the test set. In the current iteration of the model, the model output is limited to missense variants and single nucleotide substitution nonsense variants. Future development of the model may extend its prediction scope to new variant types.
[0124] Figure 8A , Figure 8B , Figure 8C and Figure 8D illustrate the prediction outputs generated by an embodiment of a population frequency model configured as described, i.e., the population frequency modeling results. Figure 8A shows a comparison of variants found in the genes MLH1 (associated with Lynch syndrome) and KMT2D (associated with Kabuki syndrome). Figure 8B shows a comparison of variants found in the genes TSC2 (associated with tuberous sclerosis complex) and LAMA2 (associated with LAMA2-related muscular dystrophy). Figure 8CShows a comparison of the genes EYS (associated with retinitis pigmentosa), TGM1 (associated with ichthyosis), CACNA1C (associated with epileptic encephalopathy, etc.), FBN1 (associated with Marfan syndrome, etc.), and DNAH11 (associated with primary ciliary dyskinesia).
[0125] All variants are plotted based on the output generated by an example of a population frequency model as described. Each variant (represented as a point in each figure) is placed along the x-axis according to the allele frequency reported in the gnomAD population database. The y-axis represents the probability that the variant is pathogenic based on the population frequency model configured as described. The gene-disease attributes for each gene are summarized in the table, where AD indicates autosomal dominant and AR indicates autosomal recessive.
[0126] These model outputs illustrate the limitations of conventional methods: Conventional methods define discrete allele frequency thresholds for variant classification and then apply them broadly to many genes as traditionally done. When using this conventional method, variants may cross the threshold with very different pathogenic probabilities (depending on the gene). Notably, variants in DNAH11, which is associated with autosomal recessive primary ciliary dyskinesia, generally have a lower pathogenic probability than variants in CACNA1C and FBN1 with similar allele frequencies, both of which are associated with autosomal dominant conditions. This may be due to DNAH11 being associated with a highly penetrant, early-onset, severe, and rare condition, but it may also be due to this large gene (4516 codons) giving rise to many different but individually rare pathogenic variants. This further highlights the challenges associated with applying a general allele frequency threshold to many genes.
[0127] Although Figure 8A , Figure 8B , Figure 8C and Figure 8D The figures in show the machine-learned relationship between variant pathogenicity probability and allele frequency for specific genes and variants generated by the population frequency model, it should be understood that in addition to allele frequency, the model inputs include other features.
[0128] Figure 8A Illustrates an example of the population frequency modeling results of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure. Figure 8A Shows: Examples of the pathogenicity probabilities that have been estimated and output by the described population frequency model for each variant in two different genes, MLH1 and KMT2D. In Figure 8AIn it, each point in the figure represents a different specific variant within a particular gene. Most variants in the gene MLH1 are plotted in region 802, while most variants in the gene KMT2D are plotted in region 804. The y-axis shows the model output (i.e., the pathogenicity probability of the variant), and the x-axis shows the allele frequency of the variant in a given human population. As the frequency of the variant in the population increases, the pathogenicity probability decreases.
[0129] More specifically, Figure 8A illustrates an example of a population frequency model configured as described, which outputs a variant pathogenicity probability that takes into account differences in genes with the same genetic pattern. For example, as shown in the table included in Figure 8A , both MLH1 and KMT2D are dominant for certain diseases, but have different severity, penetrance, and age of onset characteristics.
[0130] Although it can be expected that population frequency information should be used differently for genes with different genetic patterns, in the case of Figure 8A , both genes are dominant, but they have different other key characteristics (severity, penetrance, age of onset). In a conventional variant classification system, it is difficult to determine how much adjustment should be made to the variant pathogenicity probability for these two genes to take these differences into account.
[0131] In contrast, using a population frequency model configured as described, the model output can automatically determine the amount of adjustment to the variant pathogenicity probability to take into account the similar genetic pattern and different severity, penetrance, and age of onset characteristics of these two genes, without having access to any information about these characteristics. Although the model input does not include any information indicating that these two genes have the same genetic pattern (or different severity, penetrance, and age of onset), the model has learned through machine learning-based training and calibration (using the feature set described) that variants in these two genes should be treated differently in terms of the population frequency threshold and how much the probability needs to be adjusted.
[0132] Figure 8B illustrates another example of the population frequency modeling results of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure. Figure 8B Shows: Examples of the pathogenicity probabilities estimated and output by the described population frequency model for each variant in two different genes (TSC2 and LAMA2). In Figure 8BIn it, each point in the figure represents a different specific variant within a particular gene. Most variants in the gene LAMA2 are plotted in region 822, while most variants in the gene TSC2 are plotted in region 824. The y-axis shows the model output (i.e., the pathogenicity probability of the variant), and the x-axis shows the allele frequency of the variant in a given human population. As the frequency of the variant in the population increases, the pathogenicity probability decreases.
[0133] More specifically, Figure 8B illustrates an example of a population frequency model configured as described, which outputs the pathogenicity probabilities of different variants of two genes that have similar severity, penetrance, and disease onset properties but different inheritance patterns. As shown in the table included in Figure 8B , TSC2 is dominant for a particular disease, while LAMA2 is recessive. The model output supports the intuition that variants in genes with different inheritance patterns should follow different population frequency criteria. Although these differences can be taken into account during the conventional variant classification process by setting different thresholds for recessive and dominant genes, the commonly used thresholds are still categorical (e.g., recessive, dominant) and usually require manual adjustment.
[0134] In contrast, using a population frequency model configured as described, the model output supports the conventional intuition but without the need to be granted access to any information about the inheritance pattern. Although the model input does not include any information indicating that the two genes have different inheritance patterns, the model has learned via machine learning-based training and calibration (using the feature set described) that variants in these two genes should be treated differently in terms of population frequency thresholds and how much the probability needs to be adjusted.
[0135] Figure 8C illustrates another example of the population frequency modeling results of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure. Figure 8C shows examples of the pathogenicity probabilities estimated and output by the described population frequency model for each variant in several different genes (EYS, TGM1, CACNA1C, FBN1, and DNAH11). In Figure 8CIn it, each point in the figure represents a different specific variant within a particular gene. Most variants in the gene EYS are plotted in region 830, most variants in the gene TGM1 are plotted in region 832, most variants in the gene CACNA1C are plotted in region 834, most variants in the gene FBN1 are plotted in region 836, and most variants in the gene DNAH11 are plotted in region 838. The y-axis shows the model output (i.e., the pathogenicity probability of the variant), and the x-axis shows the allele frequency of the variant in a given human population. As the frequency of the variant in the population increases, the pathogenicity probability decreases.
[0136] More specifically, Figure 8C illustrates an example of a population frequency model configured as described, which outputs the pathogenicity probabilities of different variants of multiple different genes, and these genes have differences that are not intuitive to variant classification scientists. As shown in the table included in Figure 8C , these genes have a mixture of similarities and differences in traditional gene-disease attributes. The differences in pathogenicity probabilities between these genes are not obvious to humans. However, the population frequency model configured as described has taken into account these gene-specific differences through a machine learning-based training and calibration process (using the feature set as described). Therefore, the pathogenicity curves generated based on the model output show that variants in these genes should be treated differently in terms of population frequency thresholds.
[0137] Figure 8D illustrates another example of the population frequency modeling results of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure. Figure 8D shows an example of the pathogenicity probability that has been estimated and output by the described population frequency model for each variant in a given gene (MLH1). In Figure 8D , each point in the figure represents a different specific variant. The y-axis shows the pathogenicity probability of the variant, and the x-axis shows the allele frequency of the variant in a given human population. As the frequency of the variant in the population increases, the pathogenicity probability decreases. For a specific variant 870, the model output shows that the variant has a relatively high pathogenicity probability (i.e., the model output can indicate that the variant is within the pathogenicity range, even though the actual probability may still be less than 50%). For a specific variant 872, the model output shows that the pathogenicity probability is very low. Using a variant classification framework such as Sherloc (where point values are assigned based on evidence of pathogenicity), a benign point value will not be given to variant 870, but will be given to variant 872.
[0138] Figure 8EIllustrates another example of the population frequency modeling results of a population frequency model configured to use a logistic regression model according to some embodiments of the present disclosure. Figure 8E Shows: an example of the pathogenicity probabilities that have been estimated and output by the described population frequency model for each variant in a given gene (MLH1). In Figure 8E , each point in the graph represents a different specific variant. The y-axis shows the pathogenicity probability of the variant, and the x-axis shows the allele frequency of the variant in a given human population. As the frequency of the variant in the population increases, the pathogenicity probability decreases.
[0139] In Figure 8E , the graph is the same as the graph in Figure 8D , but highlights different variants for discussion. For a specific variant 880, the model output shows that the variant has a high pathogenicity probability. For a specific variant 882, the model output shows that the pathogenicity probability is relatively low, although the allele frequency indicates that variant 882 is relatively rare.
[0140] Additionally, Figure 8E , variants 880 and 882 illustrate the vertical distribution of variants within a single gene at a given frequency. This is due to the fact that in addition to gene-level properties, regional-level and position-level properties (e.g., gene-level features, regional-level features, and position-level features as described) are provided as model inputs to the population frequency model.
[0141] Figure 8E Shows that variant 882 at a given allele frequency is less likely to be pathogenic, but variant 880 at the same allele frequency is more likely to be pathogenic. Conventionally, it can be understood that different parts of the same gene can be more or less tolerant of variations caused by, for example, protein domains and functional sequence motifs, and such low-probability variants may affect regions in the gene that are less important than high-probability variants. However, it is difficult to resolve these differences using conventional variant classification techniques. Instead, although the domain and motif information is not provided to the population frequency model, the model configured as described has taken these differences into account through a machine learning-based training and calibration process (using the feature set as described). Thus, the pathogenicity illustration generated based on the model output shows that different variants in the same gene should be treated differently in terms of the population frequency threshold.
[0142] Similarly, embodiments of the population frequency model can not only model the expected frequency of a disease for a given gene, but also determine the association of different variants in the same gene with the same frequency to different pathogenicity probabilities. Additionally, embodiments of the model configured as described can discern different allele frequency thresholds for different regions within a gene. This is because the model is trained and calibrated using the regional properties within the gene.
[0143] Figure 8F FIG. illustrates an example of the population frequency modeling results of a population frequency model configured using a logistic regression model according to some embodiments of the present disclosure, which is integrated with a variant classification framework. In particular, Figure 8F FIG. illustrates how the accuracy (measured by negative and positive predictive values) of variant pathogenicity prediction values generated and output by embodiments of the described population frequency model is mapped to points in a variant classification system such as Sherloc. Although Figure 8F FIG. illustrates an example implementation using the Sherloc framework, the model output can be similarly adapted to incorporate other variant classification frameworks.
[0144] As Figure 8F shown, to accommodate the output of the population frequency model, four new chains of evidence were created for the population modeling data used for Sherloc, where benign and pathogenic point values were assigned to each new chain of evidence based on the associated accuracy level.
[0145] For example, variants with a very high predicted benign population modeling score (which has an NPV (negative predictive value) greater than 99%) are assigned, for example, 5 benign point values, and in the absence of conflicting information from other evidence categories, will be classified as benign variants using the Sherloc framework.
[0146] Conversely, variants with a moderate or high predicted pathogenic population modeling score (which has a PPV (positive predictive value) greater than 80%) are assigned, for example, 1 point value for pathogenicity, and will require, for example, 4 additional pathogenic point values from other independent evidence categories to reach the threshold of 5 point values within the Sherloc framework and thus be classified as pathogenic.
[0147] Population modeling using the modeling methods described herein can have a significant impact on variant classification. For example, gnomAD population frequency data can now be applied to four times as many variants as could be handled using previous methods. This expanded use of population frequency data has enabled the reclassification of 15,000 variants from VUS to benign and likely benign variants. For example, 3,146 variants have been reclassified from VUS to benign, and 11,640 variants have been reclassified from VUS to likely benign.
[0148] In addition, the population frequency model configured as described enables more efficient use of population frequency data. As a result, many variants of unknown significance (which could not be reclassified based on that data alone) are now closer to receiving a definitive classification in the future. The reclassified variants have affected over 50,000 patients. It is estimated that the population modeling method described can reduce the VUS rate for future patients by 2 - 2.5%.
[0149] Furthermore, the population modeling method described has been applied to variants from different populations and across clinical domains. At an aggregated level, this enables the identification of subpopulations with the highest frequencies for specific variants and the grouping of reclassified variants by major clinical domain.
[0150] Figure 9 Method 900 for population frequency modeling using a population frequency model configured using a logistic regression model is illustrated, in accordance with some embodiments of the present disclosure.
[0151] Method 900 is executed by processing logic that includes hardware (e.g., a processing device, circuitry, special logic, programmable logic, microcode, device hardware, an integrated circuit, etc.), software (e.g., instructions running or executing on a processing device), or a combination thereof. In some embodiments, the method is executed by components of a computing system, and in some embodiments, the components of the computing system include components or processes that may not be specifically shown in other figures Figure 9 shown in, and / or in some embodiments, the components of the computing system include components or processes shown in other figures that may not be specifically shown in Figure 9 Although shown in a particular order or sequence, the order of the processes may be modified unless otherwise stated. Accordingly, the illustrated embodiments should be understood as merely examples, and the illustrated processes may be executed in a different order, and some processes may be executed in parallel. Additionally, in various embodiments, at least one process may be omitted. Accordingly, not all processes are required in every embodiment. Other processing flows are possible.
[0152] At operation 902, the processing device applies a logistic regression model to a first population dataset of a first gene set. For variants at positions within genes located in the first gene set, the entries in the first population dataset include a feature set. The feature set includes at least one gene-level feature, at least one variant-level feature, and at least one population frequency meta-feature. The entries in the first population dataset also include a reference label indicating whether the variant is benign or pathogenic.
[0153] At operation 904, for each entry in the first population dataset, the processing device compares the variant classification prediction value output by the logistic regression model with the expected variant classification value indicated by the reference label.
[0154] At operation 906, the processing device iteratively adjusts the value of at least one parameter or coefficient of the logistic regression model until one or more performance criteria are met. For example, the output of a loss function calculated based on the variant classification prediction values output by the logistic regression model meets at least one first performance criterion (e.g., an error tolerance threshold) to produce a trained logistic regression model (e.g., the model has converged). The trained logistic regression model is capable of outputting a variant pathogenicity estimate value that meets at least one second performance criterion (e.g., one or more validation criteria).
[0155] In some embodiments, the processing device uses the trained logistic regression model to generate a prediction value as to whether the variant is benign or pathogenic. In some embodiments, the processing device provides the prediction value as to whether the variant is benign or pathogenic to the clinician's computing device for use by the clinician in formulating a patient's diagnosis.
[0156] In some embodiments, the processing device applies the trained logistic regression model to a second population dataset of multiple variants in a second plurality of genes. For each variant in the second plurality of genes, the processing device receives the variant classification prediction value output by the trained logistic regression model. In response to the variant classification prediction value meeting at least a second performance criterion, the processing device stores the variant classification prediction value associated with the variant for retrieval via at least one query.
[0157] In some embodiments, the processing device calculates a gene-level constraint, includes the gene-level constraint in at least one gene-level feature, calculates an allele frequency, includes the allele frequency in at least one position-level feature; includes a mathematical combination of the gene-level constraint and the allele frequency in at least one population frequency meta-feature, and applies the logistic regression model to a feature set including the mathematical combination of the gene-level constraint and the allele frequency. Based on the feature set including the mathematical combination of the gene-level constraint and the allele frequency, the trained logistic regression model is capable of outputting a variant pathogenicity estimate value that meets at least one second performance criterion.
[0158] In some embodiments, the processing device calculates allele frequencies by computing the binomial proportion of confidence values associated with the allele frequencies. In some embodiments, the processing device calculates a mathematical combination of gene-level constraints and allele frequencies by dividing the allele frequencies by the gene-level constraints to produce a quotient and taking the exponent of the quotient.
[0159] In some embodiments, for a gene, the processing device calculates the exponent of the ratio of synonymous missense variants to non-synonymous missense variants. The processing device includes the exponent of the ratio of synonymous missense variants to non-synonymous missense variants in at least one population frequency meta-feature. The processing device applies a logistic regression model to a feature set that includes the exponent of the ratio of synonymous missense variants to non-synonymous missense variants. Based on the feature set that includes the exponent of the ratio of synonymous missense variants to non-synonymous missense variants, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0160] In some embodiments, the processing device selects no more than thirty features as the feature set. The processing device applies a logistic regression model to the selected set of no more than thirty features. Based on the selected set of no more than thirty features, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0161] In some embodiments, the processing device calculates a fixation index. The fixation index includes sub-population frequency data. The processing device includes the fixation index in the feature set. The processing device applies a logistic regression model to the feature set that includes the fixation index. Based on the feature set that includes the fixation index, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0162] In some embodiments, for a variant, the processing device calculates a mathematical combination of sub-population frequency data and population frequency data. The processing device includes the mathematical combination of sub-population frequency data and population frequency data in at least one position-level feature. The processing device applies a logistic regression model to a feature set that includes the mathematical combination of sub-population frequency data and population frequency data. Based on the feature set that includes the mathematical combination of sub-population frequency data and population frequency data, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0163] In some embodiments, iteratively adjusting the value of at least one parameter of a logistic regression model includes: adjusting the C value of the logistic regression model until the loss meets at least one first performance criterion to produce a trained logistic regression model. In some embodiments, a processing device configures the logistic regression model using L1 regularization. In some embodiments, the processing device uses mean squared error or area under the receiver operating characteristic curve to estimate the first performance criterion. In some embodiments, the processing device uses at least one of the following to determine a second performance criterion: a decision boundary test, a feature weight test, or a comparison with a baseline variant classification value.
[0164] In some embodiments, the processing device uses the variant classification estimate output by the trained logistic regression model as an input to a variant classification framework. In some embodiments, at least one population frequency meta-feature includes the expected frequency distribution of known benign variants and known pathogenic variants within a gene.
[0165] Figure 10 Illustrated is an example computing system including a population frequency modeling system according to some embodiments of the present disclosure. In Figure 10 this embodiment, the computing system 1000 includes one or more user systems 1010, a network 1020, an application system 1030, a population frequency modeling system 1050, and a data storage system 1080.
[0166] The population frequency modeling system 1050 includes a logistic regression model 1052, a model development subsystem 1054, model performance criteria 1056, model validation criteria 1058, a training data set 1060, and a validation data set 1062. The components of the population frequency modeling system 1050 may correspond to components of similar descriptions shown in other figures and described above. For example, the logistic regression model 1052 may correspond to the logistic regression model 228 or the logistic regression model 306 or the trained logistic regression model 604. The model development subsystem 1054 may correspond to a combination of components including a gene data selection subsystem 206, a feature generation subsystem 212, a modeling and calibration subsystem 220, and a model validation subsystem 230. The model development subsystem 1054 may be configured to perform process 300 and / or process 400 and / or process 600. The model performance criteria 1056 may correspond to the model performance criteria 232. The model validation criteria 1058 may correspond to the model validation criteria 236. The training data set 1060 may correspond to the training data set 234. The validation data set 1062 may correspond to the validation data set 238.
[0167] The logistic regression model 1052 includes one or more machine learning models that are trained to determine a probability or statistical relationship between inputs and outputs using machine learning algorithms. For example, given one or more inputs, the logistic regression model 1052 outputs a label or score that can be used to classify the inputs into different categories, and the score can be used to group the inputs into clusters or rank them in a ranked list. An example of the logistic regression model 1052 is the population frequency model described above.
[0168] The model development subsystem 1054 trains one or more machine learning models of the logistic regression model 1052 by, for example, applying supervised machine learning techniques to training data that includes training examples of input data and ground truth labels. The predicted outputs of the one or more machine learning models are observed iteratively until a set of model performance criteria are met. For example, a loss function is used to quantify the difference between the predicted output and the expected output. The model performance criteria are used to determine when the one or more machine learning models have converged to provide an output that can be relied upon with a certain degree of certainty. The necessary level of certainty and performance criteria are determined based on the requirements or design of a particular implementation of the one or more machine learning models. As described above, examples of the model development subsystem 1054 include a data preparation component, a feature construction component, and a model selection component.
[0169] In some embodiments, the training data set 1060 includes training data for training the logistic regression model 1052. The training data set 1060 includes, for example, a set of input features and corresponding ground truth labels. In some embodiments, the training data set 1060 includes or is derived from a database of historical population data. As described above, examples of the training data set 1060 include variant-level data sets, gene-level data sets, and location-level data sets.
[0170] The user system 1010 includes at least one computing device, such as a personal computing device, a server, a mobile computing device, or a smart appliance. The user system 1010 includes at least one software application that includes a user interface 1012 installed on the computing device or accessible by the computing device via a network. For example, embodiments of the user interface 1012 include a graphical display screen that displays controls and graphical elements for operating and / or manipulating one or more of the logistic regression model 1052, the model development subsystem 1054, and the training data set 1060.
[0171] The user interface 1012 can be used to input data, initiate user interface events, and observe or otherwise sense outputs including pathogenicity prediction values and / or other data generated by the population frequency modeling system 1050. Examples of the user interface 1012 include a web browser, a command line interface, and a mobile app front end. As used herein, the user interface 1012 can include an application programming interface (API). The user interface 1012 can include a front-end portion of the application system 1030 used by a clinician. For example, the output of the population frequency modeling system 1050 can be transmitted to and displayed by the user interface 1012 of a computing device used by a clinician. Alternatively or additionally, another version of the user interface 1012 can include a front-end portion of the application software system 1030 used by variant scientists and / or other individuals working in the field of genetic testing. Similarly, the output of the population frequency modeling system 1050 can be transmitted to and displayed by the user interface 1012 of a computing device used by any of these individuals.
[0172] The application system 1030 is any type of application software system that provides or implements the generation, display, or manipulation of the outputs generated by the population frequency modeling system 1050. Examples of the application system 1030 include but are not limited to variant classification systems, DNA (deoxyribonucleic acid) analysis software, genetic testing software, medical testing software, healthcare management software, or any combination of any of the foregoing software.
[0173] The data storage system 1080 includes data storage devices and / or data services that store data received, used, manipulated, and generated by the application system 1030 and / or the population frequency modeling system 1050, such as training data, validation data, machine learning model parameters and coefficients, performance criteria, validation criteria, machine learning model outputs, and the like. In Figure 10 it, the data storage system 1080 includes one or more data storage devices that store unlabeled population data 1082, labeled population data 1084, and logistic regression model outputs 1086. The unlabeled population data can correspond to the unlabeled population data 202. The labeled population data 1084 can correspond to the labeled population data 204. The logistic regression model outputs 1086 can correspond to, for example, model outputs 314, prediction outputs 414, or model outputs 606. In some embodiments, the data storage system 1080 includes multiple different types of data storage and / or distributed data services. As used herein, a data service can refer to: a physical, geographical group of machines, a logical group of machines, or a single machine. For example, a data service can be a data center, a cluster, a group of clusters, or a machine.
[0174] The data storage system 1080 resides on at least one persistent and / or volatile storage device, which may reside within the same local network as at least one other device of the computing system 1000 and / or within a network remote from at least one other device of the computing system 1000. Thus, although depicted as being included in the computing system 1000, portions of the data storage system 1080 may be part of the computing system 1000 or accessed by the computing system 1000 via a network such as the network 1020.
[0175] Although not specifically shown, it should be understood that any one of the user system 1010, the application system 1030, the population frequency modeling system 1050, and the data storage system 1080 includes an interface embodied as computer program code stored in a computer memory, which, when executed, causes the computing device to implement two-way communication with any other system among the user system 1010, the application system 1030, the population frequency modeling system 1050, and the data storage system 1080 using a communication coupling mechanism. Examples of communication coupling mechanisms include network interfaces, inter-process communication (IPC) interfaces, and application programming interfaces (APIs).
[0176] Each of the user system 1010, the application system 1030, the population frequency modeling system 1050, and the data storage system 1080 is implemented using at least one computing device that is communicatively coupled to the electronic communication network 1020. Any one of the user system 1010, the application system 1030, the population frequency modeling system 1050, and the data storage system 1080 may be communicatively coupled two-way via the network 1020. The user system 1010, as well as other different user systems (not shown), may be communicatively coupled two-way to the application system 1030 and / or the population frequency modeling system 1050.
[0177] A typical user of the user system 1010 may be an administrator or an end user of the application system 1030 and / or the population frequency modeling system 1050. The user system 1010 is configured to communicate two-way with the application system 1030 and / or the population frequency modeling system 1050 via the network 1020.
[0178] The features and functions of the user system 1010, the application system 1030, the population frequency modeling system 1050, and the data storage system 1080 are implemented using computer software, hardware, or a combination of software and hardware, and these features and functions may include a combination of automated functions, data structures, and digital data, which are schematically represented in the drawings. For ease of discussion, the user system 1010, the application system 1030, the population frequency modeling system 1050, and the data storage system 1080 are in Figure 10Elements shown as separate in the figure, however, unless otherwise described, the figure is not meant to imply a requirement for separating these elements. The illustrated systems, services, and data storage devices (or their functions) of each of the user system 1010, application system 1030, population frequency modeling system 1050, and data storage system 1080 can be partitioned onto any number of physical systems, including a single physical computer system, and can communicate with each other in any suitable manner.
[0179] The network 1020 can be implemented on any medium or mechanism that provides for the exchange of data, signals, and / or instructions between the various components of the computing system 1000. Examples of the network 1020 include, but are not limited to, a local area network (LAN), a wide area network (WAN), Ethernet, or the Internet, or at least one terrestrial, satellite, or wireless link, or any combination of a number of different networks and / or communication links.
[0180] For ease of discussion, in Figure 11 aspects of the population frequency modeling system 1050 are represented as the population frequency modeling system 1150.
[0181] Figure 11 is a block diagram of an example computer system in which aspects of the present disclosure can operate. Figure 11 illustrates an example machine of the computer system 1100, in which a set of instructions can be executed to cause the machine to perform any of the methods discussed herein. In some embodiments, the computer system 1100 can correspond to a component of a networked computer system (e.g., Figure 10 the computing system 1000 in Figure 10 ), which includes, is coupled to, or utilizes a machine to execute an operating system to perform the operations described above corresponding to aspects of the population frequency modeling system 1050 in
[0182] The machine is connected (e.g., networked) to other machines in a local area network (LAN), intranet, extranet, and / or the Internet. The machine can operate as a server or client machine in a client-server network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or client machine in a cloud computing infrastructure or environment.
[0183] The machine is a personal computer (PC), smartphone, tablet, set-top box (STB), personal digital assistant (PDA), cellular phone, network device, server, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by that machine. Additionally, although a single machine is illustrated, the term "machine" should also be understood to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any of the methods discussed herein.
[0184] Example computer system 1100 includes a processing device 1102, a main memory 1104 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM), etc.), a memory 1105 (e.g., flash memory, static random access memory (SRAM), etc.), an input / output system 1110, and a data storage system 1140, which communicate with each other via a bus 1130.
[0185] The processing device 1102 represents at least one general-purpose processing device, such as a microprocessor, a central processing unit, etc. More particularly, the processing device can be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor implementing other instruction sets, or a processor implementing a combination of instruction sets. The processing device 1102 can also be at least one special-purpose processing device, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, etc. The processing device 1102 is configured to execute instructions 1112 to perform the operations and steps discussed herein.
[0186] When the processing device 1102 is executing portions of the population frequency modeling system, the instructions 1112 include those portions of the population frequency modeling system 1150. Thus, the population frequency modeling system is shown in dashed lines as part of the instructions 1112 to illustrate that, sometimes, portions of the population frequency modeling system are executed by the processing device 1102. For example, when at least some portions of the population frequency modeling system are embodied in instructions for the processing device 1102 to perform the (one or more) methods described above, some of those instructions can be read into the processing device 1102 (e.g., read into an internal cache or other memory) from the main memory 1104 and / or the data storage system 1140. However, it is not required that all of the population frequency modeling system be included in the instructions 1112 at the same time, and portions of the population frequency modeling system are stored in at least one other component of the computer system 1100 at other times (e.g., when at least one portion of the population frequency modeling system is not being executed by the processing device 1102).
[0187] The computer system 1100 further includes a network interface device 1108 that communicates via a network 1120. The network interface device 1108 provides two-way data communication coupled to the network. For example, the network interface device 1108 can be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem that provides a data communication connection to a corresponding type of telephone line. As another example, the network interface device 1108 can be a Local Area Network (LAN) card to provide a data communication connection to a compatible LAN. A wireless link can also be implemented. In any such implementation, the network interface device 1108 can send and receive electrical, electromagnetic, or optical signals carrying digital data streams representing various types of information.
[0188] The network link can provide data communication to other data devices via at least one network. For example, the network link can provide a connection to the global packet data communication network commonly referred to as the "Internet", such as via a local network to a host or data equipment operated by an Internet Service Provider (ISP). Local area networks and the Internet use electrical, electromagnetic, or optical signals that carry digital data to and from the computer system 1100.
[0189] The computer system 1100 can send messages and receive data (including program code) via the network(s) and the network interface device 1108. In the Internet example, a server can transmit request code for an application via the Internet and the network interface device 1108. The received code can be executed by the processing device 1102 when it is received, and / or stored in the data storage system 1140 or other non-volatile memory for later execution.
[0190] The input / output system 1110 includes output devices for displaying information to a computer user, such as a display, e.g., a Liquid Crystal Display (LCD) or a touchscreen display, or a speaker, a haptic device, or other forms of output devices. The input / output system 1110 can include input devices, such as alphanumeric keys and other keys configured to pass information and command selections to the processing device 1102. Alternatively or additionally, the input device can include a cursor control, such as a mouse, a trackball, or cursor direction keys, for passing direction information and command selections to the processing device 1102 and for controlling the movement of a cursor on the display. Alternatively or additionally, the input device can include a microphone, a sensor, or a sensor array for passing sensed information to the processing device 1102. For example, the sensed information can include voice commands, audio signals, geographical location information, and / or digital images.
[0191] The data storage system 1140 includes a machine-readable storage medium 1142 (also referred to as a computer-readable medium) having stored thereon at least one set of instructions 1144 or software embodying any method or function described herein. During execution by the computer system 1100, the instructions 1144 may also reside completely or at least partially within the main memory 1104 and / or within the processing device 1102, which also constitutes a machine-readable storage medium.
[0192] In one embodiment, the instructions 1144 include instructions for implementing the functionality corresponding to a population frequency modeling system (e.g., Figure 10 the population frequency modeling system 1050 in
[0193] In Figure 11 a dashed line is used to indicate that it is not required that the population frequency modeling system be fully embodied in the instructions 1112, 1114, and 1144 simultaneously. In one example, a portion of the population frequency modeling system is embodied in the instructions 1144, the instructions 1144 are read into the main memory 1104 as the instructions 1114, and a portion of the instructions 1114 is read into the processing device 1102 as the instructions 1112 for execution. In another example, some portions of the population frequency modeling system are embodied in the instructions 1144, while other portions are embodied in the instructions 1114, and still other portions are embodied in the instructions 1112.
[0194] Although the machine-readable storage medium 1142 is shown as a single medium in the exemplary embodiment, the term "machine-readable storage medium" should be understood to include a single medium or multiple media storing at least one set of instructions. The term "machine-readable storage medium" should also be understood to include any medium that is capable of storing or encoding a set of instructions for execution by a machine and that causes the machine to perform any method of the present disclosure. Thus, the term "machine-readable storage medium" should be considered to include, but not be limited to, solid-state memory, optical media, and magnetic media.
[0195] Some of the foregoing detailed descriptions have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, considered to be a self-consistent sequence of operations leading to a desired result. These operations are those requiring physical manipulation of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, combined, compared, and otherwise manipulated. For general reasons, it has proven convenient at times to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc.
[0196] However, it should be remembered that all these and similar terms are associated with appropriate physical quantities and are merely convenient labels applied to these quantities. This disclosure may refer to actions and processes of a computer system or similar electronic computing device that manipulate and transform data represented as physical (electronic) quantities within the registers and memories of the computer system into other data similarly represented as physical quantities in the computer system memory or registers or other such information storage systems.
[0197] This disclosure also relates to an apparatus for performing the operations herein. The apparatus may be specially constructed for the intended purpose, or it may comprise a general purpose computer selectively activated or reconfigured by a computer program stored in the computer. For example, a computer system such as computing system 1100 or other data processing system may implement the techniques described above in response to its processor executing a computer program (e.g., a sequence of instructions) contained in a memory or other non-transitory machine-readable storage medium. Such a computer program may be stored in a computer-readable storage medium, such as but not limited to any type of disk, including floppy disks, optical disks, CD-ROMs, and magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, or any type of medium suitable for storing electronic instructions, each coupled to the computer system bus.
[0198] The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. According to the teachings herein, various general purpose systems may be used with the program, or it may prove convenient to construct more specialized apparatus to perform the method. The structure of various such systems will be set forth in the following description. Additionally, this disclosure is not described with reference to any particular programming language. It will be understood that a variety of programming languages may be used to implement the teachings of this disclosure as described herein.
[0199] This disclosure may be provided as a computer program product or software, which may include a machine-readable medium having instructions stored thereon that may be used to program a computer system (or other electronic device) to perform a process according to this disclosure. The machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). In some embodiments, the machine-readable (e.g., computer-readable) medium includes a machine (e.g., computer) readable storage medium, such as read-only memory (“ROM”), random access memory (“RAM”), disk storage media, optical storage media, flash memory components, etc.
[0200] Illustrative aspects of the techniques disclosed herein are provided below. Embodiments of these techniques may include any aspect described herein, or any combination of any aspect described herein, or any combination of any part of any aspect described herein.
[0201] In some aspects, the techniques described herein relate to a method for configuring a machine learning model to model population frequencies of variants for variant classification, the method comprising: applying a logistic regression model to a first population dataset of a first gene set, wherein for variants at positions within genes in the first gene set, the entries in the first population dataset include a feature set and a reference label, the feature set including at least one gene-level feature, at least one variant-level feature, and at least one population frequency meta-feature, the reference label indicating whether the variant is benign or pathogenic, wherein the at least one population frequency meta-feature quantifies a predicted value of allele frequency in the gene; comparing, for each entry in the first population dataset, the variant classification predicted value output by the logistic regression model with the expected variant classification value indicated by the reference label; and iteratively adjusting the value of at least one parameter or coefficient of the logistic regression model until the variant classification predicted value output by the logistic regression model meets at least one first performance criterion to produce a trained logistic regression model, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate value that meets at least one second performance criterion.
[0202] In some aspects, the techniques described herein further relate to a method comprising: using the trained logistic regression model to generate a predicted value as to whether a variant is benign or pathogenic.
[0203] In some aspects, the techniques described herein further relate to a method comprising: providing the predicted value as to whether a variant is benign or pathogenic to a clinician for use by the clinician in formulating a diagnosis of a patient.
[0204] In some aspects, the techniques described herein further relate to a method comprising: applying the trained logistic regression model to a second population dataset of a plurality of variants in a second plurality of genes; and for each variant in the second plurality of genes, receiving the variant classification predicted value output by the trained logistic regression model, and in response to the variant classification predicted value meeting at least a second performance criterion, storing the variant classification predicted value associated with the variant for retrieval via at least one query.
[0205] In some aspects, the techniques described herein relate to a method further comprising: calculating gene-level constraints; including the gene-level constraints in at least one gene-level feature; calculating allele frequencies; including the allele frequencies in at least one variant-level feature; including a mathematical combination of the gene-level constraints and the allele frequencies in at least one population frequency meta-feature; and applying a logistic regression model to a feature set including the mathematical combination of the gene-level constraints and the allele frequencies, wherein based on the feature set including the mathematical combination of the gene-level constraints and the allele frequencies, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0206] In some aspects, the techniques described herein relate to a method further comprising: calculating allele frequencies by calculating a binomial proportion of confidence values associated with the allele frequencies.
[0207] In some aspects, the techniques described herein relate to a method further comprising: calculating a mathematical combination of the gene-level constraints and the allele frequencies by dividing the allele frequencies by the gene-level constraints to produce a quotient and taking the exponent of the quotient.
[0208] In some aspects, the techniques described herein relate to a method further comprising: calculating, for a gene, an exponent of a ratio of synonymous missense variants to non-synonymous missense variants; including the exponent of the ratio of synonymous missense variants to non-synonymous missense variants in at least one population frequency meta-feature; and applying a logistic regression model to a feature set including the exponent of the ratio of synonymous missense variants to non-synonymous missense variants, wherein based on the feature set including the exponent of the ratio of synonymous missense variants to non-synonymous missense variants, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0209] In some aspects, the techniques described herein relate to a method further comprising: selecting no more than thirty features as a feature set; and applying a logistic regression model to the selected set of no more than thirty features, wherein based on the selected set of no more than thirty features, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0210] In some aspects, the techniques described herein relate to a method further comprising: calculating a fixation index, wherein the fixation index includes sub-population frequency data; including the fixation index in a feature set; and applying a logistic regression model to the feature set including the fixation index, wherein based on the feature set including the fixation index, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0211] In some aspects, the techniques described herein relate to a method further comprising: calculating a mathematical combination of subpopulation frequency data and population frequency data for a variant; including the mathematical combination of subpopulation frequency data and population frequency data in at least one variant-level feature; and applying a logistic regression model to a feature set comprising the mathematical combination of subpopulation frequency data and population frequency data, wherein, based on the feature set comprising the mathematical combination of subpopulation frequency data and population frequency data, a trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0212] In some aspects, the techniques described herein relate to a method wherein iteratively adjusting the value of at least one parameter of the logistic regression model comprises: adjusting the C value of the logistic regression model until the loss meets at least one first performance criterion to produce a trained logistic regression model.
[0213] In some aspects, the techniques described herein relate to a method further comprising: configuring the logistic regression model using L1 regularization to model population frequencies for variant classification.
[0214] In some aspects, the techniques described herein relate to a method further comprising: estimating the first performance criterion using mean squared error or area under the receiver operating characteristic curve.
[0215] In some aspects, the techniques described herein relate to a method further comprising: determining the second performance criterion using at least one of: a decision boundary test, a feature weight test, or a comparison to a baseline variant classification value.
[0216] In some aspects, the techniques described herein relate to a method further comprising: using the variant classification estimate output by the trained logistic regression model as an input to a classification framework.
[0217] In some aspects, the techniques described herein relate to a method wherein at least one population frequency meta-feature comprises the expected frequency distribution of known benign variants and known pathogenic variants within a gene.
[0218] In some aspects, the techniques described herein relate to a system comprising: at least one processor; and at least one memory coupled to the at least one processor, wherein the at least one memory includes at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: applying a logistic regression model to a first population dataset of a first gene set, wherein for variants at positions within genes in the first gene set, an entry in the first population dataset includes a feature set and a reference label, the feature set including at least one gene-level feature, at least one variant-level feature, and at least one population frequency meta-feature, the reference label indicating whether the variant is benign or pathogenic, wherein the at least one population frequency meta-feature quantifies a predicted value of allele frequency in the gene; for each entry in the first population dataset, comparing a variant classification prediction value output by the logistic regression model with an expected variant classification value indicated by the reference label; and adjusting a value of at least one parameter or coefficient of the logistic regression model until at least one first performance criterion is satisfied to produce a trained logistic regression model, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate value that satisfies at least one second performance criterion.
[0219] In some aspects, the techniques described herein relate to a system, wherein the at least one memory further includes at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: using the trained logistic regression model to generate a prediction value as to whether a variant is benign or pathogenic.
[0220] In some aspects, the techniques described herein relate to a system, wherein the at least one memory further includes at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: providing a clinician with a prediction value as to whether a variant is benign or pathogenic for use by the clinician in formulating a diagnosis of a patient.
[0221] In some aspects, the techniques described herein relate to a system, wherein the at least one memory further includes at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: applying a trained logistic regression model to a second population dataset of variants among a second plurality of genes; and for each variant among the second plurality of genes, receiving a variant classification prediction value output by the trained logistic regression model, and in response to the variant classification prediction value meeting at least a second performance criterion, storing the variant classification prediction value associated with the variant for retrieval via at least one query.
[0222] In some aspects, the techniques described herein relate to a system, wherein the at least one memory further includes at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: calculating gene-level constraints; including the gene-level constraints in at least one gene-level feature; calculating allele frequencies; including the allele frequencies in at least one variant-level feature; including a mathematical combination of the gene-level constraints and the allele frequencies in at least one population frequency meta-feature; and applying a logistic regression model to a feature set including the mathematical combination of the gene-level constraints and the allele frequencies, wherein based on the feature set including the mathematical combination of the gene-level constraints and the allele frequencies, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0223] In some aspects, the techniques described herein relate to a system, wherein the at least one memory further includes at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: calculating allele frequencies by calculating a binomial proportion of confidence values associated with the allele frequencies.
[0224] In some aspects, the techniques described herein relate to a system, wherein the at least one memory further includes at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: calculating a mathematical combination of the gene-level constraints and the allele frequencies by dividing the allele frequencies by the gene-level constraints to produce a quotient and taking the exponent of the quotient.
[0225] In some aspects, the techniques described herein relate to a system, wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: calculating an exponent of a ratio of synonymous missense variants to non-synonymous missense variants for a gene; including the exponent of the ratio of synonymous missense variants to non-synonymous missense variants in at least one population frequency meta-feature; and applying a logistic regression model to a feature set that includes the exponent of the ratio of synonymous missense variants to non-synonymous missense variants, wherein based on the feature set that includes the exponent of the ratio of synonymous missense variants to non-synonymous missense variants, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0226] In some aspects, the techniques described herein relate to a system, wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: selecting no more than thirty features as a feature set; and applying a logistic regression model to the selected set of no more than thirty features, wherein based on the selected set of no more than thirty features, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0227] In some aspects, the techniques described herein relate to a system, wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: calculating a fixed exponent, wherein the fixed exponent includes sub-population frequency data; including the fixed exponent in a feature set; and applying a logistic regression model to the feature set that includes the fixed exponent, wherein based on the feature set that includes the fixed exponent, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0228] In some aspects, the techniques described herein relate to a system, wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: calculating a mathematical combination of subpopulation frequency data and population frequency data for a variant; including the mathematical combination of subpopulation frequency data and population frequency data in at least one variant-level feature; and applying a logistic regression model to a feature set comprising the mathematical combination of subpopulation frequency data and population frequency data, wherein, based on the feature set comprising the mathematical combination of subpopulation frequency data and population frequency data, a trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0229] In some aspects, the techniques described herein relate to a system, wherein iteratively adjusting the value of at least one parameter of a logistic regression model comprises: adjusting the C value of the logistic regression model until a loss meets at least one first performance criterion to produce a trained logistic regression model.
[0230] In some aspects, the techniques described herein relate to a system, wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: configuring a logistic regression model using L1 regularization to model population frequencies for variant classification.
[0231] In some aspects, the techniques described herein relate to a system, wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: estimating the first performance criterion using mean squared error or area under the receiver operating characteristic curve.
[0232] In some aspects, the techniques described herein relate to a system, wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: determining the second performance criterion using at least one of: a decision boundary test, a feature weight test, or a comparison to a baseline variant classification value.
[0233] In some aspects, the techniques described herein relate to a system, wherein the at least one memory further includes at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: using the variant classification estimates output by a trained logistic regression model as an input to a classification framework.
[0234] In some aspects, the techniques described herein relate to a system, wherein at least one population frequency meta-feature includes the expected frequency distributions of known benign variants and known pathogenic variants within a gene.
[0235] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium including at least one instruction that, when executed by at least one processor, causes the at least one processor to perform an operation, the operation including: applying a logistic regression model to a first population data set of a first gene set, wherein for variants at positions within genes in the first gene set, the entries in the first population data set include a feature set and a reference label, the feature set including at least one population frequency meta-feature, the reference label indicating whether the variant is benign or pathogenic, wherein the at least one population frequency meta-feature quantifies a predicted value of the allele frequency in the gene; comparing, for each entry in the first population data set, the variant classification prediction value output by the logistic regression model with the expected variant classification value indicated by the reference label; and adjusting the value of at least one parameter or coefficient of the logistic regression model until at least one first performance criterion is met to produce a trained logistic regression model, wherein the trained logistic regression model is capable of outputting variant pathogenicity estimates that satisfy at least one second performance criterion.
[0236] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium further including at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: using the trained logistic regression model to generate a prediction value as to whether a variant is benign or pathogenic.
[0237] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium further including at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: providing a clinician with a prediction value as to whether a variant is benign or pathogenic for use by the clinician in formulating a diagnosis of a patient.
[0238] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: applying a trained logistic regression model to a second population dataset of variants in a second plurality of genes; and for each variant in the second plurality of genes, receiving a variant classification prediction value output by the trained logistic regression model and, in response to the variant classification prediction value meeting at least a second performance criterion, storing the variant classification prediction value associated with the variant for retrieval via at least one query.
[0239] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: calculating gene-level constraints; calculating allele frequencies; including a mathematical combination of the gene-level constraints and the allele frequencies in at least one population frequency meta-feature; and applying a logistic regression model to a feature set including the mathematical combination of the gene-level constraints and the allele frequencies, wherein based on the feature set including the mathematical combination of the gene-level constraints and the allele frequencies, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0240] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: calculating allele frequencies by calculating a binomial proportion of confidence values associated with the allele frequencies.
[0241] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: calculating a mathematical combination of the gene-level constraints and the allele frequencies by dividing the allele frequencies by the gene-level constraints to produce a quotient and taking the exponent of the quotient.
[0242] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: calculating an exponent of a ratio of synonymous missense variants to non-synonymous missense variants for a gene; including the exponent of the ratio of synonymous missense variants to non-synonymous missense variants in at least one population frequency meta-feature; and applying a logistic regression model to a feature set including the exponent of the ratio of synonymous missense variants to non-synonymous missense variants, wherein, based on the feature set including the exponent of the ratio of synonymous missense variants to non-synonymous missense variants, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0243] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: selecting no more than thirty features as a feature set; and applying a logistic regression model to the selected set of no more than thirty features, wherein, based on the selected set of no more than thirty features, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0244] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: calculating a fixed exponent, wherein the fixed exponent includes sub-population frequency data; including the fixed exponent in a feature set; and applying a logistic regression model to the feature set including the fixed exponent, wherein, based on the feature set including the fixed exponent, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0245] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: calculating a mathematical combination of sub-population frequency data and population frequency data for a variant; and applying a logistic regression model to a feature set including the mathematical combination of sub-population frequency data and population frequency data, wherein, based on the feature set including the mathematical combination of sub-population frequency data and population frequency data, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
[0246] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium, wherein adjusting a value of at least one parameter of a logistic regression model includes: adjusting the C value of the logistic regression model until a loss meets at least one first performance criterion to produce a trained logistic regression model.
[0247] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: configuring a logistic regression model using L1 regularization to model population frequencies for variant classification.
[0248] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: estimating the first performance criterion using mean squared error or area under the receiver operating characteristic curve.
[0249] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: determining a second performance criterion using at least one of: a decision boundary test, a feature weight test, or a comparison with a baseline variant classification value.
[0250] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: using the variant classification estimate output by the trained logistic regression model as an input to a classification framework.
[0251] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable medium, wherein at least one population frequency meta-feature includes the expected frequency distribution of known benign variants and known pathogenic variants within a gene.
[0252] In some aspects, the techniques described herein relate to any one or more of the aspects, steps, components, elements, processes, or limitations that are at least one of those described in the accompanying specification or shown in the accompanying drawings.
[0253] In the foregoing specification, embodiments of the present disclosure have been described with reference to specific example embodiments thereof. Obviously, various modifications can be made thereto without departing from the broader spirit and scope of the embodiments of the present disclosure as set forth in the following claims. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense.
Claims
1. A method for configuring a machine learning model to model population frequencies for variant classification, the method comprising: Applying a logistic regression model to a first population dataset of a first gene set, wherein for variants at positions within genes in the first gene set, an entry in the first population dataset includes a feature set and a reference label, the feature set including at least one gene-level feature, at least one variant-level feature, and at least one population frequency meta-feature, the reference label indicating whether the variant is benign or pathogenic, wherein the at least one population frequency meta-feature quantifies a predicted value of allele frequency in the gene; For each entry in the first population dataset, evaluating a variant classification prediction value output by the logistic regression model based on an expected variant classification value indicated by the reference label; And Iteratively adjusting a value of at least one parameter or coefficient of the logistic regression model until an output of a loss function satisfies at least one first performance criterion to produce a trained logistic regression model, the output of the loss function being calculated based on the variant classification prediction value output by the logistic regression model, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate value that satisfies at least one second performance criterion.
2. The method according to claim 1, further comprising: Using the trained logistic regression model to generate a prediction value as to whether the variant is benign or pathogenic.
3. The method according to claim 2, further comprising: Providing the prediction value as to whether the variant is benign or pathogenic to a clinician for use by the clinician in formulating a diagnosis of a patient.
4. The method according to claim 1, further comprising: Applying the trained logistic regression model to a second population dataset of multiple variants in a second plurality of genes; And For each variant in the second plurality of genes, receiving a variant classification prediction value output by the trained logistic regression model, and in response to the variant classification prediction value satisfying at least the second performance criterion, storing the variant classification prediction value associated with the variant for retrieval via at least one query.
5. The method according to claim 1, further comprising: Calculating gene-level constraints; Including the gene-level constraints in the at least one gene-level feature; Calculating allele frequencies; Including the allele frequencies in the at least one variant-level feature; Including a mathematical combination of the gene-level constraints and the allele frequencies in the at least one population frequency meta-feature; And Applying the logistic regression model to the feature set including the mathematical combination of the gene-level constraints and the allele frequencies, wherein based on the feature set including the mathematical combination of the gene-level constraints and the allele frequencies, the trained logistic regression model is capable of outputting a variant pathogenicity estimate value that satisfies the at least one second performance criterion.
6. The method according to claim 5, further comprising: The allele frequency is calculated by computing the binomial proportion of the confidence value associated with the allele frequency.
7. The method according to claim 5, further comprising: calculating the mathematical combination of the gene-level constraint and the allele frequency by dividing the allele frequency by the gene-level constraint to produce a quotient and taking the exponent of the quotient.
8. The method according to claim 1, further comprising: calculating, for the gene, the exponent of the ratio of the synonymous missense variants to the non-synonymous missense variants; including the exponent of the ratio of the synonymous missense variants to the non-synonymous missense variants in the at least one population frequency meta-feature; and applying the logistic regression model to the feature set including the exponent of the ratio of the synonymous missense variants to the non-synonymous missense variants, wherein based on the feature set including the exponent of the ratio of the synonymous missense variants to the non-synonymous missense variants, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion.
9. The method according to claim 1, further comprising: selecting no more than thirty features as the feature set; and applying the logistic regression model to the selected set of no more than thirty features, wherein based on the selected set of no more than thirty features, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion.
10. The method according to claim 1, further comprising: calculating a fixed exponent, wherein the fixed exponent includes sub-population frequency data; including the fixed exponent in the feature set; and applying the logistic regression model to the feature set including the fixed exponent, wherein based on the feature set including the fixed exponent, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion.
11. The method according to claim 1, further comprising: calculating, for a variant, the mathematical combination of the sub-population frequency data and the population frequency data; including the mathematical combination of the sub-population frequency data and the population frequency data in the at least one variant-level feature; and applying the logistic regression model to the feature set including the mathematical combination of the sub-population frequency data and the population frequency data, wherein based on the feature set including the mathematical combination of the sub-population frequency data and the population frequency data, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion.
12. The method according to claim 1, wherein iteratively adjusting the value of at least one parameter of the logistic regression model comprises: Adjusting the C value of the logistic regression model until the output of the loss function meets the at least one first performance criterion to produce the trained logistic regression model.
13. The method according to claim 1, further comprising: configuring the logistic regression model using L1 regularization to model the population frequency for variant classification.
14. The method according to claim 1, further comprising: Estimate the first performance criterion using the mean squared error or the area under the receiver operating characteristic curve.
15. The method according to claim 1, further comprising: Determine the second performance criterion using at least one of: a decision boundary test, a feature weight test, or a comparison with a reference variant classification value.
16. The method according to claim 1, further comprising: Use the variant classification estimate output by the trained logistic regression model as an input to a variant classification framework.
17. The method according to claim 1, wherein the at least one population frequency meta-feature includes the expected frequency distributions of known benign variants and known pathogenic variants within the gene.
18. A system comprising: At least one processor; And At least one memory coupled to the at least one processor, wherein the at least one memory includes at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: Apply a logistic regression model to a first population data set of a first gene set, wherein for variants at positions within genes in the first gene set, the entries in the first population data set include a feature set and a reference label, the feature set including at least one gene-level feature, at least one variant-level feature, and at least one population frequency meta-feature, the reference label indicating whether the variant is benign or pathogenic, wherein the at least one population frequency meta-feature quantifies the predicted value of the allele frequency in the gene; For each entry in the first population data set, evaluate the variant classification prediction value output by the logistic regression model based on the expected variant classification value indicated by the reference label; and Adjust the value of at least one parameter or coefficient of the logistic regression model until at least one first performance criterion is satisfied to produce a trained logistic regression model, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that satisfies at least one second performance criterion.
19. The system according to claim 18, wherein the at least one memory further includes at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: Use the trained logistic regression model to generate a prediction value as to whether the variant is benign or pathogenic.
20. The system according to claim 19, wherein the at least one memory further includes at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: Provide the clinician with the prediction value as to whether the variant is benign or pathogenic for use by the clinician in formulating a diagnosis for a patient.
21. The system according to claim 18, wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: applying the trained logistic regression model to a second population dataset of variants among a second plurality of genes; and for each variant among the second plurality of genes, receiving a variant classification prediction value output by the trained logistic regression model and, in response to the variant classification prediction value meeting at least the second performance criterion, storing the variant classification prediction value associated with the variant for retrieval via at least one query.
22. The system according to claim 18, wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: calculating gene-level constraints; including the gene-level constraints in the at least one gene-level feature; calculating allele frequencies; including the allele frequencies in the at least one variant-level feature; including a mathematical combination of the gene-level constraints and the allele frequencies in the at least one population frequency meta-feature; and applying the logistic regression model to the feature set including the mathematical combination of the gene-level constraints and the allele frequencies, wherein based on the feature set including the mathematical combination of the gene-level constraints and the allele frequencies, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion.
23. The system according to claim 22, wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: calculating the allele frequencies by calculating a binomial proportion of confidence values associated with the allele frequencies.
24. The system according to claim 22, wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: calculating the mathematical combination of the gene-level constraints and the allele frequencies by dividing the allele frequencies by the gene-level constraints to produce a quotient and taking the exponent of the quotient.
25. The system according to claim 18, wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: calculating an exponent of a ratio of synonymous missense variants to non-synonymous missense variants for the gene; Including the exponent of the ratio of the synonymous missense variants to the non-synonymous missense variants in the at least one population frequency meta-feature; And Applying the logistic regression model to the feature set including the exponent of the ratio of the synonymous missense variants to the non-synonymous missense variants, wherein based on the feature set including the exponent of the ratio of the synonymous missense variants to the non-synonymous missense variants, the trained logistic regression model can output a variant pathogenicity estimate that meets the at least one second performance criterion.
26. The system according to claim 18, wherein the at least one memory further includes at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: Selecting no more than thirty features as the feature set; And Applying the logistic regression model to the set of no more than thirty selected features, wherein based on the set of no more than thirty selected features, the trained logistic regression model can output a variant pathogenicity estimate that meets the at least one second performance criterion.
27. The system according to claim 18, wherein the at least one memory further includes at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: Calculating a fixed exponent, wherein the fixed exponent includes sub-population frequency data; Including the fixed exponent in the feature set; And Applying the logistic regression model to the feature set including the fixed exponent, wherein based on the feature set including the fixed exponent, the trained logistic regression model can output a variant pathogenicity estimate that meets the at least one second performance criterion.
28. The system according to claim 18, wherein the at least one memory further includes at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: Calculating a mathematical combination of sub-population frequency data and population frequency data for a variant; Including the mathematical combination of sub-population frequency data and population frequency data in the at least one variant-level feature; And Applying the logistic regression model to the feature set including the mathematical combination of sub-population frequency data and population frequency data, wherein based on the feature set including the mathematical combination of sub-population frequency data and population frequency data, the trained logistic regression model can output a variant pathogenicity estimate that meets the at least one second performance criterion.
29. The system according to claim 18, wherein iteratively adjusting the value of at least one parameter of the logistic regression model comprises: Adjusting the C value of the logistic regression model until the output of the loss function meets the at least one first performance criterion to produce the trained logistic regression model.
30. The system according to claim 18, wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: Configuring the logistic regression model using L1 regularization to model population frequencies for variant classification.
31. The system according to claim 18, wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: Estimating the first performance criterion using mean squared error or area under the receiver operating characteristic curve.
32. The system according to claim 18, wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: Determining the second performance criterion using at least one of: a decision boundary test, a feature weight test, or a comparison with a reference variant classification value.
33. The system according to claim 18, wherein the at least one memory further comprises at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: Using the variant classification estimate output by the trained logistic regression model as an input to a variant classification framework.
34. The system according to claim 18, wherein the at least one population frequency meta-feature comprises the expected frequency distribution of known benign variants and known pathogenic variants within the gene.
35. At least one non-transitory machine-readable medium comprising at least one instruction that, when executed by at least one processor, causes the at least one processor to perform operations comprising: Applying a logistic regression model to a first population data set of a first gene set, wherein for variants at positions within genes in the first gene set, the entries in the first population data set comprise a feature set and a reference label, the feature set comprising at least one population frequency meta-feature, the reference label indicating whether the variant is benign or pathogenic, wherein the at least one population frequency meta-feature quantifies a predicted value of allele frequency in the gene; For each entry in the first population data set, evaluating the variant classification prediction value output by the logistic regression model based on the expected variant classification value indicated by the reference label; And Adjusting the value of at least one parameter or coefficient of the logistic regression model until at least one first performance criterion is met to produce a trained logistic regression model, wherein the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets at least one second performance criterion.
36. The at least one non-transitory machine-readable medium according to claim 35, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: Using the trained logistic regression model to generate a predicted value as to whether the variant is benign or pathogenic.
37. The at least one non-transitory machine-readable medium according to claim 36, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: Providing the clinician with the predicted value as to whether the variant is benign or pathogenic for use by the clinician in formulating a diagnosis of a patient.
38. The at least one non-transitory machine-readable medium according to claim 35, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: Applying the trained logistic regression model to a second population dataset of variants in a second plurality of genes; And For each variant in the second plurality of genes, receiving a variant classification predicted value output by the trained logistic regression model and, in response to the variant classification predicted value meeting at least the second performance criterion, storing the variant classification predicted value associated with the variant for retrieval via at least one query.
39. The at least one non-transitory machine-readable medium according to claim 35, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: Calculating gene-level constraints; Calculating allele frequencies; Including a mathematical combination of the gene-level constraints and the allele frequencies in the at least one population frequency meta-feature; And Applying the logistic regression model to the feature set including the mathematical combination of the gene-level constraints and the allele frequencies, wherein based on the feature set including the mathematical combination of the gene-level constraints and the allele frequencies, the trained logistic regression model is capable of outputting a variant pathogenicity estimate that meets the at least one second performance criterion.
40. The at least one non-transitory machine-readable medium according to claim 39, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: Calculating the allele frequency by calculating a binomial proportion of a confidence value associated with the allele frequency.
41. The at least one non-transitory machine-readable medium according to claim 39, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: Calculating the mathematical combination of the gene-level constraint and the allele frequency by: dividing the allele frequency by the gene-level constraint to produce a quotient, and taking the exponent of the quotient.
42. The at least one non-transitory machine-readable medium according to claim 35, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: Calculating an exponent of a ratio of a synonymous missense variant to a non-synonymous missense variant for the gene; Including the exponent of the ratio of the synonymous missense variant to the non-synonymous missense variant in the at least one population frequency meta-feature; And Applying the logistic regression model to the feature set including the exponent of the ratio of the synonymous missense variant to the non-synonymous missense variant, wherein based on the feature set including the exponent of the ratio of the synonymous missense variant to the non-synonymous missense variant, the trained logistic regression model can output a variant pathogenicity estimate that meets the at least one second performance criterion.
43. The at least one non-transitory machine-readable medium according to claim 35, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: Selecting no more than thirty features as the feature set; And Applying the logistic regression model to the set of the selected no more than thirty features, wherein based on the set of the selected no more than thirty features, the trained logistic regression model can output a variant pathogenicity estimate that meets the at least one second performance criterion.
44. The at least one non-transitory machine-readable medium according to claim 35, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: Calculating a fixed exponent, wherein the fixed exponent includes sub-population frequency data; Including the fixed exponent in the feature set; And Applying the logistic regression model to the feature set including the fixed exponent, wherein based on the feature set including the fixed exponent, the trained logistic regression model can output a variant pathogenicity estimate that meets the at least one second performance criterion.
45. The at least one non-transitory machine-readable medium according to claim 35, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: Calculate a mathematical combination of sub-population frequency data and population frequency data for the variants; and Apply the logistic regression model to the feature set comprising the mathematical combination of sub-population frequency data and population frequency data, wherein, based on the feature set comprising the mathematical combination of sub-population frequency data and population frequency data, the trained logistic regression model is capable of outputting an estimate of variant pathogenicity that meets the at least one second performance criterion.
46. The at least one non-transitory machine-readable medium according to claim 35, wherein adjusting a value of at least one parameter of the logistic regression model comprises: Adjust the C value of the logistic regression model until the output of the loss function meets the at least one first performance criterion to produce the trained logistic regression model.
47. The at least one non-transitory machine-readable medium according to claim 35, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: Configure the logistic regression model using L1 regularization to model population frequencies for variant classification.
48. The at least one non-transitory machine-readable medium according to claim 35, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: Estimate the at least one first performance criterion using mean squared error or area under the receiver operating characteristic curve.
49. The at least one non-transitory machine-readable medium according to claim 35, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: Determine the second performance criterion using at least one of: a decision boundary test, a feature weight test, or a comparison with a reference variant classification value.
50. The at least one non-transitory machine-readable medium according to claim 35, further comprising at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: Use the variant classification estimate output by the trained logistic regression model as an input to a variant classification framework.
51. The at least one non-transitory machine-readable medium according to claim 35, wherein the at least one population frequency meta-feature comprises the expected frequency distribution of known benign variants and known pathogenic variants within the gene.