Clinical Variant Modeling

JP2026530011APending Publication Date: 2026-09-03LABORATORY CORPORATION OF AMERICA HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026513085
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-07
Filing Date
2024-08-30
Publication Date
2026-09-03

Smart Images

  • Figure 2026530011000001_ABST
    Figure 2026530011000001_ABST
Patent Text Reader

Abstract

For example, using phenotypic data including natural language text and genotypic data, at least one machine learning model can be constructed. At least one machine learning model can generate and output a patient score. Given phenotypic data including natural language text, the patient score may include a machine learning likelihood that the patient has a genetic condition. In some embodiments, a Bayesian causal model can be constructed using the patient score and genotypic data to generate and output a variant score. The variant score may include a machine learning likelihood that a variant of a gene is pathogenic with respect to the genetic condition. The patient score and / or variant score may be provided to one or more downstream systems, devices, processes, components, or models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross-Reference to Related Applications This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 579,939 filed on August 31, 2023 and U.S. Provisional Patent Application No. 63 / 562,696 filed on March 7, 2024, and each of said U.S. provisional patent applications is incorporated herein by reference in its entirety.

[0002] The technical field to which the present application relates is genetic testing. Another technical field to which the present application relates is a machine learning-based variant classification system. Background Art

[0003] A genetic variant is a difference in DNA sequence between individuals in a population. There are many different types of variants, including structural variations, single nucleotide polymorphisms, insertion and deletion mutations, copy number variations, and translocations and inversions.

[0004] Gene sequencing technology continues to evolve rapidly. High-throughput sequencing technology has increasingly enabled genetic testing ranging from genotyping of genetic diseases to single-gene testing, gene panel testing, exome testing, genome testing, transcriptome testing and epigenetic assays. The increasing complexity of analysis and interpretation of clinical genetic testing, as well as the increasing volume of testing, have brought about new challenges in the interpretation of sequence variants.

[0005] For example, clinical molecular laboratories are increasingly detecting novel sequence variants in the process of testing patient specimens in response to the rapid increase in the number of genes associated with genetic diseases. While some phenotypes are associated with a single gene, many are associated with multiple genes.

[0006] Variant classification refers to the process of classifying gene variants based on evidence supporting or rejecting a causal relationship with a disease. The clinical significance of any given sequence variant is placed on a gradient from those that are almost certainly pathogenic to those that are almost certainly benign.

[0007] Variant classification is not a diagnosis in itself, but it can be used by clinicians to make diagnostic decisions. [Brief explanation of the drawing]

[0008] This disclosure will be better understood from the following detailed description and the accompanying drawings of various embodiments of this disclosure. The drawings are for illustrative and illustrative purposes only and should not be construed as limiting this disclosure to the specific embodiments shown.

[0009] [Figure 1] Examples of variant scoring processes according to several embodiments of this disclosure are shown.

[0010] [Figure 2] This disclosure provides examples of processes for generating patient scores using natural language processing-based models, according to several embodiments of this disclosure.

[0011] [Figure 3A] This disclosure illustrates an exemplary process for generating text scores using a natural language processing-based model, according to several embodiments of this disclosure.

[0012] [Figure 3B] This disclosure illustrates an exemplary process for generating patient scores using a natural language processing-based model, according to several embodiments of this disclosure.

[0013] [Figure 4] This disclosure illustrates an exemplary process for constructing a clinical variant model using machine learning, according to several embodiments of this disclosure.

[0014] [Figure 5A] An example of patient score calculation according to some embodiments of the present disclosure is shown.

[0015] [Figure 5B] An example of patient score calculation according to some embodiments of the present disclosure is shown.

[0016] [Figure 5C] An example of patient score calculation using machine learning feature weights according to some embodiments of the present disclosure is shown.

[0017] [Figure 6A] An example of variant score calculation using a Bayesian model according to some embodiments of the present disclosure is shown.

[0018] [Figure 6B] An example of the distribution of patient scores for a variant according to some embodiments of the present disclosure is shown.

[0019] [Figure 6D] An example of the distribution of patient scores and variant scores for classified variants according to some embodiments of the present disclosure is shown.

[0020] [Figure 6E] An example of the distribution of patient scores and variant scores for classified variants according to some embodiments of the present disclosure is shown.

[0021] [Figure 6F] An example of the distribution of patient scores and variant scores according to some embodiments of the present disclosure is shown.

[0022] [Figure 6G] An example of the distribution of patient scores and classified variants according to some embodiments of the present disclosure is shown.

[0023] [Figure 7A] Examples of patient score and variant score distributions for classified variants according to some embodiments of this disclosure are shown.

[0024] [Figure 7B] Examples of patient score and variant score distributions for classified variants according to some embodiments of this disclosure are shown.

[0025] [Figure 8A] Examples of performance data for patient score generators and variant score generators according to some embodiments of this disclosure are shown.

[0026] [Figure 8B] Examples of clinical variant modeling and classification systems according to several embodiments of this disclosure are shown.

[0027] [Figure 8C] Examples of variant classification systems, including clinical variant modeling, according to some embodiments of this disclosure are shown.

[0028] [Figure 9A] This disclosure describes methods for clinical variant modeling according to several embodiments.

[0029] [Figure 9B] This disclosure describes methods for clinical variant modeling according to several embodiments.

[0030] [Figure 9C] This disclosure describes methods for clinical variant modeling according to several embodiments.

[0031] [Figure 10] This disclosure illustrates exemplary computing systems, including clinical variant modeling, according to several embodiments of this disclosure.

[0032] [Figure 11] This aspect of the disclosure is a block diagram of an exemplary computer system in which it can operate. [Modes for carrying out the invention]

[0033] Genetic variants can be classified as pathogenic (i.e., disease-causing), benign (not disease-causing), likely pathogenic, likely benign, or of unknown significance (VUS). Today, many variants are classified as VUS, primarily because the information available about these variants is insufficient to make a proper classification. As genetic testing is increasingly adopted in healthcare for disease diagnosis and management, the field of clinical genomics is increasingly confronted with novel variants, including both common and rare novel variants, which also require classification. Simultaneously, the expansion of genetic testing brings with it a growing wealth of data, including clinical data.

[0034] Embodiments of the clinical variant modeling approach described herein address these and / or other challenges by including one or more patient score generators that generate patient scores, or variant score generators that generate variant scores based on patient scores. While this disclosure describes approaches that include both patient score generators and variant score generators, either the patient score generator or the variant score generator may be used independently of the other components. For example, in some applications, patient scores generated by a patient score generator may be useful separately and independently of variant scores generated by a variant score generator. As another example, a variant score generator may generate variant scores based on patient scores obtained from another source, such as a database or a different type of patient score generator, rather than from the patient score generators described herein.

[0035] The clinical variant modeling approaches described herein (e.g., the patient score generator and / or the variant score generator described herein) utilize diverse genotypes and anonymized clinical data from a population of genetically tested patients to classify or reclassify both novel and previously observed variants, thereby reducing variants of unknown significance and / or improving the accuracy of previous variant classifications.

[0036] The described embodiments of the clinical variant modeling approach include a variant score generator configured as a Bayesian causal model, which has proven highly accurate for incorporating clinical evidence into variant classification at large scale. The described approach can be used to improve variant classification of genes and diseases that have not yet been tested, and to leverage information at a continuously growing scale of clinical datasets. The described approach can improve variant classification and reduce uncertainty in genetic testing in a scalable manner.

[0037] The described clinical variant modeling approach is robust to dataset size. For example, the described approach can produce reliable predictions for both large datasets (e.g., 3-4 million patients) and smaller datasets. The variant score generator allows domain expertise to be leveraged in situations where data is sparse. In the case of patient score models, the curriculum learning approach improves performance in sparse regimes.

[0038] The purpose of clinical genetic testing is to assess the risk of hereditary diseases or to confirm the diagnosis of hereditary diseases. Distinguishing pathogenic DNA (deoxyribonucleic acid) variants from benign variants is a crucial challenge in the process of genetic testing. Clinical genetic testing laboratories are facing this challenge further as more individuals undergo genetic testing, encountering a growing number of novel and rare variants. These variants are not always adequately documented in published literature or the CLINVAR database and are therefore often classified as variants of unknown significance (VUS).

[0039] Clinical data is one of the strongest forms of evidence for distinguishing pathogenic variants from benign variants. For example, clinical data may include observations of variants in patients with a clearly defined disease, observations of variants in patients clearly unaffected by a clearly defined disease, variants co-separating with a clearly defined disease among affected individuals within a family, or observations of de novo occurrence of a variant. However, incorporating clinical data into variant classification often presents challenges. For example, a complete and relevant medical history is not always available at the time of genetic testing, which makes it difficult to distinguish between affected and unaffected individuals with missing data.

[0040] Secondly, many hereditary disorders include symptoms that can be associated with common sporadic diseases such as cancer and cardiovascular disease, limiting the ability to identify molecular causes and establish genetic etiology. Finally, not all hereditary disorders exhibit full penetrance, making it difficult to determine when a variant is not associated with a hereditary disorder and is therefore benign. These challenges are particularly severe when clinical data are reviewed on a case-by-case basis for variant classification. However, increased access to a larger set of clinical health information and genetic test results can help overcome these challenges.

[0041] The disclosed approach can leverage accumulated genotyping and clinical data from populations including individuals of diverse racial and ethnic backgrounds introduced to a wide range of clinical genetic testing. The datasets are large and clinically diverse: in some embodiments, they include over 100 million words of submitted clinical descriptions (e.g., personal and family history, indications for testing) about tested patients, as well as over 2 million unique variants observed across more than 3,900 genes.

[0042] The disclosed clinical variant modeling approach can maximize the usefulness of clinical datasets for improving variant classification and reducing VUS. Embodiments of the described approach for clinical variant modeling use machine learning to determine patterns from clinical data available to millions of patients and apply this learning precisely as evidence for variant classification using a Bayesian approach. The described approach can, but is not limited to, machine learning the relationships (e.g., statistical correlations) between different variables and associated outcomes, including disease prevalence, age at examination, potential phenotypic mimicry, and missing data on clinical patient records (e.g., test application forms).

[0043] Some embodiments of clinical variant modeling described herein involve two distinct but connected sequential machine learning (ML) steps. The first step involves estimating the probability that a patient undergoing genetic testing has a particular genetic condition. This probability, referred to as the patient score, incorporates clinical and demographic information from the commissioned provider, including reported signs and symptoms, ICD-10 codes, age at testing, and family history. The patient score is estimated by comparing and distinguishing the clinical profiles of patients with a positive molecular diagnosis with those with a negative molecular diagnosis. In the second step of clinical variant modeling, a Bayesian inference model learns the distribution of patient scores that may be associated with benign or pathogenic variants. The estimated probability that a variant is pathogenic is referred to as the variant score.

[0044] More specifically, the described embodiment of the clinical variant modeling approach uses a combination of natural language processing (NLP) and Bayesian inference to predict the pathogenicity of a gene variant using clinical data provided in patient records (e.g., test request forms). The resulting clinical variant modeling system is designed to learn which clinical features can distinguish patients with molecular diagnoses from genotype-negative controls.

[0045] As used herein, clinical variant modeling may refer to a modeling pipeline comprising one or more patient score generators and variant score generators, and clinical variant model may refer to a patient score generator, a variant score generator, or a combination of a patient score generator and a variant score generator. Either or both of the patient score generator and variant score generator may include one or more machine learning models. For example, as will be described in more detail below, an embodiment of the patient score generator includes at least two machine learning models, such as NLP components and a tree-based scoring model. Thus, for example, some embodiments of the clinical variant model may include up to or at least three machine learning models (e.g., a first scoring model such as an NLP-based scoring model, a second scoring model such as a tree-based scoring model, and a third scoring model such as a Bayesian inference model).

[0046] In some embodiments, the clinical variant models are state-specific (e.g., a neurofibromatosis type 1 model including NF1, or a Lynch syndrome model including MLH1, MSH2, MSH6, PMS2, and EPCAM). As a result, each clinical variant model can, in some embodiments, be trained and tested separately for specific states and their associated genes.

[0047] In some embodiments, the clinical variant model is gene-specific (for example, the clinical variant model is specific to MMR or another gene). As a result, in some embodiments, each clinical variant model can be trained and tested separately for one or more specific genes.

[0048] The described clinical variant modeling approach differs from existing methods in several ways. Firstly, existing methods for incorporating clinical data for a single gene state evaluate each patient's clinical information on a case-by-case basis. While this approach may be effective for certain pathologies, it struggles with some hereditary conditions that may be associated with common sporadic diseases (i.e., high phenotypic mimicry rates such as cancer and cardiovascular disease), exhibit incomplete penetrance, and show variable phenotypic expression. In contrast, the described clinical variant modeling approach can leverage a large clinical database from genetically tested patients (e.g., over 4 million patients), allowing for the sharing of gene variants and the differentiation between the number of patients who appear to have the disease phenotype in question and those who do not (in the case of phenotypic mimicry).

[0049] Secondly, existing methods are generally binary (e.g., the prediction either meets the diagnostic criteria or does not). In contrast, clinical variant models constructed using the described approach can calculate the probability that a given individual is affected based on clinical information on a continuous scale, and can also calculate the probability that a variant is pathogenic based on the continuous distribution of affected and unaffected states of all patients with the same variant. As a result, scientists can use available clinical data in a much more nuanced way.

[0050] Next, the described clinical variant modeling approach can learn patterns present in clinical data within a patient population cohort and apply (or generalize) those learned patterns to other patients. This capability can be particularly useful in revealing patterns and trends in clinical data that may not be apparent by other means, such as the use of abbreviations by clinicians, abbreviations, terminology in other languages, as well as the frequency of absence of clinical information (whether or not the patient is affected), age distribution, and ICD-10 code usage patterns. As a result, pathogenicity predictions can be generalized to specific but broader patient populations.

[0051] In contrast to existing methods that rely on patient data reported in the literature, the described clinical variant modeling approach relies solely on patients observed through examination. Therefore, existing methods may be biased towards cohorts of patients more likely to be reported in the literature (often more severely affected or with earlier onset than broader populations), whereas the described approach is not in this biased manner.

[0052] Additionally, because the clinical variant modeling described is a machine learning-based approach, the models can be periodically updated as more clinician-patient data and / or more variant information becomes available to the medical genetics community.

[0053] Clinical variant modeling is a machine learning approach that learns specific features (e.g., specific ICD-10 codes, free-text clinical information, free-text family history information, etc.) that best distinguish patients with molecular diagnoses from those that are genotype-negative; therefore, each feature is not necessarily weighted equally. Furthermore, the weighting of each feature can change from one genetic state to the next. A clinical variant model for a given disease state can automatically learn which features best predict pathogenicity and weight them appropriately.

[0054] A technical challenge to using clinical data for variant classification is that the format, quantity, and / or quality of information contained in patient records can vary dramatically from patient to patient. For example, some patient records may have detailed textual descriptions of indications and / or family history, while others may have only a few vague keywords within these fields, or no information at all. The described clinical variant modeling approach can accommodate the variability of information in available patient records, including missing information. Because different features are weighted differently for each gene / pathology, the impact of missing or sparse information depends on the specific information missing for a given model. Since models learn at a macro level, when a feature is missing, the model learns that there is a certain level of missing information.

[0055] The described embodiments of the clinical variant model have been trained, tested, and validated for clinical variant classification based on clinical information available for patients and variants in a dedicated database. Future plans include updating the model as the clinical cohort grows, with more genotypic and phenotypic information becoming available as additional patients are genetically tested. Future updates to the model should undergo rigorous training, testing, and clinical validation before being implemented into the variant classification process.

[0056] The described approach does not use, or rely on, external clinical data (e.g., from publications). Instead, the described clinical variant models can learn from clinical data obtained during laboratory genetic testing and apply to new patients and variants observed in the laboratory. Experimental results have shown that clinical variant models as described are better at making predictions about individuals in a cohort because these models understand data patterns in sampling and enable generalizability of predictions for specific patient populations. External clinical data from publications do not have these same data patterns in patient sampling as a specific patient population. However, evaluation of external clinical data, such as that from publications, can still be used to complement clinical variant modeling as described.

[0057] The described approach can generate meaningful / accurate patient and variant scores from clinical data provided on clinical patient records (e.g., laboratory request forms or TRFs), even when clinical patient records are incomplete or blank. This is because the clinical variant models described have been trained, tested, and validated using a large amount of clinical records.

[0058] To calculate variant scores, embodiments of the clinical variant model analyze the overall distribution of patient scores for a given variant and compare this distribution to the distribution of patient scores observed for known pathogenic variants versus known benign variants. These known variants may have a similar amount of missing or incomplete clinical information for the patients examined. Because the clinical variant model examines patterns in a large number of patients and variants, the impact of incomplete clinical information is much smaller than if the modeling were applied to a single patient or a small number of patients. Additionally, since variant scores can be calculated for all variants, even those classified as pathogenic or benign variants, the described approach can be useful, for example, in identifying potentially misclassified variants for further expert review.

[0059] The described clinical variant model can capture the distribution of patient scores across the range from unaffected to affected. In contrast, existing methods often omit clinical data from individuals who appear unaffected from the variant classification process, perhaps due to concerns that the patient's clinical records are simply incomplete. Unlike conventional approaches, the clinical variant model actually examines how often patients appear unaffected between molecular diagnostic cohorts (due to either incomplete penetrance or incomplete patient clinical records) and genotype-negative cohorts. The described approach learns how much weight a given observation should be assigned. Even if a single observation is not very meaningful, over dozens or hundreds of observations, it can be significant enough to reliably reclassify a variant.

[0060] These clinical variant models can be incorporated into variant classification processes, given their high performance in distinguishing known benign variants from known pathogenic variants for specific genes and / or gene-disease combinations. However, while seemingly well-functioning clinical variant models have been developed for more genes and conditions, these models have not been fully appraised and validated. The set of available clinical variant models can be expanded once the models have been fully evaluated for clinical validity.

[0061] To demonstrate consistency between clinical variant model predictions and known pathogenicity variants, and to help build confidence that the clinical variant model is functioning properly, pathogenicity evidence for known pathogenicity variants is collected and included in the model where available. Furthermore, documenting this evidence may be useful in the future if contradictory evidence emerges later, so that clinical scientists have all possible information to determine whether a reclassification to LP (likely pathogenic), VUS (variant of unknown significance), LB (likely benign), or B (benign) is justified. Applying pathogenicity evidence from clinical variant modeling to pathogenicity variants is expected to further increase the stability of these pathogenicity classifications.

[0062] For genes associated with multiple disease states, if the molecular mechanisms of the disease are the same (e.g., loss-of-function [LOF] or gain-of-function [GOF]), the clinical variant models described may not need to be trained separately for distinct gene-disease associations. The described clinical variant models may be able to learn that distinct combinations of clinical features are observed in patients with different states caused by variants in the same gene. Genes associated with multiple disease states (due to different molecular mechanisms or different inheritance patterns) are a complex issue that is still under investigation, but there are instances where the described modeling approaches appear to work reliably for genes associated with multiple diseases with the same mechanism but different inheritance patterns (e.g., MSH2, MSH6, MLH1, PMS2 associated with dominant LoF Lynch and recessive LoF constitutive mismatch repair deficiency). Additionally, there are instances where the described approaches appear to work reliably using a single model of genes associated with multiple diseases having the same inheritance pattern but different molecular mechanisms (e.g., LoF CASR and GoF CASR).

[0063] To ensure that only the best-performing clinical variant models are used for variant classification, the Area Under the Receiver Operating Characteristic Curve (AUROC) is calculated for each model to measure the model's performance in distinguishing between benign and pathogenic variants. In some embodiments, only models with an AUROC ≥ 0.8 are selected for further evaluation. In some embodiments, continuous improvement of the set of clinical variant models is achieved through a combination of further validation metrics and expert review. In some embodiments, additional steps are taken to establish weighting of variant scores for incorporation as evidence into variant classification frameworks such as the INVITAE Sherloc framework (see, for example, Nykamp K, Anderson M, Powers M, Garcia J, Herrera B, Ho YY, Kobayashi Y, Patil N, Thusberg J, Westbrook M; Invitae Clinical Genomics Group; Topper S. Sherloc: a comprehensive refinement of the ACMG-AMP variant classification criteria. Genet Med. 2017 Oct;19(10):1105-1117.doi:10.1038 / gim.2017.37.Epub 2017 May 11.Erratum in:Genet Med. 2020 Jan;22(1):240-242.PMID:28492532;PMCID:PMC5632818).

[0064] Embodiments of the clinical variant models (CVMs) described have now been developed and validated for, for example, 11 genetic conditions associated with 17 genes, and they demonstrate high performance in distinguishing known benign variants from known pathogenic variants (≧0.8AUROC curve). Predictions from these CVMs have been used as evidence to resolve over 1,000 unique VUSs, affecting >45,000 individuals. >99% of these reclassifications corresponded to a downgrade of VUS to benign or likely benign, while <1% upgraded to pathogenic or likely pathogenic affected approximately 160 individuals. Notably, 91% (10 / 11) of variant upgrades were in genes associated with conditions for which established guidelines for screening and treatment exist, highlighting the potential to alter the medical management of individuals and identify at-risk relatives.

[0065] This disclosure will be better understood from the following detailed description with reference to the accompanying drawings. The detailed description of the drawings is for illustrative and illustrative purposes only and should not be construed as limiting this disclosure to the specific embodiments described.

[0066] In the drawings and the following description, components that have the same name but different reference numbers may be referenced in different drawings. The use of different reference numbers in different drawings indicates that components with the same name may represent the same or different embodiments of the same component. For example, components with the same name but different reference numbers in different drawings may have the same or similar function in some embodiments, such that the description of one of those components in one drawing may apply to other components with the same name in other drawings.

[0067] Furthermore, components illustrated and described in the drawings and the following description in relation to some embodiments may be used in conjunction with other embodiments or incorporated into other embodiments. For example, components shown in a particular drawing may not be limited to use in relation to the embodiment to which that drawing relates, but may be used in conjunction with other embodiments, including embodiments shown in other drawings, or incorporated into other embodiments.

[0068] Figure 1 shows examples of variant scoring processes according to several embodiments of the present disclosure.

[0069] In Figure 1, each patient in the patient population 102 has an associated clinical patient profile 104 and genetic test results 106. The clinical patient profile 104 includes structured data such as ICD-10 codes (International Classification of Diseases codes), sex indications, age, and / or other demographic data, as well as unstructured data, such as natural language text descriptions including clinical descriptions, indications for genetic testing, and family history. The genetic test results 106 can identify one or more variants detected as a result of genetic testing. If multiple variants are identified, the genetic test results can be filtered or sorted by variant to identify positive and negative patients for each variant.

[0070] The clinical patient profile 104 and the genetic test results 106 are input to the patient score generator 108. The patient score generator 108 uses the variant data 110 to generate and output the predicted patient score 112 for the patient associated with the clinical patient profile 104 and the corresponding genetic test results 106. The variant data 110 includes reference labels or ground truth labels for variants that are known / established in the scientific community as pathogenic or benign with respect to a given genetic condition. Each patient score includes the probability that the patient has the genetic condition, based on clinical evidence extracted from the clinical patient profile 104, including unstructured text descriptions.

[0071] One embodiment of the patient score generator 108 includes a machine learning model trained to distinguish between the occurrence of affected (molecular diagnostic positive) and unaffected (genetically negative) clinical features and to estimate the probability that a patient is affected, considering only clinical evidence. The patient score generator 108 may include one or more machine learning models. One embodiment of the patient score generator 108 including two machine learning models is described with reference to Figure 2.

[0072] The patient score 112 is input to the variant score generator 114. The variant score generator 114 uses the variant data 110 and combines the patient scores 112 of the relevant patients to generate and output the variant score 116. Given the patient score, the variant score 116 estimates the probability that any given variant is the cause of one of the gene-associated diseases.

[0073] The patient score 112 and / or variant score 116 may be sent, passed, provided, or otherwise made accessible to one or more downstream systems, processes, components, models, or frameworks, such as a variant classification framework or a clinical data system. As a result, the described computing system 100 may generate and output either or both of the patient-level predictive score 112 and the pathogenicity variant-level estimate (variant score) 116. For example, the patient score 112 may be useful on its own in some applications. Examples of applications where the patient score 112 may be useful independently of the variant score include:

[0074] 1. Medical need review (reimbursement) to automatically identify patients who are likely to meet the examination / reimbursement criteria;

[0075] 2. Determine which disease state the patient is most likely to have, and as a result, determine which genes should be ordered, or which genes should be prioritized in exome / genome testing;

[0076] 3. Support variant classification by prioritizing patients who should receive a more thorough review by expert scientists, and flagging patients who do not require any further review;

[0077] 4. For the discovery of new disease / phenotypic associations, for example, clinical variant models have been able to learn that a particular phenotype correlates with a pathogenic variant in a given gene that was previously unknown or unproven;

[0078] 5. Refine specific diagnostic criteria for association; or

[0079] 6. Assist in the identification of cohorts for clinical testing or other projects.

[0080] Furthermore, while specific implementations of the patient score generator 108 and the variant score generator 114 are described, either the patient score generator 108 or the variant score generator 114 may be implemented using different approaches. For example, in some embodiments, other techniques may be used to calculate either the patient score 112 or the variant score 116.

[0081] In some specific embodiments, the development and application of clinical variant models for a particular gene follows several general steps: (1) First, a patient score is generated to represent the probability that a patient has the molecular condition of interest. This patient score is used in the next step. (2) Second, a variant score is calculated to represent the probability that the variant is pathogenic, based on the distribution of patient scores for that variant. (3) Next, the performance of the CVM (Clinical Variant Modeling) system is determined by using a holdout set of known phenotypic-genotype relationship data points. (4) A well-functioning CVM is calibrated by measuring positive and negative predictive values ​​(PPV and NPV) from the previous step and then integrated into a variant classification framework such as Sherloc with appropriate weights. (5) Optionally, a subset of variant classifications is reviewed by a panel of clinical genomics experts to ensure that the CVM is functioning as expected.

[0082] In some specific embodiments, the following details describe the implementation configuration. In these embodiments, the target dataset is defined. For example, the Monarch Disease Ontology (MonDO) can be used to discover all disease-related sets for the genes being tested. The Gene Curation Union (GenCC) database can be used to assign a set of genes and their associated inheritance patterns for each disease state.

[0083] Patient cohorts are identified. For each disease state, positive and negative (control) cohorts of patients from population 102 are identified as follows: First, affected individuals are identified as patients with a positive molecular diagnosis for at least one of the genes involved. For example, autosomal recessive association requires either compound heterozygosity or homozygosity for a pre-existing or potential pathogenic variant. The control cohort can be identified in the same manner for all conditions defined as patients with only benign variants in the genes involved.

[0084] For data preprocessing and feature selection, features used in patient score modeling can include both structured patient information (e.g., ICD-10 code, age at registration, clinical area, etc.) and unstructured text patient information (e.g., indications for testing, family history, clinician-reported ancestors). Other sources of patient clinical data, including but not limited to non-clinical data, can also be used, or alternatively, including media recordings such as clinic notes in PDF (Portable Document Format) files, EMR (Electronic Medical Record), and / or audio or video recordings of clinician notes or patient records. ICD-10 codes can be preprocessed by first shortening them to the category level and then selecting the most abundant codes in the positive cohort based on a chi-square test. For each patient's clinical record, a string combining both indications for testing and family history can be created by concatenating these substrings with special tokens to distinguish the beginning and end of each span of text.

[0085] The construction of the patient score model can follow the following sequence: A pre-trained large-scale language model (e.g., UFNLP Gatortron) can be fine-tuned across the entire corpus of labeled patients across disease states (TL1). For each disease state, TL1 is further fine-tuned to examples of that disease state (TL2). This model is used to predict the text score for all patients and then used as features along with other patient information (e.g., age at examination, ICD-10 code) to train the final model (PS). This model is then used to predict each patient's overall score, also known as the patient score.

[0086] The variant model utilizes the same variant travel used in genotype filtering for cohorts of genotype-positive and genotype-negative patients, and / or other sets of patients that provide information for interpretation, such as alternative bellwether patients. In addition, the set of patients that are informationally valuable for interpretation, i.e., bellwether patients, is determined in the following manner: For each gene included in the condition, annotated inheritance patterns (see definition of dataset above) are used to identify patients who are expected to develop the disease if the variant is the cause, patients who should not have the associated disease if the variant is benign, and patients who do not have the disease explained by another known variant in the included gene. For example, in a single-gene autosomal dominant model, bellwether patients for a variant are patients who have the variant of interest and do not have other non-benign variants in the gene. In a multiple-gene model, these patients should also not have non-benign variants in other genes.

[0087] A variant score model can be a partially pooled hierarchical Bayesian inference model for interpreting gene variants across multiple genes. It estimates parameters such as pathogenicity rate, penetrance, and gene probability based on input tensors representing genes, patient scores, and variant travel. The variant score, which is the probability that a variant is pathogenic, is sampled from a partially observed posterior predictive distribution of Bernoulli variables.

[0088] For validation, estimates of generalization performance can be obtained by evaluating the model against a holdout set of 20% of labeled variants not used in training. The metrics evaluated included area under the receiver operating characteristic (AUROC) curve, mean precision (AP), mean squared error (MSE), and classification metrics including F1 score, precision, and PPV and NPV. For high-performance models (AUROC ≥ 0.8) and high-performance genes (AUROC ≥ 0.8), variants with a posterior probability of pathogenicity ≤ 0.05 or ≥ 0.99 and two or more apparent diseased observations in unrelated individuals were designated for evidence via a variant classification framework (e.g., Sherloc point assignment).

[0089] Figure 2 shows an example of a process for generating patient scores using a natural language processing-based model, according to several embodiments of the present disclosure.

[0090] Process 200 is executed by processing logic, which includes hardware (e.g., processing devices, circuits, dedicated logic, programmable logic, microcode, device hardware, integrated circuits, etc.), software (e.g., instructions that are run on or executed on the processing device), or a combination thereof. In some embodiments, the method is executed by components of a computing system, which include components or flows shown in Figure 2, which in some embodiments may not be specifically shown in other figures, and / or in some embodiments, components or flows shown in other figures, which may not be specifically shown in Figure 2. Although shown in a specific sequence or order, the order of processes can be changed unless otherwise specified. Thus, the illustrated embodiments should be understood as examples only, the illustrated processes can be executed in a different order, and some processes can be executed in parallel. In addition, in various embodiments, at least one process can be omitted. Thus, not all processes are required in all embodiments. Other process flows are possible.

[0091] In Figure 2, process 200 begins by identifying two cohorts of patients for each gene or set of genes: 1) patients with a confirmed molecular diagnosis and 2) a control cohort. In machine learning, these patient cohorts provide reference labels to guide model training. For each cohort, a first set of clinical features 202 is identified. Features 202 include reference labels and unstructured text, such as natural language descriptions including indication descriptions and / or family history. Features 202 are extracted from clinical patient records associated with patients within the cohort.

[0092] A machine learning model training approach called curriculum learning can be used to train a language model M1 and classify patients based solely on the text descriptions provided by feature 202 (for example, instances of feature 202 include only text descriptions and reference labels). An example of curriculum learning is illustrated in more detail, for example, with reference to Figure 3A below. In general, curriculum learning may include sorting or grouping instances of feature 202 in a specific way to make model training of language model M1 more efficient. For example, using curriculum learning, instances of features with more complete text descriptions may be selected first for input to model M1, so that model M1 is trained with instances with more available unstructured text data before model M1 is trained with instances with less available unstructured text data. The modeling approaches described are not limited to the use of curriculum learning. Other machine learning training approaches can be used, including conventional supervised or semi-supervised training approaches.

[0093] The M1 model outputs a text score, which is the likelihood that a patient associated with an instance of feature 202 has the genetic condition. The text scores generated by model M1 are combined with other clinical features of the patient (e.g., structured data extracted from the patient's clinical records, excluding unstructured text descriptions) to form a second set of clinical features 204. The clinical features 204 of the second set are used to train a second model, M2. Model M2 generates a patient score, which is the final probability. If a holdout set is used, this step can use the same holdout dataset that was used for the M1 model.

[0094] Figure 3A shows an exemplary process 300 for generating a text score using a natural language processing-based model, according to some embodiments of the present disclosure.

[0095] In Figure 3A, the first model input includes a first clinical feature set 302. Examples of the first clinical features include phenotypic data associated with patient identifiers, disease identifiers, and variant identifiers. More specifically, the phenotypic data includes unstructured text such as natural language text, e.g., textual clinical descriptions, indications for clinical tests, descriptions of family history related to the disease, and / or descriptions of demographic data. The first clinical feature set 302 can contain many instances of the first clinical feature (e.g., hundreds, thousands, or more). For example, the first clinical feature set 302 can contain one instance of the first clinical feature for each patient who is a member of a selected cohort.

[0096] The first submodel M1 306 receives the first clinical feature set 302 as input. In response to the first clinical feature set 302, the first submodel M1 306 generates and outputs a text score 314. That is, for each instance of the first clinical feature, the first submodel M1 306 generates and outputs a corresponding text score 314. The text score 314 includes a machine learning prediction about whether the patient associated with the instance of the first clinical feature has the identified disease. The text score 314 is based on unstructured text extracted from the patient's clinical record, excluding other structured data that may also be included in the patient's clinical record, regarding whether the patient associated with the instance of the first clinical feature has the identified disease. The text score 314 is based on unstructured text extracted from the patient's clinical record, excluding other structured data that may also be included in the patient's clinical record.

[0097] The first submodel M1 306 comprises a machine learning algorithm 308, one or more model parameters (including, but not limited to, feature weights or coefficients such as clinical feature weights) 310, and one or more model hyperparameters 312. The first submodel M1 306, comprising the machine learning algorithm 308, one or more model parameters 310, and one or more model hyperparameters 312, can be implemented as a language model. In some implementations, the first submodel M1 306 is constructed using a neural network-based machine learning model architecture. In some implementations, the neural network-based machine learning model architecture includes one or more self-attention layers that allow the model to assign different weights to different words or phrases contained in the model input (e.g., weighting different parts of unstructured text differently).

[0098] Alternatively or additionally, the neural network architecture may include feedforward layers and residual connections, enabling the model to machine learn complex data patterns, including relationships between different words or phrases in the model input across multiple different contexts. In some implementations, the first submodel M1 306 is constructed using a transformer-based architecture that includes self-attention layers, feedforward layers, and residual connections between layers. The exact number and arrangement of each type of layer, as well as the hyperparameter values ​​used to construct the model, are determined based on the requirements of the specific design or implementation of the clinical variant modeling system.

[0099] In some examples, neural network-based machine learning model architectures include, or are based on, one or more generative transformer models, one or more generative pre-trained transformer (GPT®) models, one or more bidirectional encoder representation (BERT) models from transformers, one or more large language models (LLMs), and / or one or more other natural language processing (NLP) models.

[0100] The first submodel M1 306 is an NLP-based machine learning model trained on a dataset of natural language text extracted from patient clinical records. For example, training samples of natural language text extracted from patient forms submitted to a genetic testing service are used to train the first submodel M1 306. The size and composition of the dataset used to train the first submodel M1 306 can vary according to the requirements of the specific design or implementation form of the clinical variant modeling system. In some implementation forms, the dataset used to train the first submodel M1 306 contains hundreds of thousands to millions or more of different natural language text training samples.

[0101] In some implementations, the machine learning algorithm 308 can include a logistic function. The logistic function can model the relationship between X (input) and Y (predicted output), where the probability of Y is a linear combination of the independent variables in the input X. Mathematically, a simplified form of the logistic function can be expressed as P(X)=f(x)=1 / (1+e^(-(β_0+β_1 x))), where e is the exponential constant and β_0 and β_1 are feature coefficients. During the training of the first submodel M1 306, the logistic regression estimates the values ​​of the coefficients in the linear combination based on the feature values ​​in the training dataset and iteratively adjusts them.

[0102] In Figure 3A, the first submodel M1 306 is constructed through the supervised machine learning training, calibration, and validation processes described. During training, one or more values ​​of the model parameters (e.g., weights or feature coefficients such as clinical feature weights) can be iteratively adjusted based on the values ​​of feature inputs x in the first clinical feature set 302 to test the relative effect of a particular feature input x in the first clinical feature set 302 on a predicted outcome P(Y|X), e.g., a predicted text score 314. The values ​​of the model parameters (e.g., weights or feature coefficients such as clinical feature weights) are initialized and adjusted during model training and calibration.

[0103] The first submodel M1 306 also includes model hyperparameters 312 that are selected or tuned at a global level and are generally not modified based on a particular instance of the training data. Examples of model hyperparameters 312 may include the learning rate (the rate at which the algorithm updates its estimates), learning rate decay (the progressive decrease in the learning rate over time to accelerate learning), momentum (the direction of the next step relative to the previous step), regularization constant, the number of branches in the decision tree, and neural network nodes (the number of nodes in each hidden layer of the neural network).

[0104] The first submodel M1 306 can be configured as either a binary classifier or a scoring model. In binary classification mode, the output of the first submodel M1 306 indicates, for a given set of input features, whether the predicted outcome is pathogenic or benign, as a discrete or binary value, for example, 0 indicates benign and 1 indicates pathogenic. In scoring mode, the output of the first submodel M1 306 includes a score (e.g., a floating-point value between 0 and 1) corresponding to the continuous probability that the predicted outcome is pathogenic or benign.

[0105] During inference, the first submodel M1 306 receives a first clinical feature set 302 containing unstructured text and / or variant identifiers that the first submodel M1 306 has never seen as input before (i.e., the first submodel M1 306 has never previously analyzed a particular combination of the first clinical features), and can generate a predictive text score 314 for unknown instances of the first clinical features.

[0106] Figure 3B shows an exemplary process 350 for generating patient scores using a natural language processing-based model, according to some embodiments of the present disclosure.

[0107] In Figure 3B, the second model input includes a second clinical feature set 352. An example of a second clinical feature includes a patient score 364 and corresponding structured phenotypic data associated with a patient identifier, disease identifier, and variant identifier. More specifically, the phenotypic data of the second clinical feature set 352 includes structured data such as ICD codes, demographic data, and other structured data obtained from the patient's clinical record, excluding unstructured phenotypic data (e.g., natural language text descriptions obtained from the patient's clinical record). The second clinical feature set 352 can contain many instances of the second clinical feature (e.g., hundreds, thousands, or more). For example, the second clinical feature set 352 may contain one instance of the second clinical feature for each patient whose text score is generated by the first submodel M1 306.

[0108] The second submodel M2 356 is a machine learning-based scoring model, such as a tree-based model, e.g., a boosted tree model. The second submodel M2 356 receives the second clinical feature set 352 as input. In response to the second clinical feature set 352, the second submodel M2 356 generates and outputs a patient score 364. That is, for each instance of the second clinical feature, the second submodel M2 356 generates and outputs a corresponding patient score 364. The patient score 364 includes a machine learning prediction about whether the patient associated with the instance of the second clinical feature suffers from the identified disease. The patient score 364 is based on structured data extracted from the patient's clinical record, with the exception of unstructured text which may also be included in the patient's clinical record. For example, the patient score 364 is calculated based on the text score 314 and other structured data, with the exception of the text description used by the first submodel M1 306 to generate the text score 314.

[0109] The second submodel M2 356 includes a machine learning algorithm 358, one or more model parameters (including, but not limited to, feature coefficients or weights) 360, and one or more model hyperparameters 362. The second submodel M2 356, including the machine learning algorithm 358, one or more feature coefficients (or weights) 360, and one or more model hyperparameters 362, can be implemented using a neural network-based machine learning model architecture. For example, the second submodel M2 356 can be implemented using another type of machine learning model configured as a boost tree model, a regression model, or a classification or scoring model.

[0110] In some implementations, the second submodel M2 356 is constructed using a transformer-based architecture that includes self-attention layers, feedforward layers, and residual connections between layers. The exact number and arrangement of each type of layer, as well as the hyperparameter values ​​used to construct the model, are determined based on the requirements of the specific design or implementation of the clinical variant modeling system. The second submodel M2 356 is trained on a dataset of the second clinical features 352. For example, structured data extracted from patient forms submitted to a genetic testing service, along with text scores 314, is used to train the second submodel M2 356. The size and composition of the dataset used to train the second submodel M2 356 can vary according to the requirements of the specific design or implementation of the clinical variant modeling system. In some implementations, the dataset used to train the second submodel M2 356 contains hundreds of thousands to millions or more different training samples, and the size of the dataset corresponds to the size of the first clinical feature set 302. In other words, for each instance of the first clinical feature 302, there exists a corresponding instance of the second clinical feature 352.

[0111] In some implementations, the machine learning algorithm 358 may include one or more of the following: the logistic function, decision trees, or boosting functions (e.g., gradient boosting). The logistic function can model the relationship between X (input) and Y (predicted output), where the probability of Y is a linear combination of the independent variables in the input X. Mathematically, a simplified form of the logistic function can be expressed as P(X)=f(x)=1 / (1+e^(-(β_0+β_1 x))), where e is the exponential constant and β_0 and β_1 are feature coefficients. During the training of the second submodel M2 356, the logistic regression estimates the values ​​of the coefficients in the linear combination based on the feature values ​​in the training dataset and iteratively adjusts them.

[0112] In Figure 3B, the second submodel M2 356 is constructed through the described supervised machine learning training, calibration, and validation processes. During training, the feature coefficient values ​​can be iteratively adjusted based on the value of feature input x in feature set 352 to test the relative effect of a particular feature input x in feature set 352 on the predicted outcome P(Y|X), e.g., the predicted patient score 364. The feature coefficient values ​​are initialized and adjusted during model training and calibration.

[0113] The second submodel M2 356 also includes model hyperparameters 362 that are selected or tuned at a global level and are generally not modified based on a particular instance of the training data. Examples of model hyperparameters 362 may include the learning rate (the rate at which the algorithm updates its estimates), learning rate decay (the progressive decrease in the learning rate over time to accelerate learning), momentum (the direction of the next step relative to the previous step), regularization constant, the number of branches in the decision tree, and neural network nodes (the number of nodes in each hidden layer of the neural network).

[0114] The second submodel M2 356 can be configured as a binary classifier or a scoring model. In binary classification mode, the output of the second submodel M2 356 indicates, for a given set of input features, whether the predicted outcome is pathogenic or benign, as a discrete or binary value, for example, 0 indicates benign and 1 indicates pathogenic. In scoring mode, the output of the second submodel M2 356 includes a score (e.g., a floating-point value between 0 and 1) corresponding to the continuous probability that the predicted outcome is pathogenic or benign.

[0115] During inference, the second submodel M2 356 receives structured clinical data, text scores 314, and / or a second set of clinical features 352 that the second submodel M2 356 has never seen before as input (i.e., the second submodel M2 356 has never previously analyzed a particular combination of the second clinical features), and can generate a predicted patient score 364 for an unknown instance of the second clinical features.

[0116] Figure 4 illustrates an exemplary process for constructing a clinical variant model using machine learning, according to several embodiments of this disclosure.

[0117] In Figure 4, the model training process 400 applies a machine learning algorithm to the labeled feature set 402, for example, using supervised machine learning. For example, to train a first submodel M1 306, the labeled feature set 402 includes instances of the first clinical feature 302 and, for each instance of the clinical feature 302, a ground truth or reference label indicating whether the associated variant is a known benign, known pathogenic, or known variant of unknown significance (VUS). As another example, to train a second submodel M2 356, the labeled feature set 402 includes instances of the second clinical feature 352 and, for each instance of the clinical feature 352, a corresponding ground truth or reference label indicating whether the associated variant is a known benign, known pathogenic, or known variant of unknown significance (VUS).

[0118] In some implementations, the model training process 400 implements one or more aspects of curriculum-learning machine learning model training techniques. In curriculum learning, the machine learning model is trained in a meaningful order, for example, from easy training samples to difficult samples. For example, curriculum learning may be implemented by a sorter / scheduler component 404 to group, sort, or rank training instances according to some objective measure of difficulty, and then feed the training instances to the model in a predetermined order (for example, so that the model receives easier training instances before receiving more difficult ones). When training a first submodel M1 306, the sorter / scheduler 404 may group instances of the first clinical feature 302 based on, for example, the amount of unstructured text contained in patient clinical records, so that training instances with more unstructured text are received by the model before training instances with little or no unstructured text. When training the second submodel M1 356, the sorter / scheduler 404 can rank instances of the second clinical feature 352 based on the value of the text score 314, for example, so that training instances with higher text scores are received by the second submodel M1 356 before training instances with lower text scores.

[0119] The sorter / scheduler 404 can also, or alternatively, implement a pacing function that can control how many easier training examples are fed into the model before more difficult training examples are fed into the model, or how quickly the model is introduced to more difficult training examples.

[0120] In addition to organizing and pacing the input of training examples to the model, curriculum learning can also be used to incrementally adjust the model architecture or one or more model parameters, for example, by gradually increasing the model capacity (adding more neural units) or by gradually increasing the complexity of the model's task.

[0121] The use of curriculum learning can be particularly helpful in improving the efficiency of the model training process and enhancing the predictive reliability of a model when the amount of available data varies from training instance to training instance. In particular, curriculum learning can help solve the problem of using patient records with sparse, blank, or incomplete text descriptions for clinical variant modeling.

[0122] In the first training iteration, feature coefficients (or weights) are initialized or assigned to each feature input of a given input feature in subprocess 406. Feature coefficients are initialized, for example, by randomly setting coefficient values. The feature coefficient values ​​assigned in subprocess 406 act as weights applied to each feature to generate weighted features 410. The machine learning algorithm is applied to the weighted features in subprocess 412 to generate predicted outputs 414. Criteria or ground truth labels included in the training instance provide predicted outputs 408 for supervised machine learning. Predicted outputs 414 are evaluated in subprocess 416 by calculating a loss (or error) based on predicted outputs 408 and predicted outputs 414. The loss is calculated using a loss function such as the gradient descent algorithm.

[0123] In each iteration, the decision subprocess 418 evaluates the difference between the predicted output and the expected output by comparing the output of the loss function to a stopping condition that depends, for example, on the change in the output of the loss function. The change in the output of the loss function is compared to an error tolerance threshold (which may be referred to herein as a model performance criterion or model convergence criterion). If the model performance criterion, for example, the error tolerance threshold is not met (e.g., not greater than or equal to the threshold performance level, or exceeding the maximum allowable error value, or the loss has not stopped improving beyond the acceptable range, or the error has not stopped decreasing beyond the acceptable range threshold), model training continues for another iteration. If the loss has stopped improving beyond the acceptable threshold, the model has converged and training is terminated.

[0124] In subsequent training iterations, one or more of the feature coefficients are adjusted, the machine learning algorithm is applied to additional instances of the feature set, and the output of the machine learning algorithm is evaluated using the loss function and error tolerance as described above. The training process 400 ends when the model performance criterion, i.e., the error tolerance threshold, is met and the model converges (e.g., the output of the comparison, e.g., the loss function, is not greater than or equal to the threshold performance level, or is within or below the maximum allowable error value, or the loss has stopped improving beyond the tolerance threshold).

[0125] Figure 5A shows examples of patient score calculation according to several embodiments of the present disclosure. In Figure 5A, the clinical record associated with patient 1 for gene NF1 includes structured data (e.g., age, sex, ICD code) and unstructured data (e.g., indications and family history). Some of this structured and unstructured data is extracted from the clinical record and input into a patient score generator (e.g., a combination of models M1 and M2 described above). The patient score generator generates a patient score based on the extracted structured and unstructured data. In the example in Figure 5A, the patient score indicates that, based on the clinical data, the patient is highly likely to have the genetic condition.

[0126] Figure 5B shows examples of patient score calculation according to several embodiments of the present disclosure.

[0127] Figure 5B shows examples of patient score calculation according to several embodiments of the present disclosure. In Figure 5B, the clinical record associated with patient 2 for gene NF1 includes structured data (e.g., age, sex, ICD code) and unstructured data (e.g., indications and family history). Some of this structured and unstructured data is extracted from the clinical record and input into a patient score generator (e.g., a combination of models M1 and M2 described above). The patient score generator generates a patient score based on the extracted structured and unstructured data. In the example in Figure 5B, the patient score indicates that, based on the clinical data, the patient has a low probability of having the genetic condition.

[0128] Figure 5C shows an example of patient score calculation using machine learning feature weights according to several embodiments of the present disclosure.

[0129] To generate patient scores: (A) Firstly, for a given molecular disorder that can be defined by a single gene (e.g., neurofibromatosis type 1) or multiple genes (e.g., Lynch syndrome), clinically relevant patient information is collected for all patients with a molecular diagnosis of the condition, as well as for all patients who have only benign mutations in the gene of interest and no other molecular diagnoses (i.e., the genotype-negative cohort). The ML model learns appropriate evidence weights for the clinical information to distinguish individuals with a molecular diagnosis from genotype-negative individuals. (B) Next, the learned evidence strengths of the clinical symptoms observed in the cohort are applied to each individual in the entire cohort to generate a patient score for each individual.

[0130] Figure 6A shows an example of variant score calculation using a Bayesian model according to several embodiments of the present disclosure.

[0131] To generate a variant score, the probability that a given variant is pathogenic is derived from a set of patient scores associated with the variant. Molecular diagnostic results are used to identify associated patient cohorts for any given variant. This is done for both known pathogenic and known benign variants. In the example in Figure 6A, each stack of individual icons represents a single variant, and each icon in the stack represents the patient score of patients who have that variant. This distribution can be used to infer the probability that a set of observed variants originated from pathogenic variants (e.g., machine learning-based inference or statistical correlation). The distribution can be generated using a Bayesian inference model that considers multiple clinical genetic phenomena, including incomplete penetrance, age of onset, and phenotypic imitation. The development of this model can naturally represent the uncertainty inherent in having limited observations, which extends to its estimation of apparent variant penetrance, a fundamental signal that the model uses to classify variants. The model can be sampled posteriorly to estimate the probability that a particular variant is pathogenic.

[0132] Figure 6B shows an example of the distribution of patient scores for variants according to several embodiments of the present disclosure.

[0133] The square columns represent the distribution of patient scores for variants of a gene (e.g., MSH2). Each square in a column represents a patient score generated using clinical patient records. In the example in Figure 6B, patient 1 has a patient score of 0.76, and patient 2 has a patient score of 0.04. The higher patient score of 0.76 correlates with a higher likelihood that patient 1 has the genetic condition, while the lower patient score of 0.04 correlates with a lower likelihood that patient 2 has the genetic condition. The example in Figure 6B demonstrates that a patient score can still be generated for patient 2 even when unstructured text information is missing (the indication field is blank).

[0134] Exemplary distribution of patient scores for 11 patients with MSH2 c.942+3A>G for variant score generation. Each box in the figure represents a patient. Patients with low patient scores are shown in white, and patients with high patient scores are shown in black. Exemplary patients with high patient scores based on clinical evidence are shown on the left side of the figure, and exemplary patients with low patient scores based on clinical evidence are shown on the right side of the figure. The example in Figure 6B shows that a patient score can still be generated for a second patient even when some unstructured text information is missing (the family history field is blank).

[0135] Figure 6D shows examples of patient score and variant score distributions for classified variants according to some embodiments of the present disclosure.

[0136] Figure 6D shows an example of the distribution of patient scores for known benign and pathogenic variants in MSH2. The second ML model calculates the variant score for each variant based on the distribution of patient scores. In the example in Figure 6D, all known benign variants that are predominantly found in individuals with low patient scores have low variant scores, while all known pathogenic variants on the right, which are more concentrated in patients with high patient scores, have high variant scores.

[0137] Figure 6E shows examples of patient and variant score distributions for classified variants according to several embodiments of the present disclosure. In Figure 6E, variant scores are generated for VUS in MSH2 and can be compared with variant scores for known benign and known pathogenic variants.

[0138] Figure 6F shows examples of patient score and variant score distributions according to several embodiments of the present disclosure. Figure 6F shows an example of CVM results for NF1. In this cell plot, each stack of boxes represents a single gene variant, and each box represents a single patient. Each box is shaded according to the individual's patient score, the probability that the patient is presumed to have that condition. Patient score boxes 652 along the left y-axis of the plot represent low patient scores, and patient score boxes 654 along the right y-axis of the plot represent high patient scores. In the strip below the patient score plot, each box is shaded according to the variant score resulting from the stack of patient score observations from that variant. Boxes in region 656 represent low variant scores, and boxes in region 658 represent high variant scores. Note that the y-axis has been enlarged to allow visualization of the patient boxes on the right. The actual stacks of patient boxes on the left extend much higher than shown, as these are higher frequency variants.

[0139] Figure 6G shows an example of patient score and classified variant distributions according to several embodiments of the present disclosure. Figure 6G demonstrates that the expected distribution of patient scores for pathogenic and benign variants in the MMR gene can be machine-learned using the described approach. Using the expected distributions, a quantitative determination can be made as to whether the patient score distribution for VUS is similar to that of other pathogenic or benign variants. For example, suppose variants 680 and 682 are VUS and not known benign or known pathogenic. Given the patient score distributions for each of these variants 680 and 682, they can be appropriately placed into the distributions of known benign and known pathogenic variants based on their similarity to the patient score distributions for known benign and known pathogenic variants.

[0140] Figure 7A shows an example of the distribution of patient and variant scores for classified variants according to several embodiments of the present disclosure. In Figure 7A, patient scores for a sample of 20 benign variants and 20 pathogenic variants are plotted in a vertical stack along the y-axis, and variant scores obtained from a Bayesian NF1 clinical variant model are plotted along the x-axis. Figure 7A demonstrates that the clinical variant model accurately classifies these variants. The right side of Figure 7A shows data for a sample of VUS in the same gene, demonstrating that the model performance superficially extends to these variants.

[0141] Figure 7B shows an example of the distribution of patient and variant scores for classified variants according to several embodiments of the present disclosure. In Figure 7B, the sample of 40 VUSs suggests that the patient score pattern generalizes to these unlabeled variants. The variant scores reflect the pathogenicity prediction for the variant on the right, the benignity prediction for the variant on the left, and the uncertainty score for variants that were not sufficiently observed. Figure 7B also shows that variants suspected to be benign are observed far more frequently than variants suspected to be pathogenic, meaning that these benign designations should have a greater impact on the number of VUS reports.

[0142] Figure 8A shows examples of performance data for patient score generators and variant score generators according to several embodiments of the present disclosure.

[0143] Figure 8A plots data generated by running the described CVM pipeline for a set of over 2,300 disease-related genes. Performance at both steps of the pipeline is shown for each genetic disease-specific model. By filtering to only those achieving high performance against the validation set, 920 models are obtained, containing nearly 2,000 genes. These models originate from all clinical areas, from cardiology to oncology. As shown in Figure 8A, the CVM machine learning pipeline can predict variant pathogenicity using previously underutilized clinical evidence. Patient score models estimate the probability that a patient has the relevant genetic condition. Variant score models combine relevant patient observations to predict whether a variant is the cause of the disease. In some embodiments, there is a minimum number of patient observations required for a variant to qualify for a clinical variant model. For example, in some implementation forms, a clinical variant model, as described, may require a minimum of three observations from at least two separate families. However, in practice, for many genes, more patient observations may be required to ensure that the model reaches the confidence threshold necessary to provide evidence to variant classification frameworks such as Sherloc.

[0144] Experimental results demonstrated that the described approach performed well (AUROC ≥ 0.8) in distinguishing between pathogenic and benign variants for numerous clinical conditions (n=920) and genes (n=1,977). This model provides evidence for both pathogenicity and benignity for a large number of VUSs and has the potential to substantially reduce the number of VUS reports to which it applies.

[0145] Figure 8B shows examples of clinical variant modeling and classification systems according to several embodiments of the present disclosure.

[0146] The development and application of clinical variant models for each gene follows several general steps: (1) First, a patient score is generated to represent the probability that a patient has the molecular condition of interest. This patient score is used in the next step. (2) Second, a variant score is calculated to represent the probability that the variant is pathogenic, based on the distribution of patient scores for that variant. (3) Next, the performance of the CVM is determined by using a holdout set of known phenotypic-genotype relationship data points. (4) Then, a well-functioning model is calibrated by measuring positive and negative predictive values ​​(PPV and NPV) from the previous step, and then integrated into a variant classification framework such as Sherloc with appropriate weights. (5) Optionally, a subset of variant classifications is reviewed by a panel of clinical genomics experts to ensure that the CVM is functioning as expected. Each of these general steps is described in more detail below.

[0147] Step 1: Generate patient scores

[0148] CVM leverages clinical data to predict the pathogenicity of variants for a given genetic condition in a stepwise manner. By utilizing details found in clinician-reported data from test request forms (e.g., personal health history, family health history, age, sex, patient's race, ethnicity, and ancestry, ICD-10 code, and clinical domain of the requested test), the model is initially trained to learn a clinical picture that distinguishes patients with a positive molecular diagnosis for the target condition from genotype-negative controls (i.e., patients who do not have a VUS, LP, or P variant in the target condition and do not have a molecular diagnosis in another condition). For each patient, a patient score is generated (e.g., on a scale from 0 to 1), which is the probability that a given patient has the target genetic condition based solely on clinical information. Based on what has been learned, this information can be applied to other patients with VUS in the target condition, provided that those patients do not have a current molecular diagnosis in another gene. This is done by scoring the clinical profiles of other patients to see how similar they appear to be to positive or negative cases.

[0149] Step 2: Generating Variant Scores

[0150] Using a set of known pathogenic and benign variants for a condition, called labels, a second model learns the distribution of patient scores typical for pathogenic and benign variants. Based on what has been learned, VUS can score patient scores based on how similar their distributions appear to pathogenic and benign variants. A variant score is generated for each variant (e.g., on a scale from 0 to 1), which is the probability that a given variant is pathogenic. A higher variant score indicates a higher probability that the variant is pathogenic.

[0151] Step 3: Performance evaluation of the CVM

[0152] All CVMs are carefully screened to ensure high performance in distinguishing between pathogenic and benign variants (e.g., AUROC ≥ 0.8) by testing each model against a set of known pathogenic and benign variants that the model has not previously encountered (i.e., a 20% holdout set).

[0153] Step 4: Model Calibration

[0154] Next, models that pass the high-performance threshold are calibrated. In some embodiments, the calibrated models are then evaluated by clinical genomics experts before implementation. To date, clinical variant modeling has demonstrated high accuracy for more than 600 genes and disease states.

[0155] Step 5: Review by a clinical genomics specialist

[0156] Optionally, to gain greater reliability in the predictive output of CVM, a subset of variants may be selected for thorough review by expert clinical genomics scientists. This included sampling of variants with strong pathogenicity (≧99% PPV) and variants with strong benign (≧95% NPV) CVM predictions (e.g., 163 / 1,052). For each variant, all currently available non-CVM evidence was re-reviewed and evaluated for inconsistent data. In this review, 93% (13 / 14) of variants predicted to be pathogenic by CVM and 95% (155 / 163) of those predicted to be benign were confirmed by experts, while the remainder were preserved as VUS. Notably, experts could select the most challenging variants with benign CVM predictions to review (e.g., variants with a certain level of pathogenicity evidence or variants in the PMS2 pseudogene region). Furthermore, in several embodiments, four of the predicted pathogenic variants had at least one CLINVAR entry indicating they were likely or pathogenic, while a fifth variant predicted to be pathogenic by CVM was recently reclassified from VUS to likely pathogenic based on new family isolation data obtained after the CVM prediction was generated but before the results were reviewed. Similarly, fifteen of the predicted benign variants examined had at least one CLINVAR entry indicating they were likely or benign by another submitter.

[0157] Orthogonal verification.

[0158] To further enhance the reliability of the CVM's predictive output, we performed a coincidence analysis comparing the CVM with the Evolutionary Model of Variant Effects (EVE), a deep learning model that predicts variant pathogenicity in human 2 using orthogonal data, i.e., sequence preservation. The CVM predictions showed high agreement with the EVE predictions (90.5%).

[0159] For further reliability of the CVM predictive output for Lynch syndrome genes (MLH1, MSH2, MSH6, PMS2, EPCAM), CVM model predictions were compared with functional datasets for MLH1, MSH2, and PMS2 (e.g., multiple assays of variant effects (i.e., MAVE)). Variants with both CVM model and MAVE predictions showed high agreement (>98%).

[0160] Another mechanism for evaluating orthogonal validation of CVMs can be performed for TSC2, a gene associated with tuberous sclerosis. Specifically, exons 26 and 32 of TSC2 are absent in all known clinically relevant transcripts of gene 4 (note legacy exon nomenclature). Using only clinical evidence, CVM of TSC2 identified gene variants predicted to be pathogenic in each of the 42 exons of the gene, excluding exons 26 and 32—a remarkable agreement between the clinical model and independent molecular data.

[0161] Figure 8C shows examples of variant classification systems, including clinical variant modeling, according to several embodiments of the present disclosure.

[0162] To integrate predictions from CVMs into variant classification frameworks such as Sherloc, a two-tiered prediction can be established based on predictive performance thresholds measured by negative and positive predictive values ​​(NPV and PPV). The benign tier is defined as strong benign evidence, which is sufficient evidence to classify a variant as potentially benign without opposing evidence. The pathogenicity tier, while insufficient on its own to reach a possible pathogenicity classification, could be reached by adding variants that are either outside or within the expected pathogenicity range of gnomAD or other pathogenicity evidence. The predictive performance thresholds for these tiers were defined as (1) strong benign ≥ 95% NPV and (2) strong pathogenicity ≥ 99% PPV, respectively. The third and final tiers correspond to predictions falling between 99% PPV and less than 95% NPV, which were considered not reliable enough to assign weights within the Sherloc scoring system for the first release of CVMs.

[0163] Clinical variant modeling is incorporated into variant classification frameworks such as Sherloc. In one example, CVM predictions were integrated into Sherloc using a two-tiered prediction classification. Variants with CVM predictions having a NPV of 95% or higher were assigned 3 benign points, and variants with CVM predictions having a PPV of 99% or higher were assigned 3 pathogenic points. This clinical evidence is considered in the context of all variant classification evidence in Sherloc. Typical cutoffs for variant classification are shown in the lower right: 5 benign points for benign, 3 benign points for likely benign, 4 pathogenic points for likely pathogenic, and 5 pathogenic points for pathogenicity classification. Expert scientists in clinical genomics may override these classification thresholds if necessary due to conflicting evidence and other factors.

[0164] Generally, the number of points associated with a particular score correlates with the reliability or certainty of the prediction. The number of divisions does not need to be two; any number of divisions (including no divisions or an infinite number of divisions) can be used. The points mapped to the scores output by the model can be incorporated into another variant classification framework, such as Sherloc. Thus, the output of a clinical variant model, as described, can be applied to the Sherloc framework or any other variant traveling or classification framework.

[0165] Figure 9A illustrates a method for clinical variant modeling according to several embodiments of the present disclosure.

[0166] Method 900 is executed by processing logic, which includes hardware (e.g., processing devices, circuits, dedicated logic, programmable logic, microcode, device hardware, integrated circuits, etc.), software (e.g., instructions that are operated on or executed on the processing device), or a combination thereof. In some embodiments, the method is executed by components of a computing system, which include components or flows shown in Figure 9A, which in some embodiments may not be specifically shown in other figures, and / or in some embodiments, components or flows shown in other figures, which may not be specifically shown in Figure 9A. Although shown in a specific sequence or order, the order of processes can be changed unless otherwise specified. Thus, the illustrated embodiments should be understood as examples only, the illustrated processes can be executed in a different order, and some processes can be executed in parallel. In addition, in various embodiments, at least one process can be omitted. Thus, not all processes are required in all embodiments. Other process flows are possible.

[0167] In step 902, the processing device extracts phenotypic data from clinical records associated with the patient population and genetic status, the phenotypic data including natural language text containing at least one of the following: indications for testing, a description of family history, or demographic data.

[0168] In step 904, the processing device extracts genotype data from genetic test results associated with a genetic state, the genotype data comprising at least one variant of a gene and at least one molecular diagnostic associated with at least one variant.

[0169] In step 906, the processing device generates and outputs patient scores for a patient population using phenotypic data including natural language text and genotypic data to construct a first model including a natural language processing (NLP) based machine learning model, the patient scores include a machine learning likelihood that the patient has a genetic condition, taking into account the phenotypic data including natural language text.

[0170] In some implementations, configuring a natural language processing (NLP) based machine learning model involves configuring a first model to receive natural language text and generate and output a text score, the text score including a machine learning likelihood that a patient has a genetic condition, given the received natural language text; and configuring a second model to generate and output a patient score using a second model input including phenotypic data excluding the text score and natural language text. In some implementations, the processing device uses curriculum learning to configure the NLP-based machine learning model to generate and output the text score, and then either curriculum learning or another training approach is used to train the second model to generate and output the patient score. Curriculum learning includes using the text score to determine the order of the second model inputs and applying the second model to the second model inputs in that order. In some implementations, the processing device iteratively adjusts the values ​​of at least one clinical feature weight of the NLP-based machine learning model until the output of a loss function calculated based on patient scores output by the NLP-based machine learning model satisfies at least one first performance criterion, thereby generating a configured NLP-based machine learning model that can output patient scores that satisfy at least one second performance criterion. In some implementations, the processing device constructs the NLP-based machine learning model using positive and negative examples, where the positive examples include positive associations between molecular diagnoses and genetic states, and the negative examples include negative associations between molecular diagnoses and genetic states.

[0171] In step 908, the processing device configures a Bayesian causal model to generate and output a variant score using patient score and genotype data, the variant score including a machine learning probability that the variant is pathogenic with respect to the genetic state. In some implementations, the Bayesian causal model includes probability distributions of the patient score and the variant score.

[0172] In some implementations, the processing device uses a configured NLP-based machine learning model and a configured Bayesian causal model to generate and output predictions about whether a variant associated with an unknown clinical record is benign or pathogenic in terms of its genetic state. In some implementations, the processing device provides clinicians with predictions about whether a variant is benign or pathogenic for use in formulating patient diagnoses. In some implementations, the processing device provides predictions to at least one system, process, model, or component of a variant classification system or a clinical data system. In some implementations, the processing device uses variant scores output by a configured Bayesian causal model as input to a variant classification framework.

[0173] Figure 9B illustrates a method for clinical variant modeling according to several embodiments of the present disclosure.

[0174] Method 920 is executed by processing logic, which includes hardware (e.g., processing devices, circuits, dedicated logic, programmable logic, microcode, device hardware, integrated circuits, etc.), software (e.g., instructions that are operated on or executed on the processing device), or a combination thereof. In some embodiments, the method is executed by components of a computing system, which include components or flows shown in Figure 9B, which in some embodiments may not be specifically shown in other figures, and / or in some embodiments, components or flows shown in other figures, which may not be specifically shown in Figure 9B. Although shown in a specific sequence or order, the order of processes can be changed unless otherwise specified. Thus, the illustrated embodiments should be understood as examples only, the illustrated processes can be executed in a different order, and some processes can be executed in parallel. In addition, in various embodiments, at least one process can be omitted. Thus, not all processes are required in all embodiments. Other process flows are possible.

[0175] In step 922, the processing device extracts phenotypic data from clinical records associated with a patient population and genetic status, the phenotypic data including natural language text containing at least one of the following: indications for testing, a description of family history, or demographic data.

[0176] In step 924, the processing device generates and outputs a patient score for a patient population using phenotypic data including natural language text, the patient score includes the likelihood that the patient has a genetic condition, given the phenotypic data including natural language text.

[0177] In some implementations, the processing device configures a natural language processing (NLP)-based machine learning model to generate and output patient scores by (i) configuring a first model to generate and output text scores, the text scores including a machine learning likelihood that a patient has a genetic condition given a first model input including natural language text, and (ii) configuring a second model to generate and output patient scores using a second model input including phenotypic data excluding text scores and natural language text. In some implementations, the processing device uses curriculum learning to configure the NLP-based machine learning model to generate and output patient scores, the curriculum learning including using text scores to determine the order of the second model inputs and applying the second model to the second model inputs in that order. In some implementations, the processing device configures a natural language processing (NLP)-based machine learning model to generate and output patient scores, iteratively adjusts the value of at least one clinical feature weight of the NLP-based machine learning model until the output of a loss function calculated based on the patient scores output by the NLP-based machine learning model satisfies at least one first performance criterion, thereby generating a configured NLP-based machine learning model that can output patient scores that satisfy at least one second performance criterion. In some implementations, the processing device configures the NLP-based machine learning model using positive and negative examples, where the positive examples include positive associations between molecular diagnoses and genetic states, and the negative examples include negative associations between molecular diagnoses and genetic states.

[0178] In step 926, the processing device extracts genotype data from genetic test results associated with the patient population and genetic status, the genotype data including molecular diagnostics that include at least one variant of a gene. In step 928, the processing device configures a Bayesian causal model to generate and output a variant score using the patient score and genotype data, the variant score including a machine learning likelihood that the variant is pathogenic with respect to the genetic status. In some implementation forms, the Bayesian causal model includes probability distributions of the patient score and the variant score.

[0179] In some implementations, the processing device uses a configured NLP-based machine learning model and a configured Bayesian causal model to generate and output predictions about whether a variant associated with an unknown clinical record is benign or pathogenic in terms of its genetic state. In some implementations, the processing device provides clinicians with predictions about whether a variant is benign or pathogenic for use in formulating patient diagnoses. In some implementations, the processing device provides predictions to at least one system, process, model, or component of a variant classification system or a clinical data system. In some implementations, the processing device uses variant scores output by a configured Bayesian causal model as input to a variant classification framework.

[0180] Figure 9C illustrates a method for clinical variant modeling according to several embodiments of the present disclosure.

[0181] Method 930 is executed by processing logic, which includes hardware (e.g., processing devices, circuits, dedicated logic, programmable logic, microcode, device hardware, integrated circuits, etc.), software (e.g., instructions that are operated on or executed on the processing device), or a combination thereof. In some embodiments, the method is executed by components of a computing system, which include components or flows shown in Figure 9C, which may not be specifically shown in other figures, and / or in some embodiments, components or flows shown in other figures, which may not be specifically shown in Figure 9C. Although shown in a specific sequence or order, the order of processes can be changed unless otherwise specified. Thus, the illustrated embodiments should be understood as examples only, the illustrated processes can be executed in a different order, and some processes can be executed in parallel. In addition, in various embodiments, at least one process can be omitted. Thus, not all processes are required in all embodiments. Other process flows are possible.

[0182] In step 932, the processing device extracts phenotypic data from clinical records associated with a patient population and genetic status, the phenotypic data including natural language text containing at least one of the following: indications for testing, a description of family history, or demographic data.

[0183] In step 934, the processing device generates and outputs patient scores for a patient population using phenotypic data, including natural language text, to construct a natural language processing (NLP) based machine learning model, the patient scores include a machine learning likelihood that a patient has a genetic condition given phenotypic data, including natural language text.

[0184] In some implementations, configuring a natural language processing (NLP) based machine learning model involves configuring a first model to receive natural language text and generate and output a text score, the text score including a machine learning likelihood that a patient has a genetic condition, given the received natural language text; and configuring a second model to generate and output a patient score using a second model input including phenotypic data excluding the text score and natural language text. In some implementations, a processing device uses curriculum learning to configure an NLP-based machine learning model to generate and output a patient score, the curriculum learning including using the text score to determine the order of the second model inputs, and applying the second model to the second model inputs in that order. In some implementations, the processing device iteratively adjusts the values ​​of at least one clinical feature weight of the NLP-based machine learning model until the output of a loss function calculated based on patient scores output by the NLP-based machine learning model satisfies at least one first performance criterion, thereby generating a configured NLP-based machine learning model that can output patient scores that satisfy at least one second performance criterion. In some implementations, the processing device constructs the NLP-based machine learning model using positive and negative examples, where the positive examples include positive associations between molecular diagnoses and genetic states, and the negative examples include negative associations between molecular diagnoses and genetic states.

[0185] In step 936, the processing device extracts genotype data from genetic test results associated with a genetic state, the genotype data comprising at least one variant of a gene and at least one molecular diagnostic associated with at least one variant.

[0186] In step 938, the processing device generates and outputs a variant score using patient score and genotype data, the variant score including a machine learning likelihood that the variant is pathogenic with respect to the genetic state. In some implementations, the processing device configures a Bayesian causal model to generate and output the variant score. In some implementations, the Bayesian causal model includes probability distributions of the patient score and the variant score.

[0187] In some implementations, the processing device uses a configured NLP-based machine learning model and a configured Bayesian causal model to generate and output predictions about whether a variant associated with an unknown clinical record is benign or pathogenic in terms of its genetic state. In some implementations, the processing device provides clinicians with predictions about whether a variant is benign or pathogenic for use in formulating patient diagnoses. In some implementations, the processing device provides predictions to at least one system, process, model, or component of a variant classification system or a clinical data system. In some implementations, the processing device uses variant scores output by a configured Bayesian causal model as input to a variant classification framework.

[0188] Figure 10 shows an exemplary computing system, including clinical variant modeling, according to several embodiments of the present disclosure.

[0189] In the embodiment shown in Figure 10, the computing system 1000 includes a data selection subsystem 1002, a feature generation subsystem 1012, a training and calibration subsystem 1020, a model validation subsystem 1030, and one or more scoring models 1040. Data sources that can be accessed and used by the components of the computing system 1000 include a data store for model performance criteria 1032, one or more training datasets 1034, model validation criteria 1036, and one or more validation datasets 1038.

[0190] One or more components of the computing system 1000 may correspond to components and / or processes similarly described in other figures and / or herein. For example, one or more scoring models 1040 may correspond to a single natural language processing (NLP) based model (e.g., M1 or M2 as described above), or a combination of NLP-based models (e.g., a modeling pipeline including M1 and M2), or a Bayesian causal model, or a combination of one or more NLP-based models and a Bayesian causal model. In other words, the scoring model 1040 may include a patient score generator, a variant score generator, or both a patient score generator and a variant score generator.

[0191] For example, the scoring model 1040 may include one or more machine learning models configured to determine probabilistic or statistical relationships between inputs and outputs using machine learning algorithms. For example, given one or more inputs, the scoring model 1040 may output labels that can be used to classify the inputs into different categories, or scores that can be used to sort or rank the inputs into groups or ranked lists.

[0192] The data selection components 1004, 1006, and 1008 may correspond to one or more of the data selection components, techniques, or processes described herein. The feature generation subsystem 1012 may correspond to one or more of the feature generation components, techniques, or processes described herein. The model performance criteria 1032, training dataset 1034, model validation criteria 1036, and validation dataset 1038 may, as applicable, include data described herein as performance criteria, model validation data, training data, or validation data. The model validation subsystem 1030 may correspond to one or more of the model validation components, techniques, or processes described herein.

[0193] The training and calibration subsystem 1020 includes a dataset selection component 1022, a model training component 1024, and a model calibration component 1026. The model training component 1024 and / or model calibration component 1026 of the training and calibration subsystem 1020 can train one or more machine learning models of the scoring model 1040 by applying supervised machine learning techniques to training data, for example, including training examples of input data and reference (or ground truth) labels. The predicted outputs of one or more machine learning models are observed iteratively until a set of model performance criteria is met. For example, the difference between the predicted output and the expected output is quantified using a loss function. The model performance criteria are used to determine when one or more machine learning models have converged to provide reliable outputs with the desired degree of certainty. The required level of certainty and performance criteria is determined based on the requirements or design of the particular implementation form of one or more machine learning models.

[0194] The training dataset 1034 contains training data used to train one or more scoring models 1040 in several implementations. The training dataset 1034 includes, for example, a set of input features and corresponding reference (or ground truth) labels. In some implementations, the training dataset 1034 may include, or be derived from, one or more databases of historical population data and / or genetic testing data. Examples of training datasets 1034 include, as described above, clinical datasets, variant datasets, and / or ranked or ordered subsets used, for example, for curriculum learning.

[0195] The dataset selection component 1022 selects a dataset appropriate for training or calibrating a particular scoring model 1040. For example, if a curriculum learning approach is used to train the M2 scoring model, the dataset selection component 1022 may select training examples having the top k text scores in rank order, where k is a positive integer. Alternatively or additionally, the dataset selection component 1022 may sort or order training examples used to train the M1 scoring model based on the amount of unstructured text contained in one or more unstructured text fields of the patient clinical record. For example, the dataset selection component 1022 may group training examples according to the presence or absence of text in one or more unstructured text fields, and then order the groups such that, for example, groups with more text in the unstructured text fields are entered into the M1 model before groups with less text or no text in the unstructured text fields.

[0196] When a curriculum-based learning approach is used to train a scoring model, the model training component 1024 can determine a schedule for applying the scoring model to ranked or ordered datasets prepared by the dataset selection component 1022, and the model calibration component 1026 can provide feedback to the model training component 1024, which can then use to modify the scheduling or pacing of the training process.

[0197] Although not specifically shown in Figure 10, the computing system 1000 may include one or more user systems, which may be the same devices as the computing system 1000 or different devices. A user system includes at least one computing device, such as a personal computing device, a server, a mobile computing device, or a smart appliance. A user system includes at least one software application, which includes a user interface, installed on the computing device or accessible via a network. For example, the user interface may include a graphical display screen that displays controls and graphical elements for operating and / or manipulating the output of one or more scoring models 1040 and / or controlling one or more data selection, feature generation, training, calibration, or validation processes.

[0198] A user interface can be used to input data, initiate user interface events, and display or otherwise perceive output, including patient predictions, variant pathogenicity predictions, and / or other data generated by the clinical variant modeling system. Examples of user interfaces include web browsers, command-line interfaces, and mobile app frontends. A user interface can include an application programming interface (API). A user interface can include a frontend portion of an application system used by a clinician or scientist or laboratory technician. For example, the output of a clinical variant modeling system may be sent to and displayed on a user interface of a computing device used by a clinician. Alternatively or additionally, another version of the user interface may include a frontend portion of an application software system used by variant scientists and / or other individuals working in the field of genetic testing. Thus, the output of the clinical variant modeling system 1050 may be sent to and displayed on a user interface of a computing device used by any of these and / or other individuals.

[0199] An application system may include any type of application software system that provides or enables the generation, display, or manipulation of the output generated by a clinical variant modeling system. Examples of application systems include, but are not limited to, variant classification systems, DNA (deoxyribonucleic acid) analysis software, genetic testing software, medical testing software, healthcare management software, or any combination of any of the above.

[0200] Although not specifically shown, the data storage system may include data stores and / or data services that store data received, used, manipulated, and generated by the application system and / or clinical variant modeling system, such as training data, validation data, machine learning model parameters and coefficients, performance criteria, validation criteria, and machine learning model outputs. In some embodiments, the data storage system includes multiple different types of data storage and / or distributed data services. As used herein, a data service may refer to a physical, geographical grouping of machines, a logical grouping of machines, or a single machine. For example, a data service may be a data center, a cluster, a group of clusters, or a machine.

[0201] The data storage system resides on at least one persistent and / or volatile storage device that can reside on the same local network as at least one other device of computing system 1000 and / or on a network remote to at least one other device of computing system 1000. Thus, although shown as being included in computing system 1000, a portion of data storage system 1080 may be part of computing system 1000 or may be accessed by computing system 1000 via a network.

[0202] Although not specifically shown, it should be understood that any component of computing system 1000, when executed, may include one or more interfaces embodied as computer programming code stored in computer memory, enabling the computing device to communicate bidirectionally with any other component of computing system 1000 using communication coupling mechanisms. Examples of communication coupling mechanisms include network interfaces, inter-process communication (IPC) interfaces, and application programming interfaces (APIs).

[0203] Each component of the computing system 1000 is implemented using at least one computing device that can be communicatively coupled to one or more electronic communication networks. Any component of the computing system 1000 can be bidirectionally coupled to any other component of the computing system 1000 by the network.

[0204] A typical user of computing system 1000 may be an administrator or end user of the application system and / or clinical variant modeling system.

[0205] The characteristics and functionalities of the components of computing system 1000 may include automated functionalities, data structures, and combinations of digital data implemented using computer software, hardware, or software and hardware, and schematically represented in the diagram. Components may be shown in the diagram as separate elements for ease of explanation, but unless otherwise stated, the diagram does not imply that separation of these elements is necessary. The illustrated systems, services, and data stores (or their functionalities) can be divided across any number of physical systems, including a single physical computer system, and can communicate with each other in any appropriate manner.

[0206] Although not specifically stated, the network may be implemented on any medium or mechanism that provides for the exchange of data, signals, and / or instructions between various components of the computing system 1000. Examples of networks include, but are not limited to, a local area network (LAN), a wide area network (WAN), an Ethernet® network or the Internet, or at least one terrestrial link, satellite link, optical link or wireless link, or any number of different networks and / or communication links in combination.

[0207] Figure 11 is a block diagram of an exemplary computer system in which embodiments of the present disclosure can operate. Figure 11 shows an exemplary machine of computer system 1100 that can execute a set of instructions to cause a machine to perform any of the methods described herein. In some embodiments, computer system 1100 may correspond to a component of a networked computer system (e.g., computing system 1000 in Figure 10) that includes, is coupled to, or utilizes a machine for running an operating system to perform the operations described above, corresponding to embodiments of computing system 1000 in Figure 10.

[0208] The machine is connected to other machines in a local area network (LAN), intranet, extranet, and / or the internet (e.g., it is networked). The machine can operate as a server or client machine in a client-server network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or client machine in a cloud computing infrastructure or environment.

[0209] A machine is a personal computer (PC), smartphone, tablet PC, set-top box (STB), personal digital assistant (PDA®), mobile phone, web appliance, server, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be performed by such machine. Furthermore, although a single machine is shown, the term “machine” should also be interpreted to include any set of machines that individually or collectively execute a set (or set) of instructions to perform any of the methods described herein.

[0210] An exemplary computer system 1100 includes processing devices 1102, main memory 1104 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM) or rhombus DRAM (RDRAM), etc.), memory 1105 (e.g., flash memory, static random access memory (SRAM), etc.), input / output system 1110, and data storage system 1140, all communicating with each other via a bus 1130.

[0211] The processing device 1102 represents at least one general-purpose processing device, such as a microprocessor or a central processing unit. More specifically, the processing device may be a composite instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor implementing another instruction set, or a processor implementing a combination of instruction sets. The processing device 1102 may also be at least one dedicated processing device, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), or a network processor. The processing device 1102 is configured to execute instructions 1112 for performing the operations and steps discussed herein.

[0212] Instruction 1112 includes a portion of the clinical variant modeling system 1150 when a portion of the clinical variant modeling system is being executed by the processing device 1102. Therefore, the clinical variant modeling system 1150 is sometimes shown with a dashed line as part of instruction 1112 to indicate that a portion of the clinical variant modeling system is being executed by the processing device 1102. For example, if at least some portions of the clinical variant modeling system 1150 are embodied in instructions that cause the processing device 1102 to execute in the manner described above, some of those instructions can be loaded from the main memory 1104 and / or the data storage system 1140 into the processing device 1102 (e.g., into an internal cache or other memory). However, not all of the clinical variant modeling system 1150 needs to be included in instruction 1112 at the same time; portions of the clinical variant modeling system 1150 may be stored in at least one other component of the computer system 1100 at other times, for example, when at least some portions of the clinical variant modeling system are not being executed by the processing device 1102.

[0213] The computer system 1100 further includes a network interface device 1108 for communication over the network 1120. The network interface device 1108 provides bidirectional data communication to connect to the network. For example, the network interface device 1108 may be an Integrated Services Digital Network (ISDN®) card, a cable modem, a satellite modem, or a modem for providing data communication connectivity to a corresponding type of telephone line. As another example, the network interface device 1108 may be a local area network (LAN) card for providing data communication connectivity to a compatible LAN. A wireless link may also be implemented. In any such implementation, the network interface device 1108 can transmit and receive electrical, electromagnetic, or optical signals carrying digital data streams representing various types of information.

[0214] A network link can provide data communication to other data devices via at least one network. For example, a network link can provide a connection to a global packet data communication network, commonly referred to as the "Internet," to a host computer via a local network, or to data devices operated by an Internet Service Provider (ISP). The local network and the Internet use electrical, electromagnetic, or optical signals to carry digital data to and from the computer system 1100.

[0215] The computer system 1100 can send messages and receive data, including program code, through the network and network interface device 1108. In the example of the internet, a server can send requested code for an application program through the internet and network interface device 1108. The received code can be executed by the processing device 1102 when it is received and / or stored in the data storage system 1140 or other non-volatile storage for later execution.

[0216] The input / output system 1110 includes a display for displaying information to the computer user, such as a liquid crystal display (LCD) or a touchscreen display, or an output device such as a speaker, a haptic device, or another form of output device. The input / output system 1110 may also include an input device configured to communicate information and command selections to the processing device 1102, such as alphanumeric keys and other keys. The input device may also include, or alternatively, a cursor control such as a mouse, trackball, or cursor directional keys for communicating directional information and command selections to the processing device 1102 and for controlling cursor movement on the display. The input device may also include, or alternatively, a microphone, sensor, or array of sensors for communicating sensed information to the processing device 1102. The sensed information may include, for example, voice commands, audio signals, geographic location information, and / or digital images.

[0217] The data storage system 1140 includes a machine-readable storage medium 1142 (also known as a computer-readable medium) storing at least one instruction set 1144 or software that embodies any of the methods or functions described herein. The instructions 1144 may also be entirely or at least partially present in the main memory 1104 and / or the processing device 1102 during their execution by the computer system 1100, and the main memory 1104 and the processing device 1102 also constitute the machine-readable storage medium.

[0218] In one embodiment, instructions 1112, 1114, and 1144 include instructions for implementing functions corresponding to a clinical variant modeling system (e.g., any one or more components of the clinical variant modeling approaches and techniques described herein).

[0219] In Figure 11, dashed lines are used to indicate that the clinical variant modeling system 1150 does not need to be fully implemented simultaneously in instructions 1112, 1114, and 1144. In one example, part of the clinical variant modeling system is implemented in instruction 1144, which is loaded into main memory 1104 as instruction 1114, and part of instruction 1114 is loaded into processing device 1102 as instruction 1112 for execution. In another example, some parts of the clinical variant modeling system are implemented in instruction 1144, other parts in instruction 1114, and yet another part in instruction 1112.

[0220] Although the machine-readable storage medium 1142 is shown as a single medium in exemplary embodiments, the term “machine-readable storage medium” should be interpreted to include a single or more mediums that store at least one set of instructions. The term “machine-readable storage medium” should also be interpreted to include any medium that can store or encode a set of instructions for machine execution, causing a machine to perform any of the methods of the present disclosure. Accordingly, the term “machine-readable storage medium” should be interpreted to include, but not be limited to, solid-state memory, optical media, and magnetic media.

[0221] Some parts of the detailed description above are presented with respect to algorithms and symbolic representations of operations on data bits in computer memory. These algorithmic descriptions and representations are the methods used by those skilled in the data processing art to most effectively communicate the content of their work to others skilled in the art. An algorithm is considered herein, and also generally, to be a self-consistent set of operations that produce a desired result. An operation is one that requires the physical manipulation of physical quantities. Usually, but not always, these quantities take the form of electrical, optical, or magnetic signals that can be stored, combined, compared, and otherwise manipulated. Referring to these signals as bits, values, elements, symbols, characters, terms, numbers, etc., has sometimes proven convenient, primarily for reasons of general use.

[0222] However, it should be noted that all these and similar terms should be associated with appropriate physical quantities and are merely convenient labels applied to those quantities. This disclosure may refer to actions and processes of a computer system or similar electronic computing device that manipulate data represented as physical (electronic) quantities in the registers and memory of a computer system and convert them into other data similarly represented as physical quantities in the computer system memory or registers or other such information storage systems.

[0223] This disclosure also relates to an apparatus for performing the operations described herein. This apparatus may include a general-purpose computer that can be specifically constructed for an intended purpose or that can be selectively invoked or reconfigured by a computer program stored in the computer. For example, a computer system such as computing system 1100 or other data processing system may perform the techniques described above in response to its processor executing a computer program (e.g., a sequence of instructions) contained in memory or other non-temporary machine-readable storage medium. Such computer programs may be stored in computer-readable storage media, each coupled to a computer system bus, including, but not limited to, any type of disk including floppy disks, optical disks, CD-ROMs, and magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, or any type of medium suitable for storing electronic instructions.

[0224] The algorithms and representations presented herein are not inherently related to any particular computer or other device. Various general-purpose systems can be used with the programs in accordance with the teachings herein, or it may be convenient to construct more specialized devices for performing the methods. Various structures for these systems will appear as described below. Furthermore, this disclosure is not described with reference to any particular programming language. It will be understood that various programming languages ​​can be used to implement the teachings of this disclosure as described herein.

[0225] This disclosure may be provided as a computer program product or software that includes a machine-readable medium storing instructions that can be used to program a computer system (or other electronic device) to perform the processes described herein. The machine-readable medium includes a mechanism for storing information in a machine-readable form. In some embodiments, the machine-readable (e.g., computer-readable) medium includes machine-readable (e.g., computer) storage media such as read-only memory ("ROM"), random-access memory ("RAM"), magnetic disk storage media, optical storage media, and flash memory components.

[0226] Exemplary embodiments of the technology disclosed herein are provided below. One embodiment of the technology may include any of the embodiments described herein, any combination of any of the embodiments described herein, or any combination of any part of the embodiments described herein.

[0227] In some embodiments, the techniques described herein are methods comprising: an extraction step of extracting phenotypic data from clinical records associated with a patient population and a genetic condition, wherein the phenotypic data includes natural language text containing at least one of indications for testing, a description of family history, or demographic data; an extraction step of extracting genotype data from genetic test results associated with a genetic condition, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnostic associated with the at least one variant; an output step of configuring a natural language processing (NLP) based machine learning model to generate and output a patient score for a patient population using the phenotypic data and genotype data containing natural language text, wherein the patient score includes a machine learning likelihood that, given the phenotypic data and genotype data containing natural language text, the patient score includes a machine learning likelihood that the patient has a genetic condition; and a configuration step of configuring a Bayesian causal model to generate and output a variant score using the patient score and genotype data, wherein the variant score includes a machine learning likelihood that the variant is pathogenic with respect to the genetic condition.

[0228] In some embodiments, the techniques described herein further include the step of using a configured NLP-based machine learning model and a configured Bayesian causal model to generate and output predictions about whether a variant associated with an unknown clinical record is benign or pathogenic in terms of its genetic state.

[0229] In some embodiments, the techniques described herein further include a step of providing a clinician with a prediction of whether a variant is benign or pathogenic for use in formulating a patient's diagnosis.

[0230] In some embodiments, the techniques described herein further include a step of providing predictions to at least one system, process, model, or component of a variant classification system or a clinical data system.

[0231] In some embodiments, the techniques described herein further include the step of using variant scores output by a constructed Bayesian causal model as input to a classification framework.

[0232] In some embodiments, the techniques described herein relate to methods, the steps of which constitute a natural language processing (NLP) based machine learning model include: a step of configuring a first model to receive natural language text and generate and output a text score, the text score including a machine learning likelihood that, given the received natural language text, the patient has a genetic condition; and a step of configuring a second model to generate and output a patient score using a second model input including phenotypic data excluding the text score and natural language text.

[0233] In some embodiments, the techniques described herein further include the step of configuring an NLP-based machine learning model to generate and output patient scores using curriculum learning, wherein the curriculum learning method includes the steps of determining the order of second model inputs using text scores, and applying the second model to the second model inputs in order.

[0234] In some embodiments, the technique described herein further includes the step of generating a configured NLP-based machine learning model by iteratively adjusting the values ​​of at least one clinical feature weight of an NLP-based machine learning model until the patient score output by the NLP-based machine learning model satisfies at least one first performance criterion, wherein the configured NLP-based machine learning model can output a patient score that satisfies at least one second performance criterion.

[0235] In some embodiments, the techniques described herein further include a step of constructing an NLP-based machine learning model using positive and negative examples, where the positive examples include a positive association between a molecular diagnosis and a genetic state, and the negative examples include a negative association between a molecular diagnosis and a genetic state.

[0236] In some embodiments, the techniques described herein relate to a method in which a Bayesian causal model includes probability distributions of patient scores and variant scores.

[0237] In some embodiments, the technology described herein is a system comprising at least one processor and at least one memory coupled to the at least one processor, wherein the at least one memory, when executed by the at least one processor, extracts phenotypic data from clinical records associated with a patient population and genetic status, wherein the phenotypic data includes natural language text containing at least one of indications for testing, a description of family history, or demographic data, and generates and outputs patient scores for the patient population using the phenotypic data including natural language text. A relates to a system comprising: generating and outputting, given phenotypic data including natural language text, the likelihood that a patient has a genetic condition; extracting genotype data from a patient population and genetic test results associated with the genetic condition, wherein the genotype data includes a molecular diagnostic including at least one variant of a gene; and constructing a Bayesian causal model to generate and output a variant score using the patient score and genotype data, wherein the variant score includes a machine learning likelihood that the variant is pathogenic with respect to the genetic condition.

[0238] In some embodiments, the techniques described herein relate to a system in which, when at least one instruction is executed by at least one processor, causes at least one processor to perform at least one operation, which further includes using a configured Bayesian causal model to generate and output a prediction about whether a variant associated with an unknown clinical record is benign, of unknown significance, or pathogenic with respect to a genetic condition.

[0239] In some embodiments, the technology described herein relates to a system in which, when at least one instruction is executed by at least one processor, causes at least one processor to perform at least one operation, further comprising providing a clinician with a prediction of whether a variant is benign or pathogenic for use by the clinician in formulating a diagnosis of a patient.

[0240] In some embodiments, the techniques described herein relate to a system in which, when at least one instruction is executed by at least one processor, causes at least one processor to perform at least one operation, which further includes providing predictions to at least one system, process, model, or component of a variant classification system or a clinical data system.

[0241] In some embodiments, the techniques described herein relate to a system in which, when at least one instruction is executed by at least one processor, causes at least one processor to perform at least one operation, which further includes using the variant scores output by a configured Bayesian causal model as input to a variant classification framework.

[0242] In some embodiments, the techniques described herein relating to a system, when at least one instruction is executed by at least one processor, causes at least one processor to perform at least one operation that includes (i) configuring a first model to generate and output a text score, the text score including a machine learning likelihood that a patient has a genetic condition, given a first model input including natural language text; and (ii) configuring a second model to generate and output a patient score, using a second model input including phenotypic data excluding the text score and natural language text.

[0243] In some embodiments, the techniques described herein relate to a system, wherein at least one instruction, when executed by at least one processor, causes at least one processor to configure an NLP-based machine learning model to generate and output patient scores using curriculum learning, the curriculum learning further comprising configuring to determine the order of second model inputs using text scores, and applying the second model to the second model inputs in that order.

[0244] In some embodiments, the techniques described herein relate to a system in which, when at least one instruction is executed by at least one processor, causes at least one processor to perform at least one operation further comprising: configuring a natural language processing (NLP)-based machine learning model to generate and output a patient score; and iteratively adjusting the value of at least one clinical feature weight of the NLP-based machine learning model to generate a configured NLP-based machine learning model until the patient score output by the NLP-based machine learning model satisfies at least one first performance criterion, the configured NLP-based machine learning model being able to output a patient score that satisfies at least one second performance criterion.

[0245] In some embodiments, the techniques described herein relate to a system in which, when at least one instruction is executed by at least one processor, causes at least one processor to perform at least one operation, further comprising constructing an NLP-based machine learning model using positive and negative examples, where the positive examples include positive associations between molecular diagnoses and genetic conditions, and the negative examples include negative associations between molecular diagnoses and genetic conditions.

[0246] In some embodiments, the techniques described herein relate to a system in which a Bayesian causal model includes probability distributions of patient scores and variant scores.

[0247] In some embodiments, the techniques described herein include at least one non-temporary machine-readable storage medium containing at least one instruction, the at least one instruction, when executed by at least one processor, to extract phenotypic data from clinical records associated with a patient population and genetic status, the phenotypic data including natural language text containing at least one of indications for testing, a family history description, or a demographic description, and using the phenotypic data including the natural language text to construct a natural language processing (NLP) based machine learning model for patient scores of the patient population. The present invention relates to at least one non-temporary machine-readable storage medium that causes at least one operation to be performed, which includes generating and outputting a patient score, which, given phenotypic data including natural language text, includes a machine learning likelihood that the patient has a genetic condition; extracting genotype data from a patient population and genetic test results associated with the genetic condition, which includes a molecular diagnostic including at least one variant of a gene; and using the patient score and genotype data, generating and outputting a variant score that includes the likelihood that the variant is pathogenic with respect to the genetic condition.

[0248] In some embodiments, the techniques described herein relate to at least one non-temporary machine-readable storage medium, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, further comprising using a configured NLP-based machine learning model to generate and output a prediction about whether a variant associated with an unknown clinical record is benign or pathogenic in terms of its genetic condition.

[0249] In some embodiments, the technology described herein relates to at least one non-temporary machine-readable storage medium, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, further comprising providing a clinician with a prediction of whether a variant is benign or pathogenic for use by the clinician in formulating a diagnosis of a patient.

[0250] In some embodiments, the techniques described herein relate to at least one non-temporary machine-readable storage medium, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, further comprising providing predictions to at least one system, process, model, or component of a variant classification system or a clinical data system.

[0251] In some embodiments, the techniques described herein relate to at least one non-temporary machine-readable storage medium, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, further comprising using variant scores as input to a classification framework.

[0252] In some embodiments, the techniques described herein, relating to at least one non-temporary machine-readable storage medium, further include configuring an NLP-based machine learning model to receive natural language text and generate and output a text score, the text score including a machine learning likelihood that a patient has a genetic condition given the received natural language text; and configuring a second model to generate and output a patient score using a second model input including the text score and second phenotypic data not including natural language text.

[0253] In some embodiments, the techniques described herein, relating to at least one non-temporary machine-readable memory medium, further include configuring an NLP-based machine learning model to generate and output patient scores using curriculum learning, wherein the curriculum learning includes determining the order of second model inputs using text scores and applying the second model to the second model inputs in that order.

[0254] In some embodiments, the techniques described herein relate to at least one non-temporary machine-readable storage medium, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, which further includes iteratively adjusting the values ​​of at least one clinical feature weight of an NLP-based machine learning model to produce a configured NLP-based machine learning model until the patient score output by the NLP-based machine learning model satisfies at least one first performance criterion, the configured NLP-based machine learning model can output a patient score that satisfies at least one second performance criterion.

[0255] In some embodiments, the techniques described herein relate to at least one non-temporary machine-readable storage medium, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, further comprising constructing an NLP-based machine learning model using positive and negative examples, where the positive examples include positive associations between molecular diagnoses and genetic conditions, and the negative examples include negative associations between molecular diagnoses and genetic conditions.

[0256] In some embodiments, the techniques described herein relate to at least one non-temporary machine-readable storage medium, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, which further includes generating and outputting variant scores using a Bayesian causal model, wherein the Bayesian causal model includes probability distributions of patient scores and variant scores.

[0257] In some embodiments, the techniques described herein are methods comprising: an extraction step of extracting phenotypic data from a set of clinical records associated with a patient population and a genetic condition, wherein the phenotypic data includes natural language text containing at least one of indications for testing, a description of family history, or demographic data; an extraction step of extracting genotype data from genetic test results associated with a genetic condition, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnostic associated with the at least one variant; and a constructing step of configuring a machine learning model to generate and output a patient score for a patient population using the phenotypic data and genotype data, which include natural language text, wherein the patient score includes a machine learning likelihood that the patient has the genetic condition, given phenotypic data including natural language text.

[0258] In some embodiments, the techniques described herein are methods comprising: an extraction step of extracting genotype data from genetic test results associated with a genetic condition, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnostic associated with the at least one variant; and a constructing step of constructing a Bayesian causal model to generate and output a variant score using patient scores and genotype data, wherein the variant score includes a machine learning likelihood that the variant is pathogenic with respect to the genetic condition.

[0259] In some embodiments, the techniques described herein relate to any one or more embodiments, steps, components, elements, processes, or limitations that are described in the accompanying description and / or shown in the accompanying drawings.

[0260] Clause 1. A method comprising: an extraction step of extracting phenotypic data from clinical records associated with a patient population and a genetic condition, wherein the phenotypic data includes natural language text containing at least one of indications for testing, a description of family history, or demographic data; an extraction step of extracting genotype data from genetic test results associated with a genetic condition, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnostic associated with at least one variant; a natural language processing (NLP) based machine learning model to generate and output a patient score for a patient population using the phenotypic data and genotype data containing natural language text, wherein the patient score includes a machine learning likelihood that, given the phenotypic data containing natural language text, the patient has a genetic condition; and a constructing step of a Bayesian causal model to generate and output a variant score using the patient score and genotype data, wherein the variant score includes a machine learning likelihood that the variant is pathogenic with respect to the genetic condition.

[0261] Clause 2. The method according to Clause 1, further comprising the step of using a configured NLP-based machine learning model and a configured Bayesian causal model to generate and output a prediction about whether a variant associated with an unknown clinical record is benign or pathogenic in terms of its genetic state.

[0262] Clause 3. The method of Clause 2, further comprising the step of providing a clinician with a prediction of whether the variant is benign or pathogenic for use by the clinician in formulating a diagnosis of the patient.

[0263] Clause 4. The method according to Clause 2, further comprising the step of providing predictions to at least one system, process, model, or component of a variant classification system or a clinical data system.

[0264] Clause 5. The method according to Clause 2, further comprising the step of using the variant scores output by the constructed Bayesian causal model as input to a classification framework.

[0265] Clause 6. The method according to any one of Clauses 1 to 5, further comprising: a step of configuring a natural language processing (NLP) based machine learning model, the step of configuring a first model to receive natural language text and generate and output a text score, the text score including a machine learning likelihood that a patient has a genetic condition given the received natural language text; and a step of configuring a second model to generate and output a patient score using a second model input including phenotypic data excluding the text score and natural language text.

[0266] The method according to Clause 7, further comprising the step of configuring an NLP-based machine learning model to generate and output patient scores using curriculum learning, wherein the curriculum learning includes the steps of determining the order of second model inputs using text scores, and applying the second model to the second model inputs in order.

[0267] Clause 8. The method according to any one of Clauses 1 to 7, further comprising the step of iteratively adjusting the values ​​of at least one clinical feature weight of the NLP-based machine learning model until the patient score output by the NLP-based machine learning model satisfies at least one first performance criterion, wherein the configured NLP-based machine learning model is capable of outputting a patient score that satisfies at least one second performance criterion.

[0268] Clause 9. The method described in any one of Clauses 1 to 8, further comprising the step of constructing an NLP-based machine learning model using positive and negative examples, wherein the positive examples include a positive association between a molecular diagnosis and a genetic state, and the negative examples include a negative association between a molecular diagnosis and a genetic state.

[0269] Clause 10. A Bayesian causal model is defined in any one of Clauses 1 to 9, including the probability distributions of patient scores and variant scores.

[0270] Clause 11. A system comprising at least one processor and at least one memory coupled to the at least one processor, wherein the at least one memory, when executed by the at least one processor, extracts phenotypic data from clinical records associated with a patient population and genetic status, wherein the phenotypic data includes natural language text containing at least one of indications for testing, a description of family history, or demographic data; and generates and outputs patient scores for a patient population using the phenotypic data containing natural language text. A system comprising: taking phenotypic data including word text as given, outputting the likelihood that a patient has a genetic condition; extracting genotype data from a patient population and genetic test results associated with the genetic condition, wherein the genotype data includes a molecular diagnostic including at least one variant of a gene; and constructing a Bayesian causal model to generate and output a variant score using the patient score and genotype data, wherein the variant score includes a machine learning likelihood that the variant is pathogenic with respect to the genetic condition.

[0271] Clause 12. The system according to Clause 11, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, further comprising using a configured Bayesian causal model to generate and output a prediction about whether a variant associated with an unknown clinical record is benign, of unknown significance, or pathogenic with respect to a genetic condition.

[0272] Clause 13. The system as described in Clause 12, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, further comprising providing a clinician with a prediction of whether a variant is benign or pathogenic for use by the clinician in formulating a diagnosis of a patient.

[0273] Clause 14. The system described in Clause 12, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, further comprising providing predictions to at least one system, process, model, or component of a variant classification system or a clinical data system.

[0274] Clause 15. The system described in any one of Clauses 11 to 14, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, further comprising using the variant scores output by the configured Bayesian causal model as input to a classification framework.

[0275] The System according to any one of Clauses 11 to 15, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation comprising: (i) configuring a first model to generate and output a text score, the text score including a machine learning likelihood that a patient has a genetic condition, given a first model input including natural language text; and (ii) configuring a second model to generate and output a patient score, using a second model input including phenotypic data excluding the text score and natural language text.

[0276] The System as described in Clause 16, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, which further comprises configuring an NLP-based machine learning model to generate and output patient scores using curriculum learning, wherein the curriculum learning comprises determining the order of second model inputs using text scores, and applying the second model to the second model inputs in that order.

[0277] Clause 18. The system described in any one of Clauses 11 to 17, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation further comprising: configuring a natural language processing (NLP) based machine learning model to generate and output a patient score; and generating a configured NLP based machine learning model by iteratively adjusting the values ​​of at least one clinical feature weight of the NLP based machine learning model until the patient score output by the NLP based machine learning model satisfies at least one first performance criterion, the configured NLP based machine learning model being able to output a patient score that satisfies at least one second performance criterion.

[0278] Clause 19. A system described in any one of Clauses 11 to 18, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, further comprising constructing an NLP-based machine learning model using positive and negative examples, where the positive examples include positive associations between molecular diagnostics and genetic conditions, and the negative examples include negative associations between molecular diagnostics and genetic conditions.

[0279] Clause 20. A Bayesian causal model is a system described in any one of Clauses 11-19, including the probability distributions of patient scores and variant scores.

[0280] Clause 21. At least one non-transitory machine-readable storage medium comprising at least one instruction, wherein the at least one instruction, when executed by at least one processor, causes the at least one processor to perform at least one operation comprising: extracting phenotype data from clinical records associated with a patient population and a genetic condition, wherein the phenotype data comprises natural language text including at least one of an indication for testing, a description of family history, or a demographic description; configuring a natural language processing (NLP)-based machine learning model using the phenotype data comprising the natural language text to generate and output a patient score for a patient of the patient population, wherein the patient score comprises a machine learning likelihood that the patient has the genetic condition given the phenotype data comprising the natural language text; extracting genotype data from genetic test results associated with the patient population and the genetic condition, wherein the genotype data comprises a molecular diagnosis including at least one variant of a gene; and generating and outputting a variant score comprising a likelihood that the variant is pathogenic with respect to the genetic condition using the patient score and the genotype data.

[0281] Clause 22. The at least one non-transitory machine-readable storage medium of Clause 21, wherein the at least one instruction, when executed by the at least one processor, further causes the at least one processor to perform at least one operation comprising generating and outputting a prediction as to whether a variant associated with an unknown clinical record is benign or pathogenic with respect to the genetic condition using the configured NLP-based machine learning model.

[0282] Clause 23. The at least one non-transitory machine-readable storage medium of Clause 22, wherein the at least one instruction, when executed by the at least one processor, further causes the at least one processor to perform at least one operation comprising providing the prediction as to whether the variant is benign or pathogenic to a clinician for use by the clinician in formulating a diagnosis of a patient.

[0283] Clause 24. The at least one non-transitory machine-readable storage medium according to Clause 22, wherein when executed by at least one processor, the at least one instruction causes the at least one processor to perform at least one operation further comprising providing a prediction to at least one system, process, model, or component of at least one of a variant classification system or a clinical data system.

[0284] Clause 25. The at least one non-transitory machine-readable storage medium according to any one of Clauses 21 to 24, wherein when executed by at least one processor, the at least one instruction causes the at least one processor to perform at least one operation further comprising using a variant score as an input to a classification framework.

[0285] Clause 26. The at least one non-transitory machine-readable storage medium according to any one of Clauses 21 to 25, wherein configuring the NLP-based machine learning model further comprises: configuring a first model to receive natural language text, and generate and output a text score, the text score comprising a machine learning likelihood that a patient suffers from a genetic condition given the received natural language text; and configuring a second model to generate and output a patient score using a second model input comprising the text score and second phenotypic data that does not comprise natural language text.

[0286] Clause 27. The at least one non-transitory machine-readable storage medium according to Clause 26, wherein configuring the NLP-based machine learning model further comprises configuring the NLP-based machine learning model to generate and output a patient score using curriculum learning, the curriculum learning comprising: determining an order of second model inputs using the text score; and applying the second model to the second model inputs according to the order.

[0287] Clause 28. At least one non-temporary machine-readable storage medium as described in any one of Clauses 21 to 27, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, which further includes iteratively adjusting the values ​​of at least one clinical feature weight of an NLP-based machine learning model to produce a configured NLP-based machine learning model until the patient score output by the NLP-based machine learning model satisfies at least one first performance criterion, and the configured NLP-based machine learning model can output a patient score that satisfies at least one second performance criterion.

[0288] Clause 29. At least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, which further comprises constructing an NLP-based machine learning model using positive and negative examples, wherein the positive examples include positive associations between molecular diagnostics and genetic conditions, and the negative examples include negative associations between molecular diagnostics and genetic conditions, on at least one non-temporary machine-readable storage medium as described in any one of Clauses 21 to 28.

[0289] Clause 30. At least one non-temporary machine-readable storage medium as described in any one of Clauses 21 to 29, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, further comprising outputting a variant score using a Bayesian causal model, the Bayesian causal model including probability distributions of patient scores and variant scores.

[0290] Clause 31. A method comprising: extracting phenotypic data from clinical records associated with a patient population and a genetic condition, wherein the phenotypic data includes natural language text containing at least one of indications for testing, a description of family history, or demographic data; extracting genotype data from genetic test results associated with a genetic condition, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnostic associated with at least one variant; and configuring a machine learning model to generate and output a patient score for a patient population using the phenotypic data and genotype data, which include natural language text, wherein the patient score includes a machine learning likelihood that the patient has the genetic condition, given phenotypic data containing natural language text.

[0291] Clause 32. A method comprising: extracting genotype data from genetic test results associated with a genetic condition, wherein the genotype data comprises at least one variant of a gene and at least one molecular diagnostic associated with the at least one variant; and constructing a Bayesian causal model to generate and output a variant score using the patient score and genotype data, wherein the variant score comprises a machine learning likelihood that the variant is pathogenic with respect to the genetic condition.

[0292] Clause 33. The method described in Clause 31 or 32, including any of the preceding clauses.

[0293] In the above-mentioned specification, embodiments of the disclosure are described with reference to specific exemplary embodiments. It will be apparent that various modifications can be made without departing from the broader spirit and scope of the embodiments of the disclosure described in the following claims. The specification and drawings are therefore considered illustrative, not restrictive.

Claims

1. It is a method, A step of extracting phenotypic data from clinical records associated with a patient population and genetic status, wherein the phenotypic data includes natural language text containing at least one of the following: indications for testing, a description of family history, or demographic data; A step of extracting genotype data from the results of a genetic test associated with the aforementioned genetic condition, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnostic associated with the at least one variant; A step of configuring a natural language processing (NLP) based machine learning model to generate and output patient scores for the patient population using the natural language text and the phenotypic data including the genotype data, wherein the patient score includes a machine learning likelihood that the patient has the genetic condition, given the phenotypic data including the natural language text; A method comprising the steps of: configuring a Bayesian causal model to generate and output a variant score using the patient score and the genotype data, wherein the variant score includes a machine learning likelihood that the variant is pathogenic with respect to the genetic condition.

2. The method according to claim 1, further comprising the step of using the configured NLP-based machine learning model and the configured Bayesian causal model to generate and output a prediction as to whether a variant associated with an unknown clinical record is benign or pathogenic with respect to the genetic condition.

3. The method of claim 2, further comprising the step of providing the clinician with the prediction regarding whether the variant is benign or pathogenic for use by the clinician in formulating a diagnosis of a patient.

4. The method according to claim 2, further comprising the step of providing the prediction to at least one system, process, model, or component of a variant classification system or a clinical data system.

5. The method according to claim 2, further comprising the step of using the variant score output by the constructed Bayesian causal model as input to a variant classification framework.

6. The step of constructing the aforementioned natural language processing (NLP) based machine learning model is: The steps include: configuring a first model to receive the natural language text and generate and output a text score, wherein the text score includes a machine learning likelihood that the patient has the genetic condition, given the received natural language text; The method according to claim 1, further comprising the step of configuring a second model to generate and output the patient score using a second model input which includes the text score and the phenotypic data excluding the natural language text.

7. The method according to claim 6, further comprising the step of configuring the NLP-based machine learning model to generate and output the patient scores using curriculum learning, wherein the curriculum learning includes the steps of determining the order of the second model inputs using the text scores, and applying the second model to the second model inputs in the order.

8. The method according to claim 1, further comprising the step of generating the configured NLP-based machine learning model by iteratively adjusting the values ​​of at least one clinical feature weight of the NLP-based machine learning model until the output of a loss function calculated based on the patient score output by the NLP-based machine learning model satisfies at least one first performance criterion, wherein the configured NLP-based machine learning model can output a patient score that satisfies at least one second performance criterion.

9. The method according to claim 1, further comprising the step of constructing the NLP-based machine learning model using positive and negative examples, wherein the positive examples include a positive association between the molecular diagnosis and the genetic state, and the negative examples include a negative association between the molecular diagnosis and the genetic state.

10. The method according to claim 1, wherein the Bayesian causal model includes probability distributions of patient scores and variant scores.

11. It is a system, At least one processor; The system comprises at least one memory coupled to the at least one processor, and when the at least one memory is executed by the at least one processor, it provides to the at least one processor, Extracting phenotypic data from clinical records associated with patient populations and genetic conditions, wherein the phenotypic data includes natural language text containing at least one of the following: indications for testing, descriptions of family history, or demographic data; Using the phenotypic data including the natural language text, generate and output patient scores for the patient population, wherein the patient scores include the likelihood that a patient has the genetic condition, given the phenotypic data including the natural language text; Extracting genotype data from the patient population and the results of genetic tests associated with the genetic condition, wherein the genotype data includes a molecular diagnosis involving at least one variant of a gene; A system comprising: configuring a Bayesian causal model to generate and output a variant score using the patient score and the genotype data, wherein the variant score includes a machine learning likelihood that the variant is pathogenic with respect to the genetic condition; and at least one instruction causing the system to perform at least one action.

12. The system according to claim 11, wherein when the at least one instruction is executed by the at least one processor, the at least one processor further comprises using the configured Bayesian causal model to generate and output a prediction about whether the variant associated with the unknown clinical record is benign, of unknown significance, or pathogenic with respect to the genetic condition.

13. When the at least one instruction is executed by the at least one processor, the at least one processor will, The system according to claim 12, further comprising performing at least one operation, which includes providing the clinician with the prediction of whether the variant is benign or pathogenic for use by the clinician in formulating a diagnosis of a patient.

14. When the at least one instruction is executed by the at least one processor, the at least one processor will, The system according to claim 12, further comprising causing it to perform at least one operation which includes providing the prediction to at least one system, process, model, or component of a variant classification system or a clinical data system.

15. When the at least one instruction is executed by the at least one processor, the at least one processor will, The system according to claim 11, further comprising performing at least one operation, which is to use the variant score output by the configured Bayesian causal model as input to a variant classification framework.

16. When the at least one instruction is executed by the at least one processor, the at least one processor will, The system according to claim 11, which further comprises: (i) configuring a first model to generate and output a text score, wherein the text score includes a machine learning likelihood that the patient has the genetic condition, given a first model input including the natural language text; and (ii) configuring a second model to generate and output a patient score using a second model input including phenotypic data excluding the text score and the natural language text, thereby configuring a natural language processing (NLP) based machine learning model to generate and output the patient score.

17. When the at least one instruction is executed by the at least one processor, the at least one processor will, The system according to claim 16, wherein curriculum learning is used to configure the NLP-based machine learning model to generate and output the patient scores, the curriculum learning further comprises performing at least one operation including determining the order of the second model inputs using the text scores, and applying the second model to the second model inputs in that order.

18. When the at least one instruction is executed by the at least one processor, the at least one processor will, Configuring a natural language processing (NLP) based machine learning model to generate and output the aforementioned patient score; The system according to claim 11, further comprising performing at least one operation to generate the configured NLP-based machine learning model, which involves iteratively adjusting the values ​​of at least one clinical feature weight of the NLP-based machine learning model until the output of a loss function calculated based on the patient score output by the NLP-based machine learning model satisfies at least one first performance criterion, the configured NLP-based machine learning model being able to output a patient score that satisfies at least one second performance criterion.

19. When the at least one instruction is executed by the at least one processor, the at least one processor will, The system according to claim 18, further comprising performing at least one operation which includes constructing the NLP-based machine learning model using positive and negative examples, wherein the positive examples include a positive association between a molecular diagnosis and the genetic state, and the negative examples include a negative association between the molecular diagnosis and the genetic state.

20. The system according to claim 11, wherein the Bayesian causal model includes probability distributions of patient scores and variant scores.

21. A non-temporary machine-readable storage medium containing at least one instruction, wherein, when the at least one instruction is executed by the at least one processor, the at least one processor receives Extracting phenotypic data from clinical records associated with patient populations and genetic conditions, wherein the phenotypic data includes natural language text containing at least one of the following: indications for testing, family history, or demographic descriptions; A natural language processing (NLP) based machine learning model is configured to generate and output patient scores for the patient population using the phenotypic data including the natural language text, wherein the patient scores include a machine learning likelihood that a patient has the genetic condition, given the phenotypic data including the natural language text; Extracting genotype data from the patient population and the results of genetic tests associated with the genetic condition, wherein the genotype data includes a molecular diagnosis involving at least one variant of a gene; A non-temporary machine-readable storage medium that causes at least one operation to be performed, which includes generating and outputting a variant score that includes the likelihood that a variant is pathogenic with respect to the genetic condition, using the patient score and the genotype data.

22. When the at least one instruction is executed by the at least one processor, the at least one processor will, The at least one non-temporary machine-readable storage medium according to claim 21, which is configured to perform at least one operation, further comprising generating and outputting a prediction as to whether a variant associated with an unknown clinical record is benign or pathogenic with respect to the genetic condition, using the NLP-based machine learning model configured above.

23. When the at least one instruction is executed by the at least one processor, the at least one processor will, The at least one non-temporary machine-readable storage medium according to claim 22, which further comprises performing at least one operation, which includes providing the clinician with the prediction of whether the variant is benign or pathogenic for use by the clinician in formulating a diagnosis of a patient.

24. When the at least one instruction is executed by the at least one processor, the at least one processor will, The at least one non-temporary machine-readable storage medium according to claim 22, which causes the medium to perform at least one operation, further comprising providing the prediction to at least one system, process, model, or component of a variant classification system or a clinical data system.

25. When the at least one instruction is executed by the at least one processor, the at least one processor will, The at least one non-temporary machine-readable storage medium according to claim 21, which causes it to perform at least one operation further comprising using the variant score as input to a variant classification framework.

26. Configuring the aforementioned NLP-based machine learning model is A first model is configured to receive the natural language text and generate and output a text score, wherein the text score includes a machine learning likelihood that, given the received natural language text, the patient has the genetic condition; The at least one non-temporary machine-readable storage medium according to claim 21, further comprising configuring a second model to generate and output the patient score using a second model input which includes the text score and a second phenotypic data which does not include the natural language text.

27. Configuring the aforementioned NLP-based machine learning model is The at least one non-temporary machine-readable storage medium according to claim 26, further comprising using curriculum learning to configure the NLP-based machine learning model to generate and output the patient scores, wherein the curriculum learning includes determining the order of the second model inputs using the text scores, and applying the second model to the second model inputs in the order described above.

28. When the at least one instruction is executed by the at least one processor, the at least one processor will, At least one non-temporary machine-readable storage medium according to claim 21, further comprising performing at least one operation to generate the configured NLP-based machine learning model by iteratively adjusting the values ​​of at least one clinical feature weight of the NLP-based machine learning model until the output of a loss function calculated based on the patient score output by the NLP-based machine learning model satisfies at least one first performance criterion, the configured NLP-based machine learning model is capable of outputting a patient score that satisfies at least one second performance criterion.

29. When the at least one instruction is executed by the at least one processor, the at least one processor will, The at least one non-temporary machine-readable storage medium according to claim 21, further comprising performing at least one operation which involves constructing the NLP-based machine learning model using positive and negative examples, wherein the positive examples include a positive association between the molecular diagnosis and the genetic state, and the negative examples include a negative association between the molecular diagnosis and the genetic state.

30. When the at least one instruction is executed by the at least one processor, the at least one processor will, The at least one non-temporary machine-readable storage medium according to claim 21, which causes the Bayesian causal model to perform at least one operation, further comprising generating and outputting variant scores, wherein the Bayesian causal model includes probability distributions of patient scores and variant scores.

31. It is a method, A step of extracting phenotypic data from a set of clinical records associated with a patient population and genetic status, wherein the phenotypic data includes natural language text containing at least one of the following: indications for testing, a description of family history, or demographic data; A step of extracting genotype data from the results of a genetic test associated with the aforementioned genetic condition, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnostic associated with the at least one variant; A method comprising the step of configuring a machine learning model to generate and output a patient score for a patient population using the natural language text and the phenotypic data including the genotype data, wherein the patient score includes a machine learning likelihood that the patient has the genetic condition given the phenotypic data including the natural language text.

32. It is a method, A step of extracting genotype data from the results of a genetic test associated with a genetic state, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnostic associated with the at least one variant; A method comprising the steps of: configuring a Bayesian causal model to generate and output a variant score using patient scores and the genotype data, wherein the variant score includes a machine learning likelihood that the variant is pathogenic with respect to the genetic condition.