Clinical variation modeling

By combining machine learning and Bayesian inference with natural language processing to model clinical variants, this approach solves the challenge of classifying novel variants in gene testing, achieving more accurate variant and patient scores, and improving the reliability and diagnostic accuracy of gene testing.

CN121844386APending Publication Date: 2026-04-10LABORATORY CORPORATION OF AMERICA HOLDINGS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing genetic testing methods struggle to effectively classify novel variants, especially those not adequately documented in the literature, leading to many variants being classified as VUS (Variations of Indeterminate Significance), which increases the complexity and uncertainty of clinical genetic testing.

Method used

We employ a machine learning-based clinical variant modeling approach. By generating patient scores and variant scores, we utilize a large clinical database and a Bayesian inference model, combined with natural language processing technology, to learn patterns in the patient population and calculate the pathogenicity probability of variants, thereby reducing uncertainty and improving classification accuracy.

Benefits of technology

It improves the accuracy and reliability of variant classification, effectively distinguishes between pathogenic and benign variants, reduces VUS, supports more detailed clinical diagnostic decisions, and provides individuals with more accurate genetic risk assessments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121844386A_ABST
    Figure CN121844386A_ABST
Patent Text Reader

Abstract

Examples may configure at least one machine learning model using phenotypic data including natural language text and genotypic data. The at least one machine learning model may generate and output a patient score. The patient score may include a machine learning likelihood that the patient is affected by a genetic condition given phenotypic data including natural language text. Some embodiments may use patient scores and genotype data to configure a Bayesian causal model to generate and output variation scores. The variation score may include a machine-learned likelihood that a variation of a gene is pathogenic to a genetic condition. The patient score and / or variation score may be provided to one or more downstream systems, devices, processes, components, or models.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 579939, filed August 31, 2023, and U.S. Provisional Patent Application No. 63 / 562696, filed March 7, 2024, each of which is incorporated herein by reference in its entirety. Technical Field

[0002] This application relates to the technical field of gene testing. Another technical field it relates to is machine learning-based variant classification systems. Background Technology

[0003] Genetic variation refers to the differences in DNA sequences between individuals within a population. Many different types of variation exist, including structural variations, single nucleotide polymorphisms, insertions and deletions, copy number variations, and translocations and inversions.

[0004] Gene sequencing technology continues to advance rapidly. For genetic diseases, high-throughput sequencing technology is increasingly enabling gene testing encompassing genotyping, single genes, gene packages, exomes, genomes, transcriptomes, and epigenetics. The increasing complexity and volume of clinical gene testing analysis and interpretation are accompanied by new challenges in interpreting sequence variations.

[0005] For example, clinical molecular laboratories are increasingly detecting novel sequence variations when testing patient samples for a rapidly growing number of genes associated with genetic diseases. While some phenotypes are associated with a single gene, many phenotypes are associated with multiple genes.

[0006] Variance classification refers to the process of categorizing gene variations based on evidence that supports or refutes a causal relationship with a disease. The clinical significance of any given sequence variation decreases along a gradient, ranging from those variations that are almost certainly pathogenic to those that are almost certainly benign.

[0007] Variation classification itself is not a diagnosis, but it can be used by clinicians to make diagnostic decisions. Attached Figure Description

[0008] This disclosure will be more fully understood from the detailed description given below and from the accompanying drawings of various embodiments thereof. The drawings are for explanation and understanding only and should not be construed as limiting this disclosure to the specific embodiments shown.

[0009] Figure 1 Examples of variation scoring processes according to some embodiments of this disclosure are shown.

[0010] Figure 2An example showing a process for generating patient scores using natural language processing based models, according to some embodiments of the disclosure.

[0011] Figure 3A An example process for generating text scores using natural language processing based models, according to some embodiments of the disclosure is shown.

[0012] Figure 3B An example process for generating patient scores using natural language processing based models, according to some embodiments of the disclosure is shown.

[0013] Figure 4 An example process for configuring a clinical variant model using machine learning, according to some embodiments of the disclosure is shown.

[0014] Figure 5A An example of patient score calculation, according to some embodiments of the disclosure is shown.

[0015] Figure 5B An example of patient score calculation, according to some embodiments of the disclosure is shown.

[0016] Figure 5C An example of patient score calculation using machine learning feature weights, according to some embodiments of the disclosure is shown.

[0017] Figure 6A An example of variant score calculation using a Bayesian model, according to some embodiments of the disclosure is shown.

[0018] Figure 6B An example of distribution of patient scores for variants, according to some embodiments of the disclosure is shown.

[0019] Figure 6D An example of distribution of variant scores and patient scores for classified variants, according to some embodiments of the disclosure is shown.

[0020] Figure 6E An example of distribution of variant scores and patient scores for classified variants, according to some embodiments of the disclosure is shown.

[0021] Figure 6F An example of distribution of patient scores and variant scores, according to some embodiments of the disclosure is shown.

[0022] Figure 6G An example of distribution of classified variants and patient scores, according to some embodiments of the disclosure is shown.

[0023] Figure 7A An example of distribution of variant scores and patient scores for classified variants, according to some embodiments of the disclosure is shown.

[0024] Figure 7B Examples of the distribution of variant scores and patient scores for classified variants are shown according to some embodiments of this disclosure.

[0025] Figure 8A Examples of performance data for a patient score generator and a variant score generator according to some embodiments of this disclosure are shown.

[0026] Figure 8B Examples of clinical variation modeling and classification systems according to some embodiments of this disclosure are shown.

[0027] Figure 8C Examples of variant classification systems including clinical variant modeling are shown according to some embodiments of the present disclosure.

[0028] Figure 9A Methods for clinical variation modeling according to some embodiments of the present disclosure are shown.

[0029] Figure 9B Methods for clinical variation modeling according to some embodiments of the present disclosure are shown.

[0030] Figure 9C Methods for clinical variation modeling according to some embodiments of the present disclosure are shown.

[0031] Figure 10 An example computational system including clinical variation modeling is shown according to some embodiments of the present disclosure.

[0032] Figure 11 This is a block diagram of an example computer system in which aspects of this disclosure can be operated. Detailed Implementation

[0033] Genetic variants can be classified as pathogenic (i.e., causing disease), benign (not causing disease), possibly pathogenic, possibly benign, or of uncertain significance (VUS). Currently, many variants are classified as VUS because (among other things) there is insufficient information about them to make such a classification. As genetic testing is increasingly used in healthcare for disease diagnosis and management, the field of clinical genomics is encountering more and more novel variants that also require classification, including both common and rare novel variants. Simultaneously, the expansion of genetic testing has led to a continuously growing volume of data, including clinical data.

[0034] The embodiments of the clinical variation modeling methods described herein address these and / or other challenges by including one or more generators, such as a patient score generator that generates patient scores or a variation score generator that generates variation scores based on patient scores. While this disclosure describes methods that include both a patient score generator and a variation score generator, either the patient score generator or the variation score generator may be used independently of other components. For example, in some applications, the patient scores generated by the patient score generator may be useful independently of the variation scores generated by the variation score generator. As another example, the variation score generator may generate variation scores based on patient scores that are not obtained from a patient score generator as described herein, but from another source such as a database or a different type of patient score generator.

[0035] The clinical variant modeling methods described in this paper (e.g., including modeling pipelines such as the patient score generator and / or the variant score generator as described) utilize diverse genotypes and anonymized clinical data from a population of patients who have undergone genetic testing to classify or reclassify both novel variants and previously seen variants, thereby reducing variants of uncertain significance and / or improving the accuracy of previous variant classifications.

[0036] Implementations of the described clinical variant modeling method include a variant scoring generator configured as a Bayesian causal model, which has demonstrated high accuracy in incorporating clinical evidence into variant classification on a large scale. The described method can be used to improve variant classification for genes and diseases that have not yet been tested, leveraging information at scale as clinical datasets continue to grow. The described method can improve variant classification in a scalable manner and reduce uncertainty in genetic testing.

[0037] The described clinical variation modeling method is robust to dataset size. For example, the described method can generate reliable predictions on large datasets (e.g., three to four million patients) as well as on smaller datasets. The variation score generator enables the utilization of domain expertise in sparse data scenarios. For the patient scoring model, a curriculum learning method improves performance in sparse scenarios.

[0038] One goal of clinical genetic testing is to assess the risk of hereditary diseases or confirm a diagnosis of a hereditary disease. Distinguishing between pathogenic DNA (DNA) variants and benign variants is a significant challenge during genetic testing. This challenge is further exacerbated by the increasing number of novel and rare variants encountered by clinical genetic testing laboratories as more individuals undergo genetic testing. These variants are not always well-documented in published literature or in the CLINVAR database, and are therefore often classified as variants of indeterminate significance (VUS).

[0039] Clinical data is one of the strongest forms of evidence for distinguishing between pathogenic and benign variants. For example, clinical data may include observations of variants in patients with a well-defined disease, observations of variants in patients decisively unaffected by a well-defined disease, variants co-segregated with a well-defined disease in affected individuals within a family, or observations of the de novo occurrence of a variant. However, incorporating clinical data into variant classification often presents challenges. For instance, a complete relevant medical history is not always available at the time of genetic testing, making it challenging to distinguish between affected individuals with missing data and those who are not affected.

[0040] Second, many genetic diseases include symptoms that may be associated with common sporadic illnesses such as cancer and cardiovascular disease, limiting the ability to identify molecular causes and establish genetic etiologies. Finally, not all genetic diseases exhibit full penetrance, making it difficult to determine when a variant is not associated with a genetic disease and is therefore benign. These challenges are particularly acute when reviewing clinical data on a case-by-case basis for variant classification. However, increasing access to large collections of clinical health information and genetic testing results can help overcome these challenges.

[0041] The disclosed method leverages accumulated genotypic and clinical data from populations comprising individuals of diverse racial and ethnic backgrounds referred for extensive clinical genetic testing. The datasets are large and clinically diverse: in some embodiments, they include over one hundred million words of clinical descriptions (e.g., personal and family history, indications for testing) submitted for the patients being tested, as well as over two million unique variants observed across more than 3,900 genes.

[0042] The disclosed clinical variant modeling method maximizes the utility of clinical datasets for improving variant classification and reducing VUS. Embodiments of the described method for clinical variant modeling use machine learning to determine patterns based on clinical data available to millions of patients, and apply this learning precisely as evidence for variant classification using Bayesian methods. The described method can machine learn relationships between different variables (e.g., statistical correlations) and outcomes associated with those variables, including but not limited to disease penetrance, age at test, potential pseudophenotypes, and missing data in clinical patient records (e.g., test request forms).

[0043] Some embodiments of clinical variant modeling described herein include two distinct but sequential machine learning (ML) steps. The first step involves estimating the probability that a patient who has undergone genetic testing is affected by a specific genetic condition. This probability is called a patient score, which incorporates clinical and demographic information from one or more medical providers, including reported signs and symptoms, ICD-10 codes, age at testing, and family history. Patient scores are estimated by comparing and differentiating the clinical profiles of patients with positive molecular diagnoses from those with negative molecular diagnoses. In the second step of clinical variant modeling, a Bayesian inference model learns the distribution of patient scores that may be associated with benign or pathogenic variants. The inferred probability that a variant is pathogenic is called a variant score.

[0044] More specifically, embodiments of the described clinical variation modeling method use a combination of natural language processing (NLP) and Bayesian inference to predict the pathogenicity of genetic variations using clinical data provided in patient records (e.g., test request forms). The resulting clinical variation modeling system is designed to learn which clinical features can distinguish patients with molecular diagnoses from genotype-negative controls.

[0045] As used herein, clinical variation modeling can refer to a modeling pipeline that includes one or more generators, such as a patient score generator and a variation score generator, while a clinical variation model can refer to a patient score generator, a variation score generator, or a combination of both. Either or both of the patient score generator and the variation score generator can include one or more machine learning models. For example, as described in more detail below, embodiments of a patient score generator include at least two machine learning models, such as an NLP component and a tree-based scoring model. Therefore, for example, some embodiments of a clinical variation model can include up to or at least three machine learning models (e.g., a first scoring model such as an NLP-based scoring model, a second scoring model such as a tree-based scoring model, and a third scoring model such as a Bayesian inference model).

[0046] In some embodiments, the clinical variant model is condition-specific (e.g., a neurofibromatosis type I model including NF1, or a Lynch syndrome model including MLH1, MSH2, MSH6, PMS2, and EPCAM). Therefore, in some embodiments, each clinical variant model can be trained and tested separately for a specific condition and its associated genes(s).

[0047] In some embodiments, the clinical variant model is gene-specific (e.g., the clinical variant model is specific to the MMR or another gene). Therefore, in some embodiments, each clinical variant model can be trained and tested separately for one or more specific genes.

[0048] The described clinical variation modeling approach differs from existing methods in several ways. First, existing methods for incorporating clinical data from monogenic conditions assess each patient's clinical information on a case-by-case basis. While this approach may be effective for some conditions, it is challenging for many genetic conditions that may be associated with common sporadic diseases (i.e., high pseudophenotype rates, such as cancer and cardiovascular disease), exhibit incomplete penetrance, and display variable expression. In contrast, the described clinical variation modeling approach can leverage a large clinical database from genetically tested patients (e.g., over four million patients), taking into account the distinction between those patients who share genetic variations and appear to have the discussed disease phenotype and those who do not (in the case of pseudophenotype).

[0049] Second, existing methods are generally binary (e.g., predicting either meeting or not meeting diagnostic criteria). In contrast, clinical variant models configured using the described method can calculate the probability of a given individual being affected on a continuous scale based on clinical information, and can also calculate the probability of a variant being pathogenic based on a continuous distribution of affected and unaffected states across all patients with the same variant. Therefore, scientists can use available clinical data in a more granular way.

[0050] Next, the described clinical variation modeling approach can learn patterns present in clinical data within a patient population and apply (or generalize) those learned patterns to other patients. This capability helps reveal patterns and trends in clinical data that might not be apparent in other ways, such as the frequency of clinicians using shorthand, abbreviations, terminology in other languages, and the lack of clinical information (regardless of patient impact), age distribution, and ICD-10 coding usage patterns. Therefore, pathogenicity predictions can be generalized to specific but broader patient populations.

[0051] Compared to existing methods that rely on patient data reported in the literature, the described clinical variation modeling approach depends solely on patients observed through testing. Therefore, existing methods may be biased towards groups of patients more likely to be reported in the literature (who are often more severely affected or have earlier onset compared to patients in the broader population), whereas the described approach is not biased in this way.

[0052] Furthermore, since clinical variant modeling, as described, is based on machine learning methods, these models can be updated regularly as more clinician patient data and / or more variant information become available in the medical genetics community.

[0053] Because clinical variant modeling is a machine learning approach that learns specific features (e.g., specific ICD-10 codes, free-text clinical information, free-text family history information, etc.), each feature is not necessarily weighted equally, given that specific feature best distinguishes patients with a molecular diagnosis from those with a negative genotype. Furthermore, the weighting of each feature may vary depending on genetic status. A clinical variant model for a given condition can automatically learn which features best predict pathogenicity and weight them appropriately.

[0054] The technical challenge of using clinical data for variant classification lies in the fact that the format, quantity, and / or quality of information contained in patient records can vary significantly from patient to patient. For example, some patient records may have detailed textual explanations of indications and / or family history, while others may have only a few vague keywords or no information at all in these fields. The clinical variant modeling approach described herein can adapt to the variability of available information in patient records, including missing information. Because different features are weighted differently for each gene / condition, the impact of missing or sparse information depends on the specific information missing for a particular model. Because the model learns at a macroscopic level, it learns that some degree of missing information exists when features are missing.

[0055] Implementations of the described clinical variant models have been trained, tested, and validated for clinical variant classification based on clinical information available in a proprietary database for patients and variants. Future plans include updating these models as clinical cohorts grow and as more genotypic and phenotypic information becomes available due to additional patient genetic testing. Future updates to the models should undergo rigorous training, testing, and clinical validation before being implemented in variant classification.

[0056] The described method does not use or rely on external clinical data (e.g., from publications). Instead, the described clinical variant models learn from clinical data obtained during genetic testing procedures in the laboratory and can be applied to new patients and variants observed in the laboratory. Experimental results have shown that clinical variant models as described are better at predicting individuals within a cohort because these models understand data patterns in the sampling and take into account generalizability to predict specific patient populations. External clinical data from publications will not have the same data patterns in these patient samples as specific patient populations. However, evaluation of external clinical data (such as those from publications) can still be used to supplement clinical variant modeling as described.

[0057] Even when clinical patient records are incomplete or blank, the described method can generate meaningful / accurate patient scores and variation scores based on clinical data provided on clinical patient records (e.g., test request forms or TRFs). This is because the clinical variation model described has been trained, tested, and validated using a large number of clinical records.

[0058] To calculate variant scores, embodiments of the clinical variant model analyze the distribution of patient scores for a given variant as a whole and compare that distribution to the distribution of patient scores seen in known pathogenic variants against known benign variants. These known variants may have a similar number of missing or incomplete clinical information for the patients being tested. Because the clinical variant model looks at patterns across a large number of patients and variants, the impact of incomplete clinical information is far less than when the modeling is applied to a single patient or a small number of patients. Furthermore, because variant scores can be calculated for all variants (even those classified as pathogenic and benign), the described method can help identify variants that may be misclassified, for example, for further expert review.

[0059] The described clinical variant model captures the distribution of patient scores across the spectrum from unaffected to affected states. In contrast, existing methods often overlook clinical data from seemingly unaffected individuals during variant classification due to concerns that it might simply be due to incomplete patient clinical records. Unlike existing methods, the clinical variant model actually examines the frequency with which patients appear unaffected within both the molecular diagnostic cohort (either due to incomplete penetrance or incomplete patient clinical records) and the genotype-negative cohort. The described method learns how much weight should be assigned to a given observation. Even if a single observation is not very meaningful, it may become meaningful enough across dozens or hundreds of observations to confidently reclassify variants.

[0060] Given the high performance of these clinical variant models in distinguishing known benign from known pathogenic variants for specific genes and / or gene-disease combinations, they can be incorporated into variant classification processes. However, while seemingly promising clinical variant models have been developed for a wider range of genes and conditions, these models have not yet been fully reviewed and validated. Once these models have been thoroughly evaluated for clinical effectiveness, the set of available clinical variant models can be expanded.

[0061] To help demonstrate the consistency between clinical variant model predictions and known pathogenic variants and to build confidence that the clinical variant model is working appropriately, pathogenicity evidence for known pathogenic variants is collected and included in the modeling (where available). Furthermore, if contradictory evidence emerges later, it is documented as it may be helpful in the future, thus providing clinicians with all possible information to determine whether a reclassification as LP (likely pathogenic), VUS (variable of indeterminate significance), LB (likely benign), or B (benign) is necessary. It is anticipated that applying pathogenicity evidence from clinical variant modeling to pathogenic variants will further increase the stability of these pathogenicity classifications.

[0062] For genes associated with multiple conditions, if the molecular mechanisms of the diseases are the same (e.g., loss of function [LOF] or gain of function [GOF]), it may not be necessary to train the clinical variant models described as such separately for discrete gene-disease associations. The clinical variant models described as such can learn different combinations of clinical features seen in patients with different conditions caused by variations in the same gene. While genes associated with multiple conditions (due to different molecular mechanisms or different inheritance patterns) is a complex problem under investigation, there are examples where the described modeling approach appears to work reliably for genes associated with multiple diseases of the same mechanism but different inheritance patterns (e.g., MSH2, MSH6, MLH1, PMS2 associated with dominant LoF Lynch and recessive LoF constitution mismatch repair deficiency). Furthermore, there are examples where the described approach appears to work reliably using a single model for genes associated with multiple diseases of the same inheritance pattern but different molecular mechanisms (e.g., LoFCASR and GoF CASR).

[0063] To ensure that only the best-performing clinical variant models are used for variant classification, the area under the receiver operating characteristic curve (AUROC) is calculated for each model to measure its performance in distinguishing between benign and pathogenic variants. In some embodiments, only models with an AUROC ≥ 0.8 are selected for further evaluation. In some embodiments, the ensemble of clinical variant models is continuously improved through a combination of further validation metrics and expert review. In some embodiments, additional steps are performed to establish weighted variance scores for inclusion as evidence in a variance classification framework, such as the INVITAE Sherloc framework (see, for example, Nykamp K, Anderson M, Powers M, Garcia J, Herrera B, Ho YY, Kobayashi Y, Patil N, Thusberg J, Westbrook M; InvitaeClinicalGenomicsGroup; Topper S. Sherloc: a comprehensive refinement of the ACMG-AMP variantclassification criteria. Genet Med. 2017 Oct;19(10):1105-1117.doi: 10.1038 / gim.2017.37. Epub 2017 May 11. Erratum in: Genet Med. 2020 Jan;22(1):240-242.PMID: 28492532; PMCID: PMC5632818).

[0064] Implementations of the Clinical Variation Models (CVMs) described above have been developed and validated for eleven genetic conditions, such as those associated with seventeen genes. These implementations have demonstrated high performance (≥0.8 AUROC curve) in distinguishing known benign from known pathogenic variants. Predictions from these CVMs have been used as evidence to resolve over 1,000 unique variants (VUS), affecting >45,000 individuals. While >99% of these reclassifications correspond to VUS downgrades to benign or possibly benign, <1% are upgraded to pathogenic or possibly pathogenic, affecting ~160 individuals. Notably, 91% (10 / 11) of the variant upgrades are in genes associated with conditions for which established screening and treatment guidelines exist, highlighting the potential to alter individual healthcare management and identify high-risk relatives.

[0065] This disclosure will be more fully understood from the detailed description given below with reference to the accompanying drawings. The detailed description of the drawings is for explanation and understanding purposes only and should not be construed as limiting this disclosure to the specific embodiments described.

[0066] In the accompanying drawings and the following description, components with the same name but different reference numerals may be mentioned in different drawings. The use of different reference numerals in different drawings indicates that components with the same name may represent the same embodiment or different embodiments of the same component. For example, in some embodiments, components with the same name but different reference numerals in different drawings may have the same or similar functionality, such that the description of one of the components in one drawing can be applied to other components with the same name in other drawings.

[0067] Furthermore, the components shown and described in conjunction with some embodiments in the accompanying drawings and the following description can be used with or incorporated into other embodiments. For example, a component shown in a particular drawing is not limited to use in the embodiment to which that drawing pertains, but can be used with or incorporated into other embodiments, including those shown in other drawings.

[0068] Figure 1 Examples of variation scoring processes according to some embodiments of this disclosure are shown.

[0069] exist Figure 1 In this context, each patient in patient group 102 has an associated clinical patient profile 104 and genetic testing results 106. The clinical patient profile 104 includes: unstructured data, such as natural language text descriptions including clinical descriptions, indications for genetic testing, and family history; and structured data, such as ICD-10 codes (International Classification of Diseases codes), sex indicators, age, and / or other demographic data. The genetic testing results 106 can identify one or more variants detected as a result of the genetic testing. In the case of identifying multiple variants, the genetic testing results can be filtered or sorted by variant to identify positive and negative patients for each variant.

[0070] Clinical patient profiles 104 and genetic testing results 106 are input into a patient score generator 108. The patient score generator 108 uses variant data 110 to generate and output a predicted patient score 112 for the patient associated with the clinical patient profile 104 and the corresponding genetic testing results 106. Variance data 110 includes reference or real-world labels of variants known / established as pathogenic or benign in the scientific community for a given genetic condition. Each patient score includes the probability that the patient is affected by the genetic condition based on clinical evidence (including unstructured text descriptions) extracted from the clinical patient profile 104.

[0071] Embodiments of the patient score generator 108 include a machine learning model trained to distinguish between affected (positive molecular diagnosis) and unaffected (negative genetic) combinations of clinical characteristics, and to estimate the probability of a patient being affected given only clinical evidence. The patient score generator 108 may include one or more machine learning models. (Reference) Figure 2 An embodiment of a patient score generator 108 comprising two machine learning models is described.

[0072] Patient scores 112 are input into variant score generator 114. Variance score generator 114 uses variant data 110 and combines the relevant patient scores 112 to generate and output variant scores 116. Given patient scores, variant scores 116 estimate the probability that any given variant is the cause of any disease in a gene-related disease.

[0073] Patient scores 112 and / or variant scores 116 can be sent, transmitted, provided to one or more downstream systems, processes, components, models, or frameworks (such as variant classification frameworks or clinical data systems), or otherwise made accessible to one or more downstream systems, processes, components, models, or frameworks. Therefore, the described computational system 100 can generate and output either or both of patient-level predictive scores 112 and variant-level pathogenicity estimates (variation scores) 116. For example, in some applications, patient scores 112 may be useful on their own. Examples of applications in which patient scores 112 can be useful independently of variant scores include the following: 1. Used for medical necessity review (reimbursement) to automatically identify patients who may meet the testing / reimbursement criteria; 2. Determine which condition the patient is most likely to be affected by to determine which genes should be ordered or prioritized in exome / genome testing; 3. Assist variant classification operations by prioritizing patients who should undergo more thorough review by expert scientists and marking patients who do not require any additional review; 4. Used for the discovery of new disease / phenotype associations, for example, clinical variant models can learn that certain phenotypes are associated with pathogenic variants in a given gene that were previously unknown or unproven; 5. Improve the specific diagnostic criteria for association; or 6. To assist in the identification of groups used in clinical trials or other projects.

[0074] Furthermore, although specific implementations of patient score generator 108 and variant score generator 114 have been described, different methods can be used to implement patient score generator 108 or variant score generator 114. For example, in some embodiments, other techniques for calculating patient score 112 or variant score 116 may be used.

[0075] In some specific embodiments, the development and application of clinical variant models for specific genes follow several general steps: (1) First, a patient score is generated to represent the probability that a patient is affected by the molecular condition of interest. This patient score is used in the next step. (2) Second, based on the distribution of patient scores for the variant, a variant score is calculated to represent the probability that the variant is pathogenic. (3) Next, the performance of the CVM (Clinical Variance Modeling) system is determined by using a retained set of data points with known phenotypic-genotypic relationships. (4) A well-performing CVM is calibrated by measuring the positive and negative predictive values ​​(PPV and NPV) from the previous steps, and then the CVM is integrated into a variant classification framework (such as Sherloc) with appropriate weights. (5) Optionally, a subset of variant classification is reviewed by a panel of clinical genomics experts to ensure that the CVM performs as expected.

[0076] In some specific embodiments, the following details describe several aspects of the implementation. In those embodiments, the dataset of interest is defined. For example, Monarch Disease Ontology (MonDO) can be used to discover a set of all disease associations for tested genes. Using the Genome Management Consortium (GenCC) database, each condition is assigned a set of genes and their associated genetic patterns.

[0077] Patient cohorts were identified. For each condition, positive and negative (control) cohorts of patients from cohort 102 were identified as follows. First, affected individuals were identified as patients with a positive molecular diagnosis in at least one of the included genes. For example, autosomal recessive association would require either a pre-existing pathogenic or potentially pathogenic variant, or compound heterozygosity or homozygosity. Control cohorts could be identified in the same manner for all conditions—defined as patients with only benign variants in the included genes.

[0078] For data preprocessing and feature selection, features used for patient scoring modeling can include both structured patient information (e.g., ICD-10 codes, age at admission, clinical domain, etc.) and unstructured textual patient information (e.g., indications for testing, family history, ancestry reported by clinicians). Other sources of patient clinical data can be used additionally or alternatively, including but not limited to PDF (Portable Document Format) files, EMR (Electronic Health Record) and / or clinical notes in media recordings (such as audio or video recordings of clinician notes or patient records), and non-clinical data. ICD-10 codes can be preprocessed by first abbreviating them to the category level, and then selecting the most enriched codes in the positive group based on the chi-square test. For each patient clinical record, strings combining both indications for testing and family history can be created by concatenating these substrings using special markers to define the beginning and end of each text segment.

[0079] The configuration of the patient rating model can follow this sequence: a pre-trained large language model (e.g., UFNLP Gatortron) can be fine-tuned to fit the entire corpus of labeled patients across conditions (TL1). For each condition, TL1 is further fine-tuned to fit examples for that condition (TL2). This model is used to predict the text score for each patient, and then each patient's text score, along with other patient information (e.g., age at test, ICD-10 code), is used as features to train the final model (PS). This model is then used to predict the overall score for each patient, also known as the patient score.

[0080] The variant model utilizes the same variant labels used in genotyping screening for cohorts of genotype-positive and genotype-negative patients and / or other sets of patients (such as potential indicator patients) that provide useful information for interpretation. Furthermore, the set of patients (indicator patients) that provide useful information for interpretation is identified as follows: For each gene included in a condition, the annotated genetic pattern (see the dataset definition above) is used to identify patients who are likely to have the disease if the variant is pathogenic, patients who should not have the related disease if the variant is benign, and patients who will not have the disease that would be explained by another known variant in the included gene. For example, in a monogenic autosomal dominant model, indicator patients would be those who have the variant of interest and no other benign variants in the gene. In a polygenic model, these patients should also have no benign variants in one or more other genes.

[0081] Variance scoring models can be partially pooled, hierarchical Bayesian inference models used to explain gene variants across multiple genes. They estimate parameters such as pathogenicity, penetrance, and gene probability based on input tensors representing genes, patient scores, and variant labels. Variance scores (the probability of a variant being pathogenic) are sampled from the posterior predicted distribution of partially observed Bernoulli variables.

[0082] For validation, the model's generalization performance can be estimated by evaluating it against a retained set of 20% of labeled variants not used for training. Evaluation metrics include the area under the receiver operating characteristic curve (AUROC), mean accuracy (AP), mean squared error (MSE), and classification metrics (including Fl score, accuracy, and PPV and NPV). For high-performance models (AUROC ≥ 0.8) and high-performance genes (AUROC ≥ 0.8), variants with a posterior probability of pathogenicity ≤ 0.05 or ≥ 0.99 and ≥ 2 affected manifestation observations in irrelevant individuals are nominated as evidence via a variant classification framework (e.g., Sherloc point assignment).

[0083] Figure 2 Examples of a process for generating patient scores using a natural language processing-based model, according to some embodiments of the present disclosure, are shown.

[0084] Process 200 is executed by processing logic, which includes hardware (e.g., processing device, circuitry, dedicated logic, programmable logic, microcode, device hardware, integrated circuits, etc.), software (e.g., instructions that run or execute on the processing device), or a combination thereof. In some embodiments, the method is executed by components of a computing system, which in some embodiments includes... Figure 2 Components or processes shown in this figure that may not be specifically shown in other figures and / or components included in some embodiments that may not be shown in other figures. Figure 2 The components or processes are specifically illustrated. Although shown in a particular order or sequence, the order of these processes may be modified unless otherwise stated. Therefore, the illustrated embodiments should be understood as merely examples, and the illustrated processes may be performed in different orders, and some processes may be performed in parallel. Furthermore, at least one process may be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are possible.

[0085] exist Figure 2In this process, process 200 begins with the following steps: For each gene or gene set, two patient groups are identified—1) patients with a confirmed molecular diagnosis and 2) a control group. In machine learning, these patient groups provide reference labels to guide model training. For each group, a first set of clinical features 202 is identified. Features 202 include reference labels and unstructured text, such as natural language descriptions like indications and / or family history. Features 202 are extracted from clinical patient records associated with patients in the group.

[0086] A machine learning model training method known as curriculum learning can be used to train a language model M1 to classify patients based solely on the text descriptions provided in feature 202 (e.g., instances of feature 202 include only text descriptions and reference labels). See below for an example. Figure 3A This section describes examples of curriculum learning in more detail. Generally, curriculum learning can include sorting or grouping instances of feature 202 in a specific way to make model training of language model M1 more efficient. For example, curriculum learning can be used to select instances with features that have more complete textual descriptions to be used as input to model M1 first, so that model M1 is trained on instances with more available unstructured text data before being trained on instances with less available unstructured text data. The described modeling approach is not limited to the use of curriculum learning. Other machine learning training methods can be used, including traditional supervised or semi-supervised training methods.

[0087] Model M1 outputs a text score, which is the likelihood of a patient's genetic condition being influenced by an instance of feature 202. The text score generated by model M1 is combined with other clinical features of the patient (e.g., structured data extracted from the patient's clinical records, excluding unstructured text descriptions) to form a second set of clinical features 204. This second set of clinical features 204 is used to train a second model, M2. Model M2 produces the final probabilities (patient scores). If a reserved set is used, this step can use the same reserved dataset used for model M1.

[0088] Figure 3A An example process 300 for generating text scores using a natural language processing-based model is shown according to some embodiments of the present disclosure.

[0089] exist Figure 3AIn this model, the first model input includes a first set of clinical features 302. Instances of the first clinical features include phenotypic data associated with patient identifiers, disease identifiers, and variant identifiers. More specifically, the phenotypic data includes unstructured text, such as natural language text; for example, textual clinical descriptions, indications for clinical testing, descriptions of disease-related family history, and / or descriptions of demographic data. The first set of clinical features 302 may include a large number of instances of the first clinical features (e.g., hundreds, thousands, or more). For example, the first set of clinical features 302 may include one instance of the first clinical feature for each patient who is a member of a selected group.

[0090] The first sub-model M1 306 receives a first clinical feature set 302 as input. In response to the first clinical feature set 302, the first sub-model M1 306 generates and outputs a text score 314. That is, for each instance of the first clinical feature, the first sub-model M1 306 generates and outputs a corresponding text score 314. The text score 314 includes a machine learning prediction regarding whether the patient associated with the instance of the first clinical feature is affected by the identified disease. The text score 314 is based on unstructured text extracted from the patient's clinical records, excluding other structured data that could also be included in the patient's clinical records regarding whether the patient associated with the instance of the first clinical feature is affected by the identified disease.

[0091] The first sub-model M1 306 includes a machine learning algorithm 308, one or more model parameters (including but not limited to feature weights or coefficients, such as clinical feature weights) 310, and one or more model hyperparameters 312. The first sub-model M1 306, including the machine learning algorithm 308, one or more model parameters 310, and one or more model hyperparameters 312, can be implemented as a language model. In some implementations, a neural network-based machine learning model architecture is used to construct the first sub-model M1 306. In some implementations, the neural network-based machine learning model architecture includes one or more self-focused layers that allow the model to assign different weights to different words or phrases included in the model input (e.g., differentially weighting different parts of unstructured text).

[0092] Alternatively or additionally, the neural network architecture includes feedforward layers and residual connections, which allow the model to learn complex data patterns, including relationships between different words or phrases in multiple different contexts of the model input. In some implementations, a transformer-based architecture including self-focused layers, feedforward layers, and residual connections between these layers is used to construct the first sub-model M1 306. The exact number and arrangement of each type of layer, as well as the hyperparameter values ​​used to configure the model, are determined based on the specific design or implementation requirements of the clinical variation modeling system.

[0093] In some examples, neural network-based machine learning model architectures include or are based on one or more generative transformer models, one or more generative pre-trained transformer (GPT) models, one or more bidirectional encoder representation (BERT) models from transformers, one or more large language models (LLM) and / or one or more other natural language processing (NLP) models.

[0094] The first sub-model M1306 is an NLP-based machine learning model trained on a dataset of natural language text extracted from patients' clinical records. For example, training samples of natural language text extracted from patient forms submitted to a genetic testing service are used to train the first sub-model M1306. The size and composition of the dataset used to train the first sub-model M1306 can vary depending on the specific design or implementation requirements of the clinical variation modeling system. In some implementations, the dataset used to train the first sub-model M1306 includes hundreds of thousands to millions or more diverse natural language text training samples.

[0095] In some implementations, machine learning algorithm 308 may include a logistic function. A logistic function can model the relationship between X (input) and Y (predicted output), where the probability of Y is a linear combination of the independent variables in the input X. Mathematically, the simplified form of a logistic function can be expressed as: , where e is the exponential constant, and and These are the feature coefficients. During the training of the first sub-model M1 306, the values ​​of the coefficients in the linear combination are estimated and iteratively adjusted based on the feature values ​​in the training dataset using logistic regression.

[0096] exist Figure 3AIn this document, a first sub-model M1 306 has been configured via a described supervised machine learning training, calibration, and validation process. During training, the values ​​of one or more parameters in the model parameters (e.g., weights or feature coefficients, such as clinical feature weights) can be iteratively adjusted to test the relative impact of a specific feature input x in the first clinical feature set 302 on a predicted outcome P(Y|X) (e.g., predicted text score 314) based on the value of the feature input x in the first clinical feature set 302. The values ​​of the model parameters (e.g., weights or feature coefficients, such as clinical feature weights) are initialized and adjusted during model training and calibration.

[0097] The first sub-model M1 306 also includes model hyperparameters 312 that are selected or tuned at the global level and are generally not modified based on specific instances of the training data. Examples of model hyperparameters 312 may include learning rate (the rate at which the algorithm updates the estimate), learning rate decay (the learning rate gradually decreases over time to accelerate learning), momentum (the direction of the next step relative to the previous step), regularization constant, the number of branches in the decision tree, and neural network nodes (the number of nodes in each hidden layer of the neural network).

[0098] The first sub-model M1 306 can be configured as either a binary classifier or a scoring model. In binary classification mode, for a given set of input features, the output of the first sub-model M1 306 indicates whether the predicted outcome is pathogenic or benign as a discrete or binary value, for example, 0 indicates benign and 1 indicates pathogenic. In scoring mode, the output of the first sub-model M1 306 includes a score corresponding to a continuous probability (e.g., a floating-point value between 0 and 1, inclusive) of whether the predicted outcome is pathogenic or benign.

[0099] During inference, the first sub-model M1 306 may receive a first set of clinical features 302 including unstructured text and / or variant identifiers that the first sub-model M1 306 has not previously considered as input (i.e., this particular combination of first clinical features that the first sub-model M1 306 has not previously analyzed), and generate a predicted text score 314 for unseen instances of the first clinical features.

[0100] Figure 3B An example process 350 for generating patient scores using a natural language processing-based model is shown according to some embodiments of the present disclosure.

[0101] exist Figure 3BIn this model, the second model input includes a second set of clinical features 352. Instances of the second clinical features include patient scores 364 and corresponding structured phenotypic data associated with patient identifiers, disease identifiers, and variant identifiers. More specifically, the phenotypic data of the second set of clinical features 352 includes structured data, such as ICD codes, demographic data, and other structured data obtained from the patient's clinical records, excluding unstructured phenotypic data (e.g., excluding natural language text descriptions obtained from the patient's clinical records). The second set of clinical features 352 may include many instances of the second clinical features (e.g., hundreds, thousands, or more). For example, the second set of clinical features 352 may include one instance of the second clinical feature for each patient for whom the first sub-model M1 306 generates a text score.

[0102] The second sub-model M2 356 is a machine learning-based scoring model, such as a tree-based model, for example, a boosting tree model. The second sub-model M2 356 receives a second set of clinical features 352 as input. In response to the second set of clinical features 352, the second sub-model M2 356 generates and outputs a patient score 364. That is, for each instance of the second clinical feature, the second sub-model M2 356 generates and outputs the corresponding patient score 364. The patient score 364 includes a machine learning prediction regarding whether the patient associated with the instance of the second clinical feature is affected by the identified disease. The patient score 364 is based on structured data extracted from the patient's clinical records, excluding unstructured text that may also be included in the patient's clinical records. For example, the patient score 364 is calculated based on a text score 314 and other structured data, excluding the text description used by the first sub-model M1 306 to generate the text score 314.

[0103] The second sub-model M2 356 includes a machine learning algorithm 358, one or more model parameters (including but not limited to feature coefficients or weights) 360, and one or more model hyperparameters 362. The second sub-model M2 356, including the machine learning algorithm 358, one or more feature coefficients (or weights) 360, and one or more model hyperparameters 362, can be implemented using a neural network-based machine learning model architecture. For example, the second sub-model M2 356 can be implemented using a boosting tree model, a regression model, or another type of machine learning model configured as a classification model or a scoring model.

[0104] In some implementations, a transformer-based architecture including self-focus layers, feedforward layers, and residual connections between these layers is used to construct the second sub-model M2 356. The exact number and arrangement of each type of layer, as well as the hyperparameter values ​​used to configure the model, are determined based on the specific design or implementation requirements of the clinical variation modeling system. The second sub-model M2 356 is trained on a dataset of the second clinical feature 352. For example, the second sub-model M2 356 is trained using structured data extracted from patient forms submitted to a genetic testing service, along with text scores 314. The size and composition of the dataset used to train the second sub-model M2 356 can vary depending on the specific design or implementation requirements of the clinical variation modeling system. In some implementations, the dataset used to train the second sub-model M2 356 comprises hundreds of thousands to millions or more distinct training samples, where the size of the dataset corresponds to the size of the first clinical feature set 302. In other words, for each instance of the first clinical feature 302, there exists a corresponding instance of the second clinical feature 352.

[0105] In some implementations, the machine learning algorithm 358 may include one or more of a logistic function, a decision tree, or a boosting function (e.g., gradient boosting). A logistic function can model the relationship between X (input) and Y (predicted output), where the probability of Y is a linear combination of the independent variables in the input X. Mathematically, the simplified form of a logistic function can be represented as... , where e is the exponential constant, and and These are the feature coefficients. During the training of the second sub-model M2 356, logistic regression estimates and iteratively adjusts the values ​​of the coefficients in the linear combination based on the feature values ​​in the training dataset.

[0106] exist Figure 3B In this document, a second sub-model M2 356 has been configured via a described supervised machine learning training, calibration, and validation process. During training, the values ​​of the feature coefficients can be iteratively adjusted to test the relative impact of a specific feature input x in feature set 352 on the prediction outcome P(Y|X) (e.g., predicting patient score 364) based on the values ​​of the feature input x in feature set 352. The values ​​of the feature coefficients are initialized and adjusted during model training and calibration.

[0107] The second sub-model M2 356 also includes model hyperparameters 362 that are selected or tuned at the global level and are generally not modified based on specific instances of the training data. Examples of model hyperparameters 362 may include learning rate (the rate at which the algorithm updates the estimate), learning rate decay (the learning rate gradually decreases over time to accelerate learning), momentum (the direction of the next step relative to the previous step), regularization constant, the number of branches in the decision tree, and neural network nodes (the number of nodes in each hidden layer of the neural network).

[0108] The second sub-model M2 356 can be configured as a binary classifier or a scoring model. In binary classification mode, for a given set of input features, the output of the second sub-model M2 356 indicates whether the predicted outcome is pathogenic or benign as a discrete or binary value; for example, 0 indicates benign and 1 indicates pathogenic. In scoring mode, the output of the second sub-model M2 356 includes a score corresponding to a continuous probability (e.g., a floating-point value between 0 and 1, inclusive) of whether the predicted outcome is pathogenic or benign.

[0109] During inference, the second sub-model M2 356 may receive a set 352 of second clinical features, including structured clinical data that the second sub-model M2 356 has not previously considered as input, text scores 314, and / or variant identifiers (i.e., this particular combination of second clinical features that the second sub-model M2 356 has not previously analyzed), and generate a predicted patient score 364 for unseen instances of the second clinical features.

[0110] Figure 4 Example procedures for configuring clinical variation models using machine learning are shown according to some embodiments of this disclosure.

[0111] exist Figure 4 In the model training process 400, machine learning algorithms, such as supervised machine learning, are applied to the labeled feature set 402. For example, to train the first sub-model M1 306, the labeled feature set 402 includes instances of the first clinical feature 302, and for each instance of the clinical feature 302, a real or reference label indicating whether the associated variant is known benign, known pathogenic, or a variant of known indeterminate significance (VUS). As another example, to train the second sub-model M2 356, the labeled feature set 402 includes instances of the second clinical feature 352, and for each instance of the clinical feature 352, a corresponding real or reference label indicating whether the associated variant is known benign, known pathogenic, or a variant of known indeterminate significance (VUS).

[0112] In some implementations, the model training process 400 implements one or more aspects of a curriculum-based machine learning model training technique. In curriculum-based learning, the machine learning model is trained in a meaningful order, for example, from easy training samples to difficult samples. For example, curriculum-based learning can be implemented by a sorter / scheduler component 404 to group, sort, or rank training instances according to some objective measure of difficulty, and then feed the training instances to the model in a predetermined order (e.g., such that the model receives easier training instances before it receives more difficult ones). In the case of training the first sub-model M1306, the sorter / scheduler 404 can, for example, group instances of the first clinical feature 302 based on the amount of unstructured text contained in the patient's clinical record, such that the model receives training instances with more unstructured text before training instances with little or no unstructured text. In the case of training the second sub-model M1356, the sorter / scheduler 404 can, for example, classify the instances of the second clinical feature 352 based on the value of the text score 314, such that the second sub-model M1356 receives training instances with higher text scores before training instances with lower text scores.

[0113] The sorter / scheduler 404 may additionally or alternatively implement progress control functionality, which can control how many easy training examples are fed into the model before more difficult training examples are fed into the model, or how quickly the model is introduced into more difficult training examples.

[0114] In addition to organizing training examples into the model and controlling the input progress, or alternatively, course learning can also be used to incrementally adjust the model architecture or one or more model parameters; for example, gradually increasing the model capacity (adding more neural units) or gradually increasing the complexity of the model's task.

[0115] When the amount of available data varies across different training instances, the use of curriculum learning can be particularly helpful in improving the efficiency of the model training process and enhancing the predictive reliability of the model. In particular, curriculum learning can help address the problem of using patient records with sparse, blank, or incomplete textual descriptions for clinical variation modeling.

[0116] In the first training iteration, in subprocess 406, feature coefficients (or weights) are initialized or assigned to each feature input of a given input feature. For example, feature coefficients are initialized by randomly setting their values. The feature coefficient values ​​assigned in subprocess 406 are operated as weights applied to the corresponding features to produce weighted features 410. In subprocess 412, a machine learning algorithm is applied to the weighted features to produce a predicted output 414. Reference or ground truth labels included in the training instances provide the expected output 408 for supervised machine learning. In subprocess 416, the predicted output 414 is evaluated by calculating a loss (or error) based on the expected output 408 and the predicted output 414. The loss is calculated using a loss function such as gradient descent.

[0117] In each iteration, decision subprocess 418 evaluates the difference between the predicted and expected outputs, for example, by comparing the output of the loss function with a stopping condition that depends on the change in the output of the loss function. The change in the output of the loss function is compared with an error tolerance threshold (which may be referred to herein as the model performance criterion or model convergence criterion). If the model performance criterion (e.g., the error tolerance threshold) is not met (e.g., not greater than or equal to a threshold performance level, or exceeding the maximum permissible error value, or the loss has not stopped improving beyond the tolerance, or the error has not stopped decreasing beyond the tolerance threshold), model training continues for another iteration. If the loss has stopped improving beyond the tolerance threshold, the model has converged and training has ended.

[0118] In subsequent training iterations, the values ​​of one or more feature coefficients are adjusted, the machine learning algorithm is applied to additional instances of the feature set, and the output of the machine learning algorithm is evaluated using the loss function and error tolerance as described above. The training process ends when the model performance criteria (i.e., the error tolerance threshold) are met and the model converges (e.g., the output of the comparison (e.g., the loss function) is greater than or equal to the threshold performance level or within or below the maximum permissible error value, or the loss has stopped improving beyond the tolerance threshold).

[0119] Figure 5A Examples of patient score calculations according to some embodiments of this disclosure are shown. Figure 5A In the context of patient 1, the clinical record associated with the NF1 gene includes structured data (e.g., age, sex, ICD code) and unstructured data (e.g., indications and family history). Portions of this structured and unstructured data are extracted from the clinical record and fed into a patient scoring generator (e.g., a combination of models M1 and M2 described above). The patient scoring generator generates a patient score based on the extracted structured and unstructured data. Figure 5A In the example, the patient score indicates a high likelihood based on clinical data and the influence of the patient's genetic status.

[0120] Figure 5B Examples of patient score calculations according to some embodiments of this disclosure are shown.

[0121] Figure 5B Examples of patient score calculations according to some embodiments of this disclosure are shown. Figure 5B In the clinical record associated with patient 2, specifically for gene NF1, there are structured data (e.g., age, sex, ICD code) and unstructured data (e.g., indications and family history). Portions of these structured and unstructured data are extracted from the clinical record and fed into a patient scoring generator (e.g., a combination of models M1 and M2 described above). The patient scoring generator generates a patient score based on the extracted structured and unstructured data. Figure 5B In the example, the patient score indicator is based on clinical data and the low likelihood that the patient is affected by genetic conditions.

[0122] Figure 5C Examples of patient score calculations using machine learning feature weights are shown according to some embodiments of this disclosure.

[0123] To generate patient scores: (A) First, for a given molecular disease, which may be defined by a single gene (e.g., neurofibromatosis type I) or by multiple genes (e.g., Lynch syndrome), clinically relevant patient information is collected for all patients with a molecular diagnosis of the condition and for all patients with only benign variants in the gene of interest and no other molecular diagnoses (i.e., the genotype-negative group). The ML model learns appropriate evidence weights for segments of clinical information that distinguish individuals with molecular diagnoses from genotype-negative individuals. (B) Next, the evidence strength learned for the clinical symptoms observed in the group is applied to each individual in the entire group to generate a patient score for each individual.

[0124] Figure 6A Example 600 illustrates the calculation of a variation score using a Bayesian model according to some embodiments of this disclosure.

[0125] To generate variant scores, the probability of a given variant being pathogenic is derived from a set of patient scores associated with the variant. Molecular diagnostic results are used to identify relevant patient groups for any given variant. This is done for both known pathogenic and known benign variants. Figure 6AIn the example, each stack of person icons represents a single variant, where each icon in the stack represents a patient score of a patient with that variant. Using this distribution, the probability of a set of observations of variants resulting from pathogenic variants can be inferred (e.g., through machine learning-based inference or statistical correlation). The distribution can be generated using a Bayesian inference model that considers several clinical genetic phenomena, including incomplete penetrance, age of onset, and pseudophenotype. This model construction naturally represents the inherent uncertainty with limited observations, and this extends to its estimation of epigenetic penetrance (the basic signal the model uses to classify variants). The posterior distribution of the model can be sampled to estimate the probability that a particular variant is pathogenic.

[0126] Figure 6B Example 610 illustrates the distribution of patient scores for variants according to some embodiments of this disclosure.

[0127] The columns of the boxes represent the distribution of patient scores for genetic variants (e.g., MSH2). Each box in the column represents a patient score generated using clinical patient records. Figure 6B In the example, the first patient had a patient score of 0.76, and the second patient had a patient score of 0.04. The higher patient score of 0.76 was associated with a higher likelihood of the first patient being influenced by genetic status, while the lower patient score of 0.04 was associated with a lower likelihood of the second patient being influenced by genetic status. Figure 6B The example shows that even if unstructured text information is missing (the indicator field is blank), a patient score can still be generated for a second patient.

[0128] For the generation of variation scores, it has An example distribution of patient scores for 11 patients is shown. Each box in the figure represents a patient. Patients with low patient scores are shown in white, and patients with high patient scores are shown in black. The left side of the figure shows example patients with high patient scores based on clinical evidence, and the right side shows example patients with low patient scores based on clinical evidence. Figure 6B The example shows that even if part of the unstructured text information is missing (the family history field is blank), a patient score can still be generated for the second patient.

[0129] Figure 6D Example 630 illustrates the distribution of variant scores and patient scores for classified variants according to some embodiments of the present disclosure.

[0130] exist Figure 6D The image shows an example distribution of patient scores for known benign and pathogenic variants in MSH2. A second ML model calculates a variant score for each variant based on this distribution. Figure 6DIn the example, all known benign variants (which are predominantly found in individuals with low patient scores) have low variant scores, while all known pathogenic variants on the right (which are more enriched in patients with high patient scores) have high variant scores.

[0131] Figure 6E Example 640 illustrates the distribution of variant scores and patient scores for classified variants according to some embodiments of this disclosure. Figure 6E In this study, variant scores are generated for VUS in MSH2, and these variant scores can be compared with those of known benign and known pathogenic variants.

[0132] Figure 6F Examples of the distribution of patient scores and variant scores according to some embodiments of this disclosure are shown. Figure 6F The image shows an example of CVM results for NF1. In this cell plot, each stack of boxes represents a single gene variant, and each box represents a single patient. Each box is shaded according to the individual's patient score (the inferred probability that the patient is affected by the condition). Patient score box 652 along the upper y-axis on the left side of the plot represents a low patient score; while patient score box 654 along the upper y-axis on the right side of the plot represents a high patient score. In the bar below the patient score plot, each box is shaded according to the variant score derived from the stack of observations from the patient scores of that variant. The box in region 656 indicates a low variant score; while the box in region 658 indicates a high variant score. Note that the y-axis is magnified to visualize the patient boxes on the upper right. The actual stack of patient boxes on the upper left expands much higher than shown because these are higher frequency variants.

[0133] Figure 6G Example 660 illustrates the distribution of classified variants and patient scores according to some embodiments of this disclosure. Figure 6G This illustrates how, using the described method, machine learning can be used to predict the distribution of patient scores for pathogenic and benign variants in MMR genes. Using this predicted distribution, a quantitative determination can be made as to whether the distribution of patient scores for VUS is similar to that for other pathogenic or benign variants. For example, suppose variants 680 and 682 are VUS, rather than known benign or known pathogenic variants. Given the distribution of patient scores for each of these variants 680 and 682, they can be appropriately placed into the distributions for known benign and known pathogenic variants based on their similarity to those for known benign and known pathogenic variants.

[0134] Figure 7A Example 700 illustrates the distribution of variant scores and patient scores for classified variants according to some embodiments of this disclosure. Figure 7AIn the study, patient scores for samples of 20 benign variants and 20 pathogenic variants were plotted in a vertical stack along the y-axis, and variant scores derived from the Bayesian NF1 clinical variant model were plotted along the x-axis. Figure 7A The clinical variation model demonstrates how these variations are correctly classified. Figure 7A The right side shows data for samples of VUS in the same gene, demonstrating that the model performance superficially extends to these variants.

[0135] Figure 7B Example 710 illustrates the distribution of variant scores and patient scores for classified variants according to some embodiments of this disclosure. Figure 7B In this study, a sample of 40 VUS variants was used to suggest that the patient scoring model be extended to these unlabeled variants. The variant score reflects the pathogenicity prediction of variants on the right, the benign prediction of variants on the left, and the indeterminate score of variants in the absence of sufficient observations. Figure 7B It also shows that benign manifestation variants are observed at a much higher frequency than pathogenic manifestation variants, implying that these benign nominations should have a greater impact on the number of VUS reports.

[0136] Figure 8A Example 800 illustrates performance data of a patient score generator and a variant score generator according to some embodiments of the present disclosure.

[0137] Figure 8A The data generated by running the described CVM pipeline on a set of over 2300 disease-associated genes is plotted. For each gene-disease-specific model, performance in both steps of the pipeline is shown. Only those achieving high performance on the validation set were selected, resulting in 920 models comprising nearly 2000 genes. These models cover every clinical domain from cardiology to cancer. Figure 8A As shown, the CVM machine learning pipeline can use previously underutilized clinical evidence to predict the pathogenicity of variants. The patient scoring model estimates the probability that a patient is affected by a relevant genetic condition. The variant scoring model combines relevant patient observations to predict whether the variant is the cause of the disease. In some embodiments, there is a minimum number of patient observations required to establish a variant that meets the criteria of a clinical variant model. For example, in some implementations, a clinical variant model as described may require a minimum of three observations from at least two independent families. However, in practice, for many genes, more patient observations may be requested to ensure that the model reaches the confidence threshold required to provide evidence to a variant classification framework such as Sherloc.

[0138] In the experimental results, the described method has demonstrated high performance (AUROC ≥ 0.8) in distinguishing between pathogenic and benign variants across a large number of conditions (n=920) and genes (n=1977) across clinical domains. This model has the potential to provide both pathogenic and benign evidence for a large number of VUS cases and significantly reduces VUS reporting where the model is applied.

[0139] Figure 8B Examples of clinical variation modeling and classification systems according to some embodiments of this disclosure are shown.

[0140] The development and application of clinical variant models for each gene follow several general steps: (1) First, a patient score is generated to represent the probability that a patient is affected by the molecular condition of interest. This patient score is used in the next step. (2) Second, based on the distribution of patient scores for the variant, a variant score is calculated to represent the probability that the variant is pathogenic. (3) Next, the performance of the CVM is determined by using a retained set of data points with known phenotypic-genotypic relationships. (4) Then, the well-performing model is calibrated by measuring the positive and negative predictive values ​​(PPV and NPV) from the previous steps, and then the model is integrated into a variant classification framework (such as Sherloc) with appropriate weights. (5) Optionally, a subset of variant classification is reviewed by a panel of clinical genomics experts to ensure that the CVM performs as expected. Each of these general steps is explained in more detail below.

[0141] Step 1: Patient Score Generation CVM uses clinical data to predict the pathogenicity of variants for a given genetic condition in a stepwise manner. By leveraging details found in clinician reports from test requests (e.g., personal health history; family health history; age; sex; patient's race, ethnicity, and ancestry; ICD-10 code; and the clinical domain of the ordered test), the model is first trained to learn to distinguish between patients with a positive molecular diagnosis for the condition of interest and genotype-negative controls (i.e., patients without VUS, LP, or P variants in the condition of interest and without a molecular diagnosis in another condition). For each patient, a patient score (e.g., on a scale from 0 to 1) is generated, representing the probability that a given patient is affected by the genetic condition of interest based solely on clinical information. Based on what has been learned, this information can be applied to other patients with VUS in the condition of interest, provided those patients do not have a current molecular diagnosis in another gene. This is accomplished by scoring how similar the clinical profiles of other patients appear to positive or negative cases.

[0142] Step 2: Generate Variance Score Using a set of known pathogenic and benign variants for each condition, called labels, the second model learns the distribution of typical patient scores for pathogenic and benign variants. Based on what has been learned, VUS scores can be assigned based on how similar the distribution of patient scores for VUS appears to pathogenic and benign variants. A variant score (e.g., on a scale from 0 to 1) is generated for each variant, which is the probability that a given variant is pathogenic. The higher the variant score, the higher the probability that the variant is pathogenic.

[0143] Step 3: Performance Evaluation of CVM Each model was tested against a set of previously unseen known pathogenic and benign variants (i.e., the 20% retainer set), and all CVMs were carefully screened to ensure high performance in distinguishing between pathogenic and benign variants (e.g., AUROC ≥ 0.8).

[0144] Step 4: Model Calibration The models that pass the high-performance threshold screening are then calibrated. In some embodiments, the calibrated models are evaluated by clinical genomics experts prior to implementation. To date, clinical variant modeling has demonstrated high accuracy for over 600 genes and conditions.

[0145] Step 5: Review by clinical genomics experts Optionally, to further obtain confidence in the CVM's predictive output, a subset of variants may be selected for thorough review by expert clinical genomics scientists. This includes variants with strong pathogenicity (≥99% PPV) CVM predictions and variants with strong benign (≥95% NPV) CVM predictions (e.g., 163 / 1052). For each variant, all currently available non-CVM evidence was reviewed to assess any conflicting data. In this review, experts confirmed 93% (13 / 14) of variants predicted as pathogenic and 95% (155 / 163) of variants predicted as benign, while the remaining variants were kept as VUS. Note that experts may select the most challenging variants with benign CVM predictions for review (e.g., variants with some level of pathogenicity evidence or variants in the PMS2 pseudogenome region with benign CVM predictions). Furthermore, in some embodiments, four of the predicted pathogenic variants have at least one CLINVAR entry that is either likely or pathogenic, while a fifth variant predicted as pathogenic by the CVM is recently reclassified from VUS as likely pathogenic based on new family isolation data obtained after the CVM prediction was generated but before the review results. Similarly, 15 of the predicted benign variants under review have at least one CLINVAR entry that was labeled as likely benign or benign by another submitter.

[0146] Orthogonal verification.

[0147] To further assess the confidence level of the CVM's predictive output, a consistency analysis was performed, comparing the CVM with an evolutionary model of variant effects (EVE). The deep learning model uses orthogonal data (i.e., sequence conservation) to predict the pathogenicity of variants in humans2. The CVM predictions were highly consistent with the EVE predictions (90.5%).

[0148] To add confidence to the predicted outputs of the CVM for Lynch syndrome genes (MLH1, MSH2, MSH6, PMS2, EPCAM), the CVM model predictions were compared with functional datasets (e.g., multiple tests of variant effects or MAVE) for MLH1, MSH2, and PMS2. Variants with both CVM model predictions and MAVE predictions showed high consistency (>98%).

[0149] Another mechanism for orthogonal validation of the CVM can be performed for the TSC2 gene, which is associated with tuberous sclerosis. Specifically, exons 26 and 32 of TSC2 are absent in all known clinically relevant transcripts of gene 4 (note the legacy exon nomenclature). Using only clinical evidence, the CVM for TSC2 identified pathogenic gene variants in every one of the 42 exons of the gene except for exons 26 and 32—a significant concordance between clinical models and independent molecular data.

[0150] Figure 8C Examples of variant classification systems including clinical variant modeling are shown according to some embodiments of the present disclosure.

[0151] To integrate predictions from CVM into variant classification frameworks such as Sherloc, a two-tiered prediction system can be established based on prediction performance thresholds, such as negative and positive predictive values ​​(NPV and PPV). The benign tier is defined as strong benign evidence, which is sufficient evidence to classify a variant as possibly benign in the absence of contrary evidence. The pathogenicity tier is insufficient on its own to achieve a possibly pathogenic classification, but can achieve a definitive classification in the presence of a variant absent or within the expected pathogenicity range in gnomAD or another piece of pathogenic evidence. The prediction performance thresholds for those tiers are defined as: (1) strong benign ≥ 95% NPV, (2) strong pathogenicity ≥ 99% PPV. The third and final tiers correspond to predictions between 99% PPV and below 95% NPV, which, for the initial release of CVM, were considered insufficiently certain to be assigned weight within the Sherloc scoring system.

[0152] Clinical variant modeling is incorporated into variant classification frameworks such as Sherloc. In one example, CVM predictions are integrated into Sherloc using a two-level prediction interval. Variants with CVM predictions of ≥95% NPV are assigned 3 benign points, and variants with CVM predictions of ≥99% PPV are assigned 3 pathogenic points. This clinical evidence is considered in the context of all variant classification evidence in Sherloc. Typical cutoff values ​​for variant classification are shown in the lower right corner: 5 benign points correspond to benign, 3 benign points correspond to possibly benign, 4 pathogenic points correspond to possibly pathogenic, and 5 pathogenic points correspond to pathogenicity. Clinical genomics experts can overturn these classification thresholds when necessary due to conflicting evidence and other factors.

[0153] Generally, the number of points associated with a particular score is related to the confidence or certainty of the prediction. The number of intervals does not have to be two; any number of intervals can be used (including no intervals or an unlimited number of intervals). Those points mapped to the score in the model output can be incorporated into another variant classification framework such as Sherloc. In this way, the output of a clinical variant model as described can be applied to the Sherloc framework or to any other variant labeling or classification framework.

[0154] Figure 9A Methods for clinical variation modeling according to some embodiments of the present disclosure are shown.

[0155] Method 900 is executed by processing logic, which includes hardware (e.g., processing device, circuitry, special-purpose logic, programmable logic, microcode, device hardware, integrated circuits, etc.), software (e.g., instructions that run or execute on the processing device), or a combination thereof. In some embodiments, the method is executed by a component of a computing system, which in some embodiments includes... Figure 9A Components or processes shown in this figure that may not be specifically shown in other figures and / or components included in some embodiments that may not be shown in other figures. Figure 9A The components or processes are specifically illustrated. Although shown in a particular order or sequence, the order of these processes may be modified unless otherwise stated. Therefore, the illustrated embodiments should be understood as merely examples, and the illustrated processes may be performed in different orders, and some processes may be performed in parallel. Furthermore, at least one process may be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are possible.

[0156] In step 902, the processing device extracts phenotypic data from clinical records associated with patient populations and genetic conditions, wherein the phenotypic data includes natural language text, which includes at least one of the indicators for testing, family history descriptions, or demographic data.

[0157] In step 904, the processing device extracts genotype data from gene test results associated with a genetic condition, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnosis associated with the at least one variant.

[0158] In step 906, the processing apparatus uses phenotypic data and genotypic data, including natural language text, to configure a first model including a machine learning model based on natural language processing (NLP) to generate and output patient scores for a patient population, wherein the patient scores include machine learning likelihood of the patient’s genetic status as given phenotypic data including natural language text.

[0159] In some implementations, configuring a Natural Language Processing (NLP)-based machine learning model includes: configuring a first model to receive natural language text and generate and output text scores, wherein the text scores include a machine learning likelihood of the patient's genetic condition given the received natural language text; and configuring a second model to generate and output patient scores using the text scores and second model inputs, the second model inputs including phenotypic data excluding the natural language text. In some implementations, the processing apparatus uses curriculum learning to configure the NLP-based machine learning model to generate and output text scores, and then trains the second model to generate and output patient scores using either curriculum learning or another training method. Curriculum learning includes using the text scores to determine the order of inputs for the second model, and applying the second model to the second model inputs according to that order. In some implementations, the processing apparatus iteratively adjusts the values ​​of at least one clinical feature weights of the NLP-based machine learning model until the output of a loss function computed based on the patient scores output by the NLP-based machine learning model satisfies at least one first performance criterion, to produce a configured NLP-based machine learning model, wherein the configured NLP-based machine learning model is capable of outputting patient scores that satisfy at least one second performance criterion. In some implementations, the processing device uses positive and negative examples to configure NLP-based machine learning models, where positive examples include positive associations between molecular diagnoses and genetic conditions, and negative examples include negative associations between molecular diagnoses and genetic conditions.

[0160] In step 908, the processing device uses patient scores and genotype data to configure a Bayesian causal model to generate and output a variant score, wherein the variant score includes a machine learning probability that the variant is pathogenic with respect to the genetic condition. In some implementations, the Bayesian causal model includes probability distributions of patient scores and variant scores.

[0161] In some implementations, the processing device uses a configured NLP-based machine learning model and a configured Bayesian causal model to generate and output predictions about whether variants associated with unseen clinical records are benign or pathogenic in relation to a genetic condition. In some implementations, the processing device provides clinicians with predictions about whether a variant is benign or pathogenic for use in developing a patient diagnosis. In some implementations, the processing device provides predictions to at least one system, process, model, or component of a variant classification system or clinical data system. In some implementations, the processing device uses variant scores output by a configured Bayesian causal model as input to a variant classification framework.

[0162] Figure 9B Methods for clinical variation modeling according to some embodiments of the present disclosure are shown.

[0163] Method 920 is executed by processing logic, which includes hardware (e.g., processing device, circuitry, dedicated logic, programmable logic, microcode, device hardware, integrated circuits, etc.), software (e.g., instructions that run or execute on the processing device), or a combination thereof. In some embodiments, the method is executed by a component of a computing system, which in some embodiments includes... Figure 9B Components or processes shown in this figure that may not be specifically shown in other figures and / or components included in some embodiments that may not be shown in other figures. Figure 9B The components or processes are specifically illustrated. Although shown in a particular order or sequence, the order of these processes may be modified unless otherwise stated. Therefore, the illustrated embodiments should be understood as merely examples, and the illustrated processes may be performed in different orders, and some processes may be performed in parallel. Furthermore, at least one process may be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are possible.

[0164] In step 922, the processing device extracts phenotypic data from clinical records associated with the patient population and genetic status, wherein the phenotypic data includes natural language text, which includes at least one of the indicators for testing, family history descriptions, or demographic data. In step 924, the processing device uses phenotypic data including natural language text to generate and output patient scores for a patient population, wherein the patient scores include the likelihood of a patient being affected by genetic status given phenotypic data including natural language text.

[0165] In some implementations, the processing apparatus configures a Natural Language Processing (NLP)-based machine learning model to generate and output patient scores by: (i) configuring a first model to generate and output text scores, wherein the text scores include machine learning likelihoods of the patient's genetic condition given a first model input comprising natural language text; and (ii) configuring a second model to generate and output patient scores using second model inputs, the second model inputs including text scores and phenotypic data excluding natural language text. In some implementations, the processing apparatus uses curriculum learning to configure the NLP-based machine learning model to generate and output patient scores, wherein curriculum learning includes using the text scores to determine the order for use as input to the second model, and applying the second model to the second model inputs according to that order. In some implementations, the processing device configures a Natural Language Processing (NLP)-based machine learning model to generate and output patient scores; and iteratively adjusts the values ​​of at least one clinical feature weights of the NLP-based machine learning model until the output of a loss function calculated based on the patient scores output by the NLP-based machine learning model satisfies at least one first performance criterion, to produce a configured NLP-based machine learning model, wherein the configured NLP-based machine learning model is capable of outputting patient scores that satisfy at least one second performance criterion. In some implementations, the processing device uses positive and negative examples to configure the NLP-based machine learning model, wherein positive examples include positive associations between molecular diagnoses and genetic conditions, and negative examples include negative associations between molecular diagnoses and genetic conditions.

[0166] In step 926, the processing device extracts genotype data from gene testing results associated with the patient population and genetic condition, wherein the genotype data includes molecular diagnostics, which include at least one variant of a gene. In step 928, the processing device uses the patient score and genotype data to configure a Bayesian causal model to generate and output a variant score, wherein the variant score includes a machine learning likelihood that the variant is pathogenic to the genetic condition. In some implementations, the Bayesian causal model includes probability distributions of the patient score and the variant score.

[0167] In some implementations, the processing device uses a configured NLP-based machine learning model and a configured Bayesian causal model to generate and output predictions about whether variants associated with unseen clinical records are benign or pathogenic in relation to a genetic condition. In some implementations, the processing device provides clinicians with predictions about whether a variant is benign or pathogenic for use in developing a patient diagnosis. In some implementations, the processing device provides predictions to at least one system, process, model, or component of a variant classification system or clinical data system. In some implementations, the processing device uses variant scores output by a configured Bayesian causal model as input to a variant classification framework.

[0168] Figure 9C Methods for clinical variation modeling according to some embodiments of the present disclosure are shown.

[0169] Method 930 is executed by processing logic, which includes hardware (e.g., processing device, circuitry, special-purpose logic, programmable logic, microcode, device hardware, integrated circuits, etc.), software (e.g., instructions that run or execute on the processing device), or a combination thereof. In some embodiments, the method is executed by a component of a computing system, which in some embodiments includes... Figure 9C Components or processes shown in this figure that may not be specifically shown in other figures and / or components included in some embodiments that may not be shown in other figures. Figure 9C The components or processes are specifically illustrated. Although shown in a particular order or sequence, the order of these processes may be modified unless otherwise stated. Therefore, the illustrated embodiments should be understood as merely examples, and the illustrated processes may be performed in different orders, and some processes may be performed in parallel. Furthermore, at least one process may be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are possible.

[0170] In step 932, the processing device extracts phenotypic data from clinical records associated with the patient population and genetic status, wherein the phenotypic data includes natural language text, which includes at least one of the indicators for testing, family history descriptions, or demographic data.

[0171] In step 934, the processing device uses phenotypic data including natural language text to configure a natural language processing (NLP)-based machine learning model to generate and output patient scores for a patient population, wherein the patient scores include machine learning likelihood of the patient’s genetic status given phenotypic data including natural language text.

[0172] In some implementations, configuring a Natural Language Processing (NLP)-based machine learning model includes: configuring a first model to receive natural language text and generate and output text scores, wherein the text scores include a machine learning likelihood of the patient's genetic status given the received natural language text; and configuring a second model to generate and output patient scores using the text scores and second model inputs, the second model inputs including phenotypic data excluding the natural language text. In some implementations, the processing apparatus uses curriculum learning to configure the NLP-based machine learning model to generate and output text scores, wherein curriculum learning includes using the text scores to determine an order for the second model inputs and applying the second model to the second model inputs according to that order. In some implementations, the processing apparatus iteratively adjusts the values ​​of at least one clinical feature weights of the NLP-based machine learning model until the output of a loss function computed based on the patient scores output by the NLP-based machine learning model satisfies at least one first performance criterion, to produce a configured NLP-based machine learning model, wherein the configured NLP-based machine learning model is capable of outputting patient scores that satisfy at least one second performance criterion. In some implementations, the processing device uses positive and negative examples to configure NLP-based machine learning models, where positive examples include positive associations between molecular diagnoses and genetic conditions, and negative examples include negative associations between molecular diagnoses and genetic conditions.

[0173] In step 936, the processing device extracts genotype data from gene test results associated with a genetic condition, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnosis associated with the at least one variant.

[0174] In step 938, the processing device uses patient scores and genotype data to generate and output variant scores, wherein the variant scores include a machine learning likelihood that the variant is pathogenic to the genetic condition. In some implementations, the processing device configures a Bayesian causal model to generate and output variant scores. In some implementations, the Bayesian causal model includes probability distributions of patient scores and variant scores.

[0175] In some implementations, the processing device uses a configured NLP-based machine learning model and a configured Bayesian causal model to generate and output predictions about whether variants associated with unseen clinical records are benign or pathogenic in relation to a genetic condition. In some implementations, the processing device provides clinicians with predictions about whether a variant is benign or pathogenic for use in developing a patient diagnosis. In some implementations, the processing device provides predictions to at least one system, process, model, or component of a variant classification system or clinical data system. In some implementations, the processing device uses variant scores output by a configured Bayesian causal model as input to a variant classification framework.

[0176] Figure 10 An example computational system including clinical variation modeling is shown according to some embodiments of the present disclosure.

[0177] exist Figure 10 In this embodiment, the computing system 1000 includes a data selection subsystem 1002, a feature generation subsystem 1012, a training and calibration subsystem 1020, a model validation subsystem 1030, and one or more scoring models 1040. Data sources accessible and usable by the components of the computing system 1000 include data storage for model performance criteria 1032, one or more training datasets 1034, model validation criteria 1036, and one or more validation datasets 1038.

[0178] Any one or more components of the computing system 1000 may correspond to components and / or processes similarly described in other figures and / or herein. For example, one or more rating models 1040 may correspond to a single natural language processing (NLP) based model (e.g., M1 or M2 described above) or a combination of NLP-based models (e.g., a modeling pipeline including M1 and M2), or a Bayesian causal model, or a combination of one or more NLP-based models and a Bayesian causal model. In other words, the rating model(s) 1040 may include a patient rating generator, a variant rating generator, or both a patient rating generator and a variant rating generator.

[0179] For example, one or more rating models 1040 may include one or more machine learning models configured to use machine learning algorithms to determine a probabilistic or statistical relationship between the inputs and the outputs. For example, given one or more inputs, rating model 1040 may output labels that can be used to classify the inputs into different categories, or output ratings that can be used to sort or rank the inputs into grouped or graded lists.

[0180] Data selection components 1004, 1006, and 1008 may correspond to one or more of the data selection components, techniques, or processes described herein. Feature generation subsystem 1012 may correspond to one or more of the feature generation components, techniques, or processes described herein. Depending on the circumstances, model performance criteria 1032, training dataset(s) 1034, model validation criteria 1036, and validation dataset(s) 1038 may include data described herein as performance criteria, model validation data, training data, or validation data. Model validation subsystem 1030 may correspond to one or more of the model validation components, techniques, or processes described herein.

[0181] The training and calibration subsystem 1020 includes a dataset selection component 1022, a model training component 1024, and a model calibration component 1026. The model training component 1024 and / or the model calibration component 1026 of the training and calibration subsystem 1020 can train one or more machine learning models of the scoring model 1040 by, for example, applying supervised machine learning techniques to training data including input data and training examples with reference (or real) labels. The predicted outputs of one or more machine learning models are iteratively observed until a set of model performance criteria is satisfied. For example, a loss function is used to quantify the difference between the predicted output and the expected output. Model performance criteria are used to determine when one or more machine learning models have converged and are able to provide a reliable output with the desired degree of determinism. The necessary level of determinism and performance criteria are determined based on the requirements or design of the specific implementation of one or more machine learning models.

[0182] In some implementations, training datasets(s) 1034 include training data used to train one or more scoring models 1040. Training datasets(s) 1034 include, for example, sets of input features and corresponding reference (or true) labels. In some implementations, training datasets(s) 1034 may include, or be derived from, one or more databases of historical population data and / or genetic testing data. As described above, examples of training datasets(s) 1034 include clinical datasets, variant datasets, and / or grading or ranking subsets used for, for example, curriculum learning.

[0183] The dataset selection component 1022 selects an appropriate dataset for training or calibration of a specific scoring model 1040. For example, when training an M2 scoring model using a curriculum learning method, the dataset selection component 1022 may select training examples with the top k text scores in a hierarchical order, where k is a positive integer. Alternatively or additionally, the dataset selection component 1022 may sort or arrange the training examples used to train an M1 scoring model based on the amount of unstructured text contained in one or more unstructured text fields of a patient's clinical record. For example, the dataset selection component 1022 may group the training examples according to the presence or absence of text in one or more unstructured text fields, and then arrange the groups such that, for example, a group with more text in an unstructured text field is input into the M1 model before a group with less text or no text in an unstructured text field.

[0184] When the course learning method is used to train the scoring model, the model training component 1024 can determine the schedule for applying the scoring model to the graded or sorted dataset prepared by the dataset selection component 1022, and the model calibration component 1026 can provide feedback to the model training component 1024, which can use the feedback to change the scheduling or progress control of the training process.

[0185] Although not in Figure 10 As specifically illustrated, however, computing system 1000 may include one or more user systems, which may be the same device as computing system 1000 or different devices. A user system includes at least one computing device, such as a personal computing device, server, mobile computing device, or smart appliance. The user system includes at least one software application, including a user interface installed on the computing device or accessible via a network by the computing device. For example, the user interface may include a graphical display screen that displays controls and graphical elements for operating and / or manipulating the output of one or more scoring models 1040 and / or controlling one or more data selection, feature generation, training, calibration, or verification processes.

[0186] User interfaces can be used to input data, initiate user interface events, and view or otherwise perceive outputs including patient predictions, variant pathogenicity predictions, and / or other data generated by the clinical variant modeling system. Examples of user interfaces include web browsers, command-line interfaces, and mobile application front-ends. User interfaces may include application programming interfaces (APIs). User interfaces may include front-end portions of application systems used by clinicians, scientists, or laboratory technicians. For example, the output of the clinical variant modeling system may be transmitted to a user interface of a computing device used by a clinician, and displayed by that user interface. Alternatively or additionally, another version of the user interface may include a front-end portion of an application software system used by variant scientists and / or other individuals working in the field of genetic testing. Thus, the output of the clinical variant modeling system 1050 may be transmitted to a user interface of a computing device used by any of these and / or other individuals, and displayed by that user interface.

[0187] Application systems can include any type of application software system that provides or implements the generation, display, or manipulation of the output produced by a clinical variation modeling system. Examples of application systems include, but are not limited to, variation classification systems, DNA (deoxyribonucleic acid) analysis software, gene testing software, medical testing software, healthcare management software, or any combination of the foregoing.

[0188] Although not specifically shown, a data storage system may include data storage and / or data services that store data (such as training data, validation data, machine learning model parameters and coefficients, performance criteria, validation criteria, machine learning model outputs, etc.) received, used, manipulated, and generated by application systems and / or clinical variation modeling systems. In some embodiments, the data storage system includes multiple different types of data storage and / or distributed data services. As used herein, a data service may refer to a physical, geographical grouping of machines, a logical grouping of machines, or a single machine. For example, a data service may be a data center, a cluster, a set of clusters, or machines.

[0189] The data storage system resides on at least one persistent and / or volatile storage device in a local network that may reside within the same local network as at least one other device of the computing system 1000 and / or in a network that is remote relative to at least one other device of the computing system 1000. Therefore, although depicted as included in the computing system 1000, a portion of the data storage system 1080 may be part of the computing system 1000 or accessible by the computing system 1000 via a network.

[0190] Although not specifically shown, it should be understood that any component of the computing system 1000 may include one or more interfaces implemented as computer programming code stored in computer memory, which, when executed, causes the computing device to perform bidirectional communication with any other component of the computing system 1000 using a communication coupling mechanism. Examples of communication coupling mechanisms include network interfaces, inter-process communication (IPC) interfaces, and application programming interfaces (APIs).

[0191] Each component of the computing system 1000 is implemented using at least one computing device that can be communicatively coupled to one or more electronic communication networks. Any component of the computing system 1000 can be bidirectionally communicatively coupled to any other component of the computing system 1000 via the network.

[0192] Typical users of the Computing System 1000 can be administrators or end users of application systems and / or clinical variation modeling systems.

[0193] The components of computing system 1000 are characterized and functionalized using computer software, hardware, or a combination of software and hardware. The characteristics and functionality of the components of computing system 1000 may include combinations of automation functions, data structures, and digital data schematically illustrated in the accompanying drawings. For ease of discussion, components may be shown as independent elements in the drawings; however, unless otherwise stated, the illustrations do not imply the need to separate these elements. The illustrated systems, services, and data storage (or their functionality) may be partitioned across any number of physical systems (including a single physical computer system), and the illustrated systems, services, and data storage (or their functionality) may communicate with each other in any suitable manner.

[0194] Although not specifically shown, a network can be implemented on any medium or mechanism that facilitates the exchange of data, signals, and / or instructions between various components of the computing system 1000. Examples of networks include, but are not limited to, local area networks (LANs), wide area networks (WANs), Ethernet or the Internet, or at least one terrestrial, satellite, optical, or wireless link, or any combination of different network and / or communication links.

[0195] Figure 11 This is a block diagram of an example computer system in which aspects of this disclosure can be operated. Figure 11 An example machine is shown as computer system 1100, within which an instruction set can be executed to cause the machine to perform any of the methods discussed herein. In some embodiments, computer system 1100 may correspond to a networked computer system (e.g., Figure 10 A component of a computing system 1000, said component including, coupled to, or utilizing the machine to execute an operating system to perform operations related to... Figure 10 The above operations correspond to aspects of the computing system 1000.

[0196] The machine is connected (e.g., networked) to other machines on a local area network (LAN), intranet, extranet, and / or the Internet. The machine can operate as a peer-to-peer machine in a peer-to-peer (or distributed) network environment, or as a server or client machine in a cloud computing infrastructure or environment, or as a server or client machine in a client-server network environment.

[0197] A machine is a personal computer (PC), smartphone, tablet PC, set-top box (STB), personal digital assistant (PDA), cellular phone, web device, server, or any machine capable of executing (sequentially or otherwise) a set of instructions specifying actions to be taken by that machine. Furthermore, although a single machine is shown, the term "machine" should also be considered as any collection of machines that individually or jointly execute a set (or multiple sets of instructions) to perform any of the methods discussed herein.

[0198] Example computer system 1100 includes processing device 1102, main memory 1104 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM), etc.), memory 1105 (e.g., flash memory, static random access memory (SRAM), etc.), input / output system 1110 and data storage system 1140, which communicate with each other via bus 1130.

[0199] Processing device 1102 represents at least one general-purpose processing device, such as a microprocessor, central processing unit, etc. More specifically, the processing device may be a Complex Instruction Set Computing (CISC) microprocessor, a Reduced Instruction Set Computing (RISC) microprocessor, a Very Long Instruction Word (VLIW) microprocessor, or a processor implementing other instruction sets, or a processor implementing a combination of instruction sets. Processing device 1102 may also be at least one special-purpose processing device, such as an Application-Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), a Digital Signal Processor (DSP), a Network Processor, etc. Processing device 1102 is configured to execute instructions 1112 for performing the operations and steps discussed herein.

[0200] When processing device 1102 is executing a portion of clinical variation modeling system 1150, instructions 1112 include those portions of the clinical variation modeling system. Therefore, the clinical variation modeling system 1150 is shown in dashed lines as part of instructions 1112 to illustrate portions of the clinical variation modeling system sometimes executed by processing device 1102. For example, when at least some portion of the clinical variation modeling system 1150 is implemented in instructions causing processing device 1102 to perform one or more of the methods described above, some of those instructions may be read from main memory 1104 and / or data storage system 1140 into processing device 1102 (e.g., read into internal cache or other memory). However, it is not required that the entire clinical variation modeling system 1150 be included in instructions 1112 simultaneously, and at other times (e.g., when processing device 1102 is not executing at least one portion of the clinical variation modeling system), some portions of the clinical variation modeling system 1150 may be stored in at least one other component of computer system 1100.

[0201] Computer system 1100 further includes a network interface device 1108 that communicates via network 1120. Network interface device 1108 provides bidirectional data communication coupled to the network. For example, network interface device 1108 may be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem providing a data communication connection to a corresponding type of telephone line. As another example, network interface device 1108 may be a Local Area Network (LAN) card providing a data communication connection to a LAN-compatible network. A wireless link may also be implemented. In any such implementation, network interface device 1108 can transmit and receive electrical, electromagnetic, or optical signals carrying digital data streams representing various types of information.

[0202] A network link can provide data communication to other data devices through at least one network. For example, a network link can provide a connection to a global packet data communication network commonly known as the "Internet," such as a connection from a local network to a host computer or to a data device operated by an Internet Service Provider (ISP). The local network and the Internet use electrical, electromagnetic, or optical signals carrying digital data to and from the computer system 1100.

[0203] Computer system 1100 can send messages and receive data (including program code) via one or more networks and network interface devices 1108. In an Internet example, the server can transmit application request code via the Internet and network interface device 1108. The received code can be executed by processing device 1102 when it is received and / or stored in data storage system 1140 or other non-volatile storage device for later execution.

[0204] Input / output system 1110 includes output devices, such as a display (e.g., a liquid crystal display (LCD) or touchscreen display) for displaying information to a computer user, or speakers, haptic devices, or other forms of output devices. Input / output system 1110 may include input devices, such as alphanumeric keys and other keys configured to transmit information and command selections to processing device 1102. Alternatively or additionally, input devices may include cursor controls, such as a mouse, trackball, or cursor direction keys for transmitting directional information and command selections to processing device 1102 and for controlling cursor movement on the display. Alternatively or additionally, input devices may include microphones, sensors, or sensor arrays for transmitting sensed information to processing device 1102. For example, sensed information may include voice commands, audio signals, geographic location information, and / or digital images.

[0205] Data storage system 1140 includes machine-readable storage medium 1142 (also referred to as computer-readable medium) on which at least one set of instructions 1144 or software that implement any of the methods or functions described herein are stored. During execution by computer system 1100, instructions 1144 may also reside wholly or at least partially in main memory 1104 and / or processing device 1102, which also constitute machine-readable storage media.

[0206] In one embodiment, instructions 1112, 1114, and 1144 include instructions for implementing functionality corresponding to a clinical variation modeling system (e.g., any one or more components of the clinical variation modeling methods and techniques described herein).

[0207] exist Figure 11 The dashed lines indicate that it is not required that the clinical variation modeling system 1150 be fully contained in instructions 1112, 1114, and 1144 simultaneously. In one example, a portion of the clinical variation modeling system is contained in instruction 1144, which is read into main memory 1104 as instruction 1114, and a portion of instruction 1114 is read into processing device 1102 as instruction 1112 for execution. In another example, some portions of the clinical variation modeling system are implemented in instruction 1144, while other portions are implemented in instruction 1114 and still other portions are implemented in instruction 1112.

[0208] Although machine-readable storage medium 1142 is shown as a single medium in the exemplary embodiment, the term "machine-readable storage medium" should be considered as a single or multiple media comprising at least one set of storage instructions. The term "machine-readable storage medium" should also be considered as any medium capable of storing or encoding a set of instructions for execution by a machine and causing the machine to perform any of the methods of this disclosure. The term "machine-readable storage medium" should therefore be considered as including, but not limited to, solid-state memories, optical media, and magnetic media.

[0209] Some of the details described above have already been presented regarding the algorithms and symbolic representations of operations on data bits within computer memory. These algorithmic descriptions and representations are methods used by those skilled in the art of data processing to most effectively communicate the essence of their work to others skilled in the art. An algorithm here is generally considered to be a self-consistent sequence of operations that leads to a desired result. An operation is one that requires the physical manipulation of physical quantities. Typically, though not strictly necessary, these quantities take the form of electrical, optical, or magnetic signals that can be stored, combined, compared, and otherwise manipulated. It has sometimes proven convenient to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc., primarily for common use.

[0210] However, it should be remembered that all these and similar terms are to be associated with appropriate physical quantities and are merely convenient labels applied to those quantities. This disclosure may refer to the actions and processes of a computer system or similar electronic computing device that manipulate and convert data represented as physical (electronic) quantities within the registers and memories of the computer system into other data similarly represented as physical quantities within the computer system's memory or registers or other such information storage systems.

[0211] This disclosure also relates to an apparatus for performing the operations described herein. This apparatus may be specifically constructed for its intended purpose, or it may comprise a general-purpose computer that can be selectively activated or reconfigured by a computer program stored in the computer. For example, a computer system such as Computing System 1100 or other data processing system may implement the techniques described above in response to its processor executing a computer program (e.g., a sequence of instructions) contained in memory or other non-transitory machine-readable storage medium. Such a computer program may be stored in a computer-readable storage medium, such as, but not limited to, any type of disk (including floppy disks, optical disks, CD-ROMs and magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic cards, or optical cards) or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.

[0212] The algorithms and displays presented herein are not inherently related to any particular computer or other device. Based on the teachings herein, various general-purpose systems can be used with the program, or it can be demonstrated that it is convenient to construct more specialized devices to execute the methods. The architectures of various such systems will appear as illustrated in the description below. Furthermore, this disclosure is described without reference to any particular programming language. It will be appreciated that the teachings of this disclosure as described herein can be implemented using various programming languages.

[0213] This disclosure may be provided as a computer program product or software, which may include a machine-readable medium having instructions stored thereon, the instructions being used to program a computer system (or other electronic device) to perform processes according to this disclosure. Machine-readable media include any mechanism for storing information in a machine-readable (e.g., computer-readable) form. In some embodiments, machine-readable (e.g., computer-readable) media include machine-readable (e.g., computer-readable) storage media, such as read-only memory (“ROM”), random access memory (“RAM”), disk storage media, optical storage media, flash memory components, etc.

[0214] The following provides illustrative aspects of the technology disclosed herein. Embodiments of the technology may include any aspect of the aspects described herein, any combination of any aspect of the aspects described herein, or any combination of any part of the aspects described herein.

[0215] In some aspects, the techniques described herein relate to a method comprising: extracting phenotypic data from clinical records associated with a patient population and genetic condition, wherein the phenotypic data includes natural language text, which includes at least one of indicators for testing, family history descriptions, or demographic data; extracting genotype data from genetic test results associated with the genetic condition, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnosis associated with the at least one variant; configuring a natural language processing (NLP)-based machine learning model using the phenotypic data and genotype data including natural language text to generate and output patient scores for the patient population, wherein the patient scores include a machine learning likelihood of the patient being affected by the genetic condition given the phenotypic data including natural language text; and configuring a Bayesian causal model using the patient scores and genotype data to generate and output variant scores, wherein the variant scores include a machine learning likelihood of the variant being pathogenic to the genetic condition.

[0216] In some respects, the techniques described herein involve a method that further includes: using a configured NLP-based machine learning model and a configured Bayesian causal model to generate and output predictions about whether variants associated with unseen clinical records are benign or pathogenic in relation to a genetic condition.

[0217] In some respects, the techniques described herein involve a method that further includes providing clinicians with predictions about whether a variant is benign or pathogenic, for clinicians to use in developing a patient diagnosis.

[0218] In some respects, the techniques described herein relate to a method that further includes providing predictions to at least one system, process, model, or component of a variant classification system or clinical data system.

[0219] In some respects, the techniques described in this paper involve a method that further includes using a variation score output by a configured Bayesian causal model as input to a classification framework.

[0220] In some respects, the techniques described herein relate to a method in which configuring a natural language processing (NLP)-based machine learning model further includes: configuring a first model to receive natural language text and generate and output text scores, wherein the text scores include machine learning likelihoods of a patient’s genetic condition given the received natural language text; and configuring a second model to generate and output patient scores using the text scores and a second model input, the second model input including phenotypic data excluding the natural language text.

[0221] In some respects, the techniques described herein relate to a method that further includes: using curriculum learning to configure an NLP-based machine learning model to generate and output patient scores, wherein curriculum learning includes using text scores to determine the order of inputs for a second model, and applying the second model to the second model inputs according to that order.

[0222] In some respects, the techniques described herein relate to a method that further includes: iteratively adjusting the values ​​of at least one clinical feature weights of an NLP-based machine learning model until a patient score output by the NLP-based machine learning model meets at least one first performance criterion, to produce a configured NLP-based machine learning model, wherein the configured NLP-based machine learning model is capable of outputting a patient score that meets at least one second performance criterion.

[0223] In some respects, the techniques described herein relate to a method that further includes configuring an NLP-based machine learning model using positive and negative examples, wherein positive examples include positive associations between molecular diagnoses and genetic conditions, and negative examples include negative associations between molecular diagnoses and genetic conditions.

[0224] In some respects, the techniques described in this paper involve a method in which a Bayesian causal model includes probability distributions of patient scores and variant scores.

[0225] In some aspects, the techniques described herein relate to a system comprising: at least one processor; and at least one memory coupled to the at least one processor, wherein the at least one memory includes at least one instruction that, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: extracting phenotypic data from clinical records associated with a patient population and genetic condition, wherein the phenotypic data includes natural language text, the natural language text including at least one of indicators for testing, family history descriptions, or demographic data; using the phenotypic data including the natural language text, generating and outputting a patient score for the patient population, wherein the patient score includes a likelihood that the patient is affected by the genetic condition given the phenotypic data including the natural language text; extracting genotype data from genetic test results associated with the patient population and genetic condition, wherein the genotype data includes a molecular diagnosis, the molecular diagnosis including at least one variant of a gene; and configuring a Bayesian causal model using the patient score and genotype data to generate and output a variant score, wherein the variant score includes a machine learning likelihood that the variant is pathogenic to the genetic condition.

[0226] In some respects, the techniques described herein relate to a system in which at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: using a configured Bayesian causal model to generate and output predictions about whether a variant associated with an unseen clinical record is benign, of uncertain significance, or pathogenic to a genetic condition.

[0227] In some respects, the technology described herein relates to a system in which at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: providing a clinician with a prediction of whether the variant is benign or pathogenic for the clinician to use in formulating a diagnosis for the patient.

[0228] In some respects, the techniques described herein relate to a system in which at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: providing a prediction to at least one system, process, model, or component of a variant classification system or a clinical data system.

[0229] In some respects, the techniques described herein relate to a system in which at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: using a variation score output by a configured Bayesian causal model as input to a classification framework.

[0230] In some respects, the techniques described herein relate to a system in which at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: configuring a natural language processing (NLP)-based machine learning model to generate and output patient scores in such a way as to: (i) configuring a first model to generate and output text scores, wherein the text scores include machine learning likelihoods of the patient’s genetic condition given a first model input comprising natural language text; and (ii) configuring a second model to generate and output patient scores using a second model input comprising text scores and phenotypic data excluding natural language text.

[0231] In some respects, the techniques described herein relate to a system in which at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: using curriculum learning to configure an NLP-based machine learning model to generate and output patient scores, wherein curriculum learning includes using text scores to determine an order for input to a second model, and applying the second model to the second model input according to that order.

[0232] In some aspects, the techniques described herein relate to a system in which at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: configuring a natural language processing (NLP)-based machine learning model to generate and output a patient score; and iteratively adjusting the values ​​of at least one clinical feature weights of the NLP-based machine learning model until the patient score output by the NLP-based machine learning model meets at least one first performance criterion, to produce a configured NLP-based machine learning model, wherein the configured NLP-based machine learning model is capable of outputting a patient score that meets at least one second performance criterion.

[0233] In some respects, the techniques described herein relate to a system in which at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: configuring an NLP-based machine learning model using positive and negative examples, wherein positive examples include positive associations between molecular diagnoses and genetic conditions, and negative examples include negative associations between molecular diagnoses and genetic conditions.

[0234] In some respects, the techniques described in this paper involve a system in which a Bayesian causal model includes probability distributions of patient scores and variant scores.

[0235] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable storage medium comprising at least one instruction that, when executed by at least one processor, causes at least one processor to perform at least one operation, the at least one operation comprising: extracting phenotypic data from clinical records associated with a patient population and genetic condition, wherein the phenotypic data includes natural language text, the natural language text including at least one of indications for testing, a family history description, or a demographic description; configuring a natural language processing (NLP)-based machine learning model using the phenotypic data including the natural language text to generate and output patient scores for the patient population, wherein the patient scores include a machine learning likelihood of the patient being affected by the genetic condition given the phenotypic data including the natural language text; extracting genotype data from genetic testing results associated with the patient population and genetic condition, wherein the genotype data includes a molecular diagnosis, the molecular diagnosis including at least one variant of a gene; and using the patient scores and genotype data to generate and output a variant score, wherein the variant score includes a likelihood that the variant is pathogenic to the genetic condition.

[0236] In some respects, the techniques described herein relate to at least one non-transitory machine-readable storage medium, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: using a configured NLP-based machine learning model to generate and output a prediction of whether a variant associated with an unseen clinical record is benign or pathogenic to a genetic condition.

[0237] In some respects, the techniques described herein relate to at least one non-transitory machine-readable storage medium, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: providing a clinician with a prediction of whether the variant is benign or pathogenic for the clinician to use in formulating a diagnosis for the patient.

[0238] In some respects, the techniques described herein relate to at least one non-transitory machine-readable storage medium, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: providing a prediction to at least one system, process, model, or component of a variant classification system or clinical data system.

[0239] In some respects, the techniques described herein relate to at least one non-transitory machine-readable storage medium, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: using a mutation score as input to a classification framework.

[0240] In some respects, the techniques described herein relate to at least one non-transitory machine-readable storage medium, wherein configuring an NLP-based machine learning model further includes: configuring a first model to receive natural language text and generate and output text scores, wherein the text scores include machine learning likelihood of a patient's genetic condition given the received natural language text; and configuring a second model to generate and output patient scores using second model inputs, the second model inputs including text scores and second phenotypic data excluding natural language text.

[0241] In some respects, the techniques described herein relate to at least one non-transitory machine-readable storage medium, wherein configuring an NLP-based machine learning model further includes: using curriculum learning to configure the NLP-based machine learning model to generate and output patient scores, wherein curriculum learning includes using text scores to determine the order of inputs for a second model, and applying the second model to the second model inputs according to that order.

[0242] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable storage medium, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: iteratively adjusting the values ​​of at least one clinical feature weights of an NLP-based machine learning model until a patient score output by the NLP-based machine learning model meets at least one first performance criterion, to produce a configured NLP-based machine learning model, wherein the configured NLP-based machine learning model is capable of outputting a patient score that meets at least one second performance criterion.

[0243] In some aspects, the techniques described herein relate to at least one non-transitory machine-readable storage medium, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: configuring an NLP-based machine learning model using positive and negative examples, wherein positive examples include positive associations between molecular diagnoses and genetic conditions, and negative examples include negative associations between molecular diagnoses and genetic conditions.

[0244] In some respects, the techniques described herein relate to at least one non-transitory machine-readable storage medium, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: using a Bayesian causal model to generate and output a variant score, wherein the Bayesian causal model includes a probability distribution of patient scores and variant scores.

[0245] In some aspects, the techniques described herein relate to a method comprising: extracting phenotypic data from clinical records associated with a patient population and genetic condition, wherein the phenotypic data includes natural language text, which includes at least one of indications for testing, a family history description, or demographic data; extracting genotype data from genetic test results associated with the genetic condition, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnosis associated with the at least one variant; and configuring a machine learning model using the phenotypic data and genotype data including natural language text to generate and output patient scores for the patient population, wherein the patient scores include a machine learning likelihood of the patient's genetic condition being influenced by the phenotypic data including natural language text.

[0246] In some respects, the techniques described herein relate to a method comprising: extracting genotype data from genetic test results associated with a genetic condition, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnosis associated with the at least one variant; and configuring a Bayesian causal model using patient scores and genotype data to generate and output a variant score, wherein the variant score includes a machine learning likelihood that the variant is pathogenic to the genetic condition.

[0247] In some respects, the techniques described herein relate to any one or more aspects, steps, components, elements, processes, or limitations of at least one of those described in the appended description and / or shown in the accompanying drawings.

[0248] Clause 1. A method comprising: extracting phenotypic data from clinical records associated with a patient population and a genetic condition, wherein the phenotypic data includes natural language text, the natural language text including at least one of indications for testing, a family history description, or demographic data; extracting genotype data from genetic test results associated with the genetic condition, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnosis associated with the at least one variant; configuring a natural language processing (NLP)-based machine learning model using the phenotypic data including natural language text and the genotype data to generate and output patient scores for the patient population, wherein the patient scores include a machine learning likelihood of the patient being affected by the genetic condition given the phenotypic data including natural language text; and configuring a Bayesian causal model using the patient scores and the genotype data to generate and output variant scores, wherein the variant scores include a machine learning likelihood of the variant being pathogenic to the genetic condition.

[0249] Clause 2. The method of Clause 1 further includes: using a configured NLP-based machine learning model and a configured Bayesian causal model to generate and output predictions about whether variants associated with unseen clinical records are benign or pathogenic to a genetic condition.

[0250] Clause 3. The method of Clause 2 further includes: providing clinicians with predictions about whether the variant is benign or pathogenic for clinicians to use in developing a patient diagnosis.

[0251] Clause 4. The method of Clause 2 further includes: providing predictions to at least one system, process, model, or component of a variant classification system or clinical data system.

[0252] Clause 5. The method of Clause 2 further includes: using the variation score output by the configured Bayesian causal model as input to the classification framework.

[0253] Clause 6. The method of any of Clauses 1-5, wherein configuring a machine learning model based on natural language processing (NLP) further comprises: configuring a first model to receive natural language text and generate and output a text score, wherein the text score includes a machine learning likelihood of the patient’s genetic status given the received natural language text; and configuring a second model to generate and output a patient score using the text score and a second model input, the second model input including phenotypic data excluding the natural language text.

[0254] Clause 7. The method of Clause 6 further includes: using curriculum learning to configure an NLP-based machine learning model to generate and output patient scores, wherein curriculum learning includes using text scores to determine the order of inputs for a second model, and applying the second model to the second model inputs in that order.

[0255] Clause 8. The method of any of Clauses 1-7 further comprises: iteratively adjusting the values ​​of at least one clinical feature weight of the NLP-based machine learning model until the patient score output by the NLP-based machine learning model meets at least one first performance criterion, to produce a configured NLP-based machine learning model, wherein the configured NLP-based machine learning model is capable of outputting a patient score that meets at least one second performance criterion.

[0256] Clause 9. The method of any of the clauses 1-8 further includes: using positive and negative examples to configure an NLP-based machine learning model, wherein positive examples include positive associations between molecular diagnoses and genetic conditions, and negative examples include negative associations between molecular diagnoses and genetic conditions.

[0257] Clause 10. The method of any of the clauses 1-9, wherein the Bayesian causal model includes the probability distribution of patient scores and variant scores.

[0258] Clause 11. A system comprising: at least one processor; and at least one memory coupled to the at least one processor, wherein the at least one memory includes at least one instruction, which, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: extracting phenotypic data from clinical records associated with a patient population and a genetic condition, wherein the phenotypic data includes natural language text, the natural language text including at least one of indicators for testing, a family history description, or demographic data; generating and outputting a patient score for the patient population using the phenotypic data including the natural language text, wherein the patient score includes a likelihood that the patient is affected by the genetic condition given the phenotypic data including the natural language text; extracting genotype data from genetic test results associated with the patient population and the genetic condition, wherein the genotype data includes a molecular diagnosis, the molecular diagnosis including at least one variant of a gene; and configuring a Bayesian causal model using the patient score and the genotype data to generate and output a variant score, wherein the variant score includes a machine learning likelihood that the variant is pathogenic to the genetic condition.

[0259] Clause 12. The system of Clause 11, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: using a configured Bayesian causal model to generate and output predictions about whether a variant associated with an unseen clinical record is benign, of uncertain significance, or pathogenic to a genetic condition.

[0260] Clause 13. The system of Clause 12, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: providing a clinician with a prediction of whether the variant is benign or pathogenic for the clinician to use in formulating a diagnosis for the patient.

[0261] Clause 14. The system of Clause 12, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: providing a prediction to at least one system, process, model, or component of a variant classification system or a clinical data system.

[0262] Clause 15. A system of any of the clauses 11-14, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: using a variation score output by a configured Bayesian causal model as input to a classification framework.

[0263] Clause 16. A system of any of Clauses 11-15, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: configuring a natural language processing (NLP)-based machine learning model to generate and output a patient score in such a way as to: (i) configuring a first model to generate and output a text score, wherein the text score includes a machine learning likelihood of the patient’s genetic condition given a first model input comprising natural language text; and (ii) configuring a second model to generate and output a patient score using a second model input comprising the text score and phenotypic data excluding the natural language text.

[0264] Clause 17. The system of Clause 16, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: using curriculum learning to configure an NLP-based machine learning model to generate and output patient scores, wherein curriculum learning includes using text scores to determine an order for input to a second model, and applying the second model to the second model input according to that order.

[0265] Clause 18. A system of any of the clauses 11-17, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: configuring a natural language processing (NLP)-based machine learning model to generate and output a patient score; and iteratively adjusting the values ​​of at least one clinical feature weight of the NLP-based machine learning model until the patient score output by the NLP-based machine learning model meets at least one first performance criterion, to produce a configured NLP-based machine learning model, wherein the configured NLP-based machine learning model is capable of outputting a patient score that meets at least one second performance criterion.

[0266] Clause 19. A system of any of the clauses 11-18, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: configuring an NLP-based machine learning model using positive and negative examples, wherein positive examples comprise a positive association between molecular diagnoses and genetic conditions, and negative examples comprise a negative association between molecular diagnoses and genetic conditions.

[0267] Clause 20. A system of any of the clauses 11-19, wherein the Bayesian causal model comprises the probability distribution of patient scores and variant scores.

[0268] Clause 21. At least one non-transitory machine-readable storage medium comprising at least one instruction, which, when executed by at least one processor, causes at least one processor to perform at least one operation, the at least one operation comprising: extracting phenotypic data from clinical records associated with a patient population and genetic condition, wherein the phenotypic data includes natural language text, the natural language text including at least one of indications for testing, a family history description, or a demographic description; configuring a natural language processing (NLP)-based machine learning model using the phenotypic data including the natural language text to generate and output a patient score for the patient population, wherein the patient score includes a machine learning likelihood of the patient being affected by the genetic condition given the phenotypic data including the natural language text; extracting genotype data from genetic test results associated with the patient population and genetic condition, wherein the genotype data includes a molecular diagnosis, the molecular diagnosis including at least one variant of a gene; and using the patient score and the genotype data to generate and output a variant score, wherein the variant score includes a likelihood that the variant is pathogenic to the genetic condition.

[0269] Clause 22. At least one non-transitory machine-readable storage medium of Clause 21, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: using a configured NLP-based machine learning model to generate and output a prediction of whether a variant associated with an unseen clinical record is benign or pathogenic to a genetic condition.

[0270] Clause 23. At least one non-transitory machine-readable storage medium of Clause 22, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: providing a clinician with a prediction of whether the variant is benign or pathogenic for the clinician to use in formulating a diagnosis for the patient.

[0271] Clause 24. At least one non-transitory machine-readable storage medium of Clause 22, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: providing a prediction to at least one system, process, model, or component of a variant classification system or a clinical data system.

[0272] Clause 25. At least one non-transitory machine-readable storage medium of any of Clauses 21-24, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: using a variation score as input to a classification framework.

[0273] Clause 26. At least one non-transitory machine-readable storage medium of any of Clauses 21-25, wherein configuring an NLP-based machine learning model further comprises: configuring a first model to receive natural language text and generate and output text scores, wherein the text scores include machine learning likelihood of the patient’s genetic status given the received natural language text; and configuring a second model to generate and output patient scores using second model inputs, the second model inputs including text scores and second phenotypic data excluding natural language text.

[0274] Clause 27. At least one non-transitory machine-readable storage medium of Clause 26, wherein configuring an NLP-based machine learning model further comprises: using curriculum learning to configure the NLP-based machine learning model to generate and output patient scores, wherein curriculum learning comprises using text scores to determine the order of inputs for a second model, and applying the second model to the second model inputs in that order.

[0275] Clause 28. At least one non-transitory machine-readable storage medium of any of Clauses 21-27, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: iteratively adjusting the values ​​of at least one clinical feature weights of an NLP-based machine learning model until a patient score output by the NLP-based machine learning model satisfies at least one first performance criterion, to produce a configured NLP-based machine learning model, wherein the configured NLP-based machine learning model is capable of outputting a patient score that satisfies at least one second performance criterion.

[0276] Clause 29. At least one non-transitory machine-readable storage medium of any of Clauses 21-28, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: configuring an NLP-based machine learning model using positive and negative examples, wherein positive examples comprise a positive association between molecular diagnoses and genetic conditions, and negative examples comprise a negative association between molecular diagnoses and genetic conditions.

[0277] Clause 30. At least one non-transitory machine-readable storage medium of any of Clauses 21-29, wherein at least one instruction, when executed by at least one processor, causes at least one processor to perform at least one operation, said at least one operation further comprising: using a Bayesian causal model to generate and output a variance score, wherein the Bayesian causal model includes a probability distribution of patient scores and variance scores.

[0278] Clause 31. A method comprising: extracting phenotypic data from clinical records associated with a patient population and genetic condition, wherein the phenotypic data includes natural language text, the natural language text including at least one of indications for testing, a family history description, or demographic data; extracting genotype data from genetic test results associated with the genetic condition, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnosis associated with the at least one variant; and configuring a machine learning model using the phenotypic data and genotype data including natural language text to generate and output patient scores for the patient population, wherein the patient scores include a machine learning likelihood of the patient being affected by the genetic condition given the phenotypic data including natural language text.

[0279] Clause 32. A method comprising: extracting genotype data from genetic test results associated with a genetic condition, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnosis associated with the at least one variant; and configuring a Bayesian causal model using patient scores and genotype data to generate and output a variant score, wherein the variant score includes a machine learning likelihood that the variant is pathogenic to the genetic condition.

[0280] Article 33. The methods of Article 31 or 32, including any of the foregoing provisions.

[0281] In the foregoing description, embodiments of the present disclosure have been described with reference to specific exemplary embodiments therein. It will be apparent that various modifications can be made thereto without departing from the broader spirit and scope of the embodiments of the present disclosure as set forth in the appended claims. Therefore, the description and drawings are to be regarded as illustrative rather than restrictive.

Claims

1. A method comprising: Phenotypic data are extracted from clinical records associated with patient populations and genetic conditions, wherein the phenotypic data includes natural language text, which includes at least one of the indicators for testing, family history descriptions, or demographic data. Genotype data is extracted from gene test results associated with the genetic condition, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnosis associated with the at least one variant; Using the phenotypic data and genotype data including the natural language text, a natural language processing (NLP)-based machine learning model is configured to generate and output patient scores for the patient population, wherein the patient scores include a machine learning likelihood of the patient's genetic condition being influenced by the phenotypic data including the natural language text; and Using the patient scores and the genotype data, a Bayesian causal model is configured to generate and output a variant score, wherein the variant score includes a machine learning likelihood that the variant is pathogenic to the genetic condition.

2. The method of claim 1, further comprising: A configured NLP-based machine learning model and a configured Bayesian causal model are used to generate and output predictions about whether variants associated with unseen clinical records are benign or pathogenic to the genetic condition.

3. The method of claim 2, further comprising: Provide clinicians with predictions about whether the variant is benign or pathogenic, so that they can use these predictions to develop a diagnosis for the patient.

4. The method of claim 2, further comprising: The prediction is provided to at least one system, process, model, or component of a variant classification system or clinical data system.

5. The method of claim 2, further comprising: The variation score output by the configured Bayesian causal model is used as input to the variation classification framework.

6. The method of claim 1, wherein, Configuring the aforementioned natural language processing (NLP)-based machine learning model further includes: A first model is configured to receive the natural language text and generate and output a text score, wherein the text score includes a machine learning likelihood of the patient's genetic condition being affected given the received natural language text; and A second model is configured to generate and output the patient score using the text score and a second model input, the second model input including the phenotypic data excluding the natural language text.

7. The method of claim 6, further comprising: The NLP-based machine learning model is configured using a course of learning to generate and output the patient scores, wherein the course of learning includes using the text scores to determine the order of inputs for the second model, and applying the second model to the second model inputs according to the order.

8. The method of claim 1, further comprising: The values ​​of at least one clinical feature weights of the NLP-based machine learning model are iteratively adjusted until the output of a loss function calculated based on the patient score output by the NLP-based machine learning model satisfies at least one first performance criterion, to produce a configured NLP-based machine learning model, wherein the configured NLP-based machine learning model is capable of outputting a patient score that satisfies at least one second performance criterion.

9. The method of claim 1, further comprising: The NLP-based machine learning model is configured using positive and negative examples, wherein positive examples include positive associations between the molecular diagnosis and the genetic condition, and negative examples include negative associations between the molecular diagnosis and the genetic condition.

10. The method of claim 1, wherein, The Bayesian causal model includes the probability distributions of patient scores and variant scores.

11. A system comprising: At least one processor; as well as At least one memory coupled to the at least one processor, wherein the at least one memory includes at least one instruction, which, when executed by the at least one processor, causes the at least one processor to perform at least one operation, the at least one operation including: Phenotypic data are extracted from clinical records associated with patient populations and genetic conditions, wherein the phenotypic data includes natural language text, which includes at least one of the indicators for testing, family history descriptions, or demographic data. Using the phenotypic data including the natural language text, generate and output patient scores for the patient population, wherein the patient scores include the likelihood that the patient is affected by the genetic status given the phenotypic data including the natural language text; Genotype data is extracted from genetic test results associated with the patient population and the genetic condition, wherein the genotype data includes molecular diagnostics, the molecular diagnostics including at least one variant of the gene; and Using the patient scores and the genotype data, a Bayesian causal model is configured to generate and output a variant score, wherein the variant score includes a machine learning likelihood that the variant is pathogenic to the genetic condition.

12. The system of claim 11, wherein, When the at least one instruction is executed by the at least one processor, it causes the at least one processor to perform at least one operation, the at least one operation further comprising: A configured Bayesian causal model is used to generate and output predictions about whether variants associated with unseen clinical records are benign, of uncertain significance, or pathogenic to the genetic condition.

13. The system of claim 12, wherein, When the at least one instruction is executed by the at least one processor, it causes the at least one processor to perform at least one operation, the at least one operation further comprising: Provide clinicians with predictions about whether the variant is benign or pathogenic, so that they can use these predictions to develop a diagnosis for the patient.

14. The system of claim 12, wherein, When the at least one instruction is executed by the at least one processor, it causes the at least one processor to perform at least one operation, the at least one operation further comprising: The prediction is provided to at least one system, process, model, or component of a variant classification system or clinical data system.

15. The system of claim 11, wherein, When the at least one instruction is executed by the at least one processor, it causes the at least one processor to perform at least one operation, the at least one operation further comprising: The variance score, output by the configured Bayesian causal model, is used as input to the variance classification framework.

16. The system of claim 11, wherein, When the at least one instruction is executed by the at least one processor, it causes the at least one processor to perform at least one operation, the at least one operation further comprising: Configure a machine learning model based on natural language processing (NLP) to generate and output the patient score in the following manner: (i) configure a first model to generate and output a text score, wherein the text score includes a machine learning likelihood of the patient's genetic condition being affected given a first model input including the natural language text; and (ii) configure a second model to generate and output a patient score using a second model input, the second model input including the text score and phenotypic data excluding the natural language text.

17. The system of claim 16, wherein, When the at least one instruction is executed by the at least one processor, it causes the at least one processor to perform at least one operation, the at least one operation further comprising: The NLP-based machine learning model is configured using a course of learning to generate and output the patient scores, wherein the course of learning includes using the text scores to determine the order of inputs for the second model, and applying the second model to the second model inputs according to the order.

18. The system of claim 11, wherein, When the at least one instruction is executed by the at least one processor, it causes the at least one processor to perform at least one operation, the at least one operation further comprising: Configure a natural language processing (NLP)-based machine learning model to generate and output the patient scores; and The values ​​of at least one clinical feature weights of the NLP-based machine learning model are iteratively adjusted until the output of a loss function calculated based on the patient score output by the NLP-based machine learning model satisfies at least one first performance criterion, to produce a configured NLP-based machine learning model, wherein the configured NLP-based machine learning model is capable of outputting a patient score that satisfies at least one second performance criterion.

19. The system of claim 18, wherein, When the at least one instruction is executed by the at least one processor, it causes the at least one processor to perform at least one operation, the at least one operation further comprising: The NLP-based machine learning model is configured using positive and negative examples, wherein positive examples include positive associations between molecular diagnoses and genetic conditions, and negative examples include negative associations between molecular diagnoses and genetic conditions.

20. The system of claim 11, wherein, The Bayesian causal model includes the probability distributions of patient scores and variant scores.

21. At least one non-transitory machine-readable storage medium, comprising at least one instruction, which, when executed by at least one processor, causes the at least one processor to perform at least one operation, the at least one operation comprising: Phenotypic data are extracted from clinical records associated with patient populations and genetic conditions, wherein the phenotypic data includes natural language text, which includes at least one of the indicators for testing, family history descriptions, or demographic descriptions. Using the phenotypic data including the natural language text, configure a natural language processing (NLP)-based machine learning model to generate and output patient scores for the patient population, wherein the patient scores include machine learning likelihood of the patient's genetic condition being affected given the phenotypic data including the natural language text; Genotype data is extracted from genetic test results associated with the patient population and the genetic condition, wherein the genotype data includes molecular diagnostics, the molecular diagnostics including at least one variant of the gene; and The patient scores and genotype data are used to generate and output a variant score, wherein the variant score includes the likelihood that the variant is pathogenic to the genetic condition.

22. The at least one non-transitory machine-readable storage medium as described in claim 21, wherein, When the at least one instruction is executed by the at least one processor, it causes the at least one processor to perform at least one operation, the at least one operation further comprising: A configured NLP-based machine learning model is used to generate and output predictions about whether variants associated with unseen clinical records are benign or pathogenic to the genetic condition.

23. The at least one non-transitory machine-readable storage medium as described in claim 22, wherein, When the at least one instruction is executed by the at least one processor, it causes the at least one processor to perform at least one operation, the at least one operation further comprising: Provide clinicians with predictions about whether the variant is benign or pathogenic, so that they can use these predictions to develop a diagnosis for the patient.

24. The at least one non-transitory machine-readable storage medium as described in claim 22, wherein, When the at least one instruction is executed by the at least one processor, it causes the at least one processor to perform at least one operation, the at least one operation further comprising: The prediction is provided to at least one system, process, model, or component of a variant classification system or clinical data system.

25. The at least one non-transitory machine-readable storage medium as claimed in claim 21, wherein, When the at least one instruction is executed by the at least one processor, it causes the at least one processor to perform at least one operation, the at least one operation further comprising: The mutation score is used as input to the mutation classification framework.

26. The at least one non-transitory machine-readable storage medium as described in claim 21, wherein, Configuring the NLP-based machine learning model further includes: A first model is configured to receive the natural language text and generate and output a text score, wherein the text score includes a machine learning likelihood of the patient's genetic condition being affected given the received natural language text; and A second model is configured to generate and output the patient score using a second model input, the second model input including the text score and second phenotypic data excluding the natural language text.

27. The at least one non-transitory machine-readable storage medium as described in claim 26, wherein, Configuring the NLP-based machine learning model further includes: The NLP-based machine learning model is configured using a course of learning to generate and output the patient scores, wherein the course of learning includes using the text scores to determine the order of inputs for the second model, and applying the second model to the second model inputs according to the order.

28. The at least one non-transitory machine-readable storage medium as described in claim 21, wherein, When the at least one instruction is executed by the at least one processor, it causes the at least one processor to perform at least one operation, the at least one operation further comprising: The values ​​of at least one clinical feature weights of the NLP-based machine learning model are iteratively adjusted until the output of a loss function calculated based on the patient score output by the NLP-based machine learning model satisfies at least one first performance criterion, to produce a configured NLP-based machine learning model, wherein the configured NLP-based machine learning model is capable of outputting a patient score that satisfies at least one second performance criterion.

29. The at least one non-transitory machine-readable storage medium as described in claim 21, wherein, When the at least one instruction is executed by the at least one processor, it causes the at least one processor to perform at least one operation, the at least one operation further comprising: The NLP-based machine learning model is configured using positive and negative examples, wherein positive examples include positive associations between the molecular diagnosis and the genetic condition, and negative examples include negative associations between the molecular diagnosis and the genetic condition.

30. The at least one non-transitory machine-readable storage medium as described in claim 21, wherein, When the at least one instruction is executed by the at least one processor, it causes the at least one processor to perform at least one operation, the at least one operation further comprising: A Bayesian causal model is used to generate and output variant scores, wherein the Bayesian causal model includes the probability distribution of patient scores and variant scores.

31. A method comprising: Phenotypic data is extracted from a collection of clinical records associated with patient populations and genetic conditions, wherein the phenotypic data includes natural language text, which includes at least one of the indicators for testing, family history descriptions, or demographic data. Genotype data is extracted from gene test results associated with the genetic condition, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnosis associated with the at least one variant; and Using the phenotypic data and genotype data including the natural language text, a machine learning model is configured to generate and output patient scores for the patient population, wherein the patient scores include machine learning likelihood of the patient's genetic condition being affected given the phenotypic data including the natural language text.

32. A method comprising: Genotype data is extracted from gene test results associated with a genetic condition, wherein the genotype data includes at least one variant of a gene and at least one molecular diagnosis associated with said at least one variant; and Using patient scores and the genotype data, a Bayesian causal model is configured to generate and output a variant score, wherein the variant score includes a machine learning likelihood that the variant is pathogenic to the genetic condition.