META machine-learning models for generating improved pathogenicity scores

By employing meta machine-learning models that integrate pathogenicity scores from multiple sources with additional biological data, the challenges of inconsistent and inaccurate pathogenicity predictions in existing models are addressed, resulting in enhanced accuracy and consistency across clinical benchmarks.

WO2025137642A1PCT designated stage expired Publication Date: 2025-06-26ILLUMINA INC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/061574
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-22
Filing Date
2024-12-20
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing pathogenicity prediction models lack cross-benchmark consistency and accuracy due to noisy training data, insufficient ground-truth data, and variations in training data, leading to inconsistent and unreliable pathogenicity scores.

Method used

The development of meta machine-learning models that combine pathogenicity scores from different models, utilizing inputs such as protein structural data, conservation profiles, and protein annotations to generate refined pathogenicity scores and pathogenicity probabilities, thereby improving accuracy and consistency.

Benefits of technology

The proposed meta machine-learning models achieve improved accuracy and precision in predicting pathogenicity, providing consistent results across various clinical benchmarks and protocols, and more accurately identifying variants associated with cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024061574_26062025_PF_FP_ABST
    Figure US2024061574_26062025_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure describes methods, non-transitory-computer readable media, and systems that can combine pathogenicity scores generated by different variant pathogenicity machine-learning models for amino-acid variants at particular protein positions to determine meta pathogenicity scores that estimate the degree to which amino-acid variants are benign or pathogenic at particular protein positions. In particular, the disclosed systems can train or utilize a meta variant pathogenicity machine-learning model that processes multiple inputs, including pathogenicity scores generated by different variant pathogenicity machine-learning models, to generate refined pathogenicity scores that outperform scores of other pathogenicity prediction models. The disclosed systems can also train or utilize a meta pathogenicity probability machine-learning model that maps pathogenicity scores from variant pathogenicity machine-learning models to predict pathogenicity probabilities.
Need to check novelty before this filing date? Find Prior Art

Description

META MACHINE-LEARNING MODELS FOR GENERATING IMPROVED PATHOGENICITY SCORESCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 614,363, entitled, “META MACHINE-LEARNING MODELS FOR GENERATING IMPROVED PATHOGENICITY SCORES,” filed on December 22, 2023 (IP-2633 -PRV), which is incorporated herein by reference in its entirety.BACKGROUND

[0002] In recent years, biotechnology firms and research institutions have improved software for predicting a pathogenicity of protein or genetic variants. For instance, some existing pathogenicity prediction models generate predictions that estimate a degree to which amino-acid variants are benign or pathogenic. Such pathogenicity predictions can indicate whether an aminoacid variant is likely to cause various diseases, such as certain cancers, developmental disorders, or heart conditions. In addition to the intrinsic predictive value of such predictions, biotechnology firms and research institutions have developed downstream applications for pathogenicity predictions. For instance, pathogenicity predictions output by machine-learning models have been used to identify target variants in a population subset for new drugs as well as target variants that may be the subject of genetic editing.

[0003] While existing pathogenicity prediction models have demonstrated significant improvements in accuracy and downstream applications, existing models do not consistently generate accurate predictions across a range of different clinical benchmarks and cell-line protocols. Such clinical benchmarks and cell-line protocols may include, for instance, scores for protein variants or benign proteins in data from the Deciphering Developmental Disorders (DDD) study, the United Kingdom (UK) Biobank, cell-line experiments for Saturation Mutagenesis, Clinical Variant (ClinVar) from the National Library of Medicine, and Genomics England Variants (GELVar). While certain pathogenicity prediction models generate predictions that accurately indicate pathogenicity for variants in the UK Biobank, for instance, the same models do not accurately predict pathogenicity for certain between-protein benchmarks from DDD.

[0004] Because of differences in coverage among databases identifying clinical or other effects of amino-acid variants, many exiting databases assign variants unknown pathogenicity labels. In particular, human proteins contain hundreds of millions of potential amino-acid variants. However, because of a relatively low overall mutation rate in humans and because relatively few humans’ genomes have been sequenced, existing databases include only tens of thousands of high-quality experimental pathogenicity labels for those variants. For example, ClinVar, a publicly accessible database providing information about genetic variants, contains pathogenicity labels for only afraction of potential amino-acid variants in human proteins. Furthermore, existing databases often suffer from a scarcity of high signal data, as existing databases are either predominately populated with low-quality data or a limited amount of sufficiently high-quality data. For example, some training datasets have a disproportionate amount of information about variants in a small portion of high value or well-studied genes. Thus, variance in confidence levels is not only based on evolutionary constraints, as designed, but also on variations in training data. Primate benign labels disagree or conflict with ClinVar benign labels at around 4%, for instance, thereby providing noisy data under well-accepted standards for training machine-learning-models. Because many existing databases lack high-quality signal data for training purposes and contain some unknown pathogenicity labels, existing pathogenicity prediction models often lack sufficient ground-truth data upon which to train such models to generate accurate or complete pathogenicity predictions for various amino-acid variants.

[0005] Given that existing databases include noisy training data and various unlabeled aminoacid variants — thereby providing insufficient ground-truth data — existing pathogenicity prediction models often rely on different networks and methods to generate raw pathogenicity scores that cannot be directly compared to each other without additional processing and complicate interpretation of such scores. The choice of network or method affects how each pathogenicity prediction model processes and interprets input data. Beyond network or method choices, existing pathogenicity prediction models have experimented with and implemented different inputs and / or training data to generate predictions for diverse amino-acid variants. Adding further to the diversity of existing models, pathogenicity prediction models may be trained to predict pathogenicity for specific types of variants or a subset of variants, which leads to differences in performance between pathogenicity prediction models across variant types. Due to these and other variations, pathogenicity prediction models generate raw pathogenicity scores having unique pathogenicity score distributions. Some existing systems attempt to directly compare raw pathogenicity scores across methods by monotonically transforming the raw pathogenicity scores into a uniform score distribution or individually determining model-specific raw pathogenicity score thresholds above or below which the scores might indicate variants to be pathogenic or benign. However, transforming the raw pathogenicity scores and determining model-specific raw pathogenicity score thresholds can be computationally expensive. For example, both methods often require existing systems to process and store raw pathogenicity scores from each examined method. Consequently, raw pathogenicity scores from different pathogenicity prediction models often cannot be directly compared to one another or well leveraged to cover predictions for different types amino-acid variants (e.g., corresponding to different genes) without additional processing. Even when rawpathogenicity scores can be monotonically transformed, the raw pathogenicity scores generated by existing pathogenicity prediction models lack a consistent physical interpretation.

[0006] To address the lack of cross-benchmark consistency, more complex pathogenicity prediction models have been developed in the form of transformer machine-learning models with (i) self-attention mechanisms that process sequential input data and (ii) an ensemble of different pathogenicity prediction models that together generate combined or refined predictions. While such transformers have developed highly accurate pathogenicity predictions, in some cases, the transformers can consume considerable computer processing to generate predictions. To train either transformer or various individual models as part of an ensemble of pathogenicity prediction models, servers and other computing devices can likewise consume considerable computer processing and time. By adding further layers to the architecture of such transformers or additional models, existing models may improve accuracy but likewise further increase computer processing.

[0007] Some existing transformer machine-learning models have been trained to generate pathogenicity predictions. Such transformer machine-learning models often fail to generate pathogenicity predictions with accuracy across a range of different clinical benchmarks and cellline protocols. For example, some state-of-the-art transformer machine-learning models analyze information from multiple potential amino-acid variants at each protein position to output scores for the potential amino-acid variants. However, such transformer machine-learning models have proven impractical or difficult to train with cross-clinical-benchmark accuracy. More specifically, ground-truth training data often exhibits gaps for certain protein positions or amino-acid variants at a given protein position. The deficiencies in ground-truth training data introduce training signal noise or bias upon which existing transformer machine-learning models rely. Variations in training data introduce additional signal noise that confounds signals of interest. Because training data may contain different amounts of information, existing transformer machine-learning models generate pathogenicity predictions with varying confidence levels. For example, a transformer machinelearning model may be trained to score all 20 candidate amino-acid variants at a protein position. Combining the information for all 20 candidate amino-acid variants can create training signal noise that negatively impacts the performance of the existing transformer machine-learning model.

[0008] These, along with additional problems and issues exist in existing sequencing systems.SUMMARY

[0009] This disclosure describes one or more embodiments of systems, methods, and non- transitory computer readable storage media that solve one or more of the problems described above or provide other advantages over the art. In particular, the disclosed systems can combine pathogenicity scores generated by different variant pathogenicity machine-learning models for amino-acid variants at particular protein positions to determine meta pathogenicity scores (e.g.,refined pathogenicity scores or combined pathogenicity probabilities) that estimate the degree to which the amino-acid variants are benign or pathogenic at the particular protein positions.

[0010] As described further below, the disclosed systems can train or utilize a meta variant pathogenicity machine-learning model (e.g., a meta variant pathogenicity classifier) that processes multiple inputs to determine refined pathogenicity scores. In particular, the meta variant pathogenicity machine-learning model can process one or more of (i) pathogenicity scores generated by different variant pathogenicity machine-learning models for amino-acid variants at a target position of a protein, (ii) protein structural data for the protein, (iii) protein annotations indicating a benign-ness or pathogenicity of variants within a window of the target position, or (iv) a conservation profile of conserved amino acids at particular positions of the protein. By processing one or more of the above-listed inputs, the disclosed meta variant pathogenicity machine-learning model generates refined pathogenicity scores that outperform scores from other pathogenicity prediction models.

[0011] Independent of such a meta variant pathogenicity machine-learning model, in some implementations, the disclosed systems can train or utilize a meta pathogenicity probability machine-learning model that (i) includes pathogenicity probability models that respectively map pathogenicity scores from specific variant pathogenicity machine-learning models to pathogenicity probabilities that a target amino acid has been determined (e.g., as reflected in a clinical variant database) to be a benign or pathogenic variant and (ii) combines the pathogenicity probabilities into meta pathogenicity scores. More specifically, the disclosed systems may utilize model-specific probability models (e.g., model-specific classifiers) to map pathogenicity scores from variant pathogenicity machine-learning models to predict pathogenicity probabilities. The model-specific probability models map pathogenicity scores to pathogenicity probabilities from specific variant pathogenicity machine-learning models. The pathogenicity probability models thus generate pathogenicity probabilities that are comparable (e.g., standardized) across variant pathogenicity machine-learning models.

[0012] Additional features and advantages of one or more embodiments of the present disclosure will be set forth in the description which follows.BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The detailed description refers to the drawings briefly described below.

[0014] FIG. 1 illustrates a schematic diagram of a computing system in which a variant pathogenicity prediction system can operate in accordance with one or more embodiments of the present disclosure.

[0015] FIGS. 2A-2B illustrate the variant pathogenicity prediction system utilizing and training a meta variant pathogenicity machine-learning model to generate refined pathogenicity scores in accordance with one or more embodiments of the present disclosure.

[0016] FIG. 3 illustrates the variant pathogenicity prediction system utilizing positive-and- unlabeled (PU) learning as part of training the meta variant pathogenicity machine-learning model in accordance with one or more embodiments of the present disclosure.

[0017] FIGS. 4A-4C illustrate a meta variant pathogenicity machine-learning model demonstrating improvements in performance relative to existing variant pathogenicity machinelearning models in accordance with one or more embodiments of the present disclosure.

[0018] FIG. 5 illustrates the variant pathogenicity prediction system utilizing an all-data transformer neural network in accordance with one or more embodiments of the present disclosure.

[0019] FIG. 6 illustrates a comparison of score distributions against ClinVar benign and pathogenic labels in accordance with one or more embodiments of the present disclosure.

[0020] FIGS. 7A-7B illustrate the variant pathogenicity prediction system creating an overall distribution of model scores by combining benign histograms and pathogenic histograms in accordance with one or more embodiments of the present disclosure.

[0021] FIG. 8 illustrates a logic behind the variant pathogenicity prediction system training a pathogenicity probability model in accordance with one or more embodiments of the present disclosure.

[0022] FIGS. 9A-9B illustrate the variant pathogenicity prediction system utilizing and training a meta pathogenicity probability machine-learning model in accordance with one or more embodiments of the present disclosure.

[0023] FIG. 10 illustrates the meta variant pathogenicity machine-learning model and the meta pathogenicity probability machine-learning model demonstrating improvements in performance relative to existing variant pathogenicity machine-learning models in accordance with one or more embodiments of the present disclosure.

[0024] FIG. 11 illustrates a flowchart of a series of acts for generating a refined pathogenicity score in accordance with one or more embodiments of the present disclosure.

[0025] FIG. 12 illustrates a flowchart of a series of acts for determining a combined pathogenicity probability in accordance with one or more embodiments of the present disclosure.

[0026] FIG. 13 illustrates a block diagram of an example computing device in accordance with one or more embodiments of the present disclosure.DETAILED DESCRIPTION

[0027] This disclosure describes one or more embodiments of a variant pathogenicity prediction system that can utilize one or both of (a) a meta variant pathogenicity machine-learningmodel and (b) a meta pathogenicity probability machine-learning model. As an overview, in some cases, the variant pathogenicity prediction system uses a meta variant pathogenicity machinelearning model to combine pathogenicity scores from different machine-learning models and determine refined pathogenicity scores that estimate a degree to which amino-acid variants are benign or pathogenic at particular protein positions. By contrast, in other cases, the variant pathogenicity prediction system uses a meta pathogenicity probability machine-learning model that includes multiple pathogenicity probability models that respectively map pathogenicity scores from specific machine-learning models to pathogenicity probabilities and combines the pathogenicity probabilities at particular protein positions into meta pathogenicity scores.

[0028] To further illustrate, in some implementations of (a), the variant pathogenicity prediction system trains or utilizes a meta variant pathogenicity machine-learning model to combine pathogenicity scores from different variant pathogenicity machine-learning models to generate a refined pathogenicity score. In particular, the variant pathogenicity prediction system feeds one or more of the following inputs into a meta variant pathogenicity machine-learning model: (i) pathogenicity scores generated by various variant pathogenicity machine-earning models for a target amino-acid variant at a target position within a protein, (ii) protein structural data indicating densities among amino acids within the protein, (iv) protein annotations indicating one or more variants at the target protein position are benign or pathogenic within a population, or (iv) a conservation profile comprising data representing conservation of amino acids at protein positions of the protein from different species. Based on one or more of the inputs (i) - (v), the meta variant pathogenicity machine-learning model generates a refined pathogenicity score for the target aminoacid variant at the target protein position. As exhibiting by various figures described below, the refined pathogenicity score is more accurate than pathogenicity scores generated by existing variant pathogenicity machine-learning models.

[0029] To effectively process such inputs and improve accuracy, in certain implementations, the variant pathogenicity prediction system trains the meta variant pathogenicity machine-learning model to generate such refined pathogenicity scores. In some embodiments, for instance, the variant pathogenicity prediction system employs an expectation-maximization (EM) algorithm that trains the meta variant pathogenicity machine-learning model to generate both refined pathogenicity scores and propensity scores. In contrast to some existing models, in some embodiments, the meta variant pathogenicity machine-learning model can thus generate a dual output for each amino-acid variant. While one type of output of the meta variant pathogenicity machine-learning model takes the form of pathogenicity scores during training, another type of output from the meta variant pathogenicity machine-learning model takes the form of a propensityscore indicating a probability that a variant amino acid at a target protein position is labeled benign in a training dataset given that the variant amino acid is actually benign.

[0030] As mentioned, because of the low overall mutation rate in humans and because few humans have been sequenced, many datasets include unlabeled amino-acid variants. To compensate or adjust for such unlabeled data, in some implementations, the variant pathogenicity prediction system employs Positive-Unlabeled (PU) learning to generate new training labels. For example, during the expectation stage of the EM algorithm, in certain implementations, the variant pathogenicity prediction system combines propensity scores and the pathogenicity scores with benign labels from a database. By combining the scores with the benign labels, the variant pathogenicity prediction system generates combined pathogenicity -propensity -based scores as new training labels. During the maximization stage, the variant pathogenicity prediction system compares the combined pathogenicity -propensity scores with the propensity score and the refined pathogenicity score to generate a variant propensity loss and a variant pathogenicity loss, respectively. As described further below, the variant pathogenicity prediction system utilizes the variant propensity loss to train a propensity prediction head of the meta variant pathogenicity machine-learning model and the variant pathogenicity loss to train a pathogenicity prediction head of the meta variant pathogenicity machine-learning model. Such a dual-head meta variant pathogenicity machine-learning model can take the form of a multilayer perceptron (MLP) or other machine-learning model described below.

[0031] In some implementations, the meta variant pathogenicity machine-learning model can take the form of an all-data transformer neural network to generate a refined pathogenicity score. The all-data transformer neural network may receive several inputs. For example, the all-data transformer neural network may receive pathogenicity scores, allele counts and frequencies, mutation rates, reference-residues embeddings, and other forms of input. In comparison to existing pathogenicity prediction models that receive and process information on a per-protein-position basis (e.g., scoring all 20 amino acids per protein position at the same time), the all-data transformer neural network may receive data for a single variant at a protein position and generate an improved pathogenicity score for the single variant.

[0032] In addition, or in the alternative to (a) a meta variant pathogenicity machine-learning model, in some implementations, the variant pathogenicity prediction system utilizes a meta pathogenicity probability machine-learning model (b) that includes multiple pathogenicity probability models that facilitate generating and combining pathogenicity probabilities into meta pathogenicity scores. In particular, a given pathogenicity probability model may comprise an MLP or other machine-learning architectures (e.g., random forest) that is specific to each variant pathogenicity machine-learning model. Such a pathogenicity probability model can mappathogenicity scores from a given variant pathogenicity machine-learning model for a target aminoacid variant at a target protein position to pathogenicity probabilities to generate a combined pathogenicity probability for the target amino-acid variant. For example, the variant pathogenicity prediction system leverages pathogenicity score histograms by utilizing pathogenicity probability models to fit mapping functions that map input pathogenicity scores from variant pathogenicity machine-learning models to pathogenicity probabilities that a target amino acid has been determined (e.g., as reflected in a clinical variant database) to be a benign or pathogenic variant.

[0033] To accurately fit mapping functions, in some embodiments, the variant pathogenicity prediction system trains model-specific pathogenicity probability models. More specifically, in some cases, the variant pathogenicity prediction system trains a pathogenicity probability model (e.g., a multi-layer perceptron) to map pathogenicity scores from each variant pathogenicity machine-learning model to pathogenicity probabilities. The variant pathogenicity prediction system determines pathogenicity-probability losses between a weighted average predicted pathogenicity probability from the pathogenicity probability model and respective pathogenicity probabilities output by model-specific pathogenicity probability models. The variant pathogenicity prediction system further adjusts parameters of the pathogenicity probability model based on the pathogenicity-probability loss until convergence.

[0034] As indicated above, the variant pathogenicity prediction system provides several technical advantages relative to existing pathogenicity prediction models. For example, the variant pathogenicity prediction system improves the accuracy and precision with which existing pathogenicity prediction models generate pathogenicity predictions for amino-acid variants. As noted above, existing pathogenicity prediction models generate only raw pathogenicity scores that fail to exhibit consistent accuracy across certain clinical or other benchmarks. Unlike existing pathogenicity prediction models that often require model-specific raw pathogenicity score thresholds to interpret such raw scores, the variant pathogenicity prediction system can train and utilize one or more of (a) a meta variant pathogenicity machine-learning model that generates a refined pathogenicity score based on several inputs and (b) a meta pathogenicity probability machine-learning model that maps pathogenicity scores to pathogenicity probabilities. The following paragraphs further detail how the (a) meta variant pathogenicity machine-learning model and the (b) meta pathogenicity probability machine-learning model improve accuracy and precision relative to existing pathogenicity prediction models.

[0035] As mentioned, by training and utilizing the (a) meta variant pathogenicity machinelearning model, the variant pathogenicity prediction system makes improvements to the accuracy and precision of pathogenicity predictions. For example, by utilizing at least pathogenicity scores from several variant pathogenicity machine-learning models and one or more of conservationprofiles, protein annotations, or protein structural data, the (a) meta-variant pathogenicity machinelearning model generates accurate refined pathogenicity scores relative to existing pathogenicity prediction models. As depicted and described herein, for example, the disclosed variant pathogenicity prediction system generates refined pathogenicity scores that exhibit a consistent accuracy across clinical benchmarks and protocols not exhibited by existing pathogenicity prediction models — including pathogenicity scores that accurately predict a pathogenicity for target amino acids across the Deciphering Developmental Disorders (DDD) study, the United Kingdom (UK) Biobank, Saturation Mutagenesis, a Clinical Variant (ClinVar), and Genomics England Variants (GELVar). Furthermore, the disclosed variant pathogenicity prediction system more accurately predicts variants that cause cancer relative to existing pathogenicity prediction models. As further depicted and demonstrated by various charts and results illustrated in FIGS. 4A-4C, in some cases, such refined pathogenicity scores exhibit better performance relative to unrefined pathogenicity scores in each of the foregoing benchmarks and protocols.

[0036] Additionally, the variant pathogenicity prediction system improves the accuracy of pathogenicity predictions by utilizing PU learning as part of training the (a) meta variant pathogenicity machine-learning model. The meta variant pathogenicity machine-learning model represents a first-of-its-kind model that employs parameters tuned by both propensity scores and propensity loss. By combining a pathogenicity score and a propensity score with benign labels to create combined pathogenicity-propensity scores as training labels, the variant pathogenicity prediction system can more accurately account for unlabeled data from incomplete ground-truth training datasets. More specifically, the variant pathogenicity prediction system utilizes benign labels that have better coverage of the whole genome than existing databases. FIG. 3 below illustrates performance of PU learning relative to traditional learning methods. The variant pathogenicity prediction system improves the accuracy of the new training labels by training the meta variant pathogenicity machine-learning model utilizing an expectation-maximization algorithm.

[0037] As mentioned, the variant pathogenicity prediction system also generates more accurate pathogenicity predictions by training and utilizing a (b) meta pathogenicity probability machinelearning model. The meta pathogenicity probability machine-learning model accesses pathogenicity scores generated by variant pathogenicity machine-learning models. The meta pathogenicity probability machine-learning model processes the pathogenicity scores to generate pathogenicity probabilities that a target amino acid has been determined to be a pathogenic variant or a benign variant at a target protein position. The meta pathogenicity probability machinelearning model may determine a combined pathogenicity probability based on the pathogenicity probabilities from the different variant pathogenicity machine-learning models. The metapathogenicity probability machine-learning model can also more accurately predict variants that cause cancer. As depicted and described below in relation to FIG. 10, the disclosed meta pathogenicity probability machine-learning model outperforms existing pathogenicity prediction models across various benchmarks.

[0038] The variant pathogenicity prediction system applies a meta pathogenicity probability machine-learning model in a novel and more interpretable manner than existing, stand-alone pathogenicity prediction models because the system can combine pathogenicity scores from different model types by leveraging pathogenicity probabilities, thereby broadening the utility and applicability of the disclosed meta pathogenicity probability machine-learning model. More specifically, the variant pathogenicity prediction system may utilize a (b) meta pathogenicity probability machine-learning model to map pathogenicity scores from different variant pathogenicity machine-learning models to pathogenicity probabilities. As mentioned, existing pathogenicity prediction models typically generate pathogenicity scores or predictions that cannot be directly compared to one another. By mapping pathogenicity scores to pathogenicity probabilities, the variant pathogenicity prediction system generates pathogenicity probabilities that are standardized for direct comparison.

[0039] In some embodiments, the variant pathogenicity prediction system improves the computing efficiency with which pathogenicity prediction models adjust the accuracy pathogenicity scores for amino-acid variants. As indicated above, existing pathogenicity prediction models have increased the accuracy of pathogenicity scores in part by adding neural -network layers or more complex architecture designed for deep-learning neural networks, such as transformer machine-learning models. But such additive layers or complex architecture increases both the number of operations and computer processing executed by existing pathogenicity prediction models. The variant pathogenicity prediction system provides improvements in efficiency relative to existing pathogenicity prediction models when utilizing both the (a) meta variant pathogenicity machine-learning model and the (b) meta pathogenicity probability machine-learning model. The variant pathogenicity prediction system can train the (a) meta variant pathogenicity machinelearning model to process multiple different inputs to generate more accurate refined pathogenicity scores. Additionally, rather than adding layers or more complex architecture, in some embodiments, the variant pathogenicity prediction system efficiently improves the accuracy of pathogenicity scores by training and utilizing small multilayer perceptrons (MLPs) as the (b) meta pathogenicity probability machine-learning model to map a pathogenicity score from a variant pathogenicity machine-learning model to pathogenicity probabilities. More specifically, the variant pathogenicity prediction system trains and utilizes model-specific pathogenicity probabilitymodels to quickly and simply generate accurate combined pathogenicity probabilities without more complex neural -network layers.

[0040] As illustrated by the foregoing discussion, the present disclosure utilizes a variety of terms to describe features and advantages of the variant pathogenicity prediction system. As used herein, for example, the term “machine-learning model” refers to a computer algorithm or a collection of computer algorithms that automatically improve for a particular task through experience based on use of data. For example, a machine-learning model can utilize one or more learning techniques to improve in accuracy and / or effectiveness. Example machine-learning models include various types of decision trees (e.g., gradient boosted trees), support vector machines, Bayesian networks, or neural networks (e.g., transformer neural networks, recurrent neural networks, triangle attention neural networks).

[0041] In some cases, the variant pathogenicity prediction system uses a variant pathogenicity machine-learning model to generate, modify, or update a pathogenicity score for a target amino acid. As used herein, the term “variant pathogenicity machine-learning model” refers to a machinelearning model that generates a pathogenicity score for either a protein (e.g., protein variant) or an amino acid at a particular protein position of a protein. For example, a variant pathogenicity machine-learning model includes a machine-learning model that generates an initial pathogenicity score for a variant amino acid at a target protein position within a protein based on an amino-acid sequence for the protein. In addition to or as part of an amino-acid sequence for the protein as an input, in some cases, a variant pathogenicity machine-learning model processes other inputs, such as a multiple sequence alignment (MSA) corresponding to the protein or a reference amino-acid sequence for the protein. As indicated below, a variant pathogenicity machine-learning model can take the form of different models, including, but not limited to, a transformer machine-learning model, a convolutional neural network (CNN), a sequence-to-sequence model, a variational autoencoder (VAE), a multilayer perceptron (MLP), a recurrent neural network (RNN), a long short-term memory (LSTM), or a decision tree model.

[0042] Relatedly, as used herein, the term “pathogenicity score” refers to a measurement, numerical value, or score indicating a degree to which a protein or an amino acid at a protein position within a protein is benign or pathogenic. In some cases, for example, a pathogenicity score includes a logit or other numerical value indicating a probability of a variant amino acid at a target protein position of a protein relative to a reference amino acid at the target protein. Because a pathogenicity score can indicate a particular amino acid in a protein position is benign, in some cases, a pathogenicity score represents a fitness of the particular amino acid in the protein position. As but one example of a pathogenicity score, in some embodiments, the pathogenicity score for a target alternative amino acid (sait) at a target protein position includes a numerical value determinedfrom a usual difference of a logit for an alternative amino acid (pait) and a logit for a reference amino acid (pref) at the target protein position. More details concerning this specific example can be found in U.S. Patent Application No. 17 / 975,547, titled “Pathogenicity Language Model,” by Tobias Hamp, Anastasia Dietrich, Yibing Wu, Jeffrey Ede, and Kai-How Farh, filed on October 27, 2022, which is hereby incorporated in its entirety by reference. Other formulations of a pathogenicity score, however, can likewise be used and are described below.

[0043] As used herein, the term “propensity score” refers to a measurement, numerical value, or score indicating a probability that the target protein position is labeled benign in a dataset (e.g., training dataset) given a ground truth that the target amino acid is benign. For example, because different variants have different sampling probabilities, some protein variants may be labeled as unknown or benign. Accordingly, the propensity score indicates a probability that if a variant is indeed benign that it is labeled as benign in the training dataset.

[0044] As mentioned, the variant pathogenicity prediction system may utilize a meta variant pathogenicity machine-learning model to generate a refined pathogenicity score based on outputs from various machine-learning models. As used herein, the term “meta variant pathogenicity machine-learning model” refers to a machine-learning model that generates refined pathogenicity scores for either a protein (e.g., protein variant) or an amino acid at a particular protein position of a protein. A meta variant pathogenicity machine-learning model includes a machine-learning model that takes the outputs or predictions of multiple machine-learning models to generate a refined pathogenicity score. A meta variant pathogenicity machine-learning model can process different types of inputs such as pathogenicity scores from variant pathogenicity machine-learning models, conservation profiles, protein annotations, protein structural data, and other inputs. In some examples, a meta variant pathogenicity machine-learning model takes the form of a multilayer perceptron (MLP). In other examples, the meta variant pathogenicity machine-learning model takes the form of different models including, but not limited to, a CNN, a sequence-to- sequence model, a VAE, an RNN, an LSTM, or a decision tree model.

[0045] As suggested above, the term “refined pathogenicity score” refers to a pathogenicity score that has been adjusted or modified to classify the pathogenicity of protein variants. In particular, a refined pathogenicity score can be generated by a meta variant pathogenicity machinelearning model that combines scores from various pathogenicity predictors and one or more of protein structural data, conservation signals, or protein annotations. In some examples, a refined pathogenicity score can classify the pathogenicity of protein or amino acid variants that are labeled unknown as well as amino acid variants that are labelled as benign. A refined pathogenicity score can be a more-accurate indicator of whether a target amino acid variant is benign or pathogenic.

[0046] As mentioned, the meta variant pathogenicity machine-learning model receives various inputs. For example, the meta variant pathogenicity machine-learning model may receive a conservation profde as an input. As used herein, the term “conservation profile” comprises a representation of the degree of conservation or variability of a specific amino acid at a protein position within a protein. In particular, a conservation profile includes data for three or more amino-acid sequences from different species for a same given protein. For example, a conservation profile may comprise data representing a multiple sequence alignment (MSA) or a condensed version of an MSA for a given protein from multiple related species.

[0047] Additionally, the variant pathogenicity prediction system may input protein structural data into a meta variant pathogenicity machine-learning model. As used herein, the term “protein structural data” refers to data describing the three-dimensional (3D) arrangement of atoms and amino acids within a protein molecule. In particular, protein structural data provides a detailed representation of a protein’s spatial organization. Protein structural data can indicate a density of amino acids within a protein. Protein structural data may comprise density values indicating a an atom-number-based density or a mass-number-based density of amino acids within a protein. For example, protein structural data may comprise alpha carbon (Ccr) atom number densities, density of mass, or other densities. In some embodiments, protein structural data comprises structural data predicted based on the protein’s amino acid sequence. For example, protein structural data may comprise an AlphaFold protein structure that indicates if a part of a protein is an alpha helix, beta sheet, etc.

[0048] As used herein, the term “protein annotation” refers to data indicating a likelihood or proportion of a protein or its constituent amino acids (e.g., amino-acid variants at particular positions) (i) being present in a population or database or (ii) having a particular label in a population or database. A protein annotation can include a protein-position-based annotation or a protein-frequency-based annotation. In particular, the term protein annotation can comprise data files that include functional annotations. But, for purposes of clarity in this disclosure and to differentiate the term “protein structural data,” the term “protein annotation” is not used in this disclosure as a representation of a protein’s structure, which is addressed above. For instance, protein-position-based annotations or protein-frequency -based annotations may indicate a number of observed benign labels in a target protein position. Additionally, protein-position-based annotations or protein-frequency-based annotations can indicate a number of predicted benign labels in the target protein position. In some examples, protein-position-based annotations or protein-frequency-based annotations comprise an observed-expected ratio of observed benign labels for amino acids within a threshold number of adjacent protein positions of the target protein position compared to a total number of expected benign labels for amino acids within the thresholdnumber of adjacent protein positions. An observed-expected ratio can also or alternatively include, but is not limited to, an observed proportion of missense variants to an expected proportion relative to a number of possible variants in a defined codon window or other ratios described by Kathie Y. Sun et al., “A deep catalog of protein-coding variation in 985,830 individuals,” bioRxiv 2023.05.09.539329, available at doi: https: / / doi.org / 10.1101 / 2023.05.09.539329 (hereinafter, Sun). In some implementations, protein-position-based annotations or protein-frequency-based annotations comprise allele frequencies or the frequencies of a specific allele within a population’s gene pool.

[0049] As described further below, the variant pathogenicity prediction system may utilize benign labels as part of training the meta variant pathogenicity machine-learning model. As used herein, the term “benign label” refers to a classification label assigned to a protein or amino acid at a target protein position that is considered benign. The benign label can exist within a training dataset. In particular, a benign label is assigned to an amino acid unlikely to cause a disease in an organism (e.g., to a high degree of confidence or with a high degree of certainty). A benign label can be applied to an amino acid at a target protein position within a protein not known to cause a disease in a human or primate. The benign label can be benign more than 95% of the time (e.g., 95.8%) based on primate data. Relatedly, an amino acid may be associated with an unknown label. An unknown label indicates a protein or amino acid at a target position within a protein for which it is unknown whether the particular type of amino acid causes a disease in a human or other primate.

[0050] As used herein, the term “pathogenicity probability model” refers to a probability model that generates a pathogenicity probability. In particular, a pathogenicity probability model is trained to generate a pathogenicity probability based on a pathogenicity score from a variant pathogenicity machine-learning model. For example, a pathogenicity probability model may comprise a multilayer perceptron (MLP) specific to a variant pathogenicity machine-learning model. The pathogenicity probability model can map pathogenicity scores to pathogenicity probabilities based on a clinical variant database.

[0051] As used herein, the term “pathogenicity probability” refers to a likelihood or probability that a protein or amino acid variant is associated with a disease or pathological condition. In particular, a pathogenicity probability indicates that a target amino acid has been determined to be a pathogenic variant or a benign variant at a target protein position. For example, a pathogenicity probability includes a probability that a database (e.g., a clinical variant database or a primate variant database) comprises a benign label or a pathogenic label for a target amino acid at a target protein position. As mentioned previously, a pathogenicity probability can be generated by apathogenicity probability model based on a pathogenicity score from a variant pathogenicity machine-learning model.

[0052] As used herein, the term “combined pathogenicity probability” refers to a likelihood or probability that a protein or amino acid at a protein position within a protein is benign or pathogenic based on pathogenicity probabilities from two or more variant pathogenicity machine learning models. In particular, the combined pathogenicity probability may comprise a pathogenicity probability determined by mapping a pathogenicity score from a variant pathogenicity machine learning model to one or more pathogenicity probabilities based on a clinical variant database. The variant pathogenicity prediction system can generate combined pathogenicity probabilities that are directly comparable to each other, even if they arise from pathogenicity scores from different variant pathogenicity machine-learning models.

[0053] As used herein, the term “clinical variant database” refers to a repository that comprises information about genetic, protein, and / or amino acid variants and their clinical significance. A clinical variant database can include variant information that includes details about the locations and characteristics of amino acid variants at protein positions. Furthermore, in some examples, a clinical variant database includes assessments of the clinical significance of amino acid or protein variants. For instance, a clinical variant database may indicate whether an amino acid variant is benign or pathogenic and whether a variant is associated with a specific disease or condition. For example, a clinical variant database may include ClinVar or Humsavar (HVAR).

[0054] As further used herein, the term “target amino acid” refers to a particular type of amino acid. In particular, a target amino acid includes a particular alternate or variant residue within an amino-acid sequence corresponding to a protein. As just indicated, a target amino acid may include a particular alternate or variant residue at a target protein position within an amino-acid sequence. A target amino acid may likewise be any of 20 amino acids that are part of a protein associated with an organism (e.g., humans), such as, alanine, arginine, asparagine, aspartic acid, cysteine, etc. In other embodiments, however, a target amino acid may include any of more than 20 amino acids (e.g., 22 amino acids).

[0055] Relatedly, as used herein, the term “target protein position” refers to a particular location or order for an amino acid within an amino-acid sequence forming a polypeptide chain for a protein. In particular, a target protein position includes a numerically identified location for an amino acid in an ordered amino-acid sequence representing a protein. For example, a target protein position could include a seventh, fifty-fourth, one hundred and ninety-fifth, two hundredth, or any numbered position within an amino-acid sequence of amino acids (e.g., 300-amino acid sequence) representing a protein. In some cases, a target protein position can be represented as a number along or within a residue sequence index (e.g., depicted in accompanying figures).

[0056] The following paragraphs describe the variant pathogenicity prediction system with respect to illustrative figures that portray example embodiments and implementations. For example, FIG. 1 illustrates a schematic diagram of a computing system 100 in which a variant pathogenicity prediction system 104 operates in accordance with one or more embodiments. As illustrated, the computing system 100 includes one or more server device(s) 102 connected to a client device 110 and therapeutics analysis device(s) 114 via a network 116. While FIG. 1 shows an embodiment of the variant pathogenicity prediction system 104, this disclosure describes alternative embodiments and configurations below.

[0057] As shown in FIG. 1, the server device(s) 102, the client device 110, and the therapeutics analysis device(s) 114 are connected via the network 116. Accordingly, each of the components of the computing system 100 can communicate via the network 116. The network 116 comprises any suitable network over which computing devices can communicate. Example networks are discussed in additional detail below with respect to FIG. 13.

[0058] As indicated by FIG. 1, the therapeutics analysis device(s) 114 comprises a device for analyzing (and identifying candidate therapeutics for) amino-acid sequences corresponding to proteins and / or nucleotide sequences representing coding and non-coding genomic regions. In some embodiments, the therapeutics analysis device(s) 114 analyzes a set of amino-acid sequences or a set of nucleotide sequences from a database comprising samples exhibiting genetic diversity. From among the analyzed set of amino-acid sequences and / or analyzed set of nucleotide sequences, the therapeutics analysis device(s) 114 can identify subsets of amino-acid sequences and / or nucleotide sequences exhibiting common variant amino acids or variant nucleotides. In combination with or separate from such variant identification, the therapeutics analysis device(s) 114 can execute machine-learning models (or other models) that identify coding or non-coding genomic regions that are intolerant to variation and for which variants can cause loss or change in biological functions. For identified subsets of amino-acid sequences and / or nucleotide sequences, in some cases, the therapeutics analysis device(s) 114 identifies candidate biologies, drugs, or geneediting protocols for treatment.

[0059] In addition, or in the alternative to communicating across the network 116, in some embodiments, the therapeutics analysis device(s) 114 bypasses the network 116 and communicates directly with the server device(s) 102 or the client device 110. Additionally, as shown in FIG. 1, in one or more embodiments, the therapeutics analysis device(s) 114 includes the variant pathogenicity prediction system 104.

[0060] As further indicated by FIG. 1, the server device(s) 102 may generate, receive, analyze, store, and transmit digital data, such as data for amino-acid sequences or nucleotide sequences. As shown in FIG. 1, the therapeutics analysis device(s) 114 may send (and the server device(s) 102may receive) various data from the therapeutics analysis device(s) 114, including data representing amino-acid sequences or nucleotide sequences. The server device(s) 102 may also communicate with the client device 110. In particular, the server device(s) 102 can send data representing aminoacid sequences or nucleotide sequences (or variants thereof), pathogenicity scores, or propensity scores, to the client device 110.

[0061] Additionally, as shown in FIG. 1, the server device(s) 102 can include the variant pathogenicity prediction system 104. The variant pathogenicity prediction system 104 may utilize various models to generate pathogenicity scores and / or pathogenicity probabilities. For example, and as illustrated, the variant pathogenicity prediction system 104 includes one or more of a meta variant pathogenicity machine-learning model 106 and / or a meta pathogenicity probability machine-learning model 118. In some embodiments, the meta variant pathogenicity machinelearning model 106 an all-data transformer neural network 108. In one or more embodiments, as explained further below, the variant pathogenicity prediction system 104 combines pathogenicity scores from different models and determines refined pathogenicity scores by utilizing the meta variant pathogenicity machine-learning model 106. In some cases, the variant pathogenicity prediction system 104 can use an all-data transformer neural network 108 to utilize information about potential variants per position and across positions to inform pathogenicity probabilities. Furthermore, in some implementations, the variant pathogenicity prediction system 104 utilizes the meta pathogenicity probability machine-learning model 118 to compare score distributions from different models against benign and pathogenic labels to predict pathogenicity probabilities for scores from those different models. In addition to generating pathogenicity probabilities, the server device(s) 102 can also send data representing pathogenicity probabilities and / or propensity scores to the client device 110 for graphical visualization. The figures depicted herein and paragraphs below further illustrate such functionalities of the variant pathogenicity prediction system 104, the meta variant pathogenicity machine-learning model 106, the all-data transformer neural network 108, and the meta pathogenicity probability machine-learning model 118.

[0062] In addition or in the alternative to executing one or more of the meta variant pathogenicity machine-learning model 106, the all-data transformer neural network 108, or the meta pathogenicity probability machine-learning model 118, in some embodiments, the variant pathogenicity prediction system 104 accesses a database or table comprising pathogenicity scores. For example, in certain embodiments, the variant pathogenicity prediction system 104 may access pathogenicity scores from variant pathogenicity machine-learning models. The variant pathogenicity prediction system 104 may utilize the meta variant pathogenicity machine-learning model 106 (or the all-data transformer neural network 108) to process the pathogenicity scores to generate refined pathogenicity scores. Additionally, or alternatively, the variant pathogenicityprediction system 104 utilizes a meta pathogenicity probability machine-learning model 118 to map the pathogenicity scores to pathogenicity probabilities to generate combined pathogenicity scores. Consistent with the disclosure above and below, the table or database includes pathogenicity scores that have been precomputed from a combination of machine-learning models. For example, the table or database includes pathogenicity scores generated by the meta variant pathogenicity machine-learning model 106, the all-data transformer neural network 108, the meta pathogenicity probability machine-learning model 118, or another machine-learning model. These machinelearning models may comprise a transformer machine-learning model, a convolutional neural network (CNN), a variational autoencoder (VAE), a multilayer perceptron (MLP), a recurrent neural network (RNN), a long short-term memory (LSTM), or a decision tree model trained to generate pathogenicity scores.

[0063] In some embodiments, the server device(s) 102 comprise a distributed collection of servers where the server device(s) 102 include a number of server devices distributed across the network 116 and located in the same or different physical locations. Further, the server device(s) 102 can comprise a content server, an application server, a communication server, a web-hosting server, or another type of server.

[0064] In some cases, the server device(s) 102 is located at or near a same physical location of the therapeutics analysis device(s) 114 or remotely from the therapeutics analysis device(s) 114. Indeed, in some embodiments, the server device(s) 102 and the therapeutics analysis device(s) 114 are integrated into a same computing device. The server device(s) 102 may run software on the therapeutics analysis device(s) 114 or the variant pathogenicity prediction system 104 to generate, receive, analyze, store, and transmit digital data, such as by sending or receiving data representing amino-acid sequences or nucleotide sequences (or variants thereof), pathogenicity scores, or propensity scores. Additionally or alternatively, in some embodiments, the therapeutics analysis device(s) 114 or the variant pathogenicity prediction system 104 store and access a database or table of pathogenicity scores corresponding to particular proteins and / or protein positions.

[0065] As further illustrated and indicated in FIG. 1, the client device 110 can generate, store, receive, and send digital data. In particular, the client device 110 can receive data for amino-acid sequences or nucleotide sequences (or variants thereof), or pathogenicity scores from the server device(s) 102 and / or the therapeutics analysis device(s) 114. The client device 110 can accordingly present data concerning pathogenicity scores within a graphical user interface to a user associated with the client device 110.

[0066] The client device 110 illustrated in FIG. 1 may comprise various types of client devices. For example, in some embodiments, the client device 110 includes non-mobile devices, such as desktop computers or servers, or other types of client devices. In yet other embodiments, the clientdevice 110 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones. Additional details regarding the client device 110 are discussed below with respect to FIG. 13.

[0067] As further illustrated in FIG. 1, the client device 110 includes an analytics application 112. The analytics application 112 may be a web application or a native application stored and executed on the client device 110 (e.g., a mobile application, desktop application). The analytics application 112 can include instructions that, when executed, cause the client device 110 to receive data from the variant pathogenicity prediction system 104 and present data from the therapeutics analysis device(s) 114 and / or the server device(s) 102. Furthermore, the analytics application 112 can instruct the client device 110 to display data for pathogenicity scores, such as data for graphical visualization of pathogenicity by protein position for a two-dimensional or three-dimensional representation of a protein.

[0068] As further illustrated in FIG. 1, the variant pathogenicity prediction system 104 may be located on the client device 110 as part of the analytics application 112 or on the therapeutics analysis device(s) 114. Accordingly, in some embodiments, the variant pathogenicity prediction system 104 is implemented by (e.g., located entirely or in part) on the client device 110. As mentioned, in yet other embodiments, the variant pathogenicity prediction system 104 is implemented by one or more other components of the computing system 100, such as the therapeutics analysis device(s) 114. In particular, the variant pathogenicity prediction system 104 can be implemented in a variety of different ways across the server device(s) 102, the network 116, the client device 110, and the therapeutics analysis device(s) 114.

[0069] Though FIG. 1 illustrates the components of the computing system 100 communicating via the network 116, in certain implementations, the components of computing system 100 can also communicate directly with each other, bypassing the network. For instance, and as previously mentioned, in some implementations, the client device 110 communicates directly with the therapeutics analysis device(s) 114. Additionally, in some embodiments, the client device 110 communicates directly with the variant pathogenicity prediction system 104. Moreover, the variant pathogenicity prediction system 104 can access one or more databases housed on or accessed by the server device(s) 102 or elsewhere in the computing system 100.

[0070] As described previously, the variant pathogenicity prediction system 104 may utilize a (a) meta variant pathogenicity machine-learning model to generate refined pathogenicity scores and a (b) meta pathogenicity probability machine-learning model to generate combined pathogenicity scores. The following figures and paragraphs further detail the training, utilization, and performance improvements of the (a) meta variant pathogenicity machine-learning model and the (b) meta pathogenicity probability machine-learning model. More specifically, FIGS. 2A-6 depict details relating to a meta variant pathogenicity model in accordance with one or moreimplementations. FIGS. 6-10 and the corresponding discussion provide additional details regarding the meta pathogenicity probability machine-learning model in accordance with one or more embodiments.

[0071] As indicated above, the variant pathogenicity prediction system 104 generates refined pathogenicity scores for target amino acids at target protein positions. In accordance with one or more embodiments, FIG. 2A depicts an overview of the variant pathogenicity prediction system 104 generating such a refined pathogenicity score. FIG. 2B depicts an overview of the variant pathogenicity prediction system 104 training a meta variant pathogenicity machine-learning model to generate a refined pathogenicity score in accordance with one or more implementations of the present disclosure.

[0072] As shown in FIG. 2A, the variant pathogenicity prediction system 104 runs a meta variant pathogenicity machine-learning model 212 to generate a refined pathogenicity score 216 for a target amino acid at a target position within a protein. More specifically, the variant pathogenicity prediction system 104 can process multiple inputs to determine the refined pathogenicity score 216. As shown in FIG. 2A, the meta variant pathogenicity machine-learning model 212 processes one or more of pathogenicity scores 204 from various models of the variant pathogenicity machine-learning models 202. Additionally, the meta variant pathogenicity machine-learning model 212 can also receive, as input, a conservation profile 206 for the protein, protein annotations 208 indicating benign-ness of variants within a window of the target position, and protein structural data 210 indicating amino acid densities of the protein. Based on these inputs, the meta variant pathogenicity machine-learning model 212 generates the refined pathogenicity score 216 for the variant amino acid at the target position. The meta variant pathogenicity machinelearning model 212 performs well even when small numbers of inputs are removed. Thus, in some implementations, the variant pathogenicity prediction system 104 can input various combinations of two or more of the conservation profile 206, the protein annotations 208, or the protein structural data 210 into the meta variant pathogenicity machine-learning model 212.

[0073] As further shown in FIG. 2A, the variant pathogenicity prediction system 104 utilizes the variant pathogenicity machine-learning models 202 to generate, modify, or update the pathogenicity scores 204. The variant pathogenicity machine-learning models 202 may take the form of different models trained to generate pathogenicity scores 204. For example, the variant pathogenicity machine-learning models 202 may comprise a transformer machine-learning model, a convolutional neural network (CNN), a variational autoencoder (VAE), a multilayer perceptron (MLP), a recurrent neural network (RNN), a long short-term memory (LSTM), or a decision tree model. For example, the variant pathogenicity machine-learning models 202 may comprisePrimateAI3D, a Triangle Attention (TriAttn) neural network, an MSA transformer, or other type of pathogenicity prediction machine-learning model.

[0074] As shown in FIG. 2A, the meta variant pathogenicity machine-learning model 212 may receive, as input, the pathogenicity scores 204 for a target amino acid at a protein position within a protein. As indicated above, a pathogenicity score of the pathogenicity scores 204 indicates a degree to which the target amino acid is benign or pathogenic to an organism when located at the target protein position within a protein. As shown in FIG. 2A, the pathogenicity scores 204 can be generated by multiple, different models of the variant pathogenicity machine-learning models 202. In some cases, the variant pathogenicity machine-learning models 202 can generate pathogenicity scores for other target amino acids at the same or different target protein positions within the protein.

[0075] As indicated above, the pathogenicity scores 204 tend to exhibit inconsistent accuracy across different benchmarks. As a result, the pathogenicity scores 204 may not accurately reflect the pathogenicity of a target amino acid. More specifically, the pathogenicity scores 204 output by the variant pathogenicity machine-learning models 202 may not be accurate due to uncertainty of the variant pathogenicity machine-learning models 202 themselves or due to limitations of the data input into the conservation profile 206. In one example, pathogenicity scores 204 can be inaccurate because of variants having different sampling probabilities. Furthermore, some of the pathogenicity scores 204 may be inaccurate because of incomplete input data. Many variants are labeled as unknown as neither being pathogenic nor benign because of incomplete data. For instance, many databases include labels for tens of thousands of variants out of hundreds of millions of potential variants. Furthermore, and as previously mentioned, many variants that are labelled as benign are incorrectly labelled. For example, approximately 4% of variants that have been labelled as benign are incorrectly labelled because of noisy data.

[0076] The variant pathogenicity prediction system 104 may utilize the conservation profde 206 as additional input into the meta variant pathogenicity machine-learning model 212. The conservation profile 206 comprises data representing a multiple sequence alignment (MSA) or a condensed version of an MSA for a given protein from multiple species. For example, the conservation profile 206 includes data for three or more amino-acid sequences from different species for a same given protein. The different species may include 50, 100, 150, or other suitable number of related species, such as 100 vertebrate species.

[0077] In certain embodiments, the conservation profile 206 comprises or is input into the meta variant pathogenicity machine-learning model 212 with learned weights for each species. As indicated, in some embodiments, the conservation profile 206 comprises data representing a condensed version of such an MSA with learned weights. To condense an MSA, in someembodiments, the variant pathogenicity prediction system 104 identifies or determines, for each protein position in a given protein, a number of times each amino acid from (i) the twenty candidate amino acids occurs in the species (e.g., 100 species) and (ii) a gap token representing a position at which an aligned, non-human amino-acid sequence does not include a residue that aligns with a human amino-acid sequence, and divide the number of occurrences for each amino acid by the number of species (e.g., 100). Because of (i) the twenty candidate amino acids and (ii) the one gap token, in some embodiments, the conservation profile 206 accounts for twenty-one candidate values per position and include values that are proportional to each amino acid in an MSA column at a given position. In a condensed version, consequently, the conservation profile 206 comprises values indicating a probability of each amino acid at particular protein positions across related specific for a given protein. Accordingly, in some embodiments, the conservation profile 206 constitutes or comes in the form of a position weight matrix (PWM), a position-specific weight matrix (PSWM), or a position-specific scoring matrix (PSSM) derived from a MSA corresponding to the protein and includes an alignment of amino-acid sequences from different species (e.g., a conservation MSA for a group of primates).

[0078] As further shown in FIG. 2A, the variant pathogenicity prediction system 104 utilizes the protein annotations 208 as input into the meta variant pathogenicity machine-learning model 212. Generally, the protein annotations 208 indicate a likelihood or proportion of a protein or its constituent amino acids (e.g., amino-acid variants at particular positions) (i) being present in a population or database or (ii) having a particular label in a population or database. The protein annotations 208 can comprise an indication of benign labels at or near a protein position within a population. For example, in some implementations, the protein annotations 208 comprise an observed-expected ratio of benign labels for amino acids within a threshold number of adjacent protein positions of the target protein position compared to a total number of expected benign labels for amino acids within the threshold number of adjacent protein positions. More specifically, the variant pathogenicity prediction system 104 identifies a region or window surrounding the target protein position to compute a function of the protein annotations. In some implementations, the variant pathogenicity prediction system 104 determines the target protein position by identifying a threshold number of adjacent protein positions of a target protein position.

[0079] The variant pathogenicity prediction system 104 can identify a target protein position and determine an observed expected ratio of benign labels corresponding to one or more variants at the target protein position. In one example, the variant pathogenicity prediction system 104 determines an observed number of benign labels corresponding to one or more variants at the target protein position. The variant pathogenicity prediction system 104 further determines an expected number of benign labels for variants at the target protein position. The variant pathogenicityprediction system 104 may determine the expected number of benign labels based on animal allele frequency data (e.g., human or primate). The variant pathogenicity prediction system 104 can generate an observed-expected ratio of benign labels by dividing the observed number of benign labels by a sum of the observed number of benign labels and the expected number of benign labels. Additionally or alternatively, an observed-expected ratio can include, but is not limited to, an observed proportion of missense variants to an expected proportion relative to a number of possible variants in a defined codon window or other ratios described by Sun.

[0080] Additionally, and as mentioned, the variant pathogenicity prediction system 104 may use the protein structural data 210 as input into the meta variant pathogenicity machine-learning model 212. Generally, the protein structural data 210 comprises protein structural data around a given protein position. Protein structural data may indicate protein density, or relative location of an amino acid within a protein. For example, in some implementations, the protein structural data 210 comprises the density of amino acid numbers within the protein structure based on AlphaFold predictions. In some examples, amino acid number densities are computed from AlphaFold predicted protein structures. Local amino acid density comprises the number density of amino acids at an amino acid alpha carbon (Ca) position. Local amino acid density is used to describe the concentration or clustering of neighboring amino acids around a specific Ca atom within a protein structure. The variant pathogenicity prediction system 104 may determine the local amino acid density by determining distances (r between an amino acid Ca position and other amino acid Ca atoms. The variant pathogenicity prediction system 104 computes the densities using the following equations:

[0081] Where the sum is across all the other Ca distances. Each of the local amino acid densities is normalized by subtracting the mean, dividing by the standard deviation, and sigmoidalizing the densities.

[0082] AlphaFold predictions are further described by John Jumper et al., “Highly Accurate Protein Structure Prediction with AlphaFold,” 596 Nature 583-589 (2021) (hereinafter Jumper), and the corresponding supplementary information by John Jumper et al., “Supplementary Information for: Highly Accurate Protein Structure Prediction with AlphaFold,” both of which are hereby incorporated by reference in their entirety.

[0083] As shown, the meta variant pathogenicity machine-learning model 212 analyzes one or more of the inputs described above to generate a refined pathogenicity score 216. In some implementations, the meta variant pathogenicity machine-learning model 212 generates the refined pathogenicity score 216 using only the pathogenicity scores 204 as input. In other embodiments,the meta variant pathogenicity machine-learning model 212 generates the refined pathogenicity score 216 based on both the pathogenicity scores 204 and the protein structural data 210 as input.

[0084] As shown, the meta variant pathogenicity machine-learning model 212 generates the refined pathogenicity score 216. The refined pathogenicity score 216 for a target amino acid at a protein position within a protein is more accurate than the pathogenicity scores 204 for the same target amino acid at the same protein position within a protein. By utilizing data from the various inputs, the meta variant pathogenicity machine-learning model 212 generates the refined pathogenicity score 216 that classifies the pathogenicity of unknown protein variants. More specifically, the refined pathogenicity score indicates a degree to which a protein or amino acid at a protein position within a protein is benign or pathogenic.

[0085] In one or more implementations, the meta variant pathogenicity machine-learning model 212 generates a propensity score. As described previously, the propensity score comprises a measure indicating a probability that a target protein position is labeled benign in a training dataset given a ground truth that the target amino acid is benign. The variant pathogenicity prediction system 104 primarily utilizes a propensity score to tune a propensity prediction head of the meta variant pathogenicity machine-learning model 212 during training. In some implementations, the variant pathogenicity prediction system 104 may also utilize the meta variant pathogenicity machine-learning model 212 to generate the propensity score as an output together with a refined pathogenicity score. For instance, the propensity score may indicate the accuracy of the training dataset in identifying and labeling variants at a target protein position as benign.

[0086] As mentioned, the variant pathogenicity prediction system 104 trains the meta variant pathogenicity machine-learning model 212. FIG. 2B illustrates the variant pathogenicity prediction system 104 training the meta variant pathogenicity machine-learning model 212 to generate refined pathogenicity scores in accordance with one or more embodiments of the present disclosure. The variant pathogenicity prediction system 104 utilizes the meta variant pathogenicity machinelearning model 212 to generate a propensity score 220 and a refined pathogenicity score 222. The propensity score 220 indicates a probability that the target amino acid at the target protein position is labeled benign in a training dataset given a ground truth that the target amino acid is benign. The refined pathogenicity score classifies the pathogenicity of a protein or amino acid variant at the target protein position. The variant pathogenicity prediction system 104 combines the propensity score 220 and the refined pathogenicity score 222 with benign labels 230 to generate a combined pathogenicity-propensity scores 224 as a new training label. The variant pathogenicity prediction system 104 utilizes separate loss functions for propensity and pathogenicity to adjust parameters of the meta variant pathogenicity machine-learning model 212.

[0087] The meta variant pathogenicity machine-learning model 212 illustrated in FIGS. 2A- 2B may comprise one or two machine-learning models trained to generate the propensity score 220 and the refined pathogenicity score 222. In some embodiments, the meta variant pathogenicity machine-learning model 212 comprises a single machine-learning model with two prediction heads. A propensity prediction head can be utilized to predict the propensity score 220, and a pathogenicity prediction head can be utilized to predict the refined pathogenicity score 222. In another example, the meta variant pathogenicity machine-learning model 212 comprises two machine-learning models in which a propensity machine-learning model generates the propensity score 220 and a pathogenicity machine-learning model generates the refined pathogenicity score 222.

[0088] As part of training the meta variant pathogenicity machine-learning model 212, the variant pathogenicity prediction system 104 inputs one or more of pathogenicity scores, conservation profiles, protein annotations, and protein structural data into the meta variant pathogenicity machine-learning model 212. Based on the inputs, the meta variant pathogenicity machine-learning model 212 generates the propensity score 220. As indicated previously, the propensity score 220 indicates a probability that a variant amino acid at a target protein position is labeled benign in a training dataset given that the variant amino acid is actually benign. The variant pathogenicity prediction system 104 can utilize the propensity score 220 to create new training labels as part of training the meta variant pathogenicity machine-learning model 212.

[0089] As further shown in FIG. 2B, the variant pathogenicity prediction system 104 utilizes the meta variant pathogenicity machine-learning model 212 to generate the refined pathogenicity score 222. As mentioned previously, the refined pathogenicity score 222 comprises a score indicating a degree to which a protein or an amino acid at a protein position within a protein is benign or pathogenic. More specifically, the refined pathogenicity score 222 is more accurate than the input pathogenicity scores processed by the meta variant pathogenicity machine-learning model 212.

[0090] In some implementations, the variant pathogenicity prediction system 104 combines the propensity score 220 and the refined pathogenicity score 222 with benign labels 230. As mentioned previously, benign labels exist within a training dataset and comprise a classification label assigned to a protein or amino acid at a target protein position that is considered benign. Relatedly, an amino acid at a protein position may be associated with an unknown label that indicates that a protein or amino acid at a protein position for which it is unknown whether the given protein or amino acid causes a disease in a human or other primate.

[0091] In some implementations, the benign labels 230 comprise primate benign labels. Primate benign labels often have better coverage and comprise higher signal data than existingdatabases. In some embodiments, primate benign labels can be marginally more noisy than existing database benign labels. For example, primate benign labels may be at a ~4% disagreement with existing database levels (e.g., ClinVar benign labels) and may not be suitable for typical training thresholds. However, a similar difference exists between existing database benign labels. To illustrate, ClinVar 1+ star labels and HumsaVar labels often have a ~3% disagreement due to clinician / research error. While primate benign labels may be noisier than existing datasets, primate benign labels tend to be more extensive and therefore have better coverage of the entire genome. Thus, the use of primate benign labels as the benign labels 230 can improve the accuracy of generating pathogenicity predictions.

[0092] As mentioned, the variant pathogenicity prediction system 104 combines the propensity score 220 and the refined pathogenicity score 222 with the benign labels 230 to create combined pathogenicity-propensity scores 224 as new training labels. Some training datasets (e.g., ClinVar) contain a limited number of labels relative to an absolute number of potential variants. For instance, a training dataset may have tens of thousands of labels that classify only a fraction of 60-70 million actual potential human variants. Thus, training datasets may assign benign labels or pathogenic labels to only a small fraction of actual potential human variants. For example, pathogenic variants often have unknown labels as individuals having pathogenic variants are rarely observed in life. Training datasets may include tens of thousands of benign labels and / or pathogenic labels while tens of millions of the variants are benign in reality. Accordingly, existing pathogenicity prediction systems that rely solely on training datasets may be inaccurate because of an underrepresentation of both benign and pathogenic labels.

[0093] As mentioned, the variant pathogenicity prediction system 104 combines the propensity score 220 and the refined pathogenicity score 222 with the benign labels 230 to generate the combined pathogenicity-propensity scores 224. The combined pathogenicity-propensity scores 224 estimate how benign unknown variants are to create an updated set of training labels. More specifically, the variant pathogenicity prediction system 104 utilizes the benign labels 230 to identify variants that are labeled “unknown” in the training dataset. By combining the propensity score 220 and the refined pathogenicity score 222, the variant pathogenicity prediction system 104 can estimate how benign the unknown variants are. The variant pathogenicity prediction system 104 utilizes the combined pathogenicity -propensity scores 224 as an updated set of training labels.

[0094] The variant pathogenicity prediction system 104 utilizes the combined pathogenicitypropensity scores 224 to train the meta variant pathogenicity machine-learning model 212. More specifically, in some embodiments, the variant pathogenicity prediction system 104 utilizes the combined pathogenicity -propensity scores 224 to train both the propensity prediction head and the pathogenicity prediction head of a single model, that is, the meta variant pathogenicity machine-learning model 212. Additionally, or alternatively, the variant pathogenicity prediction system 104 utilizes the combined pathogenicity-propensity scores 224 to train the pathogenicity machinelearning model and the propensity machine-learning model.

[0095] The variant pathogenicity prediction system 104 generates a variant propensity loss 226 as part of training the propensity prediction head and / or the propensity machine-learning model of the meta variant pathogenicity machine-learning model 212. As part of determining the variant propensity loss 226, for example, the variant pathogenicity prediction system 104 determines a Positive-Unlabeled (PU) learning propensity loss and a mutation rate rank loss. The variant pathogenicity prediction system 104 generates the PU learning propensity loss by comparing the propensity score 220 and the combined pathogenicity -propensity scores 224.

[0096] As further shown in FIG. 2B, the variant pathogenicity prediction system 104 further generates the variant propensity loss 226 by determining a mutation rate rank loss. The mutation rate rank loss utilizes mutation rate data to predict the mutation rate expected to be at a particular protein variant. By utilizing the mutation rate rank loss, the variant pathogenicity prediction system 104 ensures that the propensity score 220 is consistent with the mutation rate data. For example, higher mutation rates typically correspond with higher variant likelihoods and, accordingly, higher likelihood that variants at the target protein position are benign. Because the probability that a variant at the target protein position is benign, the variant pathogenicity prediction system 104 can also imply that the propensity score 220 is higher.

[0097] When using Python syntax, for instance, the variant pathogenicity prediction system 104 uses propensity scores represented by e_logit_for_bootstrap = model_output[ “propensity: ][mutation_rate_mask] and mutation rates represented by bootstrap e logit = (model_target[ “train_annotations”][ . . . , 60:80])[ mutation rate mask] where mutation rate mask is used to select values where mutation rate info is available. For example mutation rate info can be available for amino acids that are accessible in a single nucleotide change. The propensity scores (e logit for bootstrap) and the mutation rates (bootstrap e logit) can be fed into the following rank loss computation function: else: bootstrap loss = selfget rank loss ( e logit for bootstrap, bootstrap e logit, gap=selfargs.propensity_model_guidance_diff, ratio=True,) where the input can be represented as bootstrap_loss = y_hat [:,None] * y_hat[None,:] * bootstrap_loss or bootstrap loss = bootstrap_loss.mean( ) according to Python syntax. The losses [“mutation rate loss”] = bootstrap loss. In some implementations, self. args. propensity _model_guidance_diff = 0. The following function can be used to compute the rank loss: def get_rank_loss (self, x_select, y_select, gap=0.0, ratio=False): dx_select = x_select [:, None] - x_select [None,:] dx = F.relu(dx_select). If ratio and gap: dy = (y_select [: ,None] / y_select [None, :]) < gap, else: dy = (y_select [:, None] + gap) < y_select [None,: ] rank_loss = dx*dy. The output of this function is the rank loss (rank loss).

[0098] As further shown in FIG. 2B, the variant pathogenicity prediction system 104 generates a variant pathogenicity loss 228. The variant pathogenicity prediction system 104 utilizes the variant pathogenicity loss 228 to train the pathogenicity prediction head and / or the pathogenicity machine-learning model of the meta variant pathogenicity machine-learning model 212. As shown, the variant pathogenicity prediction system 104 generates the variant pathogenicity loss 228 by determining a Positive-and-Unlabeled (PU) learning pathogenicity loss and a pathogenicity composite rank loss. The variant pathogenicity prediction system 104 determines the PU learning pathogenicity loss by comparing the refined pathogenicity score 222 with the combined pathogenicity-propensity scores 224.

[0099] As part of determining the variant pathogenicity loss 228, as indicated by FIG. 2B, the variant pathogenicity prediction system 104 determines a pathogenicity composite rank loss. The variant pathogenicity prediction system 104 generates the pathogenicity composite rank loss by combining scores from variant pathogenicity machine-learning models. In particular, the variant pathogenicity prediction system 104 may add the scores from the variant pathogenicity machinelearning models and use the added scores as a rank cluster. For example, to determine the rank cluster, in some implementations, unsupervised scores can be added to create new scores. The variant pathogenicity prediction system 104 can compare the new pathogenicity scores with the refined pathogenicity scores during training to compute the rank loss. In some embodiments, the variant pathogenicity prediction system 104 adds the pathogenicity composite rank loss to the PU learning pathogenicity loss to adjust parameters of the meta variant pathogenicity machine-learning model 212.

[0100] In some implementations, the variant pathogenicity machine-learning models comprise unsupervised machine-learning models that predict pathogenicity of protein variants. Variant pathogenicity machine-learning models can include, but are not limited to, an unsupervised MSA transformer, DeepSequence, ESM-2, Global Epistatic Model for predicting Mutational Effects (GEMME), and / or ESM-lv.

[0101] In some embodiments, the variant pathogenicity prediction system 104 utilizes mask revelation MSA transformer scores, which achieve higher performance as described by H. Gao et al., “The landscape of tolerated genetic variation in humans and primates,” Science 380, eabn8153 (2023), https: / / www.science.org / doi / 10.1126 / science.abn8197, which is hereby incorporated by reference in its entirety. The MSA transformer comprises an unsupervised protein sequence language model which uses an MSA of a query sequence instead of a single amino acid sequence as described by Yiyu Hong et al., “S-Pred: protein structural property prediction usingMSA transformer,” Sci Rep 12, 13891 (2022), which is hereby incorporated by reference in its entirety. DeepSequence comprises a probabilistic model for sequence families and can predict the effects of variants as described by Adam J. Riesselman et al., “Deep Generative Models of Genetic Variation Capture the Effects of Mutations,” 15 Nat. Methods 816-822 (2018) (hereinafter Riesselman), which is hereby incorporated by reference in its entirety. ESM-2 comprises a language model of protein sequences that can predict pathogenicity described by Xiangling Lu, et al., “Protein Language Model Predicts Mutation Pathogenicity and Clinical Prognosis,” bioRxiv 2022.09.30.510294; doi: https: / / doi.org / 10.1101 / 2022.09.30.510294 (2022), which is hereby incorporated by reference in its entirety.

[0102] GEMME comprises a method that predicts mutational outcomes by modeling the evolutionary history of natural sequences as described by Elodie Laine, et al., “GEMME: A Simple and Fast Global Epistatic Model Predicting Mutational Effects,” Mol Biol Evol. 36(11): 2604-2619 (2019), which is hereby incorporated by reference in its entirety. ESM-lv comprises a protein language model as described by Joshua Meier et al., “Language models enable zero-shot prediction of the effects of mutations on protein function,” Advances in Neural Information Processing Systems 34 (2021), which is hereby incorporated by reference in its entirety.

[0103] In some implementations, the variant pathogenicity prediction system 104 trains the meta variant pathogenicity machine-learning model 212 by utilizing an Expectation-Maximization (EM) algorithm. An EM algorithm is a general iterative optimization algorithm that can be applied to PU learning to train parameters of the PU model and make predictions on unlabeled data. More specifically, during the expectation stage of the EM algorithm, the variant pathogenicity prediction system 104 combines the propensity score 220 and the refined pathogenicity score 222 with the benign labels 230 to generate new training labels in the form of the combined pathogenicitypropensity scores 224. During the maximization stage of the EM algorithm, the variant pathogenicity prediction system 104 determines the variant propensity loss 226 and the variant pathogenicity loss 228. The variant pathogenicity prediction system 104 utilizes an adjusted version ofthe meta variant pathogenicity machine-learning model 212 to estimate a new propensity score and a new refined pathogenicity score. The variant pathogenicity prediction system 104 runs the EM algorithm until the refined pathogenicity scores for variants converge.

[0104] The utilization of PU learning in training the meta variant pathogenicity machinelearning model 212 provides several benefits relative to traditional learning methods. Utilization of PU learning can also reduce complexity relative to existing systems. In one example, PU learning reduces complexity by obviating the need for mutation rates because PU learning estimates propensity scores directly from data. PU learning does not rely on an explicit method, such as balancing the number of benign variants and unknown variants for each mutation type, tocompensate for varying mutation rates for different amino acid transitions. Instead, PU learning corrects for mutation rates from the data itself. While PU learning does not require the utilization of mutation data, in some implementations, mutation rates can be utilized as part of determining a rank loss for the meta variant pathogenicity machine-learning model 212 to improve performance.

[0105] Furthermore, PU learning can incorporate experimental data to improve the accuracy of the meta variant pathogenicity machine-learning model 212. For example, PU learning can be utilized to estimate pathogenicity labels for all unknown variants but uses benign labels where experimental data indicates a variant is benign. In some implementations, the variant pathogenicity prediction system 104 can extract experimental pathogenic labels from databases, e.g., ClinVar. In one or more embodiments, the variant pathogenicity prediction system 104 can generate experimental benign labels and / or experimental pathogenicity labels by utilizing unsupervised model scores being above certain thresholds.

[0106] As mentioned, the variant pathogenicity prediction system 104 utilizes PU learning as part of training the meta variant pathogenicity machine-learning model. PU learning provides benefits relative to traditional learning methods across several benchmarks. FIG. 3 illustrates the variant pathogenicity prediction system 104 utilizing PU learning as part of training the meta variant pathogenicity machine-learning model in accordance with one or more embodiments. FIG. 3 illustrates a chart 300 demonstrating the performance of PU learning compared to traditional learning approaches. More specifically, traditional approaches comprise handcrafted mutation rates for different possible trinucleotide transitions and other variations. Within traditional approaches, mutation rates are individually determined. The propensity is determined based on the mutation rates.

[0107] The chart 300 illustrated in FIG. 3 includes the following benchmarks: UK Bio Bank (UKBB), Assay, -logl0(DDD p val), and Local ClinVar area under the curve (AUC). As shown in the chart 300, despite being a completely unsupervised deep learning approach, PU learning achieves competitive performance in accuracy to handcrafted mutation rates in the UKBB and Assay benchmarks. Furthermore, when trained using PU learning, the meta variant pathogenicity machine-learning model is more accurate for Deciphering Developmental Disorders (DDD) database and Local ClinVar AUC relative to traditional learning methods.

[0108] As shown in FIG. 3, a meta variant pathogenicity machine-learning model trained by PU learning performs competitively in UKBB and Assay benchmarks relative to a meta variant pathogenicity machine-learning model trained using traditional learning methods. Traditional learning methods are typically resource intensive and supervised. More particularly, traditional learning methods rely on hand-crafted mutation rates for different possible mutations such as trinucleotide transitions. In the traditional learning methods depicted in FIG. 3, a user hasdetermined different mutation rates based on a predicted likelihood of given variants. For instance, more likely or common variants are more likely to be mutated and, thus, more likely to be benign. More specifically, when employing PU learning, the meta variant pathogenicity machine-learning model is competitive with hand-crafted mutation rates in accurately identifying pathogenic aminoacid variants associated with particular phenotypes represented in UKBB and in related assays. A Spearman’s rank correlation was determined for the |R| value for UKBB and assay.

[0109] As further indicated by the chart 300, the meta variant pathogenicity machine-learning model generates refined pathogenicity scores (using PU learning) that match or outperform scores from a meta variant pathogenicity machine-learning model trained using traditional learning methods in Developmental Disorders (DDD) database and Local ClinVar AUC benchmarks. In particular, PU learning enables a meta variant machine-learning model to more accurately identify variant amino acids that cause developmental disorders from the Deciphering Developmental Disorders (DDD) database — and identify control or benign acids that do not cause such developmental disorders — when utilizing PU learning than traditional learning methods. The local area under the curve (AUC) was identified for ClinVar by determining the AUC per gene and further determining an average AUC across genes. As shown, a meta variant pathogenicity machine-learning model using PU learning outperforms traditional learning for Local ClinVar AUC.

[0110] As mentioned, the meta variant pathogenicity machine-learning model performs more accurately relative to existing machine-learning models. FIGS. 4A-4C illustrate the meta variant pathogenicity machine-learning model demonstrating improvements in performance relative to existing variant pathogenicity machine-learning models in accordance with one or more embodiments. In particular, FIG. 4A illustrates the meta variant pathogenicity machine-learning model outperforming PrimateAI3D, which utilizes a unique transformer neural network, a 3D CNN, and VAE scores, PrimateAI3D, across a number of benchmarks in accordance with one or more embodiments. FIG. 4B illustrates a chart demonstrating how the meta variant pathogenicity machine-learning model generates more accurate pathogenicity probabilities than various other variant pathogenicity machine-learning models in accordance with one or more implementations. FIG. 4C illustrates the meta variant pathogenicity machine-learning model outperforming various variant pathogenicity machine-learning models in cancer prediction in accordance with one or more embodiments.[OHl] FIG. 4A illustrates a chart 400 demonstrating the performance % of the meta variant pathogenicity machine-learning model (i.e., “meta classifier”) relative to PrimateAI3D. PrimateAI3D comprises a unique transformer neural network and is described by U.S. Patent Application No. 17 / 975,536, titled “Mask Patten for Protein Language Models,” by Tobias Hamp,Anastasia Dietrich, Yibing Wu, Jeffrey Ede, and Kai-How Farh, filed on October 27, 2022; and U.S. Patent Application No. 17 / 975,547, titled “Pathogenicity Language Model,” by Tobias Hamp, Anastasia Dietrich, Yibing Wu, Jeffrey Ede, and Kai-How Farh, filed on October 27, 2022; each of which are hereby incorporated by reference in their entirety. In some implementations, PrimateAI3D, like the meta variant pathogenicity machine-learning model, is trained to score variants accessible via a single nucleotide change.

[0112] As shown in FIG. 4A, the meta variant pathogenicity machine-learning model outperforms PrimateAI3D across different benchmarks. More specifically, the meta variant pathogenicity machine-learning model 212 generates refined pathogenicity scores that outperform pathogenicity scores from PrimateAI3D across UKBB, Assay, -log!0(DDD p val), Local Clinical Variant (ClinVar) AUC, Global ClinVar AUC, and -loglO(GELVar p_val). As shown by the chart 400, the meta variant pathogenicity machine-learning model more accurately identifies pathogenic amino-acid variants associated with particular phenotypes represented in United Kingdom (UK) Biobank (together UKBB) than PrimateAI3D. Additionally, the meta variant pathogenicity machine-learning model outperforms PrimateAI3D in a related assay. A Spearman’s rank correlation was determined for the |R| value for UKBB and assay. As indicated by the chart 400, the meta variant pathogenicity machine-learning model more accurately identified variant amino acids that cause developmental disorders from the Deciphering Developmental Disorders (DDD) database — and identify control or benign acids that do not cause such developmental disorders — better than PrimateAI3D. As further shown by the ClinVar AUC bars and the Genomics England Variants (GELVar) p-values, the meta variant pathogenicity machine-learning model 212 more accurately identifies pathogenic amino acid variants in the ClinVar GELVar databases than the pathogenicity scores of the PrimateAI3D approaches.

[0113] As mentioned, the meta variant pathogenicity machine-learning model outperforms other similar machine-learning models both within protein performance and between protein performance. FIG 4B illustrates a chart 410 demonstrating that the meta variant pathogenicity machine-learning model more accurately determines the pathogenicity of variant amino acids both between proteins and within proteins. The meta variant pathogenicity machine-learning model generates more accurate refined pathogenicity scores than various other variant pathogenicity machine-learning models. More specifically, FIG. 4B demonstrates the performance of a first and second version of the meta variant pathogenicity machine-learning model — shown as Meta Class. VI and Meta Class. V2 — compared with other variant pathogenicity machine-learning models in the chart 410. The first version of the meta variant pathogenicity machine-learning model is trained using traditional learning methods while the second version of the meta variant pathogenicity machine-learning model is trained using PU learning as described above. More specifically, insome embodiments, the traditional learning method is automated machine learning via auto-skleam as described by Feurer, Matthias, et al. “Auto-skleam 2.0: Hands-free automl via meta-leaming.” The Journal of Machine learning Research 23.1 (2022): 11936-11996, which is hereby incorporated by reference in its entirety. FIG. 4B further illustrates the performance of the second version of the meta variant pathogenicity machine-learning model when trained using noncommercial methods (Meta Class. V2: All Non-Commercial) and internal methods (Meta Class. V2 Internal. As shown, FIG. 4B illustrates the performance of the meta variant pathogenicity machine-learning model compared with performances of PrimateAI3D, Triangle Attention (Tri Attn), Multiple Sequence Alignment (MSA) transformer, DeepSequence variational autoencoder (VAE), Evolutionary model of Variant Effect (EVE) VAE, GEMME, 15B Parameter ESM2, Mammalian transformer, and MultizlOO Conservation.

[0114] As shown in the chart 410, the IB Parameter Tri Attn approach utilizes a triangle attention neural network applied to a mask revelation MSA transformer that has 1 billion parameters. Similarly, the 150M Parameter Tri Attn approach utilizes a triangle attention neural network applied to a mask revelation MSA transformer that has 150 million parameters. In some implementations, the triangle attention neural network is trained to score all variants at all protein positions within a protein. Furthermore, the meta variant pathogenicity machine-learning model outperforms EVE VAE. EVE VAE comprises a model for prediction of clinical significance of human variants based on sequences of diverse organisms across evolution as described by Jonathan Frazer, et al., “Disease variant prediction with deep generative models of evolutionary data,” Nature 599(7883), 91-95 (2021), which is incorporated by reference in its entirety.

[0115] As mentioned previously, the meta variant pathogenicity machine-learning model outperforms other variant pathogenicity machine-learning models in cancer prediction. FIG. 4C illustrates a chart 420 demonstrating the meta variant pathogenicity machine-learning model outperforming other variant pathogenicity machine-learning models in cancer prediction in accordance with one or more embodiments of the present disclosure. In particular, the meta variant pathogenicity machine-learning model more accurately identifies variants that cause cancer. FIG. 4C demonstrates the performance for cancer local AUC and cancer global AUC for various machine-learning models. More specifically, the chart 420 includes PrimateAI 3D (PAI3D), one billion parameter TriAttn (IBParamTriAttnScore), Eigen raw coding, GEMME, Computer-aaided drug discover (CADD), Polymorphism Phenotyping (Polyphen) v2, Humsavar (HVAR), Missense badness Polyphen-2 and constraint (MPC) pathogenicity classifier, Variant Effect Scoring Tool (VEST) 4, LIST-S2, rare exome variant ensemble learning (REVEL), DeepSeq, DEOGEN2, PrimateAI, SIFT, fathmm-MKL, SIFT4G, Mutation Assessor, MetaLR, LTR, DANN, MutPred, M-CAP, fathmm-MKL, fathmm-XF, and MutationTaster. As shown, the meta variantpathogenicity machine-learning model outperforms all the other plotted variant pathogenicity machine-learning models in cancer prediction.

[0116] As shown in FIG. 4C, the second version of the meta variant pathogenicity machinelearning model outperforms the next best performing models on both global and local area under the curve (AUC). Cancer local AUC and cancer global AUC can be used to evaluate various models for cancer classification and prediction. A cancer local AUC is determined by determining the AUC for a specific subgroup or region of cancers and assesses a model’s ability to discriminate between cancer and non-cancer cases variants within a particular localized context. Cancer global AUC represents the AUC calculated for an entire dataset that encompasses all cancer variants and non-cancer variants.

[0117] As illustrated in FIG. 4C, version two of the meta variant pathogenicity machinelearning model trained using noncommercial methods (MetaClassifierV2NonCommercial) outperforms the best performing models in both cancer global AUC and cancer local AUC. As mentioned previously, version two of the meta variant pathogenicity machine-learning model is trained using PU learning. Version two of the meta variant pathogenicity machine-learning model trained using internal methods (MetaClassifierV2Intemal) is the next-best performing model in both cancer global AUC and cancer local AUC. Version one of the meta variant pathogenicity machine-learning model (MetaClassifier Score) trained using traditional learning methods, PrimateAI3D, and the one billion parameter TriAttn comprise the next-best-performing models in cancer global AUC and cancer local AUC.

[0118] In some implementations, the variant pathogenicity prediction system 104 utilizes an all-data transformer neural network as a form of a meta variant pathogenicity machine-learning model to generate refined pathogenicity scores for the target amino acid. FIG. 5 illustrates the variant pathogenicity prediction system 104 utilizing the all-data transformer neural network in accordance with one or more implementations of the present disclosure. The all-data transformer neural network shares similarities with embodiments of the meta variant pathogenicity machinelearning model depicted in FIGS. 2A and 2B. But the all-data transformer neural network depicted in FIG. 5 also processes information about multiple amino-acid variants per protein position and across each position within a protein to inform predictions. Furthermore, in contrast to the meta variant pathogenicity machine-learning model, the variant pathogenicity prediction system 104 partially masks inputs into the all-data transformer neural network to add a masked language modelling objective.

[0119] As mentioned, the variant pathogenicity prediction system 104 may utilize the all-data transformer neural network to generate refined pathogenicity scores for each target amino acid (e.g., each of 20 candidate amino acids) at each protein position within a protein. In particular, thevariant pathogenicity prediction system 104 enters several inputs into an all-data transformer neural network 512. In some implementations, the variant pathogenicity prediction system 104 inputs pathogenicity scores 502, allele counts and frequencies 504, mutation rates 506, reference-residues embeddings 508, and reference amino-acid embedding 510 to the all-data transformer neural network 512. The data inputs for the all-data transformer neural network 512 comprise information for each of the 20 candidate amino acids at a target protein position. In some implementations, the all-data transformer neural network 512 generates a refined pathogenicity score 514 for each of the 20 candidate amino acids at a target protein position.

[0120] As shown in FIG. 5, the variant pathogenicity prediction system 104 utilizes pathogenicity scores 502 as input into the all-data transformer neural network 512. The pathogenicity scores 502 comprise pathogenicity scores for each of the 20 candidate amino acids at the target protein position. More specifically, the variant pathogenicity prediction system 104 accesses pathogenicity scores generated by variant pathogenicity machine-learning models for each candidate amino acid at each protein position within a protein. For example, the variant pathogenicity prediction system 104 may access a first set of pathogenicity scores from a first variant pathogenicity machine-learning model and a second set of pathogenicity scores from a second variant pathogenicity machine-learning model.

[0121] Additionally, and as shown in FIG. 5, the variant pathogenicity prediction system 104 may input allele counts and frequencies 504 into the all-data transformer neural network 512. In particular, the variant pathogenicity prediction system 104 accesses allele counts comprising the number of occurrences of specific alleles at or around a target protein position within a population. Additionally, the variant pathogenicity prediction system 104 accesses allele frequencies comprising the proportion of each allele relative to the number of alleles at a given genetic locus. The frequencies 504 may come from databases such as the genome aggregation database (gnomAD), the Trans-Omics for precision Medicine (TOPMed) program, or other databases.

[0122] As further illustrated in FIG. 5, the variant pathogenicity prediction system 104 may further input the mutation rates 506 into the all-data transformer neural network 512. The mutation rates 506 indicate observed mutation rates for an amino acid at the target protein position. In particular, the variant pathogenicity prediction system 104 may access a mutation rate with synonymous rate correction into the all-data transformer neural network 512. In particular, the mutation rates 506 may comprise the frequency at which mutations occur in a given genetic sequence or locus over a specific period of time. Synonymous mutation rates refer to the rate at which synonymous mutations occur within a gene or genetic sequence. More specifically, synonymous mutations are a type of mutation that result in a change in DNA sequence but do not lead to an amino acid change in the corresponding protein.

[0123] As further shown in FIG. 5, the variant pathogenicity prediction system 104 inputs the reference-residues embeddings 508 into the all-data transformer neural network 512. Generally, the reference-residues embeddings 508 represent reference residues for the protein. Reference residues comprise digitally represented reference amino acids assembled in a sequence for a given protein. In some examples, reference residues comprise accepted values for residues making up a canonical protein amino-acid sequence, where the individual reference residues or reference amino acids are considered to be benign. The variant pathogenicity prediction system 104 can access reference-residues embeddings for each of the candidate amino acids at the target protein position. In particular, the reference-residues embeddings 508 aim to represent each residue within a protein sequence in relation to its neighboring residues or a reference point.

[0124] FIG. 5 further illustrates the variant pathogenicity prediction system 104 inputting a reference amino-acid embedding 510 into the all-data transformer neural network 512. The reference amino-acid embedding 510 represents a reference amino-acid sequence for the protein. More specifically, for a given protein sequence, an amino acid at a protein position within the reference amino-acid sequence can be calculated and used to create an embedding. The variant pathogenicity prediction system 104 access the reference amino-acid embedding 510 for each candidate amino acid at the target protein position.

[0125] In some implementations, the variant pathogenicity prediction system 104 can input additional data into the all-data transformer neural network 512. For example, the variant pathogenicity prediction system 104 may input benign labels into the all-data transformer neural network 512. In some embodiments, the variant pathogenicity prediction system 104 inputs nonhuman primate random forest scores and unique mapper scores into the all-data transformer neural network 512. More specifically, the variant pathogenicity prediction system 104 applies thresholds to the benign labels to create primate benign labels, for example.

[0126] Additionally, in some implementations, the variant pathogenicity prediction system 104 inputs local amino acid density into the all-data transformer neural network 512. Local amino acid density comprises the number density of amino acids at an amino acid alpha carbon (Ccr) position. The variant pathogenicity prediction system 104 may utilize the local amino acid density as described above. More specifically, the variant pathogenicity prediction system 104 may utilize AlphaFold predictions to generate the amino acid density.

[0127] In some implementations, the variant pathogenicity prediction system 104 utilizes a pathogenicity probability model as a meta variant pathogenicity machine-learning model. As indicated above, FIGS. 6-9B further detail the variant pathogenicity prediction system 104 utilizing and training a meta pathogenicity probability machine-learning model comprising multiplepathogenicity probability models in accordance with one or more embodiments of the present disclosure.

[0128] As mentioned, existing pathogenicity prediction models output pathogenicity scores that have different distributions and characteristics and cannot be directly compared to one another. Different variant pathogenicity machine-learning models demonstrate different strengths and weaknesses, which are reflected in benign and pathogenic distributions of their output pathogenicity scores. As shown in FIGS. 6-8 below, pathogenicity score distributions are nonlinear and unique to each variant pathogenicity machine-learning model. The variant pathogenicity prediction system 104 can leverage the benign distributions and pathogenic distributions from different variant pathogenicity machine-learning models to estimate an overall distribution of pathogenicity scores. Using this overall distribution of pathogenicity scores, the variant pathogenicity prediction system 104 can train and utilize a meta pathogenicity probability machinelearning model to map pathogenicity scores to pathogenicity probabilities that can be directly compared to each other.

[0129] FIG. 6 illustrates a comparison of pathogenicity score distributions against ClinVar benign and pathogenic labels in accordance with one or more implementations of the present disclosure. FIG. 6 illustrates differences between pathogenicity score distributions for different variant pathogenicity machine-learning models. For example, a chart 602 and a chart 604 demonstrate pathogenicity score distributions for a 1 billion parameter transformer (IBParamTransformer). A chart 606 and a chart 608 demonstrate pathogenicity score distributions for PrimateAI 3D (PAI3D). The chart 602 and the chart 606 demonstrate pathogenicity score distributions for variant pathogenicity machine-learning models utilizing ClinVar labels. The chart 604 and the chart 608 demonstrate pathogenicity score distributions for variant pathogenicity machine-learning models models trained using additional labels.

[0130] Each of the charts 602-608 depict a benign score distribution and a pathogenic score distribution. In particular, the chart 602 depicts a benign distribution 610 and a pathogenic distribution 612. The chart 604 depicts a benign distribution 614 and a pathogenic distribution 616. The chart 606 depicts a benign distribution 618 and a pathogenic distribution 620. The chart 608 depicts a benign distribution 622 and a pathogenic distribution 624.

[0131] As further shown in FIG. 6, the peaks in the charts 602-608 have different shapes. For example, the chart 602 displaying score distributions from the IBParamTransformer depicts peaks in different places than the score distributions from PrimateAI 3D. Furthermore, the benign distribution 610 from the IBParamTransformer has a higher peak than the benign distribution 618 from PrimateAI 3D shown in the chart 606. The chart 604 and the chart 608 demonstrate similar differences between IBParamTransformer scores and PrimateAI3D scores.

[0132] As shown in FIG. 6, the distribution of pathogenicity scores output by each variant pathogenicity machine-learning model is different and has its own characteristic shape. In particular, different variant pathogenicity machine-learning models demonstrate different strengths and weaknesses. These different strengths and weaknesses affect the confidence of the variant pathogenicity machine-learning models. The variant pathogenicity prediction system 104 may leverage the benign distributions and the pathogenic distributions from different variant pathogenicity machine-learning models to estimate an overall distribution of pathogenicity scores. As shown in FIG. 6, pathogenicity score distributions from different variant pathogenicity machinelearning models are incomparable. The pathogenicity prediction system 104 can train and utilize one or more pathogenicity probability models to generate standardized pathogenicity probabilities across different variant pathogenicity machine-learning models and label datasets.

[0133] As indicated above, the variant pathogenicity prediction system 104 may determine that the overall distribution of model scores should be a combination of the benign histogram and the pathogenic histogram. As mentioned, the variant pathogenicity prediction system 104 combines benign distributions and pathogenic distributions from different variant pathogenicity machinelearning models to estimate an overall distribution of pathogenicity scores. The variant pathogenicity prediction system 104 may utilize the overall distribution of pathogenicity scores to train individual pathogenicity probability models to output standardized pathogenicity probabilities. FIGS. 7A-7B illustrate creating an overall distribution of model scores by combining benign histograms and pathogenic histograms in accordance with one or more implementations of the current disclosure.

[0134] FIG. 7A illustrates the creation of a ClinVar histogram by combining a benign histogram and a pathogenicity histogram in accordance with one or more embodiments. For example, FIG. 7A demonstrates the general shape of a ClinVar histogram 702 created by combining a benign histogram 704 with a pathogenic histogram 706. More specifically, assuming that all variants are either benign or pathogenic, the combined histogram is equal to the passing of the benign histogram 704 and the pathogenic histogram 706 together with some factor p.

[0135] FIG. 7A illustrates the shapes of combined histograms with different p weights. P weights are the relative weights of benign and pathogenic histograms. FIG. 7A illustrates a ClinVar histogram 708 with a p weight of 0.45, a ClinVar histogram 710 with a p weight of 0.65, and a ClinVar histogram 712 with a p weight of 0.85. The histograms illustrated in FIG. 7A were normalized to have cumulative probabilities equaling 1 (or a constant value) before adding them together with the p weights to create a ClinVar histogram. FIG. 7A also illustrates comparisons between ClinVar histograms and histograms for pathogenicity scores from a group of differentvariant pathogenicity machine-learning models. More specifically, FIG. 7A illustrates a grouped- scores histogram 714, a grouped-scores histogram 716, and a grouped-scores histogram 718.

[0136] Based on determining that the grouped-scores histogram 716 is most similar to the ClinVar histogram 710, the variant pathogenicity prediction system 104 determines that a p weight of 0.65 is most accurate. Accordingly, the variant pathogenicity prediction system 104 may determine that about 65% of all variants are benign. The variant pathogenicity prediction system 104 can perform this analysis on different variant pathogenicity machine-learning models.

[0137] FIG. 7B illustrates comparisons between ClinVar histograms and histograms for various variant pathogenicity machine-learning models in accordance with one or more implementations of the present disclosure. As mentioned, because benign and pathogenic labels are approximately equal across all scores, grouped-scores histograms roughly matchp=0.5 ClinVar histograms and (l-p)=0.5 pathogenic weights for score distributions. However, histograms for individual variant pathogenicity machine-learning models may deviate from a p weight of 0.5. FIG. 7B illustrates comparisons between ClinVar histograms and different variant pathogenicity machine-learning models.

[0138] FIG. 7B illustrates a chart 742 portraying a ClinVar histogram 720 compared with a IB Param MSA Transformer histogram 728. As shown in the chart 742, the p weight equals 0.35 and not 0.50. A chart 740 shows a ClinVar histogram 722 compared with a DeepSequence histogram 730. The p-value for the DeepSequence histogram equals 0.50. FIG. 7B also includes a chart 738 portraying a ClinVar histogram 724 compared with a PrimateAI3D histogram 732 where the p- value equals 0.60. FIG. 7B further illustrates a chart 736 portraying a ClinVar histogram 726 compared with a IB Param Tri Attn histogram 734 where the p-value equals 0.50.

[0139] While the histograms illustrated in FIGS. 6-7B rely on ClinVar data, the variant pathogenicity prediction system 104 can utilize labels based on other clinical variant databases. For example, the variant pathogenicity prediction system 104 may utilize labels from Humsavar. Humsavar is a database that catalogs information about single nucleotide variants and single nucleotide polymorphisms in the human genome. Humsavar also includes information relating to amino acid variants at various target protein positions. Humsavar documents the clinical significance of variants and may categorize variants as pathogenic, benign, or unlabeled.

[0140] The variant pathogenicity prediction system 104 leverages determined similarities between ClinVar histograms or other database histograms (e.g., HumsaVar) to train a pathogenicity probability model. FIG. 8 illustrates logic behind the variant pathogenicity prediction system 104 training a pathogenicity probability model in accordance with one or more implementations of the present disclosure.

[0141] As shown in FIG. 8, the variant pathogenicity prediction system 104 trains a pathogenicity probability model to map pathogenicity scores to a standard pathogenicity probability. The variant pathogenicity prediction system 104 may designate any database as a standard pathogenicity probability. For example, the variant pathogenicity prediction system 104 may designate a clinically determined database, such as ClinVar or HumsaVar as the standard pathogenicity probability. Generally, the variant pathogenicity prediction system 104 maps pathogenicity scores from a variant pathogenicity machine-learning model to the standard pathogenicity probability histogram to determine the p-value resulting in the most similar score distributions.

[0142] The variant pathogenicity prediction system 104 trains a small pathogenicity probability model for each variant pathogenicity machine-learning model. FIG. 8 illustrates a logic behind the variant pathogenicity prediction system 104 training a pathogenicity probability model for a IB Parameter Transformer in accordance with one or more implementations.

[0143] FIG. 8 illustrates a chart 806 portraying a benign distribution 802 and a pathogenic distribution 804 for the IB Parameter Transformer. The x-axis of the chart 806 comprises the pathogenicity scores output by the IB Parameter Transformer. As shown by the chart 806, as the pathogenicity scores increase, the pathogenic distribution 804 shows an increase of pathogenicity.

[0144] A plot 810 displays a relative proportion of benign and pathogenic variants from pathogenicity scores across a group of variant pathogenicity machine-learning models. The plot 810 includes a benign representation 812 and a pathogenic representation 814. As shown in FIG. 8, at the far left of the plot 810, virtually all variants across the variant pathogenicity machinelearning models are benign. The plot 810 further shows a sigmoid shape such that, as the pathogenicity scores increase, the likelihood of pathogenic variants also increases.

[0145] Instead of using the experimental curve shown in the plot 810, the variant pathogenicity prediction system 104 can train small pathogenicity probability models to interpolate the roughness and produce smooth and, in some cases, monotonic transforms to predict how pathogenicity scores from each variant pathogenicity machine-learning model correspond to pathogenicity probabilities. If the transform is not smoothed due to the small number of parameters in the pathogenicity probability model that leams the transform, it could be prone to transforming to highly fluctuating values for extreme score values where there is less data to inform the fit. For example, and as shown in FIG. 8, the variant pathogenicity prediction system 104 can train a pathogenicity probability model to generate a curve shown in a plot 808. As shown in the plot 808, pathogenicity scores from the IB Param Transformer have a non-linear relationship with pathogenicity probability. The variant pathogenicity prediction system 104 trains pathogenicity probability models for each variant pathogenicity machine-learning model to determine the exact non-linearrelationship between the pathogenicity score and pathogenicity probabilities. In some implementations, the variant pathogenicity prediction system 104 determines the relationship between pathogenicity scores and pathogenicity probabilities based on Spearman correlation, AUCs, or similar benchmarks in such a way that the shape of the output score distribution is less impactful. By mapping pathogenicity scores from different variant pathogenicity machine-learning models to pathogenicity probabilities, the variant pathogenicity prediction system 104 may directly compare pathogenicity probabilities from multiple variant pathogenicity machine-learning models.

[0146] FIGS. 9A-9B illustrate the variant pathogenicity prediction system 104 utilizing and training a meta pathogenicity probability machine-learning model, respectively, in accordance with one or more implementations of the present disclosure. FIG. 9A illustrates the variant pathogenicity prediction system 104 utilizing a meta pathogenicity probability machine-learning model comprising multiple pathogenicity probability models to generate a combined pathogenicity probability in accordance with one or more embodiments of the present disclosure. FIG. 9B illustrates the variant pathogenicity prediction system 104 adjusting parameters of the pathogenicity probability model in accordance with one or more implementations of the present disclosure.

[0147] FIG. 9A illustrates the variant pathogenicity prediction system 104 utilizing a meta pathogenicity probability machine-learning model 900 to generate combined pathogenicity probabilities 908a, 908b through 908n. As shown in FIG. 9A, the variant pathogenicity prediction system 104 utilizes variant pathogenicity machine-learning models 902a, 902b through 902n to generate a pathogenicity scores 904a, 904b through 904n. The variant pathogenicity machinelearning models 902a-902n may comprise any number of variant pathogenicity machine-learning models. For example, and as shown, the variant pathogenicity machine-learning models 902a-902n may comprise a transformer machine-learning model, a convolutional neural network (CNN), a variational autoencoder (VAE), a multilayer perceptron (MLP), a recurrent neural network (RNN), a long short-term memory (LSTM), a decision tree model, a triangle attention neural network (TriAttn), or another type of neural network. More specifically, the variant pathogenicity machinelearning models 902a-902n may comprise PrimateAI 3D. As further shown in FIG. 9A, the variant pathogenicity prediction system 104 utilizes the variant pathogenicity machine-learning models 902a-902n to generate pathogenicity scores 904a-904n. As described above, the pathogenicity scores 904a-904n indicate a degree to which a target amino acid is benign or pathogenic to an organism when located at the target protein position within a protein. Different pathogenicity scores from different variant pathogenicity machine-learning models are often incomparable given that each variant pathogenicity machine-earning model has different strengths and weaknesses. For example, a variant pathogenicity machine-learning model 902a may comprise PrimateAI 3D thatgenerates a pathogenicity score 904a. A second variant pathogenicity machine-learning model 902b may comprise a triangle attention neural network (TriAttn) that generates a second pathogenicity score 904b.

[0148] The variant pathogenicity prediction system 104 utilizes the meta pathogenicity probability machine-learning model 900 to generate a combined pathogenicity probability of the combined pathogenicity probabilities 908a-908n. As shown in FIG. 9A, the meta pathogenicity probability machine-learning model 900 comprises pathogenicity probability models 906a, 906b through 906n. As mentioned, the variant pathogenicity prediction system 104 trains a unique pathogenicity probability model corresponding to each model of the variant pathogenicity machinelearning models 902a-902n. For example, in some embodiments, the variant pathogenicity prediction system 104 trains a pathogenicity probability model 906a to process the pathogenicity score 904a from the variant pathogenicity machine-learning model 902a. Further, in some embodiments, the variant pathogenicity prediction system 104 trains a second pathogenicity probability model 906b to process the pathogenicity score 904b from the variant pathogenicity machine-learning model 902b.

[0149] In some implementations, the pathogenicity probability models 906a-906n map the pathogenicity scores 904a-904n from the variant pathogenicity machine-learning models 902a- 902n to pathogenicity probabilities. The pathogenicity probabilities indicate a likelihood that the target amino acid at the target protein position has been determined to be a pathogenic variant or a benign variant. The variant pathogenicity prediction system 104 utilizes the pathogenicity probability models 906a-906n to generate pathogenicity probabilities. The pathogenicity probability models 906a-906n may also comprise small and computationally efficient machinelearning models. For example, in some implementations, the pathogenicity probability models 906a-906n comprise multilayer perceptrons (MLPs). In other embodiments, the pathogenicity probability models may comprise other models, such as random forest; however, based on initial auto machine learning experiments, MLPs are a top performer.

[0150] Based on the pathogenicity probabilities generated by the pathogenicity probability models 906a-906n, the variant pathogenicity prediction system 104 determines combined pathogenicity probabilities 908a-908n. The combined pathogenicity probabilities 908a-908n indicate a probability or likelihood that a target amino acid has been determined to be a pathogenic variant or a benign variant at the target protein position. The pathogenicity probabilities can comprise combined pathogenicity probabilities from all variant pathogenicity machine-learning models. In some examples, the pathogenicity probabilities comprise pathogenicity probabilities from a single variant pathogenicity machine-learning model (e.g., ClinVar).

[0151] As shown in FIG. 9A, the variant pathogenicity prediction system 104 utilizes the meta pathogenicity probability machine-learning model 900 to generate the combined pathogenicity probabilities 908a-908n for a target amino acid at a target position within a protein. The combined pathogenicity probabilities 908a-908n indicate a degree to which a protein or amino acid at a protein position within a protein is benign or pathogenic. The combined pathogenicity probabilities 908a-908n corresponding to the variant pathogenicity machine-learning models 902a-902n are directly comparable to each other as they are the result of mapping the pathogenicity scores to the same pathogenicity probabilities.

[0152] In some implementations, the variant pathogenicity prediction system 104 generates the combined pathogenicity probabilities 908a-908n by determining a weighted average of the pathogenicity probabilities generated by the pathogenicity probability models 906a-906n. For example, the variant pathogenicity prediction system 104 can determine a weighted average of a first pathogenicity probability and a second pathogenicity probability to generate a combined pathogenicity probability. For example, the variant pathogenicity prediction system 104 can determine a first weight for a first pathogenicity probability based on a difference between the first pathogenicity score 904a and a first pathogenicity probability. The variant pathogenicity prediction system 104 further determines a second weight for the second pathogenicity probability based on a difference between a second pathogenicity score 904b and a second pathogenicity probability. The variant pathogenicity prediction system 104 can average the first pathogenicity probability multiplied by a first weight and the second pathogenicity probability multiplied by a second weight. The variant pathogenicity prediction system 104 may determine a weighted average of any number of pathogenicity probabilities.

[0153] FIG. 9B illustrates the variant pathogenicity prediction system 104 training the pathogenicity probability model 906a of the meta pathogenicity probability machine-learning model 900 in accordance with one or more embodiments of the present disclosure. As mentioned, the variant pathogenicity prediction system 104 trains a unique pathogenicity probability model for each variant pathogenicity machine-learning model. As shown, the variant pathogenicity prediction system 104 inputs a pathogenicity score 920 from a given variant pathogenicity machine-learning model into the pathogenicity probability model 906a corresponding to the given model. The pathogenicity probability model 906a generates a predicted pathogenicity probability 924. The predicted pathogenicity probability 924 indicates a probability that a protein or amino acid at a protein position within a protein is benign or pathogenic.

[0154] As mentioned previously, the variant pathogenicity prediction system 104 trains the pathogenicity probability model 906a to map the pathogenicity score 920 to pathogenicity probabilities. The pathogenicity probabilities may comprise a combined pathogenicity probabilityof all variant pathogenicity machine-learning models. Additionally, or alternatively, the pathogenicity probability may comprise pathogenicity probabilities from one or more variant pathogenicity machine-learning models using a single cinical variant database, e.g., ClinVar or HumsaVar.

[0155] As mentioned, the variant pathogenicity prediction system 104 trains a unique pathogenicity probability model corresponding to each variant pathogenicity machine-learning model. As shown in FIG. 9B, the variant pathogenicity prediction system 104 can determine a pathogenicity -probability loss 926 between the pathogenicity probabilities 930 and the predicted pathogenicity probability 924.

[0156] In some implementations, the variant pathogenicity prediction system 104 generates a weighted sum or average of the pathogenicity probabilities 930. In some examples, the variant pathogenicity prediction system 104 combines the pathogenicity probabilities 930 arising from each of the variant pathogenicity machine-learning models 928. For example, the variant pathogenicity prediction system 104 can add or combine losses for pathogenicity scores from all of the variant pathogenicity machine-learning models 928. More specifically, in some implementations, the pathogenicity probabilities 930 comprise outputs from the two or more variant pathogenicity machine-learning models of the variant pathogenicity machine-learning models 928 trained using ClinVar labels. The variant pathogenicity prediction system 104 trains each model-specific pathogenicity probability model to map the pathogenicity scores to the pathogenicity probabilities 930.

[0157] The variant pathogenicity prediction system 104 adjusts parameters of the pathogenicity probability model 906a by utilizing a pathogenicity-probability loss 926. For example, in some implementations, the variant pathogenicity prediction system 104 maps the predicted pathogenicity probability 924 to the pathogenicity probabilities 930. The variant pathogenicity prediction system 104 applies weights to the predicted pathogenicity probability 924. Generally, predicted pathogenicity probabilities that are nearer to pathogenicity probabilities from the pathogenicity probabilities 930 are given higher weights, and predicted pathogenicity probabilities that are farther from the pathogenicity probabilities 930 are given lower weights. In some implementations, the variant pathogenicity prediction system 104 determines distances between the predicted pathogenicity probability 924 and the pathogenicity probabilities 930 by analyzing Gaussian distances or densities. The variant pathogenicity prediction system 104 adjusts the parameters of the pathogenicity probability model 906a based on the pathogenicity -probability loss 926 until convergence. In some implementations, the variant pathogenicity prediction system 104 stops convergence criteria after a specific number of iterations or until asymptotic convergence.

[0158] As further shown in FIG. 9B, the variant pathogenicity prediction system 104 utilizes the pathogenicity-probability loss 926 to adjust parameters of the pathogenicity probability model 906a. More specifically, the variant pathogenicity prediction system 104 trains the pathogenicity probability model 906a to map pathogenicity scores to pathogenicity probabilities in a uniform manner across variant pathogenicity machine-learning models.

[0159] As mentioned previously, the variant pathogenicity prediction system 104 improves accuracy of pathogenicity predictions and scores by utilizing the meta variant pathogenicity machine-learning model and the meta pathogenicity probability machine-learning model. FIG. 10 illustrates a chart 1000 demonstrating the performance % of the meta variant pathogenicity machine-learning model (Meta Classifier V2), the meta pathogenicity probability machine-learning model (Met Classifier, Probability), and the all-data transformer neural network (all-data transformer) relative to PrimateAI3D and a Triangle Attention machine-learning model, across a number of benchmarks in accordance with one or more embodiments.

[0160] As shown in FIG. 10, the meta variant pathogenicity machine-learning model (Meta Classifier V2) outperforms PrimateAI3D across different benchmarks. FIG. 10 illustrates the performance of the second version of the meta variant pathogenicity machine-learning model that the variant pathogenicity prediction system 104 trains utilizing PU learning. As shown, the meta variant pathogenicity machine-learning model generates refined pathogenicity scores that outperform pathogenicity scores from PrimateAI3D across Assay, -logl0(DDD p val), Local Clinical Variant (ClinVar) AUC, Global ClinVar AUC, and -loglO(GELVar p val). The meta variant pathogenicity machine-learning model performs competitively with PrimateAI3D at identifying pathogenic amino-acid variants associated with particular phenotypes represented in United Kingdom (UK) Biobank (together UKBB). Additionally, the meta variant pathogenicity machine-learning model outperforms PrimateAI3D in a related assay. A Spearman’s rank correlation was determined for the R2value for UKBB and assay. As indicated by the chart 1000, the meta variant pathogenicity machine-learning model more accurately identified variant amino acids that cause developmental disorders from the Deciphering Developmental Disorders (DDD) database — and identify control or benign acids that do not cause such developmental disorders — better than PrimateAI3D. As further shown by the ClinVar AUC bars and the Genomics England Variants (GELVar) p-values, the meta variant pathogenicity machine-learning model more accurately identifies pathogenic amino acid variants in the ClinVar GELVar databases than the pathogenicity scores of the PrimateAI3D approaches.

[0161] As further shown in FIG. 10, the meta pathogenicity probability machine-learning model (Met Classifier, Probability) also outperforms PrimateAI3D across various benchmarks. The meta pathogenicity probability machine-learning model generates combined pathogenicityscores that are competitive with PrimateAI3D in identifying pathogenic amino-acid variants associated with particular phenotypes represented in UKBB. The meta pathogenicity probability machine-learning model generates refined pathogenicity scores that outperform pathogenicity scores from PrimateAI3D across Assay, -loglO(DDD p val), Local Clinical Variant (ClinVar) AUC, Global ClinVar AUC, and -loglO(GELVar p val).

[0162] FIG. 10 further illustrates the performance % of the all-data transformer neural network (all-data transformer) relative to PrimateAI3D. As shown, the all-data transformer neural network generates refined pathogenicity scores that are competitive with PrimateAI3D across the following benchmarks: UKBB, Assay, -loglO(DDD p val), Local ClinVar AUC, and Global ClinVar AUC. The all-data transformer neural network outperforms PrimateAI3D in -LoglO(GELVar P val).

[0163] FIG. 10 also illustrates the performance % of a one-billion parameter triangle attention neural network relative to PrimateAI3D across several benchmarks. As shown in FIG. 10, the triangle attention neural network underperforms PrimateAI3D across the following benchmarks: UKBB, Local ClinVar AUC, and Global ClinVar AUC. The triangle attention neural network competes with or exceeds PrimateAI3D performance in Assay, -log!0(DDD P val), and - loglO(GELVar P val).

[0164] PrimateAI3D is an ensemble of 40 models trained until convergence. In contrast, the triangle attention neural network and all data transformer scores are for ensemble size 1. The meta pathogenicity probability machine-learning model scores are ensemble size 5. FIG. 10 illustrates the performance of PrimateAI3D, the triangle attention neural network, and the meta pathogenicity probability machine-learning model without training to convergence. The performance of these models likely improves upon convergence.

[0165] FIGS. 1-10, the corresponding text, and the examples provide a number of different methods, systems, devices, and non-transitory computer-readable media of the variant pathogenicity prediction system 104. In addition to the foregoing, one or more implementations can also be described in terms of flowcharts comprising acts for accomplishing a particular result, as shown in FIGS. 11-12. FIG. 11 illustrates a flowchart of a series of acts 1100 for generating a refined pathogenicity score in accordance with one or more embodiments of the present disclosure. FIG. 12 illustrates a flowchart of a series of acts 1200 for determining a combined pathogenicity probability in accordance with one or more embodiments of the present disclosure. While FIGS. 11-12 illustrate acts according to one embodiment, alternative embodiments may omit, add to, reorder, and / or modify any of the acts shown in FIGS. 11-12. The acts of FIGS. 11-12 can be performed as part of a method. Alternatively, a non-transitory computer readable storage medium can comprise instructions that, when executed by one or more processors, cause a computing device or a system to perform the acts depicted in FIGS. 11-12. In still further embodiments, a systemcomprising at least one processor and a non-transitory computer readable medium comprising instructions that, when executed by one or more processors, cause the system to perform the acts of FIGS. 11-12.

[0166] As shown in FIG. 11, the series of acts 1100 includes an act 1110 of accessing a first pathogenicity score and a second pathogenicity score, an act 1120 of accessing protein structural data, an act 1130 of providing the first pathogenicity score, the second pathogenicity score, and the protein structural data to a meta variant pathogenicity machine-learning model, and an act 1140 of generating a refined pathogenicity score. For example, the series of acts 1100 can include acts to perform any of the operations described in the following clauses:CLAUSE 1. A method comprising: accessing, for a target amino acid at a target protein position within a protein, a first pathogenicity score generated by a first variant pathogenicity machine-learning model and a second pathogenicity score generated by a second variant pathogenicity machine-learning model; accessing protein structural data indicating a density of amino acids within the protein; providing, to a meta variant pathogenicity machine-learning model, the first pathogenicity score, the second pathogenicity score, and the protein structural data; and generate, utilizing the meta variant pathogenicity machine-learning model, a refined pathogenicity score for the target amino acid at the target protein position based on the first pathogenicity score, the second pathogenicity score, and the protein structural data.CLAUSE 2. The method of clause 1, further comprising: providing, to the meta variant pathogenicity machine-learning model, a conservation profile comprising data representing conservation of amino acids at protein positions of the protein from different species; and generating, utilizing the meta variant pathogenicity machine-learning model, the refined pathogenicity score further based on the conservation profile.CLAUSE 3. The method of clause 1, further comprising: providing, from a population database and to the meta variant pathogenicity machinelearning model, protein annotations (e.g., protein-position-based annotations or protein-frequencybased annotations) indicating one or more variants at the target protein position are benign or pathogenic within a population; and generating, utilizing the meta variant pathogenicity machine-learning model, the refined pathogenicity score further based on the protein annotations.CLAUSE 4. The method of clause 1, further comprising accessing the protein structural data by accessing density values indicating a density of the amino acids (e.g., a density of alpha carbon (Ccr) atoms, a density of mass) within the protein.CLAUSE 5. The method of clause 1, further comprising generating, utilizing the meta variant pathogenicity machine-learning model, a propensity score indicating a probability that the target amino acid at the target protein position is labeled benign in a training dataset given a ground truth that the target amino acid is benign.CLAUSE 6. The method of clause 5, further comprising: combining, as part of an expectation stage of an expectation-maximization (EM) algorithm, the refined pathogenicity score and the propensity score with benign labels to generate a combined pathogenicity-propensity score as a training label; and determining, as part of a maximization stage of the EM algorithm, a variant pathogenicity loss based on a comparison of the refined pathogenicity score and the combined pathogenicitypropensity score; and adjusting parameters of the meta variant pathogenicity machine-learning model based on the variant pathogenicity loss.CLAUSE 7. The method of clause 6, further comprising: determining, as part of a maximization stage of the EM algorithm, a variant propensity loss based on a comparison of the propensity score and the combined pathogenicity -propensity score; and adjusting parameters of the meta variant pathogenicity machine-learning model based on the variant propensity loss.CLAUSE 8. The method of clause 7, wherein the variant pathogenicity loss comprises a positive-and-unlabeled (PU) learning pathogenicity loss and the variant propensity loss comprises a PU learning propensity loss.CLAUSE 9. The method of clause 7, further comprising: determining, as part of the maximization stage of the EM algorithm, a mutation rate rank loss; applying the mutation rate rank loss to the propensity score; and combining the mutation rate rank loss with the variant propensity loss.CLAUSE 10. The method of clause 7, further comprising: iteratively determining combined pathogenicity -propensity scores as part of the expectation stage until satisfying a convergence criteria for refined pathogenicity scores generated and propensity scores generated by the meta variant pathogenicity machine-learning model; and iteratively adjusting the parameters of the meta variant pathogenicity machine-learning model as part of the maximization stage until satisfying the convergence criteria.CLAUSE 11. The method of clause 10, wherein the meta variant pathogenicity machinelearning model comprises a multilayer perceptron (MLP).CLAUSE 12. The method of clause 1, wherein the first variant pathogenicity machinelearning model or the second variant pathogenicity machine-learning model comprises a transformer machine-learning model, a convolutional neural network (CNN), a variational autoencoder (VAE), a multilayer perceptron (MLP), a recurrent neural network (RNN), a long short-term memory (LSTM), or a decision tree model.CLAUSE 13. The method of clause 1, further comprising generating, utilizing an all-data transformer neural network, the refined pathogenicity score for the target amino acid by: providing, to the meta variant pathogenicity machine-learning model: a first set of pathogenicity scores generated by the first variant pathogenicity machine-learning model for each candidate amino acid at each protein position within the protein; a second set of pathogenicity scores generated by the second variant pathogenicity machine-learning model for each candidate amino acid at each protein position within the protein; allele counts and allele frequencies for candidate amino acids in protein positions of the protein; mutation rates for the protein positions of the protein; a reference-residues embedding representing reference residues for the protein; density values indicating a density of the amino acids within the protein (e.g., a local amino-acid density); a reference amino-acid embedding representing differences between amino acids in an amino-acid sequence for the protein; and generating, utilizing the meta variant pathogenicity machine-learning model, the refined pathogenicity score further based on the first set of pathogenicity scores, the second set of pathogenicity scores, the allele counts and allele frequencies, the mutation rates, the referenceresidues embedding, and the reference amino-acid embedding.CLAUSE 14. The method of clause 1, further comprising accessing the protein annotations by accessing an observed expected ratio of observed benign labels for amino acids within a threshold number of adjacent protein positions of the target protein position compared to a total number of expected benign labels for amino acids within the threshold number of adjacent protein positions.

[0167] As shown in FIG. 12, the series of acts 1200 includes an act 1210 of accessing a first pathogenicity score and a second pathogenicity score, an act 1220 of generating a first pathogenicity probability, an act 1230 of generating a second pathogenicity probability, and an act1240 of determining a combined pathogenicity probability. For example the series of acts 1200 can include acts to perform any of the operations described in the following clauses:CLAUSE 15. A method comprising: accessing, for a target amino acid at a target protein position within a protein, a first pathogenicity score generated by a first variant pathogenicity machine-learning model and a second pathogenicity score generated by a second variant pathogenicity machine-learning model; generating, by utilizing the first pathogenicity probability model to process the first pathogenicity score, a first pathogenicity probability that the target amino acid has been determined to be a pathogenic variant or a benign variant at the target protein position; generating, by utilizing the second pathogenicity probability model to process the second pathogenicity score, a second pathogenicity probability that the target amino acid has been determined to be a pathogenic variant or a benign variant at the target protein position; and determining, based on the first pathogenicity probability and the second pathogenicity probability, a combined pathogenicity probability that the target amino acid has been determined to be a pathogenic variant or a benign variant at the target protein position.CLAUSE 16. The method of clause 15, further comprising: generating the first pathogenicity probability by generating a first probability that a clinical variant database comprises a benign label or a pathogenic label for the target amino acid at the target protein position; and generating the second pathogenicity probability by generating a second probability that the clinical variant database comprises a benign label or a pathogenic label for the target amino acid at the target protein position.CLAUSE 17. The method of clause 15, further comprising: generating the first pathogenicity probability by generating a first probability that a primate variant database comprises a benign label for the target amino acid at the target protein position; and generating the second pathogenicity probability by generating a second probability that the primate variant database comprises a benign label for the target amino acid at the target protein position.CLAUSE 18. The method of clause 15, further comprising determining the combined pathogenicity probability by determining a weighted average of the first pathogenicity probability and the second pathogenicity probability.CLAUSE 19. The method of clause 17, further comprising determining the weighted average by:determining a first weight for the first pathogenicity probability based on a difference between the first pathogenicity score and the first pathogenicity probability; determining a second weight for the second pathogenicity probability based on a difference between the second pathogenicity score and the second pathogenicity probability; and averaging the first pathogenicity probability multiplied by the first weight and the second pathogenicity probability multiplied by the second weight.CLAUSE 20. The method of clause 15, wherein the first variant pathogenicity machinelearning model or the second variant pathogenicity machine-learning model comprises a transformer machine-learning model, a convolutional neural network (CNN), a variational autoencoder (VAE), a multilayer perceptron (MLP), a recurrent neural network (RNN), a long short-term memory (LSTM), or a decision tree model.CLAUSE 21. The method of clause 15, further comprising: determining a first pathogenicity-probability loss between the combined pathogenicity probability and the first pathogenicity probability; adjusting parameters of the first pathogenicity probability model based on the first pathogenicity-probability loss; determining a second pathogenicity-probability loss between the combined pathogenicity probability and the second pathogenicity probability; adjusting parameters of the second pathogenicity probability model based on the second pathogenicity-probability loss.CLAUSE 22. The method of clause 15, wherein the first pathogenicity probability model and the second pathogenicity probability model each comprise a multilayer perceptron (MLP).

[0168] The components of the variant pathogenicity prediction system 104 can include software, hardware, or both. For example, the components of the variant pathogenicity prediction system 104 can include one or more instructions stored on a computer-readable storage medium and executable by processors of one or more computing devices (e.g., the client device 110). When executed by the one or more processors, the computer-executable instructions of the variant pathogenicity prediction system 104 can cause the computing devices to perform the bubble detection methods described herein. Alternatively, the components of the variant pathogenicity prediction system 104 can comprise hardware, such as special purpose processing devices to perform a certain function or group of functions. Additionally, or alternatively, the components of the variant pathogenicity prediction system 104 can include a combination of computer-executable instructions and hardware.

[0169] Furthermore, the components of the variant pathogenicity prediction system 104 performing the functions described herein with respect to the variant pathogenicity prediction system 104 may, for example, be implemented as part of a stand-alone application, as a module of an application, as a plug-in for applications, as a library function or functions that may be called by other applications, and / or as a cloud-computing model. Thus, components of the variant pathogenicity prediction system 104 may be implemented as part of a stand-alone application on a personal computing device or a mobile device. Additionally, or alternatively, the components of the variant pathogenicity prediction system 104 may be implemented in any application that provides sequencing services including, but not limited to Illumina PrimateAI, Illumina PrimateAIlD, Illumina PrimateAI2D, Illumina PrimateAI3D, or Illumina TruSight. “Illumina,” “PrimateAI,” “PrimateAIlD,” “PrimateAI2D,” “PrimateAI3D,” and “TruSight,” are either registered trademarks or trademarks of Illumina, Inc. in the United States and / or other countries.

[0170] Embodiments of the present disclosure may comprise or utilize a special purpose or general-purpose computer including computer hardware, such as, for example, one or more processors and system memory, as discussed in greater detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more of the processes described herein may be implemented at least in part as instructions embodied in anon-transitory computer-readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein). In general, a processor (e.g., a microprocessor) receives instructions, from a non-transitory computer-readable medium, (e.g., a memory, etc.), and executes those instructions, thereby performing one or more processes, including one or more of the processes described herein.

[0171] Computer-readable media can be any available media that can be accessed by a general purpose or special purpose computer system. Computer-readable media that store computerexecutable instructions are non-transitory computer-readable storage media (devices). Computer- readable media that carry computer-executable instructions are transmission media. Thus, by way of example, and not limitation, embodiments of the disclosure can comprise at least two distinctly different kinds of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.

[0172] Non-transitory computer-readable storage media (devices) includes RAM, ROM, EEPROM, CD-ROM, solid state drives (SSDs) (e.g., based on RAM), Flash memory, phasechange memory (PCM), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store desired programcode means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer.

[0173] A “network” is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a computer, the computer properly views the connection as a transmission medium. Transmissions media can include a network and / or data links which can be used to carry desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer. Combinations of the above should also be included within the scope of computer- readable media.

[0174] Further, upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to non-transitory computer-readable storage media (devices) (or vice versa). For example, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface module (e.g., a NIC), and then eventually transferred to computer system RAM and / or to less volatile computer storage media (devices) at a computer system. Thus, it should be understood that non-transitory computer- readable storage media (devices) can be included in computer system components that also (or even primarily) utilize transmission media.

[0175] Computer-executable instructions comprise, for example, instructions and data which, when executed at a processor, cause a general-purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. In some embodiments, computer-executable instructions are executed on a general-purpose computer to turn the general-purpose computer into a special purpose computer implementing elements of the disclosure. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.

[0176] Those skilled in the art will appreciate that the disclosure may be practiced in network computing environments with many types of computer system configurations, including, personal computers, desktop computers, laptop computers, message processors, hand-held devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs,minicomputers, mainframe computers, mobile telephones, PDAs, tablets, pagers, routers, switches, and the like. The disclosure may also be practiced in distributed system environments where local and remote computer systems, which are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network, both perform tasks. In a distributed system environment, program modules may be located in both local and remote memory storage devices.

[0177] Embodiments of the present disclosure can also be implemented in cloud computing environments. In this description, “cloud computing” is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be employed in the marketplace to offer ubiquitous and convenient on-demand access to the shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then scaled accordingly.

[0178] A cloud-computing model can be composed of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and so forth. A cloud-computing model can also expose various service models, such as, for example, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (laaS). A cloud-computing model can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, and so forth. In this description and in the claims, a “cloud-computing environment” is an environment in which cloud computing is employed.

[0179] FIG. 13 illustrates a block diagram of a computing device 1300 that may be configured to perform one or more of the processes described above. One will appreciate that one or more computing devices such as the computing device 1300 may implement the variant pathogenicity prediction system 104. As shown by FIG. 13, the computing device 1300 can comprise a processor 1302, a memory 1304, a storage device 1306, an I / O interface 1308, and a communication interface 1310, which may be communicatively coupled by way of a communication infrastructure 1312. In certain embodiments, the computing device 1300 can include fewer or more components than those shown in FIG. 13. The following paragraphs describe components of the computing device 1300 shown in FIG. 13 in additional detail.

[0180] In one or more embodiments, the processor 1302 includes hardware for executing instructions, such as those making up a computer program. As an example, and not by way of limitation, to execute instructions for dynamically modifying workflows, the processor 1302 may retrieve (or fetch) the instructions from an internal register, an internal cache, the memory 1304, or the storage device 1306 and decode and execute them. The memory 1304 may be a volatile or non-volatile memory used for storing data, metadata, and programs for execution by the processor(s). The storage device 1306 includes storage, such as a hard disk, flash disk drive, or other digital storage device, for storing data or instructions for performing the methods described herein.

[0181] The I / O interface 1308 allows a user to provide input to, receive output from, and otherwise transfer data to and receive data from computing device 1300. The I / O interface 1308 may include a mouse, a keypad or a keyboard, a touch screen, a camera, an optical scanner, network interface, modem, other known I / O devices or a combination of such I / O interfaces. The I / O interface 1308 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O interface 1308 is configured to provide graphical data to a display for presentation to a user. The graphical data may be representative of one or more graphical user interfaces and / or any other graphical content as may serve a particular implementation.

[0182] The communication interface 1310 can include hardware, software, or both. In any event, the communication interface 1310 can provide one or more interfaces for communication (such as, for example, packet-based communication) between the computing device 1300 and one or more other computing devices or networks. As an example, and not by way of limitation, the communication interface 1310 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI.

[0183] Additionally, the communication interface 1310 may facilitate communications with various types of wired or wireless networks. The communication interface 1310 may also facilitate communications using various communication protocols. The communication infrastructure 1312 may also include hardware, software, or both that couples components of the computing device 1300 to each other. For example, the communication interface 1310 may use one or more networks and / or protocols to enable a plurality of computing devices connected by a particular infrastructure to communicate with each other to perform one or more aspects of the processes described herein. To illustrate, the sequencing process can allow a plurality of devices (e.g., a client device, sequencing device, and server device(s)) to exchange information such as sequencing data and error notifications.

[0184] In the foregoing specification, the present disclosure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of the present disclosure(s) are described with reference to details discussed herein, and the accompanying drawings illustrate the various embodiments. The description above and drawings are illustrative of the disclosure and are not to be construed as limiting the disclosure. Numerous specific detailsare described to provide a thorough understanding of various embodiments of the present disclosure.

[0185] The present disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. For example, the methods described herein may be performed with less or more steps / acts or the steps / acts may be performed in differing orders. Additionally, the steps / acts described herein may be repeated or performed in parallel with one another or in parallel with different instances of the same or similar steps / acts. The scope of the present application is, therefore, indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.

Claims

CLAIMSWe Claim:

1. A system comprising: at least one processor; and a non-transitory computer readable medium comprising instructions that, when executed by the at least one processor, cause the system to: access, for a target amino acid at a target protein position within a protein, a first pathogenicity score generated by a first variant pathogenicity machine-learning model and a second pathogenicity score generated by a second variant pathogenicity machine-learning model; access protein structural data indicating a density of amino acids within the protein; provide, to a meta variant pathogenicity machine-learning model, the first pathogenicity score, the second pathogenicity score, and the protein structural data; and generate, utilizing the meta variant pathogenicity machine-learning model, a refined pathogenicity score for the target amino acid at the target protein position based on the first pathogenicity score, the second pathogenicity score, and the protein structural data.

2. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: provide, to the meta variant pathogenicity machine-learning model, a conservation profile comprising data representing conservation of amino acids at protein positions of the protein from different species; and generate, utilizing the meta variant pathogenicity machine-learning model, the refined pathogenicity score further based on the conservation profile.

3. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: provide, from a population database and to the meta variant pathogenicity machine-learning model, protein annotations indicating one or more variants at the target protein position are benign or pathogenic within a population; and generate, utilizing the meta variant pathogenicity machine-learning model, the refined pathogenicity score further based on the protein annotations.

4. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to access the protein structural data by accessing density values indicating a density of the amino acids within the protein.

5. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate, utilizing the meta variant pathogenicity machine-learning model, a propensity score indicating a probability that the target amino acid at the target protein position is labeled benign in a training dataset given a ground truth that the target amino acid is benign.

6. The system of claim 5, further comprising instructions that, when executed by the at least one processor, cause the system to: combine, as part of an expectation stage of an expectation-maximization (EM) algorithm, the refined pathogenicity score and the propensity score with benign labels to generate a combined pathogenicity-propensity score as a training label; determine, as part of a maximization stage of the EM algorithm, a variant pathogenicity loss based on a comparison of the refined pathogenicity score and the combined pathogenicitypropensity score; and adjust parameters of the meta variant pathogenicity machine-learning model based on the variant pathogenicity loss.

7. The system of claim 6, further comprising instructions that, when executed by the at least one processor, cause the system to: determine, as part of a maximization stage of the EM algorithm, a variant propensity loss based on a comparison of the propensity score and the combined pathogenicity -propensity score; and adjust parameters of the meta variant pathogenicity machine-learning model based on the variant propensity loss.

8. The system of claim 7, wherein the variant pathogenicity loss comprises a positive- and-unlabeled (PU) learning pathogenicity loss and the variant propensity loss comprises a PU learning propensity loss.

9. The system of claim 7, further comprising instructions that, when executed by the at least one processor, cause the system to: determine, as part of the maximization stage of the EM algorithm, a mutation rate rank loss; applying the mutation rate rank loss to the propensity score; and combining the mutation rate rank loss with the variant propensity loss.

10. The system of claim 7, further comprising instructions that, when executed by the at least one processor, cause the system to: iteratively determine combined pathogenicity-propensity scores as part of the expectation stage until satisfying a convergence criteria for refined pathogenicity scores generated and propensity scores generated by the meta variant pathogenicity machine-learning model; anditeratively adjust the parameters of the meta variant pathogenicity machine-learning model as part of the maximization stage until satisfying the convergence criteria.

11. The system of claim 7, wherein the meta variant pathogenicity machine-learning model comprises a multilayer perceptron (MLP).

12. The system of claim 1, wherein first variant pathogenicity machine-learning model or the second variant pathogenicity machine-learning model comprises a transformer machinelearning model, a convolutional neural network (CNN), a variational autoencoder (VAE), a multilayer perceptron (MLP), a recurrent neural network (RNN), a long short-term memory (LSTM), or a decision tree model.

13. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate, utilizing an all-data transformer neural network, the refined pathogenicity score for the target amino acid by: providing, to the meta variant pathogenicity machine-learning model: a first set of pathogenicity scores generated by the first variant pathogenicity machine-learning model for each candidate amino acid at each protein position within the protein; a second set of pathogenicity scores generated by the second variant pathogenicity machine-learning model for each candidate amino acid at each protein position within the protein; allele counts and allele frequencies for candidate amino acids in protein positions of the protein; mutation rates for the protein positions of the protein; a reference-residues embedding representing reference residues for the protein; density values indicating a density of the amino acids within the protein; a reference amino-acid embedding representing a reference amino-acid sequence for the protein; and generating, utilizing the meta variant pathogenicity machine-learning model, the refined pathogenicity score further based on the first set of pathogenicity scores, the second set of pathogenicity scores, the allele counts and allele frequencies, the mutation rates, the referenceresidues embedding, and the reference amino-acid embedding.

14. A system comprising: at least one processor and a meta pathogenicity probability machine-learning model comprising a first pathogenicity probability model and a second pathogenicity probability model; anda non-transitory computer readable medium comprising instructions that, when executed by the at least one processor, cause the system to: access, for a target amino acid at a target protein position within a protein, a first pathogenicity score generated by a first variant pathogenicity machine-learning model and a second pathogenicity score generated by a second variant pathogenicity machine-learning model; generate, by utilizing the first pathogenicity probability model to process the first pathogenicity score, a first pathogenicity probability that the target amino acid has been determined to be a pathogenic variant or a benign variant at the target protein position; generate, by utilizing the second pathogenicity probability model to process the second pathogenicity score, a second pathogenicity probability that the target amino acid has been determined to be a pathogenic variant or a benign variant at the target protein position; and determine, based on the first pathogenicity probability and the second pathogenicity probability, a combined pathogenicity probability that the target amino acid has been determined to be a pathogenic variant or a benign variant at the target protein position.

15. The system of claim 14, further comprising instructions that, when executed by the at least one processor, cause the system to: generate the first pathogenicity probability by generating a first probability that a clinical variant database comprises a benign label or a pathogenic label for the target amino acid at the target protein position; and generate the second pathogenicity probability by generating a second probability that the clinical variant database comprises a benign label or a pathogenic label for the target amino acid at the target protein position.

16. The system of claim 14, further comprising instructions that, when executed by the at least one processor, cause the system to: generate the first pathogenicity probability by generating a first probability that a primate variant database comprises a benign label for the target amino acid at the target protein position; and generate the second pathogenicity probability by generating a second probability that the primate variant database comprises a benign label for the target amino acid at the target protein position.

17. The system of claim 14, further comprising instructions that, when executed by the at least one processor, cause the system to determine the combined pathogenicity probability bydetermining a weighted average of the first pathogenicity probability and the second pathogenicity probability.

18. The system of claim 17, further comprising instructions that, when executed by the at least one processor, cause the system to determine the weighted average by: determining a first weight for the first pathogenicity probability based on a difference between the first pathogenicity score and the first pathogenicity probability; determining a second weight for the second pathogenicity probability based on a difference between the second pathogenicity score and the second pathogenicity probability; and averaging the first pathogenicity probability multiplied by the first weight and the second pathogenicity probability multiplied by the second weight.

19. The system of claim 14, wherein the first variant pathogenicity machine-learning model or the second variant pathogenicity machine-learning model comprises a transformer machine-learning model, a convolutional neural network (CNN), a variational autoencoder (VAE), a multilayer perceptron (MLP), a recurrent neural network (RNN), a long short-term memory (LSTM), or a decision tree model.

20. The system of claim 14, further comprising instructions that, when executed by the at least one processor, cause the system to: determine a first pathogenicity-probability loss between the combined pathogenicity probability and the first pathogenicity probability; adjust parameters of the first pathogenicity probability model based on the first pathogenicity-probability loss; determine a second pathogenicity-probability loss between the combined pathogenicity probability and the second pathogenicity probability; and adjust parameters of the second pathogenicity probability model based on the second pathogenicity-probability loss.

21. The system of claim 14, wherein the first pathogenicity probability model and the second pathogenicity probability model each comprise a multilayer perceptron (MLP).

Citation Information

Patent Citations

  • Mask pattern for protein language models

    US20230207060A1

  • Pathogenicity language model

    US20230207061A1

  • Deep learning network for evolutionary conservation

    US20230207054A1

  • Image-based variant pathogenicity determination

    US20230245305A1

  • Deep convolutional neural networks for variant classification

    WO2019079180A1