Ensembling variant pathogenicity scores over artificial benign and unknown amino-acid sequences
By ensembling pathogenicity scores using a trained machine-learning model with diverse amino-acid sequences, the system addresses the inaccuracies and inefficiencies of existing models, achieving improved accuracy and reduced computational costs.
Patent Information
- Application Number
- PCT/US2024/061494
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-20
- Publication Date
- 2025-06-26
AI Technical Summary
Existing pathogenicity prediction models generate inaccurate pathogenicity scores due to their reliance on single context references and limited context variations, leading to computational inefficiencies and excessive resource consumption.
The system ensembles variant pathogenicity scores by generating combined scores through a trained machine-learning model, utilizing artificial and natural amino-acid sequences with benign variants to improve input data diversity and reduce computational burdens.
This approach enhances the accuracy and precision of pathogenicity predictions, reduces computational inefficiencies, and conserves resources by moving ensembling from the model space to the input space.
Smart Images

Figure US2024061494_26062025_PF_FP_ABST
Abstract
Description
ENSEMBLING VARIANT PATHOGENICITY SCORES OVER ARTIFICIAL BENIGN AND UNKNOWN AMINO-ACID SEQUENCESCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 613,930, entitled “ENSEMBLING VARIANT PATHOGENICITY SCORES OVER ARTIFICIAL BENIGN AND UNKNOWN AMINO-ACID SEQUENCES,” filed on December 22, 2023, which is incorporated herein by reference in its entirety.BACKGROUND
[0002] In recent years, biotechnology firms and research institutions have improved software for predicting a pathogenicity of protein variants or genetic variants. For instance, some existing pathogenicity prediction models generate predictions that estimate a degree to which amino-acid variants are benign or pathogenic. Such pathogenicity predictions can indicate whether an aminoacid variant is likely to cause various diseases, such as certain cancers, developmental disorders, or heart conditions. In addition to the intrinsic predictive value of such predictions, biotechnology firms and research institutions have developed downstream applications for pathogenicity predictions. For instance, pathogenicity predictions output by machine-learning models have been used to identify target variants in a population subset for new drugs as well as target variants that may be the subject of genetic editing.
[0003] To predict whether a variant amino acid at a given position is pathogenic or benign, existing machine-learning models can generate single pathogenicity scores for each candidate amino acid at a target position with a reference (canonical) amino-acid sequence. In particular, by substituting each of nineteen candidate alternative amino acid into a position within a reference amino-acid sequence and determining a corresponding nineteen different pathogenicity predictions, existing models can determine nineteen pathogenicity scores specific to each of the respective nineteen candidate alternative amino acids. Similarly, DeepMind or AlphaMissense’s approach to scoring a single variant at inference is likewise limited to a one-score-per-variant approach. Such one-score-per-variant approaches are limited to scores within the context of the reference amino-acid sequence and fails to account for other possible contexts in which a candidate amino acid may be benign or pathogenic.
[0004] While such a one-score-per-variant approach in pathogenicity prediction models has demonstrated state-of-the-art performance and significant improvements in downstream applications to date, existing models do not consistently generate accurate pathogenicity predictions. Due in part to limiting pathogenicity scores to a single context of a reference aminoacid sequence — or to limiting such scores to only a handful of possible context variations with naturally occurring variants to the reference amino-acid sequence — the one-score-per-variantapproach is susceptible to bias that can lead to inaccurate pathogenicity predictions in cases where existing pathogenicity prediction models overcorrect or overemphasize data from a single variant that may not be totally reliable. Such an overemphasis on a single variant can compromise performance of pathogenicity scores for protein variants or benign proteins in data from, for example, the United Kingdom (UK) Biobank and cell-line experiments for Saturation Mutagenesis.
[0005] In addition to the inaccuracies demonstrated by some existing machine-learning models, certain of these existing machine-learning models are also computationally inefficient. More specifically, many existing pathogenicity machine-learning models require retraining many times over to generate pathogenicity predictions for different sets of amino-acid sequences and / or for different target protein positions. The process of retraining pathogenicity prediction models, such as DeepSequence, for new sets of amino-acid sequences can take weeks or months (e.g., 3-4 months) of constant computational expenditure to achieve improved performance in predicting protein pathogenicity. Despite months-long training, some of these existing models have no publicly released scores for all human proteins. Certain existing models, such as transformer models, that provide pathogenicity scores for only various proteins consume notoriously large amounts of computing resources for pre-training to even begin evaluating performance. But developers of transformer models often test and release performance data for transformers only after a single checkpoint — that is, after training a single transformer model on multiple sequences — because testing on additional checkpoints or additional versions of a transformer model is computationally prohibitive. Existing models thus consume excessive amounts of computational resources that could otherwise be preserved with a more efficient approach. Such computational inefficiencies are especially pronounced in systems that retrain across large numbers of sets of amino-acid sequences to generate pathogenicity predictions.
[0006] These, along with additional problems and issues exist in existing sequencing systems.SUMMARY
[0007] This disclosure describes one or more embodiments of systems, methods, and non- transitory computer readable storage media that solve one or more of the problems described above or provide other advantages over the art. In particular, the disclosed systems can utilize a trained variant pathogenicity machine-learning model to generate combined pathogenicity scores for variant amino acids at target protein positions by ensembling a set of initial pathogenicity scores for such target amino-acid variants within amino-acid sequences that include benign amino-acid variants. For example, the disclosed systems improve pathogenicity predictions of existing variant pathogenicity machine-learning models by enhancing input data in the form of diversified benign variants or certain naturally occurring flanking target amino-acid variants. Specifically, in some embodiments, the disclosed systems generate or access artificial amino-acid sequences (and / ornatural amino-acid sequences) that include benign amino-acid variants in flanking regions of a target protein position within a reference amino-acid sequence. Such benign variants may include primate amino-acid variants with at least a threshold probability of being benign. In addition, the disclosed systems generate a set of pathogenicity scores for the respective artificial amino-acid sequences utilizing a variant pathogenicity machine-learning model. The disclosed systems can further ensemble over the pathogenicity scores to generate a combined pathogenicity scores from the artificial amino-acid sequences by, for example, generating a mean pathogenicity score for each candidate variant amino acid (e.g., across each of nineteen alternative amino acids at a target protein position).
[0008] In addition or in the alternative to implementing a trained variant pathogenicity machine-learning model, the disclosed systems can train a variant pathogenicity machine-learning model to facilitate determining pathogenicity scores that can be ensembled. As part of the training process, the disclosed systems can determine or identify a training amino-acid-sequence pair that includes two types of amino-acid sequences each comprising an amino acid at a target protein position (e.g., natural amino-acid sequence or unknown amino-acid sequence). In addition, the disclosed systems can provide the training amino-acid-sequence pair to a variant pathogenicity machine-learning model, where the two amino-acid sequences in the pair are processed by the model in a randomly permuted order, one after the other. Based on processing the training amino- acid-sequence pair, the variant pathogenicity machine-learning model generates a predicted pathogenicity likelihood for the amino acid at the target protein position for each of the two types of amino-acid sequences. The disclosed systems can further utilize one or more loss functions to compare predicted pathogenicity likelihoods for the amino acid at the target position within the training amino-acid-sequence pair to corresponding ground-truth-ordering values that indicate a true (e.g., observed) order of the two amino-acid sequences in the pair (e.g., the order they were provided to the variant pathogenicity machine-learning model). The disclosed systems further adjust internal parameters of the variant pathogenicity machine-learning model based on the iterative comparison of predicted pathogenicity likelihoods to improve accuracy over a series of training iterations.
[0009] Additional features and advantages of one or more embodiments of the present disclosure will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by the practice of such example embodiments.BRIEF DESCRIPTION OF THE DRAWINGS
[0010] This patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0011] The detailed description refers to the drawings briefly described below.
[0012] FIG. 1 illustrates a schematic diagram of a computing system in which a pathogenicity score ensembling system can operate in accordance with one or more embodiments.
[0013] FIG. 2 illustrates an overview of the pathogenicity score ensembling system generating a combined pathogenicity score by ensembling over pathogenicity score generated using a variant pathogenicity machine-learning model to process artificial amino-acid sequences in accordance with one or more embodiments.
[0014] FIG. 3 illustrates an example diagram for generating amino-acid sequences in accordance with one or more embodiments.
[0015] FIG. 4 illustrates an example diagram for identifying natural variant amino-acid sequences in accordance with one or more embodiments.
[0016] FIG. 5 illustrates an example diagram for generating an artificial amino-acid sequence from a natural amino-acid sequence in accordance with one or more embodiments.
[0017] FIG. 6 illustrates an example diagram for generating a combined pathogenicity score from multiple pathogenicity scores in accordance with one or more embodiments.
[0018] FIG. 7 illustrates an example diagram for determining a combined pathogenicity score based on logit differences in accordance with one or more embodiments.
[0019] FIGS. 8A-8B illustrate example graphs of experimental results for generating pathogenicity score using different ratios of candidate benign positions for substituting benign primate variants in accordance with one or more embodiments.
[0020] FIGS. 9A-9B illustrate graphs of experimental results for using different numbers of amino-acid sequences for ensembling to generate combined pathogenicity scores in accordance with one or more embodiments.
[0021] FIG. 10 illustrates an example diagram for training a variant pathogenicity machinelearning model in accordance with one or more embodiments.
[0022] FIGS. 11A-11B illustrate example diagrams for learning ordering values for training- amino-acid-sequence pairs using a training process in accordance with one or more embodiments.
[0023] FIGS. 12A-12B illustrate example plot charts of training results using different loss prefactors as part of training a variant pathogenicity machine-learning model in accordance with one or more embodiments.
[0024] FIG. 13 illustrates an example graph for training a variant pathogenicity machinelearning model using loss stratification in accordance with one or more embodiments.
[0025] FIGS. 14A-14B illustrate example graphs depicting results for training a variant pathogenicity machine-learning model using different loss prefactors across amino-acid sequences divided according to stratified losses in accordance with one or more embodiments.
[0026] FIGS. 15A-15B illustrate example plot charts depicting training results using loss stratification and different prefactors for a variant pathogenicity machine-learning model in accordance with one or more embodiments.
[0027] FIG. 16 illustrates a series of acts for generating a combined pathogenicity score by ensembling over pathogenicity score generated using a variant pathogenicity machine-learning model in accordance with one or more embodiments.
[0028] FIG. 17 illustrates a series of acts for training a variant pathogenicity machine-learning model in accordance with one or more embodiments.
[0029] FIG. 18 illustrates a block diagram of an example computing device in accordance with one or more embodiments.DETAILED DESCRIPTION
[0030] This disclosure describes one or more embodiments of a pathogenicity score ensembling system that can generate and ensemble pathogenicity scores for variant amino acids at target positions of a human amino-acid sequence. More specifically, the pathogenicity score ensembling system can ensemble or combine multiple pathogenicity scores for a candidate amino acid a single target protein position to improve the accuracy and robustness of the scores, reducing bias over systems that generate a single score per variant. To facilitate generating a set of multiple pathogenicity scores for a candidate amino acid (e.g., at a single target protein, for example, the pathogenicity score ensembling system accesses or generates a set of different input amino-acid sequences that include benign amino-acid variants at adjacent positions within flanking regions of the target protein position. The pathogenicity score ensembling system further (i) utilizes a variant pathogenicity machine-learning model to generate a set of pathogenicity scores for the candidate amino acid at a single target protein position within the different input amino-acid sequences comprising different sets of benign amino-acid variants (e.g., a separate score for each of the nineteen possible amino-acid variants at the target position) and (ii) further ensembles the set of scores to generate a combined pathogenicity score.
[0031] As part of this process, in some cases, the pathogenicity score ensembling system generates or accesses artificial amino-acid sequences to use as input data for a variant pathogenicity machine-learning model. For example, the pathogenicity score ensembling system generates an artificial amino-acid sequence by modifying a reference amino-acid sequence. The pathogenicity score ensembling system can determine or identify candidate benign positions within the reference amino-acid sequence as locations for substituting variant amino acids in for reference amino acids according to a ratio of candidate benign positions. In some cases, the candidate benign positions are within flanking regions of a target protein position with a reference amino-acid sequence. At such candidate benign positions, the pathogenicity score ensembling system can substitute one ormore benign variants from a benign variant database (e.g., by using protein positions and benign primate amino-acid variants identified by PrimateAI scores) for reference amino acids within the flanking regions of the target protein position. In some cases, the pathogenicity score ensembling system identifies amino-acid variants (e.g., primate amino-acid variants) that satisfy a threshold probability of being benign to use as candidate for performing such substitutions.
[0032] In addition or in the alternative to using artificial amino-acid sequences with benign variants as inputs for a variant pathogenicity machine-learning model, in some embodiments, the pathogenicity score ensembling system uses natural amino-acid sequences (e.g., amino-acid sequences that include benign variants that naturally occur in one or more species) from a multiple sequence alignment or a natural amino-acid database. For example, the pathogenicity score ensembling system accesses natural amino-acid sequences that include benign variants that naturally occur in one or more species to input into a variant pathogenicity machine-learning model. In certain cases, the pathogenicity score ensembling system generates an artificial amino-acid sequence by combining a natural amino-acid sequence with a benign primate amino-acid sequence. For instance, the pathogenicity score ensembling system substitutes (i) one or more benign primate variants that occur in flanking regions of a target protein position for (ii) amino-acid variants at corresponding positions within a natural amino-acid sequence.
[0033] The pathogenicity score ensembling system can further input artificial and / or natural amino-acid sequences into a variant pathogenicity machine-learning model, whereupon the variant pathogenicity machine-learning model generates a set of pathogenicity scores for a candidate amino acid at the target protein position. For example, the pathogenicity score ensembling system generates a set of nineteen pathogenicity scores, one for each possible variant amino acid at the target protein position. In some embodiments, the pathogenicity score ensembling system generates pathogenicity score by determining a logit difference that reflects or indicates a difference in probabilities between a reference amino acid and a candidate variant amino acid at a target protein position. In addition, the pathogenicity score ensembling system can ensemble the set of pathogenicity scores to generate a combined pathogenicity score by, for example, determining a mean across the pathogenicity scores (e.g., a mean of the nineteen possible logit differences).
[0034] As noted above, in one or more embodiments, the pathogenicity score ensembling system trains a variant pathogenicity machine-learning model to generate pathogenicity scores that can be used for ensembling. For example, the pathogenicity score ensembling system generates specific training input in the form of training amino-acid-sequence pairs that each include two amino-acid sequences of different types. In some embodiments, for instance, a training amino- acid-sequence pair includes a natural amino-acid sequence and an unknown amino-acid sequence.In other embodiments, a training amino-acid-sequence pair includes a reference (or canonical) amino-acid sequence and a modified (or artificial) amino-acid sequence with variants substituted at one or more protein positions.
[0035] The pathogenicity score ensembling system can input the training amino-acid-sequence pair into a variant pathogenicity machine-learning model, whereupon the variant pathogenicity machine-learning model generates two predicted pathogenicity likelihoods, one for an amino acid at a targe protein position within each of the amino-acid sequences in the training amino-acid- sequence pair. In addition, the pathogenicity score ensembling system can compare (e.g., using a loss function) the predicted pathogenicity likelihoods with a ground-truth-ordering value that indicates an actual order in which the amino-acid sequences within the training amino-acid- sequence pair were provided or fed into the variant pathogenicity machine-learning model. Based on the comparison of the predicted pathogenicity likelihoods, the pathogenicity score ensembling system can adjust or modify parameters of the variant pathogenicity machine-learning model.
[0036] Over multiple iterations, the pathogenicity score ensembling system thus trains the variant pathogenicity machine-learning model to distinguish between predicted pathogenicity likelihoods for the first amino-acid sequence and the second amino-acid sequence within the training amino-acid-sequence pair. For example, the pathogenicity score ensembling system generates a different loss for each amino-acid sequence in the pair, and over the training iterations, the different losses separate from each other more distinctly (e.g., where the loss for the natural amino-acid sequence is closer to 0 and the loss for the unknown amino-acid sequence is closer to 1) to thus distinguish between the amino-acid sequences. Accordingly, the pathogenicity score ensembling system trains the variant pathogenicity machine-learning model to generate accurate pathogenicity scores that reflect amino-acid variants within amino-acid sequences by learning to separate natural and unknown sequences (and / or to separate reference and modified sequences).
[0037] As indicated above, the pathogenicity score ensembling system provides several technical advantages relative to existing pathogenicity prediction models. For example, the pathogenicity score ensembling system improves the accuracy and precision with which pathogenicity prediction models generate pathogenicity predictions for amino-acid variants. As noted above, existing pathogenicity prediction models exhibit inaccuracies that stem from bias or overcorrection from generating a pathogenicity score from a single variant amino-acid sequence with limited context. By contrast, the pathogenicity score ensembling system can generate multiple (sets of) pathogenicity scores across a number of variant amino-acid sequences comprising a target amino acid at a single target protein position. Specifically, the pathogenicity score ensembling system can generate a pathogenicity score for each candidate amino-acid variant at a target protein position from among a set of candidate amino-acid variants (for a total of nineteen candidate amino-acid variants) within the context of a number of variant amino-acid sequences. The pathogenicity score ensembling system can further ensemble over the multiple pathogenicity scores to generate a combined pathogenicity score that more accurately reflects a pathogenicity prediction. By generating a combined pathogenicity score based on data from a larger sample of amino-acid sequences, the pathogenicity score ensembling system more accurately reflects the amino-acid diversity in a (human) population and is less biased than existing systems that rely on a single variant amino-acid sequence to generate a pathogenicity score. Consequently, as shown by FIGS. 8A-9B and described further below, the pathogenicity score ensembling system can generate more accurate pathogenicity score for particular assays, such as cell-line experiments for Saturation Mutagenesis, and / or for sequences in databases, such as the United Kingdom Biobank (UKBB).
[0038] In addition, in some embodiments, the pathogenicity score ensembling system is more flexible than existing pathogenicity prediction models. Specifically, the pathogenicity score ensembling system can ensemble pathogenicity scores using a model-agnostic approach that flexibly adapts to different pathogenicity prediction models. Indeed, while some existing models are strictly limited to generating (one-score-per-variant) pathogenicity scores according to their specific model architecture and amino-acid context, in some cases, the pathogenicity score ensembling system can ensemble pathogenicity scores from a variety of models in a variety of contexts. By implementing the ensembling techniques described herein, the pathogenicity score ensembling system can accordingly adapt to, and improve upon, pathogenicity scores generated by different machine-learning models.
[0039] The disclosed ensembling of pathogenicity scores for a candidate amino acid at a target protein position not only improves accuracy and flexibility but also expedites the computational speed and processing that would be required by existing pathogenicity prediction models to approach the same accuracy and precision of the disclosed ensembled pathogenicity scores. If the state-of-the-art pathogenicity models (e.g., PrimateAI3D) were to process and determine pathogenicity scores for a candidate amino acid at a target protein position within the context of each possible configuration of benign variants in an amino-acid sequence for a protein, the time and processing required to determine such a volume of pathogenicity scores would increase not only when implementing such a trained state-of-the-art pathogenicity model but also (as described in the background above) significantly slow the process of training separate models or an ensemble of models and consume days or months of processing time — before ensembling such pathogenicity scores together. Rather than such lengthy processing of each possible configuration of benign variants in an amino-acid sequence, the disclosed pathogenicity score ensembling system can generate accurate and precise ensembled pathogenicity scores for a candidate amino acid at a target protein position within seconds.
[0040] In addition to improved accuracy and precision, in some embodiments, the pathogenicity score ensembling system also improves computational efficiency relative to existing systems. As noted above, some existing systems consume extensive computational resources when training a variant pathogenicity machine-learning model to generate pathogenicity scores, and the training expense is compounded for each new set of training amino-acid sequences. Indeed, to ensemble pathogenicity scores, existing systems train and retrain over new training datasets for each new prediction to then have multiple predictions from which to ensemble. In the alternative to training separate models, a more traditional training approach of retraining an ensemble of machine-learning models (e.g., PrimateAI3D) with different random seeds (e.g., randomly initialized parameters) and evaluating performance of such an ensemble only on a human reference sequence has proven computationally inefficient. Rather than retraining for each new context and / or retraining an ensemble with performance evaluation on only a human reference sequence, the pathogenicity score ensembling system shifts the ensembling from the model space (which requires multiple models and multiple trainings to generate multiple predictions over which to ensemble) to the input space (which requires only a single model training process by using intelligently selected amino-acid sequences as inputs). Indeed, the pathogenicity score ensembling system can generate and / or access multiple artificial amino-acid sequences based on benign (primate) amino-acid variants and / or natural amino-acid sequences to use as input for the variant pathogenicity machine-learning model. By moving ensembling from the model space to the input space, the pathogenicity score ensembling system preserves large amounts of computational resources that existing systems expend on training, reducing the training time from months to hours in many cases (and thus saving the corresponding processing power and memory required to perform the training).
[0041] As illustrated by the foregoing discussion, the present disclosure utilizes a variety of terms to describe features and advantages of the pathogenicity score ensembling system. As used herein, for example, the term “machine-learning model” refers to a computer algorithm or a collection of computer algorithms that automatically improve for a particular task through experience based on use of data. For example, a machine learning model can utilize one or more learning techniques to improve in accuracy and / or effectiveness. Example machine learning models include various types of decision trees (e.g., gradient boosted trees), support vector machines, Bayesian networks, or neural networks (e.g., transformer neural networks, recurrent neural networks, triangle attention neural networks).
[0042] In some cases, the pathogenicity score ensembling system uses a variant pathogenicity machine-learning model to generate, modify, or update a pathogenicity score for a target amino acid (e.g., at a target protein position). As used herein, the term “variant pathogenicity machine-learning model” refers to a machine-learning model that generates a pathogenicity score for either a protein (e.g., protein variant) or an amino acid at a particular protein position of a protein or an amino-acid sequence. For example, a variant pathogenicity machine-learning model includes a machine-learning model that generates an initial pathogenicity score for a variant amino acid at a target protein position within a protein based on an amino-acid sequence for the protein. In addition to or as part of one or more amino-acid sequences an input, in some cases, a variant pathogenicity machine-learning model processes other inputs, such as a multiple sequence alignment (MSA) corresponding to the protein or a reference amino-acid sequence for the protein. As indicated below, a variant pathogenicity machine-learning model can take the form of different models, including, but not limited to, a transformer neural network, a convolutional neural network (CNN), a sequence-to-sequence model, a variational autoencoder (VAE), a multilayer perceptron (MLP), a recurrent neural network (RNN), a long short-term memory (LSTM), or a decision tree model.
[0043] Relatedly, as used herein, the term “pathogenicity score” refers to a measurement, numerical value, or score indicating a degree to which a protein or an amino acid at a protein position within a protein is benign or pathogenic. In particular, a pathogenicity score can include a score indicating a degree to which a candidate amino acid is benign or pathogenic when located at a target protein position within a specific amino-acid sequence. Accordingly, a pathogenicity score can be both specific to a candidate amino acid at a target protein position and specific to an amino-acid sequence within which the candidate amino acid is located.
[0044] In some cases, a pathogenicity score includes a logit, a logit difference, or some other numerical value indicating a probability of a variant amino acid at a target protein position of a protein relative to a reference amino acid at the target protein. Because a pathogenicity score can indicate a particular amino acid in a protein position is benign, in some cases, a pathogenicity score represents a fitness of the particular amino acid in the protein position. As but one example a pathogenicity score, in some embodiments, the pathogenicity score for a target alternative amino acid (Sait) at a target protein position includes a numerical value determined from a usual difference of a logit for an alternative amino acid (pait) and a logit for a reference amino acid (pref) at the target protein position. More details concerning this specific example can be found in U.S. Patent Application No. 17 / 975,547, entitled “Pathogenicity Language Model,” by Tobias Hamp, Anastasia Dietrich, Yibing Wu, Jeffrey Ede, and Kai-How Farh, filed on October 27, 2022, which is hereby incorporated in its entirety by reference. Other formulations of a pathogenicity score, however, can likewise be used and are described below.
[0045] Relatedly, the term “combined pathogenicity score” refers to a pathogenicity score for a protein or an amino acid at a protein position in which the score is combined or ensembled from multiple constituent pathogenicity scores. For example, a combined pathogenicity score can takethe form of a mean pathogenicity score, a max pathogenicity score, a min pathogenicity score, a median pathogenicity score, or some other combination of pathogenicity scores for a protein or an amino acid at a particular protein position. In particular, a combined pathogenicity score includes an ensembled pathogenicity score that combines pathogenicity scores for a candidate amino acid at a target protein position within the context of different amino-acid sequences. As explained below, such different amino-acid sequences may include (i) artificial amino-acid sequences comprising one or more benign primate amino-acid variants substituted in for reference amino acids at different adjacent protein positions flanking the target protein position and / or (ii) natural variant amino-acid sequences comprising one or more natural benign amino-acid variants at different adjacent protein positions flanking the target protein position.
[0046] Additionally, as used herein, the term “target protein position” refers to a particular location, coordinate, or order for an amino acid within an amino-acid sequence forming a polypeptide chain for a protein. In particular, a target protein position includes a numerically identified location for an amino acid in an ordered amino-acid sequence representing a protein. For example, a target protein position could include a seventh, fifty-fourth, one hundred and ninetyfifth, two hundredth, or any numbered position within an amino-acid sequence of amino acids (e.g., 300-amino acid sequence) representing a protein. In some cases, a target protein position can be represented as a number along or within a residue sequence index (e.g., depicted in accompanying figures).
[0047] As further indicated above, in some embodiments, the pathogenicity score ensembling system trains a variant pathogenicity machine-learning model using known benign amino acids and / or unknown amino-acid sequences. As used herein, the term “benign amino-acid variant” refers to a particular type of amino acid unlikely to cause a disease in an organism (e.g., to a high degree of confidence or with a high degree of certainty satisfying a threshold confidence). In particular, a benign amino acid includes a particular type of amino acid at a target protein position within a protein unlikely to cause a disease in a human or other primate. For instance, an amino acid labelled as a benign amino acid is benign more than 95% of the time (e.g., 95.8%) based on primate data. In some embodiments, a benign amino-acid variant satisfies a threshold allele frequency (e.g., 0.1%) to qualify as a benign variant. In these or other embodiments, a benign amino-acid variant satisfies a quality score threshold signifying that an observed primate variant amino acid is benign at the same position in a human amino-acid sequence. In some cases, the quality score threshold takes the form of a combination of a random forest score and a method score from PrimateAI3D, which model is described in U.S. Patent Application No. 17 / 703,958 entitled “EFFICIENT VOXELIZATION FOR DEEP LEARNING,” filed March 24, 2022, which is hereby incorporated by reference in its entirety.
[0048] In some cases, a benign amino-acid variant includes or refers to a “benign primate amino-acid variant” that is derived or synthesized from primate amino-acid data to include benign amino-acid variants in flanking regions of a target protein position. For instance, a benign primate amino-acid sequence can include primate amino-acid data, such as amino-acid sequences identified by ILLUMINA’s PrimateAI system or PrimateAI scores, as made available at https: / / primad.basespacelillumina.com, as described by L. Sundaram et al., “Predicting the Clinical Impact of Human Mutation with Deep Neural Networks,” Nat. Genet. 50, 1161-70 (2018) https: / / doi.org / 10.1038 / s41588-018-0167-z (hereinafter Sundaram), as described by H. Gao et al., “The landscape of tolerated genetic variation in humans and primates,” Science 380, eabn8153 (2023), DOL 10.1126 / science.abn8197 (hereinafter Gao), or as stored in Project PRJEB49549 of the European Nucleotide Archive, entitled Primate Whole Genome Sequences, by H. Gao et al. (2023), all of which are hereby incorporated by reference in their entirety. By contrast, the term “unknown amino-acid sequence” refers to a particular type of amino-acid sequence for which it is unknown whether a particular amino-acid sequence causes a disease in an organism. In particular, an unknown amino-acid sequence includes a particular type of amino-acid sequence for which it is unknown whether the particular amino-acid sequence causes a disease in a human or other primate or in another species (e.g., particular mammal).
[0049] Along these lines, as used herein, the term “artificial amino-acid sequence” refers to an amino-acid sequence that is generated, modified, or synthesized by combining amino acids from different sequences together. For example, an artificial amino-acid sequence includes benign primate amino-acid variants (e.g., as identified by ILLUMINA’s PrimateAI system or PrimateAI scores described by Sundaram or Gao) substituted in for reference amino acids in a reference amino-acid sequence at one or more adjacent positions within flanking regions (e.g., regions immediately on either side) of a target protein position. Thus, an artificial amino-acid sequence can refer to a combination of a reference amino-acid sequence and benign variants from a benign primate amino-acid sequence. In some cases, an artificial amino-acid sequence includes or refers to a combination of a natural amino-acid sequence and benign variants from a benign primate amino-acid sequence, where benign primate amino-acid variants are substituted within a natural amino-acid sequence at adjacent positions flanking a target protein position.
[0050] To this point, as used herein, the term “natural amino-acid sequence” (or “natural variant amino-acid sequence”) refers to an amino-acid sequence that occurs naturally in one or more species. For example, a natural amino-acid sequence refers to a sequence of amino acids that have been sampled or observed in primate and / or non-primate species, such as amino-acid sequences stored in one or more Universal Protein resource (UniProt) databases. In some cases, a natural amino-acid sequence includes a natural benign amino-acid variant at a protein position of areference amino-acid sequence, where a “natural benign amino-acid variant” refers to a variant amino acid that naturally occurs in one or more species and is known (with at least a threshold level of confidence) to be benign at a given protein position.
[0051] As further used herein, the term “training amino-acid-sequence pair” refers to an amino-acid sequence pair used to train a variant pathogenicity machine-learning model. For example, a training amino-acid-sequence pair can include a pair of amino-acid sequences that includes two different types of amino-acid sequences. In some cases, the disclosed systems provide a training amino-acid-sequence pair to a variant pathogenicity machine-learning model, where the model processes the pair in a randomly permuted order. Based on processing the training amino- acid-sequence pair, the variant pathogenicity machine-learning model generates a predicted pathogenicity likelihood for an amino acid at the target protein position for each of the two types of amino-acid sequences included in the training amino-acid-sequence pair.
[0052] Additionally, the term “modified amino-acid sequence” refers to an amino-acid sequence that includes substitute amino acids which replace or substitute reference amino acids in an initial or unmodified amino-acid sequence. For example, a modified amino-acid sequence includes benign primate variants substituted in positions of an amino-acid sequence, such as a benign amino-acid sequence, a non-benign amino-acid sequence, or an unknown amino-acid sequence.
[0053] The following paragraphs describe the pathogenicity score ensembling system with respect to illustrative figures that portray example embodiments and implementations. For example, FIG. 1 illustrates a schematic diagram of a computing system 100 in which a pathogenicity score ensembling system 104 operates in accordance with one or more embodiments. As illustrated, the computing system 100 includes one or more server device(s) 102 connected to a client device 108 and therapeutics analysis device(s) 114 via a network 112. While FIG. 1 shows an embodiment of the pathogenicity score ensembling system 104, this disclosure describes alternative embodiments and configurations below.
[0054] As shown in FIG. 1, the server device(s) 102, the client device 108, and the therapeutics analysis device(s) 114 are connected via the network 112. Accordingly, each of the components of the computing system 100 can communicate via the network 112. The network 112 comprises any suitable network over which computing devices can communicate. Example networks are discussed in additional detail below with respect to FIG. 18.
[0055] As indicated by FIG. 1, the therapeutics analysis device(s) 114 comprises a device for analyzing (and identifying candidate therapeutics for) amino-acid sequences corresponding to proteins and / or nucleotide sequences representing coding and non-coding genomic regions. In some embodiments, the therapeutics analysis device(s) 114 analyzes a set of amino-acid sequencesor a set of nucleotide sequences from a database comprising samples exhibiting genetic diversity. From among the analyzed set of amino-acid sequences and / or analyzed set of nucleotide sequences, the therapeutics analysis device(s) 114 can identify subsets of amino-acid sequences and / or nucleotide sequences exhibiting common variant amino acids or variant nucleotides. In combination with or separate from such variant identification, the therapeutics analysis device(s) 114 can execute machine-learning models (or other models) that identify coding or non-coding genomic regions that are intolerant to variation and for which variants can cause loss or change in biological functions. For identified subsets of amino-acid sequences and / or nucleotide sequences, in some cases, the therapeutics analysis device(s) 114 identifies candidate biologies, drugs, or geneediting protocols for treatment.
[0056] In addition, or in the alternative to communicating across the network 112, in some embodiments, the therapeutics analysis device(s) 114 bypasses the network 112 and communicates directly with the server device(s) 102 or the client device 108. Additionally, as shown in FIG. 1, in one or more embodiments, the therapeutics analysis device(s) 114 includes the pathogenicity score ensembling system 104.
[0057] As further indicated by FIG. 1, the server device(s) 102 may generate, receive, analyze, store, and transmit digital data, such as data for amino-acid sequences or nucleotide sequences. As shown in FIG. 1, the therapeutics analysis device(s) 114 may send (and the server device(s) 102 may receive) various data from the therapeutics analysis device(s) 114, including data representing amino-acid sequences or nucleotide sequences. The server device(s) 102 may also communicate with the client device 108. In particular, the server device(s) 102 can send data representing aminoacid sequences or nucleotide sequences (or variants thereof), pathogenicity scores, or combined pathogenicity scores, to the client device 108.
[0058] Additionally, as shown in FIG. 1 , the server device(s) 102 can include the pathogenicity score ensembling system 104. In one or more embodiments, as explained further below, the pathogenicity score ensembling system 104 generates and ensembles pathogenicity scores across a number of variant amino-acid sequences and trains a variant pathogenicity machine-learning model to generate such scores. As shown, in some cases, the pathogenicity score ensembling system 104 can or include a variant pathogenicity machine-learning model 106 that generates pathogenicity score based on processing input data, such as artificial amino-acid sequences that include benign amino-acid variants in flanking regions around a target protein position with a reference aminoacid sequence. Likewise, in one or more embodiments, the pathogenicity score ensembling system 104 trains the variant pathogenicity machine-learning model 106 to generate pathogenicity scores using a unique training process. In addition to identifying or generating such pathogenicity scores, the server device(s) 102 can also send data representing combined pathogenicity scores ensembledfrom multiple initial pathogenicity scores. The figures depicted herein and paragraphs below further illustrate such functionalities of the pathogenicity score ensembling system 104 and / or the variant pathogenicity machine-learning model 106.
[0059] In addition, or in the alternative, to executing the variant pathogenicity machinelearning model 106, in some embodiments, the pathogenicity score ensembling system 104 accesses a database or table comprising pathogenicity scores. For example, in certain embodiments, the pathogenicity score ensembling system 104 identifies a pathogenicity score by identifying a score within a table for a particular protein, a target protein position, and a target amino acid at the target protein position. Accordingly, such a table or database may organize pathogenicity scores according to protein, position, and target amino acid at the position. Consistent with the disclosure above and below, the table or database includes pathogenicity scores that have been precomputed by the variant pathogenicity machine-learning model 106 and which are combinable to form combined pathogenicity scores.
[0060] In some embodiments, the server device(s) 102 comprise a distributed collection of servers where the server device(s) 102 include a number of server devices distributed across the network 112 and located in the same or different physical locations. Further, the server device(s) 102 can comprise a content server, an application server, a communication server, a web-hosting server, or another type of server.
[0061] In some cases, the server device(s) 102 is located at or near a same physical location of the therapeutics analysis device(s) 114 or remotely from the therapeutics analysis device(s) 114. Indeed, in some embodiments, the server device(s) 102 and the therapeutics analysis device(s) 114 are integrated into a same computing device. The server device(s) 102 may run software on the therapeutics analysis device(s) 114 or the pathogenicity score ensembling system 104 to generate, receive, analyze, store, and transmit digital data, such as by sending or receiving data representing amino-acid sequences or nucleotide sequences (or variants thereof), pathogenicity scores, or combined pathogenicity scores. Additionally or alternatively, in some embodiments, the therapeutics analysis device(s) 114 or the pathogenicity score ensembling system 104 store and access a database or table of pathogenicity scores corresponding to particular proteins and / or protein positions.
[0062] As further illustrated and indicated in FIG. 1, the client device 108 can generate, store, receive, and send digital data. In particular, the client device 108 can receive data for amino-acid sequences or nucleotide sequences (or variants thereof), pathogenicity scores, or combined pathogenicity scores from the server device(s) 102 and / or the therapeutics analysis device(s) 114. The client device 108 can accordingly present data concerning pathogenicity scores within a graphical user interface to a user associated with the client device 108.
[0063] The client device 108 illustrated in FIG. 1 may comprise various types of client devices. For example, in some embodiments, the client device 108 includes non-mobile devices, such as desktop computers or servers, or other types of client devices. In yet other embodiments, the client device 108 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones. Additional details with regard to the client device 108 are discussed below with respect to FIG. 18.
[0064] As further illustrated in FIG. 1, the client device 108 includes a client application 110. The client application 110 may be a web application or a native application stored and executed on the client device 108 (e.g., a mobile application, desktop application). The client application 110 can be an analytics application that includes instructions that (when executed) cause the client device 108 to receive data from the pathogenicity score ensembling system 104 and present data from the therapeutics analysis device(s) 114 and / or the server device(s) 102. Furthermore, the client application 110 can instruct the client device 108 to display data for pathogenicity scores, such as data for combined pathogenicity scores in a user interface.
[0065] As further illustrated in FIG. 1, the pathogenicity score ensembling system 104 may be located on the client device 108 as part of the client application 110 or on the therapeutics analysis device(s) 114. Accordingly, in some embodiments, the pathogenicity score ensembling system 104 is implemented by (e.g., located entirely or in part) on the client device 108. As mentioned, in yet other embodiments, the pathogenicity score ensembling system 104 is implemented by one or more other components of the computing system 100, such as the therapeutics analysis device(s) 114. In particular, the pathogenicity score ensembling system 104 can be implemented in a variety of different ways across the server device(s) 102, the network 112, the client device 108, and the therapeutics analysis device(s) 114.
[0066] Though FIG. 1 illustrates the components of the computing system 100 communicating via the network 112, in certain implementations, the components of computing system 100 can also communicate directly with each other, bypassing the network. For instance, and as previously mentioned, in some implementations, the client device 108 communicates directly with the therapeutics analysis device(s) 114. Additionally, in some embodiments, the client device 108 communicates directly with the pathogenicity score ensembling system 104. Moreover, the pathogenicity score ensembling system 104 can access one or more databases housed on or accessed by the server device(s) 102 or elsewhere in the computing system 100.
[0067] As indicated above, in one or more embodiments, the pathogenicity score ensembling system 104 generates a combined pathogenicity score by ensembling over multiple pathogenicity scores. In particular, the pathogenicity score ensembling system 104 generates multiple pathogenicity scores for a candidate amino acid at a target protein position using a variant pathogenicity machine-learning model to process artificial amino-acid sequences (and / or naturalamino-acid sequences) that include benign primate amino-acid variants and the candidate amino acid. FIG. 2 illustrates an example overview of ensembling pathogenicity scores to generate a combined pathogenicity score from artificial amino-acid sequences in accordance with one or more embodiments. Additional detail regarding the various acts, models, and principles described in FIG. 2 is provided in relation to subsequent figures thereafter.
[0068] As illustrated in FIG. 2, the pathogenicity score ensembling system 104 generates or accesses artificial amino-acid sequences 208. More particularly, the pathogenicity score ensembling system 104 accesses primate data, such as a benign sequence database 202 that stores primate amino-acid sequences identified by ILLUMINA’ s PrimateAI system or PrimateAI scores described by Sundaram or Gao. For example, the pathogenicity score ensembling system 104 identifies benign primate amino-acid sequences from the benign sequence database 202, including sequences that are known to be benign and / or satisfy a threshold probability of being benign.
[0069] As also illustrated in FIG. 2, the pathogenicity score ensembling system 104 generates the artificial amino-acid sequences 208 by combining benign amino-acid variants from benign primate amino-acid sequences with a reference amino-acid sequence 204. To elaborate, the pathogenicity score ensembling system 104 identifies or receives a reference amino-acid sequence 204 as a stored amino-acid sequence for a human and / or as a sample from a particular human or other organism. To generate the artificial amino-acid sequences 208, the pathogenicity score ensembling system 104 replaces reference amino acids at various coordinates or locations within the reference amino-acid sequence 204 with benign amino-acid variants from benign primate amino-acid sequences.
[0070] Specifically, the pathogenicity score ensembling system 104 replaces amino acids in positions adjacent to, and / or within flanking regions of, a target protein position within the reference amino-acid sequence 204. In some cases, the pathogenicity score ensembling system 104 determines candidate benign positions (e.g., candidate locations or coordinates) within the reference amino-acid sequence 204 where benign variant amino acids can be substituted for reference amino acids. For instance, the pathogenicity score ensembling system 104 determines a set of candidate benign primate amino-acid variants that correspond to, or include data indicating, candidate benign positions within flanking regions of a target protein position of the reference amino-acid sequence 204. In addition, the pathogenicity score ensembling system 104 substitutes benign primate amino-acid variants for reference amino acids at the candidate benign positions (e.g., within the flanking regions). By performing such amino-acid substitutions, the pathogenicity score ensembling system 104 generates the artificial amino-acid sequences 208.
[0071] In some embodiments, the pathogenicity score ensembling system 104 generates an artificial amino-acid sequence by combining benign amino-acid variants from a benign primateamino-acid sequence with a natural amino-acid sequence 206. For example, the pathogenicity score ensembling system 104 replaces amino acids at candidate benign positions within the natural amino-acid sequence 206 (at locations in flanking regions of a target protein position with respect to the reference amino-acid sequence 204) by substituting candidate benign amino acids from the benign sequence database 202. Specifically, the pathogenicity score ensembling system 104 substitutes a benign primate variant for a reference amino acid within the natural amino-acid sequence 206. The pathogenicity score ensembling system 104 can repeat this process to generate artificial amino-acid sequences from various natural amino-acid sequences.
[0072] As further illustrated in FIG. 2, the pathogenicity score ensembling system 104 generates initial pathogenicity scores 212 from the artificial amino-acid sequences 208. More specifically, the pathogenicity score ensembling system 104 utilizes the artificial amino-acid sequences 208 as input into a variant pathogenicity machine-learning model 210. The variant pathogenicity machine-learning model 210, in turn, processes the artificial amino-acid sequences 208 to generate the initial pathogenicity scores 212 for candidate amino acids at target protein positions within the artificial amino-acid sequences 208. For instance, the variant pathogenicity machine-learning model 210 generates a first pathogenicity score for a first candidate variant amino acid at the target protein position within a first artificial amino-acid sequence. Likewise, the variant pathogenicity machine-learning model 210 generates a second pathogenicity score for a second candidate variant amino acid at the target position within the first artificial amino-acid sequence. Continuing, in some embodiments, the pathogenicity score ensembling system 104 applies the variant pathogenicity machine-learning model 210 to generate a pathogenicity score for each candidate variant amino acid at the target position within the first artificial amino-acid sequence (e.g., for a total of nineteen pathogenicity scores for the first artificial amino-acid sequence).
[0073] Because the artificial amino-acid sequences include different set of benign amino-acid variants at different adjacent protein positions flanking the target protein position, the pathogenicity score ensembling system 104 can also determine pathogenicity scores for a same candidate amino acid at a same target protein position — but in the context of different artificial amino-acid sequences. For instance, the variant pathogenicity machine-learning model 210 generates a pathogenicity score for the first candidate variant amino acid at the target protein position within a second artificial amino-acid sequence differing from the first artificial amino-acid sequence. Likewise, the variant pathogenicity machine-learning model 210 generates another pathogenicity score for the first candidate variant amino acid at the target position within a third artificial aminoacid sequence, etc.
[0074] Similarly, the pathogenicity score ensembling system 104 applies the variant pathogenicity machine-learning model 210 to generate a set of (nineteen) pathogenicity scores (onefor each candidate variant amino acid at the target position) for each artificial amino-acid sequence from among the artificial amino-acid sequences 208. Accordingly, the pathogenicity score ensembling system 104 generates the initial pathogenicity scores 212 in sets, one set per artificial amino-acid sequence. Indeed, the pathogenicity score ensembling system 104 generates a large number of the initial pathogenicity scores 212 over which to ensemble for robust, unbiased pathogenicity predictions of variant amino-acid sequences.
[0075] As further illustrated in FIG. 2, the pathogenicity score ensembling system 104 generates a combined pathogenicity score 214 from the initial pathogenicity scores 212. To elaborate, the pathogenicity score ensembling system 104 ensembles over the initial pathogenicity scores 212 for a candidate amino acid at a target protein position to combine them into the combined pathogenicity score 214 for the candidate amino acid at the target protein position within the reference amino-acid sequence 204. For example, the pathogenicity score ensembling system 104 generates or determines the combined pathogenicity score 214 by determining a mean, a weighted mean, a max, a min, a median, or some other combined form of the initial pathogenicity scores 212 for the particular candidate amino acid at the target protein position.
[0076] To illustrate, in some embodiments, the pathogenicity score ensembling system 104 utilizes the variant pathogenicity machine-learning model 210 to generate a first pathogenicity score for a candidate variant amino acid at a target protein position within the first artificial aminoacid sequence, a second pathogenicity score for the candidate variant amino acid at the target protein position within a second artificial amino-acid sequence, and an nth pathogenicity score for the candidate variant amino acid at the target protein position within an nth artificial amino-acid sequence. The pathogenicity score ensembling system 104 further determines a mean or median pathogenicity score of the first through nth pathogenicity score for the candidate variant amino acid at the target protein position to generate a combined pathogenicity score for the candidate variant amino acid at the target protein position. Because the pathogenicity score ensembling system 104 can determine pathogenicity scores for each candidate amino acid at the target protein position, the pathogenicity score ensembling system 104 can likewise determine candidate-amino-acid-specific combined pathogenicity scores for each of 19 different candidate amino acids at the target protein position based on each such candidate amino acid within the context of the artificial amino-acid sequences 208.
[0077] As mentioned above, in certain described embodiments, the pathogenicity score ensembling system 104 generates or accesses artificial amino-acid sequences to use for generating pathogenicity scores. In particular, the pathogenicity score ensembling system 104 generates an artificial amino-acid sequence by substituting benign primate amino-acid variants for reference amino acids at adjacent protein positions within flanking regions of a target protein position of areference amino-acid sequence. FIG. 3 illustrates generating artificial amino-acid sequences from a reference amino-acid sequence and benign primate amino acids in accordance with one or more embodiments.
[0078] As illustrated in FIG. 3, the pathogenicity score ensembling system 104 identifies or accesses a benign sequence database 302. In particular, the benign sequence database 302 includes or stores benign primate amino-acid variants within primate amino-acid sequences (e.g., from a PrimateAI database). In addition, the pathogenicity score ensembling system 104 accesses or receives a reference amino-acid sequence 306 that includes reference amino acids at respective protein locations. For example, the reference amino-acid sequence 306 is an amino-acid sequence from a human sample. As shown, the pathogenicity score ensembling system 104 also determines or identifies a target protein position (on the “R” or arginine position), along with flanking regions in either side of the target position within the reference amino-acid sequence 306. In some cases, the flanking regions span a particular number of amino acids (or protein positions), such as five amino acids, ten amino acids, or fifty amino acids (or some other number). In other cases, the flanking regions may span at least a first threshold number of amino acids (or protein positions) while fewer than a second threshold number of amino acids (or protein positions).
[0079] As further illustrated in FIG. 3, the pathogenicity score ensembling system 104 generates artificial amino-acid sequences 308 by combining the reference amino-acid sequence 306 and amino acids from the benign sequence database 302. For instance, the pathogenicity score ensembling system 104 generates a first artificial amino-acid sequence by substituting a first set of benign primate amino-acid variants for reference amino acids at a first set of adjacent protein positions flanking the target protein position. In addition, the pathogenicity score ensembling system 104 generates a second (and a third and so on) artificial amino-acid sequence by substituting a second set of benign primate amino-acid variants for reference amino acids at a second set of adjacent protein positions flanking the target protein position.
[0080] In some cases, the pathogenicity score ensembling system 104 can determine different adjacent protein positions flanking the target protein position as locations for inserting or substituting benign variant amino acids for reference amino acid. Indeed, as shown, the artificial amino-acid sequences 308 include different benign variant amino acids substituted at different locations within the flanking regions.
[0081] As part of generating the artificial amino-acid sequences 308, the pathogenicity score ensembling system 104 determines or identifies candidate benign positions (indicated by the dashed boxes in FIG. 3) within the reference amino-acid sequence 306 as locations where variant amino acids can be located or inserted. For instance, the pathogenicity score ensembling system 104 can identify or determine a candidate benign position within a flanking region of a target proteinposition. To identify a candidate benign position, the pathogenicity score ensembling system 104 can determine a probability that an amino-acid variant at a protein position within a flanking region is benign based on known data for the protein and / or amino-acid variant (e.g., from primate data). The pathogenicity score ensembling system 104 can further compare the probability with a probability threshold to designate a protein position within a flanking region as a candidate benign position if the probability of being benign satisfies the threshold (e.g., more than 95%).
[0082] The pathogenicity score ensembling system 104 can thus select, for substitution, candidate benign primate amino-acid variants (from the benign sequence database 302) that correspond (or map) to the candidate benign positions. As shown, the pathogenicity score ensembling system 104 generates the artificial amino-acid sequences 308 by replacing the amino acids in the candidate positions (indicated by the dashed boxes) with benign variant amino acids from the benign sequence database 302 (e.g., from PrimateAI amino-acid sequences).
[0083] As mentioned above, in certain described embodiments, the pathogenicity score ensembling system 104 can determine or identify natural variant amino-acid sequences to use as input for a variant pathogenicity machine-learning model. In particular, the pathogenicity score ensembling system 104 can generate pathogenicity scores for natural variant amino-acid sequences and / or for natural variant amino-acid sequences combined with benign primate amino-acid sequences. FIG. 4 illustrates an example diagram for generating to extracting natural variant amino-acid sequences from a multiple sequence alignment (MSA) in accordance with one or more embodiments.
[0084] As illustrated in FIG. 4, the pathogenicity score ensembling system 104 identifies or accesses a reference amino-acid sequence 402. In particular, the pathogenicity score ensembling system 104 identifies a reference amino-acid sequence 402 that includes an amino acid located at a target protein position (the “P” or proline position), along with flanking regions on either side of the target protein position. As indicated above, the target regions can have a certain size (number of amino acids or protein positions) or can be within a certain size range.
[0085] As further illustrated in FIG. 4, the pathogenicity score ensembling system 104 accesses or analyzes a multiple sequence alignment 404. The multiple sequence alignment 404 can refer to a sequence alignment of three or more (four in this case) amino-acid sequences that indicates amino acids at various protein positions for each sequence in the alignment. In particular, the pathogenicity score ensembling system 104 analyzes the multiple sequence alignment 404 that includes natural amino-acid sequences corresponding to a set of species (e.g., human, primate, and / or non-primate species). Indeed, the multiple sequence alignment 404 includes amino-acid sequences that naturally occur and / or have been sampled from and / or observed in one or more species.
[0086] From the multiple sequence alignment 404, the pathogenicity score ensembling system 104 determines or identifies natural variant amino-acid sequences 406 that correspond to the reference amino-acid sequence 402. To elaborate, the pathogenicity score ensembling system 104 identifies sequences from the multiple sequence alignment 404 that are candidates to be included in the natural variant amino-acid sequences 406 used for generating pathogenicity scores for the reference amino-acid sequence 402. For instance, the pathogenicity score ensembling system 104 removes, from the multiple sequence alignment 404 (or from consideration as input for a variant pathogenicity machine-learning model), those sequences that include a variant amino acid at the target protein position of the reference amino-acid sequence 402 (e.g., those that do not have a reference amino acid at the target protein position). As shown in FIG. 4, the pathogenicity score ensembling system 104 accordingly removes the third amino-acid sequence listed in the multiple sequence alignment 404 (as indicated by the “X” symbol) for including a variant at the target position.
[0087] In addition, the pathogenicity score ensembling system 104 selects the remaining three sequences to include in the natural variant amino-acid sequences 406. Indeed, the pathogenicity score ensembling system 104 identifies the natural variant amino-acid sequences 406 as sequences in the multiple sequence alignment 404 that include natural benign amino-acid variants acids at candidate positions (indicated by the dashed boxes) within the flanking regions (but not at the target position). The pathogenicity score ensembling system 104 can thus utilize the natural variant amino-acid sequences 406 as sequences for generating pathogenicity scores using a variant pathogenicity machine-learning model.
[0088] As just noted, in some embodiments, the pathogenicity score ensembling system 104 utilizes natural variant amino-acid sequences as inputs for generating pathogenicity scores. In particular, the pathogenicity score ensembling system 104 can use natural variant amino-acid sequences in the alternative to, or in combination with, artificial amino-acid sequences comprising benign amino-acid variants (e.g., benign primate amino-acid variants) substituted in a reference amino-acid sequence. For example, the pathogenicity score ensembling system 104 can generate an artificial amino-acid sequence by combining benign primate amino-acid variants with a natural variant amino-acid sequence. FIG. 5 illustrates an example diagram for generating an artificial amino-acid sequence using a natural variant amino-acid sequence and benign primate amino-acid variants in accordance with one or more embodiments.
[0089] As illustrated in FIG. 5, the pathogenicity score ensembling system 104 accesses or identifies a multiple sequence alignment 502. More specifically, as described above, the pathogenicity score ensembling system 104 analyzes the multiple sequence alignment 502 to identify and extract a natural variant amino-acid sequence 504 that corresponds to a referenceamino-acid sequence. For instance, the pathogenicity score ensembling system 104 identifies the natural variant amino-acid sequence 504 as an amino-acid sequence (naturally occurring in one or more species) that includes variant amino acids at candidate positions within flanking regions of a target protein position.
[0090] As further illustrated in FIG. 5, the pathogenicity score ensembling system 104 identifies or accesses a benign primate variant 506 for combining with the natural variant aminoacid sequence 504. In particular, the pathogenicity score ensembling system 104 identifies or extracts the benign primate variant 506 from among a set of benign primate amino-acid variants (e.g., amino-acid variants identified as benign by PrimateAI scores described by Sundaram or Gao) within the benign sequence database 508. The pathogenicity score ensembling system 104 thus selects the benign primate variant 506 that includes a variant amino acid indicating (or including data defining or corresponding to) a candidate benign position within the natural variant aminoacid sequence 504 (and within a flanking region of a target protein position of a reference aminoacid sequence).
[0091] As further illustrated in FIG. 5, the pathogenicity score ensembling system 104 combines the natural variant amino-acid sequence 504 and the benign primate variant 506 to generate an artificial amino-acid sequence 510. To elaborate, the pathogenicity score ensembling system 104 generates the artificial amino-acid sequence 510 by substituting or inserting the benign primate variant 506 in place of a reference amino acid within the natural variant amino-acid sequence 504 (e.g., at a candidate, adjacent protein position in a flanking region). As but one example shown in FIG. 5, the pathogenicity score ensembling system 104 generates the artificial amino-acid sequence 510 by replacing the “A” or alanine amino acid in the candidate position indicated by the dashed box with an “E” or glutamate amino acid from the benign primate variant 506.
[0092] Beyond the example shown in FIG. 5, the pathogenicity score ensembling system 104 can substitute multiple benign primate variants at different adjacent protein positions. Particularly, the pathogenicity score ensembling system 104 can identify a set of benign primate amino-acid variants at a set of adjacent protein positions within a flanking region of a target protein position. In addition, the pathogenicity score ensembling system 104 can substitute one or more of the benign primate variants at their own respective adjacent protein positions within one or more flanking regions of a target protein position. The pathogenicity score ensembling system 104 can additionally perform such substitutions for multiple amino-acid sequences, substituting benign primate variants at adjacent protein positions to generate artificial amino-acid sequences.
[0093] As mentioned above, in one or more embodiments, the pathogenicity score ensembling system 104 generates a combined pathogenicity score by ensembling over pathogenicity scoresassociated with respective candidate amino acids within the context of different amino-acid sequences. In particular, the pathogenicity score ensembling system 104 generates pathogenicity scores using a variant pathogenicity machine-learning model and ensembles the pathogenicity scores to generate a combined pathogenicity score for respective candidate amino acids at different target protein positions. FIG. 6 illustrates an example diagram for generating a combined pathogenicity score by ensembling pathogenicity scores in accordance with one or more embodiments.
[0094] As illustrated in FIG. 6, the pathogenicity score ensembling system 104 generates an initial pathogenicity score 608 using a variant pathogenicity machine-learning model 606. More specifically, the pathogenicity score ensembling system 104 inputs an artificial amino-acid sequence 602 (generated or accessed as described herein) into the variant pathogenicity machinelearning model 606, whereupon the variant pathogenicity machine-learning model 606 generates the initial pathogenicity score 608 (e.g., as a score from 0 to 1) indicating a probability or a likelihood of a candidate variant amino acid at a target protein position being pathogenic (or benign). As shown, in some embodiments, the pathogenicity score ensembling system 104 inputs, into the variant pathogenicity machine-learning model 606, a natural amino-acid sequence 604 (e.g., an amino-acid sequence naturally occurring in one or more species) comprising the candidate variant amino acid at the target protein position, and the variant pathogenicity machine-learning model 606 generates the initial pathogenicity score 608. Indeed, the variant pathogenicity machinelearning model 606 can generate a pathogenicity score for a candidate variant amino acid within the context of an artificial amino-acid sequence and / or from a natural amino-acid sequence.
[0095] In some embodiments, the variant pathogenicity machine-learning model 606 generates the initial pathogenicity score 608 by processing data according to its model architecture. For example, in some cases, the variant pathogenicity machine-learning model 606 is an autoregressive model, such as a transformer neural network based on bidirectional encoder representations from transformers (BERT). Specifically, the variant pathogenicity machine-learning model 606 can be a BERT-style transformer protein language model, such as ESM2 model, ESMfold model, ESM- 1b model, and / or ESM-lv model. In certain embodiments, the variant pathogenicity machinelearning model 606 has a different architecture, such as a variational autoencoder, a multilayer perceptron, a recurrent neural network, a long short-term memory, or a decision tree model.
[0096] In one or more embodiments, the variant pathogenicity machine-learning model 606 generates the initial pathogenicity score 608 using a masked-filling-the-blank (MFB) approach. To elaborate, the pathogenicity score ensembling system 104 generates pathogenicity scores by masking an amino acid (obfuscating its identity) at a target protein position and applying the variant pathogenicity machine-learning model 606 to predict a pathogenicity of different candidate aminoacids placed in the masked (target) position. The pathogenicity score ensembling system 104 can repeat the MFB process for different target protein positions and / or for different reference aminoacid sequences, masking positions to predict using the variant pathogenicity machine-learning model 606. By using the MFB approach, in some embodiments, the variant pathogenicity machinelearning model 606 generates nineteen different pathogenicity scores from a single input (e.g., a single variant amino-acid sequence), one for each possible alternative amino acid at the target protein position.
[0097] As further illustrated in FIG. 6, the pathogenicity score ensembling system 104 repeats the process of generating pathogenicity scores. More particularly, the pathogenicity score ensembling system 104 identifies additional artificial and / or natural amino-acid sequences to input into the variant pathogenicity machine-learning model 606. In turn, the variant pathogenicity machine-learning model 606 generates corresponding initial pathogenicity scores for each of the input artificial and / or natural amino-acid sequences. In some cases, depending on its architecture (e.g., as a transformer neural network or a variational autoencoder), the variant pathogenicity machine-learning model 606 can generate multiple pathogenicity scores for single iteration or application (e.g., by generating a set of nineteen pathogenicity scores for all possible amino-acid variants for a given input sequence and / or by ingesting multiple amino-acid sequences as input and generating corresponding pathogenicity score outputs). In other cases, the variant pathogenicity machine-learning model 606 generates a single pathogenicity score per iteration or application, and the pathogenicity score ensembling system 104 thus reapplies the variant pathogenicity machinelearning model 606 to generate multiple pathogenicity scores.
[0098] After generating initial pathogenicity scores, as further shown in FIG. 6, the pathogenicity score ensembling system 104 can perform an ensembling 610 to combine or ensemble pathogenicity scores (e.g., the initial pathogenicity score 608 and other initial pathogenicity scores) for generating the combined pathogenicity score 612. For example, the pathogenicity score ensembling system 104 ensembles over a plurality of pathogenicity scores for a candidate amino acid at a target protein position (e.g., a plurality of sets of pathogenicity scores, where each set includes nineteen scores generated from its own variant amino-acid sequence). In some embodiments, the pathogenicity score ensembling system 104 ensembles by determining a mean of pathogenicity scores for a given target protein position. In these or other embodiments, the pathogenicity score ensembling system 104 ensembles by determining a weighted mean by weighting pathogenicity scores differently according to the type of their respective source sequences (e.g., weighting pathogenicity scores generated from artificial amino-acid sequences higher than pathogenicity scores generated from natural amino-acid sequences). As other examples of the ensembling 610, the pathogenicity score ensembling system 104 can determine a maximum,a minimum, a median, or some other combination of pathogenicity scores to generate the combined pathogenicity score 612 for a given candidate amino acid at a target protein position.
[0099] As just mentioned, in some embodiments, the pathogenicity score ensembling system 104 generates a combined pathogenicity score by ensembling over multiple pathogenicity scores. In particular, in some cases, the pathogenicity score ensembling system 104 generates pathogenicity scores based on logits associated with variant amino-acid sequences and further combines logits to generate a combined pathogenicity score. FIG. 7 illustrates an example diagram for generating a combined pathogenicity score based on logits in accordance with one or more embodiments.
[0100] As illustrated in FIG. 7, in some embodiments, the pathogenicity score ensembling system 104 generates pathogenicity scores for candidate amino acids at target protein positions within variant amino-acid sequences. Indeed, as described above, the pathogenicity score ensembling system 104 utilizes a variant pathogenicity machine-learning model to generate pathogenicity scores in the form of logit differences. For example, as shown, the pathogenicity score ensembling system 104 generates or determines a logit difference 706 based on an artificial amino-acid sequence 702 for a protein and further generates or determines a logit difference 708 based on an artificial amino-acid sequence 704 for the protein.
[0101] As part of determining a logit difference (e.g., the logit difference 706 or the logit difference 708), the pathogenicity score ensembling system 104 determines a difference between a first logit and a second logit. Indeed, the pathogenicity score ensembling system 104 determines a first logit for a candidate variant amino acid (e.g., from the artificial amino-acid sequence 702) at a target position using a variant pathogenicity machine-learning model. The first logit can be represented by: log(pait), where pattrepresents a probability of the candidate variant — or an alternative variant — being pathogenic (or benign) at a target protein position. In addition, the pathogenicity score ensembling system 104 generates or determines a second logit for a reference amino acid (e.g., from a reference amino-acid sequence) at the target position using the variant pathogenicity machine-learning model. The second logit can be represented by: log, where Pref represents a probability of the reference amino acid being pathogenic (or benign) at the target protein position within a reference amino-acid sequence. Using the first and second logits, the pathogenicity score ensembling system 104 further determines pathogenicity scores in the form of the logit difference 706 and the logit difference 708, which each indicate differences between respective logits (corresponding to respective sequences), as shown in FIG. 7.
[0102] As further shown, the pathogenicity score ensembling system 104 generates a combined pathogenicity score 710. Indeed, the pathogenicity score ensembling system 104 generates the combined pathogenicity score 710 by ensembling over the logit difference 706, thelogit difference 708, and / or any further logit differences based on logits corresponding to the relevant candidate amino acid at the target protein position. For instance, the pathogenicity score ensembling system 104 determines a mean of the logit difference 706 and the logit difference 708 as the combined pathogenicity score 710. In certain embodiments, the pathogenicity score ensembling system 104 determines the combined pathogenicity score 710 in the form of a weighted mean, a max, a min, or a median of the logit differences.
[0103] As mentioned above, in certain embodiments, the pathogenicity score ensembling system 104 generates artificial amino-acid sequences to use for generating pathogenicity scores. In particular, the pathogenicity score ensembling system 104 generates artificial amino-acid sequences by substituting benign primate variant amino acids at candidate positions in flanking regions of a target protein position, as described above. In some cases, the pathogenicity score ensembling system 104 generates artificial amino-acid sequences according to an r value which represents a ratio of candidate benign positions to fill with benign primate variants. FIGS. 8A-8B illustrate example experimental results (e.g., using an ESM2 BERT-style transformer architecture with masked filling the blank) demonstrating the effects of different ratios of candidate benign positions used for generating artificial amino-acid sequences in accordance with one or more embodiments.
[0104] As illustrated in FIG. 8A, a graph 802 depicts experimental results for cell-line experiments for Saturation Mutagenesis. Saturation Mutagenesis can refer to a protein engineering technique that involves mutating a protein using a random mutagenesis process. In some cases, saturation mutagenesis involves substituting a single codon or a set of codons with all possible amino acids at a given protein position, which can result in site saturation at every protein position in the protein for a library of size [20 x (the number of residues in the protein)]. Experimenters tested different ratios of candidate benign positions (or r values), including more candidate benign positions in some experiments than others. As shown, the x-axis reflects changes in r value ratios across two different types of amino-acid sequence (and a third with no perturbations for substituting benign variant amino acids in candidate positions). As the r value changes, or as experimenters test amino-acid sequences with different numbers of candidate positions for substituting benign primate variants, the Spearman mean (or Spearman coefficient) also changes. In some cases, the Spearman mean indicates or reflects a measure of correlation between two variables (e.g., pathogenicity predictions and r values) — e.g., how well the relationship between the two variables can be described using a monotonic function — where higher values indicate a stronger correlation.
[0105] Accordingly, as shown, r values of around 0.5 result in strong correlation with pathogenicity predictions for the assay for both types of amino-acid sequences. Indeed, the depicted results are for benign sequences (which show little change) and for unknown sequenceswith different ratios of candidate positions. Accordingly, r values of 0.5 represent a good ratio of candidate positions for substituting benign primate variants to result in high correlation with (and better accuracy in predicting) pathogenicity scores for the assay. As r values increase closer to 1 (e.g., 100% of the sequence is made up of candidate benign positions), the correlation weakens for unknown sequences, and the prediction of pathogenicity scores becomes less accurate. In some cases, the pathogenicity score ensembling system 104 exhibits a 1.8% improvement on the assay relative to existing pathogenicity prediction models by taking a mean over ensembled pathogenicity scores.
[0106] As illustrated in FIG. 8B, a graph 804 depicts experimental results for amino-acid sequences that are part of the United Kingdom Biobank (UKBB), which is a long-term study beginning in the United Kingdom in 2006 investigating respective contributions of genetic predispositions and environmental exposure to the development of disease. The UKBB study, and its corresponding database, includes data from approximately 500,000 volunteers enrolled at ages forty to sixty-nine years of age and who are monitored for at least thirty years. For example, experimenters tested performance of the pathogenicity score ensembling system 104 over different candidate benign position ratios (or r values). Specifically, the graph 804 depicts results for two different types of amino-acid sequence (along with a third amino-acid sequence with no benign variant substitutions), such as benign amino-acid sequences and unknown amino-acid sequences. As mentioned, the different r values indicate different ratios of candidate benign positions for substituting benign primate variants. Thus, similar to the graph 802 in FIG. 8A, the graph 804 indicates that the pathogenicity score ensembling system 104 generates accurate pathogenicity scores for r values around 0.5 with decreases in performance for r values that approach 1 (where increasing r values means increases in number of benign variants used). In some cases, the pathogenicity score ensembling system 104 provides a 1.6% improvement over existing systems on UKBB by ensembling over pathogenicity scores to determine a combined (e.g., mean) pathogenicity score.
[0107] As mentioned above, in certain described embodiments, the pathogenicity score ensembling system 104 generates a combined pathogenicity score by ensembling over initial pathogenicity scores from both artificial amino-acid sequences and natural amino-acid sequences. Experimenters have demonstrated that increasing the total number of amino-acid sequences can improve the accuracy of determining a combined pathogenicity score by providing increased diversity and more data for ensembling. FIGS. 9A-9B illustrate example graphs of experimental results depicting the impact of changing the total number of amino-acid sequences for generating a combined pathogenicity score in accordance with one or more embodiments.
[0108] As illustrated in FIG. 9A, a graph 902 depicts experimental results for different total numbers of amino-acid sequences when determining a combined pathogenicity score across sequences from the UKBB dataset. For example, the graph 902 depicts a Spearman mean (or Spearman coefficient) that reflects or indicates a correlation between combined pathogenicity scores and numbers of amino-acid sequences. Thus, a higher Spearman mean indicates a stronger correlation and better combined pathogenicity scores. As shown, experimenters demonstrated that increasing the number of amino-acid sequences from which to ensemble pathogenicity scores generally improves the reliability of generating combined pathogenicity scores for both natural amino-acid sequences and for artificial amino-acid sequences generated by combining natural amino acids with benign variant amino acids. In some cases, the pathogenicity score ensembling system 104 exhibits the most improved accuracy (strongest correlation) at 900 sequences, for a 2.6% improvement on UKBB data relative to existing MFB-based systems.
[0109] As illustrated in FIG. 9B, a graph 904 depicts experimental results for different total numbers of amino-acid sequences when determining a combined pathogenicity score for a cell-line experiments for Saturation Mutagenesis. For example, the graph 904 depicts a Spearman mean (or Spearman coefficient) that reflects or indicates a correlation between combined pathogenicity scores and numbers of amino-acid sequences. As shown, experimenters demonstrated that increasing the number of amino-acid sequences from which to ensemble pathogenicity scores generally improves the reliability of generating combined pathogenicity scores for both natural amino-acid sequences and for artificial amino-acid sequences generated by combining natural amino acids with benign variant amino acids. In some cases, the pathogenicity score ensembling system 104 exhibits the most improved accuracy (strongest correlation) at 1225 sequences for a 3.7% improvement on the assay relative to existing MFB-based systems.
[0110] As mentioned above, in one or more embodiments, the pathogenicity score ensembling system 104 trains a variant pathogenicity machine-learning model to generate pathogenicity scores. In particular, the pathogenicity score ensembling system 104 trains a variant pathogenicity machine-learning model to generate pathogenicity score from artificial and / or natural amino-acid sequences. FIG. 10 illustrates an example diagram for training a variant pathogenicity machinelearning model in accordance with one or more embodiments.[OHl] As illustrated in FIG. 10, the pathogenicity score ensembling system 104 accesses or identifies a training amino-acid-sequence pair 1004 from a database 1002. More specifically, the pathogenicity score ensembling system 104 identifies a training amino-acid-sequence pair 1004 that includes two amino-acid sequences, such as a natural amino-acid sequence and an unknown amino-acid sequence. Instead of natural and unknown amino-acid sequences as part of the training amino-acid-sequence pair 1004, in some cases, the training amino-acid-sequence pair 1004includes a reference (or canonical) amino-acid sequence and a modified amino-acid sequence that includes amino-acid variants (e.g., benign primate variants) substituted in the reference sequence at one or more protein positions. For instance, the modified amino-acid sequence includes amino acids from a benign primate variants substituted in positions of a benign amino-acid sequence, a non-benign amino-acid sequence, or an unknown amino-acid sequence.
[0112] In addition, the pathogenicity score ensembling system 104 inputs the training amino- acid-sequence pair 1004 into a variant pathogenicity machine-learning model 1006. Specifically, the pathogenicity score ensembling system 104 inputs the training amino-acid-sequence pair 1004 in a particular order or sequence, inputting one training sequence of the training amino-acid- sequence pair 1004 before the other. The variant pathogenicity machine-learning model 1006, in turn, processes the training amino-acid-sequence pair 1004 to generate two predicted pathogenicity likelihoods (or pathogenicity scores), a predicted pathogenicity likelihood 1008 and a predicted pathogenicity likelihood 1010. In some cases, the predicted pathogenicity likelihood 1008 represents a probability or a likelihood that an amino acid (e.g., a reference amino acid or a variant amino acid) is pathogenic or benign when located at a target protein position within a first aminoacid sequence (e.g., the natural amino-acid sequence or the reference amino-acid sequence) within the training amino-acid-sequence pair 1004. Similarly, the predicted pathogenicity likelihood 1010 represents a probability or a likelihood that the amino acid (e.g., the reference amino acid or the variant amino acid) is pathogenic or benign when located at the target protein position within a second amino-acid sequence (e.g., the unknown amino-acid sequence or the modified amino-acid sequence) within the training amino-acid-sequence pair 1004. In some cases, the variant pathogenicity machine-learning model 1006 generates the predicted pathogenicity likelihoods 1008 and 1010 in parallel by co-processing the training amino-acid-sequence pair 1004, while in other embodiments the pathogenicity score ensembling system 104 generates one of the predicted pathogenicity likelihoods 1008 or 1010 before the other by processing the sequences of the training amino-acid-sequence pair 1004 in series.
[0113] As also illustrated in FIG. 10, the pathogenicity score ensembling system 104 performs a comparison 1012. More particularly, the pathogenicity score ensembling system 104 performs the comparison 1012 to compare the predicted pathogenicity likelihoods 1008 and 1010 for the training amino-acid-sequence pair 1004 with ground truth data from the database 1002. For instance, the pathogenicity score ensembling system 104 utilizes a loss function to compare the predicted pathogenicity likelihood 1008 and the predicted pathogenicity likelihood 1010 with a ground-truth-ordering value 1014. Indeed, the pathogenicity score ensembling system 104 identifies or accesses the ground-truth-ordering value 1014 from the database 1002 that stores training data that indicates the actual or observed order in which the sequences of the trainingamino-acid-sequence pair 1004 were provided or input into the variant pathogenicity machinelearning model 1006. In some cases, the ground-truth-ordering value 1014 indicates a 0 label or 1 label. In some such embodiments, the 0 label indicates a first order in which a natural amino-acid sequence or a reference amino-acid sequence is respectively input before an unknown amino-acid sequence or a modified amino-acid sequence; and the 1 label indicates a second order in which an unknown amino-acid sequence or a modified amino-acid sequence is respectively input before a natural amino-acid sequence or a reference amino-acid sequence.
[0114] In some embodiments, the pathogenicity score ensembling system 104 performs the comparison 1012 by utilizing a particular pathogenicity prediction loss function, such as a binary cross entropy (BCE) loss function. Indeed, in some such cases, the pathogenicity score ensembling system 104 uses a loss function to determine measures of loss in the form of pseudo log likelihoods. Using the pathogenicity prediction loss function, the pathogenicity score ensembling system 104 compares the predicted pathogenicity likelihoods with the ground-truth-ordering value 1014 to determine a measure of loss or error associated with the predictions. Because the amino-acid sequences that make up the training amino-acid-sequence pair 1004 have different compositions, they also have different predicted pathogenicity likelihoods, where one predicted pathogenicity likelihood for an amino acid at a target protein position (e.g., within a natural amino-acid sequence or a reference amino-acid sequence) is closer to 0, and the other predicted pathogenicity likelihood for the amino acid at the target protein position (e.g., within an unknown amino-acid sequence or a modified amino-acid sequence) is closer to 1. Indeed, using the loss function, the pathogenicity score ensembling system 104 generates or determines two measures of loss, one for each sequence within the training amino-acid-sequence pair (or for each of the predicted pathogenicity likelihoods).
[0115] In some embodiments the pathogenicity score ensembling system 104 generates the predicted pathogenicity likelihood 1008 and / or the predicted pathogenicity likelihood 1010 as scores, such as a probability between 0 and 1. For example, the pathogenicity score ensembling system 104 utilizes the variant pathogenicity machine-learning model 1006 to generates a score for each sequence of the training amino-acid-sequence pair 1004 representing a probability (e.g., between 0 and 1) that the sequence is pathogenic (or benign). In some cases, the pathogenicity score ensembling system 104 utilizes benign primate data (e.g., as identified by PrimateAI scores described by Sundaram or Gao) to generate the scores or probabilities for the predicted pathogenicity likelihood 1008 and the predicted pathogenicity likelihood 1010. In some embodiments, the pathogenicity score ensembling system 104 converts the scores to (or determines the scores in the form of) pseudo log likelihoods (PLLs) to use within the loss function of the comparison 1012.
[0116] As further illustrated in FIG. 10, the pathogenicity score ensembling system 104 performs a parameter adjustment 1016. In particular, the pathogenicity score ensembling system 104 adjusts, updates, or modifies internal parameters of the variant pathogenicity machine-learning model 1006, such as weights and biases, that impact how the variant pathogenicity machinelearning model 1006 processes or analyzes data across its various layers, neurons, or other architectural components. In some cases, the pathogenicity score ensembling system 104 adjusts the parameters of the variant pathogenicity machine-learning model 1006 to reduce a measure of loss associated with the comparison 1012 for subsequent training iterations.
[0117] In one or more implementations, the pathogenicity score ensembling system 104 repeats the process illustrated in FIG. 10 to train the variant pathogenicity machine-learning model 1006. To elaborate, the pathogenicity score ensembling system 104 inputs training amino-acid- sequence pairs, generates predicted pathogenicity likelihoods, compares the predictions with ground-truth-ordering values, and adjusts model parameters at each iteration. Thus, over a number of training iterations or epochs, the pathogenicity score ensembling system 104 iteratively adjusts parameters of the variant pathogenicity machine-learning model 1006 to reduce one or more measures of loss associated with a pathogenicity prediction loss function (e.g., until satisfying a threshold measure of loss and / or by performing a threshold number of training iterations). In some cases, the variant pathogenicity machine-learning model 1006 leams parameters that distinguish between different types of amino-acid sequences included within training amino-acid-sequence pairs by separating the losses. For instance, as training progresses, the losses associated with natural amino-acid sequences (or reference amino-acid sequences) over the iterations form a distribution closer to 0, while the losses associated with unknown amino-acid sequences (or modified amino-acid sequences) over the iterations form a distribution closer to 1.
[0118] As just mentioned, in one or more embodiments, the pathogenicity score ensembling system 104 trains a variant pathogenicity machine-learning model to accurately generate pathogenicity scores by distinguishing between such scores based on different types of amino-acid sequences. In particular, the pathogenicity score ensembling system 104 leams parameters of a variant pathogenicity machine-learning model to separate losses associated with different types of amino-acid sequences input in specific orders as part of training amino-acid-sequence pairs. FIGS. 11A-11B illustrate diagrams for learning parameters of a variant pathogenicity machine-learning model to separate loss values for two different training implementations. Specifically, FIG. 11A illustrates results for a first training implementation using training amino-acid-sequence pairs comprised of natural amino-acid sequences and unknown amino-acid sequences. Thereafter, FIG. 11B illustrates results for a second training implementation using training amino-acid-sequence pairs comprised of reference amino-acid sequences and modified amino-acid sequences.
[0119] As illustrated in FIG. 11 A, the pathogenicity score ensembling system 104 identifies a training amino-acid sequence pair that includes a natural amino-acid sequence 1102 and an unknown amino-acid sequence 1104. In addition, the pathogenicity score ensembling system 104 inputs or inserts the natural amino-acid sequence 1102 and the unknown amino-acid sequence 1104 into a variant pathogenicity machine-learning model 1106. Specifically, the pathogenicity score ensembling system 104 inputs the sequences into the variant pathogenicity machine-learning model 1106 in a particular order, where either the natural amino-acid sequence 1102 or the unknown amino-acid sequence 1104 is input first.
[0120] Based on training the variant pathogenicity machine-learning model 1106 using a loss function, as described above, the pathogenicity score ensembling system 104 leams parameters of the variant pathogenicity machine-learning model 1106 to distinguish between the natural aminoacid sequence 1102 and the unknown amino-acid sequence 1104. Indeed, as shown, the variant pathogenicity machine-learning model 1106 generates loss distributions throughout the training process, one distribution for natural amino-acid sequences (including the natural amino-acid sequence 1102) and another distribution for unknown amino-acid sequences (including the unknown amino-acid sequence 1104). As indicated, the patterned triangles beneath the natural amino-acid sequence 1102 and the unknown amino-acid sequence 1104 indicate positions where variant amino acids are located, and they correspond to the same patterns found in the loss distributions 1108 for pseudo log likelihoods. As shown in FIG. 11 A, the loss distribution for the natural amino-acid sequence 1102 is closer to 0, and the loss distribution for the unknown aminoacid sequence 1104 is closer to 1.
[0121] As illustrated in FIG. 1 IB, the pathogenicity score ensembling system 104 identifies a training amino-acid sequence pair that includes a reference amino-acid sequence 1110 (including no substituted benign variants) and a modified amino-acid sequence 1112 that include substituted benign variants. In addition, the pathogenicity score ensembling system 104 inputs or inserts the reference amino-acid sequence 1110 and the modified amino-acid sequence 1112 into a variant pathogenicity machine-learning model 1114. Specifically, the pathogenicity score ensembling system 104 inputs the sequences into the variant pathogenicity machine-learning model 1114 in a particular order, where either the reference amino-acid sequence 1110 or the modified amino-acid sequence 1112 is input first.
[0122] Based on training the variant pathogenicity machine-learning model 1114 using a loss function, as described above, the pathogenicity score ensembling system 104 leams parameters of the variant pathogenicity machine-learning model 1106 to distinguish between the reference amino-acid sequence 1110 and the modified amino-acid sequence 1112. Indeed, as shown, the variant pathogenicity machine-learning model 1106 generates loss distributions 1116 throughoutthe training process, one distribution for reference amino-acid sequences (including the reference amino-acid sequence 1110) and another distribution for modified amino-acid sequences (including the modified amino-acid sequence 1112). As indicated, the patterned triangles beneath the modified amino-acid sequence 1112 correspond to the same patterns found in the loss distributions 1116 for pseudo log likelihoods, and they indicate locations of variant amino acids. As shown in FIG. 11B, the loss distribution for the reference amino-acid sequence 1110 is closer to 0, and the loss distribution for the modified amino-acid sequence 1112 is closer to 1.
[0123] As mentioned, in certain described embodiments, the pathogenicity score ensembling system 104 trains a variant pathogenicity machine-learning model to generate pathogenicity scores. In particular, the pathogenicity score ensembling system 104 can utilize a training process that incorporates a particular loss function that uses or relies on a certain loss prefactor that impacts the learning of the variant pathogenicity machine-learning model. FIGS. 12A-12B illustrate example graphs depicting training results (e.g., for a GREMLIN-based Potts model available at https : / / github .com / sokrypton / GREMLIN_CPP / blob / master / GREMLIN_TF . ipynb) of using different loss prefactors in accordance with one or more embodiments.
[0124] As illustrated in FIG. 12 A, a graph 1202 depicts plot points for results of training using different loss prefactor values (or weights / weight multipliers) associated with different training data, where the training data includes amino-acid sequences with different fractions of substituted variants (e.g., benign primate variants), or r values (“Fraction Injected”). For example, as shown, the graph 1202 depicts or plots mean values for different training tests (e.g., ensembled pathogenicity scores generated by determining a mean across initial pathogenicity scores). Indeed, the graph 1202 depicts the relationship between mean pathogenicity scores for amino-acid sequences from the UKBB as well as amino-acid sequences from the assay of cell-line experiments for Saturation Mutagenesis, across a number of different variant pathogenicity machine-learning models trained on data with varying numbers of injected or substituted benign primate variants. As shown, the best performance from among the training tests is using a 0.06 loss prefactor which results in approximately 8.2% improvement in pathogenicity score predictions over existing systems for the assay and 4% for the UKBB data.
[0125] As illustrated in FIG. 12B, a graph 1204 depicts plot points for results of different loss prefactor values for median pathogenicity scores. As mentioned, the training data for the tests includes amino-acid sequences with different fractions of substituted variants (e.g., benign primate variants), or r values (“Fraction Injected”). For example, as shown, the graph 1204 depicts or plots median values for different training tests (e.g., ensembled pathogenicity scores generated by determining a median across pathogenicity scores). Similar to the discussion of FIG. 12A, the graph 1204 depicts the relationship between median pathogenicity scores for amino-acid sequencesfrom the UKBB as well as amino-acid sequences from the assay of cell-line experiments for Saturation Mutagenesis, across a number of different variant pathogenicity machine-learning models trained on data with varying numbers of injected or substituted benign primate variants.
[0126] As mentioned above, in certain embodiments, the pathogenicity score ensembling system 104 performs a stratification process as part of training a variant pathogenicity machinelearning model. In particular, the pathogenicity score ensembling system 104 leams to stratify input amino-acid sequences according to loss values associated with their predicted pathogenicity likelihoods (or pathogenicity scores). FIG. 13 illustrates an example graph depicting loss stratification as part of training a variant pathogenicity machine-learning model in accordance with one or more embodiments.
[0127] As illustrated in FIG. 13, a graph 1302 depicts a relationship between density of pathogenicity prediction likelihoods (or pathogenicity scores) and Potts model scores from a protein dependent Potts model based on GREMLIN. In particular, the graph 1302 depicts Potts scores that indicate the level of surety or confidence associated with model predictions. Indeed, the pathogenicity score ensembling system 104 determines loss values associated with predicted pathogenicity likelihoods for a variant pathogenicity machine-learning model and stratifies the losses by separating them to distinguish between input sequence types (e.g., natural vs. unknown or reference vs. modified). As shown, smaller (e.g., more negative) Potts model scores indicated clearer, easier predictions of pathogenicity and thus indicate more pathogenic amino-acid sequences. Conversely, larger (e.g., less negative) Potts model scores indicate less clear, more difficult predictions of pathogenicity and thus indicate les pathogenic amino-acid sequences.
[0128] To generate the stratified results of the graph 1302, the pathogenicity score ensembling system 104 performs a stratification process that involves multiple actions. For example, as part of the stratification process, the pathogenicity score ensembling system 104 identifies or selects loss values generated from all non-benign amino-acid sequences (e.g., sequences that include all possible twenty amino acids minus those that are known to be benign or satisfy a threshold probability of being benign). In addition, the pathogenicity score ensembling system 104 divides the non-benign amino-acid sequences into two groups: smallest and largest halves according to the Potts model evaluation and plots the losses for the two groups in the graph 1302. Accordingly, the pathogenicity score ensembling system 104 stratifies the groups to identify or distinguish between types of amino-acid sequences input into a variant pathogenicity machine-learning model in pairs, such as a pair of a natural amino-acid sequence with an unknown amino-acid sequence or a pair of a reference with a modified sequence.
[0129] As shown in FIG. 13, the smaller Potts values (or the smaller losses) correspond to sequences with clearer, more pathogenic predictions while the larger Potts values (or the largerlosses) correspond to sequences with less clear, less pathogenic predictions. In some cases, the pathogenicity score ensembling system 104 further rescales the loss by a particular rescaling loss prefactor (e.g., loss * 570 / protein length) to leam model parameters that generalize between the Saturation Mutagenesis assay and the UKBB data.
[0130] As just mentioned, in some embodiments, the pathogenicity score ensembling system 104 trains a variant pathogenicity machine-learning model using a stratified loss together with a loss prefactor. In particular, the pathogenicity score ensembling system 104 utilizes a specific loss prefactor that results in good performance, as determined by experimenters testing different prefactors for training a variant pathogenicity machine-learning model. FIGS 14A-14B illustrate experimental results of training a variant pathogenicity machine-learning model using stratified loss with different loss prefactors in accordance with one or more embodiments.
[0131] As illustrated in FIG. 14A, a graph 1402 depicts results for training a variant pathogenicity machine-learning model to generate pathogenicity scores from data within the UKBB. More specifically, the graph 1402 depicts how different loss prefactors impact the training of pathogenicity score prediction for variant pathogenicity machine-learning models trained using loss stratification. Indeed, as the protein length changes for different training implementations, the loss prefactor based on the protein length (e.g., 570 / protein length) also changes, and the resulting variant pathogenicity machine-learning model exhibits differences in performance based on training using the different loss prefactors. As mentioned above, the Spearman mean indicates a relatedness between two variables, such as the loss prefactor and predicted pathogenicity scores. Accordingly, as shown in the graph 1402, the performance of a trained variant pathogenicity machine-learning model comes from contrasting the natural amino-acid variants against the most pathogenic variants (e.g., the least benign variants), highlighting the importance of identifying high quality benign (primate) variants for loss stratification.
[0132] As illustrated in FIG. 14B, a graph 1404 depicts results for training a variant pathogenicity machine-learning model to generate pathogenicity scores for the Saturation Mutagenesis assay. As with the graph 1402, the graph 1404 depicts how different loss prefactors impact the training of pathogenicity score prediction for variant pathogenicity machine-learning models trained using loss stratification. Indeed, as the protein length changes for different training implementations, the loss prefactor also changes, and the resulting variant pathogenicity machinelearning model exhibits differences in performance based on training using the different loss prefactors. Accordingly, as shown in the graph 1404, the performance of a trained variant pathogenicity machine-learning model comes from contrasting the natural amino-acid variants against the most pathogenic variants (e.g., the least benign variants), highlighting the importance of identifying high quality benign (primate) variants for loss stratification.
[0133] As mentioned above, in certain described embodiments, the pathogenicity score ensembling system 104 trains a variant pathogenicity machine-learning model using loss stratification and a loss prefactor. In particular, the pathogenicity score ensembling system 104 utilizes a loss prefactor that modifies a loss associated with a predicted pathogenicity likelihood to adjust parameters of the variant pathogenicity machine-learning model for increased accuracy. FIGS. 15A-15B illustrate example experimental results for training a variant pathogenicity machine-learning model using different loss prefactors and loss stratification in accordance with one or more embodiments.
[0134] As illustrated in FIG. 15 A, a graph 1502 depicts plot points for results of training using different loss prefactors for stratified amino-acid sequences (e.g., all, largest, and smallest). Specifically, the graph 1502 depicts plot points for training a variant pathogenicity machinelearning model to generate pathogenicity scores which are ensembled to determine a mean pathogenicity score. As shown, relative to existing systems, the pathogenicity score ensembling system 104 exhibits a 3.85% performance on UKBB data and a 6.9% improvement on the Saturation Mutagenesis data for mean pathogenicity scores.
[0135] As illustrated in FIG. 15B, a graph 1504 depicts plot points for results of training using different loss prefactors for stratified amino-acid sequences (e.g., all, largest, and smallest). Specifically, the graph 1504 depicts plot points for training a variant pathogenicity machinelearning model to generate pathogenicity scores which are ensembled to determine a median pathogenicity score. As shown, relative to existing systems, the pathogenicity score ensembling system 104 exhibits an 8% performance on UKBB data and a 9.1% improvement on the Saturation Mutagenesis assay for median pathogenicity scores.
[0136] Turning now to FIG. 16, this figure illustrates a flowchart of a series of acts 1600 of generating a combined pathogenicity score by ensembling over pathogenicity scores generated for amino-acid sequences using a variant pathogenicity machine-learning model in accordance with one or more embodiments. While FIG. 16 illustrates acts according to one embodiment, alternative embodiments may omit, add to, reorder, and / or modify any of the acts shown in FIG. 16. The acts of FIG. 16 can be performed as part of a method. Alternatively, anon-transitory computer readable storage medium can comprise instructions that, when executed by one or more processors, cause a computing device or a system to perform the acts depicted in FIG. 16. In still further embodiments, a system comprising at least one processor and a non-transitory computer readable medium comprising instructions that, when executed by one or more processors, cause the system to perform the acts of FIG. 16.
[0137] As shown in FIG. 16, the series of acts 1600 include an act 1602 of accessing a first amino-acid sequence and a second amino-acid sequence. In addition, the series of acts 1600includes an act 1604 of generating a first pathogenicity score for the first amino-acid sequence. The series of acts 1600 also includes an act 1606 of generating a second pathogenicity score for the second amino-acid sequence. Further, the series of acts 1600 includes an act 1608 of ensembling the first pathogenicity score and the second pathogenicity score. In some embodiments, the series of acts 1600 includes acts to perform any of the operations described in the following clauses:CLAUSE 1. A method comprising: accessing, for a target protein, a first amino-acid sequence comprising a first set of benign amino-acid variants at adjacent protein positions flanking a target protein position and a second amino-acid sequence comprising a second set of benign amino-acid variants at adjacent protein positions flanking the target protein position, the second amino-acid sequence differing from the first amino-acid sequence; generating, utilizing a variant pathogenicity machine-learning model, a first pathogenicity score for a candidate variant amino acid at the target protein position within the first amino-acid sequence; generating, utilizing the variant pathogenicity machine-learning model, a second pathogenicity score for the candidate variant amino acid at the target protein position within the second amino-acid sequence; and ensembling the first pathogenicity score and the second pathogenicity score to generate a combined pathogenicity score for the candidate variant amino acid at the target protein position.CLAUSE 2. The method of clause 1, wherein generating the first amino-acid sequence and the second amino-acid sequence comprises: identifying a first set of benign primate amino-acid variants at a first set of adjacent protein positions flanking the target protein position and a second set of benign primate amino-acid variants at a second set of adjacent protein positions flanking the target protein position; generating a first artificial amino-acid sequence by substituting the first set of benign primate amino-acid variants for reference amino acids at the first set of adjacent protein positions of a reference amino-acid sequence for the target protein; and generating a second artificial amino-acid sequence by substituting the second set of benign primate amino-acid variants for reference amino acids at the second set of adjacent protein positions of the reference amino-acid sequence for the target protein.CLAUSE 3. The method of any of clauses 1 -2, wherein generating the first artificial aminoacid sequence and the second artificial amino-acid sequence comprises: determining a set of candidate benign primate amino-acid variants at a corresponding set of candidate benign positions within the reference amino-acid sequence at which benign variant amino acids can be substituted in for reference amino acids;selecting, from among the set of candidate benign primate amino-acid variants at the corresponding set of candidate benign positions, the first set of benign primate amino-acid variants at the first set of adjacent protein positions for the first artificial amino-acid sequence; and selecting, from among the set of candidate benign primate amino-acid variants at the corresponding set of candidate benign positions, the second set of benign primate amino-acid variants at the second set of adjacent protein positions for the second artificial amino-acid sequence.CLAUSE 4. The method of any of clauses 1-3, further comprising randomly selecting the first set of benign primate amino-acid variants at the first set of adjacent protein positions and the second set of benign primate amino-acid variants at the second set of adjacent protein positions according to a ratio of candidate benign positions used for artificial amino-acid sequences.CLAUSE 5. The method of any of clauses 1-4, further comprising identifying the first set of benign primate amino-acid variants and the second set of benign primate amino-acid variants by accessing a database of primate amino-acid sequences comprising variant amino acids that are likely benign to a function of the target protein at a given protein position from the first set of adjacent protein positions or the second set of adjacent protein positions.CLAUSE 6. The method of any of clauses 1-5, wherein generating the first amino-acid sequence and the second amino-acid sequence comprises: identifying, from a multiple sequence alignment corresponding to a set of species, a first natural variant amino-acid sequence comprising a first set of natural benign amino-acid variants at a first set of adjacent protein positions flanking the target protein position; and identifying, from the multiple sequence alignment, a second natural variant amino-acid sequence comprising a second set of natural benign amino-acid variants at a second set of adjacent protein positions flanking the target protein position.CLAUSE 7. The method of any of clauses 1-6, further comprising: removing, from the multiple sequence alignment, one or more natural variant amino-acid sequences comprising a variant amino acid at the target protein position to generate a candidate set of natural variant amino-acid sequences; and selecting, from among the candidate set of natural variant amino-acid sequences, the first natural variant amino-acid sequence and the second natural variant amino-acid sequence.CLAUSE 8. The method of any of clauses 1-7, wherein generating the first amino-acid sequence and the second amino-acid sequence comprises: identifying a set of benign primate amino-acid variants at a set of adjacent protein positions flanking the target protein position;substituting a first benign primate variant of the set of benign primate amino-acid variants for a first reference amino acid at an adjacent protein position of the first set of adjacent protein positions within the first natural variant amino-acid sequence; and substituting a second benign primate variant of the set of benign primate amino-acid variants for a second reference amino acid at an adjacent protein position from the second set of adjacent protein positions within the second natural variant amino-acid sequence.CLAUSE 9. The method of any of clauses 1-8, further comprising: generating, utilizing the variant pathogenicity machine-learning model, a first set of pathogenicity scores for a set of candidate variant amino acids at the target protein position within the first amino-acid sequence, the set of candidate variant amino acids including the candidate variant amino acid and the first set of pathogenicity scores including the first pathogenicity score; generating, utilizing the variant pathogenicity machine-learning model, a second set of pathogenicity scores for the set of candidate variant amino acids at the target protein position within the second amino-acid sequence, the second set of pathogenicity scores including the second pathogenicity score; and ensembling, from the first set of pathogenicity scores and the second set of pathogenicity scores, respective pathogenicity scores for each candidate variant amino acid of the set of candidate variant amino acids to generate a combined pathogenicity score for each candidate variant amino acid.CLAUSE 10. The method of any of clauses 1-9, further comprising generating the first pathogenicity score and the second pathogenicity score using a trained transformer neural network as part of a masked-filling-the-blank approach to generate pathogenicity scores for candidate amino acids at the target protein position.CLAUSE 11. The method of any of clauses 1-10, wherein generating the combined pathogenicity score for the candidate variant amino acid at the target protein position comprises determining a mean pathogenicity score based on the first pathogenicity score and the second pathogenicity score.CLAUSE 12. The method of any of clauses 1-11, further comprising: generating the first pathogenicity score for the candidate variant amino acid by determining a first logit difference between a logit, generated by the variant pathogenicity machine-learning model, for the candidate variant amino acid at the target protein position within the first aminoacid sequence and a logit, generated by the variant pathogenicity machine-learning model, for a reference amino acid at the target protein position within the first amino-acid sequence; generating the second pathogenicity score for the candidate variant amino acid by determining a second logit difference between a logit, generated by the variant pathogenicitymachine-learning model, for the candidate variant amino acid at the target protein position within the second amino-acid sequence and a logit, generated by the variant pathogenicity machinelearning model, for the reference amino acid at the target protein position within the second aminoacid sequence; and generating the combined pathogenicity score for the candidate variant amino acid at the target protein position by determining a mean pathogenicity score based on the first logit difference and the second logit difference.
[0138] Turning now to FIG. 17, this figure illustrates a flowchart of a series of acts 1700 of training a variant pathogenicity machine-learning model to generate pathogenicity score for aminoacid sequences using a training amino-acid-sequence pair and a ground-truth-ordering value. While FIG. 17 illustrates acts according to one embodiment, alternative embodiments may omit, add to, reorder, and / or modify any of the acts shown in FIG. 17. The acts of FIG. 17 can be performed as part of a method. Alternatively, a non-transitory computer readable storage medium can comprise instructions that, when executed by one or more processors, cause a computing device or a system to perform the acts depicted in FIG. 17. In still further embodiments, a system comprising at least one processor and a non-transitory computer readable medium comprising instructions that, when executed by one or more processors, cause the system to perform the acts of FIG. 17.
[0139] As shown in FIG. 17, the series of acts 1700 include an act 1702 of providing a training amino-acid-sequence pair to a variant pathogenicity machine-learning model. For example, the series of acts 1700 includes an act 1702 of providing a training amino-acid-sequence pair to a variant pathogenicity machine-learning model. In addition, the series of acts 1700 includes an act 1704 of generating predicted pathogenicity likelihoods for the training amino-acid-sequence pair. Further, the series of acts 1700 includes an act 1706 of determining a training loss associated with the training amino-acid-sequence pair. Additionally, the series of acts 1700 includes an act 1708 of adjusting parameters of the variant pathogenicity machine-learning model (based on the training loss). In some embodiments, the series of acts 1700 includes acts to perform any of the operations described in the following clauses:CLAUSE 13: A method comprising: providing, to a variant pathogenicity machine-learning model, a training amino-acid- sequence pair comprising a first type of amino-acid sequence and a second type of amino-acid sequence for a target protein; generating, utilizing the variant pathogenicity machine-learning model, a first predicted pathogenicity likelihood for an amino acid at a target protein position within the first type of amino-acid sequence and a second predicted pathogenicity likelihood for the amino acid at the target protein position within the second type of amino-acid sequence; determining a training loss by utilizing a pathogenicity prediction loss function to compare the first predicted pathogenicity likelihood and the second predicted pathogenicity likelihood to a ground-truth-ordering value indicating an order in which the first type of amino-acid sequence and the second type of amino-acid sequence were provided to the variant pathogenicity machinelearning model; and adjusting parameters of the variant pathogenicity machine-learning model based on the training loss.CLAUSE 14. The method of clause 13, wherein: the first type of amino-acid sequence comprises a natural amino-acid sequence that naturally occurs in one or more species; and the second type of amino-acid sequence comprises an unknown amino-acid sequence that includes the natural amino-acid sequence with non-benign amino-acid variants substituted in for reference amino acids at one or more protein positions.CLAUSE 15. The method of any of clauses 13-14, wherein: the first type of amino-acid sequence comprises a reference amino-acid sequence for the target protein; and the second type of amino-acid sequence comprises a modified amino-acid sequence with amino-acid variants substituted in for reference amino acids at one or more protein positions.CLAUSE 16. The method of any of clauses 13-15, wherein: the ground-truth-ordering value comprises a first value indicating the first type of aminoacid sequence is provided to the variant pathogenicity machine-learning model before the second type of amino-acid sequence; or the ground-truth-ordering value comprises a second value indicating the second type of amino-acid sequence is provided to the variant pathogenicity machine-learning model before the first type of amino-acid sequence.CLAUSE 17. The method of any of clauses 13-16, further comprising: generating the first predicted pathogenicity likelihood by generating a first pseudo log likelihood for the amino acid at the target protein position within the first type of amino-acid sequence; and generating the second predicted pathogenicity likelihood by generating a second pseudo log likelihood for the amino acid at the target protein position within the second type of amino-acid sequence.CLAUSE 18. The method of any of clauses 13-17, further comprising:generating the first predicted pathogenicity likelihood by generating, utilizing the variant pathogenicity machine-learning model, a first pathogenicity score for the amino acid at the target protein position within the first type of amino-acid sequence; generating the second predicted pathogenicity likelihood by generating, utilizing the variant pathogenicity machine-learning model, a second pathogenicity score for the amino acid at the target protein position within the second type of amino-acid sequence; and determining the training loss by utilizing the pathogenicity prediction loss to compare the first pathogenicity score and the second pathogenicity score to the ground-truth-ordering value.CLAUSE 19. The method of any of clauses 13-18, further comprising: providing, to the variant pathogenicity machine-learning model and from training amino- acid-sequence pairs, the first type of amino-acid sequence and the second type of amino-acid sequence according to a randomly permuted order; iteratively determining training losses utilizing the pathogenicity prediction loss function to compare predicted pathogenicity likelihoods for amino acids in the training amino-acid-sequence pairs to corresponding ground-truth-ordering values; and iteratively adjusting, based on the training losses, the parameters of the variant pathogenicity machine-learning model to determine more accurate predicted pathogenicity likelihoods that indicate distinctions between the first type of amino-acid sequence and the second type of amino-acid sequence within the training amino-acid-sequence pairs.CLAUSE 20. The method of any of clauses 13-19, wherein the pathogenicity prediction loss function comprises a binary cross entropy loss function.CLAUSE 21. The method of any of clauses 13-20, wherein the variant pathogenicity machine-learning model comprises a transformer neural network, a convolutional neural network (CNN), a sequence-to-sequence model, a variational autoencoder (VAE), a multilayer perceptron (MLP), a recurrent neural network (RNN), a long short-term memory (LSTM), or a decision tree model.CLAUSE 22. The method of any of clauses 13-21, wherein the amino acid at the target protein position comprises: a reference amino acid at the target protein position within the first type of amino-acid sequence and the second type of amino-acid sequence; or a variant amino acid at the target protein position within the first type of amino-acid sequence and the second type of amino-acid sequence.
[0140] The components of the pathogenicity score ensembling system 104 can include software, hardware, or both. For example, the components of the pathogenicity score ensembling system 104 can include one or more instructions stored on a computer-readable storage mediumand executable by processors of one or more computing devices (e.g., the client device 108). When executed by the one or more processors, the computer-executable instructions of the pathogenicity score ensembling system 104 can cause the computing devices to perform the methods described herein. Alternatively, the components of the pathogenicity score ensembling system 104 can comprise hardware, such as special purpose processing devices to perform a certain function or group of functions. Additionally, or alternatively, the components of the pathogenicity score ensembling system 104 can include a combination of computer-executable instructions and hardware.
[0141] Furthermore, the components of the pathogenicity score ensembling system 104 performing the functions described herein with respect to the pathogenicity score ensembling system 104 may, for example, be implemented as part of a stand-alone application, as a module of an application, as a plug-in for applications, as a library function or functions that may be called by other applications, and / or as a cloud-computing model. Thus, components of the pathogenicity score ensembling system 104 may be implemented as part of a stand-alone application on a personal computing device or a mobile device. Additionally, or alternatively, the components of the pathogenicity score ensembling system 104 may be implemented in any application that provides sequencing services including, but not limited to Illumina PrimateAI, Illumina PrimateAIlD, Illumina PrimateAI2D, Illumina PrimateAI3D, or Illumina TruSight. “Illumina,” “PrimateAI,” “PrimateAIlD,” “PrimateAI2D,” “PrimateAI3D,” and “TruSight,” are either registered trademarks or trademarks of Illumina, Inc. in the United States and / or other countries.
[0142] Embodiments of the present disclosure may comprise or utilize a special purpose or general-purpose computer including computer hardware, such as, for example, one or more processors and system memory, as discussed in greater detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more of the processes described herein may be implemented at least in part as instructions embodied in anon-transitory computer-readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein). In general, a processor (e.g., a microprocessor) receives instructions, from a non-transitory computer-readable medium, (e.g., a memory, etc.), and executes those instructions, thereby performing one or more processes, including one or more of the processes described herein.
[0143] Computer-readable media can be any available media that can be accessed by a general purpose or special purpose computer system. Computer-readable media that store computerexecutable instructions are non-transitory computer-readable storage media (devices). Computer- readable media that carry computer-executable instructions are transmission media. Thus, by wayof example, and not limitation, embodiments of the disclosure can comprise at least two distinctly different kinds of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.
[0144] Non-transitory computer-readable storage media (devices) includes RAM, ROM, EEPROM, CD-ROM, solid state drives (SSDs) (e.g., based on RAM), Flash memory, phasechange memory (PCM), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer.
[0145] A “network” is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a computer, the computer properly views the connection as a transmission medium. Transmissions media can include a network and / or data links which can be used to carry desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer. Combinations of the above should also be included within the scope of computer- readable media.
[0146] Further, upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to non-transitory computer-readable storage media (devices) (or vice versa). For example, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface module (e.g., a NIC), and then eventually transferred to computer system RAM and / or to less volatile computer storage media (devices) at a computer system. Thus, it should be understood that non-transitory computer- readable storage media (devices) can be included in computer system components that also (or even primarily) utilize transmission media.
[0147] Computer-executable instructions comprise, for example, instructions and data which, when executed at a processor, cause a general-purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. In some embodiments, computer-executable instructions are executed on a general-purpose computer to turn the general-purpose computer into a special purpose computer implementing elements of the disclosure. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to beunderstood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.
[0148] Those skilled in the art will appreciate that the disclosure may be practiced in network computing environments with many types of computer system configurations, including, personal computers, desktop computers, laptop computers, message processors, hand-held devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, tablets, pagers, routers, switches, and the like. The disclosure may also be practiced in distributed system environments where local and remote computer systems, which are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network, both perform tasks. In a distributed system environment, program modules may be located in both local and remote memory storage devices.
[0149] Embodiments of the present disclosure can also be implemented in cloud computing environments. In this description, “cloud computing” is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be employed in the marketplace to offer ubiquitous and convenient on-demand access to the shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then scaled accordingly.
[0150] A cloud-computing model can be composed of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and so forth. A cloud-computing model can also expose various service models, such as, for example, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (laaS). A cloud-computing model can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, and so forth. In this description and in the claims, a “cloud-computing environment” is an environment in which cloud computing is employed.
[0151] FIG. 18 illustrates a block diagram of a computing device 1800 that may be configured to perform one or more of the processes described above. One will appreciate that one or more computing devices such as the computing device 1800 may implement the pathogenicity score ensembling system 104. As shown by FIG. 18, the computing device 1800 can comprise a processor 1802, a memory 1804, a storage device 1806, an I / O interface 1808, and a communication interface 1810, which may be communicatively coupled by way of a communication infrastructure 1812. In certain embodiments, the computing device 1800 caninclude fewer or more components than those shown in FIG. 18. The following paragraphs describe components of the computing device 1800 shown in FIG. 18 in additional detail.
[0152] In one or more embodiments, the processor 1802 includes hardware for executing instructions, such as those making up a computer program. As an example, and not by way of limitation, to execute instructions for dynamically modifying workflows, the processor 1802 may retrieve (or fetch) the instructions from an internal register, an internal cache, the memory 1804, or the storage device 1806 and decode and execute them. The memory 1804 may be a volatile or nonvolatile memory used for storing data, metadata, and programs for execution by the processor(s). The storage device 1806 includes storage, such as a hard disk, flash disk drive, or other digital storage device, for storing data or instructions for performing the methods described herein.
[0153] The I / O interface 1808 allows a user to provide input to, receive output from, and otherwise transfer data to and receive data from computing device 1800. The I / O interface 1808 may include a mouse, a keypad or a keyboard, a touch screen, a camera, an optical scanner, network interface, modem, other known I / O devices or a combination of such I / O interfaces. The I / O interface 1808 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O interface 1808 is configured to provide graphical data to a display for presentation to a user. The graphical data may be representative of one or more graphical user interfaces and / or any other graphical content as may serve a particular implementation.
[0154] The communication interface 1810 can include hardware, software, or both. In any event, the communication interface 1810 can provide one or more interfaces for communication (such as, for example, packet-based communication) between the computing device 1800 and one or more other computing devices or networks. As an example, and not by way of limitation, the communication interface 1810 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI.
[0155] Additionally, the communication interface 1810 may facilitate communications with various types of wired or wireless networks. The communication interface 1810 may also facilitate communications using various communication protocols. The communication infrastructure 1812 may also include hardware, software, or both that couples components of the computing device 1800 to each other. For example, the communication interface 1810 may use one or more networks and / or protocols to enable a plurality of computing devices connected by a particular infrastructure to communicate with each other to perform one or more aspects of the processes described herein. To illustrate, the sequencing process can allow a plurality of devices (e.g., a client device,sequencing device, and server device(s)) to exchange information such as sequencing data and error notifications.
[0156] In the foregoing specification, the present disclosure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of the present disclosure(s) are described with reference to details discussed herein, and the accompanying drawings illustrate the various embodiments. The description above and drawings are illustrative of the disclosure and are not to be construed as limiting the disclosure. Numerous specific details are described to provide a thorough understanding of various embodiments of the present disclosure.
[0157] The present disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. For example, the methods described herein may be performed with less or more steps / acts or the steps / acts may be performed in differing orders. Additionally, the steps / acts described herein may be repeated or performed in parallel with one another or in parallel with different instances of the same or similar steps / acts. The scope of the present application is, therefore, indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.
Claims
CLAIMSWe Claim:
1. A system comprising: at least one processor; and a non-transitory computer readable medium comprising instructions that, when executed by the at least one processor, cause the system to: access, for a target protein, a first amino-acid sequence comprising a first set of benign amino-acid variants at adjacent protein positions flanking a target protein position and a second amino-acid sequence comprising a second set of benign amino-acid variants at adjacent protein positions flanking the target protein position, the second amino-acid sequence differing from the first amino-acid sequence; generate, utilizing a variant pathogenicity machine-learning model, a first pathogenicity score for a candidate variant amino acid at the target protein position within the first amino-acid sequence; generate, utilizing the variant pathogenicity machine-learning model, a second pathogenicity score for the candidate variant amino acid at the target protein position within the second amino-acid sequence; and ensemble the first pathogenicity score and the second pathogenicity score to generate a combined pathogenicity score for the candidate variant amino acid at the target protein position.
2. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate the first amino-acid sequence and the second amino-acid sequence by: identifying a first set of benign primate amino-acid variants at a first set of adjacent protein positions flanking the target protein position and a second set of benign primate amino-acid variants at a second set of adjacent protein positions flanking the target protein position; generating a first artificial amino-acid sequence by substituting the first set of benign primate amino-acid variants for reference amino acids at the first set of adjacent protein positions of a reference amino-acid sequence for the target protein; and generating a second artificial amino-acid sequence by substituting the second set of benign primate amino-acid variants for reference amino acids at the second set of adjacent protein positions of the reference amino-acid sequence for the target protein.
3. The system of claim 2, further comprising instructions that, when executed by the at least one processor, cause the system to generate the first artificial amino-acid sequence and the second artificial amino-acid sequence by:determining a set of candidate benign primate amino-acid variants at a corresponding set of candidate benign positions within the reference amino-acid sequence at which benign variant amino acids can be substituted in for reference amino acids; selecting, from among the set of candidate benign primate amino-acid variants at the corresponding set of candidate benign positions, the first set of benign primate amino-acid variants at the first set of adjacent protein positions for the first artificial amino-acid sequence; and selecting, from among the set of candidate benign primate amino-acid variants at the corresponding set of candidate benign positions, the second set of benign primate amino-acid variants at the second set of adjacent protein positions for the second artificial amino-acid sequence.
4. The system of claim 3, further comprising instructions that, when executed by the at least one processor, cause the system to randomly select the first set of benign primate aminoacid variants at the first set of adjacent protein positions and the second set of benign primate amino-acid variants at the second set of adjacent protein positions according to a ratio of candidate benign positions used for artificial amino-acid sequences.
5. The system of claim 2, further comprising instructions that, when executed by the at least one processor, cause the system to identify the first set of benign primate amino-acid variants and the second set of benign primate amino-acid variants by accessing a database of primate amino-acid sequences comprising variant amino acids that are likely benign to a function of the target protein at a given protein position from the first set of adjacent protein positions or the second set of adjacent protein positions.
6. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate the first amino-acid sequence and the second amino-acid sequence by: identifying, from a multiple sequence alignment corresponding to a set of species, a first natural variant amino-acid sequence comprising a first set of natural benign amino-acid variants at a first set of adjacent protein positions flanking the target protein position; and identifying, from the multiple sequence alignment, a second natural variant amino-acid sequence comprising a second set of natural benign amino-acid variants at a second set of adjacent protein positions flanking the target protein position.
7. The system of claim 6, further comprising instructions that, when executed by the at least one processor, cause the system to: remove, from the multiple sequence alignment, one or more natural variant amino-acid sequences comprising a variant amino acid at the target protein position to generate a candidate set of natural variant amino-acid sequences; andselect, from among the candidate set of natural variant amino-acid sequences, the first natural variant amino-acid sequence and the second natural variant amino-acid sequence.
8. The system of claim 6, further comprising instructions that, when executed by the at least one processor, cause the system to generate the first amino-acid sequence and the second amino-acid sequence by: identifying a set of benign primate amino-acid variants at a set of adjacent protein positions flanking the target protein position; substituting a first benign primate variant of the set of benign primate amino-acid variants for a first reference amino acid at an adjacent protein position of the first set of adjacent protein positions within the first natural variant amino-acid sequence; and substituting a second benign primate variant of the set of benign primate amino-acid variants for a second reference amino acid at an adjacent protein position from the second set of adjacent protein positions within the second natural variant amino-acid sequence.
9. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: generate, utilizing the variant pathogenicity machine-learning model, a first set of pathogenicity scores for a set of candidate variant amino acids at the target protein position within the first amino-acid sequence, the set of candidate variant amino acids including the candidate variant amino acid and the first set of pathogenicity scores including the first pathogenicity score; generate, utilizing the variant pathogenicity machine-learning model, a second set of pathogenicity scores for the set of candidate variant amino acids at the target protein position within the second amino-acid sequence, the second set of pathogenicity scores including the second pathogenicity score; and ensemble, from the first set of pathogenicity scores and the second set of pathogenicity scores, respective pathogenicity scores for each candidate variant amino acid of the set of candidate variant amino acids to generate a combined pathogenicity score for each candidate variant amino acid.
10. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate the first pathogenicity score and the second pathogenicity score using a trained transformer neural network as part of a masked-filling-the- blank approach to generate pathogenicity scores for candidate amino acids at the target protein position.
11. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate the combined pathogenicity score for thecandidate variant amino acid at the target protein position by determining a mean pathogenicity score based on the first pathogenicity score and the second pathogenicity score.
12. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to: generate the first pathogenicity score for the candidate variant amino acid by determining a first logit difference between a logit, generated by the variant pathogenicity machine-learning model, for the candidate variant amino acid at the target protein position within the first aminoacid sequence and a logit, generated by the variant pathogenicity machine-learning model, for a reference amino acid at the target protein position within the first amino-acid sequence; generate the second pathogenicity score for the candidate variant amino acid by determining a second logit difference between a logit, generated by the variant pathogenicity machine-learning model, for the candidate variant amino acid at the target protein position within the second amino-acid sequence and a logit, generated by the variant pathogenicity machinelearning model, for the reference amino acid at the target protein position within the second aminoacid sequence; and generate the combined pathogenicity score for the candidate variant amino acid at the target protein position by determining a mean pathogenicity score based on the first logit difference and the second logit difference.
13. A system comprising: at least one processor; and a non-transitory computer readable medium comprising instructions that, when executed by the at least one processor, cause the system to: provide, to a variant pathogenicity machine-learning model, a training amino-acid- sequence pair comprising a first type of amino-acid sequence and a second type of aminoacid sequence for a target protein; generate, utilizing the variant pathogenicity machine-learning model, a first predicted pathogenicity likelihood for an amino acid at a target protein position within the first type of amino-acid sequence and a second predicted pathogenicity likelihood for the amino acid at the target protein position within the second type of amino-acid sequence; determine a training loss by utilizing a pathogenicity prediction loss function to compare the first predicted pathogenicity likelihood and the second predicted pathogenicity likelihood to a ground-truth-ordering value indicating an order in which the first type of amino-acid sequence and the second type of amino-acid sequence were provided to the variant pathogenicity machine-learning model; andadjust parameters of the variant pathogenicity machine-learning model based on the training loss.
14. The system of claim 13, wherein: the first type of amino-acid sequence comprises a natural amino-acid sequence that naturally occurs in one or more species; and the second type of amino-acid sequence comprises an unknown amino-acid sequence that includes the natural amino-acid sequence with non-benign amino-acid variants substituted in for reference amino acids at one or more protein positions.
15. The system of claim 13, wherein: the first type of amino-acid sequence comprises a reference amino-acid sequence for the target protein; and the second type of amino-acid sequence comprises a modified amino-acid sequence with amino-acid variants substituted in for reference amino acids at one or more protein positions.
16. The system of claim 13, wherein: the ground-truth-ordering value comprises a first value indicating the first type of aminoacid sequence is provided to the variant pathogenicity machine-learning model before the second type of amino-acid sequence; or the ground-truth-ordering value comprises a second value indicating the second type of amino-acid sequence is provided to the variant pathogenicity machine-learning model before the first type of amino-acid sequence.
17. The system of claim 13, further comprising instructions that, when executed by the at least one processor, cause the system to: generate the first predicted pathogenicity likelihood by generating a first pseudo log likelihood for the amino acid at the target protein position within the first type of amino-acid sequence; and generate the second predicted pathogenicity likelihood by generating a second pseudo log likelihood for the amino acid at the target protein position within the second type of amino-acid sequence.
18. The system of claim 13, further comprising instructions that, when executed by the at least one processor, cause the system to: generate the first predicted pathogenicity likelihood by generating, utilizing the variant pathogenicity machine-learning model, a first pathogenicity score for the amino acid at the target protein position within the first type of amino-acid sequence;generate the second predicted pathogenicity likelihood by generating, utilizing the variant pathogenicity machine-learning model, a second pathogenicity score for the amino acid at the target protein position within the second type of amino-acid sequence; and determine the training loss by utilizing the pathogenicity prediction loss function to compare the first pathogenicity score and the second pathogenicity score to the ground-truthordering value.
19. The system of claim 13, further comprising instructions that, when executed by the at least one processor, cause the system to: provide, to the variant pathogenicity machine-learning model and from training amino- acid-sequence pairs, the first type of amino-acid sequence and the second type of amino-acid sequence according to a randomly permuted order; iteratively determine training losses utilizing the pathogenicity prediction loss function to compare predicted pathogenicity likelihoods for amino acids in the training amino-acid-sequence pairs to corresponding ground-truth-ordering values; and iteratively adjust, based on the training losses, the parameters of the variant pathogenicity machine-learning model to determine more accurate predicted pathogenicity likelihoods that indicate distinctions between the first type of amino-acid sequence and the second type of aminoacid sequence within the training amino-acid-sequence pairs.
20. The system of claim 13, wherein the pathogenicity prediction loss function comprises a binary cross entropy loss function.
21. The system of claim 13 , wherein the variant pathogenicity machine-learning model comprises a transformer neural network, a convolutional neural network (CNN), a sequence-to- sequence model, a variational autoencoder (VAE), a multilayer perceptron (MLP), a recurrent neural network (RNN), a long short-term memory (LSTM), or a decision tree model.
22. The system of claim 13, wherein the amino acid at the target protein position comprises: a reference amino acid at the target protein position within the first type of amino-acid sequence and the second type of amino-acid sequence; or a variant amino acid at the target protein position within the first type of amino-acid sequence and the second type of amino-acid sequence.
Citation Information
Patent Citations
Efficient voxelization for deep learning
US20220336057A1
Pathogenicity language model
US20230207061A1
Methods of predicting pathogenicity of genetic sequence variants
US20160371431A1
Deep convolutional neural networks for variant classification
WO2019079180A1
Inter-model prediction score recalibration
WO2023129955A1