Inter-protein and intra-protein pathogenicity models

The pathogenicity prediction system combines inter-protein and intra-protein models to generate re-scaled scores, addressing accuracy and flexibility issues, and achieves efficient, accurate pathogenicity predictions across different protein contexts.

WO2026152025A1PCT designated stage Publication Date: 2026-07-16ILLUMINA INC

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ILLUMINA INC
Filing Date
2026-01-09
Publication Date
2026-07-16

AI Technical Summary

Technical Problem

Existing pathogenicity prediction models lack accuracy and flexibility across different protein contexts, fail to consider structural and non-human data, and are computationally inefficient, leading to inconsistent and resource-intensive predictions.

Method used

A pathogenicity prediction system that combines inter-protein and intra-protein models to generate re-scaled pathogenicity scores, using linear transformations to preserve mean and standard deviation, and determines pathogenicity by comparing observed and expected variant numbers within position windows.

Benefits of technology

Improves accuracy and computational efficiency by generating pathogenicity predictions that account for both intra-protein and inter-protein contexts, outperforming existing models across clinical benchmarks and reducing computation time from hours to milliseconds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2026010828_16072026_PF_FP_ABST
    Figure US2026010828_16072026_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure describes methods, non-transitory computer-readable media, and systems that generate pathogenicity predictions for amino acids. For example, the disclosed systems can access a set of inter-protein pathogenicity scores generated by an inter-protein pathogenicity model for a set of amino acids in a target protein sequence. In some instances, the disclosed systems generate the inter-protein pathogenicity scores based on an observed number of variants and an expected number of variants. The disclosed systems further access a set of intra-protein pathogenicity scores generated by an intra-protein pathogenicity model for the set of amino acids in the target protein sequence. Using the central tendency and standard deviation of the inter-protein pathogenicity scores, the disclosed systems generate re-scaled pathogenicity scores for the set of amino acids according to a ranking of the set of intra-protein pathogenicity scores.
Need to check novelty before this filing date? Find Prior Art

Description

INTER-PROTEIN AND INTRA-PROTEIN PATHOGENICITY MODELSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 743,957, entitled, “DETERMINING AMINO-ACID PATHOGENICITY BY RESCALING SCORES FROM A COMBINATION OF PROTEIN PATHOGENICITY MODELS,” filed on January 10, 2025 (IP-2880-PRV). The aforementioned application is hereby incorporated by reference in its entirety.BACKGROUND

[0002] In recent years, biotechnology firms and research institutions have improved software for predicting a pathogenicity of protein variants or genetic variants. For instance, some existing pathogenicity prediction models generate predictions that estimate a degree to which amino-acid variants are benign or pathogenic. Such pathogenicity predictions can indicate whether an aminoacid variant is likely to cause various diseases, such as certain cancers, developmental disorders, or heart conditions. In addition to the intrinsic predictive value of such predictions, biotechnology firms and research institutions have developed downstream applications for pathogenicity predictions. For instance, pathogenicity predictions output by machine-learning models have been used to identify target variants in a population subset for new drugs as well as target variants that may be the subject of genetic editing.

[0003] While existing pathogenicity prediction models have demonstrated significant improvements in accuracy and downstream applications, some existing models do not consistently generate accurate predictions across a range of different protein contexts. For instance, some existing pathogenicity prediction models generate pathogenicity predictions that more accurately estimate the pathogenicity of amino-acid variants at different positions within the context of a single protein relative to other models — but generate such predictions that do not as accurately (or inaccurately) estimate the pathogenicity of amino-acid variants across different proteins. By contrast, other existing pathogenicity prediction models generate pathogenicity predictions that more accurately estimate the pathogenicity of amino-acid variants across different proteins — with relatively better accuracy for variants from protein to protein — but such predictions do not as accurately (or inaccurately) estimate pathogenicity at different positions within the context of a single protein relative to other models.

[0004] In addition to accuracy depending on protein context, some existing pathogenicity prediction models do not consistently generate accurate predictions across a range of clinicalAttorney Docket No. IP-2880-PCT 1 Patent Applicationbenchmarks or cell-line protocols. Such clinical benchmarks and cell-line protocols may include, for instance, scores for protein variants or benign proteins with varying accuracy based on data from the Deciphering Developmental Disorders (DDD) study, the United Kingdom (UK) Biobank, cell-line experiments for Saturation Mutagenesis, Clinical Variant (ClinVar) from the National Library of Medicine, and Genomics England Variants (GELVar). While certain pathogenicity prediction models generate predictions that accurately indicate pathogenicity for variants in the UK Biobank, for instance, the same such models do not accurately predict pathogenicity for certain between-protein benchmarks from DDD.

[0005] The differing accuracy of such existing pathogenicity prediction models depends in part on the relatively limited and inflexible data or context upon which some such models generate predictions. To predict whether an amino-acid variant at a given position is pathogenic or benign, for instance, some existing machine-learning models can generate single pathogenicity scores for each candidate amino acid at a target position with a reference (canonical) amino-acid sequence. By substituting each of nineteen candidate alternative amino acid into a position within a reference amino-acid sequence and determining a corresponding nineteen different pathogenicity predictions, existing models can determine nineteen pathogenicity scores specific to each of the respective nineteen candidate alternative amino acids. But the data input into such existing machine-learning models can be limited to data tied to a single reference amino-acid sequence of a single protein, thereby limiting the accuracy of pathogenicity predictions inflexibly to the context of a single protein. Many existing models accordingly limit their analysis to determining the pathogenicity of amino-acid variants within the context of the protein sequence in which the variants appear (e.g., the protein sequence corresponding to the nucleotide read(s) generated from the genomic sample).

[0006] In addition to limiting analysis to a protein sequence context, such existing models often fail to consider other factors that contribute to or otherwise indicate the pathogenicity of an amino-acid variant. For example, some existing pathogenicity prediction models do not consider the structure of a protein when predicting pathogenicity for an amino-acid variant at a target position. Additionally, some existing pathogenicity prediction models predict pathogenicity based only on variant data sequenced from humans, without taking into account other relevant variant data that could impact pathogenicity.

[0007] In addition to the inaccuracies and inflexibility demonstrated by some existing pathogenicity prediction models, certain of these existing machine-learning models are also computationally inefficient. More specifically, many existing pathogenicity machine-learning models require retraining many times over to generate pathogenicity predictions for different sets of amino-acid sequences and / or for different target protein positions. The process of retrainingAttorney Docket No. IP-2880-PCT 2 Patent Applicationpathogenicity prediction models, such as DeepSequence, for new sets of amino-acid sequences can take weeks or months (e.g., 3-4 months) of constant computational expenditure to achieve improved performance in predicting protein pathogenicity. Despite months-long training, some of these existing models have no publicly released scores for all human proteins. Certain existing models, such as transformer models, that provide pathogenicity scores for only various proteins consume notoriously large amounts of computing resources for pre-training to even begin evaluating performance.

[0008] In addition to such training, some existing machine-learning models include layer upon layer of complex neural -network operations or more complex architecture designed for deeplearning neural networks that can increase the number of operations and computer processing relative to other pathogenicity prediction models. Existing models can thus consume excessive amounts of computational resources that could otherwise be preserved with a more efficient approach. Indeed, while some transformer machine-learning models exhibit improved or state-of-the-art accuracy for predicting pathogenicity, such models require minutes to generate anew pathogenicity scores for a single position or hours to generate such scores for multiple positions (e.g., an entire protein sequence). Such computational inefficiencies are especially pronounced in existing models that retrain across large numbers of sets of amino-acid sequences to generate pathogenicity predictions.

[0009] These, along with additional problems and issues, exist with regard to existing pathogenicity prediction models. Systems and methods for addressing these issues will be described below.SUMMARY

[0010] This disclosure describes one or more embodiments of systems, methods, and non-transitory computer readable storage media that solve one or more of the problems described above or provide other advantages over the art. In particular, the disclosed systems can implement a flexible approach for accurately determining the pathogenicity of amino-acid variants within a protein sequence. For instance, in some implementations, the disclosed systems separate the problem of pathogenicity prediction into two sub-problems: the prediction of pathogenicity across proteins (or protein regions) and the prediction of variant pathogenicity within the same protein. The disclosed systems further use one or more specialized computer-implemented models for each sub-problem and combine their output scores to determine a final set of scores indicating pathogenicity of a set of amino acids within a target protein sequence. For example, in some cases, the disclosed systems use one or more linear transformations that preserve the mean and standard deviation of the scores indicating inter-protein pathogenicity and further preserve the relative ranksAttorney Docket No. IP-2880-PCT 3 Patent Applicationof the scores indicating intra-protein pathogenicity. In this manner, the disclosed systems can rescale pathogenicity scores based on both inter- and intra-protein pathogenicity scores to implement a more flexible approach that determines pathogenicity with improved accuracy.

[0011] To facilitate inter-protein pathogenicity scores, in some embodiments, the disclosed systems use a relatively simple and efficient model to predict the pathogenicity of a variant at a particular amino acid position. For instance, the disclosed systems can use a specialized interprotein pathogenicity model that determines (i) variants observed in sequencing data from multiple individuals for a target position and a window of proximate positions and (ii) variants expected across the target position and proximate-position window. Because differences between the observed and expected variant data can signal a selective pressure (e.g., evolutionary constraint) due to pathogenicity, the disclosed systems can generate an inter-protein pathogenicity score specific to a target protein position based on the relationship between the observed and expected variants. As explained below, the disclosed pathogenicity model can thereby determine and leverage an observed-to-expected variant relationship to improve computational efficiency compared to neural network-based approaches — without compromising accuracy.

[0012] Additional features and advantages of one or more embodiments of the present disclosure will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by the practice of such example embodiments.BRIEF DESCRIPTION OF THE DRAWINGS

[0013] This disclosure will describe one or more embodiments of the invention with additional specificity and detail by referencing the accompanying figures. The following paragraphs briefly describe those figures, in which:

[0014] FIG. 1 illustrates a schematic diagram of a computing system in which a pathogenicity prediction system can operate in accordance with one or more embodiments.

[0015] FIG. 2 illustrates the pathogenicity prediction system generates re-scaled pathogenicity scores for a set of amino acids within a target protein sequence in accordance with one or more embodiments.

[0016] FIG. 3 illustrates, the pathogenicity prediction system generating re-scaled pathogenicity scores using multiple sets of intra-protein pathogenicity models in accordance with one or more embodiments.

[0017] FIG. 4 illustrates the pathogenicity prediction system combining intra-protein pathogenicity scores and re-scaling the combined intra-protein pathogenicity scores in accordance with one or more embodiments.Attorney Docket No. IP-2880-PCT 4 Patent Application

[0018] FIGS. 5A-5B illustrate graphs showing experimental results regarding the effectiveness of the pathogenicity prediction system generating pathogenicity predictions for amino acids of a target protein sequence in accordance with one or more embodiments.

[0019] FIGS. 6A-6B illustrate additional experimental results regarding the effectiveness of the pathogenicity prediction system generating pathogenicity predictions for amino acids of a target protein sequence in accordance with one or more embodiments.

[0020] FIGS. 7A-7B illustrate graphs showing a window-based approach implemented by the pathogenicity prediction system in accordance with one or more embodiments.

[0021] FIG. 8 illustrates the pathogenicity prediction system generating an inter-protein pathogenicity score for a target position based on an observed number of variants and an expected number of variants within a position window of the target position in accordance with one or more embodiments.

[0022] FIG. 9A illustrates the pathogenicity prediction system identifying proximate positions within a position window based on a three-dimensional distance of a target position in accordance with one or more embodiments.

[0023] FIG. 9B-9C illustrates the pathogenicity prediction system determining an observed number of variants and an expected number of variants and generating an inter-protein pathogenicity score for a target position based on the observed and expected numbers of variants in accordance with one or more embodiments.

[0024] FIG. 10A-10B illustrate graphs showing experimental results regarding the effectiveness of the pathogenicity prediction system generating inter-protein pathogenicity scores for amino acids at target positions withing a target protein sequence in accordance with one or more embodiments.

[0025] FIG. 11 illustrates an additional graph showing experimental results regarding the effectiveness of the pathogenicity prediction system generating inter-protein pathogenicity scores for amino acids in target positions of a target protein sequence in accordance with one or more embodiments.

[0026] FIG. 12A illustrates an overview of the pathogenicity prediction system generating a combined pathogenicity score utilizing an inter-and-intra score combiner model in accordance with one or more embodiments.

[0027] FIG. 12B illustrates the pathogenicity prediction system generating features related to a distribution of inter-protein pathogenicity scores in accordance with one or more embodiments.

[0028] FIG. 13 illustrates a graph showing experimental results regarding the effectiveness of the pathogenicity prediction system generating combined pathogenicity scores utilizing an inter-and-intra score combiner model in accordance with one or more embodiments.Attorney Docket No. IP-2880-PCT 5 Patent Application

[0029] FIG. 14 illustrates a flowchart of a series of acts for determining pathogenicity predictions for a set of amino acids using inter-protein pathogenicity scores and intra-protein pathogenicity scores in accordance with one or more embodiments.

[0030] FIG. 15 illustrates a flowchart of a series of acts for generating an inter-protein pathogenicity score specific to a target position of a target amino acid within a target protein sequence in accordance with one or more embodiments.

[0031] FIG. 16 illustrates a flowchart of a series of acts for generating a combined pathogenicity score utilizing an inter-and-intra score combiner model in accordance with one or more embodiments.

[0032] FIG. 17 illustrates a block diagram of an example computing device in accordance with one or more embodiments.

[0033] Appendix A further illustrates one or more embodiments of the pathogenicity prediction system.DETAILED DESCRIPTION

[0034] This disclosure describes one or more embodiments of a pathogenicity prediction system that accurately determines the pathogenicity of amino acids within a target protein sequence by flexibly combining determinations of the pathogenicity of amino acids across proteins and the pathogenicity of the amino acids within the same protein. In particular, in one or more embodiments, the pathogenicity prediction system uses an inter-protein pathogenicity model and one or more intra-pathogenicity models to generate pathogenicity scores for a set of amino acids in a target protein sequence. The pathogenicity prediction system further combines the scores from the various models to generate a final set of pathogenicity scores that preserves certain properties of the constituent scores (e.g., mean, standard deviation, or relative ranking of pathogenicity). For example, in some cases, the pathogenicity prediction system uses the scores of the inter-protein pathogenicity model to rescale the scores of the intra-protein pathogenicity model(s).

[0035] To illustrate, in one or more embodiments, the pathogenicity prediction system accesses a set of inter-protein pathogenicity scores generated by an inter-protein pathogenicity model for a set of amino acids in a target protein sequence. The pathogenicity prediction system determines an inter-protein central tendency of the set of inter-protein pathogenicity scores and an inter-protein standard deviation of the set of inter-protein pathogenicity scores. Further, the pathogenicity prediction system accesses a set of intra-protein pathogenicity scores generated by an intra-protein pathogenicity model for the set of amino acids in the target protein sequence. Using the inter-protein central tendency and the inter-protein standard deviation, the pathogenicityAttorney Docket No. IP-2880-PCT 6 Patent Applicationprediction system generates re-scaled pathogenicity scores for the set of amino acids according to a ranking of the set of intra-protein pathogenicity scores.

[0036] Indeed, as mentioned above, in one or more embodiments, the pathogenicity prediction system determines the pathogenicity of a set of amino acids within a target protein sequence. For instance, in some cases, the pathogenicity prediction system predicts the pathogenicity of aminoacid variants within the target protein sequence, such as those amino-acid variants caused by missense nucleotide variants within corresponding nucleotide reads.

[0037] To illustrate, in some embodiments, the pathogenicity prediction system uses a combination of pathogenicity models to generate re-scaled pathogenicity scores for the set of amino acids. For instance, in some embodiments, the pathogenicity prediction system uses a combination of inter- and intra-protein pathogenicity models to generate the re-scaled pathogenicity scores. In combining the models, the pathogenicity prediction system can combine certain properties of the scores generated by each model within the re-scaled pathogenicity scores. For instance, in some cases, the pathogenicity prediction system combines the standard deviation and central tendency (e.g., mean) of the scores generated by the inter-protein pathogenicity model with the relative ranking of pathogenicity indicated by the scores generated by the intra-protein pathogenicity model.

[0038] Indeed, in one or more embodiments, the pathogenicity prediction system uses interprotein pathogenicity scores generated by an inter-protein pathogenicity model for the set of amino acids. In some embodiments, the pathogenicity prediction system uses the inter-protein pathogenicity model to generate the inter-protein pathogenicity scores. In some cases, however, the pathogenicity prediction system accesses pre-generated inter-protein pathogenicity scores.

[0039] Similarly, in some instances, the pathogenicity prediction system uses intra-protein pathogenicity scores generated by an intra-protein pathogenicity model for the set of amino acids. In particular, in some embodiments, the pathogenicity prediction system uses the intra-protein model to generate the intra-protein pathogenicity scores or accesses the pre-generated intra-protein pathogenicity scores.

[0040] As further mentioned, in one or more embodiments, the pathogenicity prediction system combines the inter-protein pathogenicity scores and the intra-protein pathogenicity scores to generate the re-scaled pathogenicity scores. For instance, in certain embodiments, the pathogenicity prediction system determines a central tendency and a standard deviation for the inter-protein pathogenicity scores. The pathogenicity prediction system further generates the rescaled pathogenicity scores from the intra-protein pathogenicity scores via one or more transformations using the central tendency and standard deviation of the inter-protein pathogenicity scores. In certain implementations, the pathogenicity prediction system further determines and usesAttorney Docket No. IP-2880-PCT 7 Patent Applicationthe central tendency and standard deviation of the intra-protein pathogenicity scores in generating the re-scaled pathogenicity scores.

[0041] In some cases, the pathogenicity prediction system uses intra-protein pathogenicity scores generated from a plurality of intra-protein pathogenicity models to generate the re-scaled pathogenicity scores for the set of amino acids. For example, in some cases, the pathogenicity prediction system generates a set of re-scaled pathogenicity scores from the intra-protein pathogenicity scores of each intra-protein pathogenicity model (e.g., by combining the intra-protein pathogenicity scores with the inter-protein pathogenicity scores). The pathogenicity prediction system further combines the various sets of re-scaled pathogenicity scores to generate a combined set of re-scaled pathogenicity scores. In some implementation, the pathogenicity prediction system combines the intra-protein pathogenicity scores of the various intra-protein pathogenicity models before combining with the inter-protein pathogenicity scores.

[0042] As indicated above the pathogenicity prediction system provides several technical advantages over existing pathogenicity prediction models. For example, the pathogenicity prediction system improves the accuracy and precision with which pathogenicity prediction models generate pathogenicity predictions for amino-acid variants. Unlike existing pathogenicity models that generate relatively accurate pathogenicity predictions for amino-acid variants in the context of the amino acids within a single protein or such variants across different proteins — but not both such contexts — the disclosed pathogenicity prediction system generates accurate and re-scaled pathogenicity scores for amino acids that accounts for both inter-protein and intra-protein contexts. By using a combination of (i) intra-protein pathogenicity scores generated by a model that specializes in determining pathogenicity within the context of a protein sequence within which an amino acid appears and (ii) inter-protein pathogenicity scores generated by another model that specializes in determining pathogenicity across proteins (or across protein regions), the pathogenicity prediction system combines and produces re-scaled pathogenicity scores for amino acids that accurately incorporates these various contexts and, thereby, demonstrates improved pathogenicity-prediction performance. As shown in FIGS. 5A - 5B and 6A - 6B and described below, the disclosed pathogenicity prediction system outperforms existing pathogenicity prediction models across different clinical benchmarks and cell-line protocols.

[0043] In addition and in part responsible for such improved accuracy, in some embodiments, the pathogenicity prediction system improves the flexibility and breadth of data processed for protein context and avoids the narrow analysis of some existing pathogenicity prediction models by implementing a new approach to determining pathogenicity that involves combining and scaling scores from both inter- and intra-protein pathogenicity models in a manner that preserves properties of their respective scores. By using pathogenicity scores generated by both inter- and intra-proteinAttorney Docket No. IP-2880-PCT 8 Patent Applicationpathogenicity models, the protein pathogenicity prediction system accounts for different contextual data and determines a pathogenicity of an amino acid — such as an amino-acid variant — in multiple relevant contexts — that is, the context of pathogenicity for amino acids intra to (or within) a single target protein and the context of pathogenicity of inter to (or across) different proteins. The pathogenicity prediction system further combines scores that account for data from both protein contexts such that both inter-protein pathogenicity scores and intra-protein pathogenicity scores contribute to the uniquely re-scaled pathogenicity scores. For instance, the pathogenicity prediction system implements a new approach that generates re-scaled pathogenicity scores from intra-protein pathogenicity scores using the central tendency and standard deviation of the inter-protein pathogenicity scores via one or more transformations.

[0044] Accordingly, the disclosed pathogenicity prediction system does not merely use an off-the-shelf pathogenicity model but rather introduces a first-of-its kind model that combines scores accounting for different protein contexts. By combining inter- and intra-protein pathogenicity scores based on a mean or other central tendency of inter-protein pathogenicity scores, the pathogenicity prediction system operates with unique flexibility and when compared to existing pathogenicity prediction models. For instance, the pathogenicity prediction system uses the combination of models to determine pathogenicity of an amino acid in the context of the protein sequence in which the amino acid appears and across multiple proteins (or protein regions). By further accessing and intelligently combining interprotein pathogenicity scores within an aminoacid window, as disclosed by some embodiments, the pathogenicity prediction system implements a broader, more flexible pathogenicity analysis that intelligently accounts for different protein contexts unlike existing models.

[0045] Beyond improved accuracy and flexibility, in some embodiments, the pathogenicity prediction system improves the computing efficiency with which some pathogenicity prediction models generate pathogenicity scores for amino-acid variants. As indicated above, some existing pathogenicity prediction models have increased the accuracy of pathogenicity scores in part by adding neural-network layers or more complex architecture designed for deep-leaming neural networks, such as transformer machine-learning models. When executing a Large Language Model (LLM)-based or other such transformer machine-learning model to generate pathogenicity scores in the first instance, the additive layers or complex architecture of the transformer machinelearning model increases both the number of operations and computer processing. Rather than adding layers or more complex architecture, in some embodiments, the pathogenicity prediction system efficiently improves the accuracy of pathogenicity scores by utilizing a compute-lite approach that can linearly transform intra-protein pathogenicity scores. By accessing and combining previously generated inter-protein and intra-protein pathogenicity scores for aminoAttorney Docket No. IP-2880-PCT 9 Patent Applicationacids in a target protein sequence, for example, the pathogenicity prediction system can quickly and simply transform and rescale pathogenicity scores from different protein contexts — without more complex neural-network layers — resulting in re-scaled pathogenicity scores for amino acids according to a raking of intra-protein pathogenicity scores. Indeed, the pathogenicity prediction system can generate such re-scaled pathogenicity scores in a split second (e.g., millisecond) compared to the minutes or hours required to generate anew pathogenicity scores using a transformer machine-learning model.

[0046] Independent of or in combination with re-scaling pathogenicity scores based on inter-and intra-protein pathogenicity scores, the present disclosure further describes one or more embodiments of an inter-protein pathogenicity model that can accurately determine the pathogenicity of a target amino acid at a target position within a protein sequence by comparing the number of observed variants to a neutral expectation of expected variant in “position windows” within the protein sequence. In particular, in one or more embodiments, the pathogenicity prediction system analyzes variant data from sequencing of multiple individuals (e.g., humans, primates) to determine an observed number of amino acid variants that occur within a position window including a target position and proximate positions. The pathogenicity prediction system can further estimate or project an expected number of variants within the position window based on, for example, mutation rates and / or population data. Based on the observed number of variants and the expected number of variants, the pathogenicity prediction system can accurately and efficiently generate a pathogenicity score for a variant (or other amino acids) at the target position. By so determining inter-protein pathogenicity scores for successive target positions, the pathogenicity prediction system can distinguish between short regions of a genome that are tolerant or intolerant to variation and thus identify regions in which variants are more likely to be pathogenic.

[0047] To illustrate, in one or more embodiments, the pathogenicity prediction system identifies, for a target position of a target amino acid within a target protein sequence, proximate positions for contextual amino acids within a position window of the target position. Such a target amino-acid can be, for example, an amino-acid variant caused by missense nucleotide variants or caused by other nucleotide-variant types. The pathogenicity prediction system determines, for the target position and the proximate positions, an observed number of variants within the position window. Further, the pathogenicity prediction system determines, for the target position and the proximate positions, an expected number of variants within the position window. Using the observed number of variants and the expected number of variants, the pathogenicity prediction system generates an inter-protein pathogenicity score specific to the target position.Attorney Docket No. IP-2880-PCT 10 Patent Application

[0048] To illustrate, as mentioned, the pathogenicity prediction system determines a position window of proximate positions for the target position. Indeed, in one or more embodiments, the pathogenicity prediction system uses a position window based on three-dimensional distance to the target position. For example, in some embodiments, the pathogenicity prediction system identifies proximate positions by ranking other amino acids of the target protein sequence based on three-dimensional proximity to the target position in the three-dimensional structure of the target protein sequence. The pathogenicity prediction system can further select a pre-determined number of amino acids that are closest (in three-dimensional space or otherwise in physical proximity) to the target position, for use as the proximate positions in the position window.

[0049] In addition to determining a position window, the pathogenicity prediction system can apply various methods to determine observed and expected numbers of variants within such a window. For example, in some embodiments, the pathogenicity prediction system determines the observed number of variants based on variant data derived from sequencing studies of multiple individuals, including population-level sequencing studies. Furthermore, in some instances, the pathogenicity prediction system bases the observed and / or expected variant counts on variant data from a plurality of species (e.g., humans, primates, other species). In particular, in some embodiments, the pathogenicity prediction system determines the observed number of variants and the expected number of variants based on variant data from humans and at least one additional primate species.

[0050] As for the expected variants, in some instances, the pathogenicity prediction system determines the expected number of variants using a mutation rate model. For instance, in some embodiments, the mutation rate model can flexibly predict the number of variants expected over the position window based on a projected mutation rate. The pathogenicity prediction system can take into account a variety of factors when estimating the expected number of variants, such as the trinucleotide context at a given position in the position window, methylation context, and / or the proportion of synonymous variants in the position window (in the case of determining an expected number of missense variants).

[0051] As further mentioned, in one or more embodiments, the pathogenicity prediction system can flexibly tune parameters of the inter-protein pathogenicity model to provide greater accuracy. For example, the pathogenicity prediction system can flexibly set minimum allele frequency cut-offs for counting variants in the variant data as observed variants. Furthermore, the pathogenicity prediction system can select a position-window size and / or allele frequency threshold based on the determination of whether a gene loss of function (LoF) constraint value for the target protein sequence is above or below a threshold. Additionally, in some instances, the pathogenicityAttorney Docket No. IP-2880-PCT 11 Patent Applicationprediction system scales the observed and / or expected number of variants based on synonymous variants and / or based on sequencing depth.

[0052] Having determined observed and expected numbers of variants, the pathogenicity prediction system uses a combination of the observed number of variants and the expected number of variants to generate an inter-protein pathogenicity score specific to the target position. For instance, in some embodiments, the pathogenicity prediction system uses an observed-overexpected ratio (OOER) to predict the pathogenicity of having a variant (or other amino acid) at the target position. For instance, in some cases, the pathogenicity prediction system determines an OOER based on the observed number of variants relative to the sum of the observed number of variants and the expected number of variants.

[0053] In addition to generating an inter-protein pathogenicity score for a given position, the pathogenicity prediction system can generate such scores for previous or subsequent position windows. In one or more embodiments, the pathogenicity prediction system iteratively determines an inter-protein pathogenicity score specific to each position in a target protein sequence in a “sliding window” approach. For example, the pathogenicity prediction system can determine a subsequent position window for a subsequent target position of a subsequent target amino acid within the target protein sequence; determine an observed number of variants and an expected number of variants within the subsequent position window; and generate a corresponding interprotein pathogenicity score specific to the subsequent target position based on the observed number of variants and the expected number of variants within the subsequent position window. The pathogenicity prediction system can repeat this process for each position in the target protein sequence, providing a series of inter-protein pathogenicity scores across the target protein sequence. Based on the inter-protein pathogenicity scores across the target protein sequence, the pathogenicity prediction system can determine regions that are tolerant or intolerant to variation (e.g., missense variation).

[0054] As indicated above, the pathogenicity prediction system uses an inter-protein pathogenicity model to generate inter-protein pathogenicity scores based on observed-overexpected variant ratios or other relationship that provide several technical advantages over existing pathogenicity prediction models. For example, the disclosed inter-protein pathogenicity models save memory and expedite inter-protein-pathogenicity-score determinations relative to existing pathogenicity models that use deep learning (e.g., PrimateAI) and other neural network-based strategies to determine an inter-protein pathogenicity score. By determining inter-protein pathogenicity score based on numbers of observed and expected variants for a target position, the pathogenicity prediction system determines an inter-protein pathogenicity score with more computational efficiency than a deep learning model. Indeed, the pathogenicity prediction systemAttorney Docket No. IP-2880-PCT 12 Patent Applicationcan generate an inter-protein pathogenicity score in a split second (e.g., millisecond) compared to the minutes required by a deep-leaming model to generate pathogenicity scores for a single position or the hours required by a deep-leaming model to generate inter-protein pathogenicity scores for multiple positions (e.g., positions for an entire protein sequence). Indeed, the pathogenicity prediction system’s computational efficiency multiplies relative to existing models when determining inter-protein pathogenicity scores for multiple target positions, such as scores for each position across an entire protein sequence or across multiple protein sequences. In addition, interprotein pathogenicity models based on observed and expected numbers of variants do not require the computationally-intensive training process of a deep learning model, which can consume hours to days and utilize multiple processors, memory, and other computational resources for such a training period.

[0055] Despite being more computationally efficient, the disclosed inter-protein pathogenicity model generates inter-protein pathogenicity scores with improved accuracy relative to existing models. By determining observed numbers of variants and expected numbers of variants within a position window of a target position, the pathogenicity prediction system generates an inter-protein pathogenicity score for the target position — based on the determined observed and expected numbers — with an accuracy that exceeds existing models. For example, as further described herein, the disclosed inter-protein pathogenicity models outperform other models in predicting interprotein pathogenicity of variants for various types of disorders. These improvements can further translate to improvements to accuracy for systems that generate pathogenicity scores based on a combination of inter-protein and intra-protein pathogenicity scores. As shown in FIG. 11 and described below, for instance, the disclosed pathogenicity prediction system outperforms existing pathogenicity prediction models across different clinical benchmarks and cell-line protocols.

[0056] Additionally, specific position-window types and features of embodiments of the disclosed inter-protein pathogenicity models can provide further accuracy improvements. For example, in some cases, the inter-protein pathogenicity model can determine observed and expected numbers of variants using a position window based on three-dimensional (3D) distance to a target position, for instance, using a predetermined number of contextual amino acids that are closest in three-dimensional space to the target position. Such 3D position windows not only improve accuracy but provide new data structures, such as a multidimensional database or table of observed / expected variant counts of 3D-proximate positions, that improve pathogenicityprediction technology and computational efficiency. Furthermore, as further described herein, the inclusion of variant data from non-human primates, in addition to human variant data, can also improve the accuracy of inter-protein pathogenicity score determinations based on observed and expected variant counts. The disclosure thus contemplates additional new data structures, such asAttorney Docket No. IP-2880-PCT 13 Patent Applicationa database or table of observed / expected variant counts based on variant data from a plurality of species (e.g., human and non-human primate variant data). In the context of an inter-protein pathogenicity model based on observed and expected variant counts, the addition of non-human primate variant data provides only negligible additional computational burden. As shown in FIGS.10A-10B and described below, the 3D window and non-human primate variant data features each individually contribute to improvements in the accuracy of the disclosed protein pathogenicity models.

[0057] As illustrated by the foregoing discussion, the present disclosure utilizes a variety of terms to describe features and advantages of the pathogenicity prediction system. For example, as used herein, the term “pathogenicity” refers to the ability or tendency of a biological molecule to contribute to disease within a host organism. In particular, pathogenicity can refer to the ability or tendency of a biological molecule to cause or lead to the susceptibility of disease within the host organism. For example, pathogenicity can refer to the ability or tendency of an amino acid to cause a protein including the amino acid to function in a manner that causes disease within a host organism or leads to susceptibility of the host organism to disease. More specifically, in some cases, pathogenicity refers to the ability or tendency of an amino-acid variant within a protein to change the function and / or structure of the protein such that the protein causes or leads to susceptibility to disease.

[0058] Additionally, as used herein, the term “amino acid” refers to an organic molecule that serves as a residue or building block for proteins. Accordingly, an amino acid can be one unit or residue within an amino-acid sequence corresponding to a protein. In particular, an amino acid can include an organic molecule composed of a central carbon atom attached to an amino group, a hydrogen atom, and a variable side chain that determines its properties. In some cases, an amino acid includes a biological molecule selected during translation for inclusion within a protein based on a codon of messenger ribonucleic acid (mRNA).

[0059] Further, as used herein, the term “protein sequence” (or “amino-acid sequence”) refers to a group of amino acids that constitute part or all of a protein. In particular, a protein sequence can include a group of amino acids including all amino acids in a protein or at least a subset (e.g., a sequence) of the amino acids of the protein. Relatedly, as used herein, the term “target protein sequence” (or “target amino-acid sequence”) refers to a protein sequence targeted for analysis. For instance, a target protein sequence can include (i) a sequence of amino acids for which a pathogenicity prediction is determined or (ii) a sequence of amino acids comprising target amino acid for which a pathogenicity prediction is determined. Additionally, as used herein, the term “protein position” refers to a particular location, coordinate, or order for an amino acid within aAttorney Docket No. IP-2880-PCT 14 Patent Applicationprotein sequence. For example, a protein position includes a location, coordinate, or order of an amino acid within a target protein sequence or within a reference protein sequence.

[0060] Relatedly, as used herein, the term “target position” refers to a protein position targeted for analysis and pathogenicity prediction. Accordingly, a target position is a specific type of protein position. For example, a target position can include a position for an amino acid within a target protein sequence and for which a pathogenicity prediction is determined. In some embodiments, the systems disclosed herein determine a pathogenicity score for a first target position in the sequence of amino acids in a target protein sequence, before moving to a second target position (e.g., the amino acid position subsequent to the target position) and determining a pathogenicity score for the second target position, and so forth for third, fourth, and additional target positions.

[0061] By contrast, as used herein, the term “target amino acid” refers to an amino acid at a target position. For example, a target amino acid can include a reference amino acid or a variant amino acid at the target position. Such a target amino acid can be at a position within a target protein sequence from a human, other primate, or other species.

[0062] Additionally, as used herein, the term “proximate position” refers to a protein position that is nearby a target position according to a predetermined measure of proximity. Accordingly, a proximate position is also a specific type of protein position but different from a target position. For example, a proximate position can include a position for an amino acid within a target protein sequence that is proximate to an amino acid at a target position. For example, in some embodiments, a proximate position is in proximity to a target position based on three-dimensional distance to the target position, referred to herein as “three-dimensional,” “3D,” or “3D space” proximity. In some embodiments, a proximate position is in proximity to a target position based on the number of protein positions between the target position and the proximate position in a contiguous protein sequence, also referred to herein as “one-dimensional,” “ID” or “sequence space” proximity.

[0063] By contrast, as used herein, the term “contextual amino acid” refers to an amino acid at a proximate position. For example, a contextual amino acid can include a reference amino acid or a variant amino acid at the proximate position. Such a contextual amino acid can be at a position within a target protein sequence from a human, other primate, or other species.

[0064] Relatedly, as used herein, the term “position window” refers to a series or set of protein positions within a predetermined criteria of proximity to a target position. For example, a position window may include a predetermined number of protein positions that are most proximate (whether in sequence space or 3D space) to a target amino acid relative to other protein positions outside the most-proximate protein positions. In other instances, the position window includes a flexible number of protein positions that are within a predetermined three-dimensional distance from aAttorney Docket No. IP-2880-PCT 15 Patent Applicationtarget position. As indicated above, a position window can slide or move. In some instances, for example, the position window corresponds to the target position such that, as the disclosed systems select a new target position, the disclosed systems update the position window to include a different set of protein positions that are proximate to the newly selected target position.

[0065] As used herein, the term “amino-acid variant” refers to an amino acid at a position within a target protein sequence that differs from the amino acid at the same position within the corresponding reference protein sequence. In other words, in some cases, an amino-acid variant refers to an amino acid that is of a different amino-acid type than the amino acid at the same position within the corresponding reference protein sequence.

[0066] Additionally, as used herein, the term “alternative amino acid” more broadly refers to an amino acid that differs from another amino acid. For instance, in some cases, an alternative amino acid refers to an amino-acid variant. In some instances, however, an alternative amino acid more generally refers to an amino acid that is of a different amino acid type than a given amino acid. For instance, there exists a set of known amino acid types (e.g., twenty common amino acids with more having been identified in nature). Thus, for a given amino acid of an amino acid type, an alternative amino acid can include an amino acid of one of the other known amino acid types.

[0067] Further, as used herein, the term “reference amino acid” refers to an amino acid used as a basis for comparison with another amino acid. For instance, in some cases, a reference amino acid includes an amino acid at a position within a reference protein sequence used for comparison with an amino acid at the same position within a corresponding target protein sequence. In some embodiments, a reference amino acid includes an amino acid at a position within a target protein sequence used to determine alternative amino acids for that position.

[0068] Additionally, as used herein, the term “observed number of variants” refers to a count or estimation of variants at one or more positions in sequencing data. For example, an observed number of variants includes DNA sequence variants resulting in missense or synonymous changes to the protein sequence within a position window. As indicated above, the disclosed systems can analyze variant data at a target position and / or proximate positions in a position window and determine a number of variants observed in the variant data. The variant data can include sequencing data about variant nucleotides and / or amino acids for multiple individuals. In some instances, the disclosed systems apply various criteria when determining the observed number of variants, for example by only considering missense variants, or variants that have an allele frequency above a predetermined threshold.

[0069] By contrast, as used herein, the term “expected number of variants” refers to a proj ected number of variants at one or more positions. For example, an expected number of variants includes DNA sequence variants resulting in missense or synonymous changes projected to occur within aAttorney Docket No. IP-2880-PCT 16 Patent Applicationposition window of the protein sequence. As indicated above, the disclosed systems can use a mutation rate model to predict how many variants would be observed at a target position and / or proximate positions in a position window under neutral selective pressure (e.g., if natural selection due to pathogenicity was not applicable). In some embodiments, the disclosed systems determine an expected number of missense variants for a position window using the mutation rates for all possible missense substitutions in the position window.

[0070] As used herein, the term “pathogenicity score” refers to a score that is indicative of pathogenicity. In particular, a pathogenicity score can refer to a score that indicates that pathogenicity of an amino acid (or corresponding nucleotide(s)) within a target protein sequence. For instance, in some cases, a pathogenicity score includes a score indicating the pathogenicity of an amino-acid variant included in the target protein sequence. In some instances, a pathogenicity score includes a numerical value where a relatively higher value indicates relatively higher pathogenicity and a relatively lower value indicates relatively lower pathogenicity or vice versa. In some embodiments, a pathogenicity score provides a direct measure of pathogenicity (e.g., the score indicates the level of pathogenicity). In some instances, however, a pathogenicity score provides an indirect measure of pathogenicity, such as by measuring some other characteristic related to pathogenicity (e.g., the depletion of observed variants).

[0071] Relatedly, as used herein, the term “inter-protein pathogenicity score” refers to a pathogenicity score generated by an inter-protein pathogenicity model. Accordingly, an interprotein pathogenicity score includes a pathogenicity score that quantifies pathogenicity of an amino acid with respect to conservation or depletion of amino acids across proteins, protein regions (e.g., nucleotide sequence encoding proteins), or other genomic regions (e.g., promoter regions, intergenic regions, regulatory regions, non-coding RNA regions, centromeres, telomeres). Accordingly, in some cases, an inter-protein pathogenicity score can indicate pathogenicity of an amino acid with respect one or more of protein-encoding genomic regions or non-encoding genomic regions. For example, an inter-protein pathogenicity score can include a metric, value, or other quantification that indicates a depletion of observed variants across a genomic region, such as across a genomic region comprising a nucleotide sequence encoding one or more proteins. Additionally, while much of the disclosure herein describes generating inter-protein pathogenicity scores using a machine learning model, such as a neural network, the inter-protein pathogenicity scores include observed measurements in some instances. As an example, in some embodiments, an inter-protein pathogenicity score includes a measurement of depletion of variation defined as the number of observed variants divided by the length of the genomic region. By contrast, as used herein, the term “intra-protein pathogenicity score” refers to a pathogenicity score generated by an intra-protein pathogenicity model. Accordingly, an intra-protein pathogenicity score includes aAttorney Docket No. IP-2880-PCT 17 Patent Applicationpathogenicity score that quantifies pathogenicity of an amino acid with respect to other amino acids within a target protein sequence. For example, an intra-protein pathogenicity score includes a metric, value, or other quantification that indicates a pathogenicity ranking of amino acids in a target protein sequence.

[0072] In one or more embodiments, the term “pathogenicity model” refers to a computer-implemented model or algorithm that generates pathogenicity scores. In particular, a pathogenicity model can refer to a computer-implemented model or algorithm that generates pathogenicity scores for amino acids within a target protein sequence. For instance, in some embodiments, a pathogenicity model includes a computer-implemented model that analyzes a target protein sequence and generates a pathogenicity score for one or more amino acids within the target protein sequence based on the analysis. In some cases, a pathogenicity model further generates pathogenicity scores for alternative amino acids. The pathogenicity prediction system uses various pathogenicity models in various implementations, which will be discussed in more detail below. In some embodiments, for example, the pathogenicity prediction system uses a machine learning model, such as a neural network, as a pathogenicity model.

[0073] As used herein, the term “machine learning model” refers to a computer-implemented model that can be tuned (e.g., trained) based on inputs to approximate unknown functions. In particular, in some embodiments, a machine-learning model includes a model that utilizes algorithms to leam from, and make predictions on, known data by analyzing the known data to leam to generate outputs that reflect patterns and attributes of the known data. For instance, in some cases, a machine-learning model includes, but is not limited to, a neural network (e.g., a convolutional neural network, recurrent neural network or other deep learning network), a decision tree (e.g., a gradient boosted decision tree, such as XGBoost), association rule learning, inductive logic programming, support vector learning, a Bayesian network, a regression-based model (e.g., censored regression), principal component analysis, or a combination thereof.

[0074] Further, as used herein, the term “neural network” refers to a model of interconnected artificial neurons (e.g., organized in layers) that communicate and leam to approximate complex functions and generate outputs based on inputs provided to the model. In some instances, a neural network includes one or more machine learning algorithms. Further, in some cases, a neural network includes an algorithm (or set of algorithms) that implements deep learning techniques that utilize a set of algorithms to model high-level abstractions in data. To illustrate, in some embodiments, a neural network includes a convolutional neural network, a recurrent neural network (e.g., a long short-term memory neural network), a generative adversarial network, a graph neural network, a multi-layer perceptron, or a diffusion neural network. In some embodiments, a neural network includes a combination of neural networks or neural network components.Attorney Docket No. IP-2880-PCT 18 Patent Application

[0075] Relatedly, as used herein, the term “inter-protein pathogenicity model” refers to a pathogenicity model that generates pathogenicity scores (i.e., inter-protein pathogenicity scores) based on an analysis of pathogenicity across proteins or protein regions. In particular, in some embodiments, an inter-protein pathogenicity model includes a pathogenicity model that generates pathogenicity scores for a set of amino acids within a target protein sequence based on an analysis of inter-protein pathogenicity — that is, pathogenicity across proteins or protein regions. For instance, in some embodiments, an inter-protein pathogenicity model refers to a pathogenicity model that specializes (e.g., is trained or otherwise configured) in predicting the pathogenicity of an amino acid within a target protein sequence based on a pathogenicity of the amino acid across proteins or protein regions. Additionally, as used herein, the term “intra-protein pathogenicity model” refers to a pathogenicity model that generates pathogenicity scores (i.e., intra-protein pathogenicity scores) based on an analysis of pathogenicity within a protein. In particular, in some embodiments, an intra-protein pathogenicity model includes a pathogenicity model that generates pathogenicity scores for a set of amino acids within a target protein sequence based on an analysis of intra-protein pathogenicity — that is, pathogenicity within the protein of the target protein sequence. For instance, in some embodiments, an intra-protein pathogenicity model refers to a pathogenicity model that specializes (e.g., is trained or otherwise configured) in predicting the pathogenicity of an amino acid within a target protein sequence based on a pathogenicity of the amino acid within the protein of the target protein sequence.

[0076] As used herein, the term “central tendency” refers to a statistical measure that indicates or describes a center or typical value within a dataset. For instance, a central tendency can refer to a statistical measure derived from a set of pathogenicity scores that indicates or describes a center or typical value within the set. To illustrate, in some cases, a central tendency includes a mean, a median, or a mode of a set of pathogenicity scores. As used here, the term “inter-protein central tendency” more specifically refers to a central tendency for a set of inter-protein pathogenicity scores. Similarly, as used herein, the term “intra-protein central tendency” more specifically refers to a central tendency for a set of intra-protein pathogenicity scores.

[0077] Additionally, as used herein, the term “standard deviation” refers to a statistical measure that indicates or describes the amount of variation or dispersion in a dataset. For instance, a standard deviation can refer to a statistic measure derived from a set of pathogenicity scores that indicates or describes the amount of variation or dispersion within the set. As used herein, the term “inter-protein standard deviation” more specifically refers to a standard deviation for a set of interprotein pathogenicity scores. Similarly, as used herein, the term “intra-protein standard deviation” more specifically refers to a standard deviation for a set of intra-protein pathogenicity scores.Attorney Docket No. IP-2880-PCT 19 Patent Application

[0078] As used herein, the term “re-scaled pathogenicity score” refers to a pathogenicity score that has been transformed via one or more transformations. In particular, a re-scaled pathogenicity score can refer to a pathogenicity score that has been generated from a combination of other pathogenicity scores output by one or more pathogenicity models. For instance, in some cases, a re-scaled pathogenicity score includes a pathogenicity score determined by combining at least one inter-protein pathogenicity score generated by an inter-protein pathogenicity model and at least one intra-protein pathogenicity score generated by an intra-protein pathogenicity model. In some such cases, a re-scaled pathogenicity score represents a transformed intra-protein pathogenicity score from a set of standardized intra-protein pathogenicity scores using the inter-protein central tendency and the inter-protein standard deviation. More detail regarding re-scaled pathogenicity scores and their determination will be provided below.

[0079] As used herein, a “genomic region” refers to a range of locations or positions within a genome. In certain implementations, a genomic region may be identified by an identifier for a chromosome and a particular position or positions, such as numbered positions following the identifier for a chromosome (e.g., chrl : 1234570-1234870). In various implementations, a genomic coordinate includes a range of locations or positions within a reference genome. In some cases, a genomic region is specific to a particular reference genome.

[0080] Additionally, as used herein, the term “pathogenicity ranking” includes a ranking of amino acids based on pathogenicity. In particular, a pathogenicity ranking can refer to a relative ranking of amino acids based on pathogenicity. For instance, a pathogenicity ranking of an amino acid can refer to a relative ranking of the amino acid compared to one or more other amino acids based on the pathogenicity of the amino acid and the one or more other amino acids.

[0081] Further, as used herein, the term “missense nucleotide variant” refers to a nucleotide variant within a DNA sequence that leads to an amino-acid variant. In particular, a missense nucleotide variant can refer to a nucleotide change in a DNA sequence (e.g., where a nucleotide at a position within the DNA sequence differs from the nucleotide at the same position within the corresponding reference DNA sequence) that causes a change to the amino acid incorporated within a protein during translation.

[0082] The following paragraphs describe the pathogenicity prediction system with respect to illustrative figures that portray example embodiments and implementations. FIG. 1 illustrates a schematic diagram of a computing system 100 in which a pathogenicity prediction system 106 operates in accordance with one or more embodiments. As illustrated, the computing system 100 includes one or more server device(s) 102 connected to a client device 110, therapeutics analysis device(s) 114, and database 120 via a network 108. While FIG. 1 shows an embodiment of theAttorney Docket No. IP-2880-PCT 20 Patent Applicationpathogenicity prediction system 106, this disclosure describes alternative embodiments and configurations below.

[0083] As shown in FIG. 1, the server device(s) 102, the database 120, the client device 110, and the therapeutics analysis device(s) 114 are connected via the network 108. Accordingly, each of the components of the computing system 100 can communicate via the network 108. The network 108 comprises any suitable network over which computing devices can communicate. Example networks are discussed in additional detail below with respect to FIG. 17.

[0084] As indicated by FIG. 1, the therapeutics analysis device(s) 114 comprises a device for analyzing (and identifying candidate therapeutics for) amino acid sequences corresponding to proteins and / or nucleotide sequences representing coding and non-coding genomic regions. In some embodiments, the therapeutics analysis device(s) 114 analyzes a set of amino acid sequences or a set of nucleotide sequences from a database comprising samples exhibiting genetic diversity. From among the analyzed set of amino acid sequences and / or analyzed set of nucleotide sequences, the therapeutics analysis device(s) 114 can identify subsets of amino acid sequences and / or nucleotide sequences exhibiting common variant amino acids or variant nucleotides. In combination with or separate from such variant identification, the therapeutics analysis device(s) 114 can execute machine-learning models (or other models) that identify coding or non-coding genomic regions that are intolerant to variation and for which variants can cause loss or change in biological functions. For identified subsets of amino-acid sequences and / or nucleotide sequences, in some cases, the therapeutics analysis device(s) 114 identifies candidate biologies, drugs, or geneediting protocols for treatment.

[0085] In addition, or in the alternative to communicating across the network 108, in some embodiments, the therapeutics analysis device(s) 114 bypasses the network 108 and communicates directly with the server device(s) 102 or the client device 110. Additionally, as shown in FIG. 1, in one or more embodiments, the therapeutics analysis device(s) 114 includes the pathogenicity prediction system 106.

[0086] As further indicated by FIG. 1, the server device(s) 102 may generate, receive, analyze, store, and transmit digital data, such as data for amino acid sequences or nucleotide sequences. As shown in FIG. 1, the therapeutics analysis device(s) 114 may send (and the server device(s) 102 may receive) various data from the therapeutics analysis device(s) 114, including data representing amino acid sequences or nucleotide sequences. The server device(s) 102 may also communicate with the client device 110. In particular, the server device(s) 102 can send data representing amino acid sequences or nucleotide sequences (or variants thereof), inter-protein pathogenicity scores (and / or the corresponding central tendency and standard deviation), intra-protein pathogenicityAttorney Docket No. IP-2880-PCT 21 Patent Applicationscores (and / or the corresponding central tendency and standard deviation), or re-scaled pathogenicity scores to the client device 110.

[0087] Additionally, as shown in FIG. 1 , the server device(s) 102 can include the pathogenicity prediction system 106. In one or more embodiments, as explained further below, the pathogenicity prediction system 106 analyzes a set of amino acids of a target protein sequence and generates rescaled pathogenicity scores indicating the pathogenicity of the amino acids. In particular, in some embodiments, the pathogenicity prediction system 106 accesses a set of inter-protein pathogenicity scores generated for the amino acids by an inter-protein pathogenicity model 104 and further accesses one or more sets of intra-protein pathogenicity scores generated for the amino acids by one or more intra-protein pathogenicity model(s) 105. The pathogenicity prediction system 106 further combines the inter- and intra-protein pathogenicity scores to generate the re-scaled pathogenicity scores. For instance, in some implementations, the pathogenicity prediction system 106 determines an inter-protein central tendency and an inter-protein standard deviation from the set of inter-protein pathogenicity scores. The pathogenicity prediction system 106 further transforms the set of intra-protein pathogenicity scores via one or more transformation using the inter-protein central tendency and the inter-protein standard deviation.

[0088] In some embodiments, the server device(s) 102 comprise a distributed collection of servers where the server device(s) 102 include a number of server devices distributed across the network 108 and located in the same or different physical locations. Further, the server device(s) 102 can comprise a content server, an application server, a communication server, a web-hosting server, or another type of server.

[0089] In some cases, the server device(s) 102 is located at or near a same physical location of the therapeutics analysis device(s) 114 or remotely from the therapeutics analysis device(s) 114. Indeed, in some embodiments, the server device(s) 102 and the therapeutics analysis device(s) 114 are integrated into a same computing device. The server device(s) 102 may run software on the therapeutics analysis device(s) 114 or the pathogenicity prediction system 106 to generate, receive, analyze, store, and transmit digital data, such as by sending or receiving data representing amino acid sequences or nucleotide sequences (or variants thereof), inter-protein pathogenicity scores (and / or the corresponding central tendency and standard deviation), intra-protein pathogenicity scores (and / or the corresponding central tendency and standard deviation), or re-scaled pathogenicity scores. Additionally, or alternatively, in some embodiments, the therapeutics analysis device(s) 114 or the pathogenicity prediction system 106 store and access a database or table of inter-protein pathogenicity scores (and / or the corresponding central tendency and standard deviation), intra-protein pathogenicity scores (and / or the corresponding central tendency andAttorney Docket No. IP-2880-PCT 22 Patent Applicationstandard deviation), or re-scaled pathogenicity scores corresponding to sets of amino acids from protein sequences.

[0090] As further illustrated and indicated in FIG. 1, the computing system 100 includes the database 120. The database 120 can store information, such as pathogenicity scores 128 and / or other data described herein. In some embodiments, the server device(s) 102, the client device 110, and / or the therapeutics analysis device(s) 114 communicate with the database 120 (e.g., via the network 108) to store and / or access information. In some cases, the database 120 also stores one or more models, such as the inter-protein pathogenicity model 104 and / or the intra-protein pathogenicity model(s) 105.

[0091] As further illustrated and indicated in FIG. 1, the client device 110 can generate, store, receive, send, and display digital data. In particular, the client device 110 can receive data for amino acid sequences or nucleotide sequences (or variants thereof), inter-protein pathogenicity scores (and / or the corresponding central tendency and standard deviation), intra-protein pathogenicity scores (and / or the corresponding central tendency and standard deviation), or re-scaled pathogenicity scores from the server device(s) 102 and / or the therapeutics analysis device(s) 114. The client device 110 can accordingly present data concerning inter-protein pathogenicity scores (and / or the corresponding central tendency and standard deviation), intra-protein pathogenicity scores (and / or the corresponding central tendency and standard deviation), or re-scaled pathogenicity scores (collectively, “pathogenicity score data”) within a graphical user interface to a user associated with the client device 110.

[0092] For example, the pathogenicity prediction system 106 can provide pathogenicity score data for display on a graphical user interface 124. In some cases, the pathogenicity prediction system 106 provides pathogenicity score data for display as textual or graphical elements. In some embodiments, as illustrated in FIG. 1, the pathogenicity prediction system 106 provides the pathogenicity score data to the client device 110 in the form of a lookup table or searchable database, such that the pathogenicity prediction system 106 can receive a search input including but not limited to a gene, region, variant, and so forth from the client device 110, upon which the pathogenicity prediction system 106 can provide the pathogenicity score data and other related information (for example, statistical metrics, gene data, variant data) via the graphical user interface 124 of the client device 110.

[0093] The client device 110 illustrated in FIG. 1 may comprise various types of client devices. For example, in some embodiments, the client device 110 includes non-mobile devices, such as desktop computers or servers, or other types of client devices. In yet other embodiments, the client device 110 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones. Additional details with regard to the client device 110 are discussed below with respect to FIG. 17.Attorney Docket No. IP-2880-PCT 23 Patent Application

[0094] As further illustrated in FIG. 1, the client device 110 includes an analytics application 112. The analytics application 112 may be a web application or a native application stored and executed on the client device 110 (e.g., a mobile application, desktop application). The analytics application 112 can include instructions that (when executed) cause the client device 110 to receive data from the pathogenicity prediction system 106 and present data from the therapeutics analysis device(s) 114 and / or the server device(s) 102. Furthermore, the analytics application 112 can instruct the client device 110 to display (for example, using the graphical user interface 124) data for amino acid sequences or nucleotide sequences (or variants thereof), pathogenicity scores, interprotein pathogenicity scores (and / or the corresponding central tendency and standard deviation), intra-protein pathogenicity scores (and / or the corresponding central tendency and standard deviation), or re-scaled pathogenicity scores.

[0095] As further illustrated in FIG. 1, the pathogenicity prediction system 106 may be located on the client device 110 as part of the analytics application 112 or on the therapeutics analysis device(s) 114. Accordingly, in some embodiments, the pathogenicity prediction system 106 is implemented by (e.g., located entirely or in part) on the client device 110. As mentioned, in yet other embodiments, the pathogenicity prediction system 106 is implemented by one or more other components of the computing system 100, such as the therapeutics analysis device(s) 114. In particular, the pathogenicity prediction system 106 can be implemented in a variety of different ways across the server device(s) 102, the network 108, the client device 110, and the therapeutics analysis device(s) 114.

[0096] Though FIG. 1 illustrates the components of the computing system 100 communicating via the network 108, in certain implementations, the components of computing system 100 can also communicate directly with each other, bypassing the network 108. For instance, and as previously mentioned, in some implementations, the client device 110 communicates directly with the therapeutics analysis device(s) 114. Additionally, in some embodiments, the client device 110 communicates directly with the pathogenicity prediction system 106. Moreover, the pathogenicity prediction system 106 can access one or more databases housed on or accessed by the server device(s) 102 or elsewhere in the computing system 100, for example the database 120.

[0097] As discussed above, the pathogenicity prediction system 106 can analyze a target protein sequence and generate re-scaled pathogenicity scores based on the analysis. In particular, the pathogenicity prediction system 106 can generate the re-scaled pathogenicity scores for a set of amino acids included in the target protein sequence. FIG. 2 illustrates the pathogenicity prediction system 106 generates re-scaled pathogenicity scores for a set of amino acids within a target protein sequence in accordance with one or more embodiments.Attorney Docket No. IP-2880-PCT 24 Patent Application

[0098] Indeed, as shown, the pathogenicity prediction system 106 analyzes a target protein sequence 202. The target protein sequence 202 includes a set of amino acids (labeled Al, A2, A3, A4, . . .). In some cases, the set of amino acids include a sequence of amino acids within the target protein sequence. For instance, in some cases, the set of amino acids includes a plurality of amino acids in a sequential order. In some instances, the set of amino acids represents a two-dimensional ordering of the amino acids determined from a three-dimensional structure associated with the target protein sequence 202.

[0099] In some cases, the set of amino acids of the target protein sequence 202 includes one or more amino-acid variants. For instance, the set of amino acids can include one or more aminoacid variants caused by one or more corresponding missense nucleotide variants. Thus, in some cases, the pathogenicity prediction system 106 can determine the pathogenicity of amino-acid variants present within target protein sequences.

[0100] As shown in FIG. 2, the pathogenicity prediction system 106 analyzes the target protein sequence 202 by accessing inter-protein pathogenicity scores 206 generated by an inter-protein pathogenicity model 204 for the set of amino acids of the target protein sequence 202.

[0101] For example, in one or more embodiments, the pathogenicity prediction system 106 accesses the inter-protein pathogenicity scores 206 by accessing pre-generated pathogenicity scores output by the inter-protein pathogenicity model 204. For instance, in certain cases, the pathogenicity prediction system 106 retrieves the inter-protein pathogenicity scores 206 from a (e.g., a remote or local) storage location or otherwise receives the inter-protein pathogenicity scores 206 from the inter-protein pathogenicity model 204. To illustrate, in some cases, a third-party system uses the inter-protein pathogenicity model 204 to generate the inter-protein pathogenicity scores 206, and the pathogenicity prediction system 106 receives the inter-protein pathogenicity scores 206 from the third-party system or retrieves the inter-protein pathogenicity scores 206 from a storage location at which the third-party system stored the inter-protein pathogenicity scores 206.

[0102] In some implementations, the pathogenicity prediction system 106 accesses the interprotein pathogenicity scores 206 by generating the inter-protein pathogenicity scores 206. In particular, the pathogenicity prediction system 106 can use the inter-protein pathogenicity model 204 to generate the inter-protein pathogenicity scores 206. For instance, in some embodiments, the pathogenicity prediction system 106 provides the target protein sequence 202 to the inter-protein pathogenicity model 204 and uses the inter-protein pathogenicity model 204 to generate the interprotein pathogenicity scores 206 based on its analysis of the target protein sequence 202 (e.g., the analysis of the set of amino acids included therein).Attorney Docket No. IP-2880-PCT 25 Patent Application

[0103] Similarly, as shown in FIG. 2, the pathogenicity prediction system 106 analyzes the target protein sequence 202 by accessing intra-protein pathogenicity scores 210 generated by an intra-protein pathogenicity model 208 for the set of amino acids of the target protein sequence 202.

[0104] For example, in one or more embodiments, the pathogenicity prediction system 106 accesses the intra-protein pathogenicity scores 210 by accessing pre-generated pathogenicity scores output by the intra-protein pathogenicity model 208. For instance, in certain cases, the pathogenicity prediction system 106 retrieves the intra-protein pathogenicity scores 210 from a (e.g., a remote or local) storage location or otherwise receives the intra-protein pathogenicity scores 210 from the intra-protein pathogenicity model 208. To illustrate, in some cases, a third-party system uses the intra-protein pathogenicity model 208 to generate the intra-protein pathogenicity scores 210, and the pathogenicity prediction system 106 receives the intra-protein pathogenicity scores 210 from the third-party system or retrieves the intra-protein pathogenicity scores 210 from a storage location at which the third-party system stored the intra-protein pathogenicity scores 210.

[0105] In some implementations, the pathogenicity prediction system 106 accesses the intra-protein pathogenicity scores 210 by generating the intra-protein pathogenicity scores 210. In particular, the pathogenicity prediction system 106 can use the intra-protein pathogenicity model 208 to generate the intra-protein pathogenicity scores 210. For instance, in some embodiments, the pathogenicity prediction system 106 provides the target protein sequence 202 to the intra-protein pathogenicity model 208 and uses the intra-protein pathogenicity model 208 to generate the intra-protein pathogenicity scores 210 based on its analysis of the target protein sequence 202 (e.g., the analysis of the set of amino acids included therein).

[0106] In one or more embodiments, the inter-protein pathogenicity scores 206 and / or the intra-protein pathogenicity scores 210 contain a pathogenicity score for each amino acid from the set of amino acids in the target protein sequence 202. Further, as suggested above, the inter-protein pathogenicity scores 206 and / or the intra-protein pathogenicity scores 210 can include a pathogenicity score for each alternative amino acid — that is, for each amino acid of a different amino acid type than the amino acid at a particular protein position of the target protein sequence 202. Thus, in some cases, the inter-protein pathogenicity scores 206 and / or the intra-protein pathogenicity scores 210 include a multi-dimensional data set of pathogenicity scores. As an example, when including a pathogenicity score for twenty amino acids — that is, a pathogenicity score for an amino acid at a given protein position within the target protein sequence 202 and additional pathogenicity scores for nineteen alternative amino acids of different amino acid types — the inter-protein pathogenicity scores 206 and the intra-protein pathogenicity scores 210 can each include L x 20 pathogenicity scores where L represents the length of the target protein sequence 202 (e.g., the number of amino acids in the target protein sequence 202).Attorney Docket No. IP-2880-PCT 26 Patent Application

[0107] As suggested above, the pathogenicity prediction system 106 can use various pathogenicity models or otherwise access pathogenicity scores generated by various pathogenicity models. Thus, the inter-protein pathogenicity model 204 and the intra-protein pathogenicity model 208 can differ in various implementations. For instance, in some embodiments, the inter-protein pathogenicity model 204 includes the PrimateAI model developed by Illumina, Inc. In other embodiments, the inter-protein pathogenicity model 204 includes a genome aggregation database (gnomAD) missense scores model or an observed-to-expected ratio model. Additionally, in some embodiments, the intra-protein pathogenicity model 208 includes a language protein model, such as the PrimateAI language model developed by Illumina, Inc or the evolutionary scale model (ESM-2). In other embodiments, the intra-protein pathogenicity model 208 includes the PrimateAI-3D model developed by Illumina, Inc. a, DeepSequence model, or a multiple sequence alignment (MSA) transformer.

[0108] As the models generating the pathogenicity scores can vary in different embodiments, so can the inputs used by the models to generate their scores. Indeed, while the above mentions providing the target protein sequence 202 as input to the inter-protein pathogenicity model 204 or the intra-protein pathogenicity model 208 in instances where the pathogenicity prediction system 106 uses the models to generate their respective scores, it should be noted that the inputs to the models can differ. For instance, the inputs used by the pathogenicity prediction system 106 can depend on the inputs for which the particular model is configured to process.

[0109] Further, as the inter-protein pathogenicity model 204 and / or the intra-protein pathogenicity model 208 can include various pathogenicity models in various implementations, the inter-protein pathogenicity scores 206 and / or the intra-protein pathogenicity scores 210 can include pathogenicity scores providing various indications of pathogenicity. For instance, in one or more embodiments, the inter-protein pathogenicity scores 206 include pathogenicity scores that indicate a depletion of observed variants across a genomic region. In some embodiments, the intra-protein pathogenicity scores 210 include pathogenicity scores having values that indicate a pathogenicity ranking of the set of amino acids. Indeed, as previously mentioned, the inter-protein pathogenicity scores 206 and / or the intra-protein pathogenicity model 208 can include pathogenicity scores providing a direct measure of pathogenicity (e.g., having values indicating a level of pathogenicity) or providing an indirect measure of pathogenicity (e.g., providing a measure of some other characteristic of the amino acids that is related to pathogenicity).

[0110] In some implementations, the pathogenicity prediction system 106 trains the interprotein pathogenicity model 204 to generate inter-protein pathogenicity scores and / or trains the intra-protein pathogenicity model 208 to generate intra-protein pathogenicity scores. To illustrate, in some cases, the pathogenicity prediction system 106 generalizes the inter-protein pathogenicityAttorney Docket No. IP-2880-PCT 27 Patent Applicationmodel 204 as a differentiable function f that takes arbitrary inputs and either produces one scalar per protein position and variant to be scored (e.g., re-ranking) or one scalar per protein position (rescaling). More specifically, the pathogenicity prediction system 106 can denote the input to the inter-protein pathogenicity model 204 by a matrix X where Xi is a feature vector corresponding to the position i in the target protein sequence. In some cases, / is a linear transformation. In some instances, however, / includes a more complex model, such as a multi-layer perceptron that operates on the features Xi, a convolutional neural network that operates on a window of size h+1, for example, on the input vectors X_{i-h}, X_{i-h+l}, ..., X_{i-1}, X_i, X_{i+1}, ..., X_{i+h-l}, X_{i+h}, or a transformer-based model. Here, h refers to the number of positions left or right of a central window position, which is counted excluding the central window position. In some cases, h is called the “half-window size.” The pathogenicity prediction system 106 can train the interprotein pathogenicity model 204 by backpropagating the error of predicted inter-protein pathogenicity scores generated from training data. In some cases, the pathogenicity prediction system 106 determines the error using one or more loss functions, such as one or more loss functions used in training PrimateAI developed by Illumina, Inc. For instance, the pathogenicity prediction system 106 can use the loss function(s) to determine gradients with predicted scores as input together with observed / unknown primate variants as ground truth labels. The pathogenicity prediction system 106 can then backpropagate the gradients to the update the model parameters.[OHl] As shown in FIG. 2, the pathogenicity prediction system 106 derives one or more metrics from the inter-protein pathogenicity scores 206. In particular, FIG. 2 illustrates the pathogenicity prediction system 106 determining an inter-protein central tendency 212 and an interprotein standard deviation 214 from the inter-protein pathogenicity scores 206. As further shown, the pathogenicity prediction system 106 performs an act 216 of re-scaling using the intra-protein pathogenicity scores 210, the inter-protein central tendency 212, and the inter-protein standard deviation 214 to generate re-scaled pathogenicity scores 218. For the target protein sequence 202 (e.g., for the set of amino acids within the target protein sequence 202).

[0112] In one or more embodiments, the pathogenicity prediction system 106 performs the act 216 of re-scaling using one or more transformations. For instance, the pathogenicity prediction system 106 can re-scale the intra-protein pathogenicity scores 210 via one or more transformations using the inter-protein central tendency 212 and the inter-protein standard deviation 214. For instance, in some cases, the pathogenicity prediction system 106 re-scales the intra-protein pathogenicity scores 210 using a first set of transformations and a second set of transformations that is the inverse of the first set of transformations. In one or more embodiments, the first set of transformations and / or the second set of transformations include one or more linear transformations.Attorney Docket No. IP-2880-PCT 28 Patent Application

[0113] To illustrate, in some implementations, the pathogenicity prediction system 106 uses a first set of transformations to generate a set of standardized intra-protein pathogenicity scores from the intra-protein pathogenicity scores 210. As used herein, the term “standardized intra-protein pathogenicity score” refers to an intra-protein pathogenicity score that has been standardized. In particular, a standardized intra-protein pathogenicity score can refer to an intra-protein pathogenicity score that has been standardized with respect to one or more other intra-protein pathogenicity scores via one or more transformations. To illustrate, in some cases, a standardized intra-protein pathogenicity score can include an intra-protein pathogenicity score that has been standardized via one or more transformations using an intra-protein central tendency and an intra-protein standard deviation.

[0114] Indeed, in one or more embodiments, the pathogenicity prediction system 106 determines an intra-protein central tendency and an intra-protein standard deviation from the intra-protein pathogenicity scores 210. Thus, in some cases, the pathogenicity prediction system 106 generates a set of standardized intra-protein pathogenicity scores from the intra-protein pathogenicity scores 210 using a first set of transformations, the intra-protein central tendency, and the intra-protein standard deviation as follows:s_intra_std = (s ntra — mean(sjntr )) / std(s_intr ) (1)

[0115] In function (1), sjntra represents an intra-protein pathogenicity score from the intra-protein pathogenicity scores 210, mean sjntra) represents the intra-protein central tendency (in this case indicated as a mean), std sjntra) represents the intra-protein standard deviation, and s ntra_std represents a standardized intra-protein pathogenicity score resulting from the first set of transformations. Thus, in some cases, the pathogenicity prediction system 106 transforms each intra-protein pathogenicity score from the intra-protein pathogenicity scores 210 using the first set of transformations represented by function (1) to obtain a corresponding standardized intra-protein pathogenicity score.

[0116] To further illustrate, the pathogenicity prediction system 106 can use a second set of transformations to generate the re-scaled pathogenicity scores 218 from the standardized intra-protein pathogenicity scores obtained via the first set of transformations. In some cases, the pathogenicity prediction system 106 performs the second set of transformations using the interprotein central tendency 212 and the inter-protein standard deviation 214 determined from the interprotein pathogenicity scores 206. For instance, in pathogenicity prediction system 106 can perform the second set of transformations as follows:Attorney Docket No. IP-2880-PCT 29 Patent Application

[0117] In function (2), mean sjnter) represents the inter-protein central tendency 212, std sj.nter)' represents the inter-protein standard deviation 214, and s_intr a_sr represents a rescaled pathogenicity score resulting from the second set of operations. Thus, in some cases, the pathogenicity prediction system 106 transforms each standardized intra-protein pathogenicity score determined from the intra-protein pathogenicity scores 210 using the second set of transformations represented by function (2) to obtain a corresponding re-scaled pathogenicity score.

[0118] As indicated, the second set of transformations represented by function 2 are the inverse of the first set of transformations represented by function 1. More specifically, the second set of transformations includes uses the inter-protein central tendency 212 and the inter-protein standard deviation 214 to perform the inverse of the first set of transformations, which uses the intra-protein central tendency and the intra-protein standard deviation. For instance, while the first set of transformations involves subtracting the intra-protein central tendency, the second set of transformations involves adding the inter-protein central tendency 212. Further, while the first set of transformations involves dividing by the intra-protein standard deviation, the second set of transformations involves multiplying by the inter-protein standard deviation 214.

[0119] As mentioned, the pathogenicity prediction system 106 can generate a re-scaled pathogenicity score from each standardized intra-protein pathogenicity score. Further, the pathogenicity prediction system 106 can generate a standardized intra-protein pathogenicity score from each intra-protein pathogenicity score. Thus, in some embodiments, the pathogenicity prediction system 106 can generate the re-scaled pathogenicity scores 218 by generating a multidimensional data set of re-scaled pathogenicity scores of the same dimensions as the inter-protein pathogenicity scores 206 and the intra-protein pathogenicity scores 210. As an example, where the inter-protein pathogenicity scores 206 and the intra-protein pathogenicity scores 210 each include L x 20 pathogenicity scores where L represents the length of the target protein sequence 202, the pathogenicity prediction system 106 can generate the re-scaled pathogenicity scores 218 also having L x 20 pathogenicity scores.

[0120] In generating the re-scaled pathogenicity scores 218 as described above, the pathogenicity prediction system 106 can combine the inter-protein pathogenicity scores 206 and the intra-protein pathogenicity scores 210 while preserving properties of both. For instance, the rescaled pathogenicity scores 218 preserves the inter-protein central tendency 212 and the interprotein standard deviation 214 of the inter-protein pathogenicity scores 206. Further, the re-scaled pathogenicity scores 218 preserves the relative ranking of the intra-protein pathogenicity scores 210. Indeed, the pathogenicity prediction system 106 generates the re-scaled pathogenicity scores 218 according to a ranking of the set of intra-protein pathogenicity scores. For instance, in some cases, the re-scaled pathogenicity scores 218 includes a same pathogenicity ranking of the aminoAttorney Docket No. IP-2880-PCT 30 Patent Applicationacids from the target protein sequence 202 as indicated by the intra-protein pathogenicity scores 210. As an example, where the intra-protein pathogenicity scores 210 indicates that a given amino acid is the most pathogenic out of the amino acids from the target protein sequence 202, the rescaled pathogenicity scores 218 also indicate that the given amino acid is the most pathogenic even though the value of the pathogenicity score assigned to the amino acid within the re-scaled pathogenicity scores 218 may have changed through the transformations.

[0121] Thus, the pathogenicity prediction system 106 can operate with improved flexibility when compared to many existing pathogenicity prediction models. For instance, by using both inter- and intra-protein pathogenicity scores, the pathogenicity prediction system 106 flexibly incorporates multiple contexts into its determination of pathogenicity. In particular, the pathogenicity prediction system 106 determines pathogenicity across proteins and within the same protein. As such, the pathogenicity prediction system 106 provides a more flexible view of pathogenicity. In doing so, the pathogenicity prediction system 106 also operates with improved accuracy. In particular, the pathogenicity prediction system 106 more accurately predicts the pathogenicity of amino acids in a target protein sequence, including amino-acid variants.

[0122] As previously mentioned, in some embodiments, the pathogenicity prediction system 106 generates re-scaled pathogenicity scores for amino acids in a target protein sequence using multiple sets of intra-protein pathogenicity scores generated by a plurality of intra-protein pathogenicity models. FIG. 3 illustrates, the pathogenicity prediction system 106 generating rescaled pathogenicity scores using multiple sets of intra-protein pathogenicity models in accordance with one or more embodiments.

[0123] Indeed, as shown in FIG. 3, the pathogenicity prediction system 106 analyzes a target protein sequence 302 including a set of amino acids. Additionally, as shown, the pathogenicity prediction system 106 accesses a set of inter-protein pathogenicity scores 306 generated by an interprotein pathogenicity model 304 for the amino acids of the target protein sequence 302 (e.g., by using the inter-protein pathogenicity model 304 to generate the set of inter-protein pathogenicity scores 306 or by retrieving or receiving pre-generated scores). Further, the pathogenicity prediction system 106 determines an inter-protein central tendency 308 and an inter-protein standard deviation 310 from the set of inter-protein pathogenicity scores 306.

[0124] As shown in FIG. 3, the pathogenicity prediction system 106 accesses a plurality of sets of intra-protein pathogenicity scores 314a-314n generated by a plurality of intra-protein pathogenicity models 312a-312n for the amino acids of the target protein sequence 302. For instance, in some cases, the pathogenicity prediction system 106 retrieves or receives pre-generated pathogenicity scores. In some implementations, however, the pathogenicity prediction system 106Attorney Docket No. IP-2880-PCT 31 Patent Applicationemploys the plurality of intra-protein pathogenicity models 312a-312n to generate the plurality of sets of intra-protein pathogenicity scores 314a-314n.

[0125] In one or more embodiments, each intra-protein pathogenicity model from the plurality of intra-protein pathogenicity models 312a-312n includes a different pathogenicity model. For instance, the intra-protein pathogenicity model 312a can include a first pathogenicity model generating the set of intra-protein pathogenicity scores 314a to provide a first measure of pathogenicity for the set of amino acids from the target protein sequence 302. Further, the intra-protein pathogenicity model 312n can include an wth pathogenicity model generating the set of intra-protein pathogenicity scores 314n to provide an wth measure of pathogenicity for the set of amino acids. Thus, the pathogenicity prediction system 106 can use the plurality of sets of intra-protein pathogenicity scores 314a-314n to implement a balanced measure of intra-protein pathogenicity.

[0126] Further, as shown in FIG. 3, the pathogenicity prediction system 106 performs an act 316 of re-scaling using the plurality of sets of intra-protein pathogenicity scores 314a-314n, the inter-protein central tendency 308, and the inter-protein standard deviation 310 to generate a plurality of sets of re-scaled pathogenicity scores 318a-318n. In some cases, the pathogenicity prediction system 106 generates each set of re-scaled pathogenicity scores from a corresponding set of intra-protein pathogenicity scores via the act 316. For instance, the pathogenicity prediction system 106 can generate the set of re-scaled pathogenicity scores 318a from the set of intra-protein pathogenicity scores 314a via the act 316. Similarly, the pathogenicity prediction system 106 can generate the set of re-scaled pathogenicity scores 318n from the set of intra-protein pathogenicity scores 314n via the act 316.

[0127] In one or more embodiments, the pathogenicity prediction system 106 performs the act 316 of re-scaling as described above with reference to FIG. 2. In particular, the pathogenicity prediction system 106 can generate a set of re-scaled pathogenicity scores from a corresponding set of intra-protein pathogenicity scores using the inter-protein central tendency 308 and the interprotein standard deviation 310. For instance, the pathogenicity prediction system 106 can use a first set of transformations (e.g., represented by function 1) to generate a set of standardized intra-protein pathogenicity scores from the set of intra-protein pathogenicity scores using an intra-protein central tendency and an intra-protein standard deviation determined from the set of intra-protein pathogenicity scores. The pathogenicity prediction system 106 can further use a second set of transformations (represented by function 2) to generate the set of re-scaled pathogenicity scores from the set of standardized intra-protein pathogenicity scores using the inter-protein central tendency 308 and the inter-protein standard deviation 310.Attorney Docket No. IP-2880-PCT 32 Patent Application

[0128] As shown in FIG. 3, the pathogenicity prediction system 106 further performs an act 320 of combining the plurality of sets of re-scaled pathogenicity scores 318a-318n to generate a combined set of re-scaled pathogenicity scores 322. In one or more embodiments, the pathogenicity prediction system 106 performs the act 320 by linearly combining the plurality of sets of re-scaled pathogenicity scores 318a-318n. For instance, in some cases, the pathogenicity prediction system 106 assigns a weight to each set of re-scaled pathogenicity scores (e.g., assigns the weight to each re-scaled pathogenicity score in the set) and then adds the weighted scores. In some cases, the pathogenicity prediction system 106 uses an equal weighting for each set of re-scaled pathogenicity scores. In certain embodiments, however, the pathogenicity prediction system 106 uses weightings learned via back propagation. Indeed, in some cases, the pathogenicity prediction system 106 employs the act 320 of combining as part of a machine learning process that is trained on training data, and the pathogenicity prediction system 106 leams the parameters for combining the sets of re-scaled pathogenicity scores (e.g., the weighting for each set) during the training process.

[0129] In one or more embodiments, the pathogenicity prediction system 106 combines the plurality of sets of re-scaled pathogenicity scores 318a-318n by combining re-scaled pathogenicity scores that correspond to the same amino acid in the target protein sequence 302 (e.g., that indicate a pathogenicity of the same amino acid). For instance, the pathogenicity prediction system 106 can combine the first score in each set of re-scaled pathogenicity scores to determine a first score of the combined set of re-scaled pathogenicity scores 322 due to the first score in each set of re-scaled pathogenicity scores corresponding to the first amino acid of the target protein sequence 302. Likewise, the pathogenicity prediction system 106 can combine the / 7 th score in each set of rescaled pathogenicity scores to determine an wth score of the combined set of re-scaled pathogenicity scores 322 due to the wth score in each set of re-scaled pathogenicity scores corresponding to the wth amino acid of the target protein sequence 302.

[0130] Thus, the pathogenicity prediction system 106 can generate the combined set of rescaled pathogenicity scores 322 to preserve properties of the set of inter-protein pathogenicity scores 306 and the plurality of sets of intra-protein pathogenicity scores 314a-314n. For instance, the pathogenicity prediction system 106 preserves the inter-protein central tendency 308 and the inter-protein standard deviation 310 determined from the set of inter-protein pathogenicity scores 306. Further, the pathogenicity prediction system 106 preserves a ranking indicated by the plurality of sets of intra-protein pathogenicity scores 314a-314n. For instance, by generating the combined set of re-scaled pathogenicity scores 322 via a linear combination, the pathogenicity prediction system 106 can incorporate a pathogenicity ranking of the plurality of sets of intra-protein pathogenicity scores 314a-314n that is determined by the linear combination.Attorney Docket No. IP-2880-PCT 33 Patent Application

[0131] As indicated above with respect to FIG. 3, the pathogenicity prediction system 106 can re-scale a plurality of sets of intra-protein pathogenicity scores and then combine the sets of rescaled pathogenicity scores. In some embodiments, however, the pathogenicity prediction system 106 combines the plurality of sets of intra-protein pathogenicity scores and re-scales the combined intra-protein pathogenicity scores. FIG. 4 illustrates the pathogenicity prediction system 106 combining intra-protein pathogenicity scores and re-scaling the combined intra-protein pathogenicity scores in accordance with one or more embodiments.

[0132] As shown in FIG. 4, the pathogenicity prediction system 106 accesses a plurality of sets of intra-protein pathogenicity scores 404a - 404n generated by a plurality of intra-protein pathogenicity models 402a - 402n for a set of amino acids of a target protein sequence. As further shown, the pathogenicity prediction system 106 uses forward feature selection 406 to determine a combined set of intra-protein pathogenicity scores 408 from the plurality of sets of intra-protein pathogenicity scores 404a - 404n. In some cases, the pathogenicity prediction system 106 uses the forward feature selection 406 to determine the combined set of intra-protein pathogenicity scores 408 from a subset of the sets from the plurality of sets of intra-protein pathogenicity scores 404a -404n (e.g., a subset consisting of less than all of the available sets). In some instances, the pathogenicity prediction system 106 uses the forward feature selection 406 to select a subset of intra-protein pathogenicity models from the plurality of intra-protein pathogenicity models 402a -402n and uses the intra-protein pathogenicity scores from the selected models. In other words, the pathogenicity prediction system 106 can perform the selection with respect to the intra-protein pathogenicity models or with respect to the intra-protein pathogenicity scores themselves.

[0133] It should be noted, however, that in some implementations, the pathogenicity prediction system 106 uses the intra-protein pathogenicity scores from all available intra-protein pathogenicity models rather than selecting a subset of the scores (or models). Thus, in some cases, the pathogenicity prediction system 106 omits the forward feature selection 406 and combines all available sets of intra-protein pathogenicity scores.

[0134] In one or more embodiments, the term “forward feature selection” includes a computer-implemented model or algorithm for selecting a combination of features or values from various available combinations of features or values. In particular, forward feature selection can include a computer-implemented model or algorithm that evaluates various available combinations of features or values based on some metric and selects a combination for further use based on the evaluation. Indeed, in some cases, forward feature selection involves evaluating various available combinations of intra-protein pathogenicity scores (or the intra-protein pathogenicity models generating the scores) to select a subset of the intra-protein pathogenicity scores to use as aAttorney Docket No. IP-2880-PCT 34 Patent Applicationcombination. In some cases, the pathogenicity prediction system 106 implements the forward feature selection 406 via a machine learning model, such as a neural network.

[0135] To illustrate, in some embodiments, the forward feature selection 406 involves iteratively building a combination of features (e.g., scores or models) — one feature at a time — by selecting the feature from the set of available features that maximizes the combination with respect to some metric. For instance, in some cases, the pathogenicity prediction system 106 implements the forward feature selection 406 starting with an empty set. The pathogenicity prediction system 106 further adds the most predictive feature (e.g., the feature that improves the feature combination with respect to some metric) in each iteration until a stopping criterion is met (e.g., a number of iterations is completed). Indeed, in some cases, for a given iteration, each feature that has yet to be selected is evaluated in combination with the features that were selected in previous iterations, and the feature that is determined to improve the combination with respect to some metric is then added to the combination of features.

[0136] The metric used for evaluating the available combinations via the forward feature selection 406 can be a pre-defined metric or a user-selected metric. Indeed, the pathogenicity prediction system 106 can use various metrics in various implementations. Some examples of the metric include one or more assays, prime observed variants, or proteomics scores.

[0137] Thus, the pathogenicity prediction system 106 determines the combined set of intraprotein pathogenicity scores 408. In some cases, the pathogenicity prediction system 106 combines the selected sets of pathogenicity scores by combining a given pathogenicity score from a given set of pathogenicity scores with the corresponding pathogenicity score from the other sets of pathogenicity scores. In other words, the pathogenicity prediction system 106 can combine the scores from each set that corresponds to the same amino acid from the target protein sequence 302.

[0138] Further, in some cases, the pathogenicity prediction system 106 determines the combined set of intra-protein pathogenicity scores 408 by determining a linear combination of the sets of intra-protein pathogenicity scores selected via the forward feature selection 406. For instance, the pathogenicity prediction system 106 can assign a weight to each set of intra-protein pathogenicity scores (e.g., assign the weight to each score included in the set) and combine the weighted scores. In some cases, the pathogenicity prediction system 106 uses an equal weighting for each set of intra-protein pathogenicity scores. In certain embodiments, however, the pathogenicity prediction system 106 uses weightings learned via back propagation. Indeed, as mentioned, the pathogenicity prediction system 106 can employs the forward feature selection 406 via a machine learning model, and the pathogenicity prediction system 106 can leam the parameters for combining the selected intra-protein pathogenicity scores (e.g., the weighting for each set) during the training process.Attorney Docket No. IP-2880-PCT 35 Patent Application

[0139] As FIG. 4 illustrates, the pathogenicity prediction system 106 further determines an inter-protein central tendency 410 and an inter-protein standard deviation 412. For instance, the pathogenicity prediction system 106 can determine the inter-protein central tendency 410 and the inter-protein standard deviation 412 from a set of inter-protein pathogenicity scores generated by an inter-protein pathogenicity model as described above with reference to FIG. 2.

[0140] As further shown in FIG. 4, the pathogenicity prediction system 106 performs an act 414 of re-scaling to generate the re-scaled pathogenicity scores 416 for the set of amino acids of the target protein sequence. In particular, the pathogenicity prediction system 106 performs the act 414 of re-scaling using the combined set of intra-protein pathogenicity scores 408, the inter-protein central tendency 410, and the inter-protein standard deviation 412. For instance, the pathogenicity prediction system 106 can re-scale the combined set of intra-protein pathogenicity scores 408 using the inter-protein central tendency 410 and the inter-protein standard deviation 412 via one or more transformations as described above.

[0141] Thus, in some implementations, the pathogenicity prediction system 106 generates rescaled pathogenicity scores for a set of amino acids of a target protein sequence by intelligently combining multiple sets of intra-protein pathogenicity scores. Indeed, the pathogenicity prediction system 106 can intelligently select (e.g., via a machine learning approach) from a plurality of available sets of intra-protein pathogenicity scores to combine. Further, the pathogenicity prediction system 106 transforms a combination of the selected sets of intra-protein pathogenicity scores using metrics derived from inter-protein pathogenicity scores (e.g., the inter-protein central tendency 410 and the inter-protein standard deviation 412). Thus, the pathogenicity prediction system 106 can preserve properties of the set of inter-protein pathogenicity scores as well as properties of the intelligently selected sets of intra-protein pathogenicity scores.

[0142] By incorporating multiple sets of intra-protein pathogenicity scores generated by different intra-protein pathogenicity models into its pathogenicity prediction, the pathogenicity prediction system 106 can further improve its flexibility and accuracy over existing pathogenicity prediction models. Indeed, the pathogenicity prediction system 106 can flexibility balance the scores of the different intra-protein pathogenicity models for an improved prediction result.

[0143] Indeed, as discussed, the pathogenicity prediction system 106 can generate more accurate pathogenicity predictions for amino acids when compared to many conventional systems. FIGS. 5A-5B illustrate graphs showing experimental results regarding the effectiveness of the pathogenicity prediction system 106 generating pathogenicity predictions for amino acids of a target protein sequence in accordance with one or more embodiments.

[0144] In particular, the graphs of FIGS. 5A-5B compares the performance of one or more embodiments of the pathogenicity prediction system 106 (labeled “PAI3D / SR(OER) (rank) SR”)Attorney Docket No. IP-2880-PCT 36 Patent Applicationwith several existing state-of-the-art amino-acid variant pathogenicity prediction systems, including an AlphaMissense model, a protein assembly interaction (PAI) model for three dimensions, and several variations of a concurrent logical framework (CLF) model. The graphs compare the performances of the tested models using various metrics. In particular, the graphs illustrate how the pathogenicity prediction of each tested model correlates with some other measure of pathogenicity using various metrics.

[0145] For instance, the first graph 500a of FIG. 5A compares the pathogenicity predictions of each tested model with a measure of pathogenicity determined via a set of assays, which involve lab-based measurements of protein functions given a certain mutation in the protein (e.g., given a certain amino-acid variant). For instance, in some cases, one amino acid is changed within a protein and then the protein is evaluated to measure how well the protein still works (e.g., compared to the unchanged or reference protein). Thus, the assays of the first graph 500a provide experimentally derived pathogenicity scores to which the pathogenicity predictions of the tested models are compared. The first graph 500a illustrates correlation measured via Spearman correlation.

[0146] The second graph 500b of FIG. 5A compares the pathogenicity predictions of each tested model with a measure of pathogenicity determined via one or more deciphering developmental disorders (DDD) studies. For instance, the DDD study was used with the PrimateAI model developed by Illumina, Inc. to determine how well the PrimateAI model could distinguish between variants in people with developmental disorders from variants in healthy people. As shown, the second graph 500b illustrates correlation via the p-value of the Mann-Whitney U test.

[0147] The third graph 500c of FIG. 5A compares the pathogenicity predictions of each tested model with a measure of pathogenicity determined via the ClinVar public archive. As shown, the third graph 500c illustrates the correlation using via the area under the curve (AUC) metric.

[0148] The fourth graph 500d of FIG. 5B compares the pathogenicity predictions of each tested model with a measure of pathogenicity determined via the UK Biobank (UKBB) study and illustrates the correlations via Spearman correlation. The fifth graph 500e compares the pathogenicity predictions with a measure determined via one or more autism spectrum disorder (ASD) studies and illustrates the correlation via the p-value of the Mann-Whitney U test. The sixth graph 500f also compares the pathogenicity predictions with a measure determined via one or more assays measuring protein levels in people with a variant (e.g., an amino-acid variant) in a particular gene and illustrates the correlations via Spearman correlation.

[0149] As shown by the graphs in FIGS. 5A-5B, the embodiment(s) of the pathogenicity prediction system 106 outperform the other tested models — in many cases, outperforming the other models significantly — in most situations. As shown by the third graph 500c of FIG. 5 A, many of the other models do outperform the pathogenicity prediction system 106, but these models wereAttorney Docket No. IP-2880-PCT 37 Patent Applicationtrained on the ClinVar dataset. Thus, these other models were better configured to perform well in that particular scenario. These models failed, however, to outperform the pathogenicity prediction system 106 more generally.

[0150] FIGS. 6A-6B illustrate additional experimental results regarding the effectiveness of the pathogenicity prediction system 106 generating pathogenicity predictions for amino acids of a target protein sequence in accordance with one or more embodiments. In particular, FIGS. 6A-6B illustrate additional experimental results regarding the effectiveness of one or more embodiments of the pathogenicity prediction system 106 that employ forward feature selection to combine multiple sets of intra-protein pathogenicity scores generated by different intra-protein pathogenicity models.

[0151] For instance, the graphs of FIG. 6A illustrate how embodiments of the pathogenicity prediction system 106 improves performance in both its pathogenicity scores and pathogenicity ranking with respect to several comparative measures, including the assays and the UKBB study discussed above with reference to FIGS. 5A-5B. In particular, the graphs show that performance improves through several of the initial selection iterations and levels off thereafter. Thus, the pathogenicity prediction system 106 can use forward feature selection to improve its prediction of the pathogenicity of an amino acid as well as the relative pathogenicity of the amino acid with respect to other amino acids in the same set.

[0152] FIG. 6B illustrates one example of intra-protein pathogenicity models (or corresponding scores) selected during the forward feature selection process. FIG. 6B shows that, during each iteration, the number of intra-protein pathogenicity models (or corresponding scores) selected by the pathogenicity prediction system 106 grows compared to the preceding iteration. Thus, the pathogenicity prediction system 106 uses forward feature selection to iteratively build a combination of intra-protein pathogenicity models (or corresponding intra-protein pathogenicity scores) when determining its pathogenicity predictions for a set of amino acids.

[0153] As indicated by FIG. 6B, in some instances, an intra-protein pathogenicity model that is included in one iteration may be omitted during a subsequent iteration. In other words, the pathogenicity prediction system 106 can remove a previously selected model during the forward feature selection process. For instance, in some cases, the pathogenicity prediction system 106 determines that a previously selected model no longer optimizes performance when included in the set of selected intra-protein pathogenicity models. Thus, the pathogenicity prediction system 106 can add or remove models (or their scores) as necessary during forward feature selection to obtain the best performance. In some cases, however, once the pathogenicity prediction system 106 selects a model (or its scores) during a given iteration, the pathogenicity prediction system 106 maintains that model (or its scores) during subsequent iterations.Attorney Docket No. IP-2880-PCT 38 Patent Application

[0154] In one or more embodiments, the pathogenicity prediction system 106 implements a window-based approach to determining pathogenicity predictions for amino acids of a target protein sequence. FIGS. 7A-7B illustrate graphs showing a window-based approach implemented by the pathogenicity prediction system 106 in accordance with one or more embodiments.

[0155] For instance, the graph of FIG. 7A illustrates the pathogenicity prediction system 106 determining pathogenicity predictions for amino acids within an amino acid window. In other words, the pathogenicity prediction system 106 can determine a window of a particular size (e.g., a size that fits a particular number of amino acids), and the pathogenicity prediction system 106 determines pathogenicity predictions for the amino acids within the amino acid window. Indeed, as previously suggested, the target protein sequence for which pathogenicity predictions are determined can include a subset of the amino acids of the corresponding protein. In some case, the amino acid window includes a three-dimensional window to accommodate the three-dimensional structure of the protein. In some cases, the amino acid window includes a user-defined window. In certain instances, however, the amino acid window includes a window covering a certain portion of a protein, such as covering a portion where missense mutations are possible (or likely).

[0156] The graph of FIG. 7B illustrates the pathogenicity prediction system 106 determining pathogenicity predictions for amino acids within a window defined by a range of inter-protein pathogenicity scores. Indeed, in some cases, the pathogenicity prediction system 106 limits the analysis to those amino acids that fall within the range of inter-protein pathogenicity scores. Thus, the pathogenicity prediction system 106 can omit extreme values that fall outside that range.

[0157] In certain cases, the pathogenicity prediction system 106 applies such a window when the distribution of inter-protein pathogenicity scores is not normal. As an example, where multiple peaks are found in the distribution, the pathogenicity prediction system 106 can put each peak within its own range, allowing the pathogenicity prediction system 106 to perform the pathogenicity prediction for each peak separately.

[0158] In some cases, the pathogenicity prediction system 106 uses multiple windows in its pathogenicity analysis. For instance, the pathogenicity prediction system 106 can use a first window that includes a particular number of amino acids and a second window that includes only those amino acids falling within a particular range of inter-protein pathogenicity scores. Thus, the pathogenicity prediction system 106 can implement various configurations to more specifically target particular amino acids.

[0159] As discussed above, the pathogenicity prediction system 106 can analyze a target protein sequence and generate one or more inter-protein pathogenicity score(s). In particular, the pathogenicity prediction system 106 can generate the one or more inter-protein pathogenicity score(s) for a set of amino acids included in the target protein sequence, with each inter-proteinAttorney Docket No. IP-2880-PCT 39 Patent Applicationpathogenicity score specific to a target position in the target protein sequence. In accordance with one or more embodiments, FIG. 8 illustrates an overview of the pathogenicity prediction system generating an inter-protein pathogenicity score for a target position based on an observed number of variants and an expected number of variants within a position window of the target position.

[0160] Indeed, as shown, the pathogenicity prediction system 106 analyzes a target protein sequence 802. The target protein sequence 802 includes an amino acid at a target position 801c and contextual amino acids at proximate positions 801a, 801b, 801d, 801e that are proximate to the target position 801c. As shown in FIG. 8, a position window 805 includes the target position 801c and proximate positions 801a, 801b, 801d, 801e. For instance, in some cases, the position window 805 includes a plurality of amino acids in a sequential order. In some instances, the position window 805 includes proximate positions determined from a three-dimensional structure associated with the target protein sequence 802.

[0161] In some cases, the set of amino acids of the target protein sequence 802 includes one or more variants (e.g., amino-acid variants). For instance, the set of amino acids can include one or more amino-acid variants caused by one or more corresponding missense nucleotide variants. Additionally or alternatively, the set of amino acids can include one or more amino-acid variants caused by one or more corresponding insertions or deletions (indels). Thus, in some cases, the pathogenicity prediction system 106 can determine the pathogenicity of variants present within target protein sequences.

[0162] Having identified the target and proximate positions, as shown in FIG. 8, the pathogenicity prediction system 106 determines an observed number of variants 804 among amino acids within the position window 805, e.g., for the target position 801c and the proximate positions 801a, 801b, 801d, 801e. In some cases, to determine the observed number of variants 804, the pathogenicity prediction system 106 accesses variant data 810 for each of the amino acids of the position window 805. The variant data 810 may be derived from sequencing data from multiple individuals, such as population sequencing studies and / or aggregated sequencing studies such as the Genome Aggregation Database (gnomAD™). Based on the variant data 810, the pathogenicity prediction system 106 determines an observed number of variants at each amino acid of the position window 805.

[0163] In addition to determining the observed number of variants 804, as further shown in FIG. 8, the pathogenicity prediction system 106 determines an expected number of variants 806 among amino acids within the position window 805, e.g., for the target position 801c and the proximate positions 801a, 801b, 801d, and 801e. In some cases, the pathogenicity prediction system 106 uses a mutation rate model 807 to determine an expected number of variants 806. ForAttorney Docket No. IP-2880-PCT 40 Patent Applicationexample, the mutation rate model 807 can generate an expected number of variants 806 for each of the amino acids of the position window 805, based on sequence information and other data.

[0164] Based on the observed number of variants 804 and the expected number of variants 806, as further shown in FIG. 8, the pathogenicity prediction system 106 generates an inter-protein pathogenicity score 808 specific to the target position 801c. In some embodiments, the pathogenicity prediction system 106 generates the inter-protein pathogenicity score 808 based on a relationship (for example, a ratio) between the observed number of variants 804 and the expected number of variants 806. This disclosure describes examples of such an observed-to-expected variant relationship below.

[0165] In some implementations, the pathogenicity prediction system 106 repeats one or more steps shown in FIG. 8 for a further target position in the target protein sequence 802. For example, the pathogenicity prediction system 106 can select, for example, a subsequent amino acid at a subsequent target position, such as the previously identified proximate position 80 Id. In such an iteration, the proximate position 801 d becomes the new or subsequent target position. Accordingly, the pathogenicity prediction system 106 can determine a subsequent set of proximate positions relative to the proximate position 80 Id, determine the observed number of variants and expected number of variants for the subsequent position window, and generate an inter-protein pathogenicity score specific to the subsequent target position (e.g., previously identified proximate position 80 Id). The pathogenicity prediction system 106 can repeat the process over each position in the target protein sequence 802 in a sliding window approach to generate a set of inter-protein pathogenicity scores.

[0166] As previously mentioned, in some embodiments, the pathogenicity prediction system 106 determines proximate positions for contextual amino acids within the position window based on three-dimensional distances. FIG. 9A illustrates one such example by showing the pathogenicity prediction system 106 identifying proximate positions within a position window based on a three-dimensional distance of the target position in accordance with one or more embodiments.

[0167] As shown in FIG. 9 A, a target protein sequence 900 includes a target position 901a. For example, in table 903 of FIG. 9A, the pathogenicity prediction system 106 ranks positions in the target protein sequence based on three-dimensional proximity to the target position 901a. The pathogenicity prediction system 106 identifies proximate positions 902 (individually labeled 1, 2, 3 . . . 14) of contextual amino acids based on a three-dimensional distance of the target position 901a in a three-dimensional structure of the target protein sequence 900. Although not shown in the table 903 representing the proximate positions 902 in FIG. 9A, in some embodiments, the position window further includes the target position 901a. Thus, in some embodiments, the position window includes the target position 901a and the proximate positions 902.Attorney Docket No. IP-2880-PCT 41 Patent Application

[0168] For example, in FIG. 9A, the pathogenicity prediction system 106 has a predetermined number (n) of proximate positions 902 that are included in the position window. The pathogenicity prediction system 106 can access 3D distance data 906 that includes information about the relative 3 -dimensional distance between the target position 901a and other amino acids of the target protein sequence 900. The pathogenicity prediction system 106 can use the 3D distance data 906 determine the closest n amino acids to the target position 901a. In some instances, for instance, the pathogenicity prediction system 106 ranks amino acids by 3 -dimensional distance to the target position 901a and determines the top-ranked n amino acids with the shortest 3-dimensional distance to the target position 901a.

[0169] In the example shown in FIG. 9A, the predetermined number n is 14. Accordingly, the pathogenicity prediction system 106 identifies the 14 amino acids with the shortest 3D distance to the target position 901a as the proximate positions 902 (individually labeled 1, 2, 3 . . . 14) for the position window. Positions for amino acids outside of the predetermined number n (e.g., ranked 15, 16, 17 and so forth) are not included as proximate positions 902 in the position window.

[0170] The pathogenicity prediction system 106 can determine 3D distance in various ways. In some embodiments, the three-dimensional distance between a target position and a contextual amino acid is determined based on distance between a beta carbon of the amino acid of the target position and a beta carbon of the contextual amino acid. In some embodiments, the three-dimensional distance between a target position and a contextual amino acid is determined based on distance between an alpha carbon of the amino acid of the target position and an alpha carbon of the contextual amino acid. In some embodiments (including embodiments measuring distance between beta carbons or alpha carbons), distance is measured in angstroms (A).

[0171] While FIG. 9A depicts the size of the position window at 15 protein positions (based on one target position shown as the target position 901a and fourteen proximate positions shown as the proximate positions 902), one of ordinary skill in the art will appreciate that the position window size can be adjusted to include more or fewer protein positions as needed. For example, a window size can be 5, 11, 21, 31, 41, 51, 61, 71, 81, 91, or 101 protein positions, or any number of protein positions therebetween or any range constructed from these values. In some cases, a window size is an odd number that includes one central position (e.g., the target position) and two equal lengths (e.g., h, the proximate positions or half- window size) on either side of the central position. In some embodiments, the position window size or number of protein positions, including, but not limited to, more than 100 protein positions (e.g., 200, 300) or a window size that includes protein positions for the entire length of the target protein. While overly large position windows will only capture the average constraint of the whole target protein sequence without providing information about the local sequence context around a target position, overly small positionAttorney Docket No. IP-2880-PCT 42 Patent Applicationwindows containing information from very few protein positions are likely to result in a noisy signal that fails to capture true constraint information. Optimal window sizes can be identified by testing the performance of different position window sizes on their ability to discriminate between cohorts enriched for disease-causing variants versus presumed benign variation. Additionally, in some embodiments, the pathogenicity prediction system 106 can use a variable position window size based on a minimum number of expected variants per position window.

[0172] In addition or in the alternative to 3D distance, in some embodiments, the pathogenicity prediction system 106 determines a position window size based on a gene loss of function (LoF) constraint value for the target protein sequence 900. For example, the pathogenicity prediction system 106 can access a gene loss of function (LoF) constraint value for the target protein sequence 900; determine whether the gene LoF constraint value is above or below a LoF constraint threshold; and select a position window size based on the determination of whether the gene LoF constraint value is above or below the LoF constraint threshold. For example, the pathogenicity prediction system 106 can select a first a position window size if the gene LoF constraint value is above the LoF constraint threshold and select a second a position window size if the gene LoF constraint value is below the LoF constraint threshold. In some embodiments, the pathogenicity prediction system 106 uses multiple LoF constraint thresholds to flexibly determine a position window size.

[0173] As previously mentioned, in some embodiments, the pathogenicity prediction system 106 determines, for the target position and the proximate positions, an observed number of variants and an expected number of variants within the position window. Based on the observed number of variants and the expected number of variants, the pathogenicity prediction system 106 generates an inter-protein pathogenicity score specific to the target position. In accordance with one or more embodiments, FIG. 9B illustrates the pathogenicity prediction system 106 determining an observed number of variants and an expected number of variants. In accordance with one or more embodiments, FIG. 9C illustrates the pathogenicity prediction system 106 generating an interprotein pathogenicity score based on the observed and expected number of variants from FIG. 9B.

[0174] Indeed, as shown in FIG. 9B, the pathogenicity prediction system 106 determines an observed number of variants 905 for the target position and the proximate positions in the position window. For example, the pathogenicity prediction system 106 can access variant data from genetic sequencing from a plurality of individuals. In some such cases, the pathogenicity prediction system 106 can access aggregated data from one or more large-scale sequencing studies comprising variant data from hundreds, thousands, or hundreds of thousands of individuals or more.

[0175] For example, as shown in FIG. 9B, a nucleic acid sequence for a target protein sequence includes trinucleotides 907a, 907b, 907c, which correspond to positions in the target protein sequence that are included in the position window. The pathogenicity prediction system 106 canAttorney Docket No. IP-2880-PCT 43 Patent Applicationaccess variant data including nucleic acid sequences 904a, 904b through 904n which are taken from genetic sequencing of a plurality of individuals. The pathogenicity prediction system 106 can determine that the position corresponding to trinucleotide 907c includes an amino acid variant, in particular a missense variant. In the example of FIG. 9B, the pathogenicity prediction system 106 determines that the nucleic acid sequences 904b and 904n include a reference nucleobase C in trinucleotide 907c, resulting in an alanine amino acid at the position in the target protein sequence corresponding to trinucleotide 907c. The pathogenicity prediction system 106 determines that the nucleic acid sequence 904a includes an alternative nucleobase A in trinucleotide 907c, resulting in a glutamate amino acid at the position in the target protein sequence corresponding to trinucleotide 907c. The pathogenicity prediction system 106 counts the missense variant at the position corresponding to trinucleotide 907c in determining an observed number of variants 905 for the position window. The pathogenicity prediction system 106 further determines that the positions corresponding to trinucleotides 907a, 907b do not include a variant in the variant data, and thus the pathogenicity prediction system 106 determines a count of 0 for these positions in determining an observed number of variants 905 for the position window.

[0176] In some embodiments, the pathogenicity prediction system 106 determines the observed number of variants 905 based on variant data from humans and at least one additional species. For example, in some embodiments, the variant data comprises sequencing information from humans and at least one additional primate species. In some embodiments, the variant data comprises sequencing information from humans and at least one additional mammalian species. In some embodiments, the variant data comprises sequencing information from humans and at least one additional vertebrate species. In some embodiments, the variant data comprises sequencing information from a plurality of species.

[0177] In addition to individuals from different species, the pathogenicity prediction system 106 can determine the observed number of variants 905 based on variant data from different types of sequencing data or techniques. For example, in some embodiments, the variant data comprises sequencing information from whole genome sequencing of multiple individuals. In some embodiments, the variant data comprises sequencing information from whole exome sequencing studies of multiple individuals. In some embodiments, the variant data from non-human species is filtered to select for primate missense variants from primate sequences that map to homologous human sequences with a quality metric above a predetermined threshold. Methods of mapping between species are further described in WO2023129953 A2, entitled “Variant calling without a target reference genome,” published July 6, 2023, which is hereby incorporated by reference in its entirety.Attorney Docket No. IP-2880-PCT 44 Patent Application

[0178] To determines the observed number of variants 905, the pathogenicity prediction system 106 can employ several, non-exclusive approaches. In some embodiments, for example, the pathogenicity prediction system 106 counts variants as observed if it appears in at least one individual in the aggregated variant data. Thus, in some embodiments, the maximum number of observed variants in the position window is equal to the number of possible variants in the position window and does not depend on the number of individuals sequenced. Further, in some embodiments, the pathogenicity prediction system 106 determines the observed number of variants 905 based on determining the number of missense variants (and not synonymous variants) within the position window.

[0179] As a further example, in some embodiments, the pathogenicity prediction system 106 filters for variants in the variant data with a minimum allele frequency. To illustrate, the pathogenicity prediction system 106 analyzes the variant data and determines allele frequencies of variants included in at least a portion of the variant data (e.g., included in the portion of the variant data derived from human sequences). For example, the pathogenicity prediction system 106 can select, based on sequences in the variant data, variants (e.g., human missense variants) that satisfy a predetermined allele frequency threshold by occurring with at least a minimum frequency within the variant data. The pathogenicity prediction system 106 can further determine the observed number of variants or the expected number of variants based on the selected variants, thereby excluding variants that do not satisfy the allele frequency threshold.

[0180] In some embodiments, the pathogenicity prediction system 106 determines an allele frequency threshold based on a gene loss of function (LoF) constraint value for the target protein sequence. For example, the pathogenicity prediction system 106 can access a gene LoF constraint value for the target protein sequence; determine whether the gene LoF constraint value is above or below a LoF constraint threshold; and select an allele frequency threshold based on the determination of whether the gene LoF constraint value is above or below the LoF constraint threshold. Take a couple of allele frequencies, for instance, the pathogenicity prediction system 106 can select a first allele frequency threshold if the gene LoF constraint value is above the LoF constraint threshold and select a second allele frequency threshold if the gene LoF constraint value is below the LoF constraint threshold. Additionally, the LoF constraint threshold can flexibly be tuned to improve accuracy. In some embodiments, the pathogenicity prediction system 106 uses multiple LoF constraint thresholds to more flexibly determine an allele frequency threshold.

[0181] In addition to determining a position window comprising proximate positions, the pathogenicity prediction system 106 can apply different weights to different proximate positions and / or target positions. In some instances, the pathogenicity prediction system 106 weights proximate positions differently when determining the observed number of variants 905. ForAttorney Docket No. IP-2880-PCT 45 Patent Applicationexample, the pathogenicity prediction system 106 can determine proximate-position weights for the proximate positions based on three-dimensional or one-dimensional proximity of each position relative to the target position within the position window. To illustrate, proximate positions that are closer in three-dimensional space and / or in sequence space may be weighted heavier than proximate positions that are farther away. Afterwards, the pathogenicity prediction system 106 can determine the observed number of variants according to the proximate-position weights (e.g., by multiply, dividing, or otherwise adjusting the number of observed variants at a given proximate position by the corresponding weight).

[0182] Not all protein positions will include observed variants. Accordingly, when the actual observed number of variants is 0, in some embodiments, the pathogenicity prediction system 106 can use a pseudo count as the observed number of variants. For example, the pathogenicity prediction system 106 can determine that the observed number of variants is 0 within the position window; access, based on the observed number of variants being 0, a pseudo count of observed variants; and generate the inter-protein pathogenicity score specific to the target position based on the pseudo count of observed variants and the expected number of variants. In some instances, the pseudo count is a fixed value. By contrast, in some instances, the pathogenicity prediction system 106 flexibly determines the value of the pseudo count based one or more factors relative to the target protein sequence 900, the position window, and / or the target position 901a.

[0183] As further shown in FIG. 9B, the pathogenicity prediction system 106 determines an expected number of variants 910 for the target position and the proximate positions in the position window. In some embodiments, the pathogenicity prediction system 106 uses a mutation rate model 912 to determine the expected number of variants 910. For example, in some embodiments, the pathogenicity prediction system 106 uses the mutation rate model 912 to estimate the mutation rates for possible variants in the position window. After estimating the mutation rates, the pathogenicity prediction system 106 can determine the expected number of variants 910 in the position window based on the estimated mutation rates, for example by summing the estimated mutation rates for each possible variant in the position window.

[0184] As shown in FIG. 9B, the pathogenicity prediction system 106 can use the mutation rate model 912 to determine mutation rates based on one or more of several factors as described further below. In some embodiments, the mutation rate model 912 uses or is based on a trinucleotide context factor 913 a. For example, a mutation rate can be estimated for each trinucleotide, e.g., each of 96 possible trinucleotides, for example as disclosed in K. E. Samocha et al., ‘A framework for the interpretation of de novo mutation in human disease’, Nat. Genet., vol.46, no. 9, Art. no. 9, Sep. 2014, doi: 10.1038 / ng.3050, which is hereby incorporated by reference in its entirety. A single mutation rate can be conceptualized as the probability to see a variant at theAttorney Docket No. IP-2880-PCT 46 Patent Applicationcorresponding site in one generation by mechanical means only, i.e., excluding the impact of e.g., selection pressures and sequencing issues. Importantly, this defines the mutation rates of synonymous, missense and protein-truncating variants, because all variants can be associated with a trinucleotide context. In some embodiments, to transform this mutation rate so to an expected number of variants 910, the pathogenicity prediction system 106 sums up all mutation rates of a particular type (e.g., missense variants) across the genome and then divides each mutation rate by that sum. This converts the mutation rate of a site into a probability of mutating relative to all other sites in the genome (die analogy: P=l / 6 for all 6 sides; 1 die side=l genomic site). Put differently, if only a single mutation were introduced in a genome, the probabilities reflect how likely it is that the mutation ends up at a particular site (die analogy: throw die once).

[0185] To model introducing multiple variants, the pathogenicity prediction system 106 can multiply the probabilities by the number of variants to derive an expected number of variants 910 at each site (die analogy: if the die is thrown 120 times, 120*1 / 6=20 counts are expected for each side). In some embodiments, the number of variants introduced is the number of all observed variants (e.g., missense variants). Having converted the mutation rate of each variant into an expected count, the pathogenicity prediction system 106 can derive or otherwise determine the expected number of variants in a region or other position window by summing up the expected counts of all variants of a certain type (e.g., missense) in the position window (die analogy: throw 5 dice 120 times5*120*1 / 6=100 counts expected for each side).

[0186] As shown in FIG. 9B, in some embodiments, the mutation rate model 912 uses or is based on a methylation context factor 913b. For example, the mutation rate model 912 can take into account a methylation level within a trinucleotide context (e.g., Beta values estimating methylation levels). For example, for CpG transition mutations (4 different trinucleotide contexts), the mutation rate model 912 can take into account 16 different levels of methylation. The mutation rate model 912 could alternatively use a different number of levels of methylation for classification. Thus, in some embodiments, each CpG variant is additionally classified into one of 16 subclasses, generating a total of (96-4) + (4*16)=156 different trinucleotide contexts, instead of the former 96. Further methods for modelling mutation rates based on methylation levels are described in Siwei Chen et al., ‘A genome-wide mutational constraint map quantified from variation in 76,156 human genomes’, bioRxiv, p. 2022.03.20.485034, Jan. 2022, doi: 10.1101 / 2022.03.20.485034, which is hereby incorporated by reference in its entirety.

[0187] As further shown in FIG. 9B, in some embodiments, the mutation rate model 912 uses or is based on a proportion of synonymous variants factor 913c. For example, the pathogenicity prediction system 106 can fit an estimated mutation rate to a proportion of synonymous variants observed for each trinucleotide context. In some embodiments, adjusting the estimated mutationAttorney Docket No. IP-2880-PCT 47 Patent Applicationrate (and thus the expected number of variants 910) by the proportion of synonymous variants factor 913c control for the effect of having observed most possible sites in contexts with the highest mutation rates in the population sample.

[0188] In addition or in the alternative to the factors described above, in some embodiments, the mutation rate model 912 uses or is based on one or more of the following factors: an extended nucleotide context factor 913d, germline methylation and expression levels factor 913e, transcription and replication asymmetry factor 913f, or long-range mutation rate variation factor 913g. For example, a mutation rate model that takes such factors into account is described in Seplyarskiy et al., ‘A mutation rate model at the basepair resolution identifies the mutagenic effect of polymerase III transcription,’ Nat Genet. 2023 Dec;55(12):2235-2242. doi: 10.1038 / s41588-023-01562-0, which is hereby incorporated by reference in its entirety.

[0189] The pathogenicity prediction system 106 can use the mutation rate model 912 including one or any combination of the factors described above. Furthermore, in some embodiments, the pathogenicity prediction system 106 uses a combination of different mutation rate models. For example, the pathogenicity prediction system 106 may use a first mutation rate model for some protein positions or position windows, and a second mutation model for other positions protein positions or position windows (e.g., subsequent windows), for example, when data for a preferred mutation rate model is unavailable for a particular protein position. The pathogenicity prediction system 106 can thus use the mutation rate model 912 (including sub mutation rate models) to determine an expected number of variants for a target position and proximate positions.

[0190] As indicated above, in addition to determining a position window comprising proximate positions, the pathogenicity prediction system 106 can apply different weights to different proximate positions and / or target positions. In some instances, the pathogenicity prediction system 106 weights proximate positions differently when determining the expected number of variants 910. For example, the pathogenicity prediction system 106 can determine proximate-position weights for the proximate positions based on three-dimensional or onedimensional proximity of each position relative to the target position within the position window. Afterwards, the pathogenicity prediction system 106 can determine the expected number of variants 910 according to the proximate-position weights (e.g., by multiply, dividing, or otherwise adjusting the number of expected variants at a given proximate position by the corresponding weight).

[0191] Continuing now to FIG. 9C, as further shown in this figure, the pathogenicity prediction system 106 can generate inter-protein pathogenicity score(s) 920 specific to a target position, based on the observed number of variants and the expected number of variants in the position window. Consistent with the disclosure above, in some embodiments, the pathogenicity prediction system 106 determines a ratio based on the observed number of variants (represented as O below) and theAttorney Docket No. IP-2880-PCT 48 Patent Applicationexpected number of variants (represented as E below). In some embodiments, the pathogenicity prediction system 106 determines an observed-over-expected ratio of the observed number of variants (O) relative to the sum of the observed number of variants and the expected number of variants (O + E). In some embodiments, the pathogenicity prediction system 106 generates the inter-protein pathogenicity score(s) 920 for a target position as equal to, weight adjusted from, or otherwise based on the observed number of variants in the position window divided by the sum of the observed number of variants and the expected number of variants in the position window, represented as O / (O+E).

[0192] By applying a sliding position window, in some embodiments, the pathogenicity prediction system 106 can determine an observed over expected ratio (OOER) for different target positions on a position-by-position basis along with corresponding proximate positions for each target position. For example, as shown in FIG. 9C, the pathogenicity prediction system 106 can determine a first OOER based on numbers of observed and expected variants for a first position window 925a including a first target position 927a. Subsequently, the pathogenicity prediction system 106 can determine a second OOER based on numbers of observed and expected variants for a second position window 925b including a second target position 927b. Afterwards, the pathogenicity prediction system 106 can determine a second OOER based on numbers of observed and expected variants for a third position window 925c including a third target position 927c.

[0193] As indicated by FIG. 9C, the observed and expected variant counts and OOER determinations change as the pathogenicity prediction system 106 progresses through the first position window 925a, the second position window 925b, and the third position window 925c. When applying the first position window 925a, for instance, the pathogenicity prediction system 106 determines an OOER of 0.9 specific to the first target position 927a. When applying the third position window 925c, by contrast, the pathogenicity prediction system 106 determines an OOER of 0.7 specific to the third target position 927c. The OOER is accordingly both target-position specific but also dependent on the specific observed and expected variant counts for respective proximate positions within respective position windows.

[0194] As further shown in FIG. 9C, in some embodiments, the pathogenicity prediction system 106 generates the inter-protein pathogenicity score(s) 920 for missense variants based on a formula 928. The formula 928 is reproduced below as function (3). According to function (3) below, in some embodiments, the pathogenicity prediction system 106 generates the inter-protein pathogenicity score(s) 920 for a target position as equal to or otherwise based on the observed number of missense variants in the position window divided by the sum of the observed number of missense variants and the expected number of missense variants in the position window, PpmA (Omis +E mis)-Attorney Docket No. IP-2880-PCT 49 Patent Application

[0195] In function (3) above, OOERmisrepresents an OOER for missense variants, 0misrepresents an observed number of missense variants, and Emisrepresents an expected number of missense variants. The variable OOERmis, 0mis, and Emisrepresent the same components described here in the functions below.

[0196] As discussed above, not all protein positions will include observed variants. For example, some protein regions (and thus position windows) might have 0 observed variants. However, this does not necessarily mean that they are totally intolerant to variation, and might, in some instances, merely reflect an inability to detect rare mutations in the sequencing cohort used in the sequencing data. That is, if N more people were sequenced and included in the variant data, the pathogenicity prediction system 106 might observe 1 variant in that window and have a low, but non-zero OOER. Thus, in some embodiments, the pathogenicity prediction system 106 can use a pseudo count so that the OOER for a position is still scaled relative to expected counts.

[0197] To account for such circumstances, in some embodiments, the pathogenicity prediction system 106 replaces the observed number of variants with a pseudo count if the observed number of variants is 0. In other embodiments, a pseudo count is added to all observed number of variants, as shown in function (4) below, to ensure that the OOER does not become 0.

[0198] In function (4) above, Opsorepresents a pseudo count for the observed number of variants as described herein. A small (e.g., less than 1) value may be selected for the pseudo count. An optimal value for this pseudo count may depend on the size of the sequencing cohort in the sequencing data and the position window size used.

[0199] Because observed variants can come from sequencing data, various technical artefacts can influence the number of observed variants at a given protein position or position window. Such technical artifacts can include, for example, low read depth, mis-mapping reads, reference sequence or genome errors, or segmental duplications. To account for such technical artefacts, the pathogenicity prediction system 106 can scale metrics relevant to an OOER. Accordingly, as further shown in FIG. 9C, in some embodiments, the pathogenicity prediction system 106 uses scaling 930 to scale one or more of an OOER, observed number of variants 905, expected number of variants 910, or other metric generated based on the observed number and expected number of variants in the position window.

[0200] In some embodiments, for example, the pathogenicity prediction system 106 determines an observed / expected synonymous variants scaling factor 932. In some embodiments,Attorney Docket No. IP-2880-PCT 50 Patent Applicationthe pathogenicity prediction system 106 can use an observed / expected synonymous variants scaling factor 932 to correct for sequencing and other technical errors. For example, without being bound by theory, because most synonymous variants are expected to be neutral with regard to fitness, changes in the proportion of observed / expected synonymous variants are likely due to technical reasons. In some embodiments, the pathogenicity prediction system 106 determines a ratio of observed synonymous variants to expected synonymous variants, e.g., Osyn / Esyn. Regions where Osyn / Esyn is significantly below 1.0 may indicate that a sequencing system does not accurately call variants in that region, e.g., because of highly repetitive or mutable sequence elements or low sequencing depths (e.g., due to high GC content regions). Regions where Osyn / Esyn is significantly greater than 1.0 may indicate a region with many technical artefacts such as errors from mismapping of sequencing reads or sequencing contexts such as homopolymer runs that can induce errors in sequencing chemistry.

[0201] Accordingly, in some embodiments, the pathogenicity prediction system 106 can adjust the expected number of variants 910 to a scaled count of expected missense variants based on the observed synonymous variants and the expected synonymous variants, to correct for such technical deviations. For example, the pathogenicity prediction system 106 can multiply the expected number of variants (Emis) by the ratio of observed synonymous variants to expected synonymous variants (Osyn / Esyn) prior to determining an observed over expected missense variant ratio (OOERmis).

[0202] Relatively, in some embodiments, the pathogenicity prediction system 106 determines a sequencing depth scaling factor 935. For example, the pathogenicity prediction system 106 can adjust the observed number of variants or the expected number of variants based on sequencing depth. The pathogenicity prediction system 106 can detect regions with relatively low sequencing depth and correct potential biased observed variant counts due to low sequencing coverage. In some embodiments, the pathogenicity prediction system 106 bins regions based on median sequencing depth, determines low-depth regions with a median sequencing depth below a threshold, (e.g., median sequencing depth < 30), and corrects the expected number of missense variants (Emis) for the low-depth regions using a coefficient from a linear regression fit on Osyn / Esyn versus sequencing depth. In some instances, the pathogenicity prediction system 106 adjusts the expected number of missense variants prior to determining the OOER.

[0203] Independent of accounting for such artefacts, as indicated above, the pathogenicity prediction system 106 can determine observed and / or expected variant counts based on variant data from a plurality of species. In some embodiments, for instance, to determine an OOER based on variant data from both humans and non-human primate species, the pathogenicity prediction systemAttorney Docket No. IP-2880-PCT 51 Patent Application106 determines a mean of human OOER and a non-human primate OOER, as shown in function (5) below:

[0204] In function (5) above, OOERmeanrepresents a mean (average) OOER based on both human and non-human primate variant data. OOERhrepresents an OOER based on variant data from human variant data, while OOE7?prepresents an OOER based on variant data from non-human primate data. Ohrepresents an observed number of variants from human variant data, and Ehrepresents an expected number of variants from human variant data. Oprepresents an observed number of variants from non-human primate variant data, and Eprepresents an expected number of variants from non-human primate variant data. The variables Oh, Eh, Op, and Eprepresent the same components described here in function (6) below.

[0205] By contrast, in other embodiments, to determine an OOER based on variant data from both humans and non-human primate species, the pathogenicity prediction system 106 determines an OOER based on summed observed number of variants and summed expected number of variants from human and non-human primate variant data, as shown below:

[0206] In function (6) above, OOERsummedrepresents an OOER based on summing observed / expected variants for human and non-human primate variant data.

[0207] The sum-based approach has the advantage that the contributions of the human and non-human primate data are relative to the number of variants in each cohort. For example, in embodiments where the variant data is based on a human cohort that is much larger than the non-human primate cohort, then the Oh and E values will generally be higher and carry more weight than the Opand Epvalues. Accordingly, the sum-based approach may be preferred when cohorts are of very different sizes. The mean-based approach, in contrast, can weight non-human primate and human OOERs equally despite differences in the sizes of the cohorts. The mean-based approach may be advantageous, for example, in embodiments that use a smaller non-human primate cohort but where the non-human primate variants provide a better estimate of pathogenicity. A further approach could use a weighted average where the pathogenicity prediction system 106 weights each species’ OOER relative to an inferred importance. Both the mean-based and sum-based approaches improve the accuracy OOERs from individual species. Based on experimental data, the summed approach appears to give the best overall performance.

[0208] As indicated above, the pathogenicity prediction system 106 can integrate the interprotein pathogenicity scores (e.g., the inter-protein pathogenicity score(s) 920) and related systemsAttorney Docket No. IP-2880-PCT Patent Applicationand methods in combination with any of the other systems and methods described herein. For example, in certain embodiments, the pathogenicity prediction system 106 uses OOER and the OOER-based scores as embodiments of the inter-protein pathogenicity model 204 and the interprotein pathogenicity scores 206 described above. In some such cases, the pathogenicity prediction system 106 can access, for a set of amino acids in the target protein sequence, a set of inter-protein pathogenicity scores generated based on observed numbers of variants and expected numbers of variants within corresponding position windows. The pathogenicity prediction system 106 can determine an inter-protein central tendency of the set of inter-protein pathogenicity scores and an inter-protein standard deviation of the set of inter-protein pathogenicity scores. The pathogenicity prediction system 106 can access a set of intra-protein pathogenicity scores generated by an intraprotein pathogenicity model for the set of amino acids in the target protein sequence. Furthermore, the pathogenicity prediction system 106 can generate re-scaled pathogenicity scores for the set of amino acids according to a ranking of the set of intra-protein pathogenicity scores using the interprotein central tendency and the inter-protein standard deviation.

[0209] As discussed, the systems and methods described herein can use different types of position windows. FIG. 10A illustrates graphs showing experimental results regarding the effectiveness of the pathogenicity prediction system 106 generating inter-protein pathogenicity scores for target amino acids at target positions of a target protein sequence, based on different types of position windows, in accordance with one or more embodiments.

[0210] Graphs 1002, 1004, 1006 show DDD benchmark performance for observed-overexpected ratios (OOERs) calculated from either a set number of closest amino acids in sequence (graph 1002) or 3D space (graph 1004), or from amino acids within a fixed 3D distance from the scored position (graph 1006). Performance is measured by a Wilcoxon Rank Sum Test of OOER scores for all de novo missense mutations in DDD patients versus OOER scores for all de novo missense mutations in control individuals. As indicated above, DDD refers to a deciphering developmental disorders (DDD) study dataset. For instance, the DDD study was used with the PrimateAI model developed by Illumina, Inc. to determine how well the PrimateAI model could distinguish between variants in people with developmental disorders from variants in healthy people. Scores are plotted for OOERs generated using different window sizes, different minimum allele frequencies (Min AF), and using the three different types of position windows.

[0211] As shown in FIG. 10A, position windows based on 3D distances from the target position (graphs 1004 and 1006) exhibit better performance than position windows based on proximity in sequence space (graph 1002). Of the two types of windows based on 3D distances, the position windows based on a predetermined number of amino acids (graph 1004) performed better than position windows based on a fixed radius (graph 1006).Attorney Docket No. IP-2880-PCT 53 Patent Application

[0212] Furthermore, as discussed, the systems and methods described herein can use variant data from a plurality of species to determine the observed number of variants and / or the expected number of variants. FIG. 10B illustrates graphs showing experimental results regarding the effectiveness of the pathogenicity prediction system 106 generating inter-protein pathogenicity scores for amino acids at target positions of a target protein sequence, based on including variant data for one or more additional primate species, in accordance with one or more embodiments.

[0213] Graphs 1020 and 1022 illustrate the impact of combining human and primate OOERs on DDD performance. As before, performance is measured by a Wilcoxon Rank Sum Test of OOER scores for de novo mutations in the DDD dataset versus OOER scores for de novo mutations in control individuals in known dominant developmental disorder associated genes. Scores are plotted for OOERs generated using different window sizes, a fixed minimum allele frequency and primate variant mapping quality thresholds, and using two different methods for calculating the OOER with position windows calculated from either a set number of closest amino acids in sequence (graph 1020) or 3D space (graph 1022). For the inter-protein pathogenicity scores in FIG.10B, OOERs were determined in four ways: human variant data only (labeled “Human”), nonhuman primate variant data only (labeled “Primate”), the mean of the human and non-human primate OOERs (labeled “Mean Human & Primate”), or OOERs generated after summing observed and expected counts for both human and non-human primate variant data (labeled “Summed Human + Primate”).

[0214] As shown in FIG. 10B, OOERs based on combining human and non-human primate variant data performed (either mean or summed) better than human or non-human primate alone. OOERs based on summing observed and expected counts for both human and non-human primate variant data performed the best overall, for both ID position windows (graph 1020) and 3D position windows (graph 1022).

[0215] Moreover, as discussed, the disclosed methods and systems that generate an interprotein pathogenicity score specific to the target position based on observed number of variants and expected number of variants increase the accuracy of pathogenicity predictions for clinical benchmarks over other inter-protein pathogenicity score methods and systems. FIG. 11 illustrates an additional graph showing experimental results regarding the effectiveness of the pathogenicity prediction system 106 generating inter-protein pathogenicity scores for amino acids at target positions of a target protein sequence in accordance with one or more embodiments.

[0216] In particular, FIG. 11 compares the performance of one or more embodiments of the pathogenicity prediction system 106 that uses OOERs based on human variant data only (labeled “Human OOER”) and OOERs based on human and non-human primate variant data (labeled “Human + Primate OOER”) with two other state-of-the-art amino-acid variant pathogenicityAttorney Docket No. IP-2880-PCT 54 Patent Applicationprediction systems, including an AlphaMissense model (labeled “AlphaMissense”), and an existing deep learning model (labeled “Deep Learning Model”) that uses a transformer machine-learning model. The “Human OOER” and “Human + Primate OOER” performances in FIG. 11 are from OOER scores only, based on protein position, ignoring the reference and alternate amino acids (e.g., an alanine to leucine variant would have the same score as an alanine to glutamate variant at the same protein position). The “AlphaMissense” and “Deep Learning Model” use intra-protein information and will vary depending on the reference and alternative amino acids (e.g., an alanine to leucine variant would have a different score to an alanine to glutamate variant at the same protein position).

[0217] FIG. 11 compares OOER inter-protein pathogenicity scores generated by the models for the deciphering developmental disorders (DDD) dataset and for an autism spectrum disorders (ASD) dataset. For instance, the DDD and ASD studies were previously used with the PrimateAI model developed by Illumina, Inc. to determine how well the PrimateAI model could distinguish between variants in people with developmental disorders from variants in healthy people.

[0218] As shown in FIG. 11 , both Human OOER and Human + Primate OOER outperformed AlphaMissense and the Deep-Learning Model, and Human + Primate OOER performed the best overall. Indeed, the Human + Primate OOER approach produces probabilities of pathogenicity for target positions with more accuracy relative to the DDD and ASD datasets than the alternatives.

[0219] As discussed above, the pathogenicity prediction system 106 can generate a combined pathogenicity score based on intra-protein pathogenicity score(s) and inter-protein pathogenicity score(s). In particular, the pathogenicity prediction system 106 can utilize an inter-and-intra score combiner model to generate the combined pathogenicity score(s) for a set of amino acids included in the target protein sequence, with each combined pathogenicity score specific to a target variant at a target position in the target protein sequence. In accordance with one or more embodiments, FIG. 12A illustrates an overview of the pathogenicity prediction system generating a combined pathogenicity score for a target position based on intra-protein pathogenicity score(s) and interprotein pathogenicity score(s).

[0220] Indeed, as shown, the pathogenicity prediction system 106 accesses a set of intra-protein pathogenicity scores, such as intra-protein pathogenicity scores 1202. The intra-protein pathogenicity scores 1202 can be generated for a target protein sequence using any of the methods and systems described herein. For example, the intra-protein pathogenicity scores 1202 can be generated by an intra-protein pathogenicity model for the set of amino acids in the target protein sequence.

[0221] From the intra-protein pathogenicity scores 1202, the pathogenicity prediction system 106 can generate intra-protein-score features 1210. For example, the intra-protein-score featuresAttorney Docket No. IP-2880-PCT 55 Patent Application1210 can include a score for a variant at a target position 1212. For instance, the score for a variant at a target position 1212 can be an intra-protein pathogenicity score for having a particular variant at the target position, as determined using an intra-protein pathogenicity model.

[0222] In some embodiments, the score for a variant at a target position 1212 undergoes standardization 1214. For example, the pathogenicity prediction system 106 can determine a distribution of intra-protein pathogenicity scores, determine an intra-protein central tendency (e.g., mean) and an intra-protein standard deviation, and standardize the score for a variant at a target position 1212 by subtracting the intra-protein mean divided by the intra-protein standard deviation. In some cases, the distribution of intra-protein pathogenicity scores includes scores a from a subset of the positions within the target protein sequence, such as a batch (e.g., about half). In some cases, the distribution of intra-protein pathogenicity scores includes scores across all the positions within the target protein sequence.

[0223] In some embodiments, the pathogenicity prediction system 106 removes any interprotein signal (e.g., contribution of inter-protein scores) from the score for a variant at a target position 1212. For example, in some cases, the score for a variant at the target position 1212 is determined using an intra-protein pathogenicity model that takes into account an inter-protein score or other measurement of inter-protein pathogenicity. In some cases, the pathogenicity prediction system 106 removes any inter-protein signal (e.g., from data derived from non-human, non-primate species) from the score for a variant at the target position 1212. In some embodiments, the distribution of inter-protein scores accordingly is the same for each protein.

[0224] As illustrated in FIG. 12 A, in some embodiments, the pathogenicity prediction system 106 can further generate a set of intra-protein-score features, such as the intra-protein-score features 1210, to include a score rank 1216. For example, the score rank 1216 can be a value (e.g., from 0 to 1) indicating a ranking of an intra-protein pathogenicity score. For example, the score rank 1216 can be based on a ranking of an intra-protein pathogenicity score (e.g., a ranking of score for a variant at a target position 1212) among other variants at the target position and / or among other variants at other target positions in the target protein sequence.

[0225] As further illustrated in FIG. 12A, the pathogenicity prediction system 106 can further access a set of inter-protein pathogenicity scores, such as inter-protein pathogenicity scores 1204. The inter-protein pathogenicity scores 1204 can be generated for a target protein sequence using any of the methods and systems described herein. For example, the inter-protein pathogenicity scores 1204 can be generated by an inter-protein pathogenicity model for the set of amino acids in the target protein sequence. In some cases, the inter-protein pathogenicity scores 1204 are based on observed-over-expected ratios of the observed number of variants relative to a sum of theAttorney Docket No. IP-2880-PCT 56 Patent Applicationobserved number of variants and the expected number of variants (OOERs), as further described herein.

[0226] Using the inter-protein pathogenicity scores 1204, the pathogenicity prediction system 106 can generate inter-protein-score features 1222. For example, the inter-protein-score features 1222 can include distribution feature(s) 1224. The distribution feature(s) 1224 can include one or more metrics determined from a distribution of the inter-protein pathogenicity scores 1204, for example, over a window 1206, a batch 1208, and / or over the protein sequence 1209. Further details regarding the distribution feature(s) 1224 are described below with respect to FIG. 12B. The inter-protein-score features 1222 can also include, in some instances, a score at target position 1225 — for example, the inter-protein pathogenicity score for the target position.

[0227] As illustrated in FIG. 12A, the pathogenicity prediction system 106 can, in some cases, also determine mixed intra-inter-score features 1218. For example, the pathogenicity prediction system 106 can generate one or more features that are based on a combination of both intra-protein pathogenicity score(s) and inter-protein pathogenicity score(s). In some instances, the mixed intra-inter-score features 1218 include reranked OOER(s) 1220. For example, the reranked OOER(s) 1220 can include inter-protein pathogenicity score(s) (e.g., observed-over-expected ratio(s)) that have been reranked according to a ranking of the set of intra-protein pathogenicity scores. For example, the pathogenicity prediction system 106 can rerank inter-protein pathogenicity scores 1204 (e.g., OOERs) according to a pathogenicity ranking of the intra-protein pathogenicity scores 1202.

[0228] As further illustrated in FIG. 12A, the pathogenicity prediction system 106 generates an inter-and-intra score feature tensor 1230. The inter-and-intra score feature tensor 1230 includes pathogenicity score features based on the inter-protein pathogenicity scores 1204 and the intra-protein pathogenicity scores 1202. For example, the inter-and-intra score feature tensor 1230 can include two or more of: at least one inter-protein-score feature from the inter-protein-score features 1222, at least one intra-protein-score feature of the intra-protein-score features 1210, or at least one mixed inter-intra-score feature from the mixed intra-inter-score features 1218. In some embodiments, the inter-and-intra score feature tensor 1230 is a scalar, vector, or matrix. In some cases, the inter-and-intra score feature tensor 1230 is specific for a target variant at the target position.

[0229] As further illustrated in FIG. 12A, the pathogenicity prediction system 106 feeds the inter-and-intra score feature tensor 1230 to an inter-and-intra score combiner model 1240, which can process the inter-and-intra score feature tensor 1230. In some cases, the inter-and-intra score combiner model 1240 comprises a neural network or other machine-learning model. In some cases,Attorney Docket No. IP-2880-PCT 57 Patent Applicationthe inter-and-intra score combiner model 1240 comprises a multi-layer perceptron (MLP) having a plurality of convolutional layers that feed into a fully-connected layer.

[0230] In some embodiments, the inter-and-intra score combiner model 1240 is trained by adjusting model parameters based on separate comparisons with ground-truth inter-protein pathogenicity scores and with ground-truth intra-protein pathogenicity scores. Moreover, in some embodiments, the inter-and-intra score combiner model 1240 is trained separately with human variant data and non-human primate variant data.

[0231] Accordingly, in some embodiments, training the inter-and-intra score combiner model 1240 includes accessing intra-protein pathogenicity scores (and optionally also inter-protein pathogenicity scores) generated from human variant data, generating the inter-and-intra score feature tensor 1230 based on the scores from human variant data, and generating the combined pathogenicity score 1250 with the inter-and-intra score combiner model. The pathogenicity prediction system 106 can then adjust model parameters based on separately comparing the combined pathogenicity score 1250 to a ground-truth inter-protein pathogenicity score and to a ground-truth intra-protein pathogenicity score.

[0232] Furthermore, in some embodiments, training the inter-and-intra score combiner model 1240 includes accessing intra-protein pathogenicity scores (and optionally also inter-protein pathogenicity scores) generated from non-human primate variant data, generating the inter-and-intra score feature tensor 1230 based on the scores from non-human primate variant data, and generating the combined pathogenicity score 1250 with the inter-and-intra score combiner model 1240. The pathogenicity prediction system 106 can further adjust model parameters of the inter-and-intra score combiner model 1240 based on separately comparing the combined pathogenicity score 1250 to a ground-truth inter-protein pathogenicity score and to a ground-truth intra-protein pathogenicity score, such as by determining (i) an inter loss (e.g., a mean squared error, Spearman intra-protein correlation loss, or OOER positional correlation loss) from a comparison of the combined pathogenicity score 1250 and a ground-truth inter-protein pathogenicity score and (ii) an intra loss (e.g., a mean squared error, Spearman intra-protein correlation loss, or OOER positional correlation loss) from comparison of the combined pathogenicity score 1250 and a ground-truth intra-protein pathogenicity score.

[0233] In some embodiments, however, the inter-and-intra score combiner model 1240 is trained by adjusting model parameters based on a combined comparison of the combined pathogenicity score output by the inter-and-intra score combiner model 1240 with ground-truth inter-protein pathogenicity scores and ground-truth intra-protein pathogenicity scores. For instance, the pathogenicity prediction system 106 determines a combined inter loss and intra loss (e.g., an average, weighted average of the inter loss and intra loss above). Moreover, in someAttorney Docket No. IP-2880-PCT 58 Patent Applicationembodiments, the inter- and-intra score combiner model 1240 is trained together with human variant data and non-human primate variant data. For example, in some cases, the inter-and-intra score combiner model 1240 is trained based on comparisons of the combined pathogenicity score output by the inter-and-intra score combiner model 1240 with ground-truth inter-protein pathogenicity scores and ground-truth intra-protein pathogenicity scores (including both human variant data and non-human primate variant data) together and is merged into the same loss. By adjusting model parameters from comparing the combined pathogenicity score to both groundtruth inter-protein and ground-truth intra-protein scores — where such ground-truth scores or values are derived from both human variant data and non-human primate variant data — the pathogenicity prediction system 106 can save memory and reduce training iterations. In the alternative to separately training one or more inter-and-intra score combiner models using ground-truth interprotein and ground-truth intra-protein scores derived from human variant data, for one set of training epochs, and non-human variant data, for another set of training epochs, the combined comparison described above can consolidate training using the combined comparison. Such a combined comparison saves memory and time that a computing device would otherwise consume to train one or more inter-and-intra score combiner models on ground-truth scores derived from human variant data and, separately, on ground-truth scores derived from non-human variant data.

[0234] In some cases, the pathogenicity prediction system 106 adjusts model parameters of the inter-and-intra score combiner model based on comparing (e.g., determining a loss) the combined pathogenicity score to one or more ground-truth human and / or ground-truth primate variants comprising observed and trinucleotide-matched unknown variants. Methods for determining a loss based on observed and trinucleotide-matched unknown variants are further described in H. Gao et al., “The Landscape of Tolerated Genetic Variation in Humans and Primates,” Science 380, eabn8153 (2023), which is hereby incorporated by reference in its entirety.

[0235] Returning to FIG. 12A, the pathogenicity prediction system 106 can utilize the inter-and-intra score combiner model 1240 to process the inter-and-intra score feature tensor 1230 and to generate, as output, a combined pathogenicity score 1250. In one or more embodiments, the combined pathogenicity score 1250 is specific for the target variant at the target position. For example, the combined pathogenicity score 1250 can indicate the combined (e.g., inter-and-intra) pathogenicity of having a particular variant at the target position.

[0236] By combining intra-protein and inter-protein pathogenicity scores using the inter-and-intra score combiner model 1240, the pathogenicity prediction system 106 can generate a more accurate pathogenicity prediction for having a particular variant at a target position. FIG. 13, described below, further illustrates such improvements in accuracy.Attorney Docket No. IP-2880-PCT 59 Patent Application

[0237] As mentioned, the pathogenicity prediction system 106 can generate inter-protein distribution features based on a window distribution, a batch distribution, and / or a protein sequence distribution of inter-protein pathogenicity scores. FIG. 12B illustrates generating features related to a distribution of inter-protein pathogenicity scores in accordance with one or more embodiments.

[0238] As illustrated, the pathogenicity prediction system 106 accesses the inter-protein pathogenicity scores 1204. In some cases, the inter-protein pathogenicity scores 1204 include interprotein pathogenicity scores for each position in the protein sequence. In some embodiments, the inter-protein pathogenicity scores 1204 are based on observed-over-expected ratios (OOERs) of the observed number of variants relative to a sum of the observed number of variants and the expected number of variants. The OOERs can be generated according to the methods and systems further described herein, for example, based on an observed number of variants and an expected number of variants within a position window surrounding the target position.

[0239] As further illustrated in FIG. 12B, the pathogenicity prediction system 106 can determine a per- window distribution 1262 from scores in a window 1206 of amino acid positions around the target position. The window 1206 can include a subset of the positions in the protein sequence, wherein the subset is significantly less than half of the positions in the protein sequence. The window 1206 can surround the target position (e.g., such that the target position is located in the middle of the window 1206).

[0240] Based on the per- window distribution 1262, the pathogenicity prediction system 106 can generate per- window distribution features 1272. For example, the per- window distribution features 1272 can include one or more metrics that describe the per- window distribution 1262 of inter-protein pathogenicity scores (e.g., OOERs). For example, the per-window distribution features 1272 can include central tendenc(ies) 1273a, e.g. a mean, median, and / or mode of the per-window distribution 1262. As further example(s), the per-window distribution features 1272 can include a minimum 1273b, a maximum 1273c, a mean absolute deviation 1273d, or a standard deviation (not illustrated) of the per-window distribution 1262. Additionally, the per-window distribution features 1272 can include percentile(s) 1273f of the per-window distribution 1262 (e.g., the inter-protein pathogenicity score at a given percentile), for example, quantile(s) 1273e of the per-window distribution 1262.

[0241] As further illustrated in FIG. 12B, the pathogenicity prediction system 106 can determine a per-batch distribution 1264 of inter-protein pathogenicity scores in a batch 1208 of amino acid positions in the target protein sequence. The batch 1208 can include a subset of the positions in the protein sequence, wherein the subset is about half of the positions in the protein sequence. The batch 1208 can include consecutive amino acid positions in the protein sequence. As another example, the batch 1208 can include non-consecutive amino acid positions in theAttorney Docket No. IP-2880-PCT 60 Patent Applicationprotein sequence (e.g., sampled across the protein sequence). The batch 1208 can include the target position and does not necessarily have to center on the target position. For example, the pathogenicity prediction system 106 can select the batch 1208 based on three-dimensional proximity to the target position (e.g., within a three-dimensional structure window).

[0242] The pathogenicity prediction system 106 can further generate per-batch distribution features 1274 based on the per-batch distribution 1264. For instance, the per-batch distribution features 1274 can include one or more metrics that describe the per-batch distribution 1264 of interprotein pathogenicity scores (e.g., OOERs). To enumerate examples, the per-batch distribution features 1274 can include central tendenc(ies) 1275a, e.g. a mean, median, and / or mode of the per-batch distribution 1264. As further example(s), the per-batch distribution features 1274 can include a minimum 1275b, a maximum 1275c, a mean absolute deviation 1275d, or a standard deviation (not illustrated) of the per-batch distribution 1264. In some cases, the per-batch distribution features 1274 include percentile(s) 1275f of the per-batch distribution 1264 (e.g., the OOER at a given percentile), for example, quantile(s) 1275e of the per-batch distribution 1264.

[0243] Additionally, as illustrated, the pathogenicity prediction system 106 can determine a per-protein distribution 1266 of inter-protein pathogenicity scores over the protein sequence 1209. The protein sequence 1209 can include, for example, all amino acid positions in the target protein sequence. Accordingly, the per-protein distribution 1266 can be a distribution of inter-protein pathogenicity scores (e.g., OOERs) over the protein sequence.

[0244] Using the per-protein distribution 1266, the pathogenicity prediction system 106 can generate per-protein distribution features 1276. For instance, the per-protein distribution features 1276 can include one or more metrics that describe the per-protein distribution 1266 of inter-protein pathogenicity scores (e.g., OOERs). The per-protein distribution features 1276 can include central tendenc(ies) 1278a, such as a mean, median, and / or mode OOER of the per-protein distribution 1266. The per-protein distribution features 1276 can include a minimum 1278b, a maximum 1278c, a mean absolute deviation 1278d, or a standard deviation (not illustrated) of the per-protein distribution 1266. The per-protein distribution features 1276 can include percentile(s) 1278f of the per-protein distribution 1266, for example, quantile(s) 1278e of the per-protein distribution 1266 (e.g., the inter-protein pathogenicity score at a given quantile of the per-protein distribution 1266).

[0245] By generating the inter-protein pathogenicity score distribution features for inclusion in the inter- and-intra score feature tensor 1230, the pathogenicity prediction system 106 can more accurately predict overall pathogenicity for a given variant at a target position. Moreover, by generating per- window distribution features 1272, per-batch distribution features 1274, and per-protein distribution features 1276, the pathogenicity prediction system 106 can capture the variationAttorney Docket No. IP-2880-PCT 61 Patent Applicationin OOERs in different contexts of the target position and use this information to further improve accuracy of generating the combined pathogenicity score 1250.

[0246] As mentioned, the use of an inter-and-intra score combiner model improves pathogenicity prediction accuracy. FIG. 13 illustrates a graph showing experimental results regarding the effectiveness of the pathogenicity prediction system generating combined pathogenicity scores in accordance with one or more embodiments.

[0247] In particular, the graphs of FIG. 13 compares the performance of one or more embodiments of the pathogenicity prediction system 106 (labeled “Inter- And-Intra Combiner Model PrimateAI-3D (1)” and Inter-And-Intra Combiner Model PrimateAI-3D (2)”) with several existing state-of-the-art amino-acid variant pathogenicity prediction systems, including a PrimateAI-3D Published model (e.g., a previously-published version of PrimateAI 3D), an AlphaMissense model, an external protein assembly interaction (PAI) model with shift-rescale using OOER scores (labeled “External SR With PAI”), and an embodiment of the pathogenicity prediction system 106 that uses shift-rescale instead of a combiner model (labeled “PrimateAI-SR”). The “Inter-And-Intra Combiner Model PrimateAI-3D (1)” and Inter-And-Intra Combiner Model PrimateAI-3D (2)” differed slightly from each other with respect to network size and loss weights.

[0248] The various pathogenicity prediction methods were tested for different data sets, including a deciphering developmental disorders (DDD) study dataset, an autism spectrum disorders (ASD) dataset, and a congenital heart disease cohort (CHD). The different methods were ranked by performance in each dataset, with the high rank value being given to the best method. The arithmetic mean of the ranks (“Mean rank”) was determined for each tested method’s performance across the different datasets. Graph 1310 shows mean rank for each tested method across the 2023 DDD, ASD, and CHD datasets, and graph 1320 shows the mean rank for each tested method across the 2024 DDD, ASD, and CHD datasets.

[0249] As shown in graph 1310 and graph 1320, the embodiment(s) of the pathogenicity prediction system 106 using an inter-and-intra combiner model consistently outperform the other tested models across the different data sets.

[0250] FIGS. 1-13, the corresponding text and the examples provide a number of different methods, systems, devices, and non-transitory computer-readable media of the pathogenicity prediction system 106. In addition to the foregoing, one or more embodiments can also be described in terms of flowcharts comprising acts for accomplishing particular results, as shown in FIGS. 14-16. FIGS. 14-16 may be performed with more or fewer acts. Further, the acts may be performed in different orders. Additionally, the acts described herein may be repeated or performed in parallel with one another or in parallel with different instances of the same or similar acts.Attorney Docket No. IP-2880-PCT 62 Patent Application

[0251] FIG. 14 illustrates a flowchart of a series of acts 1400 for determining pathogenicity predictions for a set of amino acids using inter-protein pathogenicity scores and intra-protein pathogenicity scores in accordance with one or more embodiments. While FIG. 14 illustrates acts according to one embodiment, alternative embodiments may omit, add to, reorder, and / or modify any of the acts shown in FIG. 14. In some implementations, the acts of FIG. 14 are performed as part of a method, such as a computer-implemented method. In some instances, a non-transitory computer-readable medium stores instructions thereon that, when executed by at least one processor, cause a computing device to perform the acts of FIG. 14. In some implementations, a system performs the acts of FIG. 14. For example, in one or more cases, a system includes at least one processor and a non-transitory computer-readable medium comprising instructions that, when executed by the at least one processor, cause the system to perform the acts of FIG. 14. Various types of processors can be used to perform the acts in various implementations, including, but not limited to, central processing units (CPUs), graphics processing units (GPUs), neural processing units (NPUs), field-programmable gate arrays (FPGAs), or application-specific integrated circuits (ASICs).

[0252] As shown in FIG. 14, the series of acts 1400 includes an act 1402 for accessing a set of inter-protein pathogenicity scores for a set of amino acids in a target protein sequence; an act 1404 for determining an inter-protein central tendency and an inter-protein standard deviation of the set of inter-protein pathogenicity scores; an act 1406 for accessing a set of intra-protein pathogenicity scores for the set of amino acids; and an act 1408 for generating re-scaled pathogenicity scores from the set of intra-protein pathogenicity scores, the inter-protein central tendency, and the interprotein standard deviation. For example, the series of acts 1400 can include acts to perform any of the operations described in the following clauses:CLAUSE 1. A computer-implemented method comprising:accessing a set of inter-protein pathogenicity scores generated by an inter-protein pathogenicity model for a set of amino acids in a target protein sequence;determining an inter-protein central tendency of the set of inter-protein pathogenicity scores and an inter-protein standard deviation of the set of inter-protein pathogenicity scores; accessing a set of intra-protein pathogenicity scores generated by an intra-protein pathogenicity model for the set of amino acids in the target protein sequence; andgenerating re-scaled pathogenicity scores for the set of amino acids according to a ranking of the set of intra-protein pathogenicity scores using the inter-protein central tendency and the interprotein standard deviation.Attorney Docket No. IP-2880-PCT 63 Patent ApplicationCLAUSE 2. The computer-implemented method of clause 1, wherein accessing the set of inter-protein pathogenicity scores comprises accessing inter-protein pathogenicity scores that indicate depletion of observed variants across a genomic region.CLAUSE 3. The computer-implemented method of clause 1 or 2, wherein accessing the set of intra-protein pathogenicity scores comprises accessing intra-protein pathogenicity scores having values that indicate a pathogenicity ranking of the set of amino acids.CLAUSE 4. The computer-implemented method of any of clauses 1-3, further comprising:determining an intra-protein central tendency of the set of intra-protein pathogenicity scores and an intra-protein standard deviation of the set of intra-protein pathogenicity scores; and generating a set of standardized intra-protein pathogenicity scores from the set of intra-protein pathogenicity scores using the intra-protein central tendency and the intra-protein standard deviation,wherein generating the re-scaled pathogenicity scores for the set of amino acids comprises generating the re-scaled pathogenicity scores from the set of standardized intra-protein pathogenicity scores using the inter-protein central tendency and the inter-protein standard deviation.CLAUSE 5. The computer-implemented method of any of clauses 1-4, wherein generating the re-scaled pathogenicity scores for the set of amino acids comprises:generating a set of standardized intra-protein pathogenicity scores from the set of intra-protein pathogenicity scores using a first set of transformations; andgenerating the re-scaled pathogenicity scores from the set of standardized intra-protein pathogenicity scores using a second set of transformations that is an inverse of the first set of transformations.CLAUSE 6. The computer-implemented method of any of clauses 1-5, wherein: accessing the set of intra-protein pathogenicity scores generated by the intra-protein pathogenicity model for the set of amino acids in the target protein sequence comprises accessing a plurality of sets of intra-protein pathogenicity scores generated by a plurality of intra-protein pathogenicity models for the set of amino acids in the target protein sequence; and generating the re-scaled pathogenicity scores for the set of amino acids using the interprotein central tendency and the inter-protein standard deviation comprises generating a plurality of sets of re-scaled pathogenicity scores from the plurality of sets of intra-protein pathogenicity scores using the inter-protein central tendency and the inter-protein standard deviation; and further comprising combining the plurality of sets of re-scaled pathogenicity scores to generate a combined set of re-scaled pathogenicity scores for the set of amino acids.Attorney Docket No. IP-2880-PCT 64 Patent ApplicationCLAUSE 7. The computer-implemented method of any of clauses 1-6, wherein generating the re-scaled pathogenicity scores for the set of amino acids in the target protein sequence comprises generating the re-scaled pathogenicity scores for a set of at least nineteen alternative amino acids at a protein position of the target protein sequence that differ from a reference amino acid at the protein position.CLAUSE 8. The computer-implemented method of clause 7, wherein the re-scaled pathogenicity scores for the set of at least nineteen alternative amino acids comprise a set of rescaled pathogenicity scores for a set of missense nucleotide variants.CLAUSE 9. The computer-implemented method of any of clauses 1-8, wherein determining the inter-protein central tendency of the set of inter-protein pathogenicity scores comprises determining an inter-protein mean, an inter-protein median, or an inter-protein mode of the set of inter-protein pathogenicity scores.CLAUSE 10. The computer-implemented method of any of clauses 1-9, wherein accessing the set of inter-protein pathogenicity scores generated by the inter-protein pathogenicity model comprises generating the set of inter-protein pathogenicity scores using the inter-protein pathogenicity model.CLAUSE 11. The computer-implemented method of any of clauses 1-10, wherein accessing the set of intra-protein pathogenicity scores generated by the intra-protein pathogenicity model comprises generating the set of intra-protein pathogenicity scores using the intra-protein pathogenicity model.CLAUSE 12. The computer-implemented method of any of clauses 1-11,wherein accessing the set of intra-protein pathogenicity scores generated by the intra-protein pathogenicity model for the set of amino acids in the target protein sequence comprises accessing a plurality of sets of intra-protein pathogenicity scores generated by a plurality of intra-protein pathogenicity models for the set of amino acids in the target protein sequence;further comprising combining the plurality of sets of intra-protein pathogenicity scores to generate a combined set of intra-protein pathogenicity scores; andwherein generating the re-scaled pathogenicity scores for the set of amino acids using the inter-protein central tendency and the inter-protein standard deviation comprises generating the rescaled pathogenicity scores from the combined set of intra-protein pathogenicity scores using the inter-protein central tendency and the inter-protein standard deviation.CLAUSE 13. The computer-implemented method of clause 12, wherein combining the plurality of sets of intra-protein pathogenicity scores to generate the combined set of intra-protein pathogenicity scores comprises generating, using forward feature selection, the combined set ofAttorney Docket No. IP-2880-PCT 65 Patent Applicationintra-protein pathogenicity scores from a subset of intra-protein pathogenicity scores from the plurality of sets of intra-protein pathogenicity scores.CLAUSE 14. The computer-implemented method of any of clauses 1-13, wherein accessing the set of inter-protein pathogenicity scores generated by the inter-protein pathogenicity model for the set of amino acids in the target protein sequence comprises accessing inter-protein pathogenicity scores generated by the inter-protein pathogenicity model for amino acids within an amino acid window of the target protein sequence.

[0253] FIG. 15 illustrates a flowchart of a series of acts 1500 for determining pathogenicity predictions for a target amino acid position using inter-protein pathogenicity scores in accordance with one or more embodiments. While FIG. 15 illustrates acts according to one embodiment, alternative embodiments may omit, add to, reorder, and / or modify any of the acts shown in FIG.15. In some implementations, the acts of FIG. 15 are performed as part of a method, such as a computer-implemented method. In some instances, a non-transitory computer-readable medium stores instructions thereon that, when executed by at least one processor, cause a computing device to perform the acts of FIG. 15. In some implementations, a system performs the acts of FIG. 15. For example, in one or more cases, a system includes at least one processor and a non-transitory computer-readable medium comprising instructions that, when executed by the at least one processor, cause the system to perform the acts of FIG. 15. Various types of processors can be used to perform the acts in various implementations, including, but not limited to, central processing units (CPUs), graphics processing units (GPUs), neural processing units (NPUs), field-programmable gate arrays (FPGAs), or application-specific integrated circuits (ASICs).

[0254] As shown in FIG. 15, the series of acts 1500 includes an act 1502 for identifying, for a target position of a target amino acid within a target protein sequence, proximate positions for contextual amino acids within a position window of the target position; an act 1504 for determining, for the target position and the proximate positions, an observed number of variants within the position window; an act 1506 for determining, for the target position and the proximate positions, an expected number of variants within the position window; and an act 1508 for generating an inter-protein pathogenicity score specific to the target position based on the observed number of variants and the expected number of variants. For example, the series of acts 1500 can include acts to perform any of the operations described in the following clauses:CLAUSE 15. A computer-implemented method comprising:identifying, for a target position of a target amino acid within a target protein sequence, proximate positions for contextual amino acids within a position window of the target position;determining, for the target position and the proximate positions, an observed number of variants within the position window;Attorney Docket No. IP-2880-PCT 66 Patent Applicationdetermining, for the target position and the proximate positions, an expected number of variants within the position window; andgenerating an inter-protein pathogenicity score specific to the target position based on the observed number of variants and the expected number of variants.CLAUSE 16. The computer-implemented method of clause 15, wherein identifying the proximate positions within the position window comprises identifying the proximate positions for the contextual amino acids based on a three-dimensional distance of the target position in a three-dimensional structure of the target protein sequence.CLAUSE 17. The computer-implemented method of clause 15 or 16, wherein identifying the proximate positions comprises identifying a predetermined number of contextual amino acids ranked by three-dimensional proximity to the target position in a three-dimensional structure of the target protein sequence.CLAUSE 18. The computer-implemented method of any of clauses 15-17, wherein generating the inter-protein pathogenicity score is based on:determining a sum of the observed number of variants and the expected number of variants; anddetermining an observed-over-expected ratio of the observed number of variants relative to the sum of the observed number of variants and the expected number of variants.CLAUSE 19. The computer-implemented method of any of clauses 15-18, comprising determining the observed number of variants and the expected number of variants based on variant data from a plurality of species.CLAUSE 20. The computer-implemented method of any of clauses 15-19, wherein determining the observed number of variants and the expected number of variants is based on variant data from humans and at least one additional primate species.CLAUSE 21. The computer-implemented method of any of clauses 15-20, further comprising:selecting, from human sequences in variant data, human missense variants satisfying an allele frequency threshold; anddetermining one or more of the observed number of variants or the expected number of variants based on the selected human missense variants.CLAUSE 22. The computer-implemented method of any of clauses 15-21, further comprising:determining the observed number of variants and the expected number of variants based on a count of observed missense variants and a count of expected missense variants within the position window; andAttorney Docket No. IP-2880-PCT 67 Patent Applicationadjusting the count of expected missense variants to a scaled count of expected missense variants based on both a count of observed synonymous variants and a count of expected synonymous variants.CLAUSE 23. The computer-implemented method of any of clauses 15-22, further comprising:accessing a gene loss of function (LoF) constraint value for the target protein sequence; determining whether the gene LoF constraint value is above or below a LoF constraint threshold; andselecting a position window size or allele frequency threshold based on the determination of whether the gene LoF constraint value is above or below the LoF constraint threshold.CLAUSE 24. The computer-implemented method of any of clauses 15-23, further comprising utilizing a sliding position window by:determining a subsequent position window for a subsequent target position of a subsequent target amino acid within the target protein sequence;determining, for the subsequent target position and corresponding proximate positions, an observed number of variants and an expected number of variants within the subsequent position window; andgenerating a corresponding inter-protein pathogenicity score specific to the subsequent target position based on the observed number of variants and the expected number of variants within the subsequent position window.CLAUSE 25. The computer-implemented method of any of clauses 15-24, further comprising:determining that the observed number of variants is 0 within the position window; accessing, based on the observed number of variants being 0, a pseudo count of observed variants; andgenerating the inter-protein pathogenicity score specific to the target position based on the pseudo count of observed variants and the expected number of variants.CLAUSE 26. The computer-implemented method of any of clauses 15-25, further comprising determining the observed number of variants or the expected number of variants by:determining proximate-position weights for the proximate positions based on three-dimensional or one-dimensional proximity of each position relative to the target position within the position window; anddetermining one or more of the observed number of variants or the expected number of variants according to the proximate-position weights.Attorney Docket No. IP-2880-PCT 68 Patent ApplicationCLAUSE 27. The computer-implemented method of any of clauses 15-26, further comprising:accessing, for a set of amino acids in the target protein sequence, a set of inter-protein pathogenicity scores generated based on observed numbers of variants and expected numbers of variants within corresponding position windows;determining an inter-protein central tendency of the set of inter-protein pathogenicity scores and an inter-protein standard deviation of the set of inter-protein pathogenicity scores; accessing a set of intra-protein pathogenicity scores generated by an intra-protein pathogenicity model for the set of amino acids in the target protein sequence; andgenerating re-scaled pathogenicity scores for the set of amino acids according to a ranking of the set of intra-protein pathogenicity scores using the inter-protein central tendency and the interprotein standard deviation.CLAUSE 28. The computer-implemented method of any of clauses 15-27, further comprising adjusting the observed number of variants or the expected number of variants based on sequencing depth.

[0255] FIG. 16 illustrates a flowchart of a series of acts 1600 for determining pathogenicity predictions for a target variant at a target amino acid position using a combined pathogenicity score in accordance with one or more embodiments. While FIG. 16 illustrates acts according to one embodiment, alternative embodiments may omit, add to, reorder, and / or modify any of the acts shown in FIG. 16. In some implementations, the acts of FIG. 16 are performed as part of a method, such as a computer-implemented method. In some instances, a non-transitory computer-readable medium stores instructions thereon that, when executed by at least one processor, cause a computing device to perform the acts of FIG. 16. In some implementations, a system performs the acts of FIG. 16. For example, in one or more cases, a system includes at least one processor and a non-transitory computer-readable medium comprising instructions that, when executed by the at least one processor, cause the system to perform the acts of FIG. 16. Various types of processors can be used to perform the acts in various implementations, including, but not limited to, central processing units (CPUs), graphics processing units (GPUs), neural processing units (NPUs), field-programmable gate arrays (FPGAs), or application-specific integrated circuits (ASICs).

[0256] As shown in FIG. 16, the series of acts 1600 includes an act 1602 for accessing a set of inter-protein pathogenicity scores; an act 1604 for accessing a set of intra-protein pathogenicity scores; an act 1606 for generating an inter-and-intra score feature tensor for a target variant at a target position; and an act 1608 for generating a combined pathogenicity score utilizing an inter-and-intra score combiner model. For example, the series of acts 1600 can include acts to perform any of the operations described in the following clauses:Attorney Docket No. IP-2880-PCT 69 Patent ApplicationCLAUSE 29. A computer-implemented method comprising:accessing, for a set of amino acids in a protein sequence, a set of inter-protein pathogenicity scores generated from observed numbers of variants and expected numbers of variants within corresponding position windows;accessing, for the set of amino acids in the protein sequence, a set of intra-protein pathogenicity scores generated by an intra-protein pathogenicity model;generating, for a target variant at a target position with the protein sequence, an inter-and-intra score feature tensor comprising pathogenicity score features based on the set of inter-protein pathogenicity scores and the set of intra-protein pathogenicity scores; andgenerating, utilizing an inter-and-intra score combiner model to process the inter-and-intra score feature tensor, a combined pathogenicity score for the target variant at the target position.CLAUSE 30. The computer-implemented method of clause 29, wherein generating the inter-and-intra score feature tensor comprising pathogenicity score features comprises combining two or more of inter-protein-score features, intra-protein-score features, or mixed inter-intra-score features.CLAUSE 31. The computer-implemented method of clause 29 or 30, wherein the pathogenicity score features comprise one or more inter-protein-score features comprising:an inter-protein central tendency of a distribution of observed-over-expected ratios of the observed number of variants relative to a sum of the observed number of variants and the expected number of variants (distribution of OOERs), wherein the distribution of OOERs is over a window of positions of the protein sequence, over a batch of positions of the protein sequence, or over the protein sequence;an inter-protein minimum of the distribution of OOERs;an inter-protein maximum of the distribution of OOERs;an inter-protein mean absolute deviation of the distribution of OOERs;an inter-protein percentile of the distribution of OOERs; oran observed-over-expected ratio of the observed number of variants relative to a sum of the observed number of variants and the expected number of variants (OOER) for the target position.CLAUSE 32. The computer-implemented method of any of clauses 29-31, wherein the pathogenicity score features comprise one or more intra-protein-score features comprising: i) a standardized intra-protein pathogenicity score using an intra-protein central tendency and intra-protein standard deviation of a distribution of intra-protein scores, wherein the distribution is over a batch of positions of the protein sequence or over the protein sequence, or ii) an intra-protein pathogenicity score ranking.CLAUSE 33. The computer-implemented method of any of clauses 29-32, wherein theAttorney Docket No. IP-2880-PCT 70 Patent Applicationpathogenicity score features comprise one or more mixed inter-intra-score features comprising, for the target position, an observed-over-expected ratio of the observed number of variants relative to a sum of the observed number of variants and the expected number of variants (OOER) reranked according to a ranking of the set of intra-protein pathogenicity scores.CLAUSE 34. The computer-implemented method of any of clauses 29-33, further comprising:accessing the set of intra-protein pathogenicity scores generated by the intra-protein pathogenicity model using human variant data;generating the inter-and-intra score feature tensor comprising pathogenicity score features based on the set of intra-protein pathogenicity scores generated using the human variant data; generating, utilizing the inter-and-intra score combiner model to process the inter-and-intra score feature tensor, the combined pathogenicity score; andadjusting model parameters of the inter-and-intra score combiner model based on (i) comparing the combined pathogenicity score to a ground-truth inter-protein pathogenicity score and (ii) comparing the combined pathogenicity score to a ground-truth intra-protein pathogenicity score.CLAUSE 35. The computer-implemented method of any of clauses 29-34, further comprising:accessing the set of intra-protein pathogenicity scores generated by the intra-protein pathogenicity model using non-human primate variant data;generating the inter-and-intra score feature tensor comprising pathogenicity score features based on the set of intra-protein pathogenicity scores generated using the non-human primate variant data;generating, utilizing the inter-and-intra score combiner model to process the inter-and-intra score feature tensor, the combined pathogenicity score; andadjusting model parameters of the inter-and-intra score combiner model based on (i) comparing the combined pathogenicity score to a ground-truth inter-protein pathogenicity score and (ii) comparing the combined pathogenicity score to a ground-truth intra-protein pathogenicity score.

[0257] CLAUSE 36. The computer-implemented method of any of clauses 29-35, wherein the inter-and-intra score combiner model comprises a neural network or other machinelearning model.

[0258] CLAUSE 37. The computer-implemented method of any of clauses 29-36, further comprising adjusting model parameters of the inter-and-intra score combiner model based onAttorney Docket No. IP-2880-PCT 71 Patent Applicationcomparing the combined pathogenicity score to one or more ground-truth human and / or primate variants comprising observed and trinucleotide-matched unknown variants.

[0259] The components of the pathogenicity prediction system 106 can include software, hardware, or both. For example, the components of the pathogenicity prediction system 106 can include one or more instructions stored on a computer-readable storage medium and executable by processors of one or more computing devices (e.g., the server device(s) 102, the client device 110, and / or the therapeutics analysis device(s) 114). When executed by the one or more processors, the computer-executable instructions of the pathogenicity prediction system 106 can cause the computing devices to perform the pathogenicity prediction methods described herein. Alternatively, the components of the pathogenicity prediction system 106 can comprise hardware, such as special purpose processing devices to perform a certain function or group of functions. Additionally, or alternatively, the components of the pathogenicity prediction system 106 can include a combination of computer-executable instructions and hardware.

[0260] Furthermore, the components of the pathogenicity prediction system 106 performing the functions described herein with respect to the pathogenicity prediction system 106 may, for example, be implemented as part of a stand-alone application, as a module of an application, as a plug-in for applications, as a library function or functions that may be called by other applications, and / or as a cloud-computing model. Thus, components of the pathogenicity prediction system 106 may be implemented as part of a stand-alone application on a personal computing device or a mobile device. Additionally, or alternatively, the components of the pathogenicity prediction system 106 may be implemented in any application that provides sequencing services including, but not limited to, Illumina PrimateAI, Illumina PrimateAI-2D, Illumina PrimateAI-3D, or Illumina TruSight software. “Illumina,” “PrimateAI,” “PrimateAI-2D,” “PrimateAI-3D,” and “TruSight,” are either registered trademarks or trademarks of Illumina, Inc. in the United States and / or other countries.

[0261] Embodiments of the present disclosure may comprise or utilize a special purpose or general-purpose computer including computer hardware, such as, for example, one or more processors and system memory, as discussed in greater detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more of the processes described herein may be implemented at least in part as instructions embodied in a non-transitory computer-readable medium and executable by one or more computing devices (e.g., any of the media content access devices described herein). In general, a processor (e.g., a microprocessor) receives instructions, from a non-transitory computer-readable medium,Attorney Docket No. IP-2880-PCT 72 Patent Application(e.g., a memory, etc.), and executes those instructions, thereby performing one or more processes, including one or more of the processes described herein.

[0262] Computer-readable media can be any available media that can be accessed by a general purpose or special purpose computer system. Computer-readable media that store computerexecutable instructions are non-transitory computer-readable storage media (devices). Computer-readable media that carry computer-executable instructions are transmission media. Thus, by way of example, and not limitation, embodiments of the disclosure can comprise at least two distinctly different kinds of computer-readable media: non-transitory computer-readable storage media (devices) and transmission media.

[0263] Non-transitory computer-readable storage media (devices) includes RAM, ROM, EEPROM, CD-ROM, solid state drives (SSDs) (e.g., based on RAM), Flash memory, phasechange memory (PCM), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer.

[0264] A “network” is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided over a network or another communications connection (either hardwired, wireless, or a combination of hardwired or wireless) to a computer, the computer properly views the connection as a transmission medium. Transmissions media can include a network and / or data links which can be used to carry desired program code means in the form of computer-executable instructions or data structures and which can be accessed by a general purpose or special purpose computer. Combinations of the above should also be included within the scope of computer-readable media.

[0265] Further, upon reaching various computer system components, program code means in the form of computer-executable instructions or data structures can be transferred automatically from transmission media to non-transitory computer-readable storage media (devices) (or vice versa). For example, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface module (e.g., a NIC), and then eventually transferred to computer system RAM and / or to less volatile computer storage media (devices) at a computer system. Thus, it should be understood that non-transitory computer-readable storage media (devices) can be included in computer system components that also (or even primarily) utilize transmission media.

[0266] Computer-executable instructions comprise, for example, instructions and data which, when executed at a processor, cause a general-purpose computer, special purpose computer, orAttorney Docket No. IP-2880-PCT 73 Patent Applicationspecial purpose processing device to perform a certain function or group of functions. In some embodiments, computer-executable instructions are executed on a general-purpose computer to turn the general-purpose computer into a special purpose computer implementing elements of the disclosure. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or even source code. Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the described features or acts described above. Rather, the described features and acts are disclosed as example forms of implementing the claims.

[0267] Those skilled in the art will appreciate that the disclosure may be practiced in network computing environments with many types of computer system configurations, including, personal computers, desktop computers, laptop computers, message processors, hand-held devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile telephones, PDAs, tablets, pagers, routers, switches, and the like. The disclosure may also be practiced in distributed system environments where local and remote computer systems, which are linked (either by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) through a network, both perform tasks. In a distributed system environment, program modules may be located in both local and remote memory storage devices.

[0268] Embodiments of the present disclosure can also be implemented in cloud computing environments. In this description, “cloud computing” is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be employed in the marketplace to offer ubiquitous and convenient on-demand access to the shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then scaled accordingly.

[0269] A cloud-computing model can be composed of various characteristics such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, and so forth. A cloud-computing model can also expose various service models, such as, for example, Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (laaS). A cloud-computing model can also be deployed using different deployment models such as private cloud, community cloud, public cloud, hybrid cloud, and so forth. In this description and in the claims, a “cloud-computing environment” is an environment in which cloud computing is employed.Attorney Docket No. IP-2880-PCT 74 Patent Application

[0270] FIG. 17 illustrates a block diagram of a computing device 1700 that may be configured to perform one or more of the processes described above. One will appreciate that one or more computing devices such as the computing device 1700 may implement the pathogenicity prediction system 106. As shown by FIG. 17, the computing device 1700 can comprise a processor 1702, a memory 1704, a storage device 1706, an I / O interface 1708, and a communication interface 1710, which may be communicatively coupled by way of a communication infrastructure 1712. In certain embodiments, the computing device 1700 can include fewer or more components than those shown in FIG. 17. The following paragraphs describe components of the computing device 1700 shown in FIG. 17 in additional detail.

[0271] In one or more embodiments, the processor 1702 includes hardware for executing instructions, such as those making up a computer program. As an example, and not by way of limitation, to execute instructions for dynamically modifying workflows, the processor 1702 may retrieve (or fetch) the instructions from an internal register, an internal cache, the memory 1704, or the storage device 1706 and decode and execute them. The memory 1704 may be a volatile or nonvolatile memory used for storing data, metadata, and programs for execution by the processor(s). The storage device 1706 includes storage, such as a hard disk, flash disk drive, or other digital storage device, for storing data or instructions for performing the methods described herein.

[0272] The I / O interface 1708 allows a user to provide input to, receive output from, and otherwise transfer data to and receive data from computing device 1700. The I / O interface 1708 may include a mouse, a keypad or a keyboard, a touch screen, a camera, an optical scanner, network interface, modem, other known I / O devices or a combination of such I / O interfaces. The I / O interface 1708 may include one or more devices for presenting output to a user, including, but not limited to, a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., display drivers), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O interface 1708 is configured to provide graphical data to a display for presentation to a user. The graphical data may be representative of one or more graphical user interfaces and / or any other graphical content as may serve a particular implementation.

[0273] The communication interface 1710 can include hardware, software, or both. In any event, the communication interface 1710 can provide one or more interfaces for communication (such as, for example, packet-based communication) between the computing device 1700 and one or more other computing devices or networks. As an example, and not by way of limitation, the communication interface 1710 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI.Attorney Docket No. IP-2880-PCT 75 Patent Application

[0274] Additionally, the communication interface 1710 may facilitate communications with various types of wired or wireless networks. The communication interface 1710 may also facilitate communications using various communication protocols. The communication infrastructure 1712 may also include hardware, software, or both that couples components of the computing device 1700 to each other. For example, the communication interface 1710 may use one or more networks and / or protocols to enable a plurality of computing devices connected by a particular infrastructure to communicate with each other to perform one or more aspects of the processes described herein. To illustrate, the sequencing process can allow a plurality of devices (e.g., a client device, sequencing device, and server device(s)) to exchange information such as sequencing data and error notifications.

[0275] In the foregoing specification, the present disclosure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of the present disclosure(s) are described with reference to details discussed herein, and the accompanying drawings illustrate the various embodiments. The description above and drawings are illustrative of the disclosure and are not to be construed as limiting the disclosure. Numerous specific details are described to provide a thorough understanding of various embodiments of the present disclosure.

[0276] The present disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not restrictive. For example, the methods described herein may be performed with less or more steps / acts or the steps / acts may be performed in differing orders. Additionally, the steps / acts described herein may be repeated or performed in parallel with one another or in parallel with different instances of the same or similar steps / acts. The scope of the present application is, therefore, indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are to be embraced within their scope.Attorney Docket No. IP-2880-PCT 76 Patent Application

Claims

CLAIMSWhat is claimed is:

1. A system comprising:at least one processor; anda non-transitory computer-readable medium comprising instructions that, when executed by the at least one processor, cause the system to:access a set of inter-protein pathogenicity scores generated by an inter-protein pathogenicity model for a set of amino acids in a target protein sequence;determine an inter-protein central tendency of the set of inter-protein pathogenicity scores and an inter-protein standard deviation of the set of inter-protein pathogenicity scores;access a set of intra-protein pathogenicity scores generated by an intra-protein pathogenicity model for the set of amino acids in the target protein sequence; and generate re-scaled pathogenicity scores for the set of amino acids according to a ranking of the set of intra-protein pathogenicity scores using the inter-protein central tendency and the inter-protein standard deviation.

2. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to access the set of inter-protein pathogenicity scores by accessing inter-protein pathogenicity scores that indicate depletion of observed variants across a genomic region.

3. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to access the set of intra-protein pathogenicity scores by accessing intra-protein pathogenicity scores having values that indicate a pathogenicity ranking of the set of amino acids.

4. The system of claim 1, further comprising instructions that, when executed by the at least one processor cause the system to:determine an intra-protein central tendency of the set of intra-protein pathogenicity scores and an intra-protein standard deviation of the set of intra-protein pathogenicity scores;generate a set of standardized intra-protein pathogenicity scores from the set of intra-protein pathogenicity scores using the intra-protein central tendency and the intra-protein standard deviation; andAttorney Docket No. IP-2880-PCT 77 Patent Applicationgenerate the re-scaled pathogenicity scores for the set of amino acids by generating the rescaled pathogenicity scores from the set of standardized intra-protein pathogenicity scores using the inter-protein central tendency and the inter-protein standard deviation.

5. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate the re-scaled pathogenicity scores for the set of amino acids by:generating a set of standardized intra-protein pathogenicity scores from the set of intra-protein pathogenicity scores using a first set of transformations; andgenerating the re-scaled pathogenicity scores from the set of standardized intra-protein pathogenicity scores using a second set of transformations that is an inverse of the first set of transformations.

6. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to:access the set of intra-protein pathogenicity scores generated by the intra-protein pathogenicity model for the set of amino acids in the target protein sequence by accessing a plurality of sets of intra-protein pathogenicity scores generated by a plurality of intra-protein pathogenicity models for the set of amino acids in the target protein sequence;generate the re-scaled pathogenicity scores for the set of amino acids using the inter-protein central tendency and the inter-protein standard deviation by generating a plurality of sets of rescaled pathogenicity scores from the plurality of sets of intra-protein pathogenicity scores using the inter-protein central tendency and the inter-protein standard deviation; andcombine the plurality of sets of re-scaled pathogenicity scores to generate a combined set of re-scaled pathogenicity scores for the set of amino acids.

7. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to generate the re-scaled pathogenicity scores for the set of amino acids in the target protein sequence by generating the re-scaled pathogenicity scores for a set of at least nineteen alternative amino acids at a protein position of the target protein sequence that differ from a reference amino acid at the protein position.

8. The system of claim 7, wherein the re-scaled pathogenicity scores for the set of at least nineteen alternative amino acids comprise a set of re-scaled pathogenicity scores for a set of missense nucleotide variants.Attorney Docket No. IP-2880-PCT 78 Patent Application9. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to determine the inter-protein central tendency of the set of inter-protein pathogenicity scores by determining an inter-protein mean, an inter-protein median, or an inter-protein mode of the set of inter-protein pathogenicity scores.

10. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to access the set of inter-protein pathogenicity scores generated by the inter-protein pathogenicity model by generating the set of inter-protein pathogenicity scores using the inter-protein pathogenicity model.

11. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to access the set of intra-protein pathogenicity scores generated by the intra-protein pathogenicity model by generating the set of intra-protein pathogenicity scores using the intra-protein pathogenicity model.

12. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to:access the set of intra-protein pathogenicity scores generated by the intra-protein pathogenicity model for the set of amino acids in the target protein sequence by accessing a plurality of sets of intra-protein pathogenicity scores generated by a plurality of intra-protein pathogenicity models for the set of amino acids in the target protein sequence;combine the plurality of sets of intra-protein pathogenicity scores to generate a combined set of intra-protein pathogenicity scores; andgenerate the re-scaled pathogenicity scores for the set of amino acids using the inter-protein central tendency and the inter-protein standard deviation by generating the re-scaled pathogenicity scores from the combined set of intra-protein pathogenicity scores using the inter-protein central tendency and the inter-protein standard deviation.

13. The system of claim 12, further comprising instructions that, when executed by the at least one processor, cause the system to combine the plurality of sets of intra-protein pathogenicity scores to generate the combined set of intra-protein pathogenicity scores by generating, using forward feature selection, the combined set of intra-protein pathogenicity scores from a subset of intra-protein pathogenicity scores from the plurality of sets of intra-protein pathogenicity scores.Attorney Docket No. IP-2880-PCT 79 Patent Application14. The system of claim 1, further comprising instructions that, when executed by the at least one processor, cause the system to access the set of inter-protein pathogenicity scores generated by the inter-protein pathogenicity model for the set of amino acids in the target protein sequence by accessing inter-protein pathogenicity scores generated by the inter-protein pathogenicity model for amino acids within an amino acid window of the target protein sequence.

15. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause a computing device to:access a set of inter-protein pathogenicity scores generated by an inter-protein pathogenicity model for a set of amino acids in a target protein sequence;determine an inter-protein central tendency of the set of inter-protein pathogenicity scores and an inter-protein standard deviation of the set of inter-protein pathogenicity scores;access a set of intra-protein pathogenicity scores generated by an intra-protein pathogenicity model for the set of amino acids in the target protein sequence; andgenerate re-scaled pathogenicity scores for the set of amino acids according to a ranking of the set of intra-protein pathogenicity scores using the inter-protein central tendency and the interprotein standard deviation.

16. The non-transitory computer-readable medium of claim 15, further comprising instructions that, when executed by at least one processor, cause the computing device to access the set of inter-protein pathogenicity scores by accessing inter-protein pathogenicity scores that indicate depletion of observed variants across a genomic region.

17. The non-transitory computer-readable medium of claim 15, further comprising instructions that, when executed by at least one processor, cause the computing device to access the set of intra-protein pathogenicity scores by accessing intra-protein pathogenicity scores having values that indicate a pathogenicity ranking of the set of amino acids.

18. A computer-implemented method comprising:accessing a set of inter-protein pathogenicity scores generated by an inter-protein pathogenicity model for a set of amino acids in a target protein sequence;determining an inter-protein central tendency of the set of inter-protein pathogenicity scores and an inter-protein standard deviation of the set of inter-protein pathogenicity scores;Attorney Docket No. IP-2880-PCT 80 Patent Applicationaccessing a set of intra-protein pathogenicity scores generated by an intra-protein pathogenicity model for the set of amino acids in the target protein sequence; andgenerating re-scaled pathogenicity scores for the set of amino acids according to a ranking of the set of intra-protein pathogenicity scores using the inter-protein central tendency and the interprotein standard deviation.

19. The computer-implemented method of claim 18, further comprising: determining an intra-protein central tendency of the set of intra-protein pathogenicity scores and an intra-protein standard deviation of the set of intra-protein pathogenicity scores; generating a set of standardized intra-protein pathogenicity scores from the set of intra-protein pathogenicity scores using the intra-protein central tendency and the intra-protein standard deviation; andgenerating the re-scaled pathogenicity scores for the set of amino acids by generating the re-scaled pathogenicity scores from the set of standardized intra-protein pathogenicity scores using the inter-protein central tendency and the inter-protein standard deviation.

20. The computer-implemented method of claim 18, wherein generating the re-scaled pathogenicity scores for the set of amino acids comprises:generating a set of standardized intra-protein pathogenicity scores from the set of intra-protein pathogenicity scores using a first set of transformations; andgenerating the re-scaled pathogenicity scores from the set of standardized intra-protein pathogenicity scores using a second set of transformations that is an inverse of the first set of transformations.

21. A system comprising:at least one processor; anda non-transitory computer-readable medium comprising instructions that, when executed by the at least one processor, cause the system to:identify, for a target position of a target amino acid within a target protein sequence, proximate positions for contextual amino acids within a position window of the target position;determine, for the target position and the proximate positions, an observed number of variants within the position window;determine, for the target position and the proximate positions, an expected number of variants within the position window; andAttorney Docket No. IP-2880-PCT 81 Patent Applicationgenerate an inter-protein pathogenicity score specific to the target position based on the observed number of variants and the expected number of variants.

22. The system of claim 21, further comprising instructions that, when executed by the at least one processor, cause the system to identify the proximate positions within the position window by identifying the proximate positions for the contextual amino acids based on a three-dimensional distance of the target position in a three-dimensional structure of the target protein sequence.

23. The system of claim 21, further comprising instructions that, when executed by the at least one processor, cause the system to identify the proximate positions by identifying a predetermined number of contextual amino acids ranked by three-dimensional proximity to the target position in a three-dimensional structure of the target protein sequence.

24. The system of claim 21, further comprising instructions that, when executed by the at least one processor, cause the system to generate the inter-protein pathogenicity score based on:determining a sum of the observed number of variants and the expected number of variants; anddetermining an observed-over-expected ratio of the observed number of variants relative to the sum of the observed number of variants and the expected number of variants.

25. The system of claim 21, further comprising instructions that, when executed by the at least one processor, cause the system to determine the observed number of variants and the expected number of variants based on variant data from a plurality of species.

26. The system of claim 21, further comprising instructions that, when executed by the at least one processor, cause the system to determine the observed number of variants and the expected number of variants based on variant data from humans and at least one additional primate species.

27. The system of claim 21, further comprising instructions that, when executed by the at least one processor, cause the system to:select, from human sequences in variant data, human missense variants satisfying an allele frequency threshold; andAttorney Docket No. IP-2880-PCT 82 Patent Applicationdetermine one or more of the observed number of variants or the expected number of variants based on the selected human missense variants.

28. The system of claim 21, further comprising instructions that, when executed by the at least one processor, cause the system to:determine the observed number of variants and the expected number of variants based on a count of observed missense variants and a count of expected missense variants within the position window; andadjust the count of expected missense variants to a scaled count of expected missense variants based on both a count of observed synonymous variants and a count of expected synonymous variants.

29. The system of claim 21, further comprising instructions that, when executed by the at least one processor, cause the system to:access a gene loss of function (LoF) constraint value for the target protein sequence; determine whether the gene LoF constraint value is above or below a LoF constraint threshold; andselect a position window size or allele frequency threshold based on the determination of whether the gene LoF constraint value is above or below the LoF constraint threshold.

30. The system of claim 21, further comprising instructions that, when executed by the at least one processor, cause the system to utilize a sliding position window by:determining a subsequent position window for a subsequent target position of a subsequent target amino acid within the target protein sequence;determining, for the subsequent target position and corresponding proximate positions, an observed number of variants and an expected number of variants within the subsequent position window; andgenerating a corresponding inter-protein pathogenicity score specific to the subsequent target position based on the observed number of variants and the expected number of variants within the subsequent position window.

31. The system of claim 21 , further comprising instructions that, when executed by the at least one processor, cause the system to:determine that the observed number of variants is 0 within the position window; access, based on the observed number of variants being 0, a pseudo count of observedAttorney Docket No. IP-2880-PCT 83 Patent Applicationvariants; andgenerate the inter-protein pathogenicity score specific to the target position based on the pseudo count of observed variants and the expected number of variants.

32. The system of claim 21, further comprising instructions that, when executed by the at least one processor, cause the system to determine the observed number of variants or the expected number of variants by:determining proximate-position weights for the proximate positions based on three-dimensional or one-dimensional proximity of each position relative to the target position within the position window; anddetermining one or more of the observed number of variants or the expected number of variants according to the proximate-position weights.

33. The system of claim 21, further comprising instructions that, when executed by the at least one processor, cause the system to:access, for a set of amino acids in the target protein sequence, a set of inter-protein pathogenicity scores generated based on observed numbers of variants and expected numbers of variants within corresponding position windows;determine an inter-protein central tendency of the set of inter-protein pathogenicity scores and an inter-protein standard deviation of the set of inter-protein pathogenicity scores;access a set of intra-protein pathogenicity scores generated by an intra-protein pathogenicity model for the set of amino acids in the target protein sequence; andgenerate re-scaled pathogenicity scores for the set of amino acids according to a ranking of the set of intra-protein pathogenicity scores using the inter-protein central tendency and the inter-protein standard deviation.

34. The system of claim 21, further comprising instructions that, when executed by the at least one processor, cause the system to adjust the observed number of variants or the expected number of variants based on sequencing depth.

35. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause a computing device to:identify, for a target position of a target amino acid within a target protein sequence, proximate positions for contextual amino acids within a position window of the target position;determine, for the target position and the proximate positions, an observed number ofAttorney Docket No. IP-2880-PCT 84 Patent Applicationvariants within the position window;determine, for the target position and the proximate positions, an expected number of variants within the position window; andgenerate an inter-protein pathogenicity score specific to the target position based on the observed number of variants and the expected number of variants.

36. The non-transitory computer-readable medium of claim 35, further comprising instructions that, when executed by at least one processor, cause the computing device to identify the proximate positions within the position window by identifying the proximate positions for the contextual amino acids based on a three-dimensional distance of the target position in a three-dimensional structure of the target protein sequence.

37. The non-transitory computer-readable medium of claim 35, further comprising instructions that, when executed by at least one processor, cause the computing device to identify the proximate positions by identifying a predetermined number of contextual amino acids ranked by three-dimensional proximity to the target position in a three-dimensional structure of the target protein sequence.

38. A computer-implemented method comprising:identifying, for a target position of a target amino acid within a target protein sequence, proximate positions for contextual amino acids within a position window of the target position;determining, for the target position and the proximate positions, an observed number of variants within the position window;determining, for the target position and the proximate positions, an expected number of variants within the position window; andgenerating an inter-protein pathogenicity score specific to the target position based on the observed number of variants and the expected number of variants.

39. The computer-implemented method of claim 38, further comprising determining the observed number of variants and the expected number of variants based on variant data from a plurality of species.

40. The computer-implemented method of claim 38, wherein generating the interprotein pathogenicity score comprises:Attorney Docket No. IP-2880-PCT 85 Patent Applicationdetermining a sum of the observed number of variants and the expected number of variants; anddetermining an observed-over-expected ratio of the observed number of variants relative to the sum of the observed number of variants and the expected number of variants.

41. A system comprising:at least one processor; anda non-transitory computer-readable medium comprising instructions that, when executed by the at least one processor, cause the system to:access, for a set of amino acids in a protein sequence, a set of inter-protein pathogenicity scores generated from observed numbers of variants and expected numbers of variants within corresponding position windows;access, for the set of amino acids in the protein sequence, a set of intra-protein pathogenicity scores generated by an intra-protein pathogenicity model;generate, for a target variant at a target position with the protein sequence, an inter- and-intra score feature tensor comprising pathogenicity score features based on the set of inter-protein pathogenicity scores and the set of intra-protein pathogenicity scores; and generate, utilizing an inter-and-intra score combiner model to process the inter-and- intra score feature tensor, a combined pathogenicity score for the target variant at the target position.

42. The system of claim 41, further comprising instructions that, when executed by the at least one processor, cause the system to generate the inter-and-intra score feature tensor comprising pathogenicity score features by combining two or more of inter-protein-score features, intra-protein-score features, or mixed inter-intra-score features.

43. The system of claim 41, wherein the pathogenicity score features comprise one or more inter-protein-score features comprising:an inter-protein central tendency of a distribution of observed-over-expected ratios of the observed number of variants relative to a sum of the observed number of variants and the expected number of variants (distribution of OOERs), wherein the distribution of OOERs is over a window of positions of the protein sequence, over a batch of positions of the protein sequence, or over the protein sequence;an inter-protein minimum of the distribution of OOERs;an inter-protein maximum of the distribution of OOERs;Attorney Docket No. IP-2880-PCT 86 Patent Applicationan inter-protein mean absolute deviation of the distribution of OOERs;an inter-protein percentile of the distribution of OOERs; oran observed-over-expected ratio of the observed number of variants relative to a sum of the observed number of variants and the expected number of variants (OOER) for the target position.

44. The system of claim 41, wherein the pathogenicity score features comprise one or more intra-protein-score features comprising: i) a standardized intra-protein pathogenicity score using an intra-protein central tendency and intra-protein standard deviation of a distribution of intra-protein scores, wherein the distribution is over a batch of positions of the protein sequence or over the protein sequence, or ii) an intra-protein pathogenicity score ranking.

45. The system of claim 41, wherein the pathogenicity score features comprise one or more mixed inter-intra-score features comprising, for the target position, an observed-overexpected ratio of the observed number of variants relative to a sum of the observed number of variants and the expected number of variants (OOER) reranked according to a ranking of the set of intra-protein pathogenicity scores.

46. The system of claim 41, further comprising instructions that, when executed by the at least one processor, cause the system to:access the set of intra-protein pathogenicity scores generated by the intra-protein pathogenicity model using human variant data;generate the inter-and-intra score feature tensor comprising pathogenicity score features based on the set of intra-protein pathogenicity scores generated using the human variant data; generate, utilizing the inter-and-intra score combiner model to process the inter-and-intra score feature tensor, the combined pathogenicity score; andadjust model parameters of the inter-and-intra score combiner model based on (i) comparing the combined pathogenicity score to a ground-truth inter-protein pathogenicity score and (ii) comparing the combined pathogenicity score to a ground-truth intra-protein pathogenicity score.

47. The system of claim 41, further comprising instructions that, when executed by the at least one processor, cause the system to:access the set of intra-protein pathogenicity scores generated by the intra-protein pathogenicity model using non-human primate variant data;generate the inter-and-intra score feature tensor comprising pathogenicity score featuresAttorney Docket No. IP-2880-PCT 87 Patent Applicationbased on the set of intra-protein pathogenicity scores generated using the non-human primate variant data;generate, utilizing the inter-and-intra score combiner model to process the inter-and-intra score feature tensor, the combined pathogenicity score; andadjust model parameters of the inter-and-intra score combiner model based on (i) comparing the combined pathogenicity score to a ground-truth inter-protein pathogenicity score and (ii) comparing the combined pathogenicity score to a ground-truth intra-protein pathogenicity score.

48. The system of claim 41, wherein the inter-and-intra score combiner model comprises a neural network or other machine-learning model.

49. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause a computing device to:access, for a set of amino acids in a protein sequence, a set of inter-protein pathogenicity scores generated from observed numbers of variants and expected numbers of variants within corresponding position windows;access, for the set of amino acids in the protein sequence, a set of intra-protein pathogenicity scores generated by an intra-protein pathogenicity model;generate, for a target variant at a target position with the protein sequence, an inter-and-intra score feature tensor comprising pathogenicity score features based on the set of inter-protein pathogenicity scores and the set of intra-protein pathogenicity scores; andgenerate, utilizing an inter-and-intra score combiner model to process the inter-and-intra score feature tensor, a combined pathogenicity score for the target variant at the target position.

50. The non-transitory computer-readable medium of claim 49, further comprising instructions that, when executed by at least one processor, cause the computing device to generate the inter-and-intra score feature tensor comprising pathogenicity score features by combining two or more of inter-protein-score features, intra-protein-score features, or mixed inter-intra-score features.

51. The non-transitory computer-readable medium of claim 49, wherein the pathogenicity score features comprise one or more inter-protein-score features comprising:an inter-protein central tendency of a distribution of observed-over-expected ratios of the observed number of variants relative to a sum of the observed number of variants and the expectedAttorney Docket No. IP-2880-PCT 88 Patent Applicationnumber of variants (distribution of OOERs), wherein the distribution of OOERs is over a window of positions of the protein sequence, over a batch of positions of the protein sequence, or over the protein sequence;an inter-protein minimum of the distribution of OOERs;an inter-protein maximum of the distribution of OOERs;an inter-protein mean absolute deviation of the distribution of OOERs;an inter-protein percentile of the distribution of OOERs; oran observed-over-expected ratio of the observed number of variants relative to a sum of the observed number of variants and the expected number of variants (OOER) for the target position.

52. The non-transitory computer-readable medium of claim 49, wherein the pathogenicity score features comprise one or more intra-protein-score features comprising: i) a standardized intra-protein pathogenicity score using an intra-protein central tendency and intraprotein standard deviation of a distribution of intra-protein scores, wherein the distribution is over a batch of positions of the protein sequence or over the protein sequence, or ii) an intra-protein pathogenicity score ranking.

53. The non-transitory computer-readable medium of claim 49, wherein the pathogenicity score features comprise one or more mixed inter-intra-score features comprising, for the target position, an observed-over-expected ratio of the observed number of variants relative to a sum of the observed number of variants and the expected number of variants (OOER) reranked according to a ranking of the set of intra-protein pathogenicity scores.

54. The non-transitory computer-readable medium of claim 49, wherein the inter-and-intra score combiner model comprises a neural network or other machine-learning model.

55. A computer-implemented method comprising:accessing, for a set of amino acids in a protein sequence, a set of inter-protein pathogenicity scores generated from observed numbers of variants and expected numbers of variants within corresponding position windows;accessing, for the set of amino acids in the protein sequence, a set of intra-protein pathogenicity scores generated by an intra-protein pathogenicity model;generating, for a target variant at a target position with the protein sequence, an inter-and-intra score feature tensor comprising pathogenicity score features based on the set of inter-protein pathogenicity scores and the set of intra-protein pathogenicity scores; andAttorney Docket No. IP-2880-PCT 89 Patent Applicationgenerating, utilizing an inter-and-intra score combiner model to process the inter-and-intra score feature tensor, a combined pathogenicity score for the target variant at the target position.

56. The computer-implemented method of claim 55, wherein generating the inter-and-intra score feature tensor comprising pathogenicity score features comprises combining two or more of inter-protein-score features, intra-protein-score features, or mixed inter-intra-score features.

57. The computer-implemented method of claim 55, wherein the pathogenicity score features comprise one or more inter-protein-score features comprising:an inter-protein central tendency of a distribution of observed-over-expected ratios of the observed number of variants relative to a sum of the observed number of variants and the expected number of variants (distribution of OOERs), wherein the distribution of OOERs is over a window of positions of the protein sequence, over a batch of positions of the protein sequence, or over the protein sequence;an inter-protein minimum of the distribution of OOERs;an inter-protein maximum of the distribution of OOERs;an inter-protein mean absolute deviation of the distribution of OOERs;an inter-protein percentile of the distribution of OOERs; oran observed-over-expected ratio of the observed number of variants relative to a sum of the observed number of variants and the expected number of variants (OOER) for the target position.

58. The computer-implemented method of claim 55, wherein the pathogenicity score features comprise one or more intra-protein-score features comprising: i) a standardized intraprotein pathogenicity score using an intra-protein central tendency and intra-protein standard deviation of a distribution of intra-protein scores, wherein the distribution is over a batch of positions of the protein sequence or over the protein sequence, or ii) an intra-protein pathogenicity score ranking.

59. The computer-implemented method of claim 55, wherein the pathogenicity score features comprise one or more mixed inter-intra-score features comprising, for the target position, an observed-over-expected ratio of the observed number of variants relative to a sum of the observed number of variants and the expected number of variants (OOER) reranked according to a ranking of the set of intra-protein pathogenicity scores.Attorney Docket No. IP-2880-PCT 90 Patent Application60. The computer-implemented method of claim 55, wherein the inter-and-intra score combiner model comprises a neural network or other machine-learning model.

61. The computer-implemented method of claim 55, further comprising adjusting model parameters of the inter-and-intra score combiner model based on comparing the combined pathogenicity score to one or more ground-truth human and / or primate variants comprising observed and trinucleotide-matched unknown variants.Attorney Docket No. IP-2880-PCT 91 Patent Application