Deep learning-based use of protein contact maps for variant pathogenicity prediction
Patent Information
- Application Number
- JP2023579793
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-28
- Filing Date
- 2022-08-04
- Publication Date
- 2025-08-14
AI Technical Summary
Current methods for predicting the pathogenicity of genetic variants rely heavily on prior knowledge and are limited by the availability of clinical data, leading to suboptimal performance and reliance on traditional algorithms.
A deep learning-based approach using protein contact maps, specifically through a neural network that incorporates tensorized protein data, including protein contact maps, to predict variant pathogenicity by leveraging transfer learning and residual blocks to generate and utilize protein contact maps for pathogenicity prediction.
The method significantly improves the accuracy of predicting pathogenic variants by capturing spatial relationships in 3D protein structures, reducing reliance on prior knowledge and enhancing the model's ability to distinguish between benign and pathogenic mutations.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] Priority Application This application claims priority to and the benefit of U.S. Nonprovisional Patent Application No. 17 / 876,501, entitled "Deep Learning-Based Use of Protein Contact Maps for Variant Pathogenicity Prediction," filed on July 28, 2022 (Attorney Docket No. ILLM 1049-2 / IP-2155-US), which also claims the benefit of U.S. Provisional Patent Application No. 63 / 229,897, entitled "Transfer Learning-Based Use of Protein Contact Maps for Variant Pathogenicity Prediction," filed on August 5, 2021 (Attorney Docket No. ILLM 1042-1 / IP-2074-PRV).
[0002] This application claims priority to and the benefit of U.S. Nonprovisional Patent Application No. 17 / 876,481, entitled "Transfer Learning-Based of Protein Contact Maps for Variant Pathogenicity Prediction," filed on July 28, 2022 (Attorney Docket No. ILLM 1042-2 / IP-2074-US), which also claims the benefit of U.S. Provisional Patent Application No. 63 / 229,897, entitled "Transfer Learning-Based Use of Protein Contact Maps for Variant Pathogenicity Prediction," filed on August 5, 2021 (Attorney Docket No. ILLM 1042-1 / IP-2074-PRV).
[0003] The priority application is incorporated herein by reference for all purposes.
[0004] The disclosed technology relates to artificial intelligence based computers and digital data processing systems and corresponding data processing methods and products for mimicking intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems), including systems for reasoning with uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. In particular, the disclosed technology relates to using deep convolutional neural networks to analyze tensorized protein data, including protein contact maps, for variant pathogenicity prediction.
[0005] Built-in The following are incorporated by reference for all purposes as if fully set forth herein:
[0006] U.S. Patent Application No. 17 / 232,056, entitled “Deep Convolutional Neural Networks to Predict Variant Pathogenicity Using Three-Dimensional (3d) Protein Structures,” filed on April 15, 2021 (Attorney Docket No. ILLM 1037-2 / IP-2051-US); Sundaram,L.et al.Predicting the clinical impact of human mutation with deep neural networks.Nat.Genet.50,1161-1170(2018); Jaganathan,K.et al.Predicting splicing from primary sequence with deep learning.Cell 176,535-548(2019); U.S. patent application Ser. No. 62 / 573,144, entitled “Training a Deep Pathogenicity Classifier Using Large-Scale Benign Training Data,” filed on October 16, 2017 (Attorney Docket No. ILLM 1000-1 / IP-1611-PRV); U.S. patent application Ser. No. 62 / 573,149, entitled “Pathogenicity Classifier Based on Deep Convolutional Neural Networks (CNNs),” filed on October 16, 2017 (Attorney Docket No. ILLM 1000-2 / IP-1612-PRV); U.S. patent application Ser. No. 62 / 573,153, entitled “Deep Semi-Supervised Learning that Generates Large-Scale Pathogenic Training Data,” filed on October 16, 2017 (Attorney Docket No. ILLM 1000-3 / IP-1613-PRV); U.S. patent application Ser. No. 62 / 582,898, entitled “Pathogenicity Classification of Genomic Data Using Deep Convolutional Neural Networks (CNNs),” filed on November 7, 2017 (Attorney Docket No. ILLM 1000-4 / IP-1618-PRV); U.S. patent application Ser. No. 16 / 160,903, entitled “Deep Learning-Based Techniques for Training Deep Convolutional Neural Networks,” filed on October 15, 2018 (Attorney Docket No. ILLM 1000-5 / IP-1611-US); U.S. patent application Ser. No. 16 / 160,986, entitled “Deep Convolutional Neural Networks for Variant Classification,” filed on October 15, 2018 (Attorney Docket No. ILLM 1000-6 / IP-1612-US); U.S. patent application Ser. No. 16 / 160,968, entitled “Semi-Supervised Learning for Training an Ensemble of Deep Convolutional Neural Networks,” filed on October 15, 2018 (Attorney Docket No. ILLM 1000-7 / IP-1613-US); and U.S. patent application Ser. No. 16 / 407,149, entitled “Deep Learning-Based Techniques for Pre-Training Deep Convolutional Neural Networks,” filed May 8, 2019 (Attorney Docket No. ILLM 1010-1 / IP-1734-US).
[0007] The present application relates to deep learning-based use of protein contact maps for variant pathogenicity prediction. [Background technology]
[0008] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, it should not be assumed that the problems mentioned in this section, or associated with the subject matter provided as background, have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which as such may also correspond to embodiments of the claimed technology.
[0009] Genomics in the broad sense, also called functional genomics, aims to characterize the function of all genomic elements of an organism by using genome-scale assays such as genome sequencing, transcriptome profiling, and proteomics. Genomics has emerged as a data-driven science and operates not by testing preconceived models and hypotheses, but by discovering novel properties from the exploration of genome-scale data. Applications of genomics include finding associations between genotypes and phenotypes, discovering biomarkers for patient stratification, predicting gene functions, and mapping biochemically active genomic regions such as transcriptional enhancers.
[0010] Genomics data is too large and complex to be mined solely by visual inspection of pairwise correlations. Instead, analytical tools are needed to support the discovery of unexpected relationships, derive novel hypotheses and models, and make predictions. Unlike some algorithms in which assumptions and domain expertise are hard-coded, machine learning algorithms are designed to automatically detect patterns in data. Thus, machine learning algorithms are well suited for data-driven science, especially genomics. However, the performance of machine learning algorithms can be highly dependent on how the data is represented, i.e., how each variable (also called a feature) is calculated. For example, to classify tumors as malignant or benign from a fluorescent microscopy image, a pre-processing algorithm can detect cells, identify cell types, and generate a list of cell counts for each cell type.
[0011] Machine learning models can take estimated cell counts, which are examples of hand-designed features, as input features to classify tumors. The central problem is that classification performance is highly dependent on the quality and relevance of these features. For example, relevant visual features such as cell morphology, distance between cells or localization within an organ are not captured in cell counts, and this incomplete representation of the data can reduce classification accuracy.
[0012] Deep learning, a subfield of machine learning, addresses this problem by embedding feature computation into the machine learning model itself, generating an end-to-end model. This result has been achieved through the development of deep neural networks, machine learning models that involve successive primitive operations that compute increasingly complex features by taking the results of previous operations as input. Deep neural networks can improve prediction accuracy by discovering relevant features of high complexity, such as cell morphology and spatial organization of cells in the above example. The construction and training of deep neural networks has been made possible by the explosion of data, advances in algorithms, and substantial increases in computing power, especially through the use of graphics processing units (GPUs).
[0013] The goal of supervised learning is to obtain a model that takes features as input and returns a prediction of a so-called target variable. An example of a supervised learning problem is the problem of predicting whether an intron will be spliced or not (target) given features on the RNA such as the presence or absence of canonical splice site sequences, the location of splicing branch points or the intron length. Learning a machine learning model refers to learning its parameters, which generally involves minimizing a loss function on the training data with the goal of making accurate predictions on unknown data.
[0014] For many supervised learning problems in computational biology, the input data can be represented as a table with multiple columns or features, each of which contains numerical or categorical data that is potentially useful for making predictions. Some input data are naturally represented as tabular features (e.g., temperature or time), while other input data must first be transformed (e.g., converting deoxyribonucleic acid (DNA) sequences to k-mer counts) using a process called feature extraction to fit into a tabular representation. For intron-splicing prediction problems, the presence or absence of canonical splice site sequences, the location of splicing branch points, and intron lengths can be preprocessed features collected in tabular form. Tabular data is the norm for a wide range of supervised machine learning models, ranging from simple linear models such as logistic regression to more flexible nonlinear models such as neural networks and many others.
[0015] Logistic regression is a binary classifier, i.e., a supervised learning model that predicts a binary target variable. Specifically, logistic regression predicts the probability of a positive class by calculating a weighted sum of input features that are mapped to the [0,1] interval using a sigmoid function, a type of activation function. The parameters of logistic regression, or other linear classifiers that use different activation functions, are the weights in the weighted sum. Linear classifiers fail when, for example, the weighted sum of input features cannot sufficiently distinguish the classes of whether an intron is spliced or not. To improve prediction performance, new input features can be added manually by transforming or combining existing features in new ways, for example, by taking exponentiations or pairwise products.
[0016] Neural networks automatically learn these nonlinear feature transformations using hidden layers, each of which can be thought of as multiple linear models with their output transformed by a nonlinear activation function such as a sigmoid function or the more common rectified linear unit (ReLU). Together, these layers organize the input features into related complex patterns, facilitating the task of distinguishing between two classes.
[0017] Deep neural networks use many hidden layers, and when each neuron receives input from all neurons in the previous layer, the layer is said to be fully connected. Neural networks are generally trained using stochastic gradient descent, an algorithm suitable for training models on very large data sets. Implementation of neural networks using modern deep learning frameworks allows rapid prototyping with different architectures and data sets. Fully connected neural networks can be used for several genomics applications, including predicting the proportion of exons spliced for a given sequence from sequence features such as the presence of splice factor binding motifs or sequence conservation, prioritizing potentially disease-causing genetic variants, and predicting cis-regulatory elements in a given genomic region using features such as chromatin marks, gene expression, and evolutionary conservation.
[0018] For effective prediction, local dependencies in spatial and longitudinal data must be considered. For example, shuffling of DNA sequences or image pixels severely disrupts information patterns. These local dependencies set spatial or longitudinal data apart from tabular data, where feature ordering is arbitrary. Consider the problem of classifying genomic regions as bound vs. unbound by a particular transcription factor, where binding regions are defined as high-confidence binding events in chromatin immunoprecipitation following sequencing (ChIP-seq) data. Transcription factors bind to DNA by recognizing sequence motifs. Fully connected layers based on sequence-derived features such as the number of k-mer instances in a sequence or position weight matrix (PWM) matches can be used for this task. Such models can generalize well to sequences with the same motif located at different positions, since k-mer or PWM instance frequencies are robust to shifting motifs in the sequence. However, they cannot recognize patterns where transcription factor binding depends on the combination of multiple motifs with distinct intervals. Furthermore, the number of possible k-mers grows exponentially with k-mer length, which poses both conservation and overfitting challenges.
[0019] A convolutional layer is a special form of a fully connected layer in which the same fully connected layer is applied locally, for example within a 6 bp window, to all sequence positions. This approach can also be viewed as scanning the sequence using multiple PWMs, for example for the transcription factors GATA1 and TAL1. By using the same model parameters across positions, the total number of parameters is dramatically reduced and the network can detect motifs at positions not seen during training. Each convolutional layer scans the sequence with several filters by generating a scalar value at every position that quantizes the match between the filter and the sequence. As in a fully connected neural network, a nonlinear activation function (typically ReLU) is applied at each layer. Then, a pooling operation is applied, which aggregates the activations in successive bins across the position axis, typically taking the maximum or average activation for each channel. Pooling reduces the effective sequence length and coarsens the signal. Subsequent convolutional layers can construct the output of the previous layer and detect whether the GATA1 motif and the TAL1 motif were present within a certain distance range. Finally, the output of the convolutional layers can be used as input to a fully connected neural network to perform the final prediction task. Thus, different types of neural network layers (e.g., fully connected and convolutional layers) can be combined within a single neural network.
[0020] Convolutional neural networks (CNNs) can predict various molecular phenotypes based solely on DNA sequence. Applications include classification of transcription factor binding sites, as well as prediction of molecular phenotypes such as chromatin features, DNA contact maps, DNA methylation, gene expression, translation efficiency, RBP binding, and microRNA (miRNA) targets. In addition to predicting molecular phenotypes from sequences, convolutional neural networks can be applied to more technical tasks traditionally addressed by hand-designed bioinformatics pipelines. For example, convolutional neural networks can predict guide RNA specificity, denoise ChIP-seq, improve Hi-C data resolution, predict laboratory origin from DNA sequences, and call genetic variants. Convolutional neural networks have also been used to model long-range dependencies in genomes. Although interacting regulatory elements may be located far apart on unfolded linear DNA sequences, these elements are often proximal in real 3D chromatin conformations. Thus, modeling of molecular phenotypes from linear DNA sequences can be improved by allowing long-range dependencies, albeit with a crude approximation of chromatin, and allowing the model to implicitly learn aspects of 3D organization such as promoter-enhancer loops. This is achieved by using dilated convolutions with receptive fields of up to 32 kb. Dilated convolutions also allow splice sites to be predicted from sequences using receptive fields of 10 kb, thereby allowing integration of gene sequences over distances as long as a typical human intron (see Jaganathan, K. et al. Predicting splicing from primary sequence with deep learning. Cell 176, 535-548 (2019)).
[0021] Different types of neural networks can be characterized by their parameter sharing schemes. For example, fully connected layers have no parameter sharing, while convolutional layers impose translational invariance by applying the same filter at every position of their input. Recurrent neural networks (RNNs) are an alternative to convolutional neural networks for processing sequential data such as DNA sequences or time series that implement a different parameter sharing scheme. Recurrent neural networks apply the same operation to each sequence element. This operation takes as input the memory of the previous sequence element and the new input. It updates the memory and optionally emits an output that is either passed to subsequent layers or used directly as a model prediction. By applying the same model at each sequence element, recurrent neural networks are invariant to position index in the processed sequence. For example, recurrent neural networks can detect open reading frames in DNA sequences regardless of their position in the sequence. This task requires the recognition of a specific sequence of input, such as a start codon followed by an in-frame stop codon.
[0022] The main advantage of recurrent neural networks over convolutional neural networks is that, in theory, they can pass on information through infinitely long sequences via memory. Furthermore, recurrent neural networks can naturally handle sequences of widely varying length, such as mRNA sequences. However, convolutional neural networks combined with various tricks (such as dilated convolutions) can reach performance comparable to or even better than recurrent neural networks for sequence modeling tasks such as audio synthesis and machine translation. Recurrent neural networks can aggregate the output of convolutional neural networks to predict single-cell DNA methylation status, RBP binding, transcription factor binding, and DNA accessibility. Furthermore, recurrent neural networks apply sequential operations, so they cannot be easily parallelized and are therefore much slower computationally than convolutional neural networks.
[0023] Although most of the human genetic code is common to all humans, each human has a unique genetic code. In some cases, the human genetic code may contain outliers, called genetic variants, that may be common among a relatively small group of individuals in the human population. For example, a particular human protein may contain a particular sequence of amino acids, but variants of that protein may differ by one amino acid in an otherwise identical particular sequence.
[0024] Gene variants can be pathogenic and can result in disease. Although most such gene variants have been depleted from the genome by natural selection, the ability to identify which gene variants are likely to be pathogenic can help researchers focus on these gene variants to gain understanding of the corresponding diseases and their diagnosis, treatment, or cure. The clinical interpretation of millions of human gene variants remains unclear. Some of the most frequent pathogenic variants are single nucleotide missense mutations that change the amino acids of proteins. However, not all missense mutations are pathogenic.
[0025] Models that can predict molecular phenotypes directly from biological sequences can be used as in silico perturbation tools to investigate the association between genetic and phenotypic variation and have emerged as new methods for quantitative trait locus identification and variant prioritization. These approaches are of great importance considering that the majority of variants identified by genome-wide association studies of complex phenotypes are non-coding, which makes it difficult to estimate their effect and contribution to the phenotype. Furthermore, linkage disequilibrium results in blocks of co-inherited variants, which makes it difficult to pinpoint individual causal variants. Thus, sequence-based deep learning models that can be used as matching tools to assess the impact of such variants provide a promising approach to find potential drivers of complex phenotypes. One example is predicting the effect of non-coding single nucleotide variants and short insertions or deletions (indels) indirectly from the difference between two variants on transcription factor binding, chromatin accessibility or gene expression prediction. Another example is predicting novel splice site generation from the quantitative effect of gene variants on sequence or splicing.
[0026] An end-to-end deep learning approach for variant effect prediction is applied to predict the pathogenicity of missense variants from protein sequence and sequence conservation data (Sundaram, L. et al. Predicting the clinical impact of human mutation with deep neural networks. Nat. Genet. 50, 1161-1170 (2018), referred to herein as "PrimateAI"). PrimateAI uses a deep neural network trained on variants of known pathogenicity with data augmentation using cross-species information. In particular, PrimateAI uses sequences of wild-type and mutant proteins to compare differences and determine the pathogenicity of the mutation using a trained deep neural network. Such an approach utilizing protein sequences for pathogenicity prediction is promising because it can avoid the circularity problem and overfitting to prior knowledge. However, the number of clinical data available in ClinVar is relatively small compared to the sufficient number of data to effectively train a deep neural network. To overcome this data scarcity, PrimateAI uses common human variants and primate-derived variants as benign data, and simulated variants based on trinucleotide context as unlabeled data.
[0027] PrimateAI outperforms conventional methods when trained directly on sequence alignments. PrimateAI learns important protein domains, conserved amino acid positions, and sequence dependencies directly from training data consisting of approximately 120,000 human samples. PrimateAI substantially outperforms other variant pathogenicity prediction tools in distinguishing benign and pathogenic de novo mutations in candidate developmental disorder genes and in reproducing prior knowledge in ClinVar. These results suggest that PrimateAI is an important step forward for variant classification tools that can reduce reliance on prior knowledge of clinical reports.
[0028] Central to protein biology is the understanding of how structural elements give rise to observed functions. The plethora of protein structural data enables the development of computational methods to systematically derive the rules governing structure-function relationships. However, the performance of these methods critically depends on the choice of protein structural representation.
[0029] Protein sites are microenvironments within a protein structure that are differentiated by their structural or functional role. A site can be defined by the local neighborhood around this location where the structure or function resides. Central to rational protein engineering is the understanding of how the structural arrangement of amino acids creates functional features within a protein site. Determination of the structural and functional roles of individual amino acids in a protein provides information to aid in the manipulation and modification of protein function. Identifying functionally or structurally important amino acids allows for focused engineering efforts, such as site-directed mutagenesis, to modify the functional properties of a target protein. Alternatively, this knowledge can help avoid engineering designs that disable the desired function.
[0030] Since it is established that structure is much more conserved than sequence, the increase in protein structural data provides an opportunity to systematically study the underlying patterns governing structure-function relationships using data-driven approaches. A fundamental aspect of any computational protein analysis is how protein structural information is represented. The performance of machine learning methods often depends more on the choice of data representation than on the machine learning algorithm used. A good representation efficiently captures the most important information, whereas a poor representation produces a noisy distribution devoid of underlying patterns.
[0031] Proteins in 3D space can be viewed as complex systems that emerged through the interactions of their constituent amino acids. This representation provides a powerful framework for revealing the general organizing principles of protein contact networks. Protein residue-residue contact prediction is the problem of predicting whether any two residues in a protein sequence are spatially close to each other in the folded 3D protein structure. By analyzing whether residue pairs in a protein sequence are in contact (i.e., close in 3D space), we can form a protein contact map.
[0032] The vast number of protein structures and recent successes of deep learning algorithms provide an opportunity to develop tools to automatically extract task-specific representations of protein structures, thus predicting variant pathogenicity using tensorized protein data, including protein contact maps, as input to deep neural networks. Summary of the Invention [Means for solving the problem]
[0033] The variant pathogenicity classifier A memory that stores (i) a reference amino acid sequence of a protein, (ii) an alternative amino acid sequence of the protein containing a variant amino acid caused by a variant nucleotide, and (iii) a protein contact map of the protein; and runtime logic having access to the memory and configured to provide (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, and (iii) the protein contact map as inputs to a first neural network, and in response to processing (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, and (iii) the protein contact map, cause the first neural network to generate as an output a pathogenicity index for the variant amino acid.
[0034] In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings, in which: [Brief description of the drawings]
[0035] [Figure 1A] FIG. 1 illustrates one embodiment of training a protein contact map generation sub-network for the task of protein contact map generation to generate a so-called "trained" protein contact map generation sub-network. [Figure 1B] FIG. 1 shows one embodiment of using transfer learning to further train a protein contact map generation sub-network trained on the task of variant pathogenicity prediction to generate a so-called "cross-trained" protein contact map generation sub-network for use in training a variant pathogenicity prediction network. [Figure 1C] FIG. 1 shows one embodiment of applying the learned variant pathogenicity prediction network in inference. [Figure 1D] A diagram of two globular proteins showing some of the contacts between them as dotted black lines with contact distances in angstroms (Å). [Figure 2A] FIG. 2 illustrates an exemplary architecture of a protein contact map generation sub-network, according to one embodiment of the disclosed technology. [Figure 2B] FIG. 2 illustrates an exemplary residual block, in accordance with one implementation of the disclosed technology. [Diagram 3] FIG. 1 illustrates an exemplary architecture of a variant pathogenicity prediction network according to one embodiment of the disclosed technology. [Figure 4] FIG. 1 shows an example of a reference amino acid sequence of a protein and an example of an alternative amino acid sequence of the protein, according to one embodiment of the disclosed technology. [Diagram 5]FIG. 1 shows one-hot encoding of a reference amino acid sequence and an alternative amino acid sequence, respectively, that are processed as input by a variant pathogenicity prediction network, according to one embodiment of the disclosed technology. [Figure 6] FIG. 1 shows an exemplary three-state secondary structure profile that is processed as input by a variant pathogenicity prediction network, according to one embodiment of the disclosed technology. [Figure 7] FIG. 1 shows an exemplary 3-state solvent accessibility profile that is processed as input by the variant pathogenicity prediction network, according to one embodiment of the disclosed technology. [Figure 8] FIG. 1 shows an exemplary position-specific frequency matrix (PSFM) processed as input by the variant pathogenicity prediction network, according to one embodiment of the disclosed technology. [Figure 9] FIG. 1 shows an exemplary position-specific scoring matrix (PSSM) that is processed as input by the variant pathogenicity prediction network, according to one embodiment of the disclosed technology. [Figure 10] FIG. 1 illustrates one embodiment for generating PSFM and PSSM. [Figure 11] FIG. 1 illustrates an exemplary PSFM encoding that is processed as input by a variant pathogenicity prediction network, according to one embodiment of the disclosed technology. [Figure 12] FIG. 1 illustrates an exemplary PSSM encoding that is processed as input by a variant pathogenicity prediction network, according to one embodiment of the disclosed technology. [Figure 13] FIG. 1 illustrates an exemplary CCMpred encoding that is processed as input by a variant pathogenicity prediction network, according to one embodiment of the disclosed technology. [Figure 14] FIG. 1 shows an example of tensorized protein data processed as input by the variant pathogenicity prediction network, according to one embodiment of the disclosed technology. [Figure 15]FIG. 2 illustrates an exemplary ground truth protein contact map used to train a protein contact map generation sub-network, according to one embodiment of the disclosed technology. [Figure 16] FIG. 2 illustrates an exemplary predicted protein contact map generated by a protein contact map generation sub-network, according to one embodiment of the disclosed technology. [Figure 17] FIG. 13 is a diagram of one implementation of the so-called "outer linkage" operation used by the protein contact map generation sub-network to convert sequence features into pairwise features. [Figure 18A] FIG. 1 represents the steps in constructing a protein contact map. [Figure 18B] FIG. 1 represents the steps in constructing a protein contact map. [Figure 18C] FIG. 1 represents the steps in constructing a protein contact map. [Figure 18D] FIG. 1 represents the steps in constructing a protein contact map. [Figure 19A] FIG. 19(b) depicts the relationship between the 2D protein contact map (FIG. 19(a)) and the corresponding 3D protein structure (FIG. 19(b)). [Figure 19B] FIG. 19(b) depicts the relationship between the 2D protein contact map (FIG. 19(a)) and the corresponding 3D protein structure (FIG. 19(b)). [Figure 19C] FIG. 19(b) depicts the relationship between the 2D protein contact map (FIG. 19(a)) and the corresponding 3D protein structure (FIG. 19(b)). [Figure 19D] FIG. 19(b) depicts the relationship between the 2D protein contact map (FIG. 19(a)) and the corresponding 3D protein structure (FIG. 19(b)). [Figure 20] FIG. 1 shows different examples of 2D protein contact maps representing the corresponding 3D protein structures. [Figure 21]FIG. 1 shows different examples of 2D protein contact maps representing the corresponding 3D protein structures. [Figure 22] FIG. 1 shows different examples of 2D protein contact maps representing the corresponding 3D protein structures. [Figure 23] FIG. 1 shows different examples of 2D protein contact maps representing the corresponding 3D protein structures. [Figure 24] FIG. 1 shows different examples of 2D protein contact maps representing the corresponding 3D protein structures. [Diagram 25] FIG. 1 shows different examples of 2D protein contact maps representing the corresponding 3D protein structures. [Figure 26] FIG. 1 shows different examples of 2D protein contact maps representing the corresponding 3D protein structures. [Figure 27A] Figure 1. Graphical illustration of the concept that pathogenic variants are distributed in a spatially distant manner along a linear / contiguous amino acid sequence, but tend to cluster in specific regions of the 3D protein structure, making protein contact maps a contribution to the task of variant pathogenicity prediction. [Figure 27B] Figure 1. Graphical illustration of the concept that pathogenic variants are distributed in a spatially distant manner along a linear / contiguous amino acid sequence, but tend to cluster in specific regions of the 3D protein structure, making protein contact maps a contribution to the task of variant pathogenicity prediction. [Figure 28] FIG. 13 illustrates a pathogenicity classifier that performs variant pathogenicity classification based at least in part on a protein contact map generated by a trained protein contact map generation sub-network. [Figure 29] FIG. 2 illustrates an exemplary network architecture of a pathogenicity classifier according to one implementation of the disclosed technology. [Diagram 30] FIG. 1 is a flow diagram for carrying out one embodiment of a computer-implemented method for variant pathogenicity prediction. [Diagram 31]FIG. 1 is a flow diagram for carrying out one embodiment of a computer-implemented method for variant pathogenicity classification. [Diagram 32] FIG. 1 shows the performance results achieved by different implementations of the variant pathogenicity prediction network for the task of variant pathogenicity prediction when applied to different test datasets. [Diagram 33] FIG. 1 shows the performance results achieved by different implementations of the pathogenicity classifier for the task of variant pathogenicity classification when applied to different test sets. [Diagram 34] FIG. 1 illustrates an example computer system that can be used to implement the disclosed techniques. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0036] The following discussion is presented to enable those skilled in the art to make and use the disclosed technology and is provided in the context of a particular application and its requirements. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the disclosed technology. Thus, the disclosed technology is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0037] The detailed description of the various embodiments can be better understood when read in conjunction with the accompanying drawings. To the extent that the figures show diagrams of functional blocks of the various embodiments, the functional blocks do not necessarily show a division between hardware circuitry. Thus, for example, one or more of the functional blocks (e.g., a module, a processor, or a memory) may be implemented in a single piece of hardware (e.g., a general-purpose signal processor or a block of random access memory, a hard disk, etc.) or in multiple pieces of hardware. Similarly, a program may be a stand-alone program, may be incorporated as a subroutine in an operating system, may be a function in an installed software package, etc. It should be understood that the various embodiments are not limited to the arrangements and instrumentalities shown in the figures.
[0038] The processing engines and databases in the figures designated as modules can be implemented in hardware or software and need not be divided in exactly the same blocks as shown in the figures. Some modules may be implemented on different processors, computers or servers, or even spread among many different processors, computers or servers. In addition, it will be understood that some of the modules can be operated in parallel or in a different order than shown in the figures without affecting the functionality achieved. The modules in the figures can also be considered as flow diagram steps in a method. Also, a module does not necessarily have all code located contiguously in memory. Some portions of code can be separated from other portions of code, with code from other modules or other functions located between them.
[0039] This section is structured as follows: First, an overview of some embodiments of the disclosed technology is provided. Then, a detailed discussion of protein contact maps is provided. This is followed by some exemplary architectural details of some transfer learning embodiments and different sub-networks working together to make variant pathogenicity predictions. This is followed by exemplary encodings of different inputs such as PSSM, PSFM, CCMPred, etc. that are processed as inputs by the different sub-networks. Below is a discussion of how 2D protein contact maps are proxies for 3D protein structure and thus contribute to solving the problem of variant pathogenicity determination. Finally, we disclose a pathogenicity classifier that trains with the disclosed transfer learning embodiments and processes protein contact maps generated by another network. Some test results are also disclosed as evidence of inventiveness and non-obviousness.
[0040] Introduction Two-dimensional (2D) protein contact maps are proxies for three-dimensional (3D) protein structure because they capture short-range contacts, medium-range contacts, and other forms of long-range contacts, as well as the 3D spatial proximity of residue pairs that are spaced apart in the protein sequence. In some proteins, certain pathogenic amino acid variants that are spaced apart in the amino acid sequence have been observed to spatially cluster in the corresponding 3D protein structure. Thus, we propose that 2D protein contact maps contribute to variant pathogenicity prediction. Specifically, we present a deep neural network trained to generate variant pathogenicity predictions as output in response to processing a 2D protein contact map as input. In one embodiment, our variant pathogenicity prediction network is constructed with a one-dimensional (1D) residual block that generates per-residue features and a 2D residual block that generates per-residue features. We also use transfer learning to generate a so-called "cross-trained" protein contact map generator. This cross-trained protein contact map generator is first trained on the task of protein contact map generation and then trained on the task of variant pathogenicity prediction.
[0041] Protein contact map prediction Proteins are represented by a collection of atoms and their coordinates in three-dimensional (3D) space. Amino acids can have various atoms such as carbon atoms, oxygen (O) atoms, nitrogen (N) atoms, and hydrogen (H) atoms. The atoms can be further classified as side chain atoms and backbone atoms. Backbone carbon atoms are the alpha carbons (C α ) atoms and β carbon (C β ) atoms.
[0042] A "protein contact map" (or simply "contact map") uses a binary two-dimensional matrix to represent the distances between all possible pairs of amino acid residues in a 3D protein structure. For two residues i and j, the ij-th element of the matrix is 1 if the two residues are closer than a given threshold, and 0 otherwise. Various contact definitions have been proposed: the distance between Cα-Cα atoms with a threshold of 6-12 Å, the distance between Cβ-Cβ atoms with a threshold of 6-12 Å (Cα is used for glycine), the distance between the centers of mass of the side chains. Figures 15, 16, 18, 19, 20, 21, 22, 23, and 24 show different examples of protein contact maps.
[0043] Protein contact maps provide a more reduced representation of a protein structure than its full 3D atomic coordinates. An advantage is that protein contact maps are invariant to rotations and translations, which makes them more easily predictable by machine learning methods. It has also been shown that under certain circumstances (e.g., low content of mispredicted contacts) it is possible to reconstruct the 3D coordinates of a protein using protein contact maps. Protein contact maps are also used for protein superposition and to describe the similarity between protein structures. They are predicted from protein sequences or calculated from a given structure.
[0044] Protein contact maps describe pairwise spatial and functional relationships of amino acids (residues) in a protein and contain important information for protein 3D structure prediction. In some embodiments, two residues of a protein are in contact if their Euclidean distance is <8 Å. The distance of two residues can be calculated using Cα or Cβ atoms that correspond to Cα or Cβ-based contacts. A protein contact map can also be considered as a binary L×L matrix, where L is the protein length. In this matrix, an element with value 1 indicates that the two corresponding residues are in contact, otherwise they are not in contact.
[0045] The 3D structure of a protein is represented as the x, y, and z coordinates of amino acid atoms, and thus contacts can be defined using a distance threshold. Figure 1D shows two globular proteins with some contacts in them as black dotted lines along with the contact distances in angstroms (Å). The α-helical protein lbkr (left) has many long-range contacts, while the β-sheet protein lc9o (right) has more short- and medium-range contacts. Contacts that occur between residues that are far apart in sequence, i.e., long-range contacts, impose strong constraints on the 3D structure of a protein and are particularly important for structural analysis, understanding the folding process, and predicting 3D structures.
[0046] In some embodiments, a minimum sequence separation in the corresponding protein sequence may also be defined such that residues that are similarly close in space are excluded. Although proteins can be better reconstructed using Cβ atoms, Cα atoms, which are backbone atoms, are widely used. The choice of distance threshold and sequence separation threshold also defines the number of contacts in the protein. At lower distance thresholds, proteins have a smaller number of contacts, and at smaller sequence separation thresholds, proteins have many local contacts. In the Critical Assessment of Techniques for Protein Structure Prediction (CASP) contest, a pair of residues is defined as contacting if the distance between their Cβ atoms is 8 Å or less, provided that they are separated by at least 5 residues in the sequence. In another example, a pair of residues is said to be in contact if their Cα atoms are at least 7 Å apart and no minimum sequence separation distance is defined.
[0047] Recognizing that contact residues that are far apart in the protein sequence but close to each other in 3D space are important for protein folding, contacts are broadly classified as short-, medium-, and long-range. Short-range contacts are contacts separated by 6-11 residues in the sequence, medium-range contacts are contacts separated by 12-23 residues, and long-range contacts are contacts separated by at least 24 residues. Long-range contacts are the most important of the three types and are also the most difficult to predict, so they are often evaluated separately. As shown in Figure 1D, depending on the 3D shape (fold), some proteins have many short-range contacts, while others have more long-range contacts.
[0048] In addition to the three categories of contacts, the total number of contacts in a protein is also important for reconstructing a 3D model of a protein. Certain proteins, such as those with long tail-like structures, have fewer contacts and are difficult to reconstruct even using true contacts, while other proteins, e.g., compact globular proteins, have many contacts and can be reconstructed with high accuracy. Another important factor of predicted contacts is the coverage of the contacts, i.e., how well the contacts are distributed across the protein's structure. A set of contacts with low coverage has most of the contacts clustered in a particular region of the structure, meaning that even if all predicted contacts are correct, more information may still be needed to reconstruct the protein with high accuracy.
[0049] FIG. 1A illustrates an embodiment of training a protein contact map generation sub-network 112 for a task of protein contact map generation 100A to generate a so-called "trained" protein contact map generation sub-network 112T. In one embodiment, the protein contact map generation sub-network 112 processes as input at least one of (i) a reference amino acid sequence (REF) 102 of the protein, (ii) a secondary structure (SS) profile 104 of the protein, (iii) a solvent accessibility (SA) profile 106 of the protein, (iv) a position-specific frequency matrix (PSFM) 108 of the protein, and (v) a position-specific scoring matrix (PSSM) 110 of the protein, and is trained to generate as output a protein contact map 114. FIG. 16 illustrates an exemplary predicted protein contact map 1600 generated by the protein contact map generation sub-network, according to one embodiment of the disclosed technology. A position-specific scoring matrix (PSSM) may also be called a position-specific weight matrix (PSWM) or a position weight matrix (PWM).
[0050] In one embodiment, the protein contact map generation sub-network 112 is trained on reference amino acid sequences of bacterial proteins (e.g., 30,000 bacterial proteins) using known protein contact maps that can be used as ground truth during training. Figure 15 shows an exemplary ground truth protein contact map 1500 used to train the protein contact map generation sub-network 112, according to one embodiment of the disclosed technology.
[0051] In some embodiments, the protein contact map generation subnetwork 112 trains using a mean squared error loss function that minimizes the error between the known protein contact maps and the protein contact maps predicted by the protein contact map generation subnetwork 112 during training. In other embodiments, the protein contact map generation subnetwork 112 trains using a mean absolute error loss function that minimizes the error between the known protein contact maps and the protein contact maps predicted by the protein contact map generation subnetwork during training.
[0052] In one embodiment, the protein contact map generation sub-network 112 is a neural network. In another embodiment, the protein contact map generation sub-network 112 uses a convolutional neural network (CNN) with multiple convolutional layers. In another embodiment, the protein contact map generation sub-network 112 uses a recurrent neural network (RNN), such as a long short-term memory network (LSTM), a bidirectional LSTM (Bi-LSTM), and a gated recurrent unit (GRU). In yet another embodiment, the protein contact map generation sub-network 112 uses both a CNN and an RNN. In yet another embodiment, the protein contact map generation sub-network 112 uses a graph convolutional neural network that models dependencies in graph-structured data. In yet another embodiment, the protein contact map generation sub-network 112 uses a variational autoencoder (VAE). In yet another embodiment, the protein contact map generation sub-network 112 uses a generative adversarial network (GAN). In yet another embodiment, the protein contact map generation sub-network 112 can also be a self-attention based language model, such as those implemented by Transformer and BERT. In yet another embodiment, the protein contact map generation sub-network 112 uses a fully connected neural network (FCNN).
[0053] In yet other embodiments, the protein contact map generation sub-network 112 can use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or expanded convolution, transposed convolution, depthwise separable convolution, pointwise convolution, 1x1 convolution, group convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. The protein contact map generation sub-network 112 can use one or more loss functions, such as logistic regression / logarithmic loss, multi-class cross-entropy / softmax loss, binary cross-entropy loss, L1 loss, L2 loss, smoothed L1 loss, and Huber loss. It can use any parallel, efficient, and compaction schemes, such as TFRecords, compression encoding (e.g., PNG), sharpening, parallel calls to map transforms, batching, prefetching, model parallel, data parallel, and synchronous / asynchronous stochastic gradient descent (SGD). The protein contact map generation sub-network 112 can include upsampling layers, downsampling layers, recursive connections, gates and gated memory units (such as LSTM or GRU), residual blocks, residual connections, highway connections, skip connections, peephole connections, activation functions (e.g., nonlinear transformation functions such as rectifying linear unit (ReLU), Leaky ReLU, exponential liner unit (ELU), sigmoid and hyperbolic tangent function (tanh)), batch normalization layers, regularization layers, dropout, pooling layers (e.g., max or average pooling), global average pooling layers, attention mechanisms, and Gaussian error linear units.
[0054] The protein contact map generating sub-network 112 may be trained using back-propagation based gradient update techniques in some implementations. Exemplary gradient descent techniques that may be used to train the protein contact map generating sub-network 112 include stochastic gradient descent (SGD), batch gradient descent, and mini-batch gradient descent. Some examples of gradient descent optimization algorithms that may be used to train the protein contact map generating sub-network 112 include Momentum, Nestorv accelerated gradient, Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, and AMSGrad. In other implementations, the protein contact map generating sub-network 112 may be trained by unsupervised learning, semi-supervised learning, self-learning, reinforcement learning, multi-task learning, multi-modal learning, transfer learning, knowledge distillation, and the like.
[0055] Transfer Learning The process of reusing or transferring learned weights from one task to another is called transfer learning. Thus, transfer learning refers to extracting learned weights from a trained base network (pre-trained model) and transferring them to another untrained target network instead of training the target network from scratch. Transfer learning can be used either (a) by using the pre-trained model as a fixed feature extractor or (b) by fine-tuning the entire model. In the former scenario, for example, the last fully connected layer (classifier layer) of the pre-trained model is replaced with a new classifier layer, and then the new classifier layer is trained on the new data set. In this way, the feature extraction layer of the pre-trained model remains fixed and only the new classifier layer is fine-tuned. In the latter scenario, the entire network, i.e., the feature extraction layer of the pre-trained model and the new classifier layer, are re-trained on the new data set by continuing the backpropagation to the feature extraction layer of the pre-trained model. In this way, all weights of the entire network are fine-tubed for the new task.
[0056] The disclosed technique first trains a protein contact map generation subnetwork 112 for the task of protein contact map generation 100A (FIG. 1A), and then retrains the trained protein contact map generation subnetwork 112T for the task of variant pathogenicity prediction 100B (FIG. 1B). The retraining includes incorporating the trained protein contact map generation subnetwork 112T into a larger variant pathogenicity prediction network 190 that includes additional subnetworks (e.g., variant encoding subnetwork 128, pathogenicity scoring subnetwork 144), and jointly training the subnetworks 128, 112T, and 144 end-to-end for the task of variant pathogenicity prediction 100B to generate a so-called "trained" variant pathogenicity prediction network 190T.
[0057] In this manner, FIG. 1A can be considered as a "pre-training" stage of the protein contact map generation subnetwork 112, in which the weights (coefficients) of the protein contact map generation subnetwork 112 are trained for the task of protein contact map generation 100A, and FIG. 1B can be considered as a "transfer learning" stage of the trained protein contact map generation subnetwork 112T, in which the trained weights of the trained protein contact map generation subnetwork 112T are further trained (or transferred) 150 for the task of variant pathogenicity prediction 100B.
[0058] Those skilled in the art will appreciate that sub-networks 128, 112T, and 144 may be arranged in any order in variant pathogenicity prediction network 190. Those skilled in the art will also appreciate that variant pathogenicity prediction network 190 may include additional layers or sub-networks.
[0059] The following description focuses on one embodiment of training the variant pathogenicity prediction network 190, in which (i) the variant encoding sub-network 128 is trained to process a first input and generate a processed representation of the first input, (ii) the trained protein contact map generation sub-network 112T is further trained to process a second input and the processed representation of the first input and generate a protein contact map, and (iii) the pathogenicity scoring sub-network 144 is trained to process the protein contact map and generate a pathogenicity prediction.
[0060] In one embodiment, the first input processed by the variant encoding sub-network 128 can include at least one of (i) an alternative amino acid sequence 120 of a protein in the training data containing a variant amino acid caused by a variant nucleotide, (ii) a primate conservation profile 122 for each amino acid of the protein, (iii) a mammalian conservation profile 124 for each amino acid of the protein, and (iv) a vertebrate conservation profile 126 for each amino acid of the protein. The resulting output generated by the variant encoding sub-network 128 in response to processing the first input is a processed representation 130 of the first input. The processed representation 130 can be a convolutional feature (or activation) in some embodiments.
[0061] In one embodiment, the second input processed by the learned protein contact map generation sub-network 112T can include at least one of: (i) a reference amino acid sequence (REF) 132 of the protein, (ii) a secondary structure (SS) profile 134 of the protein, (iii) a solvent accessibility (SA) profile 136 of the protein, (iv) a position-specific frequency matrix (PSFM) 138 of the protein, and (v) a position-specific scoring matrix (PSSM) 140 of the protein. The resulting output generated by the learned protein contact map generation sub-network 112T in response to processing the second input and the processed representation 130 of the first input is a protein contact map 142.
[0062] In one embodiment, the pathogenicity scoring sub-network 144 processes the protein contact map 142 and is trained to generate as output a pathogenicity prediction 146. The pathogenicity prediction 146 indicates the degree of pathogenicity (or benignity) of the variant amino acid in the training data.
[0063] 1C illustrates an embodiment of applying the trained variant pathogenicity prediction network 190T in inference 100C. The following description focuses on an embodiment of the trained variant pathogenicity prediction network 190T, in which (i) the trained variant encoding sub-network 128T is configured to process a first input and generate a processed representation of the first input, (ii) the "cross-trained" protein contact map generation sub-network 112CT is configured to process a second input and the processed representation of the first input and generate a protein contact map, and (iii) the trained pathogenicity scoring sub-network 144T is configured to process the protein contact map and generate a pathogenicity prediction. The term "cross-trained" refers to the notion that the protein contact map generation sub-network 112 is trained for both (a) the task of protein contact map generation 100A and (b) the task of variant pathogenicity prediction 100B.
[0064] In one embodiment, the first input processed by the learned variant encoding sub-network 128T can include at least one of: (i) an alternative amino acid sequence 160 of the protein in the inference data (e.g., an unknown protein contact map of a human protein) containing variant amino acids caused by variant nucleotides; (ii) a primate conservation profile 162 per amino acid of the protein (e.g., a PSFM determined from alignment to homologous primate sequences only); (iii) a mammalian conservation profile 164 per amino acid of the protein (e.g., a PSFM determined from alignment to homologous mammalian sequences only); and (iv) a vertebrate conservation profile 166 per amino acid of the protein (e.g., a PSFM determined from alignment to homologous vertebrate sequences only). The resulting output generated by the learned variant encoding sub-network 128T in response to processing the first input is a processed representation 170 of the first input. The processed representation 170 can be a convolutional feature (or activation) in some embodiments.
[0065] In one embodiment, the second input processed by the cross-learned protein contact map generation sub-network 112CT can include at least one of: (i) a reference amino acid sequence (REF) 172 of the protein; (ii) a secondary structure (SS) profile 174 of the protein; (iii) a solvent accessibility (SA) profile 176 of the protein; (iv) a position-specific frequency matrix (PSFM) 178 of the protein; and (v) a position-specific scoring matrix (PSSM) 180 of the protein. The resulting output generated by the cross-learned protein contact map generation sub-network 112CT in response to processing the second input and the processed representation 170 of the first input is a protein contact map 182.
[0066] In one embodiment, the trained pathogenicity scoring sub-network 144T is configured to process the protein contact map 182 and generate as output a pathogenicity prediction 184. The pathogenicity prediction 184 indicates the degree of pathogenicity (or benignity) of the variant amino acid in the inference data.
[0067] In one embodiment, the mutant encoding sub-network 128 is a neural network. In another embodiment, the mutant encoding sub-network 128 uses a convolutional neural network (CNN) with multiple convolutional layers. In another embodiment, the mutant encoding sub-network 128 uses a recurrent neural network (RNN), such as a long short-term memory network (LSTM), a bidirectional LSTM (Bi-LSTM), and a gated recurrent unit (GRU). In yet another embodiment, the mutant encoding sub-network 128 uses both CNNs and RNNs. In yet another embodiment, the mutant encoding sub-network 128 uses a graph convolutional neural network that models dependencies in graph structured data. In yet another embodiment, the mutant encoding sub-network 128 uses a variational autoencoder (VAE). In yet another embodiment, the mutant encoding sub-network 128 uses a generative adversarial network (GAN). In yet another embodiment, the mutant encoding sub-network 128 can also be a self-attention based language model, such as those implemented by Transformer and BERT. In yet another embodiment, the variant encoding sub-network 128 uses a fully connected neural network (FCNN).
[0068] In yet other implementations, the mutant encoding sub-network 128 can use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or expanded convolution, transposed convolution, depth-separable convolution, point-wise convolution, 1x1 convolution, group convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. The mutant encoding sub-network 128 can use one or more loss functions, such as logistic regression / logarithmic loss, multi-class cross-entropy / softmax loss, binary cross-entropy loss, mean squared error loss, LI loss, L2 loss, smoothed LI loss, and Huber loss. It can use any parallel, efficient, and compression schemes, such as TFRecords, compression encoding (e.g., PNG), sharpening, parallel calls to map transforms, batching, prefetching, model parallel, data parallel, and synchronous / asynchronous stochastic gradient descent (SGD). The mutant encoding sub-network 128 can include upsampling layers, downsampling layers, recursive connections, gates and gated memory units (such as LSTM or GRU), residual blocks, residual connections, highway connections, skip connections, peephole connections, activation functions (e.g., nonlinear transformation functions such as rectified linear units (ReLU), leaky ReLU, exponential linear units (ELU), sigmoid, and hyperbolic tangent functions (tanh)), batch normalization layers, regularization layers, dropout, pooling layers (e.g., max or average pooling), global average pooling layers, attention mechanisms, and Gaussian error linear units.
[0069] The mutant encoding sub-network 128 may be trained using backpropagation-based gradient update techniques in some implementations. Exemplary gradient descent techniques that can be used to train the mutant encoding sub-network 128 include stochastic gradient descent (SGD), batch gradient descent, and mini-batch gradient descent. Some examples of gradient descent optimization algorithms that can be used to train the mutant encoding sub-network 128 are Momentum, Nestorv accelerated gradient, Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, and AMSGrad. In other implementations, the mutant encoding sub-network 128 can be trained by unsupervised learning, semi-supervised learning, self-learning, reinforcement learning, multitask learning, multimodal learning, transfer learning, knowledge distillation, and the like.
[0070] In one embodiment, the pathogenicity scoring sub-network 144 is a neural network. In another embodiment, the pathogenicity scoring sub-network 144 uses a convolutional neural network (CNN) with multiple convolutional layers. In another embodiment, the pathogenicity scoring sub-network 144 uses a recurrent neural network (RNN), such as a long short-term memory network (LSTM), a bidirectional LSTM (Bi-LSTM), and a gated recurrent unit (GRU). In yet another embodiment, the pathogenicity scoring sub-network 144 uses both CNNs and RNNs. In yet another embodiment, the pathogenicity scoring sub-network 144 uses a graph convolutional neural network that models dependencies in graph structured data. In yet another embodiment, the pathogenicity scoring sub-network 144 uses a variational autoencoder (VAE). In yet another embodiment, the pathogenicity scoring sub-network 144 uses a generative adversarial network (GAN). In yet another embodiment, the pathogenicity scoring sub-network 144 can also be a self-attention based language model, such as those implemented by Transformer and BERT. In yet another embodiment, the pathogenicity scoring sub-network 144 uses a fully connected neural network (FCNN).
[0071] In yet other embodiments, the pathogenicity scoring sub-network 144 can use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or expanded convolution, transposed convolution, depth-separable convolution, point-wise convolution, 1×1 convolution, group convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. The pathogenicity scoring sub-network 144 can use one or more loss functions, such as logistic regression / logarithmic loss, multi-class cross-entropy / softmax loss, binary cross-entropy loss, mean squared error loss, L1 loss, L2 loss, smoothed L1 loss, and Huber loss. It can use any parallel, efficient, and compaction schemes, such as TFRecords, compression encoding (e.g., PNG), sharpening, parallel calls to map transforms, batching, prefetching, model parallel, data parallel, and synchronous / asynchronous stochastic gradient descent (SGD). The pathogenicity scoring sub-network 144 can include upsampling layers, downsampling layers, recursive connections, gates and gated memory units (such as LSTM or GRU), residual blocks, residual connections, highway connections, skip connections, peephole connections, activation functions (e.g., nonlinear transformation functions such as rectified linear units (ReLU), leaky ReLU, exponential linear units (ELU), sigmoid, and hyperbolic tangent functions (tanh)), batch normalization layers, regularization layers, dropout, pooling layers (e.g., max or average pooling), global average pooling layers, attention mechanisms, and Gaussian error linear units.
[0072] The pathogenicity scoring sub-network 144 may be trained using backpropagation based gradient update techniques in some implementations. Exemplary gradient descent techniques that may be used to train the pathogenicity scoring sub-network 144 include stochastic gradient descent (SGD), batch gradient descent, and mini-batch gradient descent. Some examples of gradient descent optimization algorithms that may be used to train the pathogenicity scoring sub-network 144 include Momentum, Nestorv accelerated gradient, Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, and AMSGrad. In other implementations, the pathogenicity scoring sub-network 144 may be trained by unsupervised learning, semi-supervised learning, self-learning, reinforcement learning, multitask learning, multimodal learning, transfer learning, knowledge distillation, and the like.
[0073] In one embodiment, the variant pathogenicity prediction network 190 is a neural network. In another embodiment, the variant pathogenicity prediction network 190 uses a convolutional neural network (CNN) with multiple convolutional layers. In another embodiment, the variant pathogenicity prediction network 190 uses a recurrent neural network (RNN), such as a long short-term memory network (LSTM), a bidirectional LSTM (Bi-LSTM), and a gated recurrent unit (GRU). In yet another embodiment, the variant pathogenicity prediction network 190 uses both CNNs and RNNs. In yet another embodiment, the variant pathogenicity prediction network 190 uses a graph convolutional neural network that models dependencies in graph-structured data. In yet another embodiment, the variant pathogenicity prediction network 190 uses a variational autoencoder (VAE). In yet another embodiment, the variant pathogenicity prediction network 190 uses a generative adversarial network (GAN). In yet another embodiment, the variant pathogenicity prediction network 190 can be a self-attention based language model, such as those implemented by Transformer and BERT. In yet another embodiment, the variant pathogenicity prediction network 190 uses a fully connected neural network (FCNN).
[0074] In yet other embodiments, the variant pathogenicity prediction network 190 can use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or expanded convolution, transposed convolution, depth separable convolution, point-wise convolution, 1×1 convolution, group convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. The variant pathogenicity prediction network 190 can use one or more loss functions, such as logistic regression / logarithmic loss, multiclass cross-entropy / softmax loss, binary cross-entropy loss, mean squared error loss, L1 loss, L2 loss, smoothed L1 loss, and Huber loss. It can use any parallel, efficient, and compaction schemes, such as TFRecords, compression encoding (e.g., PNG), sharpening, parallel calls to map transforms, batching, prefetching, model parallel, data parallel, and synchronous / asynchronous stochastic gradient descent (SGD). The variant pathogenicity prediction network 190 can include upsampling layers, downsampling layers, recursive connections, gates and gated memory units (such as LSTM or GRU), residual blocks, residual connections, highway connections, skip connections, peephole connections, activation functions (e.g., nonlinear transformation functions such as rectified linear units (ReLU), leaky ReLU, exponential linear units (ELU), sigmoid and hyperbolic tangent functions (tanh)), batch normalization layers, regularization layers, dropout, pooling layers (e.g., max or average pooling), global average pooling layers, attention mechanisms, and Gaussian error linear units.
[0075] The variant pathogenicity prediction network 190 may be trained using backpropagation-based gradient update techniques in some implementations. Exemplary gradient descent techniques that can be used to train the variant pathogenicity prediction network 190 include stochastic gradient descent (SGD), batch gradient descent, and mini-batch gradient descent. Some examples of gradient descent optimization algorithms that can be used to train the variant pathogenicity prediction network 190 include Momentum, Nestorv accelerated gradient, Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, and AMSGrad. In other implementations, the variant pathogenicity prediction network 190 may be trained by unsupervised learning, semi-supervised learning, self-learning, reinforcement learning, multitask learning, multimodal learning, transfer learning, knowledge distillation, and the like.
[0076] Exemplary architecture of protein contact map generation subnetwork 2A illustrates an exemplary architecture 200 of the protein contact map generation sub-network 112, according to one embodiment of the disclosed technology. In one embodiment, the inputs 202 to the protein contact map generation sub-network 112 include a reference amino acid sequence of the protein under analysis, a three-state secondary structure profile of the protein under analysis, a three-state solvent accessibility profile of the protein under analysis, a position-specific frequency matrix (PSFM) of the protein under analysis, and a position-specific scoring matrix (PSSM) of the protein under analysis. In one embodiment, the input 202 is (i) an L×20×1 matrix of one-hot encoding of the reference amino acid sequence (L is the number of amino acids in the reference amino acid sequence and 20 indicates 20 amino acid categories), (ii) an L×3×1 matrix of three-state encoding of the three-state secondary structure profile (the three states are helix, beta sheet, and coil), (iii) an L×3×1 matrix of three-state encoding of the three-state solvent accessibility profile (the three states are buried, intermediate, and exposed), (iv) an L×20×1 matrix of the PSFM, and (v) a tensor linking the L×20×1 matrix of the PSFM. According to some embodiments, the resulting linked tensor 202 is of size L×66×1.
[0077] Tensor 202 is processed by one or more initial 1D convolution layers, such as 1D convolution layers 203 and 204. In the illustrated example, 1D convolution layers 203 and 204 each have 16 convolution filters, each operating on a window of size 5×1.
[0078] The output of the second 1D convolution layer 204 is provided as an input to a 1D residual block 210, which, together with a intermediate concatenation (CT) 209, performs a series of 1D convolutions (e.g., four 1D convolutions 205, 206, 207, and 208) of the sequence features at the output of the second 1D convolution layer 204. As used herein, a concatenation operation can include combining by stitching, addition, or multiplication.
[0079] FIG. 2B shows an example of a residual block that includes two convolutional layers and two activation layers. In FIG. 2B, X I and X I+1 are the input and output of the residual block, respectively. The activation layer performs a nonlinear transformation of the input without using parameters. One example of a nonlinear transformation is the rectified linear (ReLU) activation function. f(X I ) is the X that passes through two activation layers and two convolutional layers. I Let us show the result of X. I+1 =X I +f(X I ) That is, X I+1 X I and its nonlinear transformation. I ) is X l+1 and X I Since f is equal to the difference between f and , f is called the residual function, and this logic is called the residual block (or residual network or residual sub-network).
[0080] The output of the 1D residual block 210 is denoted herein as so-called "convoluted array features" 211 with dimensionality L×n. The convoluted array features 211 are converted to a 2D matrix by so-called "outer concatenation", an operation similar to a cross product. Outer concatenation is implemented by the spatial dimension expansion layer 212. Outer concatenation converts the array features into pairwise features: v={v1,v2,..,v i ,...,v L} be the final output of the 1D residual network, i.e., the convolutional sequence features 211, where L is the protein sequence length, and v i is a feature vector that stores the output information for amino acid i. For a pair of amino acids i and j, the outer link is v i , v (i+j) / 2 and v jis converted to a single vector, which is used as one input feature for this amino acid pair. Figure 17 is an embodiment of the operation of outer coupling 1700 used by the protein contact map generation sub-network 112 to convert sequence features into pairwise features. In some embodiments, the input features for this amino acid pair also include mutual information, e.g., evolutionary coupling (EC) information calculated by CCMpred and pairwise contact potentials.
[0081] The output of the spatial dimension expansion layer 212 is denoted herein as the so-called “spatially expanded output” 213, which has dimensions L×L×2n, twice the spatial dimensions of the convolutional array features 211, which has dimensions L×n.
[0082] The spatially extended output 213, in some implementations after processing by one or more initial 2D convolution layers (e.g., 2D convolution layer 214), is provided as input to a 2D residual block 226. The 2D residual block 226, together with a intermediate concatenation (CT) 225, performs a series of 2D convolutions (e.g., ten 1D convolutions 215, 216, 217, 218, 219, 220, 221, 222, 223, and 224) of the spatially extended output 213. As used herein, a concatenation operation may include combining by stitching, addition, or multiplication. In the illustrated example, each of the 2D convolution layers 215-224 has 16 convolution filters, each operating on a window of size 5×5.
[0083] The output of the 2D residual block 226 is provided as input to one or more terminal 2D convolutional layers (e.g., 2D convolutional layer 227), which produces as output a predicted protein contact map 228. The predicted protein contact map 228 has dimensions L×L×1.
[0084] In some implementations, each convolutional layer in the 1D residual block 210 and the 2D residual block 226 is preceded by a nonlinear transformation such as ReLU. Mathematically, the output of the 1D residual block 210 is a 2D matrix with dimensions L×n, where n is the number of new features (or hidden neurons / filters) generated by the last 1D convolutional layer of the 1D residual block 210. Biologically, the 1D residual block 210 learns the sequence context of amino acids. By stacking multiple 1D convolutional layers, the 1D residual block 210 learns information in a very large sequence context.
[0085] In the 2D residual block 226, the output of the 2D convolutional layer has dimensions L×L×n, where n is the number of new features (or hidden neurons / filters) generated by the 2D convolutional layer for one amino acid pair. The 2D residual block 226 learns contact occurrence patterns with higher order correlations (e.g., 2D context of amino acid pairs).
[0086] In the 1D residual block 210, I and X I+1 represent sequence features, each of which has dimensions L × n I and L×n I+1 where L is the protein sequence length, and n I (n I+1 ) can be interpreted as the number of features or hidden neurons at each position (i.e., amino acid).
[0087] In the 2D residual block 226, I and X I+1 represent pairwise features, each of which has dimensions L×L×n I and L x L x n I+1 wherein n I (n I+1 ) can be interpreted as the number of features or hidden neurons at each position (i.e., amino acid pair). In some embodiments, a position at a higher level is assumed to carry more information, so that the condition n I ≦(n I+1) will be strengthened. I <(n I+1 ), then X I +f(X I ) when calculating X I X I+1 In some implementations, to speed up training, a batch normalization layer is added before each activation layer, which normalizes the input to the activation layer to have 0 mean and 1 standard deviation.
[0088] The number of hidden neurons / filters may vary in each convolutional layer, both in the 1D residual block 210 and the 2D residual block 226. In some implementations, each of the 1D residual block 210 and the 2D residual block 226 may also include one or more residual blocks concatenated together.
[0089] 1D and 2D convolution operations are matrix-vector multiplications. Let X and Y (with dimensions L×m and L×n, respectively) be the input and output of a 1D convolution layer, respectively. Let the window size be 2w+1, and let s=(2w+1)m. The convolution operator that transforms X to Y can be represented as a 2D matrix with dimensions n×s, denoted as C. C does not depend on the length of the protein, and each convolution layer can have a different C. X i Let be a submatrix of X centered on amino acid i (1≦i≦L) with dimension (2w+1)×m, and let Y i Let Y be the i-th row of Y. i First, X i to a vector of length s, then C and the flattened X i It can be calculated by multiplying
[0090] Exemplary architecture of variant pathogenicity prediction network 3 illustrates an exemplary architecture 300 of a variant pathogenicity prediction network 190 according to one embodiment of the disclosed technology. In the illustrated example, 1D convolutions 312 and 322 form the variant encoding sub-network 128. Also in the illustrated example, a fully connected neural network 358 forms the pathogenicity scoring sub-network 144. Also in the illustrated example, 1D convolution layers 203 and 204, 1D residual block 210, spatial dimension expansion layer 212, 2D convolution layers 214 and 227, and 2D residual block 226 form the protein contact map generation sub-network 112.
[0091] In FIG. 3, the input 306 to the protein contact map generation sub-network 112 is tensorized similar to the input 202, as described above.
[0092] In Figure 3, the input 302 to the variant encoding sub-network 128 includes an alternative amino acid sequence of the protein under analysis containing the variant amino acid caused by the variant nucleotide, a primate conservation profile per amino acid of the protein under analysis, a mammalian conservation profile per amino acid of the protein under analysis, and a vertebrate conservation profile per amino acid of the protein under analysis. In one embodiment, the input 302 is a tensor that concatenates (i) an L x 20 x 1 matrix of one-hot encoding of the alternative amino acid sequence (where L is the number of amino acids in the reference amino acid sequence and 20 indicates 20 amino acid categories), (ii) an L x 20 x 1 matrix of PSFMs determined from alignments to homologous primate sequences only, (iii) an L x 20 x 1 matrix of PSFMs determined from alignments to homologous mammalian sequences only, and (iv) an L x 20 x 1 matrix of PSFMs determined from alignments to homologous vertebrate sequences only. According to some embodiments, the resulting concatenated tensor 302 is of size L x 80 x 1.
[0093] The tensor 302 is processed by one or more 1D convolutional layers (e.g., 1D convolutions 312 and 322) of the mutant encoding sub-network 128. In the illustrated example, the 1D convolutional layers 312 and 322 each have 32 convolutional filters, each operating on a window of size 5×1.
[0094] The output of the second 1D convolutional layer 322 is referred to herein as the so-called "processed representation" 334, which is provided as an input to the 1D residual block 210 of the protein contact map generation sub-network 112. In some embodiments, the output of the second 1D convolutional layer 204 of the protein contact map generation sub-network 112 is concatenated with the processed representation 334, and the resulting concatenated output is provided as an input to the 1D residual block 210. As used herein, a concatenation operation can include combining by stitching, addition, or multiplication.
[0095] As discussed above, the 1D residual block 210 generates convolved sequence features 356. Also as discussed above, the spatial dimension expansion layer 212 generates a spatially expanded output 308. The spatially expanded output 308 is processed through an initial 2D convolution layer 214, followed by a 2D residual block 226, followed by a terminal 2D convolution layer 227 to generate a predicted protein contact map 348.
[0096] The predicted protein contact map 348 is processed through a fully connected neural network 358 (and a classification layer, such as a softmax layer, a sigmoid layer, or a tanh layer (not shown)) of the pathogenicity scoring sub-network 144 to generate a variant pathogenicity score 368.
[0097] One-hot coding 4 shows an example of a reference amino acid sequence 402 of a protein 400 and an example of an alternative amino acid sequence 412 of the protein 400 according to one embodiment of the disclosed technology. The protein 400 comprises N amino acids. The positions of the amino acids in the protein 400 are labeled 1, 2, 3...N. In the illustrated example, position 16 is a position that experiences an amino acid mutation 414 (mutation) caused by an underlying nucleotide variant. For example, for the reference amino acid sequence 402, position 1 has the reference amino acid phenylalanine (F), position 16 has the reference amino acid glycine (G) 404, and position N (e.g., the last amino acid of the reference amino acid sequence 402) has the reference amino acid leucine (L). Although not illustrated for clarity, the remaining positions in the reference amino acid sequence 402 contain various amino acids in an order specific to the protein 400. Alternative amino acid sequence 412 is identical to reference amino acid sequence 402 except for variant amino acid 414 at position 16, which contains alternative amino acid alanine (A) 414 instead of reference amino acid glycine (G) 404.
[0098] Figure 5 shows one-hot encodings 514 and 516 of a reference amino acid sequence 504 and an alternative amino acid sequence 506, respectively, that are processed as input by a variant pathogenicity prediction network 190, according to one embodiment of the disclosed technology. In Figure 8, the left-most column 502 lists the 20 amino acid categories corresponding to the 20 naturally occurring amino acids that appear in the genetic code, along with a 21st gap amino acid marker for an undetermined amino acid.
[0099] In one-hot coding, each amino acid in an amino acid sequence of size L (e.g., L=51 in FIG. 5) is coded with a 20-bit (or 21-bit including gap amino acids) binary vector, where one of the bits is hot (i.e., 1) and the other is 0. A hot bit indicates that a given amino acid position in an L-length amino acid sequence belongs to a corresponding amino acid category among the 20 amino acid categories. Also, note that one-hot coded REF 514 and one-hot coded ALT 516 differ only in the 26th vector, which corresponds to the 26th position in the reference amino acid sequence 504 and the alternative amino acid sequence 506 undergoing an amino acid variant, i.e., glycine (G) → alanine (A).
[0100] Secondary structure profile Protein secondary structure (SS) refers to the local conformation of the polypeptide backbone of a protein. There are two ordered SS states, α-helix (H) and β-sheet (B), and one disordered SS state, coil (C). Figure 6 shows an exemplary three-state secondary structure profile 600 that is processed as an input by the variant pathogenicity prediction network 190, according to one embodiment of the disclosed technology. In the illustrated example, each amino acid position in the L-length reference amino acid sequence of the protein is assigned three probabilities corresponding to the three SS states H, B, and C, respectively. In some embodiments, the three probabilities for each amino acid position sum to one.
[0101] Solvent Accessibility Profile Solvent accessibility (SA) is defined as the surface area of a residue (amino acid) that is accessible to rounded solvent during surface probing of that residue. There are three SA states: buried (B), intermediate (I), and exposed (E). FIG. 7 shows an exemplary three-state solvent accessibility profile 700 that is processed as an input by the variant pathogenicity prediction network 190, according to one embodiment of the disclosed technology. In the illustrated example, each amino acid position in an L-length reference amino acid sequence of a protein is assigned three probabilities corresponding to the three SA states B, I, and E, respectively. In some embodiments, the three probabilities for each amino acid position sum to one.
[0102] PSFM and PSSM Figure 8 shows an exemplary position specific frequency matrix (PSFM) 800 that is processed as an input by the variant pathogenicity prediction network 190, according to one embodiment of the disclosed technology. Figure 9 shows an exemplary position specific scoring matrix (PSSM) 900 that is processed as an input by the variant pathogenicity prediction network 190, according to one embodiment of the disclosed technology.
[0103] Multiple sequence alignment (MSA) is the sequence alignment of multiple homologous protein sequences to a target protein. MSA is an important step in comparative analysis and property prediction of biological sequences, since a lot of information, such as evolutionary and co-evolutionary clusters, can be generated from MSA and mapped onto the selected target sequence or protein structure.
[0104] The sequence profile of a protein sequence X of length L is an L×20 matrix in the form of either a PSSM or a PSFM. The columns of PSSM and PSFM are indexed by the alphabet of amino acids, and each row corresponds to a position in the protein sequence. PSSM and PSFM contain the substitution scores and frequencies of amino acids at different positions in the protein sequence, respectively. Each row of the PSFM is normalized to sum to 1. The sequence profile of a protein sequence X is calculated by aligning X with multiple sequences in a protein database that have statistically significant sequence similarity with X. Thus, the sequence profile contains more general evolutionary and structural information of the protein family to which the protein sequence X belongs, and thus provides valuable information for remote homology detection and fold recognition.
[0105] A protein sequence (e.g., a reference amino acid sequence of a protein, called a query sequence) can be used as a seed to search and align homogeneous sequences from a protein database (e.g., SWISSPROT), for example, using the PSI-BLAST program. The aligned sequences share some homologous segments and belong to the same protein family. The aligned sequences are further converted into two profiles to represent their homology information, PSSM and PSFM. Both PSSM and PSFM are matrices with 20 rows and L columns, where L is the total number of amino acids in the query sequence. Each column of the PSSM represents the log-likelihood of a residue substitution at the corresponding position in the query sequence. The (i,j)th entry of the PSSM matrix represents the likelihood that an amino acid at the jth position of the query sequence is mutated to amino acid type i during the evolution process. The PSFM contains the weighted observation frequency of each position of the aligned sequences. Specifically, the (i,j)th entry of the PSFM matrix represents the likelihood of having amino acid type i at position j of the query sequence.
[0106] Figure 10 illustrates one embodiment for generating PSFMs and PSSMs. Figure 11 illustrates an exemplary PSFM 1100 encoding that is processed as input by a variant pathogenicity prediction network 190, according to one embodiment of the disclosed technology. Figure 12 illustrates an exemplary PSSM 1200 encoding that is processed as input by a variant pathogenicity prediction network 190, according to one embodiment of the disclosed technology.
[0107] Given a query sequence, we first obtain its sequence profile by submitting it to PSI-BLAST to search and align homologous protein sequences from a protein database 1002 (e.g., Swiss-Prot Database). Figure 10 shows the procedure of obtaining a sequence profile using the PSI-BLAST program. The parameters h and j of PSI-BLAST are usually set to 0.001 and 3, respectively. The sequence profile of a protein encompasses its homology information with respect to the query protein sequence. In PSI-BLAST, the homology information is represented by two matrices, PSFM and PSSM. Examples of PSFM and PSSM are shown in Figures 11 and 12, respectively.
[0108] In Figure 11, the (l,u)th element (l∈{1,2,...,L i},u∈{1,2,...,20}) represents the probability of having the uth amino acid at the lth position of the query protein. For example, the probability of having the amino acid M at the 1st position of the query protein is 0.36.
[0109] In Figure 12, the (l,u)th element l∈{1,2,...,L i},u∈{1,2,...,20}) represents the likelihood score that the amino acid at the lth position of the query protein will be mutated to the uth amino acid during the evolutionary process. For example, the score for an amino acid V at the 1st position of a query protein that is mutated to H during the evolutionary process is −3, while the score for an amino acid V at the 8th position is −4.
[0110] CCMpred, which resembles coevolutionary features Evolutionary coupling analysis (ECA) utilizes MSA to identify correlations in changing (co-evolving) residue pairs using the idea that closely spaced residues mutate in sync with the evolutionary functional and structural requirements of proteins. Common ECA methods include: CCMPred, FreeContact, GREMLIN, PlmDCA, and PSICOV. These methods are useful for predicting long-range contacts in proteins with multiple sequence homologs. In some embodiments, the protein contact map generation sub-network 112 (or variant pathogenicity prediction network 190) can be configured to receive evolutionary coupling features generated from CCMPred, FreeContact, GREMLIN, PlmDCA, and / or PSICOV as input and generate a protein contact map as output.
[0111] 13 shows an exemplary CCMpred encoding 1300 processed as input by the variant pathogenicity prediction network 190 according to one embodiment of the disclosed technology. The CCMPred encoding 1300 is a predicted contact probability matrix with dimensionality of sequence length (L) by sequence length (L). The CCMMPred encoding 1300 contains the co-evolutionary contact probability / score predicted using CCMPred. The CCMMPred encoding 1300 uses pseudo-likelihood maximization (PLM) to distinguish direct connections between pairs of strings in a multiple sequence alignment from merely correlated pairs.
[0112] Tensorized protein data 14 shows an example of tensorized protein data 1400 processed as input by the variant pathogenicity prediction network 190, according to one embodiment of the disclosed technology. In one embodiment, the tensorized protein data 1400 includes solvent accessibility (SA) data 1402, PSFM data 1404, PSSM data 1406, secondary structure (SS) data 1408, an atom distance matrix 1410 (protein contact map), and CCMPredz data 1412 (normalized CCMpred matrix (L*L)). In one embodiment, the name 1414 of the protein and its amino acid sequence 1416 are also identified.
[0113] 2D protein contact maps as "proxies" for 3D protein structures A protein contact map is a two-dimensional (2D) representation of a three-dimensional (3D) protein structure. A protein contact map forms a structural fingerprint of a protein, and therefore each protein can be identified based on its protein contact map. A protein contact map provides a lot of useful information about the 3D structure of a protein. For example, a cluster of contacts represents a specific secondary structure, and also captures non-local interactions, giving clues to the tertiary structure. Secondary structure, fold topology, and side chain packing patterns can also be conveniently visualized and read off from the contact map.
[0114] Protein shapes are typically described using four levels of structural complexity: primary, secondary, tertiary, and quaternary. For some proteins, a single polypeptide chain folded into its proper 3D structure creates the final protein. Protein structures are complex systems with tens, hundreds, or even thousands of residues that interact with each other and help stabilize the tertiary structure so that a specific function can be realized in vivo. In this sense, network modeling approaches are suitable to characterize and analyze protein structures, where the residues correspond to the vertices of the network and the interactions between residues (or any other type of relationship) are represented as edges connecting the corresponding nodes. One way to conceptualize and model protein structures is to consider the contacts between atoms in amino acids as a network of interactions, regardless of the secondary structure and fold type. The contacts are naturally differentiated into two types: long-range and short-range interactions. Long-range interactions occur between residues that are far from each other in the primary structure, but are located at a much closer distance in the tertiary structure. These interactions are important to define the overall topology. Short-range interactions occur between residues that are local to each other both in the primary, secondary, and tertiary structures. In most networks, what are called nodes and links are fairly simple. Looking at a protein transition state, the Cα atoms are considered to be nodes, and a link between two nodes is established if the atoms are within 8.5 Å of each other.
[0115] Figures 18A-D represent the steps in constructing a protein contact map. The Cα atoms of each amino acid are considered as vertices of the corresponding protein contact network, as shown in Figure 18A. The distance between each pair of residues is determined using Euclidean distance, and a portion of the distance matrix is shown in Figure 18B. The diagonal in the distance matrix is always 0, since the distance between the same residues is 0. To determine whether any two residues are connected, the distance between the residues should be less than or equal to a cutoff value of 7 Å distance in the illustrated embodiment. The choice of the cutoff distance is based on the extent of non-covalent interactions that are responsible for folding the polypeptide chain into its native state. Various cutoffs ranging from 5 Å to 7 Å to 8.5 Å can be used. The protein contact map is derived using the cutoff values, which are represented in a two-dimensional binary matrix (Figure 18C). If any two residues are connected, the matrix cell value is set to 1 (black), and if they are not connected, it is set to 0 (white) (Figure 18D).
[0116] Figures 19(a)-19(d) represent the relationship between the 2D protein contact map (Figure 19(b)) and the corresponding 3D protein structure (Figure 19(a)). To construct the protein contact network (Figure 19(d)) of the 3D protein structure (Figure 19(a)), Cartesian or xyz coordinates are required, which can be obtained from the RCSB protein data bank. The secondary structure of the Trp cage miniprotein (20 amino acids) is visualized using Rasmol, an open source molecular graphics visualization tool. The protein contact map is determined using a cutoff distance of 7 Å, as shown in Figure 19(b), which indicates non-covalent interactions. The protein contact network can be represented by its adjacency matrix (Figure 19(c), i.e., a binary depiction of the protein contact map). The rows or columns of the matrix represent the nodes or vertices, and the elements in the matrix represent the links or edges. An element aij in the matrix is equal to 1 whenever there is an edge connecting vertices i and j, and equal to 0 otherwise. If the graph is undirected, the adjacency matrix is symmetric, i.e., element aij=aji for any i and j. Each element of the adjacency matrix represents a connection between two nodes. For example, node 1 is connected to nodes 2, 3, 4, and 5, so a12=a13=a14=a15=1, and for the symmetric elements a21=a31=a41=a51=1. This adjacency matrix can then be visualized as an undirected network using Pajek, a program for large-scale network analysis tools, as shown in Figure 19(d).
[0117] Figures 20, 21, 22, 23, 24, 25 and 26 show different examples of 2D protein contact maps representing the corresponding 3D protein structures.
[0118] In FIG. 20, a 3D protein structure of a protein is shown on the right and a corresponding 2D protein contact map of that protein is shown on the left. The x and y axes of the 2D protein contact map are the residues (amino acids) of the protein, i.e., L×L, where L=1500. The color coding of the 2D protein contact map indicates the spatial proximity between residue pairs. For example, a residue pair of a protein that has a distance of 0-20 angstroms (Å) between them in the 3D protein structure is shown with purple contacts in the 2D protein contact map. Similarly, as another example, a residue pair of a protein that has a distance of more than 140 Å between them in the 3D protein structure is shown with yellow contacts in the 2D protein contact map.
[0119] On the right side, Figures 21-26 show the 3D protein structure of the copper transport protein ATOX1. On the left side, Figures 21-26 show the 2D protein contact maps corresponding to the 3D protein structure of the ATOX1 protein.
[0120] Please note that in Figures 21-26, the contact values and the resulting contact patterns are indicated by a color coding scheme. According to the color coding scheme, for example, residue pairs of the ATOX1 protein that have a distance of 0-5 Å between them in the 3D protein structure are shown as black contacts in the 2D protein contact map. Similarly, as another example, residue pairs of the ATOX1 protein that have a distance of more than 25 Å between them in the 3D protein structure are shown as bright orange contacts in the 2D protein contact map.
[0121] In other words, in Figures 21-26, the 2D protein contact maps show residue pairs that are "spatially close" in the 3D protein structure as "darker shading" and residue pairs that are "spatially distant" in the 3D protein structure as "lighter shading." Also, note that certain residue pairs may be "spatially distant" in the "contiguous" amino acid sequence of the protein, but "spatially close" in the 3D protein structure, and thus their "3D spatial proximity" is represented by "darker shading" in the 2D protein contact maps.
[0122] Also note that the 2D protein contact maps in Figures 21-26 have dark diagonals. This is because the 2D protein contact map is a sequence length x sequence length matrix (i.e., L x L, where L = 66), and each "matching" instance of a residue pair at the same position / same residue results in a high contact value and therefore a dark contact pattern. Thus, for example, the 2D protein contact map will have high contact values and therefore dark contact patterns for residue pairs (1,1), (2,2), (3,3), ..., (66,66), all of which fall into and form a dark diagonal in the 2D protein contact map.
[0123] Figure 21 focuses on the region of interest spanning residues 1-11 of the ATOX1 protein. Residues 1-11 are located on the β-sheet / strand arrow of the 3D protein structure of the ATOX1 protein. This β-sheet arrow is shown in red on the right side of Figure 21.
[0124] On the left, in the cyan box, FIG. 21 highlights those contact values and resulting contact patterns in the 2D protein contact map that encode spatial distances / interactions in the 3D protein structure of the ATOX1 protein between residue pairs spanning residues 1-11. Inside the cyan box, the shading of the contact values and resulting contact patterns creates a dark diagonal and a lighter adjacent region around the dark diagonal. This indicates that there is little or no 3D interaction between the residue pairs that are far apart in the sequence spanning residues 1-11. However, one exception is residue pair (4,8) or (8,4). Although residues 4 and 8 are far apart in the sequence, they have a greater 3D spatial proximity / interaction, which is indicated by the lighter shading that corresponds to the contact values for residue pair (4,8) or (8,4) in the cyan box of FIG. 21.
[0125] Figure 22 focuses on the region of interest spanning residues 12-28 of the ATOX1 protein, which are located on an α-helix in the 3D protein structure of the ATOX1 protein. This α-helix is shown in red on the right side of Figure 22.
[0126] On the left, in the cyan box, Figure 22 highlights those contact values and resulting contact patterns in the 2D protein contact map that encode the spatial distances / interactions in the 3D protein structure of the ATOX1 protein between residue pairs spanning residues 12-28. Inside the cyan box, the contact value shades and resulting contact patterns create a "zoomed-in" dark diagonal and a "zoomed-out" lighter adjacent region around the diagonal. This indicates that there are significant 3D interactions between residue pairs separated in sequence across residues 12-28. In particular, residue pairs spanning residues 12-28 that are four residue positions apart (e.g., residue pairs (12,16) or (16,12), (20,24) or (24,20) etc.) have larger interactions.
[0127] Figure 23 focuses on the region of interest spanning residues 29-47 of the ATOX1 protein. Residues 29-47 are located on two antiparallel β-sheet arrows in the 3D protein structure of the ATOX1 protein. These antiparallel β-sheet arrows run in opposite directions and are shown in red on the right side of Figure 23.
[0128] On the left, in a cyan box, FIG. 23 highlights the spatial distances / interactions in the 3D protein structure of the ATOX1 protein between residue pairs spanning residues 29-47, and their contact values and resulting contact patterns in the 2D protein contact map. Inside the cyan box, the contact value shades and resulting contact patterns create a dark diagonal of the "cross" and a lighter adjacent region of the "four triangles" around the dark diagonal of the cross. This indicates that there is significant 3D interaction between the "opposite" residue pairs spanning residues 29-47. For example, the adjacent residue pairs spanning residues 29-47 are dark (e.g., residue pair (29,30)(30,31)), but so are the opposite or inverse residue pairs (e.g., residue pairs (29,47) and (28,46)).
[0129] Figure 24 focuses on the region of interest spanning residues 48-60 of the ATOX1 protein, which is located on another α-helix in the 3D protein structure of the ATOX1 protein. This α-helix is shown in red on the right side of Figure 24.
[0130] On the left, in the cyan box, Figure 24 highlights those contact values and resulting contact patterns in the 2D protein contact map that encode the spatial distances / interactions in the 3D protein structure of the ATOX1 protein between residue pairs spanning residues 48-60. Inside the cyan box, the contact value shades and resulting contact patterns create another "expanded" dark diagonal and a "contracted" lighter adjacent region around the expanded dark diagonal. This indicates that there are significant 3D interactions between residue pairs that are separated in sequence across residues 48-60. In particular, residue pairs spanning residues 48-60 that are four residue positions apart (e.g., residue pairs (48,52) or (52,48), (56,60) or (60,56) etc.) have larger interactions.
[0131] Figure 25 focuses on the region of interest spanning residues 61-68 of the ATOX1 protein. Residues 61-68 are located on a small β-sheet / strand in the 3D protein structure of the ATOX1 protein. This small β-sheet is shown in red on the right side of Figure 25.
[0132] On the left, in the cyan box, Figure 25 highlights those contact values and resulting contact pattern in the 2D protein contact map that encode the spatial distance / interaction in the 3D protein structure of the ATOX1 protein between the residue pairs spanning residues 61-68. Inside the cyan box, the contact value shading and resulting contact pattern creates yet another "expanded" dark diagonal and a "contracted" lighter adjacent region around the expanded dark diagonal. This indicates that there is a significant 3D interaction between the sequence-distant residue pairs spanning residues 61-68.
[0133] The cyan box in FIG. 26 indicates the considerable 3D spatial proximity / interaction between the sequence-distant residue pair (8,37) and (8,60) in the 2D protein contact map of the ATOX1 protein.
[0134] 3D protein structures, and therefore 2D protein contact maps by proxy, contribute to variant pathogenicity determination The above discussion has illustrated that 2D protein contact maps are a proxy for 3D protein structure. We now turn our attention to how 3D protein structure, and thus the proxy 3D protein contact map, contributes to variant pathogenicity determination.
[0135] Figure 27 graphically illustrates the concept that pathogenic variants are distributed in a spatially distant manner along a linear / contiguous amino acid sequence, but tend to cluster in certain regions of a 3D protein structure, making protein contact maps contribute to the task of variant pathogenicity prediction. This means that protein contact maps are particularly useful in determining variant pathogenicity, as they capture the 3D spatial proximity of sequence-distant residues that experience mutations in the 3D protein structure. Thus, the disclosed technology uses protein contact maps as an input signal for generating variant pathogenicity predictions.
[0136] Pathogenicity classifier FIG. 28 illustrates a pathogenicity classifier 2812 that performs a variant pathogenicity classification 2814 based at least in part on a protein contact map 2826 generated by the learned protein contact map generation sub-network 112T.
[0137] In one embodiment, pathogenicity classifier 2812 processes at least one of: (i) a reference amino acid sequence (REF) 2816 of the protein; (ii) an alternative amino acid sequence of the protein 2804 containing a variant amino acid caused by a variant nucleotide; (iii) a primate conservation profile per amino acid of the protein 2806 (e.g., a PSFM determined from alignment to homologous primate sequences only); (iv) a mammalian conservation profile per amino acid of the protein 2808 (e.g., a PSFM determined from alignment to homologous mammalian sequences only); (v) a vertebrate conservation profile per amino acid of the protein 2816 (e.g., a PSFM determined from alignment to homologous vertebrate sequences only); and (vi) a protein contact map 2826. The resulting output generated by pathogenicity classifier 2812 is a variant pathogenicity classification 2814.
[0138] In one embodiment, the learned protein contact map generation subnetwork 112T generates a protein contact map 2826 in response to processing at least one of: (i) a reference amino acid sequence (REF) 2816 of the protein; (ii) a secondary structure (SS) profile 2818 of the protein; (iii) a solvent accessibility (SA) profile 2820 of the protein; (iv) a position-specific frequency matrix (PSFM) 2822 of the protein; and (v) a position-specific scoring matrix (PSSM) 2824 of the protein.
[0139] In one embodiment, the pathogenicity classifier 2812 is a neural network. In another embodiment, the pathogenicity classifier 2812 uses a convolutional neural network (CNN) with multiple convolutional layers. In another embodiment, the pathogenicity classifier 2812 uses a recurrent neural network (RNN), such as a long short-term memory network (LSTM), a bidirectional LSTM (Bi-LSTM), and a gated recurrent unit (GRU). In yet another embodiment, the pathogenicity classifier 2812 uses both CNNs and RNNs. In yet another embodiment, the pathogenicity classifier 2812 uses a graph convolutional neural network that models dependencies in graph structured data. In yet another embodiment, the pathogenicity classifier 2812 uses a variational autoencoder (VAE). In yet another embodiment, the pathogenicity classifier 2812 uses a generative adversarial network (GAN). In yet another embodiment, the pathogenicity classifier 2812 can be a self-attention based language model, such as implemented by Transformer and BERT. In yet another embodiment, the pathogenicity classifier 2812 uses a fully connected neural network (FCNN).
[0140] In yet other embodiments, the pathogenicity classifier 2812 can use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or expanded convolution, transposed convolution, depth separable convolution, point-wise convolution, 1×1 convolution, group convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. The pathogenicity classifier 2812 can use one or more loss functions such as logistic regression / logarithmic loss, multi-class cross-entropy / softmax loss, binary cross-entropy loss, L1 loss, L2 loss, smoothed L1 loss, and Huber loss. It can use any parallel, efficient, and compression schemes, such as TFRecords, compression encoding (e.g., PNG), sharpening, parallel calls to map transforms, batching, pre-fetching, model parallel, data parallel, and synchronous / asynchronous stochastic gradient descent (SGD). The pathogenicity classifier 2812 can include upsampling layers, downsampling layers, recursive connections, gates and gated memory units (such as LSTM or GRU), residual blocks, residual connections, highway connections, skip connections, peephole connections, activation functions (e.g., nonlinear transformation functions such as rectified linear units (ReLU), leaky ReLU, exponential linear units (ELU), sigmoid, and hyperbolic tangent functions (tanh)), batch normalization layers, regularization layers, dropout, pooling layers (e.g., max or average pooling), global average pooling layers, attention mechanisms, and Gaussian error linear units.
[0141] In some embodiments, the pathogenicity classifier 2812 can be trained using backpropagation-based gradient update techniques. Exemplary gradient descent techniques that the pathogenicity classifier 2812 can use to train include stochastic gradient descent (SGD), batch gradient descent, and mini-batch gradient descent. Some examples of gradient descent optimization algorithms that the pathogenicity classifier 2812 can use to train include Momentum, Nestorv accelerated gradient, Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, and AMSGrad. In other embodiments, the pathogenicity classifier 2812 can be trained by unsupervised learning, semi-supervised learning, self-learning, reinforcement learning, multitask learning, multimodal learning, transfer learning, knowledge distillation, and the like.
[0142] Exemplary Architecture of a Pathogenicity Classifier 29 illustrates an exemplary network architecture 2900 of a pathogenicity classifier 2812 according to one embodiment of the disclosed technology. In one embodiment, the pathogenicity classifier 2812 includes one or more initial 1D convolutional layers 2903 and 2904, followed by a first 1D residual block 2905, followed by one or more intermediate 1D convolutional layers (e.g., 1D convolutional layer 2906), followed by a second 1D residual block 2907, followed by a spatial dimension expansion layer 2909, followed by a first 2D residual block 2915, followed by one or more terminal 2D convolutional layers (e.g., 1D convolutional layer 2916), followed by a fully connected neural network 2917, followed by a classification layer (e.g., sigmoid and softmax).
[0143] In FIG. 29, input 2911 to trained protein contact map generation sub-network 112T is tensorized similarly to input 202, as described above.
[0144] In Figure 29, inputs 2902 to pathogenicity classifier 2812 include a reference amino acid sequence of the protein under analysis, an alternative amino acid sequence of the protein under analysis containing the variant amino acid caused by the variant nucleotide, a primate conservation profile per amino acid of the protein under analysis, a mammalian conservation profile per amino acid of the protein under analysis, and a vertebrate conservation profile per amino acid of the protein under analysis. In one embodiment, input 2902 is a tensor concatenating (i) an L x 20 x 1 matrix of one-hot encoding of the reference amino acid sequence (where L is the number of amino acids in the reference amino acid sequence and 20 indicates 20 amino acid categories), (ii) an L x 20 x 1 matrix of one-hot encoding of the alternative amino acid sequence, (iii) an L x 20 x 1 matrix of PSFMs determined from alignments to homologous primate sequences only, (iv) an L x 20 x 1 matrix of PSFMs determined from alignments to homologous mammalian sequences only, and (v) an L x 20 x 1 matrix of PSFMs determined from alignments to homologous vertebrate sequences only. According to some embodiments, the resulting connectivity tensor 2902 is of size L×100×1.
[0145] Tensor 2902 is processed by initial 1D convolution layers 2903 and 2904, a first 1D residual block 2905, one or more intermediate 1D convolution layers (e.g., 1D convolution layer 2906), and a second 1D residual block 2907 to generate convolved array features 2908 (L×n). A spatial dimension expansion layer 2909 processes the convolved array features 2908 and generates a spatially expanded output 2910 (L×L×2n).
[0146] The trained protein contact map generation sub-network 112T processes the input 2911 and generates a protein contact map 2912. A binner 2913 bins the contact scores / distances in the protein contact map 2912 into distance ranges. For example, the residue-to-contact distances in the protein contact map 2912 may be binned into 25 bins such as [0-1 Å], [1-2 Å], [2-3 Å], [3-4 Å], [4-5 Å], [4-6 Å], [5-6 Å], ..., [25 Å or more]. The output of the binner 2913 is the binned distances 2914 of dimension L×L×25.
[0147] The binned distances 2914 are concatenated (CT) 2920 with the spatially extended output 2910. As used herein, a concatenation operation may include combining by stitching, addition, or multiplication. The resulting concatenated output is processed by a first 2D residual block 2915, one or more terminal 2D convolutional layers (e.g., a 1D convolutional layer 2916), a fully connected neural network 2917, and a classification layer (e.g., sigmoid or softmax (not shown)) to generate a pathogenicity score 2918.
[0148] 29, "N1=2" indicates two 1D convolutional layers in the first 1D residual block 2905, "N2=3" indicates three 1D convolutional layers in the second 1D residual block 2907, and "N3=3" indicates three 2D convolutional layers in the first 2D residual block 2915. N1, N2, and N3 can be any number in different implementations.
[0149] process FIG. 30 is a flow diagram for executing one embodiment of a computer-implemented method of variant pathogenicity prediction. In one embodiment, the flow diagram of FIG. 30 is executed by runtime logic 3000. In step 3002, the method includes storing a reference amino acid sequence of the protein and an alternative amino acid sequence of the protein containing a variant amino acid caused by a variant nucleotide. In step 3012, the method includes processing the alternative amino acid sequence and generating a processed representation of the alternative amino acid sequence. In step 3012, the method includes processing the processed representation of the reference amino acid sequence and the alternative amino acid sequence and generating a protein contact map of the protein. In step 3032, the method includes processing the protein contact map and generating a pathogenicity index of the variant amino acid.
[0150] Figure 31 is a flow diagram for executing one embodiment of a computer-implemented method of variant pathogenicity classification. In one embodiment, the flow diagram of Figure 30 is executed by runtime logic 3100. In step 3102, the method includes storing (i) a reference amino acid sequence of the protein, (ii) an alternative amino acid sequence of the protein containing a variant amino acid caused by the variant nucleotide, and (iii) a protein contact map of the protein. In step 3112, the method includes providing (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, and (iii) the protein contact map as inputs to a first neural network, and causing the first neural network to generate a pathogenicity index of the variant amino acid as an output in response to processing (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, and (iii) the protein contact map.
[0151] Performance Results as Objective Indicators of Inventiveness and Nonobviousness Figure 32 shows the performance results achieved by different implementations of the variant pathogenicity prediction network 190 for the task of variant pathogenicity prediction when applied to different test datasets. The table in Figure 32 shows the performance evaluation of five models (rows) against five evaluation criteria (i.e., five test datasets) (columns).
[0152] The first model, referred to as the "1D model," is a variant pathogenicity prediction network that uses only 1D convolutions and does not use 2D contact maps as part of its input. The 1D model can be considered the benchmark model for the purposes of this disclosure. Also, note that in FIG. 32, an ensemble of eight 1D models is used to benchmark.
[0153] The second model, called "2D Cmap+1 FC All trainable," is an implementation of a variant pathogenicity prediction network 190 having a 2D convolutional and fully connected (FC) neural network (e.g., as shown in FIG. 3 with the fully connected (FC) neural network 358 portion of the pathogenicity scoring sub-network 144). "All trainable" refers to the notion that the entire variant pathogenicity prediction network 190, including the fully connected (FC) neural network, is retrained during the end-to-end retraining step of a transfer learning implementation (e.g., the transfer learning shown in FIG. 1B).
[0154] The third model, called "2D Cmap+Conservation Input Freeze Cmap Layers," is an implementation of the variant pathogenicity prediction network 190 that has 2D convolutions and uses conservation data (e.g., PSFM, PSSM, co-evolution features) as input. "Freeze Cmap Layers" refers to the notion that the layers of the variant pathogenicity prediction network 190 that generate 2D contact maps as output (e.g., protein contact map generation sub-network 112) are not retrained and remain frozen during the end-to-end retraining step of the transfer learning implementation (e.g., transfer learning shown in FIG. 1B). Note that the protein contact map generation sub-network 112 is trained at least once as shown in FIG. 1A, but in some implementations of transfer learning, it is not retrained in FIG. 1B as part of the variant pathogenicity prediction network 190. In other implementations of transfer learning, the protein contact map generation sub-network 112 can be retrained as part of the variant pathogenicity prediction network 190.
[0155] A fourth model, called "2D Cmap+Conservation Input All trainable", is an implementation of the variant pathogenicity prediction network 190 with 2D convolutions and using conservation data (e.g., PSFM, PSSM, co-evolutionary features) as input. "All trainable" refers to the notion that the entire variant pathogenicity prediction network 190, including the variant encoding sub-network 128, the protein contact map generation sub-network 112, and the pathogenicity scoring sub-network 144, is retrained during the end-to-end retraining step of the transfer learning implementation (e.g., the transfer learning shown in FIG. 1B).
[0156] The fifth model, called "2D Cmap+Conservation Input All trainable", is an ENSEMBLE implementation of one of the mutant pathogenicity prediction networks 190 with 2D convolution and using conservation data (e.g., PSFM, PSSM, co-evolution features) as input. "Ensemble" refers to the concept that multiple instances of the mutant pathogenicity prediction network 190 process the same input separately and generate respective outputs (e.g., respective pathogenicity predictions). A final output (e.g., a final pathogenicity prediction) is generated based on the respective outputs (e.g., by averaging the respective pathogenicity predictions or by selecting the maximum one of the respective pathogenicity predictions). The multiple instances of the mutant pathogenicity prediction network 190 have different coefficient / weight values but the same architecture. In the embodiment shown in FIG. 32, the ensemble has 10 instances of the mutant pathogenicity prediction network 190. "Fully trainable" refers to the notion that the entire variant pathogenicity prediction network 190 is retrained during the end-to-end retraining step of a transfer learning embodiment (eg, the transfer learning shown in FIG. 1B).
[0157] Turning to the five evaluation criteria, the first criterion, "Accuracy on Benign Test Set," refers to the predictive accuracy of a given model on a dataset of, e.g., 10,000 benign variants, which may include benign variants, e.g., human benign variants and non-human primate benign variants (e.g., those discovered by primate AI).
[0158] The second evaluation metric, "-log(Pval) in DDD vs. controls", indicates the accuracy of a given model in identifying / segregating pathogenic variants taken from individuals with developmental disability (DDD) like Down syndrome as "pathogenic" and benign variants taken from healthy individuals (controls) as "benign" using the negative logarithm p-value (-log(Pval)) of the Wilcoxon rank sum test.
[0159] The third evaluation criterion, "-log(Pval) in 605 genes in DDD vs. controls", uses the negative logarithmic p-value (-log(Pval)) of the Wilcoxon rank sum test to indicate the accuracy of a given model in identifying / segregating pathogenic variants located in one of the "605 genes" taken from individuals with developmental disorders like Down's syndrome (DDD) and clinically known to experience pathogenic variants as "pathogenic" and identifying / segregating benign variants taken from healthy individuals (controls) as "benign".
[0160] The fourth evaluation criterion, "-log(Pval) in new DDD vs new controls," uses the negative logarithm of the p-value (-log(Pval)) of the Wilcoxon rank-sum test to indicate the accuracy of a given model in identifying / segregating pathogenic variants taken from new individuals with developmental disorders (DDDs) like Down's syndrome as "pathogenic" and identifying / segregating benign variants taken from new healthy individuals (controls) as "benign."
[0161] The fifth evaluation criterion, "-log(Pval) in 605 genes in new DDD vs new controls," uses the negative logarithmic p-value (-log(Pval)) of the Wilcoxon rank sum test to indicate the accuracy of a given model in identifying / segregating pathogenic variants taken from new individuals who have a developmental disorder like Down's syndrome (DDD) and are located in one of the "605 genes" clinically known to experience pathogenic variants as "pathogenic," and identifying / segregating benign variants taken from new healthy individuals (controls) as "benign."
[0162] Turning to the performance results of the five models on the five evaluation criteria (i.e., five test datasets), the fifth model, i.e., the "ENSEMBLE 2D Cmap+Conservation Input All trainable" model, outperforms all other models. This is demonstrated by the 90.7% prediction accuracy of the fifth model in predicting benign variants as "benign" in the 10,000 benign variant test dataset, and also by the higher p-value. A higher p-value indicates that a given model better separates / distinguishes pathogenic / disease-causing / harmful DDD variants from benign control variants, thereby demonstrating better model performance.
[0163] FIG. 33 shows the performance results achieved by different implementations of the pathogenicity classifier for the task of variant pathogenicity classification when applied to different test sets.
[0164] The table in Figure 33 shows the performance evaluation of six models (rows) against two evaluation criteria (i.e., two test data sets) (columns). The use of 2D contact maps (e.g., with a sixth model) is also evaluated against no use of a 2D model.
[0165] The first test dataset, "Accuracy in Benign Test Set," is a dataset of benign variants, e.g., 10,000 benign variants, which may include human benign variants and non-human primate benign variants (e.g., as discovered by PrimateAI). The second test dataset, "-log(Pval) in DDD vs. Controls," shows the accuracy of a given model in identifying / segregating pathogenic variants taken from individuals with developmental disorders (DDD) such as Down's syndrome as "pathogenic" and benign variants taken from healthy individuals (controls) as "benign" using the negative logarithmic p-value (-log(Pval)) of the Wilcoxon rank sum test. Also note that in FIG. 33, each of the six models is implemented as an ensemble of eight instances. In other implementations, a different number of instances can be used.
[0166] The first model, referred to as the "1D model," is a variant pathogenicity prediction network that uses only 1D convolutions and does not use 2D contact maps as part of its input. The 1D model can be considered the benchmark model for the purposes of this disclosure.
[0167] The five 2D models (rows 2-6), i.e., five different implementations of the pathogenicity classifier 2812, differ in their respective architectures, with different numbers of residual blocks in different residual block sets N1, N2, and N3, the use of fully connected layers versus none, and different filter sizes (e.g., 5x2 vs. 2x5).
[0168] As can be seen in FIG. 33, the pathogenicity classifier 2812 using the 2D contact map as input features, i.e., the sixth model, has better performance on average.
[0169] Computer Systems 34 is an exemplary computer system 3400 that can be used to implement the disclosed techniques. The computer system 3400 includes at least one central processing unit (CPU) 3472 that communicates with a number of peripheral devices via a bus subsystem 3455. These peripheral devices can include, for example, a storage subsystem 3410 including memory devices and a file storage subsystem 3436, a user interface input device 3438, a user interface output device 3476, and a network interface subsystem 3474. The input and output devices enable user interaction with the computer system 3400. The network interface subsystem 3474 provides an interface to external networks, including interfaces to corresponding interface devices in other computer systems.
[0170] In one embodiment, the pathogenicity classifier 2104 is communicatively linked to a memory subsystem 3410 and a user interface input device 3438 .
[0171] The user interface input devices 3438 can include pointing devices such as a keyboard, a mouse, a trackball, a touch pad, or a graphics tablet, a scanner, a touch screen integrated into a display, audio input devices such as a voice recognition system and a microphone, as well as other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and manners for inputting information into the computer system 3400.
[0172] The user interface output devices 3476 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as an LED display, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide non-visual displays such as an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and manners for outputting information from the computer system 3400 to a user or to another machine or computer system.
[0173] The storage subsystem 3410 stores programming and data constructs that provide the functionality of some or all of the modules and methods described herein. These software modules are generally executed by the processor 3478.
[0174] The processor 3478 may be a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and / or a coarse-grained reconfigurable architecture (CGRA). The processor 3478 may be hosted by a deep learning cloud platform such as Google Cloud Platform™, Xilinx™, and Cirrascale™. Examples of processor 3478 include Google's Tensor Processing Unit (TPU)™, rackmount solutions such as the GX4 Rackmount Series™, GX34 Rackmount Series™, NVIDIA DGX-1™, Microsoft's Stratix V FPGA™, Graphcore's Intelligent Processor Unit (IPU)™, Qualcomm's Zeroth Platform™ with Snapdragon processors™, NVIDIA's Volta™, NVIDIA's DRIVE PX™, NVIDIA's JETSON TX1 / TX2 MODULE™, Intel's Nirvana™, Movidius VPU™, Fujitsu DPI™, ARM's DynamicIQ™, IBM TrueNorth™, Lambda GPU Server with Testa V100s™, and others.
[0175] The memory subsystem 3422 used in the storage subsystem 3410 may include multiple memories including a main random access memory (RAM) 3432 for storing instructions and data during program execution, and a read only memory (ROM) 3434 in which fixed instructions are stored. The file storage subsystem 3436 may provide persistent storage for program and data files and may include a hard disk drive, associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of a particular embodiment may be stored by the file storage subsystem 3436 in the storage subsystem 3410 or in other machines accessible by the processor.
[0176] The bus subsystem 3455 provides a mechanism for allowing the various components and subsystems of the computer system 3400 to communicate with each other as intended. Although the bus subsystem 3455 is shown generally as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0177] The computer system 3400 itself can be of various types, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a loosely distributed set of loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the description of computer system 3400 shown in Figure 34 is intended only as a specific example for purposes of illustrating a preferred embodiment of the present invention. Many other configurations of computer system 3400 can have more or fewer components than the computer system shown in Figure 34.
[0178] As used herein, "logic" may be implemented in the form of a computer product including a non-transitory computer-readable storage medium with program code usable by a computer to perform the method steps described herein. "Logic" may be implemented in the form of an apparatus including a memory and at least one processor coupled to the memory and operative to perform the exemplary method steps. "Logic" may be implemented in the form of a means for performing one or more of the method steps described herein, which may include (i) hardware modules, (ii) software modules running on one or more hardware processors, or (iii) a combination of hardware and software modules, any of (i)-(iii) implementing the particular techniques described herein, and the software modules being stored in a computer-readable storage medium (or multiple such media). In one embodiment, logic implements a data processing function. Logic may be a general-purpose, single-core or multi-core processor with a computer program that specifies the function, a digital signal processor with a computer program, configurable logic such as an FPGA with a configuration file, a special-purpose circuit such as a state machine, or any combination thereof. Also, a computer program product may embody the computer program and configuration file portions of logic.
[0179] item The disclosed technology can be implemented as a system, a method, or a product. One or more features of the embodiments can be combined with the basic embodiment. Non-mutually exclusive embodiments are taught as combinable. One or more features of the embodiments can be combined with other embodiments. The present disclosure will periodically inform users of these options. The omission from some embodiments of the enumeration of repeating these options should not be interpreted as limiting the combinations taught in the preceding section. These descriptions are incorporated herein by reference to each of the following embodiments.
[0180] One or more embodiments and items of the disclosed technology, or elements thereof, can be implemented in the form of a computer product including a non-transitory computer-readable storage medium with computer usable program code for performing the illustrated method steps. Furthermore, one or more embodiments and items of the disclosed technology, or elements thereof, can be implemented in the form of an apparatus including a memory and at least one processor coupled to the memory and operative to perform the illustrated method steps. Furthermore, in another aspect, one or more embodiments and items of the disclosed technology, or elements thereof, can be implemented in the form of a means for performing one or more of the method steps described herein, which means can include (i) a hardware module, (ii) a software module running on one or more hardware processors, or (iii) a combination of hardware and software modules, any of (i)-(iii) implementing the particular technology described herein, and the software module is stored in a computer-readable storage medium (or multiple such media).
[0181] The items described in this section can be combined as features. For the sake of brevity, combinations of features are not listed separately and are not repeated for each base set of features. The reader will understand how features identified in the items described in this section can be easily combined with sets of basic features identified as embodiments in other sections of this application. These items are not meant to be mutually exclusive, exhaustive, or limiting, and the disclosed technology is not limited to these items, but rather includes all possible combinations, modifications, and variations within the scope of the technology described in the claims and its equivalents.
[0182] Other implementations of the items described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the items described in this section. Yet another implementation of the items described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the items described in this section.
[0183] The present inventors disclose the following items:
[0184] Item Set 1 1. A variant pathogenicity prediction network, comprising: A memory that stores a reference amino acid sequence of the protein and an alternative amino acid sequence of the protein containing a variant amino acid caused by a variant nucleotide; a variant encoding sub-network having access to a memory and configured to process alternative amino acid sequences and generate processed representations of the alternative amino acid sequences; a protein contact map generation sub-network in communication with the variant encoding sub-network, the protein contact map generation sub-network configured to process the processed representations of the reference amino acid sequence and the alternative amino acid sequence and to generate a protein contact map for the protein; A variant pathogenicity prediction network comprising: a pathogenicity scoring sub-network in communication with a protein contact map generation sub-network, the pathogenicity scoring sub-network configured to process the protein contact map and generate a pathogenicity index for the variant amino acid. 2. The memory further stores a primate conservation profile for each amino acid of the protein; 2. The variant pathogenicity prediction network of item 1, wherein the processed representations of the alternative amino acid sequences are generated by the variant encoding sub-network in response to processing the alternative amino acid sequences and the primate conservation profiles for each amino acid. 3. The memory further stores a mammalian conservation profile for each amino acid of the protein; 3. The variant pathogenicity prediction network of claim 1 or 2, wherein the processed representations of the alternative amino acid sequences are generated by the variant encoding sub-network in response to processing the alternative amino acid sequences and the mammalian conservation profiles for each amino acid. 4. The memory further stores a vertebrate conservation profile for each amino acid of the protein; 4. The variant pathogenicity prediction network of any of items 1 to 3, wherein the processed representations of the alternative amino acid sequences are generated by the variant encoding sub-network in response to processing the alternative amino acid sequences and the vertebrate conservation profiles for each amino acid. 5. The variant pathogenicity prediction network of any of items 1 to 4, wherein the processed representations of the alternative amino acid sequences are generated by the variant encoding sub-network in response to processing the alternative amino acid sequences, the primate conservation profile per amino acid, the mammalian conservation profile per amino acid, and the vertebrate conservation profile per amino acid. 6. The variant pathogenicity prediction network of any of items 1 to 5, wherein the processed representations of the alternative amino acid sequences are generated by the variant encoding sub-network in response to processing the alternative amino acid sequences, the primate conservation profiles per amino acid, and the mammalian conservation profiles per amino acid. 7. The variant pathogenicity prediction network of any of items 1 to 6, wherein the processed representation of the alternative amino acid sequence is generated by the variant encoding sub-network in response to processing the alternative amino acid sequence, the primate conservation profile per amino acid, and the vertebrate conservation profile per amino acid. 8. The variant pathogenicity prediction network of any of items 1 to 7, wherein the processed representation of the alternative amino acid sequence is generated by the variant encoding sub-network in response to processing the alternative amino acid sequence, the mammalian conservation profile per amino acid, and the vertebrate conservation profile per amino acid. 9. The memory further stores a secondary structure profile for each amino acid of the protein; 9. The variant pathogenicity prediction network according to any one of items 1 to 8, wherein the protein contact map of the protein is generated by the protein contact map generation sub-network in response to processing the reference amino acid sequence and the secondary structure profile for each amino acid. 10. The memory further stores a solvent accessibility profile for each amino acid of the protein; 10. The variant pathogenicity prediction network according to any one of items 1 to 9, wherein the protein contact map of the protein is generated by the protein contact map generation sub-network in response to processing the reference amino acid sequence and the solvent accessibility profile for each amino acid. 11. The memory further stores a position-specific frequency matrix for each amino acid of the protein; 11. The variant pathogenicity prediction network according to any one of items 1 to 10, wherein the protein contact map of the protein is generated by the protein contact map generation sub-network in response to processing the reference amino acid sequence and the position-specific frequency matrix for each amino acid. 12. The memory further stores a position-specific scoring matrix for each amino acid of the protein; 12. The variant pathogenicity prediction network according to any one of items 1 to 11, wherein the protein contact map of the protein is generated by the protein contact map generation sub-network in response to processing the reference amino acid sequence and the position-specific scoring matrix for each amino acid. 13. The variant pathogenicity prediction network according to any of items 1 to 12, wherein the protein contact map of the protein is generated by the protein contact map generation sub-network in response to processing the reference amino acid sequence, the secondary structure profile per amino acid, the solvent accessibility profile per amino acid, the position-specific frequency matrix per amino acid, and the position-specific scoring matrix per amino acid. 14. The variant pathogenicity prediction network according to any one of items 1 to 13, wherein the protein contact map of the protein is generated by the protein contact map generation subnetwork in response to processing the reference amino acid sequence, the secondary structure profile per amino acid, and the solvent accessibility profile per amino acid. 15. The variant pathogenicity prediction network according to any one of items 1 to 14, wherein the protein contact map of the protein is generated by the protein contact map generation sub-network in response to processing the reference amino acid sequence, the secondary structure profile for each amino acid, and the position-specific frequency matrix for each amino acid. 16. The variant pathogenicity prediction network according to any one of items 1 to 15, wherein the protein contact map of the protein is generated by a protein contact map generation sub-network in response to processing the reference amino acid sequence, the secondary structure profile for each amino acid, and the position-specific scoring matrix for each amino acid. 17. The variant pathogenicity prediction network according to any of items 1 to 16, wherein the protein contact map of the protein is generated by the protein contact map generation subnetwork in response to processing the reference amino acid sequence, the per-amino acid solvent accessibility profile, and the per-amino acid position-specific frequency matrix. 18. The variant pathogenicity prediction network according to any of items 1 to 17, wherein the protein contact map of the protein is generated by the protein contact map generation subnetwork in response to processing the reference amino acid sequence, the per-amino acid solvent accessibility profile, and the per-amino acid position-specific scoring matrix. 19. The variant pathogenicity prediction network according to any of items 1 to 18, wherein the protein contact map of the protein is generated by the protein contact map generation subnetwork in response to processing the reference amino acid sequence, the position-specific frequency matrix per amino acid, and the position-specific scoring matrix per amino acid. 20. The variant pathogenicity prediction network according to any one of items 1 to 19, wherein the protein contact map of the protein is generated by the protein contact map generation subnetwork in response to processing the reference amino acid sequence, the secondary structure profile per amino acid, the solvent accessibility profile per amino acid, and the position-specific frequency matrix per amino acid. 21. The variant pathogenicity prediction network according to any of items 1 to 20, wherein the protein contact map of the protein is generated by a protein contact map generation subnetwork in response to processing the reference amino acid sequence, the secondary structure profile per amino acid, the solvent accessibility profile per amino acid, and the position-specific scoring matrix per amino acid. 22. The variant pathogenicity prediction network according to any of items 1 to 21, wherein the processed representation of the alternative amino acid sequence is provided as an input to a first layer of a protein contact map generation sub-network. 23. The variant pathogenicity prediction network according to any of items 1 to 22, wherein the processed representations of the alternative amino acid sequences are provided as input to one or more intermediate layers of the protein contact map generation sub-network. 24. The variant pathogenicity prediction network according to any of items 1 to 23, wherein the processed representation of the alternative amino acid sequence is provided as input to a final layer of the protein contact map generation sub-network. 25. The variant pathogenicity prediction network according to any of items 1 to 24, wherein the processed representations of the alternative amino acid sequences are combined (e.g., concatenated, summed) with inputs to the protein contact map generation sub-network. 26. The variant pathogenicity prediction network according to any of items 1 to 25, wherein the processed representation of the alternative amino acid sequence is combined (e.g., concatenated, summed) with one or more intermediate outputs of the protein contact map generation sub-network. 27. The variant pathogenicity prediction network according to any of items 1 to 26, wherein the processed representation of the alternative amino acid sequence is combined (e.g., concatenated, summed) with the final output of the protein contact map generation sub-network. 28. The variant pathogenicity prediction network according to any one of items 1 to 27, wherein the reference amino acid sequence has L amino acids. 29. The variant pathogenicity prediction network according to any of items 1 to 28, wherein the reference amino acid sequence is characterized as a one-hot coding matrix of size L×C, where C denotes 20 amino acid categories. 30. The variant pathogenicity prediction network according to any one of items 1 to 29, wherein the primate conservation profile for each amino acid is a size of L×C. 31. The variant pathogenicity prediction network according to any of items 1 to 30, wherein the mammalian conservation profile for each amino acid is a size of L×C. 32. The variant pathogenicity prediction network according to any one of items 1 to 31, wherein the vertebrate conservation profile for each amino acid is a size of L×C. 33. The variant pathogenicity prediction network according to any of items 1 to 32, wherein the secondary structure profile for each amino acid is characterized as a three-state encoding matrix of size L×S, where S indicates the three secondary structure states. 34. The mutant pathogenicity prediction network according to any of items 1 to 33, wherein the solvent accessibility profile for each amino acid is characterized as a three-state encoding matrix of size L×A, where A indicates the three solvent accessibility states. 35. The variant pathogenicity prediction network according to any of items 1 to 34, wherein the position-specific scoring matrix for each amino acid is of size L×C. 36. The variant pathogenicity prediction network according to any one of items 1 to 35, wherein the position-specific frequency matrix for each amino acid is of size L×C. 37. A variant pathogenicity prediction network according to any one of items 1 to 36, wherein the variant encoding sub-network is a first convolutional neural network. 38. A variant pathogenicity prediction network according to any of items 1 to 37, wherein the first convolutional neural network comprises one or more one-dimensional (1D) convolutional layers. 39. The variant pathogenicity prediction network according to any one of items 1 to 38, wherein the protein contact map generation sub-network is a second convolutional neural network. 40. The variant pathogenicity prediction network of any of items 1 to 39, wherein the second convolutional neural network includes (i) one or more 1D convolutional layers, followed by (ii) one or more residual blocks with 1D convolutions, followed by (iii) a spatial dimension expansion layer, followed by (iv) one or more residual blocks with two-dimensional (2D) convolutions, followed by (v) one or more 2D convolutional layers. 41. The mutant pathogenicity prediction network according to any one of items 1 to 40, wherein the spatial dimensionality (e.g., width × height) of the input processed by the first 1D convolutional layer in the one or more 1D convolutional layers of the second convolutional neural network is L × 1. 42. A variant pathogenicity prediction network according to any one of items 1 to 41, wherein the depth dimensionality of the input processed by the first 1D convolutional layer is D (e.g., 66), and D = C + S + A + C + C. 43. A variant pathogenicity prediction network according to any of items 1 to 42, wherein the output of the final residual block in one or more residual blocks with 1D convolutions of the second convolutional neural network is processed by a spatial dimension expansion layer to generate a spatially expanded output. 44. The variant pathogenicity prediction network according to any one of items 1 to 43, wherein the spatial dimensionality of the spatially extended output is L × L. 45. A variant pathogenicity prediction network according to any of items 1 to 44, wherein the spatial dimension expansion layer is configured to apply a cross product to the output of the final residual block to generate a spatially expanded output. 46. A variant pathogenicity prediction network according to any of items 1 to 45, wherein the spatially extended output is processed by the first residual block in one or more residual blocks with 2D convolutions of a second convolutional neural network. 47. A variant pathogenicity prediction network according to any of items 1 to 46, wherein the total dimensionality of the protein contact map generated by the last 2D convolutional layer in the one or more 2D convolutional layers of the second convolutional neural network is L × L × 1. 48. A variant pathogenicity prediction network according to any of items 1 to 47, wherein the protein contact map generation subnetwork is pre-trained against a reference amino acid sequence of a bacterial protein with a known protein contact map. 49. The mutant pathogenicity prediction network according to any of items 1 to 48, wherein the protein contact map generation subnetwork is pre-trained using a mean squared error loss function that minimizes the error between a known protein contact map and a protein contact map predicted by the protein contact map generation subnetwork during pre-training. 50. The mutant pathogenicity prediction network according to any of items 1 to 49, wherein the protein contact map generation subnetwork is pre-trained using a mean absolute error loss function that minimizes the error between a known protein contact map and a protein contact map predicted by the protein contact map generation subnetwork during pre-training. 51. The variant pathogenicity prediction network according to any of items 1 to 50, wherein the protein contact map generation sub-network is pre-trained to generate a protein contact map as an output in response to processing a reference amino acid sequence and at least one of a per-amino acid secondary structure profile, a per-amino acid solvent accessibility profile, a per-amino acid position-specific scoring matrix, and a per-amino acid position-specific frequency matrix. 52. The pathogenicity scoring subnetwork is jointly trained end-to-end with the pre-trained protein contact map generation subnetwork and the variant encoding subnetwork to generate a pathogenicity index of the variant amino acid as an output in response to processing the protein contact map; Protein contact maps are a reference amino acid sequence and at least one of a secondary structure profile for each amino acid, a solvent accessibility profile for each amino acid, a position-specific scoring matrix for each amino acid, and a position-specific frequency matrix for each amino acid; and 52. The variant pathogenicity prediction network of any of items 1 to 51, generated by the pre-trained protein contact map generation sub-network in response to processing the processed representation generated by the variant encoding sub-network in response to processing the alternative amino acid sequences and at least one of a primate conservation profile per amino acid, a mammalian conservation profile per amino acid, and a vertebrate conservation profile per amino acid. 53. A variant pathogenicity prediction network according to any one of items 1 to 52, wherein the pre-trained protein contact map generation sub-network remains frozen during training of the variant encoding sub-network and the pathogenicity scoring sub-network and is not re-trained. 54. A variant pathogenicity prediction network according to any one of items 1 to 53, wherein the variant encoding sub-network, the protein contact map generation sub-network, and the pathogenicity scoring sub-network are arranged as a single neural network. 55. A variant pathogenicity prediction network according to any of items 1 to 54, wherein multiple trained instances of a single neural network are used as an ensemble for variant pathogenicity prediction during inference. 56. A variant pathogenicity prediction network according to any of items 1 to 55, wherein the pathogenicity scoring sub-network is a fully connected network. 57. A variant pathogenicity prediction network according to any of items 1 to 56, wherein the pathogenicity scoring sub-network includes a pathogenicity index generation layer (e.g., sigmoid, softmax) that generates a pathogenicity index. 58. Storing a reference amino acid sequence of a protein and an alternative amino acid sequence of the protein containing a mutant amino acid caused by a mutant nucleotide; processing the alternative amino acid sequences to generate processed representations of the alternative amino acid sequences; processing the processed representations of the reference amino acid sequence and the alternative amino acid sequence to generate a protein contact map for the protein; and processing the protein contact map and generating a pathogenicity index for the variant amino acid. 59. The method further comprising storing a primate conservation profile for each amino acid of the protein; 59. The computer implementation of claim 58, wherein the processed representation of the alternative amino acid sequence is generated in response to processing the alternative amino acid sequence and the primate conservation profile for each amino acid. 60. The method further comprising storing a mammalian conservation profile for each amino acid of the protein; 60. The computer-implemented method of claim 58 or 59, wherein the processed representation of the alternative amino acid sequence is generated in response to processing the alternative amino acid sequence and the mammalian conservation profile for each amino acid. 61. The method further comprising storing a vertebrate conservation profile for each amino acid of the protein; 61. The computer-implemented method of any of items 58 to 60, wherein the processed representations of the alternative amino acid sequences are generated in response to processing the alternative amino acid sequences and the vertebrate conservation profiles for each amino acid. 62. The computer-implemented method according to any of items 58 to 61, wherein the processed representation of the alternative amino acid sequence is generated in response to processing the alternative amino acid sequence, the primate conservation profile per amino acid, the mammalian conservation profile per amino acid, and the vertebrate conservation profile per amino acid. 63. The computer-implemented method according to any of items 58 to 62, wherein the processed representation of the alternative amino acid sequence is generated in response to processing the alternative amino acid sequence, the primate conservation profile for each amino acid, and the mammalian conservation profile for each amino acid. 64. The computer-implemented method of any of items 58 to 63, wherein the processed representation of the alternative amino acid sequence is generated in response to processing the alternative amino acid sequence, the primate conservation profile for each amino acid, and the vertebrate conservation profile for each amino acid. 65. The computer-implemented method according to any of items 58 to 64, wherein the processed representation of the alternative amino acid sequence is generated in response to processing the alternative amino acid sequence, the mammalian conservation profile for each amino acid, and the vertebrate conservation profile for each amino acid. 66. The method according to claim 1, further comprising storing a secondary structure profile for each amino acid of the protein; 66. The computer-implemented method of any of items 58 to 65, wherein the protein contact map of the protein is generated in response to processing the reference amino acid sequence and the secondary structure profile for each amino acid. 67. The method according to claim 1, further comprising storing a solvent accessibility profile for each amino acid of the protein; 67. The computer-implemented method of any of items 58 to 66, wherein the protein contact map of the protein is generated in response to processing the reference amino acid sequence and the solvent accessibility profile for each amino acid. 68. The method further includes storing a position-specific frequency matrix for each amino acid of the protein; 68. The computer-implemented method of any of items 58 to 67, wherein the protein contact map of the protein is generated in response to processing the reference amino acid sequence and a position-specific frequency matrix for each amino acid. 69. The method further comprising storing a position-specific scoring matrix for each amino acid of the protein; 69. The computer-implemented method of any of items 58 to 68, wherein the protein contact map of the protein is generated in response to processing the reference amino acid sequence and a position-specific scoring matrix for each amino acid. 70. The computer-implemented method of any of items 58 to 69, wherein the protein contact map of the protein is generated in response to processing a reference amino acid sequence, a secondary structure profile per amino acid, a solvent accessibility profile per amino acid, a position-specific frequency matrix per amino acid, and a position-specific scoring matrix per amino acid. 71. The computer-implemented method of any of items 58 to 70, wherein the protein contact map of the protein is generated in response to processing a reference amino acid sequence, a secondary structure profile for each amino acid, and a solvent accessibility profile for each amino acid. 72. The computer-implemented method according to any of items 58 to 71, wherein the protein contact map of the protein is generated in response to processing a reference amino acid sequence, a secondary structure profile for each amino acid, and a position-specific frequency matrix for each amino acid. 73. The computer-implemented method of any of items 58 to 72, wherein the protein contact map of the protein is generated in response to processing a reference amino acid sequence, a secondary structure profile for each amino acid, and a position-specific scoring matrix for each amino acid. 74. The computer-implemented method according to any of items 58 to 73, wherein the protein contact map of the protein is generated in response to processing a reference amino acid sequence, a solvent accessibility profile for each amino acid, and a position-specific frequency matrix for each amino acid. 75. The computer-implemented method of any of items 58 to 74, wherein the protein contact map of the protein is generated in response to processing a reference amino acid sequence, a solvent accessibility profile for each amino acid, and a position-specific scoring matrix for each amino acid. 76. The computer-implemented method of any of items 58 to 75, wherein the protein contact map of the protein is generated in response to processing a reference amino acid sequence, a position-specific frequency matrix for each amino acid, and a position-specific scoring matrix for each amino acid. 77. The computer-implemented method of any of items 58 to 76, wherein the protein contact map of the protein is generated in response to processing a reference amino acid sequence, a secondary structure profile for each amino acid, a solvent accessibility profile for each amino acid, and a position-specific frequency matrix for each amino acid. 78. The computer-implemented method according to any of items 58 to 77, wherein the protein contact map of the protein is generated in response to processing a reference amino acid sequence, a secondary structure profile for each amino acid, a solvent accessibility profile for each amino acid, and a position-specific scoring matrix for each amino acid. 79. A non-transitory computer-readable storage medium storing computer program instructions for predicting pathogenicity of a variant, the instructions, when executed on a processor, performing: storing a reference amino acid sequence of the protein and an alternative amino acid sequence of the protein containing a variant amino acid caused by the variant nucleotide; processing the alternative amino acid sequences to generate processed representations of the alternative amino acid sequences; processing the processed representations of the reference amino acid sequence and the alternative amino acid sequence to generate a protein contact map for the protein; and processing the protein contact map to generate a pathogenicity index for the variant amino acid. 80. Performing the method further comprises storing a primate conservation profile for each amino acid of the protein; 80. The non-transitory computer-readable storage medium of item 79, wherein the processed representation of the alternative amino acid sequence is generated in response to processing the alternative amino acid sequence and the primate conservation profile for each amino acid. 81. Performing the method further comprises storing a mammalian conservation profile for each amino acid of the protein; 81. The non-transitory computer-readable storage medium of item 79 or 80, wherein the processed representation of the alternative amino acid sequence is generated in response to processing the alternative amino acid sequence and the mammalian conservation profile for each amino acid. 82. Performing the method further comprises storing a vertebrate conservation profile for each amino acid of the protein; 82. The non-transitory computer-readable storage medium of any of items 79 to 81, wherein the processed representation of the alternative amino acid sequence is generated in response to processing the alternative amino acid sequence and the vertebrate conservation profile for each amino acid. 83. The non-transitory computer-readable storage medium according to any of items 79 to 82, wherein the processed representation of the alternative amino acid sequence is generated in response to processing the alternative amino acid sequence, the primate conservation profile per amino acid, the mammalian conservation profile per amino acid, and the vertebrate conservation profile per amino acid. 84. The non-transitory computer-readable storage medium according to any of items 79 to 83, wherein the processed representation of the alternative amino acid sequence is generated in response to processing the alternative amino acid sequence, the primate conservation profile for each amino acid, and the mammalian conservation profile for each amino acid. 85. The non-transitory computer-readable storage medium of any of items 79 to 84, wherein the processed representation of the alternative amino acid sequence is generated in response to processing the alternative amino acid sequence, the primate conservation profile for each amino acid, and the vertebrate conservation profile for each amino acid. 86. The non-transitory computer-readable storage medium according to any of items 79 to 85, wherein the processed representation of the alternative amino acid sequence is generated in response to processing the alternative amino acid sequence, the mammalian conservation profile for each amino acid, and the vertebrate conservation profile for each amino acid. 87. Performing the method further comprises storing a secondary structure profile for each amino acid of the protein; 87. The non-transitory computer-readable storage medium of any of items 79 to 86, wherein the protein contact map of the protein is generated in response to processing the reference amino acid sequence and the secondary structure profile for each amino acid. 88. Performing the method further comprises storing a solvent accessibility profile for each amino acid of the protein; 88. The non-transitory computer-readable storage medium of any of items 79 to 87, wherein the protein contact map of the protein is generated in response to processing the reference amino acid sequence and the solvent accessibility profile for each amino acid. 89. Performing the method further comprises storing a position-specific frequency matrix for each amino acid of the protein; 89. The non-transitory computer-readable storage medium of any of items 79 to 88, wherein the protein contact map of the protein is generated in response to processing the reference amino acid sequence and the position-specific frequency matrix for each amino acid. 90. Performing the method further comprises storing a position-specific scoring matrix for each amino acid of the protein; 90. The non-transitory computer-readable storage medium of any of items 79 to 89, wherein the protein contact map of the protein is generated in response to processing the reference amino acid sequence and the position-specific scoring matrix for each amino acid. 91. The non-transitory computer-readable storage medium according to any of items 79 to 90, wherein the protein contact map of the protein is generated in response to processing a reference amino acid sequence, a secondary structure profile for each amino acid, a solvent accessibility profile for each amino acid, a position-specific frequency matrix for each amino acid, and a position-specific scoring matrix for each amino acid. 92. The non-transitory computer-readable storage medium according to any of items 79 to 91, wherein the protein contact map of the protein is generated in response to processing a reference amino acid sequence, a secondary structure profile for each amino acid, and a solvent accessibility profile for each amino acid. 93. The non-transitory computer-readable storage medium according to any one of items 79 to 92, wherein the protein contact map of the protein is generated in response to processing a reference amino acid sequence, a secondary structure profile for each amino acid, and a position-specific frequency matrix for each amino acid. 94. The non-transitory computer-readable storage medium according to any of items 79 to 93, wherein the protein contact map of the protein is generated in response to processing a reference amino acid sequence, a secondary structure profile for each amino acid, and a position-specific scoring matrix for each amino acid. 95. The non-transitory computer-readable storage medium according to any of items 79 to 94, wherein the protein contact map of the protein is generated in response to processing a reference amino acid sequence, a solvent accessibility profile for each amino acid, and a position-specific frequency matrix for each amino acid. 96. The non-transitory computer-readable storage medium according to any of items 79 to 95, wherein the protein contact map of the protein is generated in response to processing a reference amino acid sequence, a solvent accessibility profile for each amino acid, and a position-specific scoring matrix for each amino acid. 97. The non-transitory computer-readable storage medium of any of items 79 to 96, wherein the protein contact map of the protein is generated in response to processing a reference amino acid sequence, a position-specific frequency matrix for each amino acid, and a position-specific scoring matrix for each amino acid. 98. The non-transitory computer-readable storage medium according to any of items 79 to 97, wherein the protein contact map of the protein is generated in response to processing a reference amino acid sequence, a secondary structure profile for each amino acid, a solvent accessibility profile for each amino acid, and a position-specific frequency matrix for each amino acid. 99. The non-transitory computer-readable storage medium of any of items 79 to 98, wherein the protein contact map of the protein is generated in response to processing a reference amino acid sequence, a secondary structure profile for each amino acid, a solvent accessibility profile for each amino acid, and a position-specific scoring matrix for each amino acid. 100. A system comprising a variant pathogenicity determiner configured to determine pathogenicity of a variant causing an amino acid variant in a protein based on processing a protein contact map of the protein. 101. A computer-implemented method comprising determining pathogenicity of a variant causing an amino acid variant in a protein based on processing a protein contact map of the protein. 102. A non-transitory computer-readable storage medium storing computer program instructions for predicting pathogenicity of a variant, the instructions, when executed on a processor, performing: A non-transitory computer readable storage medium implementing a method, comprising determining pathogenicity of a variant causing an amino acid variant in a protein based on processing a protein contact map of the protein.
[0185] Item Set 2 1. A variant pathogenicity classifier comprising: A memory that stores (i) a reference amino acid sequence of the protein, (ii) an alternative amino acid sequence of the protein containing a variant amino acid caused by a variant nucleotide, and (iii) a protein contact map of the protein; and runtime logic having access to a memory and configured to provide (i) a reference amino acid sequence, (ii) an alternative amino acid sequence, and (iii) a protein contact map as inputs to a first neural network, and in response to processing (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, and (iii) the protein contact map, cause the first neural network to generate as an output a pathogenicity index for the variant amino acid. 2. The memory stores a primate conservation profile for each amino acid of the protein, a mammalian conservation profile for each amino acid of the protein, and a vertebrate conservation profile for each amino acid of the protein; 2. The variant pathogenicity classifier of claim 1, wherein the runtime logic is further configured to provide (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, (iii) the protein contact map, (iv) the primate conservation profile per amino acid, (v) the mammalian conservation profile per amino acid, and (vi) the vertebrate conservation profile per amino acid as inputs to the first neural network, and in response to processing (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, (iii) the protein contact map, (iv) the primate conservation profile per amino acid, (v) the mammalian conservation profile per amino acid, and (vi) the vertebrate conservation profile per amino acid, cause the first neural network to generate as output a pathogenicity index for the variant amino acid. 3. The variant pathogenicity classifier of item 1 or 2, wherein the reference amino acid sequence has L amino acids and the alternative amino acid sequence has L amino acids. 4. The variant pathogenicity classifier of any of items 1 to 3, wherein the reference amino acid sequence is characterized as a reference one-hot coding matrix of size L×C, where C represents 20 amino acid categories, and the alternative amino acid sequence is characterized as an alternative one-hot coding matrix of size L×C. 5. The variant pathogenicity classifier according to any of items 1 to 4, wherein the primate conserved profile per amino acid is LxC in size, the mammalian conserved profile per amino acid is LxC in size, and the vertebrate conserved profile per amino acid is LxC in size. 6. The mutant pathogenicity classifier according to any one of items 1 to 5, wherein the first neural network is a first convolutional neural network. 7. The variant pathogenicity classifier according to any of items 1 to 6, wherein the first convolutional neural network includes: (i) one or more one-dimensional (1D) convolutional layers, followed by (ii) a first set of residual blocks having 1D convolutions, followed by (iii) a second set of residual blocks having 1D convolutions, followed by (iv) a spatial dimension expansion layer, followed by (v) a first set of residual blocks having two-dimensional (2D) convolutions, followed by (vi) one or more 2D convolutional layers, followed by (vii) one or more fully connected layers, followed by (viii) a pathogenicity indicator generation layer. 8. The variant pathogenicity classifier of any of items 1 to 7, wherein the spatial dimensionality (e.g., width × height) of the input processed by a first 1D convolutional layer in the one or more 1D convolutional layers is L × 1. 9. The variant pathogenicity classifier of any of items 1 to 8, wherein the depth dimensionality of the input processed by the first 1D convolution is D (e.g., 100), where D=C+C+C+C+C. 10. A variant pathogenicity classifier according to any of items 1 to 9, wherein the first set of residual blocks having 1D convolutions has N1 residual blocks (e.g., N1=2, 3, 4, 5), the second set of residual blocks having 1D convolutions has N2 residual blocks (e.g., N2=2, 3, 4, 5), and the first set of residual blocks having 2D convolutions has N3 residual blocks (e.g., N3=2, 3, 4, 5). 11. A variant pathogenicity classifier according to any of items 1 to 10, wherein the output of a final residual block in a second set of residual blocks having 1D convolutions is processed by a spatial dimension expansion layer to generate a spatially expanded output. 12. The variant pathogenicity classifier of any of items 1 to 11, wherein the spatial dimensionality expansion layer is configured to apply a cross product to the output of the final residual block to generate a spatially expanded output. 13. The variant pathogenicity classifier according to any of items 1 to 12, wherein the spatial dimensionality of the spatially extended output is L×L. 14. The variant pathogenicity classifier of any of items 1 to 13, wherein the spatially extended output is combined (e.g., concatenated, summed) with the protein contact map to generate an intermediate combined output. 15. A variant pathogenicity classifier according to any of items 1 to 14, wherein the intermediate synthesis output is processed by a first residual block in a first set of residual blocks having a 2D convolution. 16. The variant pathogenicity classifier of any of items 1 to 15, wherein the protein contact map is provided as an input to a first layer of the first neural network. 17. The variant pathogenicity classifier of any of items 1 to 16, wherein the protein contact map is provided as an input to one or more intermediate layers of the first neural network. 18. The variant pathogenicity classifier of any of items 1 to 17, wherein the protein contact map is provided as an input to a final layer of the first neural network. 19. The variant pathogenicity classifier of any of items 1 to 18, wherein the protein contact map is combined (e.g., concatenated, summed) with an input to the first neural network. 20. The variant pathogenicity classifier of any of items 1 to 19, wherein the protein contact map is combined (e.g., concatenated, summed) with one or more intermediate outputs of the first neural network. 21. The variant pathogenicity classifier according to any of items 1 to 20, wherein the protein contact map is combined (e.g., concatenated, summed) with the final output of the first neural network. 22. The variant pathogenicity classifier of any of items 1 to 21, wherein the protein contact map is generated by a second neural network in response to processing (i) the reference amino acid sequence and at least one of (ii) a protein secondary structure profile per amino acid, (iii) a solvent accessibility profile per amino acid, (iv) a position-specific scoring matrix per amino acid, and (v) a position-specific frequency matrix per amino acid. 23. The variant pathogenicity classifier according to any of items 1 to 22, wherein the protein contact map has a total dimensionality of L×L×K (e.g., K=10, 15, 20, 25). 24. The variant pathogenicity classifier according to any of items 1 to 23, wherein the second neural network is a second convolutional neural network. 25. The variant pathogenicity classifier of any of items 1 to 24, wherein the second convolutional neural network includes (i) one or more 1D convolutional layers followed by (ii) one or more residual blocks with 1D convolutions followed by (iii) a spatial dimensionality expansion layer followed by (iv) one or more residual blocks with 2D convolutions followed by (v) one or more 2D convolutional layers. 26. A variant pathogenicity classifier according to any of items 1 to 25, wherein the first convolutional neural network uses convolutional filters of different filter sizes (e.g., 5x2, 2x5). 27. The variant pathogenicity classifier according to any of items 1 to 26, wherein the first convolutional neural network does not include one or more fully connected layers. 28. A variant pathogenicity classifier according to any of items 1 to 27, wherein multiple trained instances of the first neural network are used as an ensemble for variant pathogenicity prediction during inference. 29. The variant pathogenicity classifier of any of items 1 to 28, wherein the first and second sets of residual blocks with 1D convolutions perform a series of 1D convolution transformations of 1D sequence features in (i) a reference amino acid sequence, (ii) an alternative amino acid sequence, and at least one of (iii) a primate conservation profile per amino acid, (iv) a mammalian conservation profile per amino acid, and (v) a vertebrate conservation profile per amino acid. 30. A variant pathogenicity classifier according to any of items 1 to 29, wherein the first set of residual blocks with 2D convolutions performs a series of 2D convolution transformations of (i) the protein contact map and (ii) the 2D spatial features in the intermediate composite output. 31. The variant pathogenicity classifier according to any of items 1 to 30, wherein a first set of residual blocks with 2D convolutions extracts spatial interactions from the protein contact map for pathogenic associations between amino acids of the protein that are closer in the three-dimensional (3D) structure of the protein than the reference amino acid sequence and the alternative amino acid sequence. 32. Storing (i) a reference amino acid sequence of a protein, (ii) an alternative amino acid sequence of the protein containing a mutant amino acid caused by a mutant nucleotide, and (iii) a protein contact map of the protein; 1. A computer-implemented method for variant pathogenicity classification, comprising: providing as input to a first neural network: (i) a reference amino acid sequence; (ii) an alternative amino acid sequence; and (iii) a protein contact map; and causing the first neural network to generate as output a pathogenicity index for the variant amino acid in response to processing the (i) reference amino acid sequence, (ii) the alternative amino acid sequence, and (iii) the protein contact map. 33. Storing a primate conservation profile for each amino acid of the protein, a mammalian conservation profile for each amino acid of the protein, and a vertebrate conservation profile for each amino acid of the protein; 33. The computer-implemented method of claim 32, further comprising: providing as input to the first neural network (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, (iii) the protein contact map, (iv) the primate conservation profile per amino acid, (v) the mammalian conservation profile per amino acid, and (vi) the vertebrate conservation profile per amino acid; and causing the first neural network to generate as output a pathogenicity index for the variant amino acid in response to processing the (i) reference amino acid sequence, (ii) the alternative amino acid sequence, (iii) the protein contact map, (iv) the primate conservation profile per amino acid, (v) the mammalian conservation profile per amino acid, and (vi) the vertebrate conservation profile per amino acid. 34. The computer-implemented method of item 32 or 33, wherein the reference amino acid sequence has L amino acids and the alternative amino acid sequence has L amino acids. 35. The computer-implemented method of any of items 32 to 34, wherein the reference amino acid sequence is characterized as a reference one-hot coding matrix of size L×C, where C denotes 20 amino acid categories, and the alternative amino acid sequence is characterized as an alternative one-hot coding matrix of size L×C. 36. The computer-implemented method according to any of items 32 to 35, wherein the primate conservation profile per amino acid is LxC in size, the mammalian conservation profile per amino acid is LxC in size, and the vertebrate conservation profile per amino acid is LxC in size. 37. The computer-implemented method of any of items 32 to 36, wherein the first neural network is a first convolutional neural network. 38. The computer-implemented method of any of items 32 to 37, wherein the first convolutional neural network includes (i) one or more one-dimensional (1D) convolutional layers, followed by (ii) a first set of residual blocks having 1D convolutions, followed by (iii) a second set of residual blocks having 1D convolutions, followed by (iv) a spatial dimension expansion layer, followed by (v) a first set of residual blocks having two-dimensional (2D) convolutions, followed by (vi) one or more 2D convolutional layers, followed by (vii) one or more fully connected layers, followed by (viii) a pathogenicity indicator generation layer. 39. The computer-implemented method of any of items 32 to 38, wherein the spatial dimensionality (e.g., width × height) of the input processed by a first 1D convolutional layer in the one or more 1D convolutional layers is L × 1. 40. The computer-implemented method of any of items 32 to 39, wherein the depth dimensionality of the input processed by the first 1D convolution is D (e.g., 100), and D = C + C + C + C + C. 41. The computer-implemented method of any of items 32 to 40, wherein a first set of residual blocks having 1D convolution has N1 residual blocks (e.g., N1=2, 3, 4, 5), a second set of residual blocks having 1D convolution has N2 residual blocks (e.g., N2=2, 3, 4, 5), and a first set of residual blocks having 2D convolution has N3 residual blocks (e.g., N3=2, 3, 4, 5). 42. A computer-implemented method according to any of items 32 to 41, wherein the output of a final residual block in a second set of residual blocks having 1D convolution is processed by a spatial dimension enhancement layer to generate a spatially enhanced output. 43. A computer-implemented method according to any of items 32 to 42, wherein the spatial dimension enhancement layer is configured to apply a cross product to the output of the final residual block to generate a spatially enhanced output. 44. A computer-implemented method according to any one of items 32 to 43, wherein the spatial dimensionality of the spatially extended output is L×L. 45. A computer-implemented method according to any of items 32 to 44, wherein the spatially extended output is combined (e.g., concatenated, summed) with the protein contact map to generate an intermediate combined output. 46. A computer-implemented method according to any of items 32 to 45, wherein the intermediate synthesis output is processed by a first residual block in a first set of residual blocks having a 2D convolution. 47. A computer-implemented method according to any of items 32 to 46, wherein the protein contact map is provided as an input to a first layer of a first neural network. 48. The computer-implemented method of any of items 32 to 47, wherein the protein contact map is provided as input to one or more intermediate layers of the first neural network. 49. The computer-implemented method of any of items 32 to 48, wherein the protein contact map is provided as input to a final layer of the first neural network. 50. A computer-implemented method according to any of items 32 to 49, wherein the protein contact map is combined (e.g., concatenated, summed) with an input to the first neural network. 51. A computer-implemented method according to any of items 32 to 50, wherein the protein contact map is combined (e.g., concatenated, summed) with one or more intermediate outputs of the first neural network. 52. A computer-implemented method according to any of items 32 to 51, wherein the protein contact map is combined (e.g., concatenated, summed) with the final output of the first neural network. 53. The computer-implemented method of any of items 32 to 52, wherein the protein contact map is generated by a second neural network in response to processing (i) the reference amino acid sequence and at least one of (ii) a protein secondary structure profile per amino acid, (iii) a solvent accessibility profile per amino acid, (iv) a position-specific scoring matrix per amino acid, and (v) a position-specific frequency matrix per amino acid. 54. A computer-implemented method according to any of items 32 to 53, wherein the protein contact map has a total dimensionality of L×L×K (e.g., K=10, 15, 20, 25). 55. The computer-implemented method of any of items 32 to 54, wherein the second neural network is a second convolutional neural network. 56. The computer-implemented method of any of items 32 to 55, wherein the second convolutional neural network includes (i) one or more 1D convolutional layers followed by (ii) one or more residual blocks with 1D convolutions followed by (iii) a spatial dimension expansion layer followed by (iv) one or more residual blocks with 2D convolutions followed by (v) one or more 2D convolutional layers. 57. The computer-implemented method of any of items 32 to 56, wherein the first convolutional neural network uses convolutional filters of different filter sizes (e.g., 5x2, 2x5). 58. The computer-implemented method of any of items 32 to 57, wherein the first convolutional neural network does not include one or more fully connected layers. 59. The computer-implemented method of any of items 32 to 58, wherein multiple trained instances of the first neural network are used as an ensemble for variant pathogenicity prediction during inference. 60. The computer-implemented method of any of paragraphs 32 to 59, wherein the first and second sets of residual blocks having 1D convolutions perform a series of 1D convolution transformations of 1D sequence features in (i) a reference amino acid sequence, (ii) an alternative amino acid sequence, and at least one of (iii) a primate conservation profile per amino acid, (iv) a mammalian conservation profile per amino acid, and (v) a vertebrate conservation profile per amino acid. 61. A computer-implemented method according to any of items 32 to 60, wherein the first set of residual blocks with 2D convolutions performs a series of 2D convolution transformations of (i) the protein contact map and (ii) the 2D spatial features in the intermediate composite output. 62. The computer-implemented method of any of items 32 to 61, wherein a first set of residual blocks with 2D convolutions extracts spatial interactions from the protein contact map for pathogenic associations between amino acids of the protein that are closer in the three-dimensional (3D) structure of the protein than the reference amino acid sequence and the alternative amino acid sequence. 63. A non-transitory computer-readable storage medium storing computer program instructions for classifying pathogenicity of a variant, the instructions, when executed on a processor, performing: storing (i) a reference amino acid sequence of the protein, (ii) an alternative amino acid sequence of the protein containing a variant amino acid caused by a variant nucleotide, and (iii) a protein contact map of the protein; 1. A non-transitory computer-readable storage medium implementing a method comprising: providing as input to a first neural network (i) a reference amino acid sequence, (ii) an alternative amino acid sequence, and (iii) a protein contact map; and causing the first neural network to generate as output a pathogenicity index for the variant amino acid in response to processing the (i) reference amino acid sequence, (ii) the alternative amino acid sequence, and (iii) the protein contact map. 64. Carrying out the method includes storing a primate conservation profile for each amino acid of the protein, a mammalian conservation profile for each amino acid of the protein, and a vertebrate conservation profile for each amino acid of the protein; 64. The non-transitory computer-readable storage medium of claim 63, further comprising: providing (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, (iii) the protein contact map, (iv) the primate conservation profile per amino acid, (v) the mammalian conservation profile per amino acid, and (vi) the vertebrate conservation profile per amino acid as inputs to the first neural network; and in response to processing (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, (iii) the protein contact map, (iv) the primate conservation profile per amino acid, (v) the mammalian conservation profile per amino acid, and (vi) the vertebrate conservation profile per amino acid, causing the first neural network to generate as output a pathogenicity index for the variant amino acid. 65. The non-transitory computer-readable storage medium according to item 63 or 64, wherein the reference amino acid sequence has L amino acids and the alternative amino acid sequence has L amino acids. 66. The non-transitory computer-readable storage medium of any of items 63 to 65, wherein the reference amino acid sequence is characterized as a reference one-hot coding matrix of size L×C, where C represents 20 amino acid categories, and the alternative amino acid sequence is characterized as an alternative one-hot coding matrix of size L×C. 67. The non-transitory computer-readable storage medium according to any of items 63 to 66, wherein the primate conservation profile for each amino acid is L×C in size, the mammalian conservation profile for each amino acid is L×C in size, and the vertebrate conservation profile for each amino acid is L×C in size. 68. The non-transitory computer-readable storage medium of any of items 63 to 67, wherein the first neural network is a first convolutional neural network. 69. The non-transitory computer-readable storage medium of any of items 63 to 68, wherein the first convolutional neural network includes (i) one or more one-dimensional (1D) convolutional layers, followed by (ii) a first set of residual blocks having 1D convolutions, followed by (iii) a second set of residual blocks having 1D convolutions, followed by (iv) a spatial dimension expansion layer, followed by (v) a first set of residual blocks having two-dimensional (2D) convolutions, followed by (vi) one or more 2D convolutional layers, followed by (vii) one or more fully connected layers, followed by (viii) a pathogenicity indicator generation layer. 70. The non-transitory computer-readable storage medium of any of items 63 to 69, wherein the spatial dimensionality (e.g., width × height) of an input processed by a first 1D convolutional layer in the one or more 1D convolutional layers is L × 1. 71. A non-transitory computer-readable storage medium according to any of items 63 to 70, wherein the depth dimensionality of the input processed by the first 1D convolution is D (e.g., 100), where D = C + C + C + C + C. 72. A non-transitory computer-readable storage medium according to any of items 63 to 71, wherein a first set of residual blocks having 1D convolution has N1 residual blocks (e.g., N1 = 2, 3, 4, 5), a second set of residual blocks having 1D convolution has N2 residual blocks (e.g., N2 = 2, 3, 4, 5), and a first set of residual blocks having 2D convolution has N3 residual blocks (e.g., N3 = 2, 3, 4, 5). 73. A non-transitory computer-readable storage medium according to any of items 63 to 72, wherein the output of a final residual block in a second set of residual blocks having 1D convolution is processed by a spatial dimension enhancement layer to generate a spatially enhanced output. 74. A non-transitory computer-readable storage medium according to any of items 63 to 73, wherein the spatial dimension enhancement layer is configured to apply a cross product to the output of the final residual block to generate a spatially enhanced output. 75. A non-transitory computer-readable storage medium according to any one of items 63 to 74, wherein the spatial dimensionality of the spatially extended output is L×L. 76. A non-transitory computer-readable storage medium according to any of items 63 to 75, wherein the spatially extended output is combined (e.g., concatenated, summed) with the protein contact map to generate an intermediate combined output. 77. A non-transitory computer-readable storage medium according to any of items 63 to 76, wherein the intermediate synthesis output is processed by a first residual block in a first set of residual blocks having a 2D convolution. 78. The non-transitory computer-readable storage medium of any of items 63 to 77, wherein the protein contact map is provided as an input to a first layer of a first neural network. 79. The non-transitory computer-readable storage medium of any of items 63 to 78, wherein the protein contact map is provided as an input to one or more intermediate layers of the first neural network. 80. The non-transitory computer-readable storage medium of any of items 63 to 79, wherein the protein contact map is provided as an input to a final layer of the first neural network. 81. A non-transitory computer-readable storage medium according to any of items 63 to 80, wherein the protein contact map is combined (e.g., concatenated, summed) with the input to the first neural network. 82. The non-transitory computer-readable storage medium of any of items 63 to 81, wherein the protein contact map is combined (e.g., concatenated, summed) with one or more intermediate outputs of the first neural network. 83. A non-transitory computer-readable storage medium according to any of items 63 to 82, wherein the protein contact map is combined (e.g., concatenated, summed) with the final output of the first neural network. 84. The non-transitory computer-readable storage medium of any of items 63 to 83, wherein the protein contact map is generated by a second neural network in response to processing (i) the reference amino acid sequence and at least one of (ii) a protein secondary structure profile for each amino acid, (iii) a solvent accessibility profile for each amino acid, (iv) a position-specific scoring matrix for each amino acid, and (v) a position-specific frequency matrix for each amino acid. 85. The non-transitory computer-readable storage medium according to any of items 63 to 84, wherein the protein contact map has a total dimensionality of L×L×K (e.g., K=10, 15, 20, 25). 86. The non-transitory computer-readable storage medium of any of items 63 to 85, wherein the second neural network is a second convolutional neural network. 87. The non-transitory computer-readable storage medium of any of items 63 to 86, wherein the second convolutional neural network includes (i) one or more 1D convolutional layers, followed by (ii) one or more residual blocks having 1D convolutions, followed by (iii) a spatial dimension expansion layer, followed by (iv) one or more residual blocks having 2D convolutions, followed by (v) one or more 2D convolutional layers. 88. A non-transitory computer-readable storage medium according to any of items 63 to 87, wherein the first convolutional neural network uses convolutional filters of different filter sizes (e.g., 5x2, 2x5). 89. A non-transitory computer-readable storage medium according to any of items 63 to 88, wherein the first convolutional neural network does not include one or more fully connected layers. 90. A non-transitory computer-readable storage medium according to any of items 63 to 89, wherein multiple trained instances of the first neural network are used as an ensemble for variant pathogenicity prediction during inference. 91. The non-transitory computer-readable storage medium according to any of items 63 to 90, wherein the first and second sets of residual blocks having 1D convolutions perform a series of 1D convolution transformations of 1D sequence features in (i) a reference amino acid sequence, (ii) an alternative amino acid sequence, and at least one of (iii) a primate conservation profile per amino acid, (iv) a mammalian conservation profile per amino acid, and (v) a vertebrate conservation profile per amino acid. 92. A non-transitory computer-readable storage medium according to any of items 63 to 91, wherein the first set of residual blocks having 2D convolutions performs a series of 2D convolution transformations of (i) the protein contact map and (ii) the 2D spatial features in the intermediate composite output. 93. The non-transitory computer-readable storage medium of any of items 63 to 92, wherein the first set of residual blocks having 2D convolutions extracts spatial interactions from the protein contact map for pathogenic associations between amino acids of the protein that are closer in the three-dimensional (3D) structure of the protein than the reference amino acid sequence and the alternative amino acid sequence.
[0186] Although the present invention has been disclosed with reference to the above-mentioned preferred embodiments and examples, it should be understood that these examples are intended in an illustrative and not a limiting sense. Modifications and combinations may easily occur to those skilled in the art, and such modifications and combinations are considered to be within the spirit of the present invention and the scope of the following claims. [Explanation of symbols]
[0187] 3400 Computer System 3410 Storage Subsystem 3422 Memory Subsystem 3432 Main Random Access Memory 3434 Read-Only Memory 3436 File Storage Subsystem 3438 User Interface Input Devices 3455 Bus Subsystem 3472 Central Processing Unit 3474 Network Interface Subsystem 3476 User Interface Output Device 3478 Processor
Claims
1. A system comprising: a variant pathogenicity classifier; a memory storing (i) a reference amino acid sequence of a protein, (ii) alternative amino acid sequences of the protein containing variant amino acids caused by variant nucleotides, and (iii) a protein contact map of the protein; and runtime logic having access to the memory and configured to provide (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, and (iii) the protein contact map as inputs to a first neural network, and to cause the first neural network to generate as output a pathogenicity index for the variant amino acid in response to processing (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, and (iii) the protein contact map.
2. the memory stores a primate conservation profile for each amino acid of the protein, a mammalian conservation profile for each amino acid of the protein, and a vertebrate conservation profile for each amino acid of the protein; 2. The system of claim 1, wherein the runtime logic is further configured to provide (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, (iii) the protein contact map, (iv) the primate conservation profile for each amino acid, (v) the mammalian conservation profile for each amino acid, and (vi) the vertebrate conservation profile for each amino acid as inputs to the first neural network, and to cause the first neural network to generate the pathogenicity index of the variant amino acid as an output in response to processing (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, (iii) the protein contact map, (iv) the primate conservation profile for each amino acid, (v) the mammalian conservation profile for each amino acid, and (vi) the vertebrate conservation profile for each amino acid.
3. The reference amino acid sequence has L amino acids, and the alternative amino acid sequence has L amino acids, 2. The system of claim 1, wherein the reference amino acid sequence is characterized as a reference one-hot encoding matrix of size L×C, where C represents 20 amino acid categories, and the alternative amino acid sequences are characterized as alternative one-hot encoding matrices of size L×C.
4. The reference amino acid sequence has L amino acids, and the alternative amino acid sequence has L amino acids, 3. The system of claim 2, wherein the primate conservation profile for each amino acid is L x C in size, the mammalian conservation profile for each amino acid is L x C in size, and the vertebrate conservation profile for each amino acid is L x C in size.
5. The first neural network is a first convolutional neural network; 2. The system of claim 1, wherein the first convolutional neural network includes: (i) one or more one-dimensional (1D) convolutional layers followed by (ii) a first set of residual blocks with 1D convolutions, followed by (iii) a second set of residual blocks with 1D convolutions, followed by (iv) a spatial dimension expansion layer, followed by (v) a first set of residual blocks with two-dimensional (2D) convolutions, followed by (vi) one or more 2D convolutional layers, followed by (vii) one or more fully connected layers, followed by (viii) a pathogenicity indicator generation layer.
6. 6. The system of claim 5, wherein the spatial dimensionality of the input processed by a first 1D convolutional layer in the one or more 1D convolutional layers is L×1.
7. 7. The system of claim 6, wherein a depth dimensionality of the input processed by the first 1D convolutional layer is D, where D=C+C+C+C+C+C, the first set of residual blocks with 1D convolutions has N residual blocks, the second set of residual blocks with 1D convolutions has N residual blocks, and the first set of residual blocks with 2D convolutions has N residual blocks.
8. 6. The system of claim 5, wherein an output of a final residual block in the second set of residual blocks having a 1D convolution is processed by the spatial dimension enhancement layer to generate a spatially enhanced output.
9. The system of claim 8 , wherein the spatial dimension enhancement layer is configured to apply a cross product to the output of the final residual block to generate the spatially enhanced output.
10. The system of claim 8 , wherein the spatial dimensionality of the spatially extended output is L×L.
11. The system of claim 8 , wherein the spatially extended output is combined with the protein contact map to generate an intermediate composite output.
12. The system of claim 11 , wherein the intermediate composite output is processed by a first residual block in the first set of residual blocks with a 2D convolution.
13. A system described in any of claims 1 to 12, wherein the protein contact map is generated by a second neural network in response to processing (i) the reference amino acid sequence, and at least one of (ii) a protein secondary structure profile for each amino acid, (iii) a solvent accessibility profile for each amino acid, (iv) a position-specific scoring matrix for each amino acid, and (v) a position-specific frequency matrix for each amino acid.
14. A method for identifying a protein comprising: storing (i) a reference amino acid sequence of a protein; (ii) an alternative amino acid sequence of the protein containing a variant amino acid caused by a variant nucleotide; and (iii) a protein contact map of the protein; 1. A computer-implemented method for variant pathogenicity classification, comprising: providing as input to a first neural network (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, and (iii) the protein contact map; and causing the first neural network to generate as output a pathogenicity index for the variant amino acid in response to processing (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, and (iii) the protein contact map.
15. 15. The computer-implemented method of claim 14, wherein the protein contact map is generated by a second neural network in response to processing (i) the reference amino acid sequence and at least one of (ii) a protein secondary structure profile for each amino acid, (iii) a solvent accessibility profile for each amino acid, (iv) a position-specific scoring matrix for each amino acid, and (v) a position-specific frequency matrix for each amino acid.
16. 15. The computer-implemented method of claim 14, wherein the protein contact map has a total dimensionality of LxLxK.
17. The method of claim 16, wherein the second neural network is a second convolutional neural network.
16. The computer-implemented method of claim 15, wherein the second convolutional neural network includes: (i) one or more 1D convolutional layers followed by (ii) one or more residual blocks with 1D convolutions followed by (iii) a spatial dimension enhancement layer followed by (iv) one or more residual blocks with 2D convolutions followed by (v) one or more 2D convolutional layers.
18. 18. The computer-implemented method of any of claims 14 to 17, wherein multiple trained instances of the first neural network are used as an ensemble for variant pathogenicity prediction during inference.
19. 1. A non-transitory computer-readable storage medium storing computer program instructions for classifying the pathogenicity of a variant, the computer program instructions, when executed on a processor, comprising: Storing (i) a reference amino acid sequence of a protein, (ii) an alternative amino acid sequence of the protein containing variant amino acids caused by variant nucleotides, and (iii) a protein contact map of the protein; 10. A non-transitory computer-readable storage medium implementing a method comprising: (i) providing the reference amino acid sequence, (ii) the alternative amino acid sequence, and (iii) the protein contact map as inputs to a first neural network; and causing the first neural network to generate as output a pathogenicity index for the variant amino acid in response to processing (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, and (iii) the protein contact map.
20. Performing the method includes storing a primate conservation profile for each amino acid of the protein, a mammalian conservation profile for each amino acid of the protein, and a vertebrate conservation profile for each amino acid of the protein; 20. The non-transitory computer-readable storage medium of claim 19, further comprising: providing (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, (iii) the protein contact map, (iv) the primate conservation profile for each amino acid, (v) the mammalian conservation profile for each amino acid, and (vi) the vertebrate conservation profile for each amino acid as inputs to the first neural network; and causing the first neural network to generate the pathogenicity index of the variant amino acid as output in response to processing (i) the reference amino acid sequence, (ii) the alternative amino acid sequence, (iii) the protein contact map, (iv) the primate conservation profile for each amino acid, (v) the mammalian conservation profile for each amino acid, and (vi) the vertebrate conservation profile for each amino acid.